System

The system improves video production efficiency and quality by automating media data analysis, generation, and user corrections, addressing the inefficiencies of conventional methods.

JP2026019830APending Publication Date: 2026-02-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024121578
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Conventional methods for video production struggle with visualizing the flow and development of videos, leading to inefficient and inconsistent creation of summary videos and storyboards, especially when dealing with a wide range of material.

Method used

A system that includes receiving media data, analyzing it to extract features, generating a video summary or storyboard, encoding the result, and allowing user corrections, with the ability to detect specific objects and improve quality.

Benefits of technology

Enhances the efficiency and quality of video production by automating the analysis and editing process, ensuring consistency between visuals and story.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026019830000001_ABST
    Figure 2026019830000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: The system includes a means for receiving media data, a means for analyzing the received media data and extracting features, a means for generating a summary video image or a storyboard on the basis of the extracted features, and a means for encoding the generated summary video image or storyboard and transmitting it to a user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In video production, it is important to visualize the flow and development of the video and improve the consistency between the visuals and the story during the production process. However, with conventional methods, this is difficult and requires a lot of time and effort. In particular, when there is a wide range of material, creating summary videos and storyboards can be extremely inefficient, and the quality can be inconsistent. There is a need to solve this problem and provide a means for efficient and easy video production. [Means for solving the problem]

[0005] The present invention provides a system including a means for receiving media data, a means for analyzing the received media data and extracting features, a means for generating a video summary or storyboard based on the extracted features, and a means for encoding the generated video summary or storyboard and transmitting it to a user terminal. Furthermore, this system includes a means for detecting specific objects in the received media data, thereby improving the quality of the video summary or storyboard. Furthermore, by including a means for a user to make corrections to the video summary or storyboard after checking it and upload it back to the system, adjustments can be made according to the user's intentions, resulting in more efficient video production and improved quality.

[0006] "Media data" refers to various types of digital data such as images, video, and audio.

[0007] The term "receiving means" refers to a device or program that has the function of obtaining media data from a user via a network.

[0008] "Means for analyzing and extracting features" refers to a device or program that has the function of analyzing the information contained in media data and extracting important elements or patterns.

[0009] "Abridged video" refers to a short video that extracts only certain important scenes from the original video.

[0010] A "storyboard" is a series of drawings that are used to show the visual image and story development of each scene in a film.

[0011] "Generating means" refers to a device or program that has the function of generating a new summary video or storyboard based on the analysis results.

[0012] "Means for Encoding" means a device or program capable of storing and compressing the generated video summary or storyboard in a digital file format.

[0013] "User terminal" refers to a digital device used by a user, such as a computer, smartphone, or tablet.

[0014] "Transmitting means" refers to a device or program that has the function of transferring encoded data to a user terminal.

[0015] "Means for detecting objects" refers to a device or program that has the ability to identify and extract specific elements or portions within media data.

[0016] "Means for making modifications" refers to a device or program that has the function of allowing a user to change or adjust the content of the generated summary video or storyboard.

[0017] "Means for uploading data back to the system" refers to a device or program that has the function of sending corrected data back to the system. [Brief explanation of the drawings]

[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0020] First, the terms used in the following description will be explained.

[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0026] [First embodiment]

[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0039] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a summary video or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[0040] Server roles and processing flow

[0041] 1. Receiving media data

[0042] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[0043] 2. Data analysis and feature extraction

[0044] The server analyzes the received media data using AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted to text using speech recognition technology, and sentiment analysis is performed.

[0045] 3. Generate a summary video or storyboard

[0046] The server generates a video summary or storyboard based on the extracted features. In the case of a video summary, it combines frames from key scenes to create short video clips, and in the case of a storyboard, it organizes images according to the storyboard frames.

[0047] 4. Data Encoding

[0048] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[0049] 5. Send to user device

[0050] The server sends the encoded data to the user terminal using an appropriate protocol (e.g., SSL / TLS) to ensure security and efficiency.

[0051] Terminal roles and processing flow

[0052] 1. Providing UI

[0053] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[0054] 2. Sending and Receiving Data

[0055] The device sends the media data uploaded by the user to the server, checks the integrity of the data when it is sent, and prompts the user to correct it if necessary, and receives the processed data returned from the server and displays it appropriately.

[0056] User roles and operation flow

[0057] 1. Upload your data

[0058] Users upload the media data (images, videos, audio) they have collected to the system via their terminals. When uploading, they also enter related information (title, summary, etc.).

[0059] 2. Check and correct the processing results

[0060] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[0061] Specific examples

[0062] A concrete example of creating a movie trailer

[0063] 1. Data upload (user)

[0064] Users upload multiple video files containing scenes from their movies to the system.

[0065] 2. Data analysis (server)

[0066] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[0067] 3. Summary video generation (server)

[0068] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[0069] 4. Receiving and checking results (user)

[0070] The user can check the generated trailer on their device and, if they are not satisfied with the footage, make any necessary corrections.

[0071] 5. Corrections and final confirmation (user)

[0072] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[0073] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[0074] The processing flow will be explained below.

[0075] Step 1:

[0076] A user uploads media data (images, videos, audio) from a device to the system. Specifically, the user opens a file selection dialog using the device's UI and specifies the file to upload. By pressing the upload button, the selected file is sent to the server.

[0077] Step 2:

[0078] The server stores the media data received from the user in storage. During this process, the server analyzes the file metadata (file name, size, type) and records it in a log.

[0079] Step 3:

[0080] The server performs an initial analysis of the received media data. Specifically, it checks the file format and scans the content, and then selects the appropriate analysis module. For image data, it uses the object detection module, and for video data, it uses the scene segmentation module.

[0081] Step 4:

[0082] The server uses the selected analysis module to extract features from the media data, for example, detecting important frames and scenes in video data, identifying key objects and scenes in image data, and converting speech to text and detecting specific keywords and emotions in audio data.

[0083] Step 5:

[0084] The server generates a summary video or storyboard based on the extracted feature information. In the case of a summary video, the extracted important scenes are arranged along a timeline and edited to a specified length. In the case of a storyboard, important scenes are arranged in each frame of the storyboard and explanations are added.

[0085] Step 6:

[0086] The server encodes the generated summary or storyboard into a standard file format, for example MP4 for the summary and JPEG or PDF for the storyboard, adding metadata (scene descriptions, timestamps) during the process.

[0087] Step 7:

[0088] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[0089] Step 8:

[0090] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[0091] Step 9:

[0092] The user checks the downloaded summary video or storyboard, and if there are any problems or omissions, they can correct them using the editing tool. Once the corrections are complete, they can upload it back to the system.

[0093] Step 10:

[0094] The server receives the modified data, re-analyzes it, and encodes it, incorporating any appropriate modifications to generate the final summary or storyboard. The resulting data is then sent back to the user's device, completing the process.

[0095] Example 1

[0096] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0097] Conventional media data analysis and editing work is often performed manually, resulting in problems of time and effort. Furthermore, the accuracy and consistency of analysis results often depend on specialized knowledge and are therefore poor. To solve these problems, the present invention aims to automate the reception, analysis, summary video or storyboard generation, encoding, and transmission of media data, enabling efficient and accurate processing.

[0098] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0099] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a video summary or storyboard based on the extracted features, means for encoding the generated video summary or storyboard into a standard file format, and means for transmitting the encoded video summary or storyboard to a user terminal. This makes it possible to automate the analysis and editing of media data and generate video summary and storyboards with high accuracy and efficiency.

[0100] "Media data" is a general term for data including image data, video data, and audio data.

[0101] The term "receiving means" refers to a function or device that can transmit media data from a user to a server and receive it on the server side.

[0102] "Analysis and feature extraction means" refers to systems or algorithms that analyze received media data and identify significant elements or patterns.

[0103] "Video summary" refers to a short video clip created by combining important scenes extracted from the analyzed media data.

[0104] A "storyboard" refers to a group of images that visually represent the storyboard or development of a scene in a video.

[0105] "Standard file formats" are commonly used file formats, such as MP4 for videos and JPEG for images.

[0106] "Means for encoding" refers to the process or tools used to convert the generated summary or storyboard into an appropriate file format.

[0107] "Means for transmitting to a user terminal" refers to a function or system that transmits encoded data from a server to a user's device.

[0108] "Object detection means" refers to techniques or algorithms used to identify specific objects or people within media data.

[0109] "Means for re-uploading to the system" refers to a function or process that allows a user to send a revised summary video or storyboard to the server again.

[0110] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a video summary or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[0111] Server roles and processing flow

[0112] Receiving media data

[0113] The server receives media data (images, videos, audio) uploaded by users through HTTP requests. The server-side program sets up an API endpoint using a web framework such as Python's Flask or Django and waits for the user's upload operation. The received data is stored in a temporary storage location (for example, a cloud storage service).

[0114] Data analysis and feature extraction

[0115] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. For image data, object detection and segmentation are performed, and for video data, important scenes and frames are detected. For audio data, the Google Cloud Speech-to-Text service is used to convert it to text and perform sentiment analysis.

[0116] Generate a summary video or storyboard

[0117] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips. For a storyboard, it organizes images based on the storyboard. This process can be done using, for example, FFmpeg or the Python Imaging Library (PIL).

[0118] Data Encoding

[0119] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files into MP4 or JPEG format, and adds the necessary metadata (scene descriptions and timestamps).

[0120] Send to user terminal

[0121] The server sends the encoded data to the user's terminal, using the SSL / TLS protocol to ensure security, and returns the data to the terminal via an HTTP request.

[0122] Terminal roles and processing flow

[0123] Providing a UI

[0124] The device provides a user interface (UI) for uploading media data, checking the processed results, and editing. The UI is built using front-end technologies such as HTML, CSS, and JavaScript, and is designed to be intuitive for users to operate.

[0125] Sending and Receiving Data

[0126] The device sends the media data selected by the user to the server, checks the integrity of the data, and receives the processed data from the server and displays it on the UI.

[0127] User roles and operation flow

[0128] Uploading data

[0129] Users upload their collected media data to the system using the device's UI, and also enter related information such as title and summary when uploading.

[0130] Checking and correcting the processing results

[0131] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system.

[0132] Specific examples

[0133] A concrete example of creating a movie trailer

[0134] 1. Data upload (user)

[0135] Users upload multiple video files containing scenes from their movies to the system.

[0136] 2. Data analysis (server)

[0137] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[0138] 3. Summary video generation (server)

[0139] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[0140] 4. Receiving and checking results (user)

[0141] The user can check the generated trailer on the terminal and make any necessary corrections if they are not satisfied.

[0142] 5. Corrections and final confirmation (user)

[0143] The user can add text and narration to the trailer and make adjustments to create the final trailer.

[0144] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[0145] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0146] Step 1:

[0147] Receiving media data (server)

[0148] The server receives media data such as images, videos, and audio from the user via HTTP requests. Specifically, an API endpoint is set up using a web framework such as Flask or Django, and awaits the user's upload operation. The input is the media data uploaded by the user, which is then temporarily stored on the server (for example, in a cloud storage service).

[0149] Step 2:

[0150] Data analysis and feature extraction (server)

[0151] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. The input is the stored media data, and for image data, object detection and segmentation are performed. For video data, important scenes and frames are detected. For audio data, text conversion and sentiment analysis are performed using the Google Cloud Speech-to-Text service. The output is extracted feature information.

[0152] Step 3:

[0153] Summary video or storyboard generation (server)

[0154] The server generates a video summary or storyboard based on the extracted feature information. For a video summary, it combines frames of important scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard. This process uses FFmpeg and the Python Imaging Library (PIL). The input is the feature information, and the output is the generated video summary or storyboard.

[0155] Step 4:

[0156] Data Encoding (Server)

[0157] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files to MP4 or JPEG format and adding metadata (scene descriptions and timestamps) if necessary. The input is the generated summary or storyboard, and the output is the encoded file.

[0158] Step 5:

[0159] Send to user terminal (server)

[0160] The server sends the encoded data to the user's device, using the SSL / TLS protocol to ensure security. The input is the encoded file, and the output is the data received on the user's device.

[0161] Step 6:

[0162] UI provision (terminal)

[0163] The terminal provides the user with a user interface (UI) for uploading media data, checking the processing results, and editing them. The UI is designed to be intuitive using HTML, CSS, JavaScript, etc. Specifically, it provides a file selection button and a viewer that displays the processing results.

[0164] Step 7:

[0165] Sending and receiving data (terminal)

[0166] The device sends the media data selected by the user to the server, verifies the integrity of the data, and receives the processing results returned from the server and displays them on the UI. The input is the media data selected by the user and the processing result data from the server, and the output is the processing result displayed on the user device.

[0167] Step 8:

[0168] Data Upload (User)

[0169] Users upload collected media data to the system using the device's UI. The input is the media data and related information (title, summary, etc.), which is sent to the server and stored. The output is the media data stored in a temporary storage location on the server.

[0170] Step 9:

[0171] Check and correct the processing results (user)

[0172] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system. The input is the summary video and storyboard sent from the server and the corrected data, and the output is the corrected summary video and storyboard.

[0173] Through these steps, the system is able to efficiently and accurately carry out the entire process of analyzing media data, generating summary videos and storyboards, and encoding and transmitting.

[0174] (Application example 1)

[0175] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0176] In conventional media data processing systems, users have to extract important scenes from long videos and generate video summaries, which requires a great deal of time and effort. Furthermore, there is a lack of an efficient way to share the generated summaries. Furthermore, there are issues with the security of the analyzed data and the ease with which users can edit them.

[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0178] In this invention, the server includes a means for receiving media data, a means for analyzing the received media data to extract features, and a means for generating a video summary or storyboard based on the extracted features. This enables automatic extraction of important scenes and generation of a video summary quickly and efficiently. The server also includes a means for encoding the generated video summary or storyboard and transmitting it to a user terminal, a means for using a protocol for securely transmitting the encoded data to the user terminal, a means for providing a user interface and supporting uploading of media data, confirmation of processing results, and editing, and a means for displaying the generated video summary on a smart device and sharing it on a social networking service. This allows users to easily review, modify, and securely share the generated video summary.

[0179] "Media data" refers to information stored in digital form, such as video, audio, and images.

[0180] "Analysis" refers to a series of processes that extract features from digital data and convert it into an understandable form.

[0181] A "feature" is a unique or important part of media data, including a particular part of a scene, an object, or an audio.

[0182] "Abridged video" refers to a video clip that has been shortened by extracting important scenes from a longer video.

[0183] A storyboard is a series of sketches or images that visually illustrate a plan for a film production or presentation.

[0184] "Encoding" refers to the process of converting digital data into a particular format so that it can be stored or transmitted.

[0185] A "user terminal" is an electronic device that can be directly operated by a user, and includes smartphones, tablets, computers, etc.

[0186] A "protocol" refers to the rules and procedures that define data communication between different electronic devices.

[0187] "User interface" refers to the means, screen display, and input operations that allow a user to interact with a system.

[0188] "SNS" is an abbreviation for social networking service, and refers to a platform where people can share information and interact with each other online.

[0189] "Smart devices" is a general term for modern electronic devices with internet connectivity, including smartphones, tablets, and smartwatches.

[0190] A specific embodiment of the present invention will be described. The present invention relates to a system that extracts important scenes from a long video shot by a user and automatically generates a video summary. Below, each component of the system and its processing procedure will be described.

[0191] Overall system configuration

[0192] This system consists of multiple elements, including a user terminal, a server, and a user interface. The user terminal is a portable device such as a smartphone or tablet. The server handles the main processing such as analyzing media data and generating video summaries, and a cloud server with a stable connection is suitable.

[0193] Hardware and Software

[0194] Hardware: smartphones, servers, cloud services

[0195] Software: Flask (web framework), OpenCV (image processing library), moviepy (video editing library), HTTP request library

[0196] Data analysis and video summary generation

[0197] 1. Receiving media data: The application on the user's device uploads the video the user has taken to the server using an HTTP request. In this case, Flask is used to set up an API endpoint and wait for the data to be received.

[0198] 2. Data Analysis:

[0199] The server analyzes the received video data. It uses OpenCV to sequentially read the video frames and extract important scenes from each specific frame. This analysis uses AI algorithms and machine learning models to identify scene changes and important events.

[0200] 3. Summary video generation:

[0201] Based on the analysis results, the server uses moviepy to generate video clips from important frames and stitches them together to create a summary video.

[0202] 4. Encoding and Transmission:

[0203] The server encodes the generated summary video into a common format such as MP4 and transmits it securely to the user terminal using an appropriate protocol such as SSL / TLS.

[0204] User Interface and Sharing

[0205] The user device provides a user interface for displaying the generated video summary. This interface is intuitive and designed to allow users to easily check and edit the video summary. It also includes buttons for sharing the generated video summary on social media.

[0206] Specific examples

[0207] In a specific usage scenario, when a user uploads a video (60 minutes) of a family trip, the system automatically generates a 5-minute summary video including smiling scenes and landmarks. The user can review this summary video and, if they like it, share it on social media. For example, the prompt text might look like this:

[0208] "Generate a 5-minute summary video from a 60-minute family trip, including funny moments and landmarks."

[0209] This prompt allows the generative AI model to suggest and quickly create the summary video the user desires.

[0210] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0211] Step 1:

[0212] Receiving media data:

[0213] The user device uploads the video file they have taken to the server using an HTTP request. At this time, the user selects the video file through the application and starts sending it. The server receives the request at an API endpoint set up using Flask and saves the video file on the server. The input is the video file uploaded by the user, and the output is the video file saved on the server.

[0214] Step 2:

[0215] Data Analysis:

[0216] The server analyzes the stored video files. It uses OpenCV to sequentially read the video frames and calculate the differences between each frame to extract important scenes. Specifically, it identifies important frames using algorithms such as inter-frame motion, color change, and object detection. The input is the video file stored on the server, and the output is a list of important frames.

[0217] Step 3:

[0218] Feature extraction:

[0219] The server uses AI algorithms and machine learning models to extract scene and event features from key frames, such as face recognition and object detection, and tag each frame. The input is a list of key frames, and the output is a set of frames with extracted features.

[0220] Step 4:

[0221] Summary video generation:

[0222] The server uses MoviePy to generate sub-clips for each frame from the feature-extracted frames and stitch them together to create a summary video. Specifically, it extracts short clips from the original video based on the timestamps of important frames and concatenates them into a single video. The input is a set of feature-extracted frames, and the output is a summary video.

[0223] Step 5:

[0224] Encoding:

[0225] The server encodes the generated summary video into a common video format (e.g., MP4 format). This process involves compressing the video and adding metadata. The input is the summary video, and the output is the encoded summary video file.

[0226] Step 6:

[0227] Data transmission:

[0228] The server sends the encoded summary video to the user's device using a secure protocol such as SSL / TLS to ensure data integrity and security. The input is the encoded summary video file, and the output is the video file sent to the user's device.

[0229] Step 7:

[0230] UI display and confirmation:

[0231] The user terminal displays the received summary video through a user interface. The user can visually check the generated summary video and make corrections on the screen as necessary. The input is the summary video file sent to the user terminal, and the output is the summary video displayed on the user interface.

[0232] Step 8:

[0233] Share on social media:

[0234] The user device provides a function that allows users to share the generated summary video on social media using the share button on the screen. This allows users to easily share the summary video with friends and followers. The input is the summary video displayed on the user interface, and the output is the summary video shared on the social media.

[0235] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0236] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[0237] Server roles and processing flow

[0238] 1. Receiving media data

[0239] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[0240] 2. Data analysis and feature extraction

[0241] The server analyzes the received media data. This analysis uses AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted into text using speech recognition technology, and sentiment analysis is performed. It also includes a method for detecting specific objects in the received media data.

[0242] 3. Generate a summary video or storyboard

[0243] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard frames.

[0244] 4. Data Encoding

[0245] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[0246] 5. Emotional Engine Adjustment

[0247] The server receives the emotional data provided by the user and analyzes it using an emotion engine. Based on the analysis results, it adjusts the process of generating a summary video or storyboard. Specifically, it prioritizes the selection of specific scenes or frames that correspond to the user's emotional state, optimizing the content of the video or storyboard.

[0248] 6. Sending to the user terminal

[0249] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[0250] Terminal roles and processing flow

[0251] 1. Providing UI

[0252] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[0253] 2. Sending and Receiving Data

[0254] The device sends the media data uploaded by the user to the server. It checks the integrity of the data when sending it and prompts the user to make corrections if necessary. It receives the processing result data returned from the server and displays it appropriately. It also sends the user's emotional data (e.g., voice tone, facial recognition data) to the emotion engine.

[0255] User roles and operation flow

[0256] 1. Upload your data

[0257] Users upload the media data (images, videos, audio) they have collected to the system via their devices. When uploading, they also enter related information (title, summary, etc.) and emotional data.

[0258] 2. Check and correct the processing results

[0259] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[0260] Specific examples

[0261] A concrete example of creating a movie trailer

[0262] 1. Data upload (user)

[0263] Users upload multiple video files containing scenes from their movies to the system, and also provide data for emotion recognition (e.g., emotions and voice data while watching).

[0264] 2. Data analysis (server)

[0265] The server analyzes the uploaded video files and extracts the characteristics of each scene. It detects important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that have generated a large number of positive reactions.

[0266] 3. Summary video generation (server)

[0267] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[0268] 4. Receiving and checking results (user)

[0269] The user can then review the generated trailer on their device. If they are not satisfied with the video, they can make any necessary corrections. Further adjustments can also be made based on the emotional data.

[0270] 5. Corrections and final confirmation (user)

[0271] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[0272] In this way, the present invention utilizes user emotion data to not only improve the efficiency of creating video summaries and storyboards in video production and increase the consistency between the visuals and the story, but also automatically generate optimal content that corresponds to the viewer's emotions.

[0273] The processing flow will be explained below.

[0274] Step 1:

[0275] Users upload media data (images, videos, audio) from their devices to the system. Specifically, they open a file selection dialog using the device's UI and specify the files to upload. In addition, emotion recognition data (e.g., facial expression data, voice tone data) is also uploaded at the same time.

[0276] Step 2:

[0277] The device sends the media data and emotion recognition data specified by the user to the server, which checks the integrity of the files and converts or compresses the data as necessary.

[0278] Step 3:

[0279] The server stores the media data received from the user in storage. The server-side program analyzes the file format and metadata (file name, size, type) and records them in a log.

[0280] Step 4:

[0281] The server performs an initial analysis of the received media data, applying object detection algorithms to image data, scene detection algorithms to video data, and speech recognition algorithms to audio data.

[0282] Step 5:

[0283] The server uses an emotion engine to analyze the received emotion recognition data, which involves identifying patterns in the user's facial expressions and vocal tone and identifying corresponding emotions (e.g., joy, surprise, sadness).

[0284] Step 6:

[0285] The server extracts important frames and scenes based on the analysis results, for example, flagging a scene in which the user expresses joy as an important scene.

[0286] Step 7:

[0287] The server generates a video summary or storyboard based on the extracted features and emotion data. For video summaries, key scenes are combined to create short video clips, and emotion data is incorporated to create a video that is appealing to users. For storyboards, selected scenes are arranged in frame order and explanatory text is added.

[0288] Step 8:

[0289] The server encodes the generated summary video or storyboard into a standard file format (e.g., MP4, JPEG), adding metadata (scene summary, emotion tags) during the process.

[0290] Step 9:

[0291] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[0292] Step 10:

[0293] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[0294] Step 11:

[0295] The user checks the downloaded summary video or storyboard and, if necessary, makes corrections using the editing tools on the device, making final adjustments that reflect the emotion data.

[0296] Step 12:

[0297] The user then uploads the revised data back into the system, where the server reanalyzes and encodes it to generate the final summary video or storyboard, which is then sent back to the user's device to complete the process.

[0298] Through the above steps, the present invention utilizes user emotion data to improve the efficiency and quality of video production and automatically generate optimal content that appeals to the viewer's emotions.

[0299] Example 2

[0300] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0301] Conventional media data analysis systems have not been able to fully utilize the emotional data provided by users, resulting in summary videos and storyboards that often do not match the user's emotional state. Furthermore, prioritizing specific scenes or frames, as well as the need to modify and re-upload content, are cumbersome. As a result, there are issues with the quality of the generated content and user satisfaction.

[0302] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0303] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a summary video or storyboard based on the extracted features, means for encoding the generated summary video or storyboard and transmitting it to a user terminal, means for analyzing user emotion data and optimizing the content of the summary video or storyboard based on the results, and means for preferentially selecting specific scenes or frames based on the user emotion data. This enables automatic generation of content that matches the user's emotional state, thereby improving user satisfaction.

[0304] "Media data" refers to all digital content uploaded by users, including images, videos, and audio.

[0305] The "receiving means" refers to a device or mechanism that acquires media data sent from a user via a network and stores the data in storage.

[0306] "Means for analyzing and extracting features" refers to a device or mechanism that analyzes received media data using AI algorithms or machine learning models to detect significant elements or patterns within the data.

[0307] "Means for generating video summaries or storyboards" refers to a device or mechanism that combines key scenes or elements based on the analysis results to create short video clips or visual materials in the form of storyboards.

[0308] "Means for encoding" refers to a device or mechanism that converts the generated video summary or storyboard into a standard digital file format (e.g., MP4, JPEG) and assigns metadata.

[0309] "Means for transmitting to a user terminal" refers to a device or mechanism for transmitting encoded data over a network to provide a link accessible to the user.

[0310] "Means for analyzing emotional data" refers to a device or mechanism for analyzing emotion-related data provided by a user (e.g., vocal tone, facial expression) to identify the user's emotional state.

[0311] "Means for optimizing the content of a video summary or storyboard" refers to a device or mechanism that adjusts the arrangement of scenes and frames in a video summary or storyboard based on analyzed emotional data, and generates optimal content tailored to the user's emotional state.

[0312] "Means for preferentially selecting specific scenes or frames" refers to a device or mechanism that, based on the results of emotional data analysis, preferentially selects scenes or frames that have received a large number of positive reactions from users.

[0313] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[0314] Server roles and processing flow

[0315] 1. Receiving media data

[0316] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user to upload. This process can use cloud storage such as Amazon S3.

[0317] 2. Data analysis and feature extraction

[0318] The server analyzes the received media data. It uses TensorFlow to perform object detection on image data, OpenCV to perform scene segmentation on video data, and Google Cloud Speech-to-Text API to convert audio data into text. It also performs sentiment analysis on all formats. For example, it uses the YOLO model for image analysis, keyframe extraction for video analysis, and sentiment analysis APIs for audio analysis. It extracts features of each media format from the analyzed data and handles emotional data as well.

[0319] 3. Generate a summary video or storyboard

[0320] The server generates a video summary or storyboard based on the analyzed data. When generating a video summary, it combines the extracted keyframes and important scenes to create a 30-second clip, for example. When generating a storyboard, it organizes the detected images and scenes into a storyboard format. FFmpeg is used to generate the video summary, and PIL (Python Imaging Library) is used to generate the storyboard.

[0321] 4. Data Encoding

[0322] The server encodes the generated summary video and storyboard into standard formats (MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process. FFmpeg commands are used to encode the generated video into MP4 format and also generate a JSON file containing scene timestamps and descriptions.

[0323] 5. Emotional Engine Adjustment

[0324] The server analyzes the emotional data received from the user using an emotion engine and adjusts the process of generating summary videos and storyboards. The emotion engine uses, for example, IBM Watson Emotion Analysis API to analyze the user's emotional data. Based on the results, it prioritizes scenes with positive emotions, for example.

[0325] 6. Sending to the user terminal

[0326] The server generates a link to the encoded data and sends it to the user's device as an HTTP response. The server uses a cloud storage service to generate the link, and obtains a public link for the generated file.

[0327] Terminal roles and processing flow

[0328] 1. Providing UI

[0329] The device provides an intuitive user interface for users to upload media data. It includes an upload button and a preview button. React.js is used to display a dialog for selecting and uploading media files. After the user selects a file, they press the "Upload" button, and the file is sent to the server.

[0330] 2. Sending and Receiving Data

[0331] The device sends the media data uploaded by the user to the server and receives the processing results. It checks the integrity of the data before sending and prompts for re-entry if necessary. It also sends the user's emotional data. It adds the selected file to a FormData object and sends a POST request to the server using Axios. Once processing is complete, it receives a link sent by the server and displays it to the user.

[0332] User roles and operation flow

[0333] 1. Upload your data

[0334] The user selects image, video, and audio data from their own device and uploads it to the system. At the same time, they also input emotion data. They select a file using the file selection dialog in the user interface and press the "Upload" button. They input text and information that reflects emotions (e.g., facial expression recognition data) as emotion data.

[0335] 2. Check and correct the processing results

[0336] The user watches and checks the summary video and storyboard sent from the server. They check the quality and content and make corrections if they are dissatisfied. They click a link in the device's browser to play the summary video or display the storyboard in a viewer. If necessary, they press the "Edit" button to move to the correction screen and make adjustments.

[0337] Example (creating a movie trailer)

[0338] 1. Data upload (user)

[0339] Users upload video files containing movie scenes and emotional data recorded while watching the movie to the system.

[0340] 2. Data analysis (server)

[0341] The server extracts features from video files to detect important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that evoke positive emotional responses.

[0342] 3. Summary video generation (server)

[0343] The server combines the selected scenes to generate a 30-second summary video. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[0344] 4. Receiving and checking results (user)

[0345] The user views the generated trailer and makes corrections if they are not satisfied.

[0346] 5. Corrections and final confirmation (user)

[0347] The user makes any necessary modifications (e.g., adding text or narration) to complete the final trailer.

[0348] Prompt Sentence Examples

[0349] "Generate a 30-second trailer based on the video and emotional data provided by the user. Please prioritize scenes with a high number of positive reactions from users."

[0350] As described above, the present invention utilizes user emotion data to improve the efficiency of creating video summaries and storyboards in video production, improve the consistency between visuals and storylines, and automatically generate optimal content that corresponds to the viewer's emotions.

[0351] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0352] Step 1:

[0353] Receiving media data

[0354] The server receives media data (images, videos, audio) uploaded by the user. The inputs include the file data sent by the user and an HTTP POST request. The server receives the HTTP POST request at its API endpoint (e.g., / upload) and saves the media data to storage (e.g., Amazon S3). The output provides a confirmation message that the save was successful and the path to the file.

[0355] Specific behavior:

[0356] 1. The endpoint receives an HTTP request.

[0357] 2. Analyze the file data and generate the appropriate storage path.

[0358] 3. Save the file to storage (e.g. Amazon S3).

[0359] 4. A confirmation message confirming the save is complete is returned as an HTTP response.

[0360] Step 2:

[0361] Data analysis and feature extraction

[0362] The server analyzes the stored media data. The input includes file data read from storage. Object detection is performed on image data using TensorFlow, scene segmentation is performed on video data using OpenCV, and text conversion is performed on audio data using the Google Cloud Speech-to-Text API. Sentiment analysis is also performed. Output includes extracted feature data, textual audio data, and sentiment analysis results.

[0363] Specific behavior:

[0364] 1. Read media data from storage.

[0365] 2. TensorFlow and YOLO model are used for image analysis to detect objects.

[0366] 3. For video analysis, OpenCV is used to detect scene changes and extract keyframes.

[0367] 4. The voice data is converted to text using the Google Cloud Speech-to-Text API and sent to the sentiment analysis API for sentiment analysis.

[0368] 5. Save the extracted feature data, text, and emotion data and pass them to the next processing step.

[0369] Step 3:

[0370] Generate a summary video or storyboard

[0371] The server generates a video summary or storyboard based on the analyzed feature data. The inputs include the feature data and emotion data from the analysis results. To generate the video summary, important scenes and frames are combined and a clip is created using FFmpeg. To generate the storyboard, PIL (Python Imaging Library) is used to organize images in storyboard format. The output includes the generated video summary and storyboard.

[0372] Specific behavior:

[0373] 1. Load the analysis results and select important scenes and frames.

[0374] 2. Use FFmpeg to combine the selected scenes and generate a 30-second clip.

[0375] 3. For storyboarding, we used PIL to organize the detected images into a storyboard format.

[0376] 4. Temporarily save the generated summary video and storyboard.

[0377] Step 4:

[0378] Data Encoding

[0379] The server encodes the generated summary and storyboard into standard formats (MP4, JPEG). The inputs are the summary and storyboard data. FFmpeg is used for encoding, and metadata (scene descriptions and timestamps) is added. The output contains the encoded files and the metadata.

[0380] Specific behavior:

[0381] 1. Load the saved summary video or storyboard.

[0382] 2. Encode to MP4 format using FFmpeg.

[0383] 3. Descriptions and timestamps for each scene are generated in JSON format and added as metadata.

[0384] 4. Save the encoded file.

[0385] Step 5:

[0386] Emotional engine regulation

[0387] The server adjusts the generated content based on the emotional data provided by the user. The inputs include the user's emotional data and the generated summary video and storyboard. An emotion engine (e.g., IBM Watson Emotion Analysis) analyzes the emotional state and prioritizes scenes with a high number of positive reactions. The output includes the adjusted summary video and storyboard.

[0388] Specific behavior:

[0389] 1. Analyze emotion data using an emotion engine.

[0390] 2. Based on the analysis results, the arrangement of scenes and frames in the summary video and storyboard is readjusted.

[0391] 3. Save the reworked summary footage and storyboard.

[0392] Step 6:

[0393] Send to user terminal

[0394] The server sends the encoded data to the user's device. The input includes the encoded file and its metadata. It saves the file in cloud storage and generates a link to send to the user. The output includes the link sent to the user's device.

[0395] Specific behavior:

[0396] 1. Save the encoded file to cloud storage.

[0397] 2. Generate a public link and send it to the user's device as an HTTP response.

[0398] 3. The user clicks on the link to download or watch the content.

[0399] (Application example 2)

[0400] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0401] In modern content distribution services, there is a demand for methods to provide customized video summaries and storyboards based on individual emotions to improve the user viewing experience. However, conventional systems have difficulty generating content that fully understands and reflects the viewer's emotions, making it difficult to improve viewer satisfaction.

[0402] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0403] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for evaluating the analyzed media data based on the user's emotions using an emotion engine, means for generating a summary video or storyboard based on the extracted features and the emotion evaluation, and means for encoding the generated summary video or storyboard and transmitting it to the user terminal. This enables more personalized content to be automatically generated in accordance with the user's emotion data, improving the viewing experience.

[0404] "Media data" refers to information expressed in a digital format, such as audio, images, or video.

[0405] An "emotion engine" refers to an algorithm or system that analyzes a user's emotions and adjusts or generates media data based on the results.

[0406] "Means of feature extraction" refers to techniques and methods for analyzing and identifying specific information or patterns from media data and extracting that information.

[0407] "Abridged video" refers to a video clip that has been shortened by extracting important scenes or frames from the original long video data.

[0408] A "storyboard" is a visual arrangement of images and text along the frame of a storyboard to organize the flow of a story.

[0409] "Encoding means" refers to the techniques or methods used to convert digital data into a particular form or format.

[0410] "User terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.

[0411] A "specific object" refers to a specific person, object, scene, or other element within the media data to be analyzed.

[0412] "Means for making modifications" refers to interfaces and tools that allow a user to make changes or improvements to the generated video summary or storyboard.

[0413] Server roles and processing flow

[0414] The server has a means to receive media data. For example, it uses the web application framework Flask to set up an API endpoint that receives HTTP requests from users. Media data such as audio, images, and videos are uploaded to this endpoint. The received media data is saved to disk.

[0415] The server then analyzes the received media data and extracts features. This analysis is performed using AI algorithms powered by TensorFlow. For example, image data is subjected to object detection and segmentation, video data is subjected to the detection of important frames and scenes, and audio data is subjected to text conversion using speech recognition technology. Through these analyses, specific objects contained in the data are also detected.

[0416] The server then uses an emotion engine to evaluate the analyzed media data based on the user's emotions. This allows it to extract specific scenes and frames that match the user's emotions, and generates a summary video and storyboard based on these. The generated summary video and storyboard are then converted into a standard file format (e.g., MP4, JPEG) through an encoding process.

[0417] Finally, the server generates a link to send the generated summary video and storyboard to the user's device and sends it as an HTTP response, allowing the user to check the results on their own device.

[0418] Terminal roles and processing flow

[0419] The device provides a user interface (UI) and supports uploading media data, checking the processed results, and editing. The UI is designed to be intuitive, particularly for easy visual confirmation of video and storyboards. The device also has the function of sending the media data uploaded by the user to the server.

[0420] It also has a function to receive the processing result data returned from the server and display it appropriately.It also includes a function to send the user's emotion data (e.g., voice tone, facial recognition data) to the emotion engine.

[0421] User roles and operation flow

[0422] Users upload the media data (audio, images, video) they have collected to the system via their terminal. When uploading, they also enter related information (title, summary, etc.) and emotional data. They then check the quality and content of the summary video and storyboard returned from the server. If necessary, they make corrections to the received data and re-upload it to the system.

[0423] Specific examples

[0424] For example, if a user wants to make a movie trailer more emotional, the process would be as follows:

[0425] 1. Data upload (user)

[0426] A user uploads a video file of a movie.

[0427] Input emotional data (e.g., emotional ratings during viewing).

[0428] 2. Data analysis and video summary generation (server)

[0429] The server analyzes the media data, extracts specific scenes, and generates a summary video based on the emotion data.

[0430] 3. Receive and check the results (terminal)

[0431] A download link for the generated summary video is generated and transmitted to the user terminal.

[0432] The user reviews the generated movie trailer.

[0433] Prompt Sentence Examples

[0434] "Scene analysis of movie XYZ. Provide media data for scenes that evoke positive emotions. Create a custom trailer."

[0435] As a result, the present invention enables more personalized content to be automatically generated in accordance with the user's emotional data, thereby improving the viewing experience.

[0436] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0437] Step 1:

[0438] Data upload (user)

[0439] Users upload media data (audio, images, videos) to the system using their own devices. Specifically, they select the data through the device interface and enter additional related information (title, summary) and emotional data (emotion rating, etc.). The media data and its related information are then sent to the server.

[0440] Input: User-provided audio, image, and video data, as well as related information and emotional data

[0441] Output: Media data and related information sent to the server

[0442] Step 2:

[0443] Receiving and storing media data (server)

[0444] The server receives media data uploaded by users and saves it to disk via HTTP requests. An API endpoint is set up using a web framework such as Flask, where the media data is saved.

[0445] Input: User uploaded media data

[0446] Output: Media data stored on the server's disk

[0447] Step 3:

[0448] Data analysis and feature extraction (server)

[0449] The server analyzes the received media data and extracts features. Using TensorFlow, the AI ​​model performs image and voice recognition, object detection, scene extraction, etc. It also includes a means to detect specific objects in each media data.

[0450] Input: Stored media data

[0451] Output: Extracted feature data (key scenes, audio text, detected objects)

[0452] Step 4:

[0453] Evaluation by emotion engine (server)

[0454] The emotion engine evaluates the user's emotions based on the analyzed media data. For example, an emotion recognition AI model analyzes the user's emotional data regarding a specific scene or object and outputs an evaluation score.

[0455] Input: Extracted feature data and user emotion data

[0456] Output: Score data based on the user's emotional evaluation

[0457] Step 5:

[0458] Summary video or storyboard generation (server)

[0459] The server generates a summary video or storyboard based on the feature data and emotion evaluation, stitching together important scenes and frames and composing optimal content taking into account the emotion score.

[0460] Input: feature data and emotion evaluation scores

[0461] Output: Generated summary video or storyboard

[0462] Step 6:

[0463] Encoding and sending (server)

[0464] The generated summary video or storyboard is encoded and converted into a standard file format (e.g., MP4, JPEG). A download link is generated for the encoded file to be sent to the user's device, and the link is sent as an HTTP response.

[0465] Input: Generated summary video or storyboard

[0466] Output: Encoded file and download link

[0467] Step 7:

[0468] Check and correct the results (user and device)

[0469] The user checks the summary video and storyboard generated on their device. If the video does not meet their expectations, they can make corrections and upload the corrected data back to the server. The device supports this correction process.

[0470] Input: Encoded file and download link

[0471] Output: Confirmed summary video and storyboard, and correction data as needed

[0472] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0473] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0474] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0475] [Second embodiment]

[0476] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0477] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0478] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0479] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0480] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0481] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0482] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0483] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0484] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0485] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0486] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0487] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0488] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a summary video or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[0489] Server roles and processing flow

[0490] 1. Receiving media data

[0491] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[0492] 2. Data analysis and feature extraction

[0493] The server analyzes the received media data using AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted to text using speech recognition technology, and sentiment analysis is performed.

[0494] 3. Generate a summary video or storyboard

[0495] The server generates a video summary or storyboard based on the extracted features. In the case of a video summary, it combines frames from key scenes to create short video clips, and in the case of a storyboard, it organizes images according to the storyboard frames.

[0496] 4. Data Encoding

[0497] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[0498] 5. Send to user device

[0499] The server sends the encoded data to the user terminal using an appropriate protocol (e.g., SSL / TLS) to ensure security and efficiency.

[0500] Terminal roles and processing flow

[0501] 1. Providing UI

[0502] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[0503] 2. Sending and Receiving Data

[0504] The device sends the media data uploaded by the user to the server, checks the integrity of the data when it is sent, and prompts the user to correct it if necessary, and receives the processed data returned from the server and displays it appropriately.

[0505] User roles and operation flow

[0506] 1. Upload your data

[0507] Users upload the media data (images, videos, audio) they have collected to the system via their terminals. When uploading, they also enter related information (title, summary, etc.).

[0508] 2. Check and correct the processing results

[0509] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[0510] Specific examples

[0511] A concrete example of creating a movie trailer

[0512] 1. Data upload (user)

[0513] Users upload multiple video files containing scenes from their movies to the system.

[0514] 2. Data analysis (server)

[0515] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[0516] 3. Summary video generation (server)

[0517] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[0518] 4. Receiving and checking results (user)

[0519] The user can check the generated trailer on their device and, if they are not satisfied with the footage, make any necessary corrections.

[0520] 5. Corrections and final confirmation (user)

[0521] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[0522] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[0523] The processing flow will be explained below.

[0524] Step 1:

[0525] A user uploads media data (images, videos, audio) from a device to the system. Specifically, the user opens a file selection dialog using the device's UI and specifies the file to upload. By pressing the upload button, the selected file is sent to the server.

[0526] Step 2:

[0527] The server stores the media data received from the user in storage. During this process, the server analyzes the file metadata (file name, size, type) and records it in a log.

[0528] Step 3:

[0529] The server performs an initial analysis of the received media data. Specifically, it checks the file format and scans the content, and then selects the appropriate analysis module. For image data, it uses the object detection module, and for video data, it uses the scene segmentation module.

[0530] Step 4:

[0531] The server uses the selected analysis module to extract features from the media data, for example, detecting important frames and scenes in video data, identifying key objects and scenes in image data, and converting speech to text and detecting specific keywords and emotions in audio data.

[0532] Step 5:

[0533] The server generates a summary video or storyboard based on the extracted feature information. In the case of a summary video, the extracted important scenes are arranged along a timeline and edited to a specified length. In the case of a storyboard, important scenes are arranged in each frame of the storyboard and explanations are added.

[0534] Step 6:

[0535] The server encodes the generated summary or storyboard into a standard file format, for example MP4 for the summary and JPEG or PDF for the storyboard, adding metadata (scene descriptions, timestamps) during the process.

[0536] Step 7:

[0537] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[0538] Step 8:

[0539] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[0540] Step 9:

[0541] The user checks the downloaded summary video or storyboard, and if there are any problems or omissions, they can correct them using the editing tool. Once the corrections are complete, they can upload it back to the system.

[0542] Step 10:

[0543] The server receives the modified data, re-analyzes it, and encodes it, incorporating any appropriate modifications to generate the final summary or storyboard. The resulting data is then sent back to the user's device, completing the process.

[0544] Example 1

[0545] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0546] Conventional media data analysis and editing work is often performed manually, resulting in problems of time and effort. Furthermore, the accuracy and consistency of analysis results often depend on specialized knowledge and are therefore poor. To solve these problems, the present invention aims to automate the reception, analysis, summary video or storyboard generation, encoding, and transmission of media data, enabling efficient and accurate processing.

[0547] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0548] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a video summary or storyboard based on the extracted features, means for encoding the generated video summary or storyboard into a standard file format, and means for transmitting the encoded video summary or storyboard to a user terminal. This makes it possible to automate the analysis and editing of media data and generate video summary and storyboards with high accuracy and efficiency.

[0549] "Media data" is a general term for data including image data, video data, and audio data.

[0550] The term "receiving means" refers to a function or device that can transmit media data from a user to a server and receive it on the server side.

[0551] "Analysis and feature extraction means" refers to systems or algorithms that analyze received media data and identify significant elements or patterns.

[0552] "Video summary" refers to a short video clip created by combining important scenes extracted from the analyzed media data.

[0553] A "storyboard" refers to a group of images that visually represent the storyboard or development of a scene in a video.

[0554] "Standard file formats" are commonly used file formats, such as MP4 for videos and JPEG for images.

[0555] "Means for encoding" refers to the process or tools used to convert the generated summary or storyboard into an appropriate file format.

[0556] "Means for transmitting to a user terminal" refers to a function or system that transmits encoded data from a server to a user's device.

[0557] "Object detection means" refers to techniques or algorithms used to identify specific objects or people within media data.

[0558] "Means for re-uploading to the system" refers to a function or process that allows a user to send a revised summary video or storyboard to the server again.

[0559] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a video summary or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[0560] Server roles and processing flow

[0561] Receiving media data

[0562] The server receives media data (images, videos, audio) uploaded by users through HTTP requests. The server-side program sets up an API endpoint using a web framework such as Python's Flask or Django and waits for the user's upload operation. The received data is stored in a temporary storage location (for example, a cloud storage service).

[0563] Data analysis and feature extraction

[0564] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. For image data, object detection and segmentation are performed, and for video data, important scenes and frames are detected. For audio data, the Google Cloud Speech-to-Text service is used to convert it to text and perform sentiment analysis.

[0565] Generate a summary video or storyboard

[0566] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips. For a storyboard, it organizes images based on the storyboard. This process can be done using, for example, FFmpeg or the Python Imaging Library (PIL).

[0567] Data Encoding

[0568] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files into MP4 or JPEG format, and adds the necessary metadata (scene descriptions and timestamps).

[0569] Send to user terminal

[0570] The server sends the encoded data to the user's terminal, using the SSL / TLS protocol to ensure security, and returns the data to the terminal via an HTTP request.

[0571] Terminal roles and processing flow

[0572] Providing a UI

[0573] The device provides a user interface (UI) for uploading media data, checking the processed results, and editing. The UI is built using front-end technologies such as HTML, CSS, and JavaScript, and is designed to be intuitive for users to operate.

[0574] Sending and Receiving Data

[0575] The device sends the media data selected by the user to the server, checks the integrity of the data, and receives the processed data from the server and displays it on the UI.

[0576] User roles and operation flow

[0577] Uploading data

[0578] Users upload their collected media data to the system using the device's UI, and also enter related information such as title and summary when uploading.

[0579] Checking and correcting the processing results

[0580] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system.

[0581] Specific examples

[0582] A concrete example of creating a movie trailer

[0583] 1. Data upload (user)

[0584] Users upload multiple video files containing scenes from their movies to the system.

[0585] 2. Data analysis (server)

[0586] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[0587] 3. Summary video generation (server)

[0588] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[0589] 4. Receiving and checking results (user)

[0590] The user can check the generated trailer on the terminal and make any necessary corrections if they are not satisfied.

[0591] 5. Corrections and final confirmation (user)

[0592] The user can add text and narration to the trailer and make adjustments to create the final trailer.

[0593] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[0594] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0595] Step 1:

[0596] Receiving media data (server)

[0597] The server receives media data such as images, videos, and audio from the user via HTTP requests. Specifically, an API endpoint is set up using a web framework such as Flask or Django, and awaits the user's upload operation. The input is the media data uploaded by the user, which is then temporarily stored on the server (for example, in a cloud storage service).

[0598] Step 2:

[0599] Data analysis and feature extraction (server)

[0600] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. The input is the stored media data, and for image data, object detection and segmentation are performed. For video data, important scenes and frames are detected. For audio data, text conversion and sentiment analysis are performed using the Google Cloud Speech-to-Text service. The output is extracted feature information.

[0601] Step 3:

[0602] Summary video or storyboard generation (server)

[0603] The server generates a video summary or storyboard based on the extracted feature information. For a video summary, it combines frames of important scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard. This process uses FFmpeg and the Python Imaging Library (PIL). The input is the feature information, and the output is the generated video summary or storyboard.

[0604] Step 4:

[0605] Data Encoding (Server)

[0606] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files to MP4 or JPEG format and adding metadata (scene descriptions and timestamps) if necessary. The input is the generated summary or storyboard, and the output is the encoded file.

[0607] Step 5:

[0608] Send to user terminal (server)

[0609] The server sends the encoded data to the user's device, using the SSL / TLS protocol to ensure security. The input is the encoded file, and the output is the data received on the user's device.

[0610] Step 6:

[0611] UI provision (terminal)

[0612] The terminal provides the user with a user interface (UI) for uploading media data, checking the processing results, and editing them. The UI is designed to be intuitive using HTML, CSS, JavaScript, etc. Specifically, it provides a file selection button and a viewer that displays the processing results.

[0613] Step 7:

[0614] Sending and receiving data (terminal)

[0615] The device sends the media data selected by the user to the server, verifies the integrity of the data, and receives the processing results returned from the server and displays them on the UI. The input is the media data selected by the user and the processing result data from the server, and the output is the processing result displayed on the user device.

[0616] Step 8:

[0617] Data Upload (User)

[0618] Users upload collected media data to the system using the device's UI. The input is the media data and related information (title, summary, etc.), which is sent to the server and stored. The output is the media data stored in a temporary storage location on the server.

[0619] Step 9:

[0620] Check and correct the processing results (user)

[0621] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system. The input is the summary video and storyboard sent from the server and the corrected data, and the output is the corrected summary video and storyboard.

[0622] Through these steps, the system is able to efficiently and accurately carry out the entire process of analyzing media data, generating summary videos and storyboards, and encoding and transmitting.

[0623] (Application example 1)

[0624] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0625] In conventional media data processing systems, users have to extract important scenes from long videos and generate video summaries, which requires a great deal of time and effort. Furthermore, there is a lack of an efficient way to share the generated summaries. Furthermore, there are issues with the security of the analyzed data and the ease with which users can edit them.

[0626] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0627] In this invention, the server includes a means for receiving media data, a means for analyzing the received media data to extract features, and a means for generating a video summary or storyboard based on the extracted features. This enables automatic extraction of important scenes and generation of a video summary quickly and efficiently. The server also includes a means for encoding the generated video summary or storyboard and transmitting it to a user terminal, a means for using a protocol for securely transmitting the encoded data to the user terminal, a means for providing a user interface and supporting uploading of media data, confirmation of processing results, and editing, and a means for displaying the generated video summary on a smart device and sharing it on a social networking service. This allows users to easily review, modify, and securely share the generated video summary.

[0628] "Media data" refers to information stored in digital form, such as video, audio, and images.

[0629] "Analysis" refers to a series of processes that extract features from digital data and convert it into an understandable form.

[0630] A "feature" is a unique or important part of media data, including a particular part of a scene, an object, or an audio.

[0631] "Abridged video" refers to a video clip that has been shortened by extracting important scenes from a longer video.

[0632] A storyboard is a series of sketches or images that visually illustrate a plan for a film production or presentation.

[0633] "Encoding" refers to the process of converting digital data into a particular format so that it can be stored or transmitted.

[0634] A "user terminal" is an electronic device that can be directly operated by a user, and includes smartphones, tablets, computers, etc.

[0635] A "protocol" refers to the rules and procedures that define data communication between different electronic devices.

[0636] "User interface" refers to the means, screen display, and input operations that allow a user to interact with a system.

[0637] "SNS" is an abbreviation for social networking service, and refers to a platform where people can share information and interact with each other online.

[0638] "Smart devices" is a general term for modern electronic devices with internet connectivity, including smartphones, tablets, and smartwatches.

[0639] A specific embodiment of the present invention will be described. The present invention relates to a system that extracts important scenes from a long video shot by a user and automatically generates a video summary. Below, each component of the system and its processing procedure will be described.

[0640] Overall system configuration

[0641] This system consists of multiple elements, including a user terminal, a server, and a user interface. The user terminal is a portable device such as a smartphone or tablet. The server handles the main processing such as analyzing media data and generating video summaries, and a cloud server with a stable connection is suitable.

[0642] Hardware and Software

[0643] Hardware: smartphones, servers, cloud services

[0644] Software: Flask (web framework), OpenCV (image processing library), moviepy (video editing library), HTTP request library

[0645] Data analysis and video summary generation

[0646] 1. Receiving media data: The application on the user's device uploads the video the user has taken to the server using an HTTP request. In this case, Flask is used to set up an API endpoint and wait for the data to be received.

[0647] 2. Data Analysis:

[0648] The server analyzes the received video data. It uses OpenCV to sequentially read the video frames and extract important scenes from each specific frame. This analysis uses AI algorithms and machine learning models to identify scene changes and important events.

[0649] 3. Summary video generation:

[0650] Based on the analysis results, the server uses moviepy to generate video clips from important frames and stitches them together to create a summary video.

[0651] 4. Encoding and Transmission:

[0652] The server encodes the generated summary video into a common format such as MP4 and transmits it securely to the user terminal using an appropriate protocol such as SSL / TLS.

[0653] User Interface and Sharing

[0654] The user device provides a user interface for displaying the generated video summary. This interface is intuitive and designed to allow users to easily check and edit the video summary. It also includes buttons for sharing the generated video summary on social media.

[0655] Specific examples

[0656] In a specific usage scenario, when a user uploads a video (60 minutes) of a family trip, the system automatically generates a 5-minute summary video including smiling scenes and landmarks. The user can review this summary video and, if they like it, share it on social media. For example, the prompt text might look like this:

[0657] "Generate a 5-minute summary video from a 60-minute family trip, including funny moments and landmarks."

[0658] This prompt allows the generative AI model to suggest and quickly create the summary video the user desires.

[0659] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0660] Step 1:

[0661] Receiving media data:

[0662] The user device uploads the video file they have taken to the server using an HTTP request. At this time, the user selects the video file through the application and starts sending it. The server receives the request at an API endpoint set up using Flask and saves the video file on the server. The input is the video file uploaded by the user, and the output is the video file saved on the server.

[0663] Step 2:

[0664] Data Analysis:

[0665] The server analyzes the stored video files. It uses OpenCV to sequentially read the video frames and calculate the differences between each frame to extract important scenes. Specifically, it identifies important frames using algorithms such as inter-frame motion, color change, and object detection. The input is the video file stored on the server, and the output is a list of important frames.

[0666] Step 3:

[0667] Feature extraction:

[0668] The server uses AI algorithms and machine learning models to extract scene and event features from key frames, such as face recognition and object detection, and tag each frame. The input is a list of key frames, and the output is a set of frames with extracted features.

[0669] Step 4:

[0670] Summary video generation:

[0671] The server uses MoviePy to generate sub-clips for each frame from the feature-extracted frames and stitch them together to create a summary video. Specifically, it extracts short clips from the original video based on the timestamps of important frames and concatenates them into a single video. The input is a set of feature-extracted frames, and the output is a summary video.

[0672] Step 5:

[0673] Encoding:

[0674] The server encodes the generated summary video into a common video format (e.g., MP4 format). This process involves compressing the video and adding metadata. The input is the summary video, and the output is the encoded summary video file.

[0675] Step 6:

[0676] Data transmission:

[0677] The server sends the encoded summary video to the user's device using a secure protocol such as SSL / TLS to ensure data integrity and security. The input is the encoded summary video file, and the output is the video file sent to the user's device.

[0678] Step 7:

[0679] UI display and confirmation:

[0680] The user terminal displays the received summary video through a user interface. The user can visually check the generated summary video and make corrections on the screen as necessary. The input is the summary video file sent to the user terminal, and the output is the summary video displayed on the user interface.

[0681] Step 8:

[0682] Share on social media:

[0683] The user device provides a function that allows users to share the generated summary video on social media using the share button on the screen. This allows users to easily share the summary video with friends and followers. The input is the summary video displayed on the user interface, and the output is the summary video shared on the social media.

[0684] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0685] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[0686] Server roles and processing flow

[0687] 1. Receiving media data

[0688] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[0689] 2. Data analysis and feature extraction

[0690] The server analyzes the received media data. This analysis uses AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted into text using speech recognition technology, and sentiment analysis is performed. It also includes a method for detecting specific objects in the received media data.

[0691] 3. Generate a summary video or storyboard

[0692] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard frames.

[0693] 4. Data Encoding

[0694] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[0695] 5. Emotional Engine Adjustment

[0696] The server receives the emotional data provided by the user and analyzes it using an emotion engine. Based on the analysis results, it adjusts the process of generating a summary video or storyboard. Specifically, it prioritizes the selection of specific scenes or frames that correspond to the user's emotional state, optimizing the content of the video or storyboard.

[0697] 6. Sending to the user terminal

[0698] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[0699] Terminal roles and processing flow

[0700] 1. Providing UI

[0701] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[0702] 2. Sending and Receiving Data

[0703] The device sends the media data uploaded by the user to the server. It checks the integrity of the data when sending it and prompts the user to make corrections if necessary. It receives the processing result data returned from the server and displays it appropriately. It also sends the user's emotional data (e.g., voice tone, facial recognition data) to the emotion engine.

[0704] User roles and operation flow

[0705] 1. Upload your data

[0706] Users upload the media data (images, videos, audio) they have collected to the system via their devices. When uploading, they also enter related information (title, summary, etc.) and emotional data.

[0707] 2. Check and correct the processing results

[0708] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[0709] Specific examples

[0710] A concrete example of creating a movie trailer

[0711] 1. Data upload (user)

[0712] Users upload multiple video files containing scenes from their movies to the system, and also provide data for emotion recognition (e.g., emotions and voice data while watching).

[0713] 2. Data analysis (server)

[0714] The server analyzes the uploaded video files and extracts the characteristics of each scene. It detects important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that have generated a large number of positive reactions.

[0715] 3. Summary video generation (server)

[0716] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[0717] 4. Receiving and checking results (user)

[0718] The user can then review the generated trailer on their device. If they are not satisfied with the video, they can make any necessary corrections. Further adjustments can also be made based on the emotional data.

[0719] 5. Corrections and final confirmation (user)

[0720] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[0721] In this way, the present invention utilizes user emotion data to not only improve the efficiency of creating video summaries and storyboards in video production and increase the consistency between the visuals and the story, but also automatically generate optimal content that corresponds to the viewer's emotions.

[0722] The processing flow will be explained below.

[0723] Step 1:

[0724] Users upload media data (images, videos, audio) from their devices to the system. Specifically, they open a file selection dialog using the device's UI and specify the files to upload. In addition, emotion recognition data (e.g., facial expression data, voice tone data) is also uploaded at the same time.

[0725] Step 2:

[0726] The device sends the media data and emotion recognition data specified by the user to the server, which checks the integrity of the files and converts or compresses the data as necessary.

[0727] Step 3:

[0728] The server stores the media data received from the user in storage. The server-side program analyzes the file format and metadata (file name, size, type) and records them in a log.

[0729] Step 4:

[0730] The server performs an initial analysis of the received media data, applying object detection algorithms to image data, scene detection algorithms to video data, and speech recognition algorithms to audio data.

[0731] Step 5:

[0732] The server uses an emotion engine to analyze the received emotion recognition data, which involves identifying patterns in the user's facial expressions and vocal tone and identifying corresponding emotions (e.g., joy, surprise, sadness).

[0733] Step 6:

[0734] The server extracts important frames and scenes based on the analysis results, for example, flagging a scene in which the user expresses joy as an important scene.

[0735] Step 7:

[0736] The server generates a video summary or storyboard based on the extracted features and emotion data. For video summaries, key scenes are combined to create short video clips, and emotion data is incorporated to create a video that is appealing to users. For storyboards, selected scenes are arranged in frame order and explanatory text is added.

[0737] Step 8:

[0738] The server encodes the generated summary video or storyboard into a standard file format (e.g., MP4, JPEG), adding metadata (scene summary, emotion tags) during the process.

[0739] Step 9:

[0740] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[0741] Step 10:

[0742] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[0743] Step 11:

[0744] The user checks the downloaded summary video or storyboard and, if necessary, makes corrections using the editing tools on the device, making final adjustments that reflect the emotion data.

[0745] Step 12:

[0746] The user then uploads the revised data back into the system, where the server reanalyzes and encodes it to generate the final summary video or storyboard, which is then sent back to the user's device to complete the process.

[0747] Through the above steps, the present invention utilizes user emotion data to improve the efficiency and quality of video production and automatically generate optimal content that appeals to the viewer's emotions.

[0748] Example 2

[0749] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0750] Conventional media data analysis systems have not been able to fully utilize the emotional data provided by users, resulting in summary videos and storyboards that often do not match the user's emotional state. Furthermore, prioritizing specific scenes or frames, as well as the need to modify and re-upload content, are cumbersome. As a result, there are issues with the quality of the generated content and user satisfaction.

[0751] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0752] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a summary video or storyboard based on the extracted features, means for encoding the generated summary video or storyboard and transmitting it to a user terminal, means for analyzing user emotion data and optimizing the content of the summary video or storyboard based on the results, and means for preferentially selecting specific scenes or frames based on the user emotion data. This enables automatic generation of content that matches the user's emotional state, thereby improving user satisfaction.

[0753] "Media data" refers to all digital content uploaded by users, including images, videos, and audio.

[0754] The "receiving means" refers to a device or mechanism that acquires media data sent from a user via a network and stores the data in storage.

[0755] "Means for analyzing and extracting features" refers to a device or mechanism that analyzes received media data using AI algorithms or machine learning models to detect significant elements or patterns within the data.

[0756] "Means for generating video summaries or storyboards" refers to a device or mechanism that combines key scenes or elements based on the analysis results to create short video clips or visual materials in the form of storyboards.

[0757] "Means for encoding" refers to a device or mechanism that converts the generated video summary or storyboard into a standard digital file format (e.g., MP4, JPEG) and assigns metadata.

[0758] "Means for transmitting to a user terminal" refers to a device or mechanism for transmitting encoded data over a network to provide a link accessible to the user.

[0759] "Means for analyzing emotional data" refers to a device or mechanism for analyzing emotion-related data provided by a user (e.g., vocal tone, facial expression) to identify the user's emotional state.

[0760] "Means for optimizing the content of a video summary or storyboard" refers to a device or mechanism that adjusts the arrangement of scenes and frames in a video summary or storyboard based on analyzed emotional data, and generates optimal content tailored to the user's emotional state.

[0761] "Means for preferentially selecting specific scenes or frames" refers to a device or mechanism that, based on the results of emotional data analysis, preferentially selects scenes or frames that have received a large number of positive reactions from users.

[0762] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[0763] Server roles and processing flow

[0764] 1. Receiving media data

[0765] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user to upload. This process can use cloud storage such as Amazon S3.

[0766] 2. Data analysis and feature extraction

[0767] The server analyzes the received media data. It uses TensorFlow to perform object detection on image data, OpenCV to perform scene segmentation on video data, and Google Cloud Speech-to-Text API to convert audio data into text. It also performs sentiment analysis on all formats. For example, it uses the YOLO model for image analysis, keyframe extraction for video analysis, and sentiment analysis APIs for audio analysis. It extracts features of each media format from the analyzed data and handles emotional data as well.

[0768] 3. Generate a summary video or storyboard

[0769] The server generates a video summary or storyboard based on the analyzed data. When generating a video summary, it combines the extracted keyframes and important scenes to create a 30-second clip, for example. When generating a storyboard, it organizes the detected images and scenes into a storyboard format. FFmpeg is used to generate the video summary, and PIL (Python Imaging Library) is used to generate the storyboard.

[0770] 4. Data Encoding

[0771] The server encodes the generated summary video and storyboard into standard formats (MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process. FFmpeg commands are used to encode the generated video into MP4 format and also generate a JSON file containing scene timestamps and descriptions.

[0772] 5. Emotional Engine Adjustment

[0773] The server analyzes the emotional data received from the user using an emotion engine and adjusts the process of generating summary videos and storyboards. The emotion engine uses, for example, IBM Watson Emotion Analysis API to analyze the user's emotional data. Based on the results, it prioritizes scenes with positive emotions, for example.

[0774] 6. Sending to the user terminal

[0775] The server generates a link to the encoded data and sends it to the user's device as an HTTP response. The server uses a cloud storage service to generate the link, and obtains a public link for the generated file.

[0776] Terminal roles and processing flow

[0777] 1. Providing UI

[0778] The device provides an intuitive user interface for users to upload media data. It includes an upload button and a preview button. React.js is used to display a dialog for selecting and uploading media files. After the user selects a file, they press the "Upload" button, and the file is sent to the server.

[0779] 2. Sending and Receiving Data

[0780] The device sends the media data uploaded by the user to the server and receives the processing results. It checks the integrity of the data before sending and prompts for re-entry if necessary. It also sends the user's emotional data. It adds the selected file to a FormData object and sends a POST request to the server using Axios. Once processing is complete, it receives a link sent by the server and displays it to the user.

[0781] User roles and operation flow

[0782] 1. Upload your data

[0783] The user selects image, video, and audio data from their own device and uploads it to the system. At the same time, they also input emotion data. They select a file using the file selection dialog in the user interface and press the "Upload" button. They input text and information that reflects emotions (e.g., facial expression recognition data) as emotion data.

[0784] 2. Check and correct the processing results

[0785] The user watches and checks the summary video and storyboard sent from the server. They check the quality and content and make corrections if they are dissatisfied. They click a link in the device's browser to play the summary video or display the storyboard in a viewer. If necessary, they press the "Edit" button to move to the correction screen and make adjustments.

[0786] Example (creating a movie trailer)

[0787] 1. Data upload (user)

[0788] Users upload video files containing movie scenes and emotional data recorded while watching the movie to the system.

[0789] 2. Data analysis (server)

[0790] The server extracts features from video files to detect important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that evoke positive emotional responses.

[0791] 3. Summary video generation (server)

[0792] The server combines the selected scenes to generate a 30-second summary video. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[0793] 4. Receiving and checking results (user)

[0794] The user views the generated trailer and makes corrections if they are not satisfied.

[0795] 5. Corrections and final confirmation (user)

[0796] The user makes any necessary modifications (e.g., adding text or narration) to complete the final trailer.

[0797] Prompt Sentence Examples

[0798] "Generate a 30-second trailer based on the video and emotional data provided by the user. Please prioritize scenes with a high number of positive reactions from users."

[0799] As described above, the present invention utilizes user emotion data to improve the efficiency of creating video summaries and storyboards in video production, improve the consistency between visuals and storylines, and automatically generate optimal content that corresponds to the viewer's emotions.

[0800] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0801] Step 1:

[0802] Receiving media data

[0803] The server receives media data (images, videos, audio) uploaded by the user. The inputs include the file data sent by the user and an HTTP POST request. The server receives the HTTP POST request at its API endpoint (e.g., / upload) and saves the media data to storage (e.g., Amazon S3). The output provides a confirmation message that the save was successful and the path to the file.

[0804] Specific behavior:

[0805] 1. The endpoint receives an HTTP request.

[0806] 2. Analyze the file data and generate the appropriate storage path.

[0807] 3. Save the file to storage (e.g. Amazon S3).

[0808] 4. A confirmation message confirming the save is complete is returned as an HTTP response.

[0809] Step 2:

[0810] Data analysis and feature extraction

[0811] The server analyzes the stored media data. The input includes file data read from storage. Object detection is performed on image data using TensorFlow, scene segmentation is performed on video data using OpenCV, and text conversion is performed on audio data using the Google Cloud Speech-to-Text API. Sentiment analysis is also performed. Output includes extracted feature data, textual audio data, and sentiment analysis results.

[0812] Specific behavior:

[0813] 1. Read media data from storage.

[0814] 2. TensorFlow and YOLO model are used for image analysis to detect objects.

[0815] 3. For video analysis, OpenCV is used to detect scene changes and extract keyframes.

[0816] 4. The voice data is converted to text using the Google Cloud Speech-to-Text API and sent to the sentiment analysis API for sentiment analysis.

[0817] 5. Save the extracted feature data, text, and emotion data and pass them to the next processing step.

[0818] Step 3:

[0819] Generate a summary video or storyboard

[0820] The server generates a video summary or storyboard based on the analyzed feature data. The inputs include the feature data and emotion data from the analysis results. To generate the video summary, important scenes and frames are combined and a clip is created using FFmpeg. To generate the storyboard, PIL (Python Imaging Library) is used to organize images in storyboard format. The output includes the generated video summary and storyboard.

[0821] Specific behavior:

[0822] 1. Load the analysis results and select important scenes and frames.

[0823] 2. Use FFmpeg to combine the selected scenes and generate a 30-second clip.

[0824] 3. For storyboarding, we used PIL to organize the detected images into a storyboard format.

[0825] 4. Temporarily save the generated summary video and storyboard.

[0826] Step 4:

[0827] Data Encoding

[0828] The server encodes the generated summary and storyboard into standard formats (MP4, JPEG). The inputs are the summary and storyboard data. FFmpeg is used for encoding, and metadata (scene descriptions and timestamps) is added. The output contains the encoded files and the metadata.

[0829] Specific behavior:

[0830] 1. Load the saved summary video or storyboard.

[0831] 2. Encode to MP4 format using FFmpeg.

[0832] 3. Descriptions and timestamps for each scene are generated in JSON format and added as metadata.

[0833] 4. Save the encoded file.

[0834] Step 5:

[0835] Emotional engine regulation

[0836] The server adjusts the generated content based on the emotional data provided by the user. The inputs include the user's emotional data and the generated summary video and storyboard. An emotion engine (e.g., IBM Watson Emotion Analysis) analyzes the emotional state and prioritizes scenes with a high number of positive reactions. The output includes the adjusted summary video and storyboard.

[0837] Specific behavior:

[0838] 1. Analyze emotion data using an emotion engine.

[0839] 2. Based on the analysis results, the arrangement of scenes and frames in the summary video and storyboard is readjusted.

[0840] 3. Save the reworked summary footage and storyboard.

[0841] Step 6:

[0842] Send to user terminal

[0843] The server sends the encoded data to the user's device. The input includes the encoded file and its metadata. It saves the file in cloud storage and generates a link to send to the user. The output includes the link sent to the user's device.

[0844] Specific behavior:

[0845] 1. Save the encoded file to cloud storage.

[0846] 2. Generate a public link and send it to the user's device as an HTTP response.

[0847] 3. The user clicks on the link to download or watch the content.

[0848] (Application example 2)

[0849] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0850] In modern content distribution services, there is a demand for methods to provide customized video summaries and storyboards based on individual emotions to improve the user viewing experience. However, conventional systems have difficulty generating content that fully understands and reflects the viewer's emotions, making it difficult to improve viewer satisfaction.

[0851] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0852] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for evaluating the analyzed media data based on the user's emotions using an emotion engine, means for generating a summary video or storyboard based on the extracted features and the emotion evaluation, and means for encoding the generated summary video or storyboard and transmitting it to the user terminal. This enables more personalized content to be automatically generated in accordance with the user's emotion data, improving the viewing experience.

[0853] "Media data" refers to information expressed in a digital format, such as audio, images, or video.

[0854] An "emotion engine" refers to an algorithm or system that analyzes a user's emotions and adjusts or generates media data based on the results.

[0855] "Means of feature extraction" refers to techniques and methods for analyzing and identifying specific information or patterns from media data and extracting that information.

[0856] "Abridged video" refers to a video clip that has been shortened by extracting important scenes or frames from the original long video data.

[0857] A "storyboard" is a visual arrangement of images and text along the frame of a storyboard to organize the flow of a story.

[0858] "Encoding means" refers to the techniques or methods used to convert digital data into a particular form or format.

[0859] "User terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.

[0860] A "specific object" refers to a specific person, object, scene, or other element within the media data to be analyzed.

[0861] "Means for making modifications" refers to interfaces and tools that allow a user to make changes or improvements to the generated video summary or storyboard.

[0862] Server roles and processing flow

[0863] The server has a means to receive media data. For example, it uses the web application framework Flask to set up an API endpoint that receives HTTP requests from users. Media data such as audio, images, and videos are uploaded to this endpoint. The received media data is saved to disk.

[0864] The server then analyzes the received media data and extracts features. This analysis is performed using AI algorithms powered by TensorFlow. For example, image data is subjected to object detection and segmentation, video data is subjected to the detection of important frames and scenes, and audio data is subjected to text conversion using speech recognition technology. Through these analyses, specific objects contained in the data are also detected.

[0865] The server then uses an emotion engine to evaluate the analyzed media data based on the user's emotions. This allows it to extract specific scenes and frames that match the user's emotions, and generates a summary video and storyboard based on these. The generated summary video and storyboard are then converted into a standard file format (e.g., MP4, JPEG) through an encoding process.

[0866] Finally, the server generates a link to send the generated summary video and storyboard to the user's device and sends it as an HTTP response, allowing the user to check the results on their own device.

[0867] Terminal roles and processing flow

[0868] The device provides a user interface (UI) and supports uploading media data, checking the processed results, and editing. The UI is designed to be intuitive, particularly for easy visual confirmation of video and storyboards. The device also has the function of sending the media data uploaded by the user to the server.

[0869] It also has a function to receive the processing result data returned from the server and display it appropriately.It also includes a function to send the user's emotion data (e.g., voice tone, facial recognition data) to the emotion engine.

[0870] User roles and operation flow

[0871] Users upload the media data (audio, images, video) they have collected to the system via their terminal. When uploading, they also enter related information (title, summary, etc.) and emotional data. They then check the quality and content of the summary video and storyboard returned from the server. If necessary, they make corrections to the received data and re-upload it to the system.

[0872] Specific examples

[0873] For example, if a user wants to make a movie trailer more emotional, the process would be as follows:

[0874] 1. Data upload (user)

[0875] A user uploads a video file of a movie.

[0876] Input emotional data (e.g., emotional ratings during viewing).

[0877] 2. Data analysis and video summary generation (server)

[0878] The server analyzes the media data, extracts specific scenes, and generates a summary video based on the emotion data.

[0879] 3. Receive and check the results (terminal)

[0880] A download link for the generated summary video is generated and transmitted to the user terminal.

[0881] The user reviews the generated movie trailer.

[0882] Prompt Sentence Examples

[0883] "Scene analysis of movie XYZ. Provide media data for scenes that evoke positive emotions. Create a custom trailer."

[0884] As a result, the present invention enables more personalized content to be automatically generated in accordance with the user's emotional data, thereby improving the viewing experience.

[0885] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0886] Step 1:

[0887] Data upload (user)

[0888] Users upload media data (audio, images, videos) to the system using their own devices. Specifically, they select the data through the device interface and enter additional related information (title, summary) and emotional data (emotion rating, etc.). The media data and its related information are then sent to the server.

[0889] Input: User-provided audio, image, and video data, as well as related information and emotional data

[0890] Output: Media data and related information sent to the server

[0891] Step 2:

[0892] Receiving and storing media data (server)

[0893] The server receives media data uploaded by users and saves it to disk via HTTP requests. An API endpoint is set up using a web framework such as Flask, where the media data is saved.

[0894] Input: User uploaded media data

[0895] Output: Media data stored on the server's disk

[0896] Step 3:

[0897] Data analysis and feature extraction (server)

[0898] The server analyzes the received media data and extracts features. Using TensorFlow, the AI ​​model performs image and voice recognition, object detection, scene extraction, etc. It also includes a means to detect specific objects in each media data.

[0899] Input: Stored media data

[0900] Output: Extracted feature data (key scenes, audio text, detected objects)

[0901] Step 4:

[0902] Evaluation by emotion engine (server)

[0903] The emotion engine evaluates the user's emotions based on the analyzed media data. For example, an emotion recognition AI model analyzes the user's emotional data regarding a specific scene or object and outputs an evaluation score.

[0904] Input: Extracted feature data and user emotion data

[0905] Output: Score data based on the user's emotional evaluation

[0906] Step 5:

[0907] Summary video or storyboard generation (server)

[0908] The server generates a summary video or storyboard based on the feature data and emotion evaluation, stitching together important scenes and frames and composing optimal content taking into account the emotion score.

[0909] Input: feature data and emotion evaluation scores

[0910] Output: Generated summary video or storyboard

[0911] Step 6:

[0912] Encoding and sending (server)

[0913] The generated summary video or storyboard is encoded and converted into a standard file format (e.g., MP4, JPEG). A download link is generated for the encoded file to be sent to the user's device, and the link is sent as an HTTP response.

[0914] Input: Generated summary video or storyboard

[0915] Output: Encoded file and download link

[0916] Step 7:

[0917] Check and correct the results (user and device)

[0918] The user checks the summary video and storyboard generated on their device. If the video does not meet their expectations, they can make corrections and upload the corrected data back to the server. The device supports this correction process.

[0919] Input: Encoded file and download link

[0920] Output: Confirmed summary video and storyboard, and correction data as needed

[0921] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0922] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0923] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0924] [Third embodiment]

[0925] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0926] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0927] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0928] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0929] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0930] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0931] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0932] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0933] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0934] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0935] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0936] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0937] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a summary video or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[0938] Server roles and processing flow

[0939] 1. Receiving media data

[0940] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[0941] 2. Data analysis and feature extraction

[0942] The server analyzes the received media data using AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted to text using speech recognition technology, and sentiment analysis is performed.

[0943] 3. Generate a summary video or storyboard

[0944] The server generates a video summary or storyboard based on the extracted features. In the case of a video summary, it combines frames from key scenes to create short video clips, and in the case of a storyboard, it organizes images according to the storyboard frames.

[0945] 4. Data Encoding

[0946] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[0947] 5. Send to user device

[0948] The server sends the encoded data to the user terminal using an appropriate protocol (e.g., SSL / TLS) to ensure security and efficiency.

[0949] Terminal roles and processing flow

[0950] 1. Providing UI

[0951] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[0952] 2. Sending and Receiving Data

[0953] The device sends the media data uploaded by the user to the server, checks the integrity of the data when it is sent, and prompts the user to correct it if necessary, and receives the processed data returned from the server and displays it appropriately.

[0954] User roles and operation flow

[0955] 1. Upload your data

[0956] Users upload the media data (images, videos, audio) they have collected to the system via their terminals. When uploading, they also enter related information (title, summary, etc.).

[0957] 2. Check and correct the processing results

[0958] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[0959] Specific examples

[0960] A concrete example of creating a movie trailer

[0961] 1. Data upload (user)

[0962] Users upload multiple video files containing scenes from their movies to the system.

[0963] 2. Data analysis (server)

[0964] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[0965] 3. Summary video generation (server)

[0966] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[0967] 4. Receiving and checking results (user)

[0968] The user can check the generated trailer on their device and, if they are not satisfied with the footage, make any necessary corrections.

[0969] 5. Corrections and final confirmation (user)

[0970] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[0971] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[0972] The processing flow will be explained below.

[0973] Step 1:

[0974] A user uploads media data (images, videos, audio) from a device to the system. Specifically, the user opens a file selection dialog using the device's UI and specifies the file to upload. By pressing the upload button, the selected file is sent to the server.

[0975] Step 2:

[0976] The server stores the media data received from the user in storage. During this process, the server analyzes the file metadata (file name, size, type) and records it in a log.

[0977] Step 3:

[0978] The server performs an initial analysis of the received media data. Specifically, it checks the file format and scans the content, and then selects the appropriate analysis module. For image data, it uses the object detection module, and for video data, it uses the scene segmentation module.

[0979] Step 4:

[0980] The server uses the selected analysis module to extract features from the media data, for example, detecting important frames and scenes in video data, identifying key objects and scenes in image data, and converting speech to text and detecting specific keywords and emotions in audio data.

[0981] Step 5:

[0982] The server generates a summary video or storyboard based on the extracted feature information. In the case of a summary video, the extracted important scenes are arranged along a timeline and edited to a specified length. In the case of a storyboard, important scenes are arranged in each frame of the storyboard and explanations are added.

[0983] Step 6:

[0984] The server encodes the generated summary or storyboard into a standard file format, for example MP4 for the summary and JPEG or PDF for the storyboard, adding metadata (scene descriptions, timestamps) during the process.

[0985] Step 7:

[0986] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[0987] Step 8:

[0988] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[0989] Step 9:

[0990] The user checks the downloaded summary video or storyboard, and if there are any problems or omissions, they can correct them using the editing tool. Once the corrections are complete, they can upload it back to the system.

[0991] Step 10:

[0992] The server receives the modified data, re-analyzes it, and encodes it, incorporating any appropriate modifications to generate the final summary or storyboard. The resulting data is then sent back to the user's device, completing the process.

[0993] Example 1

[0994] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0995] Conventional media data analysis and editing work is often performed manually, resulting in problems of time and effort. Furthermore, the accuracy and consistency of analysis results often depend on specialized knowledge and are therefore poor. To solve these problems, the present invention aims to automate the reception, analysis, summary video or storyboard generation, encoding, and transmission of media data, enabling efficient and accurate processing.

[0996] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0997] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a video summary or storyboard based on the extracted features, means for encoding the generated video summary or storyboard into a standard file format, and means for transmitting the encoded video summary or storyboard to a user terminal. This makes it possible to automate the analysis and editing of media data and generate video summary and storyboards with high accuracy and efficiency.

[0998] "Media data" is a general term for data including image data, video data, and audio data.

[0999] The term "receiving means" refers to a function or device that can transmit media data from a user to a server and receive it on the server side.

[1000] "Analysis and feature extraction means" refers to systems or algorithms that analyze received media data and identify significant elements or patterns.

[1001] "Video summary" refers to a short video clip created by combining important scenes extracted from the analyzed media data.

[1002] A "storyboard" refers to a group of images that visually represent the storyboard or development of a scene in a video.

[1003] "Standard file formats" are commonly used file formats, such as MP4 for videos and JPEG for images.

[1004] "Means for encoding" refers to the process or tools used to convert the generated summary or storyboard into an appropriate file format.

[1005] "Means for transmitting to a user terminal" refers to a function or system that transmits encoded data from a server to a user's device.

[1006] "Object detection means" refers to techniques or algorithms used to identify specific objects or people within media data.

[1007] "Means for re-uploading to the system" refers to a function or process that allows a user to send a revised summary video or storyboard to the server again.

[1008] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a video summary or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[1009] Server roles and processing flow

[1010] Receiving media data

[1011] The server receives media data (images, videos, audio) uploaded by users through HTTP requests. The server-side program sets up an API endpoint using a web framework such as Python's Flask or Django and waits for the user's upload operation. The received data is stored in a temporary storage location (for example, a cloud storage service).

[1012] Data analysis and feature extraction

[1013] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. For image data, object detection and segmentation are performed, and for video data, important scenes and frames are detected. For audio data, the Google Cloud Speech-to-Text service is used to convert it to text and perform sentiment analysis.

[1014] Generate a summary video or storyboard

[1015] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips. For a storyboard, it organizes images based on the storyboard. This process can be done using, for example, FFmpeg or the Python Imaging Library (PIL).

[1016] Data Encoding

[1017] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files into MP4 or JPEG format, and adds the necessary metadata (scene descriptions and timestamps).

[1018] Send to user terminal

[1019] The server sends the encoded data to the user's terminal, using the SSL / TLS protocol to ensure security, and returns the data to the terminal via an HTTP request.

[1020] Terminal roles and processing flow

[1021] Providing a UI

[1022] The device provides a user interface (UI) for uploading media data, checking the processed results, and editing. The UI is built using front-end technologies such as HTML, CSS, and JavaScript, and is designed to be intuitive for users to operate.

[1023] Sending and Receiving Data

[1024] The device sends the media data selected by the user to the server, checks the integrity of the data, and receives the processed data from the server and displays it on the UI.

[1025] User roles and operation flow

[1026] Uploading data

[1027] Users upload their collected media data to the system using the device's UI, and also enter related information such as title and summary when uploading.

[1028] Checking and correcting the processing results

[1029] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system.

[1030] Specific examples

[1031] A concrete example of creating a movie trailer

[1032] 1. Data upload (user)

[1033] Users upload multiple video files containing scenes from their movies to the system.

[1034] 2. Data analysis (server)

[1035] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[1036] 3. Summary video generation (server)

[1037] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[1038] 4. Receiving and checking results (user)

[1039] The user can check the generated trailer on the terminal and make any necessary corrections if they are not satisfied.

[1040] 5. Corrections and final confirmation (user)

[1041] The user can add text and narration to the trailer and make adjustments to create the final trailer.

[1042] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[1043] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1044] Step 1:

[1045] Receiving media data (server)

[1046] The server receives media data such as images, videos, and audio from the user via HTTP requests. Specifically, an API endpoint is set up using a web framework such as Flask or Django, and awaits the user's upload operation. The input is the media data uploaded by the user, which is then temporarily stored on the server (for example, in a cloud storage service).

[1047] Step 2:

[1048] Data analysis and feature extraction (server)

[1049] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. The input is the stored media data, and for image data, object detection and segmentation are performed. For video data, important scenes and frames are detected. For audio data, text conversion and sentiment analysis are performed using the Google Cloud Speech-to-Text service. The output is extracted feature information.

[1050] Step 3:

[1051] Summary video or storyboard generation (server)

[1052] The server generates a video summary or storyboard based on the extracted feature information. For a video summary, it combines frames of important scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard. This process uses FFmpeg and the Python Imaging Library (PIL). The input is the feature information, and the output is the generated video summary or storyboard.

[1053] Step 4:

[1054] Data Encoding (Server)

[1055] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files to MP4 or JPEG format and adding metadata (scene descriptions and timestamps) if necessary. The input is the generated summary or storyboard, and the output is the encoded file.

[1056] Step 5:

[1057] Send to user terminal (server)

[1058] The server sends the encoded data to the user's device, using the SSL / TLS protocol to ensure security. The input is the encoded file, and the output is the data received on the user's device.

[1059] Step 6:

[1060] UI provision (terminal)

[1061] The terminal provides the user with a user interface (UI) for uploading media data, checking the processing results, and editing them. The UI is designed to be intuitive using HTML, CSS, JavaScript, etc. Specifically, it provides a file selection button and a viewer that displays the processing results.

[1062] Step 7:

[1063] Sending and receiving data (terminal)

[1064] The device sends the media data selected by the user to the server, verifies the integrity of the data, and receives the processing results returned from the server and displays them on the UI. The input is the media data selected by the user and the processing result data from the server, and the output is the processing result displayed on the user device.

[1065] Step 8:

[1066] Data Upload (User)

[1067] Users upload collected media data to the system using the device's UI. The input is the media data and related information (title, summary, etc.), which is sent to the server and stored. The output is the media data stored in a temporary storage location on the server.

[1068] Step 9:

[1069] Check and correct the processing results (user)

[1070] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system. The input is the summary video and storyboard sent from the server and the corrected data, and the output is the corrected summary video and storyboard.

[1071] Through these steps, the system is able to efficiently and accurately carry out the entire process of analyzing media data, generating summary videos and storyboards, and encoding and transmitting.

[1072] (Application example 1)

[1073] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1074] In conventional media data processing systems, users have to extract important scenes from long videos and generate video summaries, which requires a great deal of time and effort. Furthermore, there is a lack of an efficient way to share the generated summaries. Furthermore, there are issues with the security of the analyzed data and the ease with which users can edit them.

[1075] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1076] In this invention, the server includes a means for receiving media data, a means for analyzing the received media data to extract features, and a means for generating a video summary or storyboard based on the extracted features. This enables automatic extraction of important scenes and generation of a video summary quickly and efficiently. The server also includes a means for encoding the generated video summary or storyboard and transmitting it to a user terminal, a means for using a protocol for securely transmitting the encoded data to the user terminal, a means for providing a user interface and supporting uploading of media data, confirmation of processing results, and editing, and a means for displaying the generated video summary on a smart device and sharing it on a social networking service. This allows users to easily review, modify, and securely share the generated video summary.

[1077] "Media data" refers to information stored in digital form, such as video, audio, and images.

[1078] "Analysis" refers to a series of processes that extract features from digital data and convert it into an understandable form.

[1079] A "feature" is a unique or important part of media data, including a particular part of a scene, an object, or an audio.

[1080] "Abridged video" refers to a video clip that has been shortened by extracting important scenes from a longer video.

[1081] A storyboard is a series of sketches or images that visually illustrate a plan for a film production or presentation.

[1082] "Encoding" refers to the process of converting digital data into a particular format so that it can be stored or transmitted.

[1083] A "user terminal" is an electronic device that can be directly operated by a user, and includes smartphones, tablets, computers, etc.

[1084] A "protocol" refers to the rules and procedures that define data communication between different electronic devices.

[1085] "User interface" refers to the means, screen display, and input operations that allow a user to interact with a system.

[1086] "SNS" is an abbreviation for social networking service, and refers to a platform where people can share information and interact with each other online.

[1087] "Smart devices" is a general term for modern electronic devices with internet connectivity, including smartphones, tablets, and smartwatches.

[1088] A specific embodiment of the present invention will be described. The present invention relates to a system that extracts important scenes from a long video shot by a user and automatically generates a video summary. Below, each component of the system and its processing procedure will be described.

[1089] Overall system configuration

[1090] This system consists of multiple elements, including a user terminal, a server, and a user interface. The user terminal is a portable device such as a smartphone or tablet. The server handles the main processing such as analyzing media data and generating video summaries, and a cloud server with a stable connection is suitable.

[1091] Hardware and Software

[1092] Hardware: smartphones, servers, cloud services

[1093] Software: Flask (web framework), OpenCV (image processing library), moviepy (video editing library), HTTP request library

[1094] Data analysis and video summary generation

[1095] 1. Receiving media data: The application on the user's device uploads the video the user has taken to the server using an HTTP request. In this case, Flask is used to set up an API endpoint and wait for the data to be received.

[1096] 2. Data Analysis:

[1097] The server analyzes the received video data. It uses OpenCV to sequentially read the video frames and extract important scenes from each specific frame. This analysis uses AI algorithms and machine learning models to identify scene changes and important events.

[1098] 3. Summary video generation:

[1099] Based on the analysis results, the server uses moviepy to generate video clips from important frames and stitches them together to create a summary video.

[1100] 4. Encoding and Transmission:

[1101] The server encodes the generated summary video into a common format such as MP4 and transmits it securely to the user terminal using an appropriate protocol such as SSL / TLS.

[1102] User Interface and Sharing

[1103] The user device provides a user interface for displaying the generated video summary. This interface is intuitive and designed to allow users to easily check and edit the video summary. It also includes buttons for sharing the generated video summary on social media.

[1104] Specific examples

[1105] In a specific usage scenario, when a user uploads a video (60 minutes) of a family trip, the system automatically generates a 5-minute summary video including smiling scenes and landmarks. The user can review this summary video and, if they like it, share it on social media. For example, the prompt text might look like this:

[1106] "Generate a 5-minute summary video from a 60-minute family trip, including funny moments and landmarks."

[1107] This prompt allows the generative AI model to suggest and quickly create the summary video the user desires.

[1108] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1109] Step 1:

[1110] Receiving media data:

[1111] The user device uploads the video file they have taken to the server using an HTTP request. At this time, the user selects the video file through the application and starts sending it. The server receives the request at an API endpoint set up using Flask and saves the video file on the server. The input is the video file uploaded by the user, and the output is the video file saved on the server.

[1112] Step 2:

[1113] Data Analysis:

[1114] The server analyzes the stored video files. It uses OpenCV to sequentially read the video frames and calculate the differences between each frame to extract important scenes. Specifically, it identifies important frames using algorithms such as inter-frame motion, color change, and object detection. The input is the video file stored on the server, and the output is a list of important frames.

[1115] Step 3:

[1116] Feature extraction:

[1117] The server uses AI algorithms and machine learning models to extract scene and event features from key frames, such as face recognition and object detection, and tag each frame. The input is a list of key frames, and the output is a set of frames with extracted features.

[1118] Step 4:

[1119] Summary video generation:

[1120] The server uses MoviePy to generate sub-clips for each frame from the feature-extracted frames and stitch them together to create a summary video. Specifically, it extracts short clips from the original video based on the timestamps of important frames and concatenates them into a single video. The input is a set of feature-extracted frames, and the output is a summary video.

[1121] Step 5:

[1122] Encoding:

[1123] The server encodes the generated summary video into a common video format (e.g., MP4 format). This process involves compressing the video and adding metadata. The input is the summary video, and the output is the encoded summary video file.

[1124] Step 6:

[1125] Data transmission:

[1126] The server sends the encoded summary video to the user's device using a secure protocol such as SSL / TLS to ensure data integrity and security. The input is the encoded summary video file, and the output is the video file sent to the user's device.

[1127] Step 7:

[1128] UI display and confirmation:

[1129] The user terminal displays the received summary video through a user interface. The user can visually check the generated summary video and make corrections on the screen as necessary. The input is the summary video file sent to the user terminal, and the output is the summary video displayed on the user interface.

[1130] Step 8:

[1131] Share on social media:

[1132] The user device provides a function that allows users to share the generated summary video on social media using the share button on the screen. This allows users to easily share the summary video with friends and followers. The input is the summary video displayed on the user interface, and the output is the summary video shared on the social media.

[1133] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1134] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[1135] Server roles and processing flow

[1136] 1. Receiving media data

[1137] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[1138] 2. Data analysis and feature extraction

[1139] The server analyzes the received media data. This analysis uses AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted into text using speech recognition technology, and sentiment analysis is performed. It also includes a method for detecting specific objects in the received media data.

[1140] 3. Generate a summary video or storyboard

[1141] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard frames.

[1142] 4. Data Encoding

[1143] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[1144] 5. Emotional Engine Adjustment

[1145] The server receives the emotional data provided by the user and analyzes it using an emotion engine. Based on the analysis results, it adjusts the process of generating a summary video or storyboard. Specifically, it prioritizes the selection of specific scenes or frames that correspond to the user's emotional state, optimizing the content of the video or storyboard.

[1146] 6. Sending to the user terminal

[1147] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[1148] Terminal roles and processing flow

[1149] 1. Providing UI

[1150] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[1151] 2. Sending and Receiving Data

[1152] The device sends the media data uploaded by the user to the server. It checks the integrity of the data when sending it and prompts the user to make corrections if necessary. It receives the processing result data returned from the server and displays it appropriately. It also sends the user's emotional data (e.g., voice tone, facial recognition data) to the emotion engine.

[1153] User roles and operation flow

[1154] 1. Upload your data

[1155] Users upload the media data (images, videos, audio) they have collected to the system via their devices. When uploading, they also enter related information (title, summary, etc.) and emotional data.

[1156] 2. Check and correct the processing results

[1157] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[1158] Specific examples

[1159] A concrete example of creating a movie trailer

[1160] 1. Data upload (user)

[1161] Users upload multiple video files containing scenes from their movies to the system, and also provide data for emotion recognition (e.g., emotions and voice data while watching).

[1162] 2. Data analysis (server)

[1163] The server analyzes the uploaded video files and extracts the characteristics of each scene. It detects important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that have generated a large number of positive reactions.

[1164] 3. Summary video generation (server)

[1165] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[1166] 4. Receiving and checking results (user)

[1167] The user can then review the generated trailer on their device. If they are not satisfied with the video, they can make any necessary corrections. Further adjustments can also be made based on the emotional data.

[1168] 5. Corrections and final confirmation (user)

[1169] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[1170] In this way, the present invention utilizes user emotion data to not only improve the efficiency of creating video summaries and storyboards in video production and increase the consistency between the visuals and the story, but also automatically generate optimal content that corresponds to the viewer's emotions.

[1171] The processing flow will be explained below.

[1172] Step 1:

[1173] Users upload media data (images, videos, audio) from their devices to the system. Specifically, they open a file selection dialog using the device's UI and specify the files to upload. In addition, emotion recognition data (e.g., facial expression data, voice tone data) is also uploaded at the same time.

[1174] Step 2:

[1175] The device sends the media data and emotion recognition data specified by the user to the server, which checks the integrity of the files and converts or compresses the data as necessary.

[1176] Step 3:

[1177] The server stores the media data received from the user in storage. The server-side program analyzes the file format and metadata (file name, size, type) and records them in a log.

[1178] Step 4:

[1179] The server performs an initial analysis of the received media data, applying object detection algorithms to image data, scene detection algorithms to video data, and speech recognition algorithms to audio data.

[1180] Step 5:

[1181] The server uses an emotion engine to analyze the received emotion recognition data, which involves identifying patterns in the user's facial expressions and vocal tone and identifying corresponding emotions (e.g., joy, surprise, sadness).

[1182] Step 6:

[1183] The server extracts important frames and scenes based on the analysis results, for example, flagging a scene in which the user expresses joy as an important scene.

[1184] Step 7:

[1185] The server generates a video summary or storyboard based on the extracted features and emotion data. For video summaries, key scenes are combined to create short video clips, and emotion data is incorporated to create a video that is appealing to users. For storyboards, selected scenes are arranged in frame order and explanatory text is added.

[1186] Step 8:

[1187] The server encodes the generated summary video or storyboard into a standard file format (e.g., MP4, JPEG), adding metadata (scene summary, emotion tags) during the process.

[1188] Step 9:

[1189] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[1190] Step 10:

[1191] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[1192] Step 11:

[1193] The user checks the downloaded summary video or storyboard and, if necessary, makes corrections using the editing tools on the device, making final adjustments that reflect the emotion data.

[1194] Step 12:

[1195] The user then uploads the revised data back into the system, where the server reanalyzes and encodes it to generate the final summary video or storyboard, which is then sent back to the user's device to complete the process.

[1196] Through the above steps, the present invention utilizes user emotion data to improve the efficiency and quality of video production and automatically generate optimal content that appeals to the viewer's emotions.

[1197] Example 2

[1198] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1199] Conventional media data analysis systems have not been able to fully utilize the emotional data provided by users, resulting in summary videos and storyboards that often do not match the user's emotional state. Furthermore, prioritizing specific scenes or frames, as well as the need to modify and re-upload content, are cumbersome. As a result, there are issues with the quality of the generated content and user satisfaction.

[1200] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1201] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a summary video or storyboard based on the extracted features, means for encoding the generated summary video or storyboard and transmitting it to a user terminal, means for analyzing user emotion data and optimizing the content of the summary video or storyboard based on the results, and means for preferentially selecting specific scenes or frames based on the user emotion data. This enables automatic generation of content that matches the user's emotional state, thereby improving user satisfaction.

[1202] "Media data" refers to all digital content uploaded by users, including images, videos, and audio.

[1203] The "receiving means" refers to a device or mechanism that acquires media data sent from a user via a network and stores the data in storage.

[1204] "Means for analyzing and extracting features" refers to a device or mechanism that analyzes received media data using AI algorithms or machine learning models to detect significant elements or patterns within the data.

[1205] "Means for generating video summaries or storyboards" refers to a device or mechanism that combines key scenes or elements based on the analysis results to create short video clips or visual materials in the form of storyboards.

[1206] "Means for encoding" refers to a device or mechanism that converts the generated video summary or storyboard into a standard digital file format (e.g., MP4, JPEG) and assigns metadata.

[1207] "Means for transmitting to a user terminal" refers to a device or mechanism for transmitting encoded data over a network to provide a link accessible to the user.

[1208] "Means for analyzing emotional data" refers to a device or mechanism for analyzing emotion-related data provided by a user (e.g., vocal tone, facial expression) to identify the user's emotional state.

[1209] "Means for optimizing the content of a video summary or storyboard" refers to a device or mechanism that adjusts the arrangement of scenes and frames in a video summary or storyboard based on analyzed emotional data, and generates optimal content tailored to the user's emotional state.

[1210] "Means for preferentially selecting specific scenes or frames" refers to a device or mechanism that, based on the results of emotional data analysis, preferentially selects scenes or frames that have received a large number of positive reactions from users.

[1211] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[1212] Server roles and processing flow

[1213] 1. Receiving media data

[1214] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user to upload. This process can use cloud storage such as Amazon S3.

[1215] 2. Data analysis and feature extraction

[1216] The server analyzes the received media data. It uses TensorFlow to perform object detection on image data, OpenCV to perform scene segmentation on video data, and Google Cloud Speech-to-Text API to convert audio data into text. It also performs sentiment analysis on all formats. For example, it uses the YOLO model for image analysis, keyframe extraction for video analysis, and sentiment analysis APIs for audio analysis. It extracts features of each media format from the analyzed data and handles emotional data as well.

[1217] 3. Generate a summary video or storyboard

[1218] The server generates a video summary or storyboard based on the analyzed data. When generating a video summary, it combines the extracted keyframes and important scenes to create a 30-second clip, for example. When generating a storyboard, it organizes the detected images and scenes into a storyboard format. FFmpeg is used to generate the video summary, and PIL (Python Imaging Library) is used to generate the storyboard.

[1219] 4. Data Encoding

[1220] The server encodes the generated summary video and storyboard into standard formats (MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process. FFmpeg commands are used to encode the generated video into MP4 format and also generate a JSON file containing scene timestamps and descriptions.

[1221] 5. Emotional Engine Adjustment

[1222] The server analyzes the emotional data received from the user using an emotion engine and adjusts the process of generating summary videos and storyboards. The emotion engine uses, for example, IBM Watson Emotion Analysis API to analyze the user's emotional data. Based on the results, it prioritizes scenes with positive emotions, for example.

[1223] 6. Sending to the user terminal

[1224] The server generates a link to the encoded data and sends it to the user's device as an HTTP response. The server uses a cloud storage service to generate the link, and obtains a public link for the generated file.

[1225] Terminal roles and processing flow

[1226] 1. Providing UI

[1227] The device provides an intuitive user interface for users to upload media data. It includes an upload button and a preview button. React.js is used to display a dialog for selecting and uploading media files. After the user selects a file, they press the "Upload" button, and the file is sent to the server.

[1228] 2. Sending and Receiving Data

[1229] The device sends the media data uploaded by the user to the server and receives the processing results. It checks the integrity of the data before sending and prompts for re-entry if necessary. It also sends the user's emotional data. It adds the selected file to a FormData object and sends a POST request to the server using Axios. Once processing is complete, it receives a link sent by the server and displays it to the user.

[1230] User roles and operation flow

[1231] 1. Upload your data

[1232] The user selects image, video, and audio data from their own device and uploads it to the system. At the same time, they also input emotion data. They select a file using the file selection dialog in the user interface and press the "Upload" button. They input text and information that reflects emotions (e.g., facial expression recognition data) as emotion data.

[1233] 2. Check and correct the processing results

[1234] The user watches and checks the summary video and storyboard sent from the server. They check the quality and content and make corrections if they are dissatisfied. They click a link in the device's browser to play the summary video or display the storyboard in a viewer. If necessary, they press the "Edit" button to move to the correction screen and make adjustments.

[1235] Example (creating a movie trailer)

[1236] 1. Data upload (user)

[1237] Users upload video files containing movie scenes and emotional data recorded while watching the movie to the system.

[1238] 2. Data analysis (server)

[1239] The server extracts features from video files to detect important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that evoke positive emotional responses.

[1240] 3. Summary video generation (server)

[1241] The server combines the selected scenes to generate a 30-second summary video. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[1242] 4. Receiving and checking results (user)

[1243] The user views the generated trailer and makes corrections if they are not satisfied.

[1244] 5. Corrections and final confirmation (user)

[1245] The user makes any necessary modifications (e.g., adding text or narration) to complete the final trailer.

[1246] Prompt Sentence Examples

[1247] "Generate a 30-second trailer based on the video and emotional data provided by the user. Please prioritize scenes with a high number of positive reactions from users."

[1248] As described above, the present invention utilizes user emotion data to improve the efficiency of creating video summaries and storyboards in video production, improve the consistency between visuals and storylines, and automatically generate optimal content that corresponds to the viewer's emotions.

[1249] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1250] Step 1:

[1251] Receiving media data

[1252] The server receives media data (images, videos, audio) uploaded by the user. The inputs include the file data sent by the user and an HTTP POST request. The server receives the HTTP POST request at its API endpoint (e.g., / upload) and saves the media data to storage (e.g., Amazon S3). The output provides a confirmation message that the save was successful and the path to the file.

[1253] Specific behavior:

[1254] 1. The endpoint receives an HTTP request.

[1255] 2. Analyze the file data and generate the appropriate storage path.

[1256] 3. Save the file to storage (e.g. Amazon S3).

[1257] 4. A confirmation message confirming the save is complete is returned as an HTTP response.

[1258] Step 2:

[1259] Data analysis and feature extraction

[1260] The server analyzes the stored media data. The input includes file data read from storage. Object detection is performed on image data using TensorFlow, scene segmentation is performed on video data using OpenCV, and text conversion is performed on audio data using the Google Cloud Speech-to-Text API. Sentiment analysis is also performed. Output includes extracted feature data, textual audio data, and sentiment analysis results.

[1261] Specific behavior:

[1262] 1. Read media data from storage.

[1263] 2. TensorFlow and YOLO model are used for image analysis to detect objects.

[1264] 3. For video analysis, OpenCV is used to detect scene changes and extract keyframes.

[1265] 4. The voice data is converted to text using the Google Cloud Speech-to-Text API and sent to the sentiment analysis API for sentiment analysis.

[1266] 5. Save the extracted feature data, text, and emotion data and pass them to the next processing step.

[1267] Step 3:

[1268] Generate a summary video or storyboard

[1269] The server generates a video summary or storyboard based on the analyzed feature data. The inputs include the feature data and emotion data from the analysis results. To generate the video summary, important scenes and frames are combined and a clip is created using FFmpeg. To generate the storyboard, PIL (Python Imaging Library) is used to organize images in storyboard format. The output includes the generated video summary and storyboard.

[1270] Specific behavior:

[1271] 1. Load the analysis results and select important scenes and frames.

[1272] 2. Use FFmpeg to combine the selected scenes and generate a 30-second clip.

[1273] 3. For storyboarding, we used PIL to organize the detected images into a storyboard format.

[1274] 4. Temporarily save the generated summary video and storyboard.

[1275] Step 4:

[1276] Data Encoding

[1277] The server encodes the generated summary and storyboard into standard formats (MP4, JPEG). The inputs are the summary and storyboard data. FFmpeg is used for encoding, and metadata (scene descriptions and timestamps) is added. The output contains the encoded files and the metadata.

[1278] Specific behavior:

[1279] 1. Load the saved summary video or storyboard.

[1280] 2. Encode to MP4 format using FFmpeg.

[1281] 3. Descriptions and timestamps for each scene are generated in JSON format and added as metadata.

[1282] 4. Save the encoded file.

[1283] Step 5:

[1284] Emotional engine regulation

[1285] The server adjusts the generated content based on the emotional data provided by the user. The inputs include the user's emotional data and the generated summary video and storyboard. An emotion engine (e.g., IBM Watson Emotion Analysis) analyzes the emotional state and prioritizes scenes with a high number of positive reactions. The output includes the adjusted summary video and storyboard.

[1286] Specific behavior:

[1287] 1. Analyze emotion data using an emotion engine.

[1288] 2. Based on the analysis results, the arrangement of scenes and frames in the summary video and storyboard is readjusted.

[1289] 3. Save the reworked summary footage and storyboard.

[1290] Step 6:

[1291] Send to user terminal

[1292] The server sends the encoded data to the user's device. The input includes the encoded file and its metadata. It saves the file in cloud storage and generates a link to send to the user. The output includes the link sent to the user's device.

[1293] Specific behavior:

[1294] 1. Save the encoded file to cloud storage.

[1295] 2. Generate a public link and send it to the user's device as an HTTP response.

[1296] 3. The user clicks on the link to download or watch the content.

[1297] (Application example 2)

[1298] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1299] In modern content distribution services, there is a demand for methods to provide customized video summaries and storyboards based on individual emotions to improve the user viewing experience. However, conventional systems have difficulty generating content that fully understands and reflects the viewer's emotions, making it difficult to improve viewer satisfaction.

[1300] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1301] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for evaluating the analyzed media data based on the user's emotions using an emotion engine, means for generating a summary video or storyboard based on the extracted features and the emotion evaluation, and means for encoding the generated summary video or storyboard and transmitting it to the user terminal. This enables more personalized content to be automatically generated in accordance with the user's emotion data, improving the viewing experience.

[1302] "Media data" refers to information expressed in a digital format, such as audio, images, or video.

[1303] An "emotion engine" refers to an algorithm or system that analyzes a user's emotions and adjusts or generates media data based on the results.

[1304] "Means of feature extraction" refers to techniques and methods for analyzing and identifying specific information or patterns from media data and extracting that information.

[1305] "Abridged video" refers to a video clip that has been shortened by extracting important scenes or frames from the original long video data.

[1306] A "storyboard" is a visual arrangement of images and text along the frame of a storyboard to organize the flow of a story.

[1307] "Encoding means" refers to the techniques or methods used to convert digital data into a particular form or format.

[1308] "User terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.

[1309] A "specific object" refers to a specific person, object, scene, or other element within the media data to be analyzed.

[1310] "Means for making modifications" refers to interfaces and tools that allow a user to make changes or improvements to the generated video summary or storyboard.

[1311] Server roles and processing flow

[1312] The server has a means to receive media data. For example, it uses the web application framework Flask to set up an API endpoint that receives HTTP requests from users. Media data such as audio, images, and videos are uploaded to this endpoint. The received media data is saved to disk.

[1313] The server then analyzes the received media data and extracts features. This analysis is performed using AI algorithms powered by TensorFlow. For example, image data is subjected to object detection and segmentation, video data is subjected to the detection of important frames and scenes, and audio data is subjected to text conversion using speech recognition technology. Through these analyses, specific objects contained in the data are also detected.

[1314] The server then uses an emotion engine to evaluate the analyzed media data based on the user's emotions. This allows it to extract specific scenes and frames that match the user's emotions, and generates a summary video and storyboard based on these. The generated summary video and storyboard are then converted into a standard file format (e.g., MP4, JPEG) through an encoding process.

[1315] Finally, the server generates a link to send the generated summary video and storyboard to the user's device and sends it as an HTTP response, allowing the user to check the results on their own device.

[1316] Terminal roles and processing flow

[1317] The device provides a user interface (UI) and supports uploading media data, checking the processed results, and editing. The UI is designed to be intuitive, particularly for easy visual confirmation of video and storyboards. The device also has the function of sending the media data uploaded by the user to the server.

[1318] It also has a function to receive the processing result data returned from the server and display it appropriately.It also includes a function to send the user's emotion data (e.g., voice tone, facial recognition data) to the emotion engine.

[1319] User roles and operation flow

[1320] Users upload the media data (audio, images, video) they have collected to the system via their terminal. When uploading, they also enter related information (title, summary, etc.) and emotional data. They then check the quality and content of the summary video and storyboard returned from the server. If necessary, they make corrections to the received data and re-upload it to the system.

[1321] Specific examples

[1322] For example, if a user wants to make a movie trailer more emotional, the process would be as follows:

[1323] 1. Data upload (user)

[1324] A user uploads a video file of a movie.

[1325] Input emotional data (e.g., emotional ratings during viewing).

[1326] 2. Data analysis and video summary generation (server)

[1327] The server analyzes the media data, extracts specific scenes, and generates a summary video based on the emotion data.

[1328] 3. Receive and check the results (terminal)

[1329] A download link for the generated summary video is generated and transmitted to the user terminal.

[1330] The user reviews the generated movie trailer.

[1331] Prompt Sentence Examples

[1332] "Scene analysis of movie XYZ. Provide media data for scenes that evoke positive emotions. Create a custom trailer."

[1333] As a result, the present invention enables more personalized content to be automatically generated in accordance with the user's emotional data, thereby improving the viewing experience.

[1334] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1335] Step 1:

[1336] Data upload (user)

[1337] Users upload media data (audio, images, videos) to the system using their own devices. Specifically, they select the data through the device interface and enter additional related information (title, summary) and emotional data (emotion rating, etc.). The media data and its related information are then sent to the server.

[1338] Input: User-provided audio, image, and video data, as well as related information and emotional data

[1339] Output: Media data and related information sent to the server

[1340] Step 2:

[1341] Receiving and storing media data (server)

[1342] The server receives media data uploaded by users and saves it to disk via HTTP requests. An API endpoint is set up using a web framework such as Flask, where the media data is saved.

[1343] Input: User uploaded media data

[1344] Output: Media data stored on the server's disk

[1345] Step 3:

[1346] Data analysis and feature extraction (server)

[1347] The server analyzes the received media data and extracts features. Using TensorFlow, the AI ​​model performs image and voice recognition, object detection, scene extraction, etc. It also includes a means to detect specific objects in each media data.

[1348] Input: Stored media data

[1349] Output: Extracted feature data (key scenes, audio text, detected objects)

[1350] Step 4:

[1351] Evaluation by emotion engine (server)

[1352] The emotion engine evaluates the user's emotions based on the analyzed media data. For example, an emotion recognition AI model analyzes the user's emotional data regarding a specific scene or object and outputs an evaluation score.

[1353] Input: Extracted feature data and user emotion data

[1354] Output: Score data based on the user's emotional evaluation

[1355] Step 5:

[1356] Summary video or storyboard generation (server)

[1357] The server generates a summary video or storyboard based on the feature data and emotion evaluation, stitching together important scenes and frames and composing optimal content taking into account the emotion score.

[1358] Input: feature data and emotion evaluation scores

[1359] Output: Generated summary video or storyboard

[1360] Step 6:

[1361] Encoding and sending (server)

[1362] The generated summary video or storyboard is encoded and converted into a standard file format (e.g., MP4, JPEG). A download link is generated for the encoded file to be sent to the user's device, and the link is sent as an HTTP response.

[1363] Input: Generated summary video or storyboard

[1364] Output: Encoded file and download link

[1365] Step 7:

[1366] Check and correct the results (user and device)

[1367] The user checks the summary video and storyboard generated on their device. If the video does not meet their expectations, they can make corrections and upload the corrected data back to the server. The device supports this correction process.

[1368] Input: Encoded file and download link

[1369] Output: Confirmed summary video and storyboard, and correction data as needed

[1370] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1371] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1372] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1373] [Fourth embodiment]

[1374] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1375] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1376] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1377] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1378] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1379] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1380] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1381] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1382] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1383] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1384] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1385] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1386] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1387] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a summary video or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[1388] Server roles and processing flow

[1389] 1. Receiving media data

[1390] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[1391] 2. Data analysis and feature extraction

[1392] The server analyzes the received media data using AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted to text using speech recognition technology, and sentiment analysis is performed.

[1393] 3. Generate a summary video or storyboard

[1394] The server generates a video summary or storyboard based on the extracted features. In the case of a video summary, it combines frames from key scenes to create short video clips, and in the case of a storyboard, it organizes images according to the storyboard frames.

[1395] 4. Data Encoding

[1396] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[1397] 5. Send to user device

[1398] The server sends the encoded data to the user terminal using an appropriate protocol (e.g., SSL / TLS) to ensure security and efficiency.

[1399] Terminal roles and processing flow

[1400] 1. Providing UI

[1401] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[1402] 2. Sending and Receiving Data

[1403] The device sends the media data uploaded by the user to the server, checks the integrity of the data when it is sent, and prompts the user to correct it if necessary, and receives the processed data returned from the server and displays it appropriately.

[1404] User roles and operation flow

[1405] 1. Upload your data

[1406] Users upload the media data (images, videos, audio) they have collected to the system via their terminals. When uploading, they also enter related information (title, summary, etc.).

[1407] 2. Check and correct the processing results

[1408] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[1409] Specific examples

[1410] A concrete example of creating a movie trailer

[1411] 1. Data upload (user)

[1412] Users upload multiple video files containing scenes from their movies to the system.

[1413] 2. Data analysis (server)

[1414] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[1415] 3. Summary video generation (server)

[1416] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[1417] 4. Receiving and checking results (user)

[1418] The user can check the generated trailer on their device and, if they are not satisfied with the footage, make any necessary corrections.

[1419] 5. Corrections and final confirmation (user)

[1420] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[1421] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[1422] The processing flow will be explained below.

[1423] Step 1:

[1424] A user uploads media data (images, videos, audio) from a device to the system. Specifically, the user opens a file selection dialog using the device's UI and specifies the file to upload. By pressing the upload button, the selected file is sent to the server.

[1425] Step 2:

[1426] The server stores the media data received from the user in storage. During this process, the server analyzes the file metadata (file name, size, type) and records it in a log.

[1427] Step 3:

[1428] The server performs an initial analysis of the received media data. Specifically, it checks the file format and scans the content, and then selects the appropriate analysis module. For image data, it uses the object detection module, and for video data, it uses the scene segmentation module.

[1429] Step 4:

[1430] The server uses the selected analysis module to extract features from the media data, for example, detecting important frames and scenes in video data, identifying key objects and scenes in image data, and converting speech to text and detecting specific keywords and emotions in audio data.

[1431] Step 5:

[1432] The server generates a summary video or storyboard based on the extracted feature information. In the case of a summary video, the extracted important scenes are arranged along a timeline and edited to a specified length. In the case of a storyboard, important scenes are arranged in each frame of the storyboard and explanations are added.

[1433] Step 6:

[1434] The server encodes the generated summary or storyboard into a standard file format, for example MP4 for the summary and JPEG or PDF for the storyboard, adding metadata (scene descriptions, timestamps) during the process.

[1435] Step 7:

[1436] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[1437] Step 8:

[1438] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[1439] Step 9:

[1440] The user checks the downloaded summary video or storyboard, and if there are any problems or omissions, they can correct them using the editing tool. Once the corrections are complete, they can upload it back to the system.

[1441] Step 10:

[1442] The server receives the modified data, re-analyzes it, and encodes it, incorporating any appropriate modifications to generate the final summary or storyboard. The resulting data is then sent back to the user's device, completing the process.

[1443] Example 1

[1444] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1445] Conventional media data analysis and editing work is often performed manually, resulting in problems of time and effort. Furthermore, the accuracy and consistency of analysis results often depend on specialized knowledge and are therefore poor. To solve these problems, the present invention aims to automate the reception, analysis, summary video or storyboard generation, encoding, and transmission of media data, enabling efficient and accurate processing.

[1446] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1447] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a video summary or storyboard based on the extracted features, means for encoding the generated video summary or storyboard into a standard file format, and means for transmitting the encoded video summary or storyboard to a user terminal. This makes it possible to automate the analysis and editing of media data and generate video summary and storyboards with high accuracy and efficiency.

[1448] "Media data" is a general term for data including image data, video data, and audio data.

[1449] The term "receiving means" refers to a function or device that can transmit media data from a user to a server and receive it on the server side.

[1450] "Analysis and feature extraction means" refers to systems or algorithms that analyze received media data and identify significant elements or patterns.

[1451] "Video summary" refers to a short video clip created by combining important scenes extracted from the analyzed media data.

[1452] A "storyboard" refers to a group of images that visually represent the storyboard or development of a scene in a video.

[1453] "Standard file formats" are commonly used file formats, such as MP4 for videos and JPEG for images.

[1454] "Means for encoding" refers to the process or tools used to convert the generated summary or storyboard into an appropriate file format.

[1455] "Means for transmitting to a user terminal" refers to a function or system that transmits encoded data from a server to a user's device.

[1456] "Object detection means" refers to techniques or algorithms used to identify specific objects or people within media data.

[1457] "Means for re-uploading to the system" refers to a function or process that allows a user to send a revised summary video or storyboard to the server again.

[1458] The present invention relates to a system in which a user and a server work together to receive media data, analyze it, generate a video summary or storyboard, encode it, and transmit it to a user terminal. Specific embodiments of the present invention will be described below.

[1459] Server roles and processing flow

[1460] Receiving media data

[1461] The server receives media data (images, videos, audio) uploaded by users through HTTP requests. The server-side program sets up an API endpoint using a web framework such as Python's Flask or Django and waits for the user's upload operation. The received data is stored in a temporary storage location (for example, a cloud storage service).

[1462] Data analysis and feature extraction

[1463] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. For image data, object detection and segmentation are performed, and for video data, important scenes and frames are detected. For audio data, the Google Cloud Speech-to-Text service is used to convert it to text and perform sentiment analysis.

[1464] Generate a summary video or storyboard

[1465] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips. For a storyboard, it organizes images based on the storyboard. This process can be done using, for example, FFmpeg or the Python Imaging Library (PIL).

[1466] Data Encoding

[1467] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files into MP4 or JPEG format, and adds the necessary metadata (scene descriptions and timestamps).

[1468] Send to user terminal

[1469] The server sends the encoded data to the user's terminal, using the SSL / TLS protocol to ensure security, and returns the data to the terminal via an HTTP request.

[1470] Terminal roles and processing flow

[1471] Providing a UI

[1472] The device provides a user interface (UI) for uploading media data, checking the processed results, and editing. The UI is built using front-end technologies such as HTML, CSS, and JavaScript, and is designed to be intuitive for users to operate.

[1473] Sending and Receiving Data

[1474] The device sends the media data selected by the user to the server, checks the integrity of the data, and receives the processed data from the server and displays it on the UI.

[1475] User roles and operation flow

[1476] Uploading data

[1477] Users upload their collected media data to the system using the device's UI, and also enter related information such as title and summary when uploading.

[1478] Checking and correcting the processing results

[1479] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system.

[1480] Specific examples

[1481] A concrete example of creating a movie trailer

[1482] 1. Data upload (user)

[1483] Users upload multiple video files containing scenes from their movies to the system.

[1484] 2. Data analysis (server)

[1485] The server analyzes the uploaded video files, extracts the characteristics of each scene, and detects important scenes and impactful frames.

[1486] 3. Summary video generation (server)

[1487] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion.

[1488] 4. Receiving and checking results (user)

[1489] The user can check the generated trailer on the terminal and make any necessary corrections if they are not satisfied.

[1490] 5. Corrections and final confirmation (user)

[1491] The user can add text and narration to the trailer and make adjustments to create the final trailer.

[1492] In this way, the present invention can improve the efficiency of creating video summaries and storyboards in video production, and improve the consistency between the visuals and the story.

[1493] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1494] Step 1:

[1495] Receiving media data (server)

[1496] The server receives media data such as images, videos, and audio from the user via HTTP requests. Specifically, an API endpoint is set up using a web framework such as Flask or Django, and awaits the user's upload operation. The input is the media data uploaded by the user, which is then temporarily stored on the server (for example, in a cloud storage service).

[1497] Step 2:

[1498] Data analysis and feature extraction (server)

[1499] The server analyzes the received media data using AI algorithms such as TensorFlow and OpenCV. The input is the stored media data, and for image data, object detection and segmentation are performed. For video data, important scenes and frames are detected. For audio data, text conversion and sentiment analysis are performed using the Google Cloud Speech-to-Text service. The output is extracted feature information.

[1500] Step 3:

[1501] Summary video or storyboard generation (server)

[1502] The server generates a video summary or storyboard based on the extracted feature information. For a video summary, it combines frames of important scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard. This process uses FFmpeg and the Python Imaging Library (PIL). The input is the feature information, and the output is the generated video summary or storyboard.

[1503] Step 4:

[1504] Data Encoding (Server)

[1505] The server encodes the generated summary or storyboard into a standard file format, for example using the FFmpeg library to convert the media files to MP4 or JPEG format and adding metadata (scene descriptions and timestamps) if necessary. The input is the generated summary or storyboard, and the output is the encoded file.

[1506] Step 5:

[1507] Send to user terminal (server)

[1508] The server sends the encoded data to the user's device, using the SSL / TLS protocol to ensure security. The input is the encoded file, and the output is the data received on the user's device.

[1509] Step 6:

[1510] UI provision (terminal)

[1511] The terminal provides the user with a user interface (UI) for uploading media data, checking the processing results, and editing them. The UI is designed to be intuitive using HTML, CSS, JavaScript, etc. Specifically, it provides a file selection button and a viewer that displays the processing results.

[1512] Step 7:

[1513] Sending and receiving data (terminal)

[1514] The device sends the media data selected by the user to the server, verifies the integrity of the data, and receives the processing results returned from the server and displays them on the UI. The input is the media data selected by the user and the processing result data from the server, and the output is the processing result displayed on the user device.

[1515] Step 8:

[1516] Data Upload (User)

[1517] Users upload collected media data to the system using the device's UI. The input is the media data and related information (title, summary, etc.), which is sent to the server and stored. The output is the media data stored in a temporary storage location on the server.

[1518] Step 9:

[1519] Check and correct the processing results (user)

[1520] The user checks the summary video and storyboard sent back from the server and makes any necessary corrections. The corrected data is then uploaded back to the system. The input is the summary video and storyboard sent from the server and the corrected data, and the output is the corrected summary video and storyboard.

[1521] Through these steps, the system is able to efficiently and accurately carry out the entire process of analyzing media data, generating summary videos and storyboards, and encoding and transmitting.

[1522] (Application example 1)

[1523] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1524] In conventional media data processing systems, users have to extract important scenes from long videos and generate video summaries, which requires a great deal of time and effort. Furthermore, there is a lack of an efficient way to share the generated summaries. Furthermore, there are issues with the security of the analyzed data and the ease with which users can edit them.

[1525] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1526] In this invention, the server includes a means for receiving media data, a means for analyzing the received media data to extract features, and a means for generating a video summary or storyboard based on the extracted features. This enables automatic extraction of important scenes and generation of a video summary quickly and efficiently. The server also includes a means for encoding the generated video summary or storyboard and transmitting it to a user terminal, a means for using a protocol for securely transmitting the encoded data to the user terminal, a means for providing a user interface and supporting uploading of media data, confirmation of processing results, and editing, and a means for displaying the generated video summary on a smart device and sharing it on a social networking service. This allows users to easily review, modify, and securely share the generated video summary.

[1527] "Media data" refers to information stored in digital form, such as video, audio, and images.

[1528] "Analysis" refers to a series of processes that extract features from digital data and convert it into an understandable form.

[1529] A "feature" is a unique or important part of media data, including a particular part of a scene, an object, or an audio.

[1530] "Abridged video" refers to a video clip that has been shortened by extracting important scenes from a longer video.

[1531] A storyboard is a series of sketches or images that visually illustrate a plan for a film production or presentation.

[1532] "Encoding" refers to the process of converting digital data into a particular format so that it can be stored or transmitted.

[1533] A "user terminal" is an electronic device that can be directly operated by a user, and includes smartphones, tablets, computers, etc.

[1534] A "protocol" refers to the rules and procedures that define data communication between different electronic devices.

[1535] "User interface" refers to the means, screen display, and input operations that allow a user to interact with a system.

[1536] "SNS" is an abbreviation for social networking service, and refers to a platform where people can share information and interact with each other online.

[1537] "Smart devices" is a general term for modern electronic devices with internet connectivity, including smartphones, tablets, and smartwatches.

[1538] A specific embodiment of the present invention will be described. The present invention relates to a system that extracts important scenes from a long video shot by a user and automatically generates a video summary. Below, each component of the system and its processing procedure will be described.

[1539] Overall system configuration

[1540] This system consists of multiple elements, including a user terminal, a server, and a user interface. The user terminal is a portable device such as a smartphone or tablet. The server handles the main processing such as analyzing media data and generating video summaries, and a cloud server with a stable connection is suitable.

[1541] Hardware and Software

[1542] Hardware: smartphones, servers, cloud services

[1543] Software: Flask (web framework), OpenCV (image processing library), moviepy (video editing library), HTTP request library

[1544] Data analysis and video summary generation

[1545] 1. Receiving media data: The application on the user's device uploads the video the user has taken to the server using an HTTP request. In this case, Flask is used to set up an API endpoint and wait for the data to be received.

[1546] 2. Data Analysis:

[1547] The server analyzes the received video data. It uses OpenCV to sequentially read the video frames and extract important scenes from each specific frame. This analysis uses AI algorithms and machine learning models to identify scene changes and important events.

[1548] 3. Summary video generation:

[1549] Based on the analysis results, the server uses moviepy to generate video clips from important frames and stitches them together to create a summary video.

[1550] 4. Encoding and Transmission:

[1551] The server encodes the generated summary video into a common format such as MP4 and transmits it securely to the user terminal using an appropriate protocol such as SSL / TLS.

[1552] User Interface and Sharing

[1553] The user device provides a user interface for displaying the generated video summary. This interface is intuitive and designed to allow users to easily check and edit the video summary. It also includes buttons for sharing the generated video summary on social media.

[1554] Specific examples

[1555] In a specific usage scenario, when a user uploads a video (60 minutes) of a family trip, the system automatically generates a 5-minute summary video including smiling scenes and landmarks. The user can review this summary video and, if they like it, share it on social media. For example, the prompt text might look like this:

[1556] "Generate a 5-minute summary video from a 60-minute family trip, including funny moments and landmarks."

[1557] This prompt allows the generative AI model to suggest and quickly create the summary video the user desires.

[1558] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1559] Step 1:

[1560] Receiving media data:

[1561] The user device uploads the video file they have taken to the server using an HTTP request. At this time, the user selects the video file through the application and starts sending it. The server receives the request at an API endpoint set up using Flask and saves the video file on the server. The input is the video file uploaded by the user, and the output is the video file saved on the server.

[1562] Step 2:

[1563] Data Analysis:

[1564] The server analyzes the stored video files. It uses OpenCV to sequentially read the video frames and calculate the differences between each frame to extract important scenes. Specifically, it identifies important frames using algorithms such as inter-frame motion, color change, and object detection. The input is the video file stored on the server, and the output is a list of important frames.

[1565] Step 3:

[1566] Feature extraction:

[1567] The server uses AI algorithms and machine learning models to extract scene and event features from key frames, such as face recognition and object detection, and tag each frame. The input is a list of key frames, and the output is a set of frames with extracted features.

[1568] Step 4:

[1569] Summary video generation:

[1570] The server uses MoviePy to generate sub-clips for each frame from the feature-extracted frames and stitch them together to create a summary video. Specifically, it extracts short clips from the original video based on the timestamps of important frames and concatenates them into a single video. The input is a set of feature-extracted frames, and the output is a summary video.

[1571] Step 5:

[1572] Encoding:

[1573] The server encodes the generated summary video into a common video format (e.g., MP4 format). This process involves compressing the video and adding metadata. The input is the summary video, and the output is the encoded summary video file.

[1574] Step 6:

[1575] Data transmission:

[1576] The server sends the encoded summary video to the user's device using a secure protocol such as SSL / TLS to ensure data integrity and security. The input is the encoded summary video file, and the output is the video file sent to the user's device.

[1577] Step 7:

[1578] UI display and confirmation:

[1579] The user terminal displays the received summary video through a user interface. The user can visually check the generated summary video and make corrections on the screen as necessary. The input is the summary video file sent to the user terminal, and the output is the summary video displayed on the user interface.

[1580] Step 8:

[1581] Share on social media:

[1582] The user device provides a function that allows users to share the generated summary video on social media using the share button on the screen. This allows users to easily share the summary video with friends and followers. The input is the summary video displayed on the user interface, and the output is the summary video shared on the social media.

[1583] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1584] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[1585] Server roles and processing flow

[1586] 1. Receiving media data

[1587] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user's upload operation.

[1588] 2. Data analysis and feature extraction

[1589] The server analyzes the received media data. This analysis uses AI algorithms tailored to each media format. For image data, object detection and segmentation are performed, and for video data, important frames and scenes are detected. Audio data is converted into text using speech recognition technology, and sentiment analysis is performed. It also includes a method for detecting specific objects in the received media data.

[1590] 3. Generate a summary video or storyboard

[1591] The server generates a video summary or storyboard based on the extracted features. For a video summary, it combines frames from key scenes to create short video clips, and for a storyboard, it organizes images according to the storyboard frames.

[1592] 4. Data Encoding

[1593] The server encodes the generated video summary or storyboard and converts it into a standard file format (e.g. MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process.

[1594] 5. Emotional Engine Adjustment

[1595] The server receives the emotional data provided by the user and analyzes it using an emotion engine. Based on the analysis results, it adjusts the process of generating a summary video or storyboard. Specifically, it prioritizes the selection of specific scenes or frames that correspond to the user's emotional state, optimizing the content of the video or storyboard.

[1596] 6. Sending to the user terminal

[1597] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[1598] Terminal roles and processing flow

[1599] 1. Providing UI

[1600] The terminal provides a user interface (UI) that supports uploading media data, checking the processing results, and editing. The UI has an intuitive design, and is especially designed to make it easy to visually check the video and storyboard.

[1601] 2. Sending and Receiving Data

[1602] The device sends the media data uploaded by the user to the server. It checks the integrity of the data when sending it and prompts the user to make corrections if necessary. It receives the processing result data returned from the server and displays it appropriately. It also sends the user's emotional data (e.g., voice tone, facial recognition data) to the emotion engine.

[1603] User roles and operation flow

[1604] 1. Upload your data

[1605] Users upload the media data (images, videos, audio) they have collected to the system via their devices. When uploading, they also enter related information (title, summary, etc.) and emotional data.

[1606] 2. Check and correct the processing results

[1607] The user checks the summary video and storyboard returned from the server. When checking, they play it back on their device and check the quality and content. If necessary, they make corrections to the received data and upload it back to the system.

[1608] Specific examples

[1609] A concrete example of creating a movie trailer

[1610] 1. Data upload (user)

[1611] Users upload multiple video files containing scenes from their movies to the system, and also provide data for emotion recognition (e.g., emotions and voice data while watching).

[1612] 2. Data analysis (server)

[1613] The server analyzes the uploaded video files and extracts the characteristics of each scene. It detects important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that have generated a large number of positive reactions.

[1614] 3. Summary video generation (server)

[1615] The server combines the detected scenes to generate a 30-second summary video (trailer) that is highly effective for promotion. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[1616] 4. Receiving and checking results (user)

[1617] The user can then review the generated trailer on their device. If they are not satisfied with the video, they can make any necessary corrections. Further adjustments can also be made based on the emotional data.

[1618] 5. Corrections and final confirmation (user)

[1619] The user can then make adjustments to the trailer to add text and narration, resulting in a final trailer.

[1620] In this way, the present invention utilizes user emotion data to not only improve the efficiency of creating video summaries and storyboards in video production and increase the consistency between the visuals and the story, but also automatically generate optimal content that corresponds to the viewer's emotions.

[1621] The processing flow will be explained below.

[1622] Step 1:

[1623] Users upload media data (images, videos, audio) from their devices to the system. Specifically, they open a file selection dialog using the device's UI and specify the files to upload. In addition, emotion recognition data (e.g., facial expression data, voice tone data) is also uploaded at the same time.

[1624] Step 2:

[1625] The device sends the media data and emotion recognition data specified by the user to the server, which checks the integrity of the files and converts or compresses the data as necessary.

[1626] Step 3:

[1627] The server stores the media data received from the user in storage. The server-side program analyzes the file format and metadata (file name, size, type) and records them in a log.

[1628] Step 4:

[1629] The server performs an initial analysis of the received media data, applying object detection algorithms to image data, scene detection algorithms to video data, and speech recognition algorithms to audio data.

[1630] Step 5:

[1631] The server uses an emotion engine to analyze the received emotion recognition data, which involves identifying patterns in the user's facial expressions and vocal tone and identifying corresponding emotions (e.g., joy, surprise, sadness).

[1632] Step 6:

[1633] The server extracts important frames and scenes based on the analysis results, for example, flagging a scene in which the user expresses joy as an important scene.

[1634] Step 7:

[1635] The server generates a video summary or storyboard based on the extracted features and emotion data. For video summaries, key scenes are combined to create short video clips, and emotion data is incorporated to create a video that is appealing to users. For storyboards, selected scenes are arranged in frame order and explanatory text is added.

[1636] Step 8:

[1637] The server encodes the generated summary video or storyboard into a standard file format (e.g., MP4, JPEG), adding metadata (scene summary, emotion tags) during the process.

[1638] Step 9:

[1639] The server generates a link to transfer the encoded data to the user's device, which is sent as an HTTP response to the user's device.

[1640] Step 10:

[1641] The terminal displays the received link on a user interface, and the user can click the link to download or view the generated summary video or storyboard.

[1642] Step 11:

[1643] The user checks the downloaded summary video or storyboard and, if necessary, makes corrections using the editing tools on the device, making final adjustments that reflect the emotion data.

[1644] Step 12:

[1645] The user then uploads the revised data back into the system, where the server reanalyzes and encodes it to generate the final summary video or storyboard, which is then sent back to the user's device to complete the process.

[1646] Through the above steps, the present invention utilizes user emotion data to improve the efficiency and quality of video production and automatically generate optimal content that appeals to the viewer's emotions.

[1647] Example 2

[1648] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1649] Conventional media data analysis systems have not been able to fully utilize the emotional data provided by users, resulting in summary videos and storyboards that often do not match the user's emotional state. Furthermore, prioritizing specific scenes or frames, as well as the need to modify and re-upload content, are cumbersome. As a result, there are issues with the quality of the generated content and user satisfaction.

[1650] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1651] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for generating a summary video or storyboard based on the extracted features, means for encoding the generated summary video or storyboard and transmitting it to a user terminal, means for analyzing user emotion data and optimizing the content of the summary video or storyboard based on the results, and means for preferentially selecting specific scenes or frames based on the user emotion data. This enables automatic generation of content that matches the user's emotional state, thereby improving user satisfaction.

[1652] "Media data" refers to all digital content uploaded by users, including images, videos, and audio.

[1653] The "receiving means" refers to a device or mechanism that acquires media data sent from a user via a network and stores the data in storage.

[1654] "Means for analyzing and extracting features" refers to a device or mechanism that analyzes received media data using AI algorithms or machine learning models to detect significant elements or patterns within the data.

[1655] "Means for generating video summaries or storyboards" refers to a device or mechanism that combines key scenes or elements based on the analysis results to create short video clips or visual materials in the form of storyboards.

[1656] "Means for encoding" refers to a device or mechanism that converts the generated video summary or storyboard into a standard digital file format (e.g., MP4, JPEG) and assigns metadata.

[1657] "Means for transmitting to a user terminal" refers to a device or mechanism for transmitting encoded data over a network to provide a link accessible to the user.

[1658] "Means for analyzing emotional data" refers to a device or mechanism for analyzing emotion-related data provided by a user (e.g., vocal tone, facial expression) to identify the user's emotional state.

[1659] "Means for optimizing the content of a video summary or storyboard" refers to a device or mechanism that adjusts the arrangement of scenes and frames in a video summary or storyboard based on analyzed emotional data, and generates optimal content tailored to the user's emotional state.

[1660] "Means for preferentially selecting specific scenes or frames" refers to a device or mechanism that, based on the results of emotional data analysis, preferentially selects scenes or frames that have received a large number of positive reactions from users.

[1661] The present invention incorporates an emotion engine into a system in which a user and a server work together to receive and analyze media data, generate a summary video or storyboard, encode it, transmit it to a user terminal, and make adjustments based on the user's emotion recognition. Specific embodiments of the present invention will be described below.

[1662] Server roles and processing flow

[1663] 1. Receiving media data

[1664] The server receives media data (images, videos, audio) uploaded by the user. The server-side program sets up an API endpoint to receive HTTP requests and waits for the user to upload. This process can use cloud storage such as Amazon S3.

[1665] 2. Data analysis and feature extraction

[1666] The server analyzes the received media data. It uses TensorFlow to perform object detection on image data, OpenCV to perform scene segmentation on video data, and Google Cloud Speech-to-Text API to convert audio data into text. It also performs sentiment analysis on all formats. For example, it uses the YOLO model for image analysis, keyframe extraction for video analysis, and sentiment analysis APIs for audio analysis. It extracts features of each media format from the analyzed data and handles emotional data as well.

[1667] 3. Generate a summary video or storyboard

[1668] The server generates a video summary or storyboard based on the analyzed data. When generating a video summary, it combines the extracted keyframes and important scenes to create a 30-second clip, for example. When generating a storyboard, it organizes the detected images and scenes into a storyboard format. FFmpeg is used to generate the video summary, and PIL (Python Imaging Library) is used to generate the storyboard.

[1669] 4. Data Encoding

[1670] The server encodes the generated summary video and storyboard into standard formats (MP4, JPEG), adding metadata (scene descriptions and timestamps) during the process. FFmpeg commands are used to encode the generated video into MP4 format and also generate a JSON file containing scene timestamps and descriptions.

[1671] 5. Emotional Engine Adjustment

[1672] The server analyzes the emotional data received from the user using an emotion engine and adjusts the process of generating summary videos and storyboards. The emotion engine uses, for example, IBM Watson Emotion Analysis API to analyze the user's emotional data. Based on the results, it prioritizes scenes with positive emotions, for example.

[1673] 6. Send to user terminal

[1674] The server generates a link to the encoded data and sends it to the user's device as an HTTP response. The server uses a cloud storage service to generate the link, and obtains a public link for the generated file.

[1675] Terminal roles and processing flow

[1676] 1. Providing UI

[1677] The device provides an intuitive user interface for users to upload media data. It includes an upload button and a preview button. React.js is used to display a dialog for selecting and uploading media files. After the user selects a file, they press the "Upload" button, and the file is sent to the server.

[1678] 2. Sending and Receiving Data

[1679] The device sends the media data uploaded by the user to the server and receives the processing results. It checks the integrity of the data before sending and prompts for re-entry if necessary. It also sends the user's emotional data. It adds the selected file to a FormData object and sends a POST request to the server using Axios. Once processing is complete, it receives a link sent by the server and displays it to the user.

[1680] User roles and operation flow

[1681] 1. Upload your data

[1682] The user selects image, video, and audio data from their own device and uploads it to the system. At the same time, they also input emotion data. They select a file using the file selection dialog in the user interface and press the "Upload" button. They input text and information that reflects emotions (e.g., facial expression recognition data) as emotion data.

[1683] 2. Check and correct the processing results

[1684] The user watches and checks the summary video and storyboard sent from the server. They check the quality and content and make corrections if they are dissatisfied. They click a link in the device's browser to play the summary video or display the storyboard in a viewer. If necessary, they press the "Edit" button to move to the correction screen and make adjustments.

[1685] Example (creating a movie trailer)

[1686] 1. Data upload (user)

[1687] Users upload video files containing movie scenes and emotional data recorded while watching the movie to the system.

[1688] 2. Data analysis (server)

[1689] The server extracts features from video files to detect important scenes and impactful frames. The emotion engine analyzes the user's emotional data and prioritizes scenes that evoke positive emotional responses.

[1690] 3. Summary video generation (server)

[1691] The server combines the selected scenes to generate a 30-second summary video. It considers the emotional data and arranges scenes that correspond to the user's emotions.

[1692] 4. Receiving and checking results (user)

[1693] The user views the generated trailer and makes corrections if they are not satisfied.

[1694] 5. Corrections and final confirmation (user)

[1695] The user makes any necessary modifications (e.g., adding text or narration) to complete the final trailer.

[1696] Prompt Sentence Examples

[1697] "Generate a 30-second trailer based on the video and emotional data provided by the user. Please prioritize scenes with a high number of positive reactions from users."

[1698] As described above, the present invention utilizes user emotion data to improve the efficiency of creating video summaries and storyboards in video production, improve the consistency between visuals and storylines, and automatically generate optimal content that corresponds to the viewer's emotions.

[1699] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1700] Step 1:

[1701] Receiving media data

[1702] The server receives media data (images, videos, audio) uploaded by the user. The inputs include the file data sent by the user and an HTTP POST request. The server receives the HTTP POST request at its API endpoint (e.g., / upload) and saves the media data to storage (e.g., Amazon S3). The output provides a confirmation message that the save was successful and the path to the file.

[1703] Specific behavior:

[1704] 1. The endpoint receives an HTTP request.

[1705] 2. Analyze the file data and generate the appropriate storage path.

[1706] 3. Save the file to storage (e.g. Amazon S3).

[1707] 4. A confirmation message confirming the save is complete is returned as an HTTP response.

[1708] Step 2:

[1709] Data analysis and feature extraction

[1710] The server analyzes the stored media data. The input includes file data read from storage. Object detection is performed on image data using TensorFlow, scene segmentation is performed on video data using OpenCV, and text conversion is performed on audio data using the Google Cloud Speech-to-Text API. Sentiment analysis is also performed. Output includes extracted feature data, textual audio data, and sentiment analysis results.

[1711] Specific behavior:

[1712] 1. Read media data from storage.

[1713] 2. TensorFlow and YOLO model are used for image analysis to detect objects.

[1714] 3. For video analysis, OpenCV is used to detect scene changes and extract keyframes.

[1715] 4. The voice data is converted to text using the Google Cloud Speech-to-Text API and sent to the sentiment analysis API for sentiment analysis.

[1716] 5. Save the extracted feature data, text, and emotion data and pass them to the next processing step.

[1717] Step 3:

[1718] Generate a summary video or storyboard

[1719] The server generates a video summary or storyboard based on the analyzed feature data. The inputs include the feature data and emotion data from the analysis results. To generate the video summary, important scenes and frames are combined and a clip is created using FFmpeg. To generate the storyboard, PIL (Python Imaging Library) is used to organize images in storyboard format. The output includes the generated video summary and storyboard.

[1720] Specific behavior:

[1721] 1. Load the analysis results and select important scenes and frames.

[1722] 2. Use FFmpeg to combine the selected scenes and generate a 30-second clip.

[1723] 3. For storyboarding, we used PIL to organize the detected images into a storyboard format.

[1724] 4. Temporarily save the generated summary video and storyboard.

[1725] Step 4:

[1726] Data Encoding

[1727] The server encodes the generated summary and storyboard into standard formats (MP4, JPEG). The inputs are the summary and storyboard data. FFmpeg is used for encoding, and metadata (scene descriptions and timestamps) is added. The output contains the encoded files and the metadata.

[1728] Specific behavior:

[1729] 1. Load the saved summary video or storyboard.

[1730] 2. Encode to MP4 format using FFmpeg.

[1731] 3. Descriptions and timestamps for each scene are generated in JSON format and added as metadata.

[1732] 4. Save the encoded file.

[1733] Step 5:

[1734] Emotional engine regulation

[1735] The server adjusts the generated content based on the emotional data provided by the user. The inputs include the user's emotional data and the generated summary video and storyboard. An emotion engine (e.g., IBM Watson Emotion Analysis) analyzes the emotional state and prioritizes scenes with a high number of positive reactions. The output includes the adjusted summary video and storyboard.

[1736] Specific behavior:

[1737] 1. Analyze emotion data using an emotion engine.

[1738] 2. Based on the analysis results, the arrangement of scenes and frames in the summary video and storyboard is readjusted.

[1739] 3. Save the reworked summary footage and storyboard.

[1740] Step 6:

[1741] Send to user terminal

[1742] The server sends the encoded data to the user's device. The input includes the encoded file and its metadata. It saves the file in cloud storage and generates a link to send to the user. The output includes the link sent to the user's device.

[1743] Specific behavior:

[1744] 1. Save the encoded file to cloud storage.

[1745] 2. Generate a public link and send it to the user's device as an HTTP response.

[1746] 3. The user clicks on the link to download or watch the content.

[1747] (Application example 2)

[1748] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1749] In modern content distribution services, there is a demand for methods to provide customized video summaries and storyboards based on individual emotions to improve the user viewing experience. However, conventional systems have difficulty generating content that fully understands and reflects the viewer's emotions, making it difficult to improve viewer satisfaction.

[1750] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1751] In this invention, the server includes means for receiving media data, means for analyzing the received media data and extracting features, means for evaluating the analyzed media data based on the user's emotions using an emotion engine, means for generating a summary video or storyboard based on the extracted features and the emotion evaluation, and means for encoding the generated summary video or storyboard and transmitting it to the user terminal. This enables more personalized content to be automatically generated in accordance with the user's emotion data, improving the viewing experience.

[1752] "Media data" refers to information expressed in a digital format, such as audio, images, or video.

[1753] An "emotion engine" refers to an algorithm or system that analyzes a user's emotions and adjusts or generates media data based on the results.

[1754] "Means of feature extraction" refers to techniques and methods for analyzing and identifying specific information or patterns from media data and extracting that information.

[1755] "Abridged video" refers to a video clip that has been shortened by extracting important scenes or frames from the original long video data.

[1756] A "storyboard" is a visual arrangement of images and text along the frame of a storyboard to organize the flow of a story.

[1757] "Encoding means" refers to the techniques or methods used to convert digital data into a particular form or format.

[1758] "User terminal" refers to a device used by a user, such as a computer, smartphone, or tablet.

[1759] A "specific object" refers to a specific person, object, scene, or other element within the media data to be analyzed.

[1760] "Means for making modifications" refers to interfaces and tools that allow a user to make changes or improvements to the generated video summary or storyboard.

[1761] Server roles and processing flow

[1762] The server has a means to receive media data. For example, it uses the web application framework Flask to set up an API endpoint that receives HTTP requests from users. Media data such as audio, images, and videos are uploaded to this endpoint. The received media data is saved to disk.

[1763] The server then analyzes the received media data and extracts features. This analysis is performed using AI algorithms powered by TensorFlow. For example, image data is subjected to object detection and segmentation, video data is subjected to the detection of important frames and scenes, and audio data is subjected to text conversion using speech recognition technology. Through these analyses, specific objects contained in the data are also detected.

[1764] The server then uses an emotion engine to evaluate the analyzed media data based on the user's emotions. This allows it to extract specific scenes and frames that match the user's emotions, and generates a summary video and storyboard based on these. The generated summary video and storyboard are then converted into a standard file format (e.g., MP4, JPEG) through an encoding process.

[1765] Finally, the server generates a link to send the generated summary video and storyboard to the user's device and sends it as an HTTP response, allowing the user to check the results on their own device.

[1766] Terminal roles and processing flow

[1767] The device provides a user interface (UI) and supports uploading media data, checking the processed results, and editing. The UI is designed to be intuitive, particularly for easy visual confirmation of video and storyboards. The device also has the function of sending the media data uploaded by the user to the server.

[1768] It also has a function to receive the processing result data returned from the server and display it appropriately.It also includes a function to send the user's emotion data (e.g., voice tone, facial recognition data) to the emotion engine.

[1769] User roles and operation flow

[1770] Users upload the media data (audio, images, video) they have collected to the system via their terminal. When uploading, they also enter related information (title, summary, etc.) and emotional data. They then check the quality and content of the summary video and storyboard returned from the server. If necessary, they make corrections to the received data and re-upload it to the system.

[1771] Specific examples

[1772] For example, if a user wants to make a movie trailer more emotional, the process would be as follows:

[1773] 1. Data upload (user)

[1774] A user uploads a video file of a movie.

[1775] Input emotional data (e.g., emotional ratings during viewing).

[1776] 2. Data analysis and video summary generation (server)

[1777] The server analyzes the media data, extracts specific scenes, and generates a summary video based on the emotion data.

[1778] 3. Receive and check the results (terminal)

[1779] A download link for the generated summary video is generated and transmitted to the user terminal.

[1780] The user reviews the generated movie trailer.

[1781] Prompt Sentence Examples

[1782] "Scene analysis of movie XYZ. Provide media data for scenes that evoke positive emotions. Create a custom trailer."

[1783] As a result, the present invention enables more personalized content to be automatically generated in accordance with the user's emotional data, thereby improving the viewing experience.

[1784] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1785] Step 1:

[1786] Data upload (user)

[1787] Users upload media data (audio, images, videos) to the system using their own devices. Specifically, they select the data through the device interface and enter additional related information (title, summary) and emotional data (emotion rating, etc.). The media data and its related information are then sent to the server.

[1788] Input: User-provided audio, image, and video data, as well as related information and emotional data

[1789] Output: Media data and related information sent to the server

[1790] Step 2:

[1791] Receiving and storing media data (server)

[1792] The server receives media data uploaded by users and saves it to disk via HTTP requests. An API endpoint is set up using a web framework such as Flask, where the media data is saved.

[1793] Input: User uploaded media data

[1794] Output: Media data stored on the server's disk

[1795] Step 3:

[1796] Data analysis and feature extraction (server)

[1797] The server analyzes the received media data and extracts features. Using TensorFlow, the AI ​​model performs image and voice recognition, object detection, scene extraction, etc. It also includes a means to detect specific objects in each media data.

[1798] Input: Stored media data

[1799] Output: Extracted feature data (key scenes, audio text, detected objects)

[1800] Step 4:

[1801] Evaluation by emotion engine (server)

[1802] The emotion engine evaluates the user's emotions based on the analyzed media data. For example, an emotion recognition AI model analyzes the user's emotional data regarding a specific scene or object and outputs an evaluation score.

[1803] Input: Extracted feature data and user emotion data

[1804] Output: Score data based on the user's emotional evaluation

[1805] Step 5:

[1806] Summary video or storyboard generation (server)

[1807] The server generates a summary video or storyboard based on the feature data and emotion evaluation, stitching together important scenes and frames and composing optimal content taking into account the emotion score.

[1808] Input: feature data and emotion evaluation scores

[1809] Output: Generated summary video or storyboard

[1810] Step 6:

[1811] Encoding and sending (server)

[1812] The generated summary video or storyboard is encoded and converted into a standard file format (e.g., MP4, JPEG). A download link is generated for the encoded file to be sent to the user's device, and the link is sent as an HTTP response.

[1813] Input: Generated summary video or storyboard

[1814] Output: Encoded file and download link

[1815] Step 7:

[1816] Check and correct the results (user and device)

[1817] The user checks the summary video and storyboard generated on their device. If the video does not meet their expectations, they can make corrections and upload the corrected data back to the server. The device supports this correction process.

[1818] Input: Encoded file and download link

[1819] Output: Confirmed summary video and storyboard, and correction data as needed

[1820] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1821] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1822] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1823] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1824] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1825] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1826] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1827] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1828] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1829] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1830] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1831] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1832] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1833] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1834] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1835] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1836] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1837] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1838] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1839] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1840] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1841] The following is further disclosed regarding the above embodiment.

[1842] (Claim 1)

[1843] means for receiving media data;

[1844] means for analyzing received media data to extract features;

[1845] a means for generating a summary video or storyboard based on the extracted features;

[1846] The system includes means for encoding the generated summary video or storyboard and transmitting it to a user terminal.

[1847] (Claim 2)

[1848] 10. The system of claim 1, further comprising means for detecting a particular object in the received media data.

[1849] (Claim 3)

[1850] 10. The system of claim 1, further comprising means for allowing a user to make corrections to the reviewed summary video or storyboard and upload it back to the system.

[1851] "Example 1"

[1852] (Claim 1)

[1853] means for receiving media data;

[1854] means for analyzing received media data to extract features;

[1855] a means for generating a summary video or storyboard based on the extracted features;

[1856] means for encoding the generated video summary or storyboard into a standard file format;

[1857] A system including means for transmitting the encoded video summary or storyboard to a user terminal.

[1858] (Claim 2)

[1859] 10. The system of claim 1, further comprising: means for detecting, from the received media data, an object corresponding to each media format.

[1860] (Claim 3)

[1861] 10. The system of claim 1, further comprising means for allowing a user to make corrections to the reviewed summary video or storyboard and upload it back to the system.

[1862] "Application Example 1"

[1863] (Claim 1)

[1864] means for receiving media data;

[1865] means for analyzing received media data to extract features;

[1866] a means for generating a summary video or storyboard based on the extracted features;

[1867] means for encoding the generated summary video or storyboard and transmitting it to a user terminal;

[1868] means for using a protocol to securely transmit the encoded data to a user terminal;

[1869] a means for providing a user interface to support uploading media data, checking the processing results, and editing;

[1870] The system displays the generated summary video on a smart device and includes a means to share it on social media.

[1871] (Claim 2)

[1872] 10. The system of claim 1, further comprising means for detecting specific objects in the received media data and means for applying AI algorithms to identify significant scenes and frames.

[1873] (Claim 3)

[1874] 2. The system according to claim 1, further comprising means for allowing the user to make corrections to the summary video or storyboard confirmed by the user and upload it again to the system.

[1875] "Example 2: Combining Emotion Engines"

[1876] (Claim 1)

[1877] means for receiving media data;

[1878] means for analyzing received media data to extract features;

[1879] a means for generating a summary video or storyboard based on the extracted features;

[1880] means for encoding the generated summary video or storyboard and transmitting it to a user terminal;

[1881] A means for analyzing the emotion data of the user and optimizing the content of the summary video or storyboard based on the result of the analysis;

[1882] The system includes a means for preferentially selecting specific scenes or frames based on the user's emotional data.

[1883] (Claim 2)

[1884] 10. The system of claim 1, further comprising means for detecting a particular object in the received media data.

[1885] (Claim 3)

[1886] 10. The system of claim 1, further comprising means for allowing a user to make corrections to the reviewed summary video or storyboard and upload it back to the system.

[1887] "Application example 2 when combining emotion engines"

[1888] (Claim 1)

[1889] means for receiving media data;

[1890] means for analyzing received media data to extract features;

[1891] means for evaluating the analyzed media data based on the user's emotions using an emotion engine;

[1892] a means for generating a summary video or storyboard based on the extracted features and emotion evaluation;

[1893] The system includes means for encoding the generated summary video or storyboard and transmitting it to a user terminal.

[1894] (Claim 2)

[1895] 10. The system of claim 1, further comprising means for detecting a particular object in the received media data.

[1896] (Claim 3)

[1897] 10. The system of claim 1, further comprising means for allowing a user to make corrections to the reviewed summary video or storyboard and upload it back to the system. [Explanation of symbols]

[1898] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for receiving media data; means for analyzing received media data to extract features; a means for generating a summary video or storyboard based on the extracted features; The system includes means for encoding the generated summary video or storyboard and transmitting it to a user terminal.

2. The system of claim 1 , further comprising means for detecting a particular object in the received media data.

3. 2. The system of claim 1, further comprising means for allowing a user to make corrections to the video summary or storyboard that the user has confirmed and upload the corrections to the system again.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A