System
The system addresses the challenge of creating high-quality videos by automatically extracting editing elements from popular videos and using generative AI to generate and edit videos efficiently, allowing users to adapt to trends with user corrections.
Patent Information
- Application Number
- JP2024130400
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Creating high-quality videos requires significant time and effort, and existing systems struggle to efficiently incorporate popular editing elements and quickly adapt to changing trends.
A system that collects metadata from popular videos, analyzes editing elements, and uses generative AI to automatically generate videos based on user instructions, allowing for efficient and flexible video production with user corrections.
Enables users to efficiently create high-quality videos that adapt to trends by automatically extracting and applying editing elements, with the ability to make corrections and re-edit as needed.
Smart Images

Figure 2026028102000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, with the spread of video platforms, many content creators are seeking methods for creating high-quality videos that are in line with trends. However, the editing work required to create high-quality videos requires a great deal of time and effort, so there is a need for a means to efficiently generate high-quality videos. There is also a growing need for flexible video production systems that can quickly respond to changes in video trends. As a means of solving this problem, the objective of this invention is to provide technology that extracts editing elements from popular videos on video platforms and automatically generates and edits videos. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for collecting metadata of popular videos from a video platform, a means for analyzing the collected metadata to extract editing elements for the video, a means for coding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, a means for receiving video composition and editing instructions from a user, and a generating artificial intelligence means for automatically generating videos based on the received composition and editing instructions. This system enables users to efficiently generate high-quality videos and quickly respond to changing trends. Furthermore, the system includes a means for evaluating the quality of the generated video based on received video viewing data and updating and improving the editing elements in accordance with the evaluation, and a means for users to input correction instructions for the generated video and re-edit the video based on the instructions, thereby achieving flexible, high-quality video production.
[0006] A "video platform" is an online service that enables users to upload, watch, and share videos.
[0007] "Metadata" refers to information related to a video, specifically data such as the number of views, ratings, comments, tags, and descriptions.
[0008] "Editing elements" are elements that determine the composition and direction of a video, and specifically include subtitles, background music, the frequency and timing of cuts, etc.
[0009] "Generative AI" refers to artificial intelligence technology that has the ability to automatically generate and edit videos based on specified instructions.
[0010] "User" refers to an individual or organization that uses the video generation system.
[0011] "Composition and Editing Instructions" means the requirements and desired editing details specified by the user regarding the production of the video.
[0012] "Viewing data" refers to data that indicates the user's viewing behavior with respect to the generated video, and includes, for example, the number of plays, viewing time, viewer feedback, and the like.
[0013] "Editing instructions" refer to specific edit instructions for changes or additions that the user makes to the generated video. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[0036] System Overview
[0037] The system consists of the following main components:
[0038] server
[0039] Terminal
[0040] Generative artificial intelligence (AI)
[0041] System program and processing flow
[0042] The system's program is designed to operate in cooperation with the server, terminals, and artificial intelligence for generation. The roles and specific operations of each entity are explained below.
[0043] server
[0044] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[0045] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[0046] Terminal
[0047] Users use their own devices to input the structure and editing instructions for the video they want to create in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0048] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[0049] Generative artificial intelligence (AI)
[0050] The artificial intelligence (AI) for generation is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video.
[0051] The generated video is provided to the user, who can then make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edits are then sent back to the server, where the AI used for generation re-edits the video. Once edits are complete, the video is provided to the user again.
[0052] Specific examples
[0053] Below are some specific usage examples.
[0054] 1. Data Collection and Analysis
[0055] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[0056] 2. Entering user instructions
[0057] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[0058] 3. Video Generation
[0059] The server receives the user's instructions and sends them to the AI for generation, which then edits the video based on the received instructions and provides it to the user.
[0060] 4. Minor corrections
[0061] The user checks the generated video and makes any necessary corrections, such as "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends the correction instructions back to the AI, which then re-edits the video.
[0062] In this way, the present invention enables users to efficiently generate high-quality videos, realizing video production that quickly responds to trends.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] The server uses the video platform's API to collect metadata for popular videos, including the number of views, ratings, comments, tags, and descriptions.
[0066] Step 2:
[0067] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements for the video, such as background music selection, subtitle display format, and cut timing.
[0068] Step 3:
[0069] The server encodes the extracted editing elements and converts them into an instruction format that can be understood by a generative artificial intelligence (AI), which includes specific editing operations and parameters.
[0070] Step 4:
[0071] Users use their devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0072] Step 5:
[0073] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[0074] Step 6:
[0075] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[0076] Step 7:
[0077] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[0078] Step 8:
[0079] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[0080] Step 9:
[0081] The user can check the generated video and input any necessary corrections into the device, such as "move the subtitles a little higher" or "lower the volume of the background music."
[0082] Step 10:
[0083] The terminal sends the user's correction instructions to the server, which include specific changes.
[0084] Step 11:
[0085] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[0086] Step 12:
[0087] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[0088] Step 13:
[0089] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[0090] This series of steps allows users to efficiently generate high-quality videos and keep up with trends.
[0091] Example 1
[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0093] In recent years, creating popular videos on video platforms requires advanced editing skills and a large amount of time. This makes it difficult for ordinary users to effectively create high-quality videos. Furthermore, due to a lack of video editing skills, it is difficult to create videos that attract viewers' attention. To solve this problem, a system that allows users to easily create high-quality videos is needed.
[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0095] In this invention, the server includes means for collecting information on highly rated videos from a video sharing service, means for analyzing the collected information to extract editing elements for the videos, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, means for receiving video composition and editing instructions from a user, generating artificial intelligence means for automatically generating videos based on the received composition and editing instructions, means for a user to input correction instructions for the generated videos, and means for re-editing the generated videos based on the correction instructions. This enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[0096] A "video sharing service" is a platform on the Internet that allows users to upload, watch, and share videos.
[0097] A "highly rated video" is a video that has received many views, high ratings, and positive comments on a video sharing service.
[0098] "Information" refers to data such as video metadata, number of views, number of ratings, number of comments, tags, and description.
[0099] "Analysis" is the process of assigning meaning to collected information and recognizing patterns according to a specific purpose.
[0100] "Editing elements" are specific components related to video editing, such as background music selection, subtitle style, and frequency of cuts.
[0101] "Extraction" is the act of extracting specific editing elements from the analyzed information.
[0102] "Encoding" is the process of converting the extracted editing elements into a format that can be understood by the generating artificial intelligence.
[0103] "Generative AI" is an AI system that automatically generates videos based on specified composition and editing elements.
[0104] "User" refers to the general user who uses the system to generate and modify videos.
[0105] "Composition" refers to the overall components of a video, such as the length of the video, the length of the intro, the timing of subtitles, and the type of background music used.
[0106] "Input" is the act of transmitting instructions or corrections from the user to the system.
[0107] "Re-editing" refers to additional editing work performed on an already generated video based on correction instructions from the user.
[0108] This invention is a system that automatically collects and analyzes the editing elements of popular videos on video sharing services and generates high-quality videos based on the results. The system includes the following main components: a server, a terminal, and artificial intelligence (AI) for generation.
[0109] server
[0110] The server collects metadata of highly rated videos from a video sharing service (e.g., YouTube). Specifically, the server periodically obtains the number of views, ratings, comments, tags, descriptions, etc. of each video using an API, and stores this information in a database. The collected metadata is then analyzed using a machine learning algorithm (e.g., scikit-learn) to extract common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos. The extracted editing elements are encoded into a format understandable by the generative artificial intelligence. This encoded data is then used as editing instructions for the videos that are generated later.
[0111] Terminal
[0112] The user uses a device (e.g., a PC or smartphone) to input video composition and editing instructions in text format. Specific examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends the user's input to the server, which then sends editing instructions to the artificial intelligence used for generation based on these instructions.
[0113] Generative artificial intelligence (AI)
[0114] Generative artificial intelligence (AI) plays a role in automatically generating videos according to editing instructions sent from the server. For example, the AI uses video editing software (e.g., Adobe Premiere Pro API) or a video library (e.g., FFmpeg) to apply the specified composition and production elements to generate a video. The generated video is then provided to the user via the server, who can review the video on their device and make minor corrections as needed. For example, corrections could include "changing the position of the subtitles" or "adjusting the volume of the background music." The server then sends these corrections back to the generative AI, which then re-edits the video. The final, edited video is then provided to the user via the server again.
[0115] Specific examples
[0116] Below are some specific usage examples.
[0117] 1. Data Collection and Analysis
[0118] The server collects metadata from video sharing services and analyzes and extracts editing elements of popular videos, such as background music and subtitle styles common to videos with over 1,000,000 views.
[0119] 2. Entering user instructions
[0120] The user enters information in text format on the device, such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, the type of background music to be used, etc. An example of a specific prompt is, "The video is 3 minutes long, and the intro will be displayed for the first 30 seconds. Subtitles will be displayed every minute, and upbeat background music will be used."
[0121] 3. Video Generation
[0122] The server receives the user's instructions and sends them to the AI generator, which then edits the video based on the instructions and provides it to the user via the server.
[0123] 4. Minor corrections
[0124] The user checks the generated video and instructs the AI to make any necessary corrections. For example, the server may request that the user increase the font size of the subtitles or decrease the volume of the background music. The server then sends the correction instructions back to the AI, which then re-edits the video.
[0125] In this way, the present invention enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[0126] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0127] System processing flow
[0128] Step 1:
[0129] The server collects metadata of highly rated videos from the video sharing service. Specifically, it uses the API to obtain information such as the number of views, ratings, comments, tags, and descriptions of videos in JSON format. To obtain this data, it uses a Python library (e.g., requests). The API endpoint and authentication information are required as input, and the collected metadata is obtained as output.
[0130] Step 2:
[0131] The server analyzes the collected metadata and uses machine learning algorithms (e.g., scikit-learn) to extract editing elements. Specifically, it applies clustering and classification algorithms to recognize the characteristics of popular videos. The collected metadata is required as input, and the output is the editing elements common to popular videos (e.g., background music selection, subtitle style, frequency of cuts, etc.).
[0132] Step 3:
[0133] The server encodes the extracted edit elements into a format that can be understood by the generative artificial intelligence (AI). This encoding process converts the edit elements into XML and then back into JSON format. The input is the edit elements obtained as a result of the analysis, and the output is the encoded data.
[0134] Step 4:
[0135] The device sends the video composition and editing instructions entered by the user (e.g., "video length is 3 minutes," "first 30 seconds is intro," "display subtitles every minute," "use upbeat background music") to the server. The input requires the text instructions entered by the user, and the output is the instruction data to be sent to the server.
[0136] Step 5:
[0137] The server analyzes the editing instructions received from the user, converts them into an appropriate format, and sends them to the AI for generation. The input requires instruction data from the user, and the output is the editing instructions sent to the AI for generation.
[0138] Step 6:
[0139] Generative artificial intelligence (AI) automatically generates videos based on the editing instructions it receives. This process uses video editing software (e.g., Adobe Premiere Pro API) and video libraries (e.g., FFmpeg). The input is the editing instructions sent from the server, and the output is the generated video file.
[0140] Step 7:
[0141] The server sends the generated video to the user's device. The input requires a video file provided by the AI for generation, and the output is the video data sent to the user's device.
[0142] Step 8:
[0143] The user checks the generated video and inputs correction instructions as necessary. For example, specific corrections can be input, such as "changing the display position of subtitles" or "adjusting the volume of background music." The input requires the generated video and correction instructions, and the output is correction instruction data.
[0144] Step 9:
[0145] The terminal transmits the correction instructions input by the user to the server. The input requires correction instruction data from the user, and the output obtains the correction instruction data to be transmitted to the server.
[0146] Step 10:
[0147] The server then sends the received correction instructions back to the AI generator, which then re-edits the video. The input requires correction instructions from the user, and the output is a re-edited video file.
[0148] Step 11:
[0149] The server then sends the re-edited video back to the user's device. The input is the re-edited video file, and the output is the final video data sent to the user's device.
[0150] (Application example 1)
[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0152] Modern content distribution services require the rapid generation of high-quality videos that capture viewers' attention. However, video production is time-consuming and labor-intensive, and incorporating popular editing elements is particularly difficult. Furthermore, there are few ways for users to reflect their desired composition and editing, and the process of making corrections or re-editing generated videos is cumbersome. A method that solves these issues and enables efficient, high-quality video generation is needed.
[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0154] In this invention, the server includes a means for collecting metadata of popular videos from a video platform, a means for analyzing the collected metadata to extract editing elements for the video, and a means for coding the extracted editing elements and providing them in the form of instructions to a generating AI. This allows users to input desired video composition and editing instructions using their own devices, and the generating AI automatically generates videos based on those instructions. The system also includes a means for linking multiple video clips and a means for users to input correction instructions for generated videos via a smart device and re-edit them. This allows users to efficiently generate high-quality videos and flexibly correct and re-edit them.
[0155] A "video platform" is an internet-based service that hosts video content and enables users to upload, view, and share it.
[0156] "Metadata" is data that includes information about the attributes and structure of the data, and specifically refers to the number of times a video has been played, the number of ratings, the number of comments, tags, descriptions, etc.
[0157] "Editing elements" refer to specific elements and techniques used in the production and editing of a video, including background music selection, subtitle style, and frequency of cuts.
[0158] "Generative AI" is an AI system that has the ability to automatically generate videos based on specified data using pre-trained models and algorithms.
[0159] A "smart device" is an electronic device such as a mobile phone, tablet, or smart glasses that has internet connectivity and allows users to input and operate information.
[0160] A "prompt sentence" refers to a text-based input sentence used to input specific instructions and conditions to a generative AI model, and includes specific instructions necessary for generating and editing videos.
[0161] "Viewing data" refers to data relating to the behavior and evaluation of a user when watching a video, and includes the number of views, viewing time, ratings, comments, and the like.
[0162] "Modification instructions" refer to instructions input in text format that the user wishes to make changes or adjustments to the generated video.
[0163] A "generative model" refers to an algorithm or mathematical model that has the ability to generate various data based on specific conditions or instructions.
[0164] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[0165] System Overview
[0166] The system consists of the following main components:
[0167] server
[0168] Terminal
[0169] Generative artificial intelligence (AI)
[0170] server
[0171] The server is responsible for the following functions:
[0172] 1. Metadata collection
[0173] Using APIs from video platforms, metadata (number of views, number of ratings, number of comments, tags, descriptions, etc.) of popular videos is periodically collected.
[0174] 2. Extracting editing elements
[0175] The collected metadata is analyzed and common editing elements (such as background music selection, subtitle style, and frequency of cuts) among popular videos are extracted using a machine learning algorithm.
[0176] 3. Coding and Data Conversion
[0177] The extracted editing elements are coded and converted into a format that can be understood by a generative artificial intelligence (AI). This coded data is then used as editing instructions for the generated video.
[0178] Terminal
[0179] The user uses a device (smartphone or PC) to perform the following operations:
[0180] 1. Enter editing instructions
[0181] You can enter detailed instructions for video composition and editing in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0182] 2. Enter correction instructions
[0183] Check the generated video and enter corrections as necessary (e.g., change the position of subtitles, adjust the volume of background music).
[0184] Generative artificial intelligence (AI)
[0185] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server:
[0186] 1. Automatic video generation
[0187] Use standard video editing software or libraries (e.g., MoviePy) to generate a video by applying the specified structure and production elements.
[0188] 2. Video Linking
[0189] Combine multiple video clips and apply any effects or edits you want.
[0190] 3. Updates
[0191] The system receives correction instructions from the user and re-edits the video. Through this process, the final video is generated according to the user's wishes.
[0192] Specific examples
[0193] Data collection and analysis
[0194] 1. The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos. For example, it can extract the background music and subtitle styles common to videos with over 1,000,000 views.
[0195] Entering user instructions
[0196] 2. The user enters information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on the device.
[0197] example:
[0198] The video is 3 minutes long
[0199] The first 30 seconds is an intro
[0200] Display subtitles every minute
[0201] Use upbeat background music
[0202] Video Generation
[0203] 3. The server receives the user's instructions and sends them to the AI for generation. The AI then edits the video based on the received instructions and provides it to the user.
[0204] Minor correction
[0205] 4. The user reviews the generated video and makes any necessary corrections. For example, "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends these corrections to the AI, which then re-edits the video.
[0206] This allows users to quickly generate high-quality videos efficiently, realizing video production that quickly responds to trends.
[0207] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0208] Step 1:
[0209] The server periodically collects metadata of popular videos from the video platform using APIs. Specifically, it obtains data such as the number of views, ratings, comments, tags, and descriptions. This provides detailed information about popular videos as input, and the collected metadata as output.
[0210] Step 2:
[0211] The server extracts editing elements of popular videos by analyzing the collected metadata. Specifically, it uses a machine learning algorithm to analyze data such as the number of views, number of ratings, and tags of each video, and detects common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.). This uses the metadata as input and provides the extracted editing elements as output.
[0212] Step 3:
[0213] The server uses a means to encode the extracted editing elements and convert them into a format understandable by the generating artificial intelligence. Specifically, the server converts the extracted editing elements into a specific format (e.g., JSON format) and provides them to the generating artificial intelligence in the form of instructions. This uses the editing elements as input and generates coded data as output.
[0214] Step 4:
[0215] The user uses their own device to input the video's structure and editing instructions in text format. Examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The user's editing instructions are then input to the device. Text editing instructions are generated as output.
[0216] Step 5:
[0217] The server uses a generating AI to automatically generate a video based on the composition and editing instructions received from the user. Specifically, the server analyzes the editing instructions and creates a video by applying the specified composition and production elements according to the instructions. This uses the user's editing instructions as input and produces a generated video as output.
[0218] Step 6:
[0219] The generative AI uses a means of concatenating multiple video clips, for example using the MoviePy library, and applies a specified order and effects, taking the video clips as input and generating the concatenated video as output.
[0220] Step 7:
[0221] The user checks the generated video and uses a means to input correction instructions. Specifically, while playing the generated video, the user inputs correction instructions in text format, such as "change the position of the subtitles" or "adjust the volume of the background music." This inputs the correction instructions into the terminal, and generates correction instructions as output.
[0222] Step 8:
[0223] The server uses a means to re-edit the generated video based on the correction instructions. Specifically, the server analyzes the correction instructions from the user, and the generating AI re-edits the video to generate a final video that reflects the corrections. This uses the correction instructions as input and generates the final video as output.
[0224] By following these steps, users can efficiently generate high-quality videos and then modify and re-edit them as needed.
[0225] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0226] This invention is a system that automatically extracts and analyzes the editing elements of popular videos on video platforms and automatically generates high-quality videos based on the results. Furthermore, this invention improves the user experience by combining it with an emotion engine that recognizes user emotions.
[0227] System Overview
[0228] The system consists of the following main components:
[0229] server
[0230] Terminal
[0231] Generative artificial intelligence (AI)
[0232] Emotion Engine
[0233] System program and processing flow
[0234] The system's program is designed to have the server, terminal, generative artificial intelligence (AI), and emotion engine work together. The roles and specific operations of each entity are explained below.
[0235] server
[0236] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[0237] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[0238] Terminal
[0239] Users use their own devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0240] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[0241] Generative artificial intelligence (AI)
[0242] The generative artificial intelligence (AI) is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video. The generated video is then provided to the user, who can make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edit instructions are then sent back to the server, and the generative AI re-edits the video. Once edits are complete, the video is provided to the user again.
[0243] Emotion Engine
[0244] The emotion engine has the ability to analyze emotions from the user's facial expressions, voice, text input, etc. It recognizes emotions in real time while the user is watching a video or entering editing instructions, and analyzes that data.
[0245] The emotion data obtained from the emotion engine is sent to a server. The server uses this data to adjust editing instructions for the artificial intelligence. For example, if a user expresses positive emotion in a particular part of the video, the server can edit it to emphasize that part. If a user expresses negative emotion, the server can issue instructions to modify or delete that element.
[0246] Specific examples
[0247] Below are some specific usage examples.
[0248] 1. Data Collection and Analysis
[0249] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[0250] 2. Entering user instructions
[0251] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[0252] 3. Collecting Emotional Data
[0253] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[0254] 4. Video Generation
[0255] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[0256] 5. Minor edits and emotional reflection
[0257] The user can check the generated video and indicate any necessary corrections. The emotion engine also analyzes the user's emotions during the editing process, and the server readjusts the edited content based on that information.
[0258] Through this series of steps, it becomes possible to efficiently generate high-quality videos that reflect the user's emotions.
[0259] The processing flow will be explained below.
[0260] Step 1:
[0261] The server uses the video platform's API to collect metadata about popular videos, including the number of views, ratings, comments, tags, descriptions, and other data.
[0262] Step 2:
[0263] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements from the video, such as background music selection, subtitle display format, and timing of cuts.
[0264] Step 3:
[0265] The server encodes the extracted editing elements and converts them into instructions that can be understood by the generative artificial intelligence (AI), including specific editing operations and parameters.
[0266] Step 4:
[0267] Users use their own devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0268] Step 5:
[0269] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[0270] Step 6:
[0271] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[0272] Step 7:
[0273] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[0274] Step 8:
[0275] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[0276] Step 9:
[0277] The emotion engine analyzes the user's facial expressions, voice, text input, etc. in real time while the user is watching a video, and collects emotional data.
[0278] Step 10:
[0279] The emotion engine uses the analyzed emotional data to identify parts of the video where the user had a positive or negative reaction. For example, a smile on the user's face is judged as positive, while a furrowed brow is judged as negative.
[0280] Step 11:
[0281] The server receives the emotion data sent from the emotion engine and reflects it in the artificial intelligence for generation, for example, by issuing instructions to edit parts that show positive emotions or to correct parts that show negative emotions.
[0282] Step 12:
[0283] The user can then review the generated video and make any necessary corrections, such as moving the subtitles a little higher or lowering the volume of the background music, via the device.
[0284] Step 13:
[0285] The terminal sends the user's correction instructions to the server, which include specific changes and desired corrections.
[0286] Step 14:
[0287] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[0288] Step 15:
[0289] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[0290] Step 16:
[0291] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[0292] This series of steps enables efficient generation of high-quality videos that reflect user emotions and realizes video production that is in line with trends.
[0293] Example 2
[0294] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0295] Conventional video generation systems have had difficulty in properly reflecting the editing elements desired by users. Furthermore, they lacked video generation functionality that took user emotions into account, making it impossible to improve the quality of the user experience. This resulted in problems such as a decline in the quality of the generated videos and a decline in user satisfaction.
[0296] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0297] In this invention, the server includes means for collecting viewing data from a video providing service, means for analyzing the collected viewing data and extracting video editing elements, means for encoding the extracted editing elements and providing them to the generation AI in the form of instructions, and emotion recognition means for analyzing the user's emotions and adjusting the editing instructions based on that data, thereby enabling the efficient automatic generation of high-quality videos that reflect the user's emotions.
[0298] "Viewing data" refers to information collected from video provision services, such as the number of views, ratings, comments, tags, and descriptions.
[0299] "Editing elements" are elements that show specific patterns or trends, such as background music selection, subtitle style, and frequency of cuts, extracted by analyzing viewing data.
[0300] "Generative artificial intelligence" refers to algorithms and programs that automatically generate videos based on instruction-style data provided by the server.
[0301] The "emotion recognition means" is a system that analyzes emotions in real time from the user's facial expressions, voice, and text input, and feeds that data back to the artificial intelligence used for generation.
[0302] "Encoding" is the process of converting extracted editorial elements into a format that can be understood by a generative artificial intelligence.
[0303] A "correction instruction" is a request for specific changes or adjustments that the user makes to the generated video.
[0304] "Re-editing" is the process in which the generating AI re-edits the video based on the user's correction instructions.
[0305] This invention is a system that automatically generates high-quality videos based on viewing data collected from video provision services. This system consists of a server, a terminal, artificial intelligence for generation, and emotion recognition means.
[0306] The server collects viewing data from video-providing services (e.g., video streaming platforms) using APIs. The collected data includes the number of views, ratings, comments, tags, and descriptions. The server then analyzes this data using machine learning algorithms to extract specific patterns and trends. For example, it can identify common editing elements (such as background music selection, subtitle style, and frequency of cuts) among videos with high views. The results of this analysis are coded as extracted editing elements and converted into a format understandable by the generative AI.
[0307] Users use their devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends these inputs to the server, which converts them into appropriate editing instructions and sends them to the AI generator.
[0308] The artificial intelligence (AI) for generation automatically generates videos based on editing instructions sent from the server. This AI generates videos by applying the specified composition and production elements using video editing software such as Adobe Premiere Pro or FFmpeg. The generated videos are provided to users, who can play and check them on their devices.
[0309] The emotion recognition means analyzes the emotions of the user in real time while watching the video or inputting editing instructions. Emotional data is acquired from the user's facial expressions, voice, and text input, and sent to the server. This emotional data is fed back to the generating AI and used to emphasize or modify specific parts of the video. For example, editing instructions are issued to emphasize parts where the user smiles more and modify parts where the user frowns.
[0310] As an example, consider the following prompt:
[0311] The video is 3 minutes long
[0312] "The first 30 seconds are the intro"
[0313] "Show subtitles every minute"
[0314] "Use upbeat background music"
[0315] This allows us to automatically generate videos that meet the user's preferences, and further improve the user experience by utilizing emotion recognition.
[0316] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0317] Step 1:
[0318] The server collects viewing data from the video provider. Specifically, it periodically obtains metadata such as the number of views, ratings, comments, tags, and descriptions using an API. It sends an API request as input and obtains viewing data as output. This data is stored in the server's database.
[0319] Step 2:
[0320] The server analyzes the collected viewing data. It uses machine learning algorithms to analyze the data and extract certain patterns and trends, such as popular background music, subtitle styles, or frequency of cuts. It uses the viewing data collected in step 1 as input and obtains the analysis results as output. These analysis results are stored on the server for use in later steps.
[0321] Step 3:
[0322] The server encodes the analysis results and converts the extracted editing elements into a format that the AI can understand. Specifically, it converts the data into JSON format and prepares it for passing to the AI. It uses the analysis results as input and generates coded editing elements as output.
[0323] Step 4:
[0324] Users input video composition and editing instructions in text format using their own devices. Examples of input include "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device collects the user's instructions as input and sends them to the server.
[0325] Step 5:
[0326] The server receives instructions from the user. After receiving these instructions, it analyzes them, converts them into appropriate editing instructions, and sends them to the AI generator. The input is the user's instructions, and the output is instructions converted into a format that can be sent to the AI generator.
[0327] Step 6:
[0328] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server. It uses video editing software (e.g., Adobe Premiere Pro or FFmpeg) to edit the videos as instructed. It receives the editing instructions as input and provides the generated video as output to the user.
[0329] Step 7:
[0330] The user checks the generated video and inputs any necessary corrections through the device. For example, this includes instructions such as "change the display position of the subtitles to the bottom right" or "reduce the background music volume by 20%." The device collects the corrections as input and sends them back to the server.
[0331] Step 8:
[0332] The server receives the correction instructions and sends them to the AI for generation. The received instructions are analyzed and sent back to the AI as appropriate re-editing instructions. The input is the user's correction instructions, and the output is the re-editing instructions.
[0333] Step 9:
[0334] The generative artificial intelligence (AI) re-edits the video based on the correction instructions. The corrections are reflected again using video editing software to generate the final video. The correction instructions are received as input, and the corrected video is provided to the user as output.
[0335] Step 10:
[0336] The emotion recognition means analyzes the user's facial expressions, voice, and text input to obtain emotion data. It performs the analysis in real time while the user is watching the video and sends the results to the server. It collects the user's real-time data as input and sends the analyzed emotion data as output to the server.
[0337] Step 11:
[0338] The server sends additional editing instructions to the AI based on the emotion data. For example, it may emphasize parts that show positive emotions and modify or delete parts that show negative emotions. The server receives emotion data as input and sends additional editing instructions to the AI as output.
[0339] (Application example 2)
[0340] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0341] In modern virtual stores, it is important to generate efficient, high-quality promotional videos to attract customer attention. However, existing methods require a significant amount of manual effort and time for video editing and generation, and it is difficult to provide optimal promotional videos that reflect customers' real-time emotions. Furthermore, it is difficult to immediately reflect users' feedback while they are watching, making it difficult to improve user satisfaction.
[0342] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting metadata of popular videos from a video platform, means for analyzing the collected metadata and extracting editing elements for the video, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, emotion recognition means for collecting real-time emotional data from users, means for adjusting the editing content of the video using the collected emotional data, and means for viewing and editing the video using a device such as smart glasses or a head-mounted display as a user terminal. This makes it possible to automatically generate optimal promotional videos that reflect the user's emotions and effectively improve user satisfaction and interest in the virtual store.
[0343] A "video platform" is a general term for services that allow various users to upload, watch, and share videos online.
[0344] "Metadata" refers to data related to a video, including information such as the number of views, ratings, comments, tags, and descriptions.
[0345] "Generative AI" refers to an AI system designed to automatically generate videos based on collected data and user instructions.
[0346] "Emotion recognition means" refers to technology or equipment that analyzes a user's facial expressions, voice, etc. to detect emotions in real time.
[0347] The term "user terminal" refers to a device used by a user, and in the present invention particularly includes devices such as smart glasses and head-mounted displays.
[0348] "Emotion Data" refers to information collected by an emotion recognition means about a user's real-time emotional state.
[0349] "Editing elements" are elements that indicate optimal patterns and techniques for creating and editing videos, and include specific elements such as background music, subtitle style, and frequency of cuts.
[0350] "Promotional videos" refer to videos created for the purpose of promoting products or services within a virtual store.
[0351] A "virtual store" refers to a virtual store that exists on the Internet or in a virtual reality space, providing an environment in which users can visually view products and services.
[0352] The present invention relates to a system for automatically generating promotional videos in a virtual store and optimizing the videos based on real-time visitor sentiment. Specific embodiments are described below.
[0353] overview
[0354] The system consists of the following main components:
[0355] server
[0356] Devices (smart glasses, head-mounted displays, etc.)
[0357] Generative artificial intelligence (AI)
[0358] Emotion Recognition Engine
[0359] Hardware and software used
[0360] 1. Server:
[0361] The server uses an API from a video platform (e.g., YouTube) to collect metadata for popular videos, obtaining information such as the number of views, ratings, comments, tags, and descriptions.
[0362] Based on the collected metadata, machine learning libraries (e.g., scikit-learn, TensorFlow) are used to analyze specific patterns and trends and extract editorial elements.
[0363] The extracted editing elements are coded and converted into a format that can be understood by the generative artificial intelligence.
[0364] 2. User Device:
[0365] The user inputs the video composition and editing instructions in text format through a terminal (smart glasses or a head-mounted display).
[0366] Real-time emotional data is collected while the user is watching through the camera and microphone built into the device.
[0367] 3. Generative Artificial Intelligence (AI):
[0368] The artificial intelligence for generation automatically generates videos based on editing instructions and emotional data sent from the server.
[0369] The generated video is provided to the user once, and modifications can be made as needed.
[0370] 4. Emotion Recognition Engine:
[0371] The emotion recognition engine analyzes the user's facial expressions and voice in real time to collect emotional data.
[0372] The collected emotional data is sent to a server and used by the generative AI to adjust the editing content of the video.
[0373] Program Processing Overview
[0374] Data collection and analysis:
[0375] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[0376] Enter user instructions:
[0377] The user inputs information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used on the device.
[0378] Emotion data collection:
[0379] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[0380] Video Generation:
[0381] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[0382] Tweaks and Emotions:
[0383] The user checks the generated video and makes any necessary corrections. The emotion engine also analyzes the user's emotions during the corrections, and the server re-edits the video based on those. For example, the user can input correction instructions such as "change the display position of the subtitles" or "adjust the volume of the background music."
[0384] Specific examples
[0385] Prompt Sentence Examples
[0386] API endpoint: https: / / api.videoplatform.com / metadata
[0387] Parameters: Over 1,000,000 views
[0388] User instructions:
[0389] Length: 3 minutes
[0390] Intro length: 30 seconds
[0391] Subtitles: Every minute
[0392] BGM: Upbeat
[0393] Emotional Data:
[0394] Video Frame: {...}
[0395] Analysis results: Smiling a lot indicates positive emotions, frowning a lot indicates negative emotions
[0396] In this way, the present invention can improve the user experience in a virtual store by automatically generating an optimal promotional video that reflects the user's emotions.
[0397] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0398] Step 1:
[0399] The server collects metadata of popular videos from a video platform using an API. The input is the API endpoint and parameters of the video platform, and the output is the collected metadata (number of views, number of ratings, number of comments, etc.).
[0400] Step 2:
[0401] The server analyzes the collected metadata and extracts the editing elements of the video. The input is the metadata, and the output is the extracted editing elements (e.g., type of background music, subtitle style, frequency of cuts, etc.). Machine learning libraries (e.g., scikit-learn, TensorFlow) are used for the analysis.
[0402] Step 3:
[0403] The server encodes the extracted editing elements and converts them into a format that can be understood by the AI for generation. The input is the editing elements, and the output is coded editing instructions. The specific operation is to map the editing elements to the corresponding code.
[0404] Step 4:
[0405] The user inputs video composition and editing instructions in text format through a device (smart glasses or a head-mounted display). The input is text format instructions, and the output is instruction data sent to the server.
[0406] Step 5:
[0407] The server sends data to the generating AI based on the composition and editing instructions received from the user, where the input is the user's instruction data and coded editing instructions, and the output is the data sent to the generating AI.
[0408] Step 6:
[0409] The artificial intelligence for generation automatically generates videos based on the received composition and editing instructions. The input is data sent from the server, and the output is the generated video file. Specific operations involve the use of video editing software and libraries.
[0410] Step 7:
[0411] While a user is watching a video generated on a device, the emotion recognition means analyzes the user's facial expressions and voice to collect emotion data in real time. The input is the user's facial expressions and voice, and the output is the collected emotion data.
[0412] Step 8:
[0413] The server analyzes the collected emotion data and adjusts the editing content of the video. The input is emotion data, and the output is adjusted editing instructions. Specifically, editing is performed to emphasize positive emotions.
[0414] Step 9:
[0415] The user inputs correction instructions for the video. The input is the user's correction instructions, and the output is correction instruction data sent to the server.
[0416] Step 10:
[0417] The server sends instructions for re-editing to the AI based on the correction instructions and the collected emotion data. The input is the correction instruction data and emotion data, and the output is the data for re-editing.
[0418] Step 11:
[0419] The artificial intelligence for generation re-edits the video based on the re-editing instructions. The input is the data for re-editing, and the output is the corrected video file. Specifically, it edits the specified parts.
[0420] This series of steps enables the system to provide optimal promotional videos within the virtual store that reflect the user's emotions and feedback.
[0421] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0422] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0423] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0424] [Second embodiment]
[0425] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0426] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0427] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0428] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0429] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0430] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0431] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0432] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0433] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0434] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0435] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0436] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0437] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[0438] System Overview
[0439] The system consists of the following main components:
[0440] server
[0441] Terminal
[0442] Generative artificial intelligence (AI)
[0443] System program and processing flow
[0444] The system's program is designed to operate in cooperation with the server, terminals, and artificial intelligence for generation. The roles and specific operations of each entity are explained below.
[0445] server
[0446] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[0447] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[0448] Terminal
[0449] Users use their own devices to input the structure and editing instructions for the video they want to create in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0450] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[0451] Generative artificial intelligence (AI)
[0452] The artificial intelligence (AI) for generation is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video.
[0453] The generated video is provided to the user, who can then make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edits are then sent back to the server, where the AI used for generation re-edits the video. Once edits are complete, the video is provided to the user again.
[0454] Specific examples
[0455] Below are some specific usage examples.
[0456] 1. Data Collection and Analysis
[0457] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[0458] 2. Entering user instructions
[0459] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[0460] 3. Video Generation
[0461] The server receives the user's instructions and sends them to the AI for generation, which then edits the video based on the received instructions and provides it to the user.
[0462] 4. Minor corrections
[0463] The user checks the generated video and makes any necessary corrections, such as "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends the correction instructions back to the AI, which then re-edits the video.
[0464] In this way, the present invention enables users to efficiently generate high-quality videos, realizing video production that quickly responds to trends.
[0465] The processing flow will be explained below.
[0466] Step 1:
[0467] The server uses the video platform's API to collect metadata for popular videos, including the number of views, ratings, comments, tags, and descriptions.
[0468] Step 2:
[0469] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements for the video, such as background music selection, subtitle display format, and cut timing.
[0470] Step 3:
[0471] The server encodes the extracted editing elements and converts them into an instruction format that can be understood by a generative artificial intelligence (AI), which includes specific editing operations and parameters.
[0472] Step 4:
[0473] Users use their devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0474] Step 5:
[0475] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[0476] Step 6:
[0477] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[0478] Step 7:
[0479] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[0480] Step 8:
[0481] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[0482] Step 9:
[0483] The user can check the generated video and input any necessary corrections into the device, such as "move the subtitles a little higher" or "lower the volume of the background music."
[0484] Step 10:
[0485] The terminal sends the user's correction instructions to the server, which include specific changes.
[0486] Step 11:
[0487] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[0488] Step 12:
[0489] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[0490] Step 13:
[0491] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[0492] This series of steps allows users to efficiently generate high-quality videos and keep up with trends.
[0493] Example 1
[0494] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0495] In recent years, creating popular videos on video platforms requires advanced editing skills and a large amount of time. This makes it difficult for ordinary users to effectively create high-quality videos. Furthermore, due to a lack of video editing skills, it is difficult to create videos that attract viewers' attention. To solve this problem, a system that allows users to easily create high-quality videos is needed.
[0496] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0497] In this invention, the server includes means for collecting information on highly rated videos from a video sharing service, means for analyzing the collected information to extract editing elements for the videos, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, means for receiving video composition and editing instructions from a user, generating artificial intelligence means for automatically generating videos based on the received composition and editing instructions, means for a user to input correction instructions for the generated videos, and means for re-editing the generated videos based on the correction instructions. This enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[0498] A "video sharing service" is a platform on the Internet that allows users to upload, watch, and share videos.
[0499] A "highly rated video" is a video that has received many views, high ratings, and positive comments on a video sharing service.
[0500] "Information" refers to data such as video metadata, number of views, number of ratings, number of comments, tags, and description.
[0501] "Analysis" is the process of assigning meaning to collected information and recognizing patterns according to a specific purpose.
[0502] "Editing elements" are specific components related to video editing, such as background music selection, subtitle style, and frequency of cuts.
[0503] "Extraction" is the act of extracting specific editing elements from the analyzed information.
[0504] "Encoding" is the process of converting the extracted editing elements into a format that can be understood by the generating artificial intelligence.
[0505] "Generative AI" is an AI system that automatically generates videos based on specified composition and editing elements.
[0506] "User" refers to the general user who uses the system to generate and modify videos.
[0507] "Composition" refers to the overall components of a video, such as the length of the video, the length of the intro, the timing of subtitles, and the type of background music used.
[0508] "Input" is the act of transmitting instructions or corrections from the user to the system.
[0509] "Re-editing" refers to additional editing work performed on an already generated video based on correction instructions from the user.
[0510] This invention is a system that automatically collects and analyzes the editing elements of popular videos on video sharing services and generates high-quality videos based on the results. The system includes the following main components: a server, a terminal, and artificial intelligence (AI) for generation.
[0511] server
[0512] The server collects metadata of highly rated videos from a video sharing service (e.g., YouTube). Specifically, the server periodically obtains the number of views, ratings, comments, tags, descriptions, etc. of each video using an API, and stores this information in a database. The collected metadata is then analyzed using a machine learning algorithm (e.g., scikit-learn) to extract common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos. The extracted editing elements are encoded into a format understandable by the generative artificial intelligence. This encoded data is then used as editing instructions for the videos that are generated later.
[0513] Terminal
[0514] The user uses a device (e.g., a PC or smartphone) to input video composition and editing instructions in text format. Specific examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends the user's input to the server, which then sends editing instructions to the artificial intelligence used for generation based on these instructions.
[0515] Generative artificial intelligence (AI)
[0516] Generative artificial intelligence (AI) plays a role in automatically generating videos according to editing instructions sent from the server. For example, the AI uses video editing software (e.g., Adobe Premiere Pro API) or a video library (e.g., FFmpeg) to apply the specified composition and production elements to generate a video. The generated video is then provided to the user via the server, who can review the video on their device and make minor corrections as needed. For example, corrections could include "changing the position of the subtitles" or "adjusting the volume of the background music." The server then sends these corrections back to the generative AI, which then re-edits the video. The final, edited video is then provided to the user via the server again.
[0517] Specific examples
[0518] Below are some specific usage examples.
[0519] 1. Data Collection and Analysis
[0520] The server collects metadata from video sharing services and analyzes and extracts editing elements of popular videos, such as background music and subtitle styles common to videos with over 1,000,000 views.
[0521] 2. Entering user instructions
[0522] The user enters information in text format on the device, such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, the type of background music to be used, etc. An example of a specific prompt is, "The video is 3 minutes long, and the intro will be displayed for the first 30 seconds. Subtitles will be displayed every minute, and upbeat background music will be used."
[0523] 3. Video Generation
[0524] The server receives the user's instructions and sends them to the AI generator, which then edits the video based on the instructions and provides it to the user via the server.
[0525] 4. Minor corrections
[0526] The user checks the generated video and instructs the AI to make any necessary corrections. For example, the server may request that the user increase the font size of the subtitles or decrease the volume of the background music. The server then sends the correction instructions back to the AI, which then re-edits the video.
[0527] In this way, the present invention enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[0528] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0529] System processing flow
[0530] Step 1:
[0531] The server collects metadata of highly rated videos from the video sharing service. Specifically, it uses the API to obtain information such as the number of views, ratings, comments, tags, and descriptions of videos in JSON format. To obtain this data, it uses a Python library (e.g., requests). The API endpoint and authentication information are required as input, and the collected metadata is obtained as output.
[0532] Step 2:
[0533] The server analyzes the collected metadata and uses machine learning algorithms (e.g., scikit-learn) to extract editing elements. Specifically, it applies clustering and classification algorithms to recognize the characteristics of popular videos. The collected metadata is required as input, and the output is the editing elements common to popular videos (e.g., background music selection, subtitle style, frequency of cuts, etc.).
[0534] Step 3:
[0535] The server encodes the extracted edit elements into a format that can be understood by the generative artificial intelligence (AI). This encoding process converts the edit elements into XML and then back into JSON format. The input is the edit elements obtained as a result of the analysis, and the output is the encoded data.
[0536] Step 4:
[0537] The device sends the video composition and editing instructions entered by the user (e.g., "video length is 3 minutes," "first 30 seconds is intro," "display subtitles every minute," "use upbeat background music") to the server. The input requires the text instructions entered by the user, and the output is the instruction data to be sent to the server.
[0538] Step 5:
[0539] The server analyzes the editing instructions received from the user, converts them into an appropriate format, and sends them to the AI for generation. The input requires instruction data from the user, and the output is the editing instructions sent to the AI for generation.
[0540] Step 6:
[0541] Generative artificial intelligence (AI) automatically generates videos based on the editing instructions it receives. This process uses video editing software (e.g., Adobe Premiere Pro API) and video libraries (e.g., FFmpeg). The input is the editing instructions sent from the server, and the output is the generated video file.
[0542] Step 7:
[0543] The server sends the generated video to the user's device. The input requires a video file provided by the AI for generation, and the output is the video data sent to the user's device.
[0544] Step 8:
[0545] The user checks the generated video and inputs correction instructions as necessary. For example, specific corrections can be input, such as "changing the display position of subtitles" or "adjusting the volume of background music." The input requires the generated video and correction instructions, and the output is correction instruction data.
[0546] Step 9:
[0547] The terminal transmits the correction instructions input by the user to the server. The input requires correction instruction data from the user, and the output obtains the correction instruction data to be transmitted to the server.
[0548] Step 10:
[0549] The server then sends the received correction instructions back to the AI generator, which then re-edits the video. The input requires correction instructions from the user, and the output is a re-edited video file.
[0550] Step 11:
[0551] The server then sends the re-edited video back to the user's device. The input is the re-edited video file, and the output is the final video data sent to the user's device.
[0552] (Application example 1)
[0553] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0554] Modern content distribution services require the rapid generation of high-quality videos that capture viewers' attention. However, video production is time-consuming and labor-intensive, and incorporating popular editing elements is particularly difficult. Furthermore, there are few ways for users to reflect their desired composition and editing, and the process of making corrections or re-editing generated videos is cumbersome. A method that solves these issues and enables efficient, high-quality video generation is needed.
[0555] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0556] In this invention, the server includes a means for collecting metadata of popular videos from a video platform, a means for analyzing the collected metadata to extract editing elements for the video, and a means for coding the extracted editing elements and providing them in the form of instructions to a generating AI. This allows users to input desired video composition and editing instructions using their own devices, and the generating AI automatically generates videos based on those instructions. The system also includes a means for linking multiple video clips and a means for users to input correction instructions for generated videos via a smart device and re-edit them. This allows users to efficiently generate high-quality videos and flexibly correct and re-edit them.
[0557] A "video platform" is an internet-based service that hosts video content and enables users to upload, view, and share it.
[0558] "Metadata" is data that includes information about the attributes and structure of the data, and specifically refers to the number of times a video has been played, the number of ratings, the number of comments, tags, descriptions, etc.
[0559] "Editing elements" refer to specific elements and techniques used in the production and editing of a video, including background music selection, subtitle style, and frequency of cuts.
[0560] "Generative AI" is an AI system that has the ability to automatically generate videos based on specified data using pre-trained models and algorithms.
[0561] A "smart device" is an electronic device such as a mobile phone, tablet, or smart glasses that has internet connectivity and allows users to input and operate information.
[0562] A "prompt sentence" refers to a text-based input sentence used to input specific instructions and conditions to a generative AI model, and includes specific instructions necessary for generating and editing videos.
[0563] "Viewing data" refers to data relating to the behavior and evaluation of a user when watching a video, and includes the number of views, viewing time, ratings, comments, and the like.
[0564] "Modification instructions" refer to instructions input in text format that the user wishes to make changes or adjustments to the generated video.
[0565] A "generative model" refers to an algorithm or mathematical model that has the ability to generate various data based on specific conditions or instructions.
[0566] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[0567] System Overview
[0568] The system consists of the following main components:
[0569] server
[0570] Terminal
[0571] Generative artificial intelligence (AI)
[0572] server
[0573] The server is responsible for the following functions:
[0574] 1. Metadata collection
[0575] Using APIs from video platforms, metadata (number of views, number of ratings, number of comments, tags, descriptions, etc.) of popular videos is periodically collected.
[0576] 2. Extracting editing elements
[0577] The collected metadata is analyzed and common editing elements (such as background music selection, subtitle style, and frequency of cuts) among popular videos are extracted using a machine learning algorithm.
[0578] 3. Coding and Data Conversion
[0579] The extracted editing elements are coded and converted into a format that can be understood by a generative artificial intelligence (AI). This coded data is then used as editing instructions for the generated video.
[0580] Terminal
[0581] The user uses a device (smartphone or PC) to perform the following operations:
[0582] 1. Enter editing instructions
[0583] You can enter detailed instructions for video composition and editing in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0584] 2. Enter correction instructions
[0585] Check the generated video and enter corrections as necessary (e.g., change the position of subtitles, adjust the volume of background music).
[0586] Generative artificial intelligence (AI)
[0587] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server:
[0588] 1. Automatic video generation
[0589] Use standard video editing software or libraries (e.g., MoviePy) to generate a video by applying the specified structure and production elements.
[0590] 2. Video Linking
[0591] Combine multiple video clips and apply any effects or edits you want.
[0592] 3. Updates
[0593] The system receives correction instructions from the user and re-edits the video. Through this process, the final video is generated according to the user's wishes.
[0594] Specific examples
[0595] Data collection and analysis
[0596] 1. The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos. For example, it can extract the background music and subtitle styles common to videos with over 1,000,000 views.
[0597] Entering user instructions
[0598] 2. The user enters information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on the device.
[0599] example:
[0600] The video is 3 minutes long
[0601] The first 30 seconds is an intro
[0602] Display subtitles every minute
[0603] Use upbeat background music
[0604] Video Generation
[0605] 3. The server receives the user's instructions and sends them to the AI for generation. The AI then edits the video based on the received instructions and provides it to the user.
[0606] Minor correction
[0607] 4. The user reviews the generated video and makes any necessary corrections. For example, "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends these corrections to the AI, which then re-edits the video.
[0608] This allows users to quickly generate high-quality videos efficiently, realizing video production that quickly responds to trends.
[0609] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0610] Step 1:
[0611] The server periodically collects metadata of popular videos from the video platform using APIs. Specifically, it obtains data such as the number of views, ratings, comments, tags, and descriptions. This provides detailed information about popular videos as input, and the collected metadata as output.
[0612] Step 2:
[0613] The server extracts editing elements of popular videos by analyzing the collected metadata. Specifically, it uses a machine learning algorithm to analyze data such as the number of views, number of ratings, and tags of each video, and detects common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.). This uses the metadata as input and provides the extracted editing elements as output.
[0614] Step 3:
[0615] The server uses a means to encode the extracted editing elements and convert them into a format understandable by the generating artificial intelligence. Specifically, the server converts the extracted editing elements into a specific format (e.g., JSON format) and provides them to the generating artificial intelligence in the form of instructions. This uses the editing elements as input and generates coded data as output.
[0616] Step 4:
[0617] The user uses their own device to input the video's structure and editing instructions in text format. Examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The user's editing instructions are then input to the device. Text editing instructions are generated as output.
[0618] Step 5:
[0619] The server uses a generating AI to automatically generate a video based on the composition and editing instructions received from the user. Specifically, the server analyzes the editing instructions and creates a video by applying the specified composition and production elements according to the instructions. This uses the user's editing instructions as input and produces a generated video as output.
[0620] Step 6:
[0621] The generative AI uses a means of concatenating multiple video clips, for example using the MoviePy library, and applies a specified order and effects, taking the video clips as input and generating the concatenated video as output.
[0622] Step 7:
[0623] The user checks the generated video and uses a means to input correction instructions. Specifically, while playing the generated video, the user inputs correction instructions in text format, such as "change the position of the subtitles" or "adjust the volume of the background music." This inputs the correction instructions into the terminal, and generates correction instructions as output.
[0624] Step 8:
[0625] The server uses a means to re-edit the generated video based on the correction instructions. Specifically, the server analyzes the correction instructions from the user, and the generating AI re-edits the video to generate a final video that reflects the corrections. This uses the correction instructions as input and generates the final video as output.
[0626] By following these steps, users can efficiently generate high-quality videos and then modify and re-edit them as needed.
[0627] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0628] This invention is a system that automatically extracts and analyzes the editing elements of popular videos on video platforms and automatically generates high-quality videos based on the results. Furthermore, this invention improves the user experience by combining it with an emotion engine that recognizes user emotions.
[0629] System Overview
[0630] The system consists of the following main components:
[0631] server
[0632] Terminal
[0633] Generative artificial intelligence (AI)
[0634] Emotion Engine
[0635] System program and processing flow
[0636] The system's program is designed to have the server, terminal, generative artificial intelligence (AI), and emotion engine work together. The roles and specific operations of each entity are explained below.
[0637] server
[0638] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[0639] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[0640] Terminal
[0641] Users use their own devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0642] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[0643] Generative artificial intelligence (AI)
[0644] The generative artificial intelligence (AI) is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video. The generated video is then provided to the user, who can make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edit instructions are then sent back to the server, and the generative AI re-edits the video. Once edits are complete, the video is provided to the user again.
[0645] Emotion Engine
[0646] The emotion engine has the ability to analyze emotions from the user's facial expressions, voice, text input, etc. It recognizes emotions in real time while the user is watching a video or entering editing instructions, and analyzes that data.
[0647] The emotion data obtained from the emotion engine is sent to a server. The server uses this data to adjust editing instructions for the artificial intelligence. For example, if a user expresses positive emotion in a particular part of the video, the server can edit it to emphasize that part. If a user expresses negative emotion, the server can issue instructions to modify or delete that element.
[0648] Specific examples
[0649] Below are some specific usage examples.
[0650] 1. Data Collection and Analysis
[0651] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[0652] 2. Entering user instructions
[0653] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[0654] 3. Collecting Emotional Data
[0655] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[0656] 4. Video Generation
[0657] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[0658] 5. Minor edits and emotional reflection
[0659] The user can check the generated video and indicate any necessary corrections. The emotion engine also analyzes the user's emotions during the editing process, and the server readjusts the edited content based on that information.
[0660] Through this series of steps, it becomes possible to efficiently generate high-quality videos that reflect the user's emotions.
[0661] The processing flow will be explained below.
[0662] Step 1:
[0663] The server uses the video platform's API to collect metadata about popular videos, including the number of views, ratings, comments, tags, descriptions, and other data.
[0664] Step 2:
[0665] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements from the video, such as background music selection, subtitle display format, and timing of cuts.
[0666] Step 3:
[0667] The server encodes the extracted editing elements and converts them into instructions that can be understood by the generative artificial intelligence (AI), including specific editing operations and parameters.
[0668] Step 4:
[0669] Users use their own devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0670] Step 5:
[0671] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[0672] Step 6:
[0673] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[0674] Step 7:
[0675] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[0676] Step 8:
[0677] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[0678] Step 9:
[0679] The emotion engine analyzes the user's facial expressions, voice, text input, etc. in real time while the user is watching a video, and collects emotional data.
[0680] Step 10:
[0681] The emotion engine uses the analyzed emotional data to identify parts of the video where the user had a positive or negative reaction. For example, a smile on the user's face is judged as positive, while a furrowed brow is judged as negative.
[0682] Step 11:
[0683] The server receives the emotion data sent from the emotion engine and reflects it in the artificial intelligence for generation, for example, by issuing instructions to edit parts that show positive emotions or to correct parts that show negative emotions.
[0684] Step 12:
[0685] The user can then review the generated video and make any necessary corrections, such as moving the subtitles a little higher or lowering the volume of the background music, via the device.
[0686] Step 13:
[0687] The terminal sends the user's correction instructions to the server, which include specific changes and desired corrections.
[0688] Step 14:
[0689] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[0690] Step 15:
[0691] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[0692] Step 16:
[0693] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[0694] This series of steps enables efficient generation of high-quality videos that reflect user emotions and realizes video production that is in line with trends.
[0695] Example 2
[0696] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0697] Conventional video generation systems have had difficulty in properly reflecting the editing elements desired by users. Furthermore, they lacked video generation functionality that took user emotions into account, making it impossible to improve the quality of the user experience. This resulted in problems such as a decline in the quality of the generated videos and a decline in user satisfaction.
[0698] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0699] In this invention, the server includes means for collecting viewing data from a video providing service, means for analyzing the collected viewing data and extracting video editing elements, means for encoding the extracted editing elements and providing them to the generation AI in the form of instructions, and emotion recognition means for analyzing the user's emotions and adjusting the editing instructions based on that data, thereby enabling the efficient automatic generation of high-quality videos that reflect the user's emotions.
[0700] "Viewing data" refers to information collected from video provision services, such as the number of views, ratings, comments, tags, and descriptions.
[0701] "Editing elements" are elements that show specific patterns or trends, such as background music selection, subtitle style, and frequency of cuts, extracted by analyzing viewing data.
[0702] "Generative artificial intelligence" refers to algorithms and programs that automatically generate videos based on instruction-style data provided by the server.
[0703] The "emotion recognition means" is a system that analyzes emotions in real time from the user's facial expressions, voice, and text input, and feeds that data back to the artificial intelligence used for generation.
[0704] "Encoding" is the process of converting extracted editorial elements into a format that can be understood by a generative artificial intelligence.
[0705] A "correction instruction" is a request for specific changes or adjustments that the user makes to the generated video.
[0706] "Re-editing" is the process in which the generating AI re-edits the video based on the user's correction instructions.
[0707] This invention is a system that automatically generates high-quality videos based on viewing data collected from video provision services. This system consists of a server, a terminal, artificial intelligence for generation, and emotion recognition means.
[0708] The server collects viewing data from video-providing services (e.g., video streaming platforms) using APIs. The collected data includes the number of views, ratings, comments, tags, and descriptions. The server then analyzes this data using machine learning algorithms to extract specific patterns and trends. For example, it can identify common editing elements (such as background music selection, subtitle style, and frequency of cuts) among videos with high views. The results of this analysis are coded as extracted editing elements and converted into a format understandable by the generative AI.
[0709] Users use their devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends these inputs to the server, which converts them into appropriate editing instructions and sends them to the AI generator.
[0710] The artificial intelligence (AI) for generation automatically generates videos based on editing instructions sent from the server. This AI generates videos by applying the specified composition and production elements using video editing software such as Adobe Premiere Pro or FFmpeg. The generated videos are provided to users, who can play and check them on their devices.
[0711] The emotion recognition means analyzes the emotions of the user in real time while watching the video or inputting editing instructions. Emotional data is acquired from the user's facial expressions, voice, and text input, and sent to the server. This emotional data is fed back to the generating AI and used to emphasize or modify specific parts of the video. For example, editing instructions are issued to emphasize parts where the user smiles more and modify parts where the user frowns.
[0712] As an example, consider the following prompt:
[0713] The video is 3 minutes long
[0714] "The first 30 seconds are the intro"
[0715] "Show subtitles every minute"
[0716] "Use upbeat background music"
[0717] This allows us to automatically generate videos that meet the user's preferences, and further improve the user experience by utilizing emotion recognition.
[0718] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0719] Step 1:
[0720] The server collects viewing data from the video provider. Specifically, it periodically obtains metadata such as the number of views, ratings, comments, tags, and descriptions using an API. It sends an API request as input and obtains viewing data as output. This data is stored in the server's database.
[0721] Step 2:
[0722] The server analyzes the collected viewing data. It uses machine learning algorithms to analyze the data and extract certain patterns and trends, such as popular background music, subtitle styles, or frequency of cuts. It uses the viewing data collected in step 1 as input and obtains the analysis results as output. These analysis results are stored on the server for use in later steps.
[0723] Step 3:
[0724] The server encodes the analysis results and converts the extracted editing elements into a format that the AI can understand. Specifically, it converts the data into JSON format and prepares it for passing to the AI. It uses the analysis results as input and generates coded editing elements as output.
[0725] Step 4:
[0726] Users input video composition and editing instructions in text format using their own devices. Examples of input include "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device collects the user's instructions as input and sends them to the server.
[0727] Step 5:
[0728] The server receives instructions from the user. After receiving these instructions, it analyzes them, converts them into appropriate editing instructions, and sends them to the AI generator. The input is the user's instructions, and the output is instructions converted into a format that can be sent to the AI generator.
[0729] Step 6:
[0730] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server. It uses video editing software (e.g., Adobe Premiere Pro or FFmpeg) to edit the videos as instructed. It receives the editing instructions as input and provides the generated video as output to the user.
[0731] Step 7:
[0732] The user checks the generated video and inputs any necessary corrections through the device. For example, this includes instructions such as "change the display position of the subtitles to the bottom right" or "reduce the background music volume by 20%." The device collects the corrections as input and sends them back to the server.
[0733] Step 8:
[0734] The server receives the correction instructions and sends them to the AI for generation. The received instructions are analyzed and sent back to the AI as appropriate re-editing instructions. The input is the user's correction instructions, and the output is the re-editing instructions.
[0735] Step 9:
[0736] The generative artificial intelligence (AI) re-edits the video based on the correction instructions. The corrections are reflected again using video editing software to generate the final video. The correction instructions are received as input, and the corrected video is provided to the user as output.
[0737] Step 10:
[0738] The emotion recognition means analyzes the user's facial expressions, voice, and text input to obtain emotion data. It performs the analysis in real time while the user is watching the video and sends the results to the server. It collects the user's real-time data as input and sends the analyzed emotion data as output to the server.
[0739] Step 11:
[0740] The server sends additional editing instructions to the AI based on the emotion data. For example, it may emphasize parts that show positive emotions and modify or delete parts that show negative emotions. The server receives emotion data as input and sends additional editing instructions to the AI as output.
[0741] (Application example 2)
[0742] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0743] In modern virtual stores, it is important to generate efficient, high-quality promotional videos to attract customer attention. However, existing methods require a significant amount of manual effort and time for video editing and generation, and it is difficult to provide optimal promotional videos that reflect customers' real-time emotions. Furthermore, it is difficult to immediately reflect users' feedback while they are watching, making it difficult to improve user satisfaction.
[0744] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting metadata of popular videos from a video platform, means for analyzing the collected metadata and extracting editing elements for the video, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, emotion recognition means for collecting real-time emotional data from users, means for adjusting the editing content of the video using the collected emotional data, and means for viewing and editing the video using a device such as smart glasses or a head-mounted display as a user terminal. This makes it possible to automatically generate optimal promotional videos that reflect the user's emotions and effectively improve user satisfaction and interest in the virtual store.
[0745] A "video platform" is a general term for services that allow various users to upload, watch, and share videos online.
[0746] "Metadata" refers to data related to a video, including information such as the number of views, ratings, comments, tags, and descriptions.
[0747] "Generative AI" refers to an AI system designed to automatically generate videos based on collected data and user instructions.
[0748] "Emotion recognition means" refers to technology or equipment that analyzes a user's facial expressions, voice, etc. to detect emotions in real time.
[0749] The term "user terminal" refers to a device used by a user, and in the present invention particularly includes devices such as smart glasses and head-mounted displays.
[0750] "Emotion Data" refers to information collected by an emotion recognition means about a user's real-time emotional state.
[0751] "Editing elements" are elements that indicate optimal patterns and techniques for creating and editing videos, and include specific elements such as background music, subtitle style, and frequency of cuts.
[0752] "Promotional videos" refer to videos created for the purpose of promoting products or services within a virtual store.
[0753] A "virtual store" refers to a virtual store that exists on the Internet or in a virtual reality space, providing an environment in which users can visually view products and services.
[0754] The present invention relates to a system for automatically generating promotional videos in a virtual store and optimizing the videos based on real-time visitor sentiment. Specific embodiments are described below.
[0755] overview
[0756] The system consists of the following main components:
[0757] server
[0758] Devices (smart glasses, head-mounted displays, etc.)
[0759] Generative artificial intelligence (AI)
[0760] Emotion Recognition Engine
[0761] Hardware and software used
[0762] 1. Server:
[0763] The server uses an API from a video platform (e.g., YouTube) to collect metadata for popular videos, obtaining information such as the number of views, ratings, comments, tags, and descriptions.
[0764] Based on the collected metadata, machine learning libraries (e.g., scikit-learn, TensorFlow) are used to analyze specific patterns and trends and extract editorial elements.
[0765] The extracted editing elements are coded and converted into a format that can be understood by the generative artificial intelligence.
[0766] 2. User Device:
[0767] The user inputs the video composition and editing instructions in text format through a terminal (smart glasses or a head-mounted display).
[0768] Real-time emotional data is collected while the user is watching through the camera and microphone built into the device.
[0769] 3. Generative Artificial Intelligence (AI):
[0770] The artificial intelligence for generation automatically generates videos based on editing instructions and emotional data sent from the server.
[0771] The generated video is provided to the user once, and modifications can be made as needed.
[0772] 4. Emotion Recognition Engine:
[0773] The emotion recognition engine analyzes the user's facial expressions and voice in real time to collect emotional data.
[0774] The collected emotional data is sent to a server and used by the generative AI to adjust the editing content of the video.
[0775] Program Processing Overview
[0776] Data collection and analysis:
[0777] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[0778] Enter user instructions:
[0779] The user inputs information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used on the device.
[0780] Emotion data collection:
[0781] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[0782] Video Generation:
[0783] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[0784] Tweaks and Emotions:
[0785] The user checks the generated video and makes any necessary corrections. The emotion engine also analyzes the user's emotions during the corrections, and the server re-edits the video based on those. For example, the user can input correction instructions such as "change the display position of the subtitles" or "adjust the volume of the background music."
[0786] Specific examples
[0787] Prompt Sentence Examples
[0788] API endpoint: https: / / api.videoplatform.com / metadata
[0789] Parameters: Over 1,000,000 views
[0790] User instructions:
[0791] Length: 3 minutes
[0792] Intro length: 30 seconds
[0793] Subtitles: Every minute
[0794] BGM: Upbeat
[0795] Emotional Data:
[0796] Video Frame: {...}
[0797] Analysis results: Smiling a lot indicates positive emotions, frowning a lot indicates negative emotions
[0798] In this way, the present invention can improve the user experience in a virtual store by automatically generating an optimal promotional video that reflects the user's emotions.
[0799] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0800] Step 1:
[0801] The server collects metadata of popular videos from a video platform using an API. The input is the API endpoint and parameters of the video platform, and the output is the collected metadata (number of views, number of ratings, number of comments, etc.).
[0802] Step 2:
[0803] The server analyzes the collected metadata and extracts the editing elements of the video. The input is the metadata, and the output is the extracted editing elements (e.g., type of background music, subtitle style, frequency of cuts, etc.). Machine learning libraries (e.g., scikit-learn, TensorFlow) are used for the analysis.
[0804] Step 3:
[0805] The server encodes the extracted editing elements and converts them into a format that can be understood by the AI for generation. The input is the editing elements, and the output is coded editing instructions. The specific operation is to map the editing elements to the corresponding code.
[0806] Step 4:
[0807] The user inputs video composition and editing instructions in text format through a device (smart glasses or a head-mounted display). The input is text format instructions, and the output is instruction data sent to the server.
[0808] Step 5:
[0809] The server sends data to the generating AI based on the composition and editing instructions received from the user, where the input is the user's instruction data and coded editing instructions, and the output is the data sent to the generating AI.
[0810] Step 6:
[0811] The artificial intelligence for generation automatically generates videos based on the received composition and editing instructions. The input is data sent from the server, and the output is the generated video file. Specific operations involve the use of video editing software and libraries.
[0812] Step 7:
[0813] While a user is watching a video generated on a device, the emotion recognition means analyzes the user's facial expressions and voice to collect emotion data in real time. The input is the user's facial expressions and voice, and the output is the collected emotion data.
[0814] Step 8:
[0815] The server analyzes the collected emotion data and adjusts the editing content of the video. The input is emotion data, and the output is adjusted editing instructions. Specifically, editing is performed to emphasize positive emotions.
[0816] Step 9:
[0817] The user inputs correction instructions for the video. The input is the user's correction instructions, and the output is correction instruction data sent to the server.
[0818] Step 10:
[0819] The server sends instructions for re-editing to the AI based on the correction instructions and the collected emotion data. The input is the correction instruction data and emotion data, and the output is the data for re-editing.
[0820] Step 11:
[0821] The artificial intelligence for generation re-edits the video based on the re-editing instructions. The input is the data for re-editing, and the output is the corrected video file. Specifically, it edits the specified parts.
[0822] This series of steps enables the system to provide optimal promotional videos within the virtual store that reflect the user's emotions and feedback.
[0823] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0824] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0825] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0826] [Third embodiment]
[0827] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0828] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0829] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0830] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0831] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0832] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0833] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0834] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0835] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0836] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0837] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0838] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0839] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[0840] System Overview
[0841] The system consists of the following main components:
[0842] server
[0843] Terminal
[0844] Generative artificial intelligence (AI)
[0845] System program and processing flow
[0846] The system's program is designed to operate in cooperation with the server, terminals, and artificial intelligence for generation. The roles and specific operations of each entity are explained below.
[0847] server
[0848] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[0849] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[0850] Terminal
[0851] Users use their own devices to input the structure and editing instructions for the video they want to create in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0852] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[0853] Generative artificial intelligence (AI)
[0854] The artificial intelligence (AI) for generation is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video.
[0855] The generated video is provided to the user, who can then make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edits are then sent back to the server, where the AI used for generation re-edits the video. Once edits are complete, the video is provided to the user again.
[0856] Specific examples
[0857] Below are some specific usage examples.
[0858] 1. Data Collection and Analysis
[0859] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[0860] 2. Entering user instructions
[0861] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[0862] 3. Video Generation
[0863] The server receives the user's instructions and sends them to the AI for generation, which then edits the video based on the received instructions and provides it to the user.
[0864] 4. Minor corrections
[0865] The user checks the generated video and makes any necessary corrections, such as "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends the correction instructions back to the AI, which then re-edits the video.
[0866] In this way, the present invention enables users to efficiently generate high-quality videos, realizing video production that quickly responds to trends.
[0867] The processing flow will be explained below.
[0868] Step 1:
[0869] The server uses the video platform's API to collect metadata for popular videos, including the number of views, ratings, comments, tags, and descriptions.
[0870] Step 2:
[0871] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements for the video, such as background music selection, subtitle display format, and cut timing.
[0872] Step 3:
[0873] The server encodes the extracted editing elements and converts them into an instruction format that can be understood by a generative artificial intelligence (AI), which includes specific editing operations and parameters.
[0874] Step 4:
[0875] Users use their devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0876] Step 5:
[0877] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[0878] Step 6:
[0879] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[0880] Step 7:
[0881] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[0882] Step 8:
[0883] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[0884] Step 9:
[0885] The user can check the generated video and input any necessary corrections into the device, such as "move the subtitles a little higher" or "lower the volume of the background music."
[0886] Step 10:
[0887] The terminal sends the user's correction instructions to the server, which include specific changes.
[0888] Step 11:
[0889] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[0890] Step 12:
[0891] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[0892] Step 13:
[0893] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[0894] This series of steps allows users to efficiently generate high-quality videos and keep up with trends.
[0895] Example 1
[0896] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0897] In recent years, creating popular videos on video platforms requires advanced editing skills and a large amount of time. This makes it difficult for ordinary users to effectively create high-quality videos. Furthermore, due to a lack of video editing skills, it is difficult to create videos that attract viewers' attention. To solve this problem, a system that allows users to easily create high-quality videos is needed.
[0898] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0899] In this invention, the server includes means for collecting information on highly rated videos from a video sharing service, means for analyzing the collected information to extract editing elements for the videos, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, means for receiving video composition and editing instructions from a user, generating artificial intelligence means for automatically generating videos based on the received composition and editing instructions, means for a user to input correction instructions for the generated videos, and means for re-editing the generated videos based on the correction instructions. This enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[0900] A "video sharing service" is a platform on the Internet that allows users to upload, watch, and share videos.
[0901] A "highly rated video" is a video that has received many views, high ratings, and positive comments on a video sharing service.
[0902] "Information" refers to data such as video metadata, number of views, number of ratings, number of comments, tags, and description.
[0903] "Analysis" is the process of assigning meaning to collected information and recognizing patterns according to a specific purpose.
[0904] "Editing elements" are specific components related to video editing, such as background music selection, subtitle style, and frequency of cuts.
[0905] "Extraction" is the act of extracting specific editing elements from the analyzed information.
[0906] "Encoding" is the process of converting the extracted editing elements into a format that can be understood by the generating artificial intelligence.
[0907] "Generative AI" is an AI system that automatically generates videos based on specified composition and editing elements.
[0908] "User" refers to the general user who uses the system to generate and modify videos.
[0909] "Composition" refers to the overall components of a video, such as the length of the video, the length of the intro, the timing of subtitles, and the type of background music used.
[0910] "Input" is the act of transmitting instructions or corrections from the user to the system.
[0911] "Re-editing" refers to additional editing work performed on an already generated video based on correction instructions from the user.
[0912] This invention is a system that automatically collects and analyzes the editing elements of popular videos on video sharing services and generates high-quality videos based on the results. The system includes the following main components: a server, a terminal, and artificial intelligence (AI) for generation.
[0913] server
[0914] The server collects metadata of highly rated videos from a video sharing service (e.g., YouTube). Specifically, the server periodically obtains the number of views, ratings, comments, tags, descriptions, etc. of each video using an API, and stores this information in a database. The collected metadata is then analyzed using a machine learning algorithm (e.g., scikit-learn) to extract common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos. The extracted editing elements are encoded into a format understandable by the generative artificial intelligence. This encoded data is then used as editing instructions for the videos that are generated later.
[0915] Terminal
[0916] The user uses a device (e.g., a PC or smartphone) to input video composition and editing instructions in text format. Specific examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends the user's input to the server, which then sends editing instructions to the artificial intelligence used for generation based on these instructions.
[0917] Generative artificial intelligence (AI)
[0918] Generative artificial intelligence (AI) plays a role in automatically generating videos according to editing instructions sent from the server. For example, the AI uses video editing software (e.g., Adobe Premiere Pro API) or a video library (e.g., FFmpeg) to apply the specified composition and production elements to generate a video. The generated video is then provided to the user via the server, who can review the video on their device and make minor corrections as needed. For example, corrections could include "changing the position of the subtitles" or "adjusting the volume of the background music." The server then sends these corrections back to the generative AI, which then re-edits the video. The final, edited video is then provided to the user via the server again.
[0919] Specific examples
[0920] Below are some specific usage examples.
[0921] 1. Data Collection and Analysis
[0922] The server collects metadata from video sharing services and analyzes and extracts editing elements of popular videos, such as background music and subtitle styles common to videos with over 1,000,000 views.
[0923] 2. Entering user instructions
[0924] The user enters information in text format on the device, such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, the type of background music to be used, etc. An example of a specific prompt is, "The video is 3 minutes long, and the intro will be displayed for the first 30 seconds. Subtitles will be displayed every minute, and upbeat background music will be used."
[0925] 3. Video Generation
[0926] The server receives the user's instructions and sends them to the AI generator, which then edits the video based on the instructions and provides it to the user via the server.
[0927] 4. Minor corrections
[0928] The user checks the generated video and instructs the AI to make any necessary corrections. For example, the server may request that the user increase the font size of the subtitles or decrease the volume of the background music. The server then sends the correction instructions back to the AI, which then re-edits the video.
[0929] In this way, the present invention enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[0930] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0931] System processing flow
[0932] Step 1:
[0933] The server collects metadata of highly rated videos from the video sharing service. Specifically, it uses the API to obtain information such as the number of views, ratings, comments, tags, and descriptions of videos in JSON format. To obtain this data, it uses a Python library (e.g., requests). The API endpoint and authentication information are required as input, and the collected metadata is obtained as output.
[0934] Step 2:
[0935] The server analyzes the collected metadata and uses machine learning algorithms (e.g., scikit-learn) to extract editing elements. Specifically, it applies clustering and classification algorithms to recognize the characteristics of popular videos. The collected metadata is required as input, and the output is the editing elements common to popular videos (e.g., background music selection, subtitle style, frequency of cuts, etc.).
[0936] Step 3:
[0937] The server encodes the extracted edit elements into a format that can be understood by the generative artificial intelligence (AI). This encoding process converts the edit elements into XML and then back into JSON format. The input is the edit elements obtained as a result of the analysis, and the output is the encoded data.
[0938] Step 4:
[0939] The device sends the video composition and editing instructions entered by the user (e.g., "video length is 3 minutes," "first 30 seconds is intro," "display subtitles every minute," "use upbeat background music") to the server. The input requires the text instructions entered by the user, and the output is the instruction data to be sent to the server.
[0940] Step 5:
[0941] The server analyzes the editing instructions received from the user, converts them into an appropriate format, and sends them to the AI for generation. The input requires instruction data from the user, and the output is the editing instructions sent to the AI for generation.
[0942] Step 6:
[0943] Generative artificial intelligence (AI) automatically generates videos based on the editing instructions it receives. This process uses video editing software (e.g., Adobe Premiere Pro API) and video libraries (e.g., FFmpeg). The input is the editing instructions sent from the server, and the output is the generated video file.
[0944] Step 7:
[0945] The server sends the generated video to the user's device. The input requires a video file provided by the AI for generation, and the output is the video data sent to the user's device.
[0946] Step 8:
[0947] The user checks the generated video and inputs correction instructions as necessary. For example, specific corrections can be input, such as "changing the display position of subtitles" or "adjusting the volume of background music." The input requires the generated video and correction instructions, and the output is correction instruction data.
[0948] Step 9:
[0949] The terminal transmits the correction instructions input by the user to the server. The input requires correction instruction data from the user, and the output obtains the correction instruction data to be transmitted to the server.
[0950] Step 10:
[0951] The server then sends the received correction instructions back to the AI generator, which then re-edits the video. The input requires correction instructions from the user, and the output is a re-edited video file.
[0952] Step 11:
[0953] The server then sends the re-edited video back to the user's device. The input is the re-edited video file, and the output is the final video data sent to the user's device.
[0954] (Application example 1)
[0955] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0956] Modern content distribution services require the rapid generation of high-quality videos that capture viewers' attention. However, video production is time-consuming and labor-intensive, and incorporating popular editing elements is particularly difficult. Furthermore, there are few ways for users to reflect their desired composition and editing, and the process of making corrections or re-editing generated videos is cumbersome. A method that solves these issues and enables efficient, high-quality video generation is needed.
[0957] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0958] In this invention, the server includes a means for collecting metadata of popular videos from a video platform, a means for analyzing the collected metadata to extract editing elements for the video, and a means for coding the extracted editing elements and providing them in the form of instructions to a generating AI. This allows users to input desired video composition and editing instructions using their own devices, and the generating AI automatically generates videos based on those instructions. The system also includes a means for linking multiple video clips and a means for users to input correction instructions for generated videos via a smart device and re-edit them. This allows users to efficiently generate high-quality videos and flexibly correct and re-edit them.
[0959] A "video platform" is an internet-based service that hosts video content and enables users to upload, view, and share it.
[0960] "Metadata" is data that includes information about the attributes and structure of the data, and specifically refers to the number of times a video has been played, the number of ratings, the number of comments, tags, descriptions, etc.
[0961] "Editing elements" refer to specific elements and techniques used in the production and editing of a video, including background music selection, subtitle style, and frequency of cuts.
[0962] "Generative AI" is an AI system that has the ability to automatically generate videos based on specified data using pre-trained models and algorithms.
[0963] A "smart device" is an electronic device such as a mobile phone, tablet, or smart glasses that has internet connectivity and allows users to input and operate information.
[0964] A "prompt sentence" refers to a text-based input sentence used to input specific instructions and conditions to a generative AI model, and includes specific instructions necessary for generating and editing videos.
[0965] "Viewing data" refers to data relating to the behavior and evaluation of a user when watching a video, and includes the number of views, viewing time, ratings, comments, and the like.
[0966] "Modification instructions" refer to instructions input in text format that the user wishes to make changes or adjustments to the generated video.
[0967] A "generative model" refers to an algorithm or mathematical model that has the ability to generate various data based on specific conditions or instructions.
[0968] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[0969] System Overview
[0970] The system consists of the following main components:
[0971] server
[0972] Terminal
[0973] Generative artificial intelligence (AI)
[0974] server
[0975] The server is responsible for the following functions:
[0976] 1. Metadata collection
[0977] Using APIs from video platforms, metadata (number of views, number of ratings, number of comments, tags, descriptions, etc.) of popular videos is periodically collected.
[0978] 2. Extracting editing elements
[0979] The collected metadata is analyzed and common editing elements (such as background music selection, subtitle style, and frequency of cuts) among popular videos are extracted using a machine learning algorithm.
[0980] 3. Coding and Data Conversion
[0981] The extracted editing elements are coded and converted into a format that can be understood by a generative artificial intelligence (AI). This coded data is then used as editing instructions for the generated video.
[0982] Terminal
[0983] The user uses a device (smartphone or PC) to perform the following operations:
[0984] 1. Enter editing instructions
[0985] You can enter detailed instructions for video composition and editing in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[0986] 2. Enter correction instructions
[0987] Check the generated video and enter corrections as necessary (e.g., change the position of subtitles, adjust the volume of background music).
[0988] Generative artificial intelligence (AI)
[0989] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server:
[0990] 1. Automatic video generation
[0991] Use standard video editing software or libraries (e.g., MoviePy) to generate a video by applying the specified structure and production elements.
[0992] 2. Video Linking
[0993] Combine multiple video clips and apply any effects or edits you want.
[0994] 3. Updates
[0995] The system receives correction instructions from the user and re-edits the video. Through this process, the final video is generated according to the user's wishes.
[0996] Specific examples
[0997] Data collection and analysis
[0998] 1. The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos. For example, it can extract the background music and subtitle styles common to videos with over 1,000,000 views.
[0999] Entering user instructions
[1000] 2. The user enters information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on the device.
[1001] example:
[1002] The video is 3 minutes long
[1003] The first 30 seconds is an intro
[1004] Display subtitles every minute
[1005] Use upbeat background music
[1006] Video Generation
[1007] 3. The server receives the user's instructions and sends them to the AI for generation. The AI then edits the video based on the received instructions and provides it to the user.
[1008] Minor correction
[1009] 4. The user reviews the generated video and makes any necessary corrections. For example, "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends these corrections to the AI, which then re-edits the video.
[1010] This allows users to quickly generate high-quality videos efficiently, realizing video production that quickly responds to trends.
[1011] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1012] Step 1:
[1013] The server periodically collects metadata of popular videos from the video platform using APIs. Specifically, it obtains data such as the number of views, ratings, comments, tags, and descriptions. This provides detailed information about popular videos as input, and the collected metadata as output.
[1014] Step 2:
[1015] The server extracts editing elements of popular videos by analyzing the collected metadata. Specifically, it uses a machine learning algorithm to analyze data such as the number of views, number of ratings, and tags of each video, and detects common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.). This uses the metadata as input and provides the extracted editing elements as output.
[1016] Step 3:
[1017] The server uses a means to encode the extracted editing elements and convert them into a format understandable by the generating artificial intelligence. Specifically, the server converts the extracted editing elements into a specific format (e.g., JSON format) and provides them to the generating artificial intelligence in the form of instructions. This uses the editing elements as input and generates coded data as output.
[1018] Step 4:
[1019] The user uses their own device to input the video's structure and editing instructions in text format. Examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The user's editing instructions are then input to the device. Text editing instructions are generated as output.
[1020] Step 5:
[1021] The server uses a generating AI to automatically generate a video based on the composition and editing instructions received from the user. Specifically, the server analyzes the editing instructions and creates a video by applying the specified composition and production elements according to the instructions. This uses the user's editing instructions as input and produces a generated video as output.
[1022] Step 6:
[1023] The generative AI uses a means of concatenating multiple video clips, for example using the MoviePy library, and applies a specified order and effects, taking the video clips as input and generating the concatenated video as output.
[1024] Step 7:
[1025] The user checks the generated video and uses a means to input correction instructions. Specifically, while playing the generated video, the user inputs correction instructions in text format, such as "change the position of the subtitles" or "adjust the volume of the background music." This inputs the correction instructions into the terminal, and generates correction instructions as output.
[1026] Step 8:
[1027] The server uses a means to re-edit the generated video based on the correction instructions. Specifically, the server analyzes the correction instructions from the user, and the generating AI re-edits the video to generate a final video that reflects the corrections. This uses the correction instructions as input and generates the final video as output.
[1028] By following these steps, users can efficiently generate high-quality videos and then modify and re-edit them as needed.
[1029] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1030] This invention is a system that automatically extracts and analyzes the editing elements of popular videos on video platforms and automatically generates high-quality videos based on the results. Furthermore, this invention improves the user experience by combining it with an emotion engine that recognizes user emotions.
[1031] System Overview
[1032] The system consists of the following main components:
[1033] server
[1034] Terminal
[1035] Generative artificial intelligence (AI)
[1036] Emotion Engine
[1037] System program and processing flow
[1038] The system's program is designed to have the server, terminal, generative artificial intelligence (AI), and emotion engine work together. The roles and specific operations of each entity are explained below.
[1039] server
[1040] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[1041] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[1042] Terminal
[1043] Users use their own devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[1044] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[1045] Generative artificial intelligence (AI)
[1046] The generative artificial intelligence (AI) is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video. The generated video is then provided to the user, who can make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edit instructions are then sent back to the server, and the generative AI re-edits the video. Once edits are complete, the video is provided to the user again.
[1047] Emotion Engine
[1048] The emotion engine has the ability to analyze emotions from the user's facial expressions, voice, text input, etc. It recognizes emotions in real time while the user is watching a video or entering editing instructions, and analyzes that data.
[1049] The emotion data obtained from the emotion engine is sent to a server. The server uses this data to adjust editing instructions for the artificial intelligence. For example, if a user expresses positive emotion in a particular part of the video, the server can edit it to emphasize that part. If a user expresses negative emotion, the server can issue instructions to modify or delete that element.
[1050] Specific examples
[1051] Below are some specific usage examples.
[1052] 1. Data Collection and Analysis
[1053] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[1054] 2. Entering user instructions
[1055] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[1056] 3. Collecting Emotional Data
[1057] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[1058] 4. Video Generation
[1059] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[1060] 5. Minor edits and emotional reflection
[1061] The user can check the generated video and indicate any necessary corrections. The emotion engine also analyzes the user's emotions during the editing process, and the server readjusts the edited content based on that information.
[1062] Through this series of steps, it becomes possible to efficiently generate high-quality videos that reflect the user's emotions.
[1063] The processing flow will be explained below.
[1064] Step 1:
[1065] The server uses the video platform's API to collect metadata about popular videos, including the number of views, ratings, comments, tags, descriptions, and other data.
[1066] Step 2:
[1067] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements from the video, such as background music selection, subtitle display format, and timing of cuts.
[1068] Step 3:
[1069] The server encodes the extracted editing elements and converts them into instructions that can be understood by the generative artificial intelligence (AI), including specific editing operations and parameters.
[1070] Step 4:
[1071] Users use their own devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[1072] Step 5:
[1073] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[1074] Step 6:
[1075] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[1076] Step 7:
[1077] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[1078] Step 8:
[1079] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[1080] Step 9:
[1081] The emotion engine analyzes the user's facial expressions, voice, text input, etc. in real time while the user is watching a video, and collects emotional data.
[1082] Step 10:
[1083] The emotion engine uses the analyzed emotional data to identify parts of the video where the user had a positive or negative reaction. For example, a smile on the user's face is judged as positive, while a furrowed brow is judged as negative.
[1084] Step 11:
[1085] The server receives the emotion data sent from the emotion engine and reflects it in the artificial intelligence for generation, for example, by issuing instructions to edit parts that show positive emotions or to correct parts that show negative emotions.
[1086] Step 12:
[1087] The user can then review the generated video and make any necessary corrections, such as moving the subtitles a little higher or lowering the volume of the background music, via the device.
[1088] Step 13:
[1089] The terminal sends the user's correction instructions to the server, which include specific changes and desired corrections.
[1090] Step 14:
[1091] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[1092] Step 15:
[1093] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[1094] Step 16:
[1095] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[1096] This series of steps enables efficient generation of high-quality videos that reflect user emotions and realizes video production that is in line with trends.
[1097] Example 2
[1098] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1099] Conventional video generation systems have had difficulty in properly reflecting the editing elements desired by users. Furthermore, they lacked video generation functionality that took user emotions into account, making it impossible to improve the quality of the user experience. This resulted in problems such as a decline in the quality of the generated videos and a decline in user satisfaction.
[1100] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1101] In this invention, the server includes means for collecting viewing data from a video providing service, means for analyzing the collected viewing data and extracting video editing elements, means for encoding the extracted editing elements and providing them to the generation AI in the form of instructions, and emotion recognition means for analyzing the user's emotions and adjusting the editing instructions based on that data, thereby enabling the efficient automatic generation of high-quality videos that reflect the user's emotions.
[1102] "Viewing data" refers to information collected from video provision services, such as the number of views, ratings, comments, tags, and descriptions.
[1103] "Editing elements" are elements that show specific patterns or trends, such as background music selection, subtitle style, and frequency of cuts, extracted by analyzing viewing data.
[1104] "Generative artificial intelligence" refers to algorithms and programs that automatically generate videos based on instruction-style data provided by the server.
[1105] The "emotion recognition means" is a system that analyzes emotions in real time from the user's facial expressions, voice, and text input, and feeds that data back to the artificial intelligence used for generation.
[1106] "Encoding" is the process of converting extracted editorial elements into a format that can be understood by a generative artificial intelligence.
[1107] A "correction instruction" is a request for specific changes or adjustments that the user makes to the generated video.
[1108] "Re-editing" is the process in which the generating AI re-edits the video based on the user's correction instructions.
[1109] This invention is a system that automatically generates high-quality videos based on viewing data collected from video provision services. This system consists of a server, a terminal, artificial intelligence for generation, and emotion recognition means.
[1110] The server collects viewing data from video-providing services (e.g., video streaming platforms) using APIs. The collected data includes the number of views, ratings, comments, tags, and descriptions. The server then analyzes this data using machine learning algorithms to extract specific patterns and trends. For example, it can identify common editing elements (such as background music selection, subtitle style, and frequency of cuts) among videos with high views. The results of this analysis are coded as extracted editing elements and converted into a format understandable by the generative AI.
[1111] Users use their devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends these inputs to the server, which converts them into appropriate editing instructions and sends them to the AI generator.
[1112] The artificial intelligence (AI) for generation automatically generates videos based on editing instructions sent from the server. This AI generates videos by applying the specified composition and production elements using video editing software such as Adobe Premiere Pro or FFmpeg. The generated videos are provided to users, who can play and check them on their devices.
[1113] The emotion recognition means analyzes the emotions of the user in real time while watching the video or inputting editing instructions. Emotional data is acquired from the user's facial expressions, voice, and text input, and sent to the server. This emotional data is fed back to the generating AI and used to emphasize or modify specific parts of the video. For example, editing instructions are issued to emphasize parts where the user smiles more and modify parts where the user frowns.
[1114] As an example, consider the following prompt:
[1115] The video is 3 minutes long
[1116] "The first 30 seconds are the intro"
[1117] "Show subtitles every minute"
[1118] "Use upbeat background music"
[1119] This allows us to automatically generate videos that meet the user's preferences, and further improve the user experience by utilizing emotion recognition.
[1120] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1121] Step 1:
[1122] The server collects viewing data from the video provider. Specifically, it periodically obtains metadata such as the number of views, ratings, comments, tags, and descriptions using an API. It sends an API request as input and obtains viewing data as output. This data is stored in the server's database.
[1123] Step 2:
[1124] The server analyzes the collected viewing data. It uses machine learning algorithms to analyze the data and extract certain patterns and trends, such as popular background music, subtitle styles, or frequency of cuts. It uses the viewing data collected in step 1 as input and obtains the analysis results as output. These analysis results are stored on the server for use in later steps.
[1125] Step 3:
[1126] The server encodes the analysis results and converts the extracted editing elements into a format that the AI can understand. Specifically, it converts the data into JSON format and prepares it for passing to the AI. It uses the analysis results as input and generates coded editing elements as output.
[1127] Step 4:
[1128] Users input video composition and editing instructions in text format using their own devices. Examples of input include "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device collects the user's instructions as input and sends them to the server.
[1129] Step 5:
[1130] The server receives instructions from the user. After receiving these instructions, it analyzes them, converts them into appropriate editing instructions, and sends them to the AI generator. The input is the user's instructions, and the output is instructions converted into a format that can be sent to the AI generator.
[1131] Step 6:
[1132] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server. It uses video editing software (e.g., Adobe Premiere Pro or FFmpeg) to edit the videos as instructed. It receives the editing instructions as input and provides the generated video as output to the user.
[1133] Step 7:
[1134] The user checks the generated video and inputs any necessary corrections through the device. For example, this includes instructions such as "change the display position of the subtitles to the bottom right" or "reduce the background music volume by 20%." The device collects the corrections as input and sends them back to the server.
[1135] Step 8:
[1136] The server receives the correction instructions and sends them to the AI for generation. The received instructions are analyzed and sent back to the AI as appropriate re-editing instructions. The input is the user's correction instructions, and the output is the re-editing instructions.
[1137] Step 9:
[1138] The generative artificial intelligence (AI) re-edits the video based on the correction instructions. The corrections are reflected again using video editing software to generate the final video. The correction instructions are received as input, and the corrected video is provided to the user as output.
[1139] Step 10:
[1140] The emotion recognition means analyzes the user's facial expressions, voice, and text input to obtain emotion data. It performs the analysis in real time while the user is watching the video and sends the results to the server. It collects the user's real-time data as input and sends the analyzed emotion data as output to the server.
[1141] Step 11:
[1142] The server sends additional editing instructions to the AI based on the emotion data. For example, it may emphasize parts that show positive emotions and modify or delete parts that show negative emotions. The server receives emotion data as input and sends additional editing instructions to the AI as output.
[1143] (Application example 2)
[1144] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1145] In modern virtual stores, it is important to generate efficient, high-quality promotional videos to attract customer attention. However, existing methods require a significant amount of manual effort and time for video editing and generation, and it is difficult to provide optimal promotional videos that reflect customers' real-time emotions. Furthermore, it is difficult to immediately reflect users' feedback while they are watching, making it difficult to improve user satisfaction.
[1146] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting metadata of popular videos from a video platform, means for analyzing the collected metadata and extracting editing elements for the video, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, emotion recognition means for collecting real-time emotional data from users, means for adjusting the editing content of the video using the collected emotional data, and means for viewing and editing the video using a device such as smart glasses or a head-mounted display as a user terminal. This makes it possible to automatically generate optimal promotional videos that reflect the user's emotions and effectively improve user satisfaction and interest in the virtual store.
[1147] A "video platform" is a general term for services that allow various users to upload, watch, and share videos online.
[1148] "Metadata" refers to data related to a video, including information such as the number of views, ratings, comments, tags, and descriptions.
[1149] "Generative AI" refers to an AI system designed to automatically generate videos based on collected data and user instructions.
[1150] "Emotion recognition means" refers to technology or equipment that analyzes a user's facial expressions, voice, etc. to detect emotions in real time.
[1151] The term "user terminal" refers to a device used by a user, and in the present invention particularly includes devices such as smart glasses and head-mounted displays.
[1152] "Emotion Data" refers to information collected by an emotion recognition means about a user's real-time emotional state.
[1153] "Editing elements" are elements that indicate optimal patterns and techniques for creating and editing videos, and include specific elements such as background music, subtitle style, and frequency of cuts.
[1154] "Promotional videos" refer to videos created for the purpose of promoting products or services within a virtual store.
[1155] A "virtual store" refers to a virtual store that exists on the Internet or in a virtual reality space, providing an environment in which users can visually view products and services.
[1156] The present invention relates to a system for automatically generating promotional videos in a virtual store and optimizing the videos based on real-time visitor sentiment. Specific embodiments are described below.
[1157] overview
[1158] The system consists of the following main components:
[1159] server
[1160] Devices (smart glasses, head-mounted displays, etc.)
[1161] Generative artificial intelligence (AI)
[1162] Emotion Recognition Engine
[1163] Hardware and software used
[1164] 1. Server:
[1165] The server uses an API from a video platform (e.g., YouTube) to collect metadata for popular videos, obtaining information such as the number of views, ratings, comments, tags, and descriptions.
[1166] Based on the collected metadata, machine learning libraries (e.g., scikit-learn, TensorFlow) are used to analyze specific patterns and trends and extract editorial elements.
[1167] The extracted editing elements are coded and converted into a format that can be understood by the generative artificial intelligence.
[1168] 2. User Device:
[1169] The user inputs the video composition and editing instructions in text format through a terminal (smart glasses or a head-mounted display).
[1170] Real-time emotional data is collected while the user is watching through the camera and microphone built into the device.
[1171] 3. Generative Artificial Intelligence (AI):
[1172] The artificial intelligence for generation automatically generates videos based on editing instructions and emotional data sent from the server.
[1173] The generated video is provided to the user once, and modifications can be made as needed.
[1174] 4. Emotion Recognition Engine:
[1175] The emotion recognition engine analyzes the user's facial expressions and voice in real time to collect emotional data.
[1176] The collected emotional data is sent to a server and used by the generative AI to adjust the editing content of the video.
[1177] Program Processing Overview
[1178] Data collection and analysis:
[1179] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[1180] Enter user instructions:
[1181] The user inputs information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used on the device.
[1182] Emotion data collection:
[1183] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[1184] Video Generation:
[1185] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[1186] Tweaks and Emotions:
[1187] The user checks the generated video and makes any necessary corrections. The emotion engine also analyzes the user's emotions during the corrections, and the server re-edits the video based on those. For example, the user can input correction instructions such as "change the display position of the subtitles" or "adjust the volume of the background music."
[1188] Specific examples
[1189] Prompt Sentence Examples
[1190] API endpoint: https: / / api.videoplatform.com / metadata
[1191] Parameters: Over 1,000,000 views
[1192] User instructions:
[1193] Length: 3 minutes
[1194] Intro length: 30 seconds
[1195] Subtitles: Every minute
[1196] BGM: Upbeat
[1197] Emotional Data:
[1198] Video Frame: {...}
[1199] Analysis results: Smiling a lot indicates positive emotions, frowning a lot indicates negative emotions
[1200] In this way, the present invention can improve the user experience in a virtual store by automatically generating an optimal promotional video that reflects the user's emotions.
[1201] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1202] Step 1:
[1203] The server collects metadata of popular videos from a video platform using an API. The input is the API endpoint and parameters of the video platform, and the output is the collected metadata (number of views, number of ratings, number of comments, etc.).
[1204] Step 2:
[1205] The server analyzes the collected metadata and extracts the editing elements of the video. The input is the metadata, and the output is the extracted editing elements (e.g., type of background music, subtitle style, frequency of cuts, etc.). Machine learning libraries (e.g., scikit-learn, TensorFlow) are used for the analysis.
[1206] Step 3:
[1207] The server encodes the extracted editing elements and converts them into a format that can be understood by the AI for generation. The input is the editing elements, and the output is coded editing instructions. The specific operation is to map the editing elements to the corresponding code.
[1208] Step 4:
[1209] The user inputs video composition and editing instructions in text format through a device (smart glasses or a head-mounted display). The input is text format instructions, and the output is instruction data sent to the server.
[1210] Step 5:
[1211] The server sends data to the generating AI based on the composition and editing instructions received from the user, where the input is the user's instruction data and coded editing instructions, and the output is the data sent to the generating AI.
[1212] Step 6:
[1213] The artificial intelligence for generation automatically generates videos based on the received composition and editing instructions. The input is data sent from the server, and the output is the generated video file. Specific operations involve the use of video editing software and libraries.
[1214] Step 7:
[1215] While a user is watching a video generated on a device, the emotion recognition means analyzes the user's facial expressions and voice to collect emotion data in real time. The input is the user's facial expressions and voice, and the output is the collected emotion data.
[1216] Step 8:
[1217] The server analyzes the collected emotion data and adjusts the editing content of the video. The input is emotion data, and the output is adjusted editing instructions. Specifically, editing is performed to emphasize positive emotions.
[1218] Step 9:
[1219] The user inputs correction instructions for the video. The input is the user's correction instructions, and the output is correction instruction data sent to the server.
[1220] Step 10:
[1221] The server sends instructions for re-editing to the AI based on the correction instructions and the collected emotion data. The input is the correction instruction data and emotion data, and the output is the data for re-editing.
[1222] Step 11:
[1223] The artificial intelligence for generation re-edits the video based on the re-editing instructions. The input is the data for re-editing, and the output is the corrected video file. Specifically, it edits the specified parts.
[1224] This series of steps enables the system to provide optimal promotional videos within the virtual store that reflect the user's emotions and feedback.
[1225] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1226] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1227] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1228] [Fourth embodiment]
[1229] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1230] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1231] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1232] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1233] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1234] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1235] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1236] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1237] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1238] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1239] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1240] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1241] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1242] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[1243] System Overview
[1244] The system consists of the following main components:
[1245] server
[1246] Terminal
[1247] Generative artificial intelligence (AI)
[1248] System program and processing flow
[1249] The system's program is designed to operate in cooperation with the server, terminals, and artificial intelligence for generation. The roles and specific operations of each entity are explained below.
[1250] server
[1251] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[1252] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[1253] Terminal
[1254] Users use their own devices to input the structure and editing instructions for the video they want to create in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[1255] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[1256] Generative artificial intelligence (AI)
[1257] The artificial intelligence (AI) for generation is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video.
[1258] The generated video is provided to the user, who can then make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edits are then sent back to the server, where the AI used for generation re-edits the video. Once edits are complete, the video is provided to the user again.
[1259] Specific examples
[1260] Below are some specific usage examples.
[1261] 1. Data Collection and Analysis
[1262] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[1263] 2. Entering user instructions
[1264] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[1265] 3. Video Generation
[1266] The server receives the user's instructions and sends them to the AI for generation, which then edits the video based on the received instructions and provides it to the user.
[1267] 4. Minor corrections
[1268] The user checks the generated video and makes any necessary corrections, such as "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends the correction instructions back to the AI, which then re-edits the video.
[1269] In this way, the present invention enables users to efficiently generate high-quality videos, realizing video production that quickly responds to trends.
[1270] The processing flow will be explained below.
[1271] Step 1:
[1272] The server uses the video platform's API to collect metadata for popular videos, including the number of views, ratings, comments, tags, and descriptions.
[1273] Step 2:
[1274] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements for the video, such as background music selection, subtitle display format, and cut timing.
[1275] Step 3:
[1276] The server encodes the extracted editing elements and converts them into an instruction format that can be understood by a generative artificial intelligence (AI), which includes specific editing operations and parameters.
[1277] Step 4:
[1278] Users use their devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[1279] Step 5:
[1280] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[1281] Step 6:
[1282] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[1283] Step 7:
[1284] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[1285] Step 8:
[1286] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[1287] Step 9:
[1288] The user can check the generated video and input any necessary corrections into the device, such as "move the subtitles a little higher" or "lower the volume of the background music."
[1289] Step 10:
[1290] The terminal sends the user's correction instructions to the server, which include specific changes.
[1291] Step 11:
[1292] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[1293] Step 12:
[1294] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[1295] Step 13:
[1296] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[1297] This series of steps allows users to efficiently generate high-quality videos and keep up with trends.
[1298] Example 1
[1299] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1300] In recent years, creating popular videos on video platforms requires advanced editing skills and a large amount of time. This makes it difficult for ordinary users to effectively create high-quality videos. Furthermore, due to a lack of video editing skills, it is difficult to create videos that attract viewers' attention. To solve this problem, a system that allows users to easily create high-quality videos is needed.
[1301] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1302] In this invention, the server includes means for collecting information on highly rated videos from a video sharing service, means for analyzing the collected information to extract editing elements for the videos, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, means for receiving video composition and editing instructions from a user, generating artificial intelligence means for automatically generating videos based on the received composition and editing instructions, means for a user to input correction instructions for the generated videos, and means for re-editing the generated videos based on the correction instructions. This enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[1303] A "video sharing service" is a platform on the Internet that allows users to upload, watch, and share videos.
[1304] A "highly rated video" is a video that has received many views, high ratings, and positive comments on a video sharing service.
[1305] "Information" refers to data such as video metadata, number of views, number of ratings, number of comments, tags, and description.
[1306] "Analysis" is the process of assigning meaning to collected information and recognizing patterns according to a specific purpose.
[1307] "Editing elements" are specific components related to video editing, such as background music selection, subtitle style, and frequency of cuts.
[1308] "Extraction" is the act of extracting specific editing elements from the analyzed information.
[1309] "Encoding" is the process of converting the extracted editing elements into a format that can be understood by the generating artificial intelligence.
[1310] "Generative AI" is an AI system that automatically generates videos based on specified composition and editing elements.
[1311] "User" refers to the general user who uses the system to generate and modify videos.
[1312] "Composition" refers to the overall components of a video, such as the length of the video, the length of the intro, the timing of subtitles, and the type of background music used.
[1313] "Input" is the act of transmitting instructions or corrections from the user to the system.
[1314] "Re-editing" refers to additional editing work performed on an already generated video based on correction instructions from the user.
[1315] This invention is a system that automatically collects and analyzes the editing elements of popular videos on video sharing services and generates high-quality videos based on the results. The system includes the following main components: a server, a terminal, and artificial intelligence (AI) for generation.
[1316] server
[1317] The server collects metadata of highly rated videos from a video sharing service (e.g., YouTube). Specifically, the server periodically obtains the number of views, ratings, comments, tags, descriptions, etc. of each video using an API, and stores this information in a database. The collected metadata is then analyzed using a machine learning algorithm (e.g., scikit-learn) to extract common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos. The extracted editing elements are encoded into a format understandable by the generative artificial intelligence. This encoded data is then used as editing instructions for the videos that are generated later.
[1318] Terminal
[1319] The user uses a device (e.g., a PC or smartphone) to input video composition and editing instructions in text format. Specific examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends the user's input to the server, which then sends editing instructions to the artificial intelligence used for generation based on these instructions.
[1320] Generative artificial intelligence (AI)
[1321] Generative artificial intelligence (AI) plays a role in automatically generating videos according to editing instructions sent from the server. For example, the AI uses video editing software (e.g., Adobe Premiere Pro API) or a video library (e.g., FFmpeg) to apply the specified composition and production elements to generate a video. The generated video is then provided to the user via the server, who can review the video on their device and make minor corrections as needed. For example, corrections could include "changing the position of the subtitles" or "adjusting the volume of the background music." The server then sends these corrections back to the generative AI, which then re-edits the video. The final, edited video is then provided to the user via the server again.
[1322] Specific examples
[1323] Below are some specific usage examples.
[1324] 1. Data Collection and Analysis
[1325] The server collects metadata from video sharing services and analyzes and extracts editing elements of popular videos, such as background music and subtitle styles common to videos with over 1,000,000 views.
[1326] 2. Entering user instructions
[1327] The user enters information in text format on the device, such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, the type of background music to be used, etc. An example of a specific prompt is, "The video is 3 minutes long, and the intro will be displayed for the first 30 seconds. Subtitles will be displayed every minute, and upbeat background music will be used."
[1328] 3. Video Generation
[1329] The server receives the user's instructions and sends them to the AI generator, which then edits the video based on the instructions and provides it to the user via the server.
[1330] 4. Minor corrections
[1331] The user checks the generated video and instructs the AI to make any necessary corrections. For example, the server may request that the user increase the font size of the subtitles or decrease the volume of the background music. The server then sends the correction instructions back to the AI, which then re-edits the video.
[1332] In this way, the present invention enables users to efficiently generate and edit high-quality videos without requiring advanced editing skills.
[1333] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1334] System processing flow
[1335] Step 1:
[1336] The server collects metadata of highly rated videos from the video sharing service. Specifically, it uses the API to obtain information such as the number of views, ratings, comments, tags, and descriptions of videos in JSON format. To obtain this data, it uses a Python library (e.g., requests). The API endpoint and authentication information are required as input, and the collected metadata is obtained as output.
[1337] Step 2:
[1338] The server analyzes the collected metadata and uses machine learning algorithms (e.g., scikit-learn) to extract editing elements. Specifically, it applies clustering and classification algorithms to recognize the characteristics of popular videos. The collected metadata is required as input, and the output is the editing elements common to popular videos (e.g., background music selection, subtitle style, frequency of cuts, etc.).
[1339] Step 3:
[1340] The server encodes the extracted edit elements into a format that can be understood by the generative artificial intelligence (AI). This encoding process converts the edit elements into XML and then back into JSON format. The input is the edit elements obtained as a result of the analysis, and the output is the encoded data.
[1341] Step 4:
[1342] The device sends the video composition and editing instructions entered by the user (e.g., "video length is 3 minutes," "first 30 seconds is intro," "display subtitles every minute," "use upbeat background music") to the server. The input requires the text instructions entered by the user, and the output is the instruction data to be sent to the server.
[1343] Step 5:
[1344] The server analyzes the editing instructions received from the user, converts them into an appropriate format, and sends them to the AI for generation. The input requires instruction data from the user, and the output is the editing instructions sent to the AI for generation.
[1345] Step 6:
[1346] Generative artificial intelligence (AI) automatically generates videos based on the editing instructions it receives. This process uses video editing software (e.g., Adobe Premiere Pro API) and video libraries (e.g., FFmpeg). The input is the editing instructions sent from the server, and the output is the generated video file.
[1347] Step 7:
[1348] The server sends the generated video to the user's device. The input requires a video file provided by the AI for generation, and the output is the video data sent to the user's device.
[1349] Step 8:
[1350] The user checks the generated video and inputs correction instructions as necessary. For example, specific corrections can be input, such as "changing the display position of subtitles" or "adjusting the volume of background music." The input requires the generated video and correction instructions, and the output is correction instruction data.
[1351] Step 9:
[1352] The terminal transmits the correction instructions input by the user to the server. The input requires correction instruction data from the user, and the output obtains the correction instruction data to be transmitted to the server.
[1353] Step 10:
[1354] The server then sends the received correction instructions back to the AI generator, which then re-edits the video. The input requires correction instructions from the user, and the output is a re-edited video file.
[1355] Step 11:
[1356] The server then sends the re-edited video back to the user's device. The input is the re-edited video file, and the output is the final video data sent to the user's device.
[1357] (Application example 1)
[1358] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1359] Modern content distribution services require the rapid generation of high-quality videos that capture viewers' attention. However, video production is time-consuming and labor-intensive, and incorporating popular editing elements is particularly difficult. Furthermore, there are few ways for users to reflect their desired composition and editing, and the process of making corrections or re-editing generated videos is cumbersome. A method that solves these issues and enables efficient, high-quality video generation is needed.
[1360] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1361] In this invention, the server includes a means for collecting metadata of popular videos from a video platform, a means for analyzing the collected metadata to extract editing elements for the video, and a means for coding the extracted editing elements and providing them in the form of instructions to a generating AI. This allows users to input desired video composition and editing instructions using their own devices, and the generating AI automatically generates videos based on those instructions. The system also includes a means for linking multiple video clips and a means for users to input correction instructions for generated videos via a smart device and re-edit them. This allows users to efficiently generate high-quality videos and flexibly correct and re-edit them.
[1362] A "video platform" is an internet-based service that hosts video content and enables users to upload, view, and share it.
[1363] "Metadata" is data that includes information about the attributes and structure of the data, and specifically refers to the number of times a video has been played, the number of ratings, the number of comments, tags, descriptions, etc.
[1364] "Editing elements" refer to specific elements and techniques used in the production and editing of a video, including background music selection, subtitle style, and frequency of cuts.
[1365] "Generative AI" is an AI system that has the ability to automatically generate videos based on specified data using pre-trained models and algorithms.
[1366] A "smart device" is an electronic device such as a mobile phone, tablet, or smart glasses that has internet connectivity and allows users to input and operate information.
[1367] A "prompt sentence" refers to a text-based input sentence used to input specific instructions and conditions to a generative AI model, and includes specific instructions necessary for generating and editing videos.
[1368] "Viewing data" refers to data relating to the behavior and evaluation of a user when watching a video, and includes the number of views, viewing time, ratings, comments, and the like.
[1369] "Modification instructions" refer to instructions input in text format that the user wishes to make changes or adjustments to the generated video.
[1370] A "generative model" refers to an algorithm or mathematical model that has the ability to generate various data based on specific conditions or instructions.
[1371] The present invention is a system that automatically extracts and analyzes the editing elements of popular videos on a video platform and automatically generates high-quality videos based on the results. Specific embodiments for implementing this system are described below.
[1372] System Overview
[1373] The system consists of the following main components:
[1374] server
[1375] Terminal
[1376] Generative artificial intelligence (AI)
[1377] server
[1378] The server is responsible for the following functions:
[1379] 1. Metadata collection
[1380] Using APIs from video platforms, metadata (number of views, number of ratings, number of comments, tags, descriptions, etc.) of popular videos is periodically collected.
[1381] 2. Extracting editing elements
[1382] The collected metadata is analyzed and common editing elements (such as background music selection, subtitle style, and frequency of cuts) among popular videos are extracted using a machine learning algorithm.
[1383] 3. Coding and Data Conversion
[1384] The extracted editing elements are coded and converted into a format that can be understood by a generative artificial intelligence (AI). This coded data is then used as editing instructions for the generated video.
[1385] Terminal
[1386] The user uses a device (smartphone or PC) to perform the following operations:
[1387] 1. Enter editing instructions
[1388] You can enter detailed instructions for video composition and editing in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[1389] 2. Enter correction instructions
[1390] Check the generated video and enter corrections as necessary (e.g., change the position of subtitles, adjust the volume of background music).
[1391] Generative artificial intelligence (AI)
[1392] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server:
[1393] 1. Automatic video generation
[1394] Use standard video editing software or libraries (e.g., MoviePy) to generate a video by applying the specified structure and production elements.
[1395] 2. Video Linking
[1396] Combine multiple video clips and apply any effects or edits you want.
[1397] 3. Updates
[1398] The system receives correction instructions from the user and re-edits the video. Through this process, the final video is generated according to the user's wishes.
[1399] Specific examples
[1400] Data collection and analysis
[1401] 1. The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos. For example, it can extract the background music and subtitle styles common to videos with over 1,000,000 views.
[1402] Entering user instructions
[1403] 2. The user enters information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on the device.
[1404] example:
[1405] The video is 3 minutes long
[1406] The first 30 seconds is an intro
[1407] Display subtitles every minute
[1408] Use upbeat background music
[1409] Video Generation
[1410] 3. The server receives the user's instructions and sends them to the AI for generation. The AI then edits the video based on the received instructions and provides it to the user.
[1411] Minor correction
[1412] 4. The user reviews the generated video and makes any necessary corrections. For example, "increase the font size of the subtitles" or "lower the volume of the background music." The server then sends these corrections to the AI, which then re-edits the video.
[1413] This allows users to quickly generate high-quality videos efficiently, realizing video production that quickly responds to trends.
[1414] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1415] Step 1:
[1416] The server periodically collects metadata of popular videos from the video platform using APIs. Specifically, it obtains data such as the number of views, ratings, comments, tags, and descriptions. This provides detailed information about popular videos as input, and the collected metadata as output.
[1417] Step 2:
[1418] The server extracts editing elements of popular videos by analyzing the collected metadata. Specifically, it uses a machine learning algorithm to analyze data such as the number of views, number of ratings, and tags of each video, and detects common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.). This uses the metadata as input and provides the extracted editing elements as output.
[1419] Step 3:
[1420] The server uses a means to encode the extracted editing elements and convert them into a format understandable by the generating artificial intelligence. Specifically, the server converts the extracted editing elements into a specific format (e.g., JSON format) and provides them to the generating artificial intelligence in the form of instructions. This uses the editing elements as input and generates coded data as output.
[1421] Step 4:
[1422] The user uses their own device to input the video's structure and editing instructions in text format. Examples include instructions such as "video length should be 3 minutes," "the first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The user's editing instructions are then input to the device. Text editing instructions are generated as output.
[1423] Step 5:
[1424] The server uses a generating AI to automatically generate a video based on the composition and editing instructions received from the user. Specifically, the server analyzes the editing instructions and creates a video by applying the specified composition and production elements according to the instructions. This uses the user's editing instructions as input and produces a generated video as output.
[1425] Step 6:
[1426] The generative AI uses a means of concatenating multiple video clips, for example using the MoviePy library, and applies a specified order and effects, taking the video clips as input and generating the concatenated video as output.
[1427] Step 7:
[1428] The user checks the generated video and uses a means to input correction instructions. Specifically, while playing the generated video, the user inputs correction instructions in text format, such as "change the position of the subtitles" or "adjust the volume of the background music." This inputs the correction instructions into the terminal, and generates correction instructions as output.
[1429] Step 8:
[1430] The server uses a means to re-edit the generated video based on the correction instructions. Specifically, the server analyzes the correction instructions from the user, and the generating AI re-edits the video to generate a final video that reflects the corrections. This uses the correction instructions as input and generates the final video as output.
[1431] By following these steps, users can efficiently generate high-quality videos and then modify and re-edit them as needed.
[1432] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1433] This invention is a system that automatically extracts and analyzes the editing elements of popular videos on video platforms and automatically generates high-quality videos based on the results. Furthermore, this invention improves the user experience by combining it with an emotion engine that recognizes user emotions.
[1434] System Overview
[1435] The system consists of the following main components:
[1436] server
[1437] Terminal
[1438] Generative artificial intelligence (AI)
[1439] Emotion Engine
[1440] System program and processing flow
[1441] The system's program is designed to have the server, terminal, generative artificial intelligence (AI), and emotion engine work together. The roles and specific operations of each entity are explained below.
[1442] server
[1443] The server collects metadata of popular videos from a video platform (e.g., YouTube). This includes using an API to periodically retrieve information such as the number of views, ratings, comments, tags, and descriptions. The server analyzes the collected metadata and uses machine learning algorithms to extract specific patterns and trends. This allows for common editing elements (e.g., background music selection, subtitle style, frequency of cuts, etc.) common to popular videos.
[1444] The server then encodes the extracted editing elements and converts them into a format that can be understood by a generative artificial intelligence (AI). This encoded data is then used as editing instructions for the generated video.
[1445] Terminal
[1446] Users use their own devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[1447] The terminal transmits user input to the server, which converts it into appropriate editing instructions and sends them to the generating artificial intelligence.
[1448] Generative artificial intelligence (AI)
[1449] The generative artificial intelligence (AI) is responsible for automatically generating videos based on editing instructions sent from the server. This AI uses video editing software and libraries to apply the specified structure and production elements to generate the video. The generated video is then provided to the user, who can make minor edits as needed via their device. For example, they can input edit instructions such as "change the position of the subtitles" or "adjust the volume of the background music." These edit instructions are then sent back to the server, and the generative AI re-edits the video. Once edits are complete, the video is provided to the user again.
[1450] Emotion Engine
[1451] The emotion engine has the ability to analyze emotions from the user's facial expressions, voice, text input, etc. It recognizes emotions in real time while the user is watching a video or entering editing instructions, and analyzes that data.
[1452] The emotion data obtained from the emotion engine is sent to a server. The server uses this data to adjust editing instructions for the artificial intelligence. For example, if a user expresses positive emotion in a particular part of the video, the server can edit it to emphasize that part. If a user expresses negative emotion, the server can issue instructions to modify or delete that element.
[1453] Specific examples
[1454] Below are some specific usage examples.
[1455] 1. Data Collection and Analysis
[1456] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[1457] 2. Entering user instructions
[1458] Users enter information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used in text format on their device.
[1459] 3. Collecting Emotional Data
[1460] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[1461] 4. Video Generation
[1462] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[1463] 5. Minor edits and emotional reflection
[1464] The user can check the generated video and indicate any necessary corrections. The emotion engine also analyzes the user's emotions during the editing process, and the server readjusts the edited content based on that information.
[1465] Through this series of steps, it becomes possible to efficiently generate high-quality videos that reflect the user's emotions.
[1466] The processing flow will be explained below.
[1467] Step 1:
[1468] The server uses the video platform's API to collect metadata about popular videos, including the number of views, ratings, comments, tags, descriptions, and other data.
[1469] Step 2:
[1470] The server analyzes the collected metadata and uses machine learning algorithms to extract editing elements from the video, such as background music selection, subtitle display format, and timing of cuts.
[1471] Step 3:
[1472] The server encodes the extracted editing elements and converts them into instructions that can be understood by the generative artificial intelligence (AI), including specific editing operations and parameters.
[1473] Step 4:
[1474] Users use their own devices to input video composition and editing elements in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music."
[1475] Step 5:
[1476] The terminal transmits the user's input to the server, which includes the specific video structure and editing instructions desired by the user.
[1477] Step 6:
[1478] The server receives input data from the user and issues instructions to the AI generator based on the data, including editing elements to satisfy the user's request.
[1479] Step 7:
[1480] The artificial intelligence (AI) for generation automatically generates videos based on instructions received from the server, specifically applying specified background music, displaying subtitles, setting the timing of cuts, and other editing operations.
[1481] Step 8:
[1482] The terminal provides the generated video to the user, who watches the video and inputs correction instructions as necessary.
[1483] Step 9:
[1484] The emotion engine analyzes the user's facial expressions, voice, text input, etc. in real time while the user is watching a video, and collects emotional data.
[1485] Step 10:
[1486] The emotion engine uses the analyzed emotional data to identify parts of the video where the user had a positive or negative reaction. For example, a smile on the user's face is judged as positive, while a furrowed brow is judged as negative.
[1487] Step 11:
[1488] The server receives the emotion data sent from the emotion engine and reflects it in the artificial intelligence for generation, for example, by issuing instructions to edit parts that show positive emotions or to correct parts that show negative emotions.
[1489] Step 12:
[1490] The user can then review the generated video and make any necessary corrections, such as moving the subtitles a little higher or lowering the volume of the background music, via the device.
[1491] Step 13:
[1492] The terminal sends the user's correction instructions to the server, which include specific changes and desired corrections.
[1493] Step 14:
[1494] The server receives the correction instructions and sends them back to the AI generator, which then re-edits the video based on the instructions.
[1495] Step 15:
[1496] The generating artificial intelligence (AI) re-edits the video based on the correction instructions and generates the corrected video.
[1497] Step 16:
[1498] The terminal again provides the corrected video to the user, and this process is repeated until the necessary corrections are completed.
[1499] This series of steps enables efficient generation of high-quality videos that reflect user emotions and realizes video production that is in line with trends.
[1500] Example 2
[1501] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1502] Conventional video generation systems have had difficulty in properly reflecting the editing elements desired by users. Furthermore, they lacked video generation functionality that took user emotions into account, making it impossible to improve the quality of the user experience. This resulted in problems such as a decline in the quality of the generated videos and a decline in user satisfaction.
[1503] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1504] In this invention, the server includes means for collecting viewing data from a video providing service, means for analyzing the collected viewing data and extracting video editing elements, means for encoding the extracted editing elements and providing them to the generation AI in the form of instructions, and emotion recognition means for analyzing the user's emotions and adjusting the editing instructions based on that data, thereby enabling the efficient automatic generation of high-quality videos that reflect the user's emotions.
[1505] "Viewing data" refers to information collected from video provision services, such as the number of views, ratings, comments, tags, and descriptions.
[1506] "Editing elements" are elements that show specific patterns or trends, such as background music selection, subtitle style, and frequency of cuts, extracted by analyzing viewing data.
[1507] "Generative artificial intelligence" refers to algorithms and programs that automatically generate videos based on instruction-style data provided by the server.
[1508] The "emotion recognition means" is a system that analyzes emotions in real time from the user's facial expressions, voice, and text input, and feeds that data back to the artificial intelligence used for generation.
[1509] "Encoding" is the process of converting extracted editorial elements into a format that can be understood by a generative artificial intelligence.
[1510] A "correction instruction" is a request for specific changes or adjustments that the user makes to the generated video.
[1511] "Re-editing" is the process in which the generating AI re-edits the video based on the user's correction instructions.
[1512] This invention is a system that automatically generates high-quality videos based on viewing data collected from video provision services. This system consists of a server, a terminal, artificial intelligence for generation, and emotion recognition means.
[1513] The server collects viewing data from video-providing services (e.g., video streaming platforms) using APIs. The collected data includes the number of views, ratings, comments, tags, and descriptions. The server then analyzes this data using machine learning algorithms to extract specific patterns and trends. For example, it can identify common editing elements (such as background music selection, subtitle style, and frequency of cuts) among videos with high views. The results of this analysis are coded as extracted editing elements and converted into a format understandable by the generative AI.
[1514] Users use their devices to input video composition and editing instructions in text format, such as "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device sends these inputs to the server, which converts them into appropriate editing instructions and sends them to the AI generator.
[1515] The artificial intelligence (AI) for generation automatically generates videos based on editing instructions sent from the server. This AI generates videos by applying the specified composition and production elements using video editing software such as Adobe Premiere Pro or FFmpeg. The generated videos are provided to users, who can play and check them on their devices.
[1516] The emotion recognition means analyzes the emotions of the user in real time while watching the video or inputting editing instructions. Emotional data is acquired from the user's facial expressions, voice, and text input, and sent to the server. This emotional data is fed back to the generating AI and used to emphasize or modify specific parts of the video. For example, editing instructions are issued to emphasize parts where the user smiles more and modify parts where the user frowns.
[1517] As an example, consider the following prompt:
[1518] The video is 3 minutes long
[1519] "The first 30 seconds are the intro"
[1520] "Show subtitles every minute"
[1521] "Use upbeat background music"
[1522] This allows us to automatically generate videos that meet the user's preferences, and further improve the user experience by utilizing emotion recognition.
[1523] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1524] Step 1:
[1525] The server collects viewing data from the video provider. Specifically, it periodically obtains metadata such as the number of views, ratings, comments, tags, and descriptions using an API. It sends an API request as input and obtains viewing data as output. This data is stored in the server's database.
[1526] Step 2:
[1527] The server analyzes the collected viewing data. It uses machine learning algorithms to analyze the data and extract certain patterns and trends, such as popular background music, subtitle styles, or frequency of cuts. It uses the viewing data collected in step 1 as input and obtains the analysis results as output. These analysis results are stored on the server for use in later steps.
[1528] Step 3:
[1529] The server encodes the analysis results and converts the extracted editing elements into a format that the AI can understand. Specifically, it converts the data into JSON format and prepares it for passing to the AI. It uses the analysis results as input and generates coded editing elements as output.
[1530] Step 4:
[1531] Users input video composition and editing instructions in text format using their own devices. Examples of input include "video length should be 3 minutes," "first 30 seconds should be an intro," "display subtitles every minute," and "use upbeat background music." The device collects the user's instructions as input and sends them to the server.
[1532] Step 5:
[1533] The server receives instructions from the user. After receiving these instructions, it analyzes them, converts them into appropriate editing instructions, and sends them to the AI generator. The input is the user's instructions, and the output is instructions converted into a format that can be sent to the AI generator.
[1534] Step 6:
[1535] The generative artificial intelligence (AI) automatically generates videos based on editing instructions sent from the server. It uses video editing software (e.g., Adobe Premiere Pro or FFmpeg) to edit the videos as instructed. It receives the editing instructions as input and provides the generated video as output to the user.
[1536] Step 7:
[1537] The user checks the generated video and inputs any necessary corrections through the device. For example, this includes instructions such as "change the display position of the subtitles to the bottom right" or "reduce the background music volume by 20%." The device collects the corrections as input and sends them back to the server.
[1538] Step 8:
[1539] The server receives the correction instructions and sends them to the AI for generation. The received instructions are analyzed and sent back to the AI as appropriate re-editing instructions. The input is the user's correction instructions, and the output is the re-editing instructions.
[1540] Step 9:
[1541] The generative artificial intelligence (AI) re-edits the video based on the correction instructions. The corrections are reflected again using video editing software to generate the final video. The correction instructions are received as input, and the corrected video is provided to the user as output.
[1542] Step 10:
[1543] The emotion recognition means analyzes the user's facial expressions, voice, and text input to obtain emotion data. It performs the analysis in real time while the user is watching the video and sends the results to the server. It collects the user's real-time data as input and sends the analyzed emotion data as output to the server.
[1544] Step 11:
[1545] The server sends additional editing instructions to the AI based on the emotion data. For example, it may emphasize parts that show positive emotions and modify or delete parts that show negative emotions. The server receives emotion data as input and sends additional editing instructions to the AI as output.
[1546] (Application example 2)
[1547] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1548] In modern virtual stores, it is important to generate efficient, high-quality promotional videos to attract customer attention. However, existing methods require a significant amount of manual effort and time for video editing and generation, and it is difficult to provide optimal promotional videos that reflect customers' real-time emotions. Furthermore, it is difficult to immediately reflect users' feedback while they are watching, making it difficult to improve user satisfaction.
[1549] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for collecting metadata of popular videos from a video platform, means for analyzing the collected metadata and extracting editing elements for the video, means for encoding the extracted editing elements and providing them in the form of instructions to a generating artificial intelligence, emotion recognition means for collecting real-time emotional data from users, means for adjusting the editing content of the video using the collected emotional data, and means for viewing and editing the video using a device such as smart glasses or a head-mounted display as a user terminal. This makes it possible to automatically generate optimal promotional videos that reflect the user's emotions and effectively improve user satisfaction and interest in the virtual store.
[1550] A "video platform" is a general term for services that allow various users to upload, watch, and share videos online.
[1551] "Metadata" refers to data related to a video, including information such as the number of views, ratings, comments, tags, and descriptions.
[1552] "Generative AI" refers to an AI system designed to automatically generate videos based on collected data and user instructions.
[1553] "Emotion recognition means" refers to technology or equipment that analyzes a user's facial expressions, voice, etc. to detect emotions in real time.
[1554] The term "user terminal" refers to a device used by a user, and in the present invention particularly includes devices such as smart glasses and head-mounted displays.
[1555] "Emotion Data" refers to information collected by an emotion recognition means about a user's real-time emotional state.
[1556] "Editing elements" are elements that indicate optimal patterns and techniques for creating and editing videos, and include specific elements such as background music, subtitle style, and frequency of cuts.
[1557] "Promotional videos" refer to videos created for the purpose of promoting products or services within a virtual store.
[1558] A "virtual store" refers to a virtual store that exists on the Internet or in a virtual reality space, providing an environment in which users can visually view products and services.
[1559] The present invention relates to a system for automatically generating promotional videos in a virtual store and optimizing the videos based on real-time visitor sentiment. Specific embodiments are described below.
[1560] overview
[1561] The system consists of the following main components:
[1562] server
[1563] Devices (smart glasses, head-mounted displays, etc.)
[1564] Generative artificial intelligence (AI)
[1565] Emotion Recognition Engine
[1566] Hardware and software used
[1567] 1. Server:
[1568] The server uses an API from a video platform (e.g., YouTube) to collect metadata for popular videos, obtaining information such as the number of views, ratings, comments, tags, and descriptions.
[1569] Based on the collected metadata, machine learning libraries (e.g., scikit-learn, TensorFlow) are used to analyze specific patterns and trends and extract editorial elements.
[1570] The extracted editing elements are coded and converted into a format that can be understood by the generative artificial intelligence.
[1571] 2. User Device:
[1572] The user inputs the video composition and editing instructions in text format through a terminal (smart glasses or a head-mounted display).
[1573] Real-time emotional data is collected while the user is watching through the camera and microphone built into the device.
[1574] 3. Generative Artificial Intelligence (AI):
[1575] The artificial intelligence for generation automatically generates videos based on editing instructions and emotional data sent from the server.
[1576] The generated video is provided to the user once, and modifications can be made as needed.
[1577] 4. Emotion Recognition Engine:
[1578] The emotion recognition engine analyzes the user's facial expressions and voice in real time to collect emotional data.
[1579] The collected emotional data is sent to a server and used by the generative AI to adjust the editing content of the video.
[1580] Program Processing Overview
[1581] Data collection and analysis:
[1582] The server collects metadata from video platforms and analyzes and extracts editing elements of popular videos, such as "common background music and subtitle styles in videos with over 1,000,000 views."
[1583] Enter user instructions:
[1584] The user inputs information such as the length of the video, the length of the intro, the timing of the subtitles to be displayed, and the type of background music to be used on the device.
[1585] Emotion data collection:
[1586] While the user is watching the video, the emotion engine recognizes the user's facial expressions and voice to collect emotional data. For example, frequent smiles indicate positive emotions, while frequent frowns indicate negative emotions.
[1587] Video Generation:
[1588] The server receives user instructions and emotion data and sends them to the AI for generation. The AI then edits the video based on the received instructions and emotion data and provides it to the user.
[1589] Tweaks and Emotions:
[1590] The user checks the generated video and makes any necessary corrections. The emotion engine also analyzes the user's emotions during the corrections, and the server re-edits the video based on those. For example, the user can input correction instructions such as "change the display position of the subtitles" or "adjust the volume of the background music."
[1591] Specific examples
[1592] Prompt Sentence Examples
[1593] API endpoint: https: / / api.videoplatform.com / metadata
[1594] Parameters: Over 1,000,000 views
[1595] User instructions:
[1596] Length: 3 minutes
[1597] Intro length: 30 seconds
[1598] Subtitles: Every minute
[1599] BGM: Upbeat
[1600] Emotional Data:
[1601] Video Frame: {...}
[1602] Analysis results: Smiling a lot indicates positive emotions, frowning a lot indicates negative emotions
[1603] In this way, the present invention can improve the user experience in a virtual store by automatically generating an optimal promotional video that reflects the user's emotions.
[1604] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1605] Step 1:
[1606] The server collects metadata of popular videos from a video platform using an API. The input is the API endpoint and parameters of the video platform, and the output is the collected metadata (number of views, number of ratings, number of comments, etc.).
[1607] Step 2:
[1608] The server analyzes the collected metadata and extracts the editing elements of the video. The input is the metadata, and the output is the extracted editing elements (e.g., type of background music, subtitle style, frequency of cuts, etc.). Machine learning libraries (e.g., scikit-learn, TensorFlow) are used for the analysis.
[1609] Step 3:
[1610] The server encodes the extracted editing elements and converts them into a format that can be understood by the AI for generation. The input is the editing elements, and the output is coded editing instructions. The specific operation is to map the editing elements to the corresponding code.
[1611] Step 4:
[1612] The user inputs video composition and editing instructions in text format through a device (smart glasses or a head-mounted display). The input is text format instructions, and the output is instruction data sent to the server.
[1613] Step 5:
[1614] The server sends data to the generating AI based on the composition and editing instructions received from the user, where the input is the user's instruction data and coded editing instructions, and the output is the data sent to the generating AI.
[1615] Step 6:
[1616] The artificial intelligence for generation automatically generates videos based on the received composition and editing instructions. The input is data sent from the server, and the output is the generated video file. Specific operations involve the use of video editing software and libraries.
[1617] Step 7:
[1618] While a user is watching a video generated on a device, the emotion recognition means analyzes the user's facial expressions and voice to collect emotion data in real time. The input is the user's facial expressions and voice, and the output is the collected emotion data.
[1619] Step 8:
[1620] The server analyzes the collected emotion data and adjusts the editing content of the video. The input is emotion data, and the output is adjusted editing instructions. Specifically, editing is performed to emphasize positive emotions.
[1621] Step 9:
[1622] The user inputs correction instructions for the video. The input is the user's correction instructions, and the output is correction instruction data sent to the server.
[1623] Step 10:
[1624] The server sends instructions for re-editing to the AI based on the correction instructions and the collected emotion data. The input is the correction instruction data and emotion data, and the output is the data for re-editing.
[1625] Step 11:
[1626] The artificial intelligence for generation re-edits the video based on the re-editing instructions. The input is the data for re-editing, and the output is the corrected video file. Specifically, it edits the specified parts.
[1627] This series of steps enables the system to provide optimal promotional videos within the virtual store that reflect the user's emotions and feedback.
[1628] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1629] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1630] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1631] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1632] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1633] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1634] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1635] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1636] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1637] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1638] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1639] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1640] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1641] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1642] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1643] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1644] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1645] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1646] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1647] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1648] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1649] The following is further disclosed regarding the above embodiment.
[1650] (Claim 1)
[1651] A means of collecting metadata of popular videos from video platforms;
[1652] A means for analyzing the collected metadata and extracting editing elements of the video;
[1653] A means for coding the extracted editing elements and providing them to a generating artificial intelligence in the form of instructions;
[1654] means for receiving video composition and editing instructions from a user;
[1655] A system including a generative artificial intelligence means for automatically generating video based on received composition and editing instructions.
[1656] (Claim 2)
[1657] 2. The system according to claim 1, further comprising means for evaluating the quality of the generated video based on the received video viewing data, and updating and improving the editing elements in accordance with the evaluation.
[1658] (Claim 3)
[1659] A means for a user to input correction instructions for the generated video;
[1660] 2. The system according to claim 1, further comprising means for re-editing the generated video based on the correction instructions.
[1661] "Example 1"
[1662] (Claim 1)
[1663] A means of collecting information on highly rated videos from video sharing services,
[1664] A means of analyzing the collected information and extracting video editing elements;
[1665] A means for encoding the extracted editing elements and providing them to a generating artificial intelligence in the form of instructions;
[1666] means for receiving video composition and editing instructions from a user;
[1667] artificial intelligence generating means for automatically generating a video based on the received composition and editing instructions;
[1668] A means for a user to input correction instructions for the generated video;
[1669] a means for re-editing the generated video based on the correction instructions;
[1670] A system including:
[1671] (Claim 2)
[1672] 2. The system according to claim 1, further comprising means for evaluating the quality of the generated video based on the received video viewing data, and updating and improving the editing elements in accordance with the evaluation.
[1673] (Claim 3)
[1674] 10. The system of claim 1, further comprising means for a user to input additional corrections or modifications to the re-edited video in real time.
[1675] "Application Example 1"
[1676] (Claim 1)
[1677] A means of collecting metadata of popular videos from video platforms;
[1678] A means for analyzing the collected metadata and extracting editing elements of the video;
[1679] A means for coding the extracted editing elements and providing them to a generating artificial intelligence in the form of instructions;
[1680] means for receiving video composition and editing instructions from a user;
[1681] A generating artificial intelligence means for automatically generating a video based on the received composition and editing instructions;
[1682] a means for concatenating a plurality of video clips;
[1683] A means for a user to input correction instructions for the generated video through a smart device;
[1684] The system includes a means for re-editing the generated video based on correction instructions.
[1685] (Claim 2)
[1686] The system according to claim 1, wherein the system evaluates the quality of the generated video based on the received video viewing data, and updates and improves the editing elements in accordance with the evaluation.
[1687] (Claim 3)
[1688] 10. The system of claim 1, which provides prompt sentences to input to the generative AI model.
[1689] "Example 2: Combining Emotion Engines"
[1690] (Claim 1)
[1691] A means for collecting viewing data from a video providing service;
[1692] A means of analyzing the collected viewing data and extracting video editing elements;
[1693] A means for coding the extracted editing elements and providing them to a generating artificial intelligence in the form of instructions;
[1694] means for receiving video composition and editing instructions from a user;
[1695] artificial intelligence generating means for automatically generating a video based on the received composition and editing instructions;
[1696] The system includes an emotion recognition means that analyzes the user's emotions and adjusts editing instructions based on that data.
[1697] (Claim 2)
[1698] The system according to claim 1, further comprising means for evaluating and updating and improving editing elements based on emotional data collected while a user is viewing the generated video.
[1699] (Claim 3)
[1700] A means for a user to input correction instructions for the generated video;
[1701] 2. The system according to claim 1, further comprising means for re-editing the generated video based on the correction instructions.
[1702] "Application example 2 when combining emotion engines"
[1703] (Claim 1)
[1704] A means of collecting metadata of popular videos from video platforms;
[1705] A means for analyzing the collected metadata and extracting editing elements of the video;
[1706] A means for coding the extracted editing elements and providing them to a generating artificial intelligence in the form of instructions;
[1707] means for receiving video composition and editing instructions from a user;
[1708] artificial intelligence generating means for automatically generating a video based on the received composition and editing instructions;
[1709] emotion recognition means for collecting real-time emotion data of a user;
[1710] A means to use collected emotional data to adjust the editing of videos;
[1711] A system that includes a means for viewing and modifying videos using devices such as smart glasses or head-mounted displays as user terminals.
[1712] (Claim 2)
[1713] 2. The system according to claim 1, further comprising means for evaluating the quality of the generated video based on the received viewing data and emotion data of the video, and updating and improving the editing elements in accordance with the evaluation.
[1714] (Claim 3)
[1715] A means for a user to input correction instructions for the generated video;
[1716] 2. The system according to claim 1, further comprising means for re-editing the generated video based on the correction instructions and the emotion data. [Explanation of symbols]
[1717] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of collecting metadata of popular videos from video platforms; A means for analyzing the collected metadata and extracting editing elements of the video; A means for coding the extracted editing elements and providing them to a generating artificial intelligence in the form of instructions; means for receiving video composition and editing instructions from a user; A system including a generative artificial intelligence means for automatically generating video based on received composition and editing instructions.
2. The system according to claim 1, further comprising means for evaluating the quality of the generated video based on the received viewing data of the video, and updating and improving the editing elements in response to the evaluation.
3. A means for a user to input correction instructions for the generated video; 2. The system according to claim 1, further comprising means for re-editing the generated video based on the correction instructions.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A