System
The system uses AI to analyze and edit user videos, extracting viral elements and updating based on feedback, addressing the challenge of creating engaging content efficiently and effectively.
Patent Information
- Application Number
- JP2024131391
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-02-20
AI Technical Summary
Creators face significant challenges in efficiently creating viral videos, as the process is time-consuming and requires specialized knowledge to tailor edits to individual user preferences, making it difficult for average users to produce engaging content.
A system utilizing generated artificial intelligence to analyze highly viewed videos, extract buzzworthy elements, automatically edit user-uploaded videos, provide edited results, and update AI based on user feedback to optimize future edits.
Enables users to create high-quality, viral videos with reduced effort by automating the editing process and improving accuracy over time through learning from user feedback.
Smart Images

Figure 2026028775000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, with the spread of video sharing platforms, many creators have been spending a great deal of time and effort on filming, editing, and uploading videos. This process requires creators to effectively incorporate elements that will create a "viral" video, which places a significant burden on creators. Furthermore, it is difficult to create personalized edits tailored to individual users' preferences. For this reason, there is a need for a system that allows users to effectively create viral videos without the hassle. [Means for solving the problem]
[0005] The system of the present invention comprises the following means:
[0006] 1. A method of using generated artificial intelligence to analyze a large number of highly viewed videos and extract elements that will create a buzz.
[0007] 2. A means of receiving and analyzing videos uploaded by users.
[0008] 3. A means for automatically editing the received video based on the extracted buzzworthy elements.
[0009] 4. A means of providing edited videos to users and receiving feedback.
[0010] 5. A means for updating the generated artificial intelligence based on said feedback.
[0011] This allows users to simply upload their videos and automatically create videos incorporating elements that will effectively create buzz. Furthermore, as the AI continues to learn based on feedback, it becomes possible to edit videos that are optimized to suit the user's preferences, significantly reducing the burden on creators.
[0012] "Generated artificial intelligence" is a program designed to learn from a large amount of data and perform specific tasks automatically.
[0013] A "highly viewed video" is a video that has received a large number of views on a video sharing platform.
[0014] "Buzz elements" are characteristics or factors that are considered effective in increasing the number of views of a video.
[0015] "Receiving" is the act of the server receiving data sent by the user.
[0016] "Analysis" is the act of examining the content of received data or video in detail and extracting its features and patterns.
[0017] "Automatic editing" means that the program edits the video according to specified rules and patterns without human intervention.
[0018] "Providing" refers to the act of showing the edited video to users or allowing them to download it.
[0019] "Feedback" is the act of collecting reactions, opinions, and evaluations from users.
[0020] "Updating" is the act of AI improving its algorithms and models based on new data and feedback.
[0021] A "system" is a general term for a series of components or programs that combine multiple means and elements to achieve a specific function. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0024] First, the terms used in the following description will be explained.
[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0030] [First embodiment]
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0043] ---
[0044] This invention is a system that uses the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: video analysis, element extraction, automatic editing, provision, and updating the AI based on feedback.
[0045] Program processing and explanation
[0046] Uploading videos
[0047] User:
[0048] Users upload video files from their own devices using the system's dedicated web interface or application, and can also enter a brief description and keywords for the video.
[0049] Receiving and storing videos
[0050] server:
[0051] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[0052] Extracting buzzworthy elements
[0053] server:
[0054] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted.
[0055] Analysis and editing of user videos
[0056] server:
[0057] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music features through audio analysis. Then, based on the extracted viral elements, it automatically performs the following edits:
[0058] Automatically cut unnecessary scenes
[0059] Highlighting the highlights
[0060] Adding sound effects and music
[0061] Inserting text and titles
[0062] Automatic thumbnail generation
[0063] Providing edited results
[0064] server:
[0065] To provide users with the completed edited video, a dedicated preview link is generated. This link is then sent to the user, allowing them to view the edited video on the preview screen. Suggested thumbnails and titles are also displayed at the same time.
[0066] Receiving feedback and updating the AI
[0067] User:
[0068] Users can view the edited video via a preview link and provide feedback, including requested corrections and a satisfaction rating.
[0069] server:
[0070] The server receives user feedback and incorporates it into the generation AI as training data. This updates the AI algorithm so that subsequent edits are more suited to the user's preferences. This feedback loop gradually improves the system's accuracy and user satisfaction.
[0071] ---
[0072] Specific examples
[0073] Example 1: Travel video editing
[0074] User:
[0075] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[0076] server:
[0077] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[0078] User:
[0079] Check the edited results and provide feedback such as "I would like the text color in the thumbnail to be changed."
[0080] server:
[0081] The text color is changed based on the user's instructions, and the final video is provided to the user. The AI continues to learn based on the feedback.
[0082] Example 2: Editing a gadget review video
[0083] User:
[0084] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[0085] server:
[0086] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0087] User:
[0088] We will review the edited video and provide feedback if we are satisfied, and if there are any requests for re-editing, we will provide specific instructions on what needs to be corrected.
[0089] server:
[0090] Based on the feedback, we make any necessary corrections and provide the final video. We also update the AI based on the feedback information to improve accuracy.
[0091] ---
[0092] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[0093] The processing flow will be explained below.
[0094] ---
[0095] Step 1: User uploads a video
[0096] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[0097] Step 2: The server receives and stores the video
[0098] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[0099] Step 3: Extract elements that will make the server buzz
[0100] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[0101] Step 4: The server analyzes the user's video
[0102] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[0103] Step 5: The server will automatically edit the video
[0104] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[0105] Cutting out unnecessary scenes
[0106] Highlighting the highlights
[0107] Adding sound effects and music
[0108] Inserting text and titles
[0109] Automatic thumbnail generation
[0110] Step 6: The server serves the edits
[0111] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[0112] Step 7: User reviews the video and provides feedback
[0113] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[0114] Step 8: The server updates the AI based on the feedback
[0115] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[0116] ---
[0117] The above is a specific explanation of the processing steps of an automatic video editing system using generative AI.
[0118] Example 1
[0119] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0120] Conventional video editing systems require users to edit videos manually, which requires a great deal of time and effort. Furthermore, determining whether a video will go viral and determining the optimal editing method requires specialized knowledge, making it difficult for average users to use. Furthermore, the manual improvement process after receiving feedback is inefficient, making it difficult to improve user satisfaction.
[0121] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0122] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated AI and extracting viral elements, means for receiving and saving video files uploaded by users, means for automatically editing the received videos based on the extracted viral elements, means for generating a preview link of the edited results and providing it to the user and receiving feedback, and means for updating the generated AI based on the feedback. This automates the editing work that users would otherwise do manually, enabling optimal editing to increase the likelihood of a video going viral. Furthermore, by updating the AI based on user feedback, editing accuracy can be improved from the next time onwards, thereby continuously increasing user satisfaction.
[0123] 1. "Generated artificial intelligence" refers to algorithms that learn from a large amount of data and then use the results to automatically perform specific tasks.
[0124] 2. "Highly viewed video" means a video that has been viewed by a large number of viewers on an online platform.
[0125] 3. "Buzz elements" refer to the features and factors that increase the number of views of a video, including the video content, editing patterns, thumbnails, titles, keywords, etc.
[0126] 4. "User" refers to an individual or corporation that uses this system to upload videos and receive edited results.
[0127] 5. "Video File" means a digital file containing video and associated audio data.
[0128] 6. "Storage means" refers to a method or device for temporarily or permanently retaining video files in storage.
[0129] 7. "Editing means" means an algorithm or program that automatically processes video, such as cutting out unnecessary scenes, adding sound effects, or inserting text.
[0130] 8. "Preview Link" means a URL that allows users to view the edited video online.
[0131] 9. "Feedback" means any opinions or requests for improvements regarding edits provided by a User.
[0132] 10. "Updating means" means a method or device for correcting or improving the generated AI algorithm based on the feedback received.
[0133] This invention is a system that utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: uploading videos, receiving and saving them, extracting viral elements, analyzing and editing them, providing the editing results, receiving feedback, and updating the AI.
[0134] First, a user uploads a video file from their device through a dedicated web interface or application, and enters the video title, description, and keywords. For example, the title might be "Summer Hawaii Trip" and the keywords "beach, scenery."
[0135] The device sends the selected video file to the system server. The server receives the video file sent from the device and temporarily stores it in the system's storage. It analyzes the video's metadata (resolution, format, file size, etc.) and records that information in a database. For example, it may be analyzed that the video's resolution is 1920x1080 pixels.
[0136] Next, the generative AI model run by the server analyzes a large amount of video data with high past views. Specifically, it references a dataset including the video content, editing patterns, thumbnails, titles, and keywords to extract elements that will create buzz. For example, in a video containing the keywords "beach" and "scenery," it detects the points where many viewers played the video as scenes of waterfronts.
[0137] The server analyzes the videos uploaded by users, tags each scene in a timeline, and extracts important lines and background music through audio analysis.Then, based on the extracted viral elements, it performs the following editing:
[0138] Automatically cut unnecessary scenes
[0139] Highlighting the highlights
[0140] Adding sound effects and music
[0141] Inserting text and titles
[0142] Automatic thumbnail generation
[0143] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[0144] To provide the user with the completed edited video, the server generates a dedicated preview link, which is sent to the user via email or app notification. The preview screen allows the user to view the edited video, along with suggested thumbnails and titles.
[0145] Users can view the edited video through a preview link and provide feedback, including requests for corrections and a satisfaction rating. For example, a user could send feedback such as, "Please change the text color in the thumbnail to blue."
[0146] The server receives user feedback and feeds it into the generation AI as training data. This routine updates the AI algorithm and improves its accuracy so that the next edit will be more tailored to the user's preferences.
[0147] Specific examples
[0148] Example 1: Travel video editing
[0149] The user uploads a video they shot during their trip. They enter a brief description of their trip, "Summer trip to Hawaii," and the keywords "beach, scenery." The server analyzes the video and audio from the travel video and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generation AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails. The user reviews the edited results and provides feedback, such as "I'd like the text color in the thumbnail changed." The server changes the text color based on the instructions and provides the final video to the user. The AI continues to learn based on this feedback.
[0150] Example 2: Editing a gadget review video
[0151] A user uploads a review video of a new gadget. They enter "technology, review, new product" as keywords. The server analyzes the content of the gadget review video and extracts scenes that highlight the product's unique functions and features. The generation AI extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails. The user reviews the edited video and provides feedback if satisfied. If there is a request for re-editing, the user can specify specific corrections. The server makes the necessary corrections based on the feedback and provides the final video. The AI is updated based on the feedback to improve accuracy.
[0152] Prompt Sentence Examples
[0153] 1. "I'm uploading a video of a beach in Hawaii that I took during my trip. Please auto-edit it. The keywords are 'beach' and 'scenery.'"
[0154] 2. "I'm going to upload a video reviewing a new gadget. I need a catchy title and thumbnail. The keywords are 'technology,' 'review,' and 'new product.'"
[0155] This allows users to easily create high-quality videos that are likely to receive many views, reducing the burden on users and increasing the chances of a video's success.
[0156] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0157] Step 1: Upload your video
[0158] User:
[0159] A user uploads a video file from their own device through a dedicated web interface or application. As input, the user specifies the video file, title, description, and keywords. For example, the user enters the title "Summer Hawaii Trip" and the keywords "Beach, Scenery." As output, the specified video file is sent to the system server.
[0160] Step 2: Receive and save the video
[0161] server:
[0162] The server receives video files sent from the device. The inputs are the video files and their metadata (resolution, format, file size, etc.). This is analyzed, and the metadata is temporarily saved in the system's storage along with the video files. The metadata is recorded in a database. For example, the data that the video resolution is 1920x1080 pixels is recorded. The saved video files and metadata are obtained as output.
[0163] Step 3: Identifying viral elements
[0164] server:
[0165] The generative AI model run by the server analyzes large amounts of video data that have recorded high numbers of views in the past. The input is a dataset of videos with high numbers of views in the past. This analysis extracts buzzworthy elements from the dataset, which includes the video content, editing patterns, thumbnails, titles, and keywords. For example, in a video containing the keywords "beach" and "scenery," the points where many viewers played the video are detected as scenes of waterfronts. The output is a list of the extracted buzzworthy elements.
[0166] Step 4: Analyze and edit user videos
[0167] server:
[0168] The server analyzes videos uploaded by users. As input, it receives the uploaded video file and a list of extracted viral elements. It tags each scene in a timeline and extracts important dialogue and background music features through audio analysis. It then performs the following edits and generates an edited video file as output:
[0169] Automatically cut unnecessary scenes
[0170] Highlighting the highlights
[0171] Adding sound effects and music
[0172] Inserting text and titles
[0173] Automatic thumbnail generation
[0174] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[0175] Step 5: Submit your edits
[0176] server:
[0177] The server generates a dedicated preview link to provide the user with the edited video. The inputs are the edited video file and the thumbnail and title suggested by the generation AI. The output is a notification message containing the preview link sent to the user. For example, the email notifying the user includes the message "Check out the preview of your new video" and a link.
[0178] Step 6: Receive feedback and update the AI
[0179] User:
[0180] Users can view the edited video through a preview link and provide feedback. The input is feedback or requests for improvements. For example, a user might send feedback such as, "Please change the text color in the thumbnail to blue." The output is reflected in the system.
[0181] server:
[0182] The server receives feedback from the user and incorporates it into the generation AI as learning data. The input is the user's feedback information, and the AI algorithm is updated based on this. The output is an updated AI algorithm, which improves the accuracy of future edits. Specifically, the text color of the thumbnail is changed to blue and provided to the user again. This feedback information is also saved as reference data for the next edit.
[0183] (Application example 1)
[0184] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0185] In recent years, with the spread of video distribution services, there has been a demand for ways to improve the quality of video content created by individuals and companies and increase the number of views. In particular, it is not easy for individuals to easily upload videos from mobile devices such as smartphones and effectively edit them. Furthermore, there are not enough methods in place to reflect user feedback and use it in editing the next video. Therefore, providing an efficient system for automatically creating videos that are likely to go viral is a challenge.
[0186] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0187] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence to extract buzzworthy elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzzworthy elements, means for providing the edited videos to users as preview links and receiving feedback, and means for updating the generated artificial intelligence based on the feedback, thereby enabling users to efficiently and easily create videos that are likely to go viral and continuously improve the quality through feedback.
[0188] "Generated artificial intelligence" refers to algorithms or systems that are trained to analyze large amounts of video data and learn patterns to edit and extract elements from videos.
[0189] "Buzz elements" are specific visual, audio, editing patterns, and other characteristics that significantly increase the number of views and viewer response of a video.
[0190] "Receiving" is the process in which the system takes in the video file uploaded by the user and temporarily stores it for processing.
[0191] "Analysis" is the process of analyzing the content of a video and identifying its features and patterns, including elements such as video, audio, and scene composition.
[0192] "Editing" is the process of changing, adding, or deleting the content of a video, and includes emphasizing specific scenes, cutting unnecessary scenes, adding sound effects, music, etc.
[0193] The "preview link" is a temporary URL provided to the user so that the user can check the video after editing is complete.
[0194] "Feedback" refers to the opinions and ratings users provide on edited videos, which are used to improve the AI model.
[0195] "Updating" is the process by which the generated AI learns from new data and feedback to improve the accuracy of the next video analysis and editing.
[0196] The "server" is a computer system that receives videos from users, analyzes, edits, stores, and processes feedback.
[0197] This invention relates to a system that uses generated artificial intelligence to automatically edit videos to make them more likely to go viral. This system automatically performs a series of processes, including video analysis, element extraction, automatic editing, provision, and updating the artificial intelligence based on feedback.
[0198] Uploading and saving videos
[0199] Users upload videos using a dedicated application on their smartphones or other devices. The videos are sent to a server and stored immediately upon receipt. Users can also enter a brief description of the video and keywords.
[0200] Extracting and analyzing buzzworthy elements
[0201] The server uses the generated AI to analyze a large number of highly viewed videos and extract elements that will create buzz, including video content, editing patterns, music selection, thumbnails, titles, keywords, etc. Based on this, videos uploaded by users are analyzed.
[0202] Automatic editing function
[0203] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, it performs the following edits:
[0204] Cutting out unnecessary scenes
[0205] Highlighting the highlights
[0206] Adding sound effects and music
[0207] Inserting text and titles
[0208] Automatic thumbnail generation
[0209] Preview and Feedback
[0210] Once the video is edited, it will be provided to the user as a dedicated preview link. Users can use the preview link to check the edited results and provide feedback. For example, users can upload a video of their "summer trip to Hawaii" and edit it to incorporate elements that will create buzz (unique scenery, lively music, catchy title). Users can also upload a review video of a new gadget and edit it to highlight unique product features.
[0211] AI Updates
[0212] The feedback provided by users is stored on a server and then incorporated into the generated AI, which then updates the AI algorithm to ensure that future edits are more accurate and tailored to the user's preferences.
[0213] Details of the hardware and software you will be using
[0214] The system uses servers equipped with high-performance CPUs and GPUs. The main software used is Django (a Python framework), FFmpeg (video processing), and TensorFlow (AI model). Django is used to manage the reception, storage, analysis, editing, and preview link generation of uploaded videos. FFmpeg is used for video analysis and editing, while TensorFlow is used to train and update the generative AI model.
[0215] (Examples of specific examples and prompts)
[0216] Specific examples
[0217] Example 1: A user uploads a video of their "Summer Trip to Hawaii" that they shot during their trip and edits it to emphasize beautiful scenery and upbeat music.
[0218] Example 2: A user uploads a review video for a new gadget and edits it to highlight the product's unique features.
[0219] Prompt Sentence Examples
[0220] Please provide us with an automatically edited video of your "Summer Trip to Hawaii" with elements that will create buzz (unique scenery, lively music, catchy title).
[0221] In this way, users can easily create high-quality videos that are likely to go viral. This system reduces the burden on creators and increases the success rate of videos.
[0222] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0223] Step 1:
[0224] The user starts the application, selects a video file, and uploads it. The user can also enter a description and keywords for the video. This input data (video file and metadata) is sent to the server. The server receives the sent data and temporarily saves it in storage. This process saves the video file and metadata on the server.
[0225] Step 2:
[0226] The server analyzes the metadata of the video file and records it in a database. As a result of the analysis, the video format, resolution, file size, etc. are extracted. This information is stored in the database and used for subsequent processing.
[0227] Step 3:
[0228] The server uses the generative AI model to analyze a large number of highly viewed videos stored in a database and extract viral elements. This extraction process identifies features such as video content, editing patterns, music selection, thumbnails, titles, and keywords. The extracted viral elements are stored as internal data in the AI model.
[0229] Step 4:
[0230] The server uses a generative AI model to analyze videos uploaded by users. It tags each scene along the video's timeline and extracts important lines and background music through audio analysis. Based on the results of this analysis, the next step is automatic editing.
[0231] Step 5:
[0232] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, the following editing is performed:
[0233] Cut Unwanted Scenes: Automatically remove unwanted scenes from your video.
[0234] Highlighting the highlights: Highlight important scenes.
[0235] Add sound effects and music: Add sound effects and background music that suit your video.
[0236] Insert text or title: Insert catchy text or title.
[0237] Auto-generate thumbnails: Automatically generate visually appealing thumbnails.
[0238] Step 6:
[0239] A preview link is generated for the completed edited video and provided to the user. The user can access this preview link through the application and check the edited results. This process generates the preview link and notifies the user.
[0240] Step 7:
[0241] Users can review the edited video and provide feedback, including ratings and specific suggestions for correction, which is then sent to the server.
[0242] Step 8:
[0243] The server receives feedback from users and updates the generative AI model. Based on the feedback data, the AI model's algorithm is adjusted to improve the accuracy of video analysis and editing in future videos. This feedback loop continuously improves the performance of the entire system.
[0244] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0245] ---
[0246] This invention utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral, and also combines it with an emotion engine that recognizes the user's emotions. This system automatically performs a series of processes including video analysis, element extraction, automatic editing, and the collection and utilization of emotional feedback.
[0247] Program processing and explanation
[0248] Uploading videos
[0249] User:
[0250] Users upload video files from their own devices using the system's dedicated web interface or application, and can enter a brief description and keywords for the video.
[0251] Receiving and storing videos
[0252] server:
[0253] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[0254] Extracting buzzworthy elements
[0255] server:
[0256] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted and used to generate patterns.
[0257] Analysis and editing of user videos
[0258] server:
[0259] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music through audio analysis. It then automatically edits the video based on the extracted viral elements, as follows:
[0260] Cutting out unnecessary scenes
[0261] Highlighting the highlights
[0262] Adding sound effects and music
[0263] Inserting text and titles
[0264] Automatic thumbnail generation
[0265] Submitting edits and gathering feedback
[0266] server:
[0267] The edited video, along with the suggested thumbnail and title, is provided to the user via a preview link that is generated for the user. The user is then notified of this link, allowing them to view the edited video on the preview screen.
[0268] User:
[0269] Visit the preview link provided to see the edited video and provide feedback, including any correction requests and a satisfaction rating.
[0270] Use of emotion engine
[0271] server:
[0272] The server uses an emotion engine to recognize the user's emotions in real time while watching videos. The emotion engine analyzes the user's facial expressions, tone of voice, and reactions to collect emotional data such as joy, surprise, sadness, and excitement.
[0273] Utilizing Emotional Feedback and Updating AI
[0274] server:
[0275] Based on the emotional feedback collected by the emotion engine, the generative AI adjusts the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This allows for further personalization.
[0276] User:
[0277] Review the final edit and provide further feedback as needed, including requests for re-edits and additional emotional feedback.
[0278] server:
[0279] Based on the feedback, necessary corrections are made and the final video is provided to the user. The generation AI and emotion engine are updated based on the feedback information to improve accuracy.
[0280] Specific examples
[0281] Example 1: Travel video editing
[0282] User:
[0283] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[0284] server:
[0285] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[0286] User:
[0287] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[0288] server:
[0289] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[0290] Example 2: Editing a gadget review video
[0291] User:
[0292] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[0293] server:
[0294] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0295] User:
[0296] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[0297] server:
[0298] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[0299] ---
[0300] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[0301] The processing flow will be explained below.
[0302] ---
[0303] Step 1: User uploads a video
[0304] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[0305] Step 2: The server receives and stores the video
[0306] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[0307] Step 3: Extract elements that will make the server buzz
[0308] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[0309] Step 4: The server analyzes the user's video
[0310] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[0311] Step 5: The server will automatically edit the video
[0312] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[0313] Cutting out unnecessary scenes
[0314] Highlighting the highlights
[0315] Adding sound effects and music
[0316] Inserting text and titles
[0317] Automatic thumbnail generation
[0318] Step 6: The server serves the edits
[0319] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[0320] Step 7: User reviews the video and provides feedback
[0321] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[0322] Step 8: The server updates the AI based on the feedback
[0323] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[0324] Step 9: The server starts the emotion engine that recognizes the user's emotions.
[0325] Server: When a user watches a video, the emotion engine is activated and analyzes the user's facial expressions, tone of voice, and reactions in real time. Through the collected data, the server recognizes the user's emotions (happiness, surprise, sadness, excitement, etc.).
[0326] Step 10: The server collects and uses emotional feedback
[0327] Server: Adjusts editing patterns based on the emotional feedback collected by the emotion engine. For example, if a user expresses strong positive emotions in a particular scene, the server emphasizes that scene. Conversely, if the user expresses negative emotions, the server deletes or edits that scene.
[0328] Specific examples
[0329] Example 1: Travel video editing
[0330] User:
[0331] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[0332] server:
[0333] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[0334] User:
[0335] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[0336] server:
[0337] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[0338] Example 2: Editing a gadget review video
[0339] User:
[0340] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[0341] server:
[0342] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0343] User:
[0344] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[0345] server:
[0346] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[0347] ---
[0348] These are the processing steps of an automatic video editing system that uses generative AI combined with an emotion engine.
[0349] Example 2
[0350] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0351] On modern video distribution platforms, videos need to be properly edited and have elements that capture viewers' interest in order to be viewed by a large audience. However, video editing is time-consuming and requires specialized knowledge, placing a heavy burden on many content creators. Furthermore, it is difficult to reflect viewer emotions and feedback in real time, limiting the improvement of video quality. There is a need for a system that can solve these problems and automatically create videos with a high probability of going viral.
[0352] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing a large number of highly viewed videos using generated artificial intelligence and extracting buzzworthy elements; means for receiving and analyzing videos uploaded by users; means for automatically editing the received videos based on the extracted buzzworthy elements; means for providing the edited videos to users and receiving feedback; means for updating the generated artificial intelligence based on the feedback and user emotional data; means for recognizing users' emotions while watching videos in real time and collecting the emotional data; and means for adjusting the video editing pattern using the emotional data. This allows users to automatically create videos that are likely to go viral without requiring specialized knowledge and personalize them in real time based on viewer emotions.
[0353] "Generated artificial intelligence" is a data processing system that is trained on past data and is a model that performs advanced analysis and predictions on new data.
[0354] "Highly viewed videos" refer to video content that has recorded a large number of views on a video distribution platform.
[0355] "Buzz elements" refer to the editing points and content features necessary for a video to be well-received by viewers and widely distributed.
[0356] "User" refers to an individual or corporation that uses this system, uploads videos, and receives editing services from the system.
[0357] "Uploading" refers to the act of a user sending a video file from their device to a server via the Internet.
[0358] "Analysis" refers to the process of examining the content, structure, metadata, etc. of a video in detail and extracting important information.
[0359] "Automatic editing" refers to the process of cutting video, adding music, or other edits based on specific algorithms and rules without user input.
[0360] "Feedback" refers to the action of a user providing opinions or requests for improvement regarding a video edited by the system.
[0361] "Emotion data" is data that expresses, as numerical values or indices, emotions such as joy, sadness, and surprise that a user shows while watching a video.
[0362] "Real-time recognition" refers to the ability to instantly analyze and record the user's emotions at each moment while watching a video.
[0363] "Editing pattern" refers to a series of editing techniques or styles in video editing, such as cutting, rearranging scenes, and adding sound effects.
[0364] "Personalization" is a method of individually optimizing services and content based on individual user preferences and feedback.
[0365] MODE FOR CARRYING OUT THE INVENTION
[0366] Uploading videos
[0367] User:
[0368] Users upload video files from their own devices using a dedicated web interface or application. In this step, they enter a brief description of the video (e.g., "Summer trip to Hawaii") and keywords (e.g., "beach, scenery").
[0369] Receiving and storing videos
[0370] server:
[0371] The server receives video files uploaded by users and temporarily stores them in the system's storage. During this process, the video format and metadata (resolution, format, file size, etc.) are analyzed and recorded in a database. The server uses a high-performance computer system and database system to efficiently analyze and edit video data.
[0372] Extracting buzzworthy elements
[0373] server:
[0374] Using a generative AI model, information such as footage, thumbnails, titles, and keywords is collected and analyzed from a database of videos with high past views. This allows elements for buzz (e.g., emphasis points, editing patterns, use of sound effects, etc.) to be extracted. The generative AI model is trained using deep learning techniques based on a large dataset.
[0375] Analysis and editing of user videos
[0376] server:
[0377] The server analyzes the video uploaded by the user scene by scene and tags important scenes. Next, it performs audio analysis to extract important dialogue and background music characteristics. Based on this, it performs the following automatic editing operations:
[0378] Cutting out unnecessary scenes
[0379] Highlighting the highlights
[0380] Adding sound effects and music
[0381] Inserting text and titles
[0382] Automatic thumbnail generation
[0383] Submitting edits and gathering feedback
[0384] server:
[0385] The completed edited video, along with the suggested thumbnail and title, is generated as a preview link for the user and notified to the user.
[0386] User:
[0387] Users can access the preview link to view the edited video and provide feedback on their satisfaction and corrections, including comments on the editing process and specific requests for improvements.
[0388] Use of emotion engine
[0389] server:
[0390] While the user is watching the preview, the server uses an emotion engine to recognize the user's emotions in real time, and collects and analyzes the data. Emotion data includes happiness, surprise, sadness, excitement, etc., obtained through facial recognition and voice analysis.
[0391] Utilizing Emotional Feedback and Updating AI
[0392] server:
[0393] The server adjusts the editing patterns of the generative AI model based on the emotion data collected by the emotion engine and user feedback, enabling more personalized video editing by emphasizing scenes with strong positive emotions and editing or deleting scenes that show negative emotions.
[0394] User:
[0395] Users can review the final edits and provide additional feedback, which improves the quality of the resulting video.
[0396] Specific examples
[0397] Example 1: Travel video editing
[0398] A user uploads a video taken during a summer trip to Hawaii and enters "Summer trip to Hawaii" and the keywords "beach, scenery" as a brief description.
[0399] The server analyzes the video and audio of travel videos and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generative AI model emphasizes scenic scenes, adds lively music, and automatically uses beautiful beach images as thumbnails.
[0400] The user reviews the edited results and provides emotional feedback on the edits, such as requesting that a particular scene be emphasized if it moved them.
[0401] The server adjusts the edit based on the emotional feedback and delivers the final video, allowing the generative AI model and emotion engine to learn and improve over time.
[0402] Example 2: Editing a gadget review video
[0403] A user uploads a video review of a new gadget and enters the keywords "technology, review, new product."
[0404] The server analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the products. The generative AI model extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0405] Users review the edited video and provide emotional feedback, highlighting scenes where they express surprise at a particular product feature.
[0406] The server adjusts the editing content based on the emotional feedback and provides the final video, thereby improving the accuracy of the generative AI model and emotion engine.
[0407] This system allows users to easily create personalized, high-quality videos that are likely to go viral, even without specialized editing skills, by reflecting viewers' real-time reactions.
[0408] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0409] Step 1: Select and upload your video
[0410] User:
[0411] The user selects a video file from their device and uploads it through the system's dedicated web interface or application. The information entered in this step is the video file (e.g., "Hawaii Trip.mp4"), a video description (e.g., "Summer Hawaii Trip"), and keywords (e.g., "Beach, Scenery"). The server receives this data and prepares it for the next step.
[0412] Step 2: Receive and save the video
[0413] server:
[0414] The server receives video files uploaded by users and temporarily stores them in the system's storage. The input is the video file uploaded by the user and its associated metadata. The server analyzes the metadata, such as the video format, resolution, format, and file size, and records it in a database as output. This step makes it easier to manage and access the videos.
[0415] Step 3: Identifying viral elements
[0416] server:
[0417] The server uses a generative AI model to analyze past videos with high view counts stored in a database. The input is data on past videos with high view counts, and the analysis targets video content, thumbnails, titles, keywords, etc. The server extracts buzz-generating elements from these and updates the generative AI model. The output is patterns of buzz-generating elements (e.g., timing of noteworthy scenes, use of sound effects).
[0418] Step 4: Analyzing user videos
[0419] server:
[0420] The server analyzes videos uploaded by users for each scene and assigns specific tags. The input is the user's video file, and the server divides the scenes using a timeline and performs audio analysis to extract important dialogue and musical features. The output is tagged scene information and extracted audio data.
[0421] Step 5: Auto Edit
[0422] server:
[0423] The server automatically edits the user's video based on the extracted viral elements. The input is the analyzed scene information and viral element patterns. This process involves cutting out unnecessary scenes, emphasizing highlights, adding sound effects and music, inserting text and titles, and automatically generating thumbnails. The output is an edited video file.
[0424] Step 6: Submit your edits and get feedback
[0425] server:
[0426] The server generates a preview link for the completed video, along with the suggested thumbnail and title, and notifies the user. The input is the edited video file, and the output is the preview link.
[0427] User:
[0428] The user accesses the provided preview link to view the edited video. The input is the preview link, and the user provides feedback on satisfaction and corrections. The output is the feedback information sent to the system.
[0429] Step 7: Real-time emotion recognition
[0430] server:
[0431] The server uses an emotion engine to recognize and collect data in real time about the emotions users express while watching videos. The input is the user's viewing data (facial expressions and voice), which the emotion engine analyzes. The output is emotional data such as joy, surprise, sadness, and excitement.
[0432] Step 8: Leverage emotional feedback and update the AI
[0433] server:
[0434] The server adjusts the editing patterns of the generative AI model based on the emotional data collected by the emotion engine and feedback from users. The input is emotional data and feedback, and the server uses this to consider whether to emphasize or delete specific scenes. The output is an updated editing pattern and model.
[0435] Step 9: Final edits and user confirmation
[0436] server:
[0437] The server then re-edits the video based on the revised editing pattern and generates the final video. The input is the updated editing pattern, and the output is the final edited video file.
[0438] User:
[0439] The user then reviews the final edit and provides further feedback if necessary. The input is the final video, and the output is the feedback sent back to the server.
[0440] In this way, by going through a series of processing steps, users can automatically generate videos that are likely to go viral and obtain personalized content based on emotional data.
[0441] (Application example 2)
[0442] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0443] When creating video content, users need to spend time and effort on editing in order to increase the number of views. However, finding effective editing techniques and buzzworthy elements is difficult, placing a significant burden on users without specialized knowledge. Furthermore, while there is a demand for personalized content that reflects the emotions of viewers, there is a lack of means to achieve this. A system that solves these issues and allows anyone to easily create high-quality, buzzworthy videos is needed.
[0444] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0445] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence and extracting buzz-generating elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzz-generating elements, means for providing the edited videos to users and receiving feedback, means for updating the generated artificial intelligence based on the feedback, means for recognizing user emotions in real time using an emotion engine and collecting emotional feedback, and means for adjusting editing patterns based on the emotional feedback. This enables users, even without technical knowledge, to efficiently create and provide videos that reflect viewer emotions and are likely to go viral.
[0446] "Generated artificial intelligence" is artificial intelligence that learns from data and is used to automate or optimize specific tasks.
[0447] A "highly viewed video" is video content that has been viewed by many users in a short period of time.
[0448] "Buzz elements" are characteristic features that make a video more likely to be shared and spread by many viewers.
[0449] "User-uploaded videos" are video files that users submit to the system from their devices.
[0450] "Means for analysis" refers to technologies or processes that provide the functionality to analyze the content of a video and extract specific elements or patterns.
[0451] "Automatic editing methods" refer to technologies and algorithms that use analyzed data to edit videos without human intervention.
[0452] "Means for receiving feedback" is a function for collecting ratings and comments from users.
[0453] "Means of updating" refers to the ability to improve artificial intelligence models and algorithms based on collected feedback.
[0454] The "emotion engine" is a technology that analyzes emotions from a user's facial expressions and voice in real time.
[0455] "Emotional feedback" is data based on the emotional reactions of users when they watch videos.
[0456] "Means for adjusting editing patterns" refers to a function that changes the method and content of video editing based on collected emotional data.
[0457] (Mode for carrying out the invention)
[0458] The system for implementing this invention is mainly composed of a server, a user's device, a generative AI model, and an emotion engine. Each step and its specific implementation method are described below.
[0459] The server receives video files uploaded by users and stores them in storage. The received video files are analyzed for their metadata (resolution, format, file size, etc.) and recorded in a database.
[0460] The server then analyzes a large number of highly viewed videos using a generative AI model. This analysis process extracts elements such as video content, editing patterns, thumbnails, titles, and keywords to identify buzzworthy elements. The generative AI model is trained using past video data and viewer feedback.
[0461] When the server analyzes a user's video, it tags each scene based on the timeline and extracts important lines and background music through audio analysis. Based on the extracted buzzworthy elements, it then cuts out unnecessary scenes, emphasizes highlights, adds sound effects and music, inserts text and titles, and automatically generates thumbnails.
[0462] Once the edits are complete, a preview link is generated and provided to the user, where the user can review the edited video and provide feedback, including requests for corrections and a satisfaction rating.
[0463] Furthermore, the emotion engine analyzes the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data such as joy, surprise, sadness, and excitement, thereby obtaining emotional feedback.
[0464] The server uses this emotional feedback to update the AI and adjust the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This feedback loop results in personalized videos and improved quality.
[0465] The main hardware used includes:
[0466] Video editing server (equipped with high-performance CPU / GPU)
[0467] Smartphone or PC (user interface)
[0468] The main software used includes:
[0469] moviepy (video editing library)
[0470] SentimentAnalyzer (custom model for sentiment analysis)
[0471] VideoEditorAI (generative AI model)
[0472] Specific examples
[0473] Example 1: Travel video editing
[0474] Users upload videos they have taken during their trip. For example, a "summer trip to Hawaii" includes scenes of the beach, scenery, and local culture. The server then selects highlights based on these scenes and adds appropriate music and sound effects. Further editing is performed based on user feedback and emotional feedback to create the final video.
[0475] Prompt Sentence Examples
[0476] "Analyze video files and generate optimal editing patterns based on elements of past viral videos."
[0477] "Improve your video editing process and create personalized edits based on user emotional feedback data."
[0478] This system enables users, even without technical knowledge, to efficiently create and distribute videos that reflect viewers' emotions and are likely to go viral.
[0479] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0480] Step 1:
[0481] The user selects a video file and uploads it through the system's dedicated web interface or application, and can enter a brief description and keywords for the video. The entered video file and metadata are then sent to the server.
[0482] Input: Video file, description, keywords
[0483] Output: Video files and metadata sent to the server
[0484] Step 2:
[0485] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video file's metadata (resolution, format, file size, etc.) and records it in a database.
[0486] Input: Received video files and metadata
[0487] Output: Metadata of video files recorded in a database
[0488] Step 3:
[0489] The server uses a generative AI model to analyze large amounts of video data that has recorded high numbers of views in the past. During this analysis process, elements such as the content of the video, editing patterns, thumbnails, titles, and keywords are extracted to identify buzzworthy elements.
[0490] Input: Data of past videos with high views
[0491] Output: Extracted buzzworthy elements
[0492] Step 4:
[0493] The server analyzes the video files uploaded by users, tagging each scene along the timeline and extracting important lines and background music characteristics through audio analysis.
[0494] Input: A video file uploaded by the user
[0495] Output: Scene-specific tagged data, audio analysis results
[0496] Step 5:
[0497] The server automatically edits the video by cutting out unnecessary scenes and emphasizing highlights based on the extracted buzzworthy elements, adding sound effects and music, inserting text and titles, and automatically generating thumbnails.
[0498] Input: Extracted buzz elements, tagged scene data, audio analysis results
[0499] Output: Edited video file
[0500] Step 6:
[0501] The server generates a preview link containing the edited video file and provides it to the user, allowing the user to view the edited video and provide feedback.
[0502] Input: Edited video file
[0503] Output: Preview link provided to the user
[0504] Step 7:
[0505] Users can access the provided preview link to view the edited video, and provide feedback by requesting corrections or rating their satisfaction on the feedback screen.
[0506] Input: Preview link, user feedback
[0507] Output: Feedback data sent to the server
[0508] Step 8:
[0509] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data, which is then stored as emotional feedback such as happiness, surprise, sadness, and excitement.
[0510] Input: User's facial expression, voice
[0511] Output: Collected emotional feedback data
[0512] Step 9:
[0513] The server updates the generative AI and adjusts the editing patterns based on the collected feedback data and emotional feedback, enhancing the quality of the video by highlighting scenes that show positive emotions and deleting or editing scenes that show negative emotions.
[0514] Input: Feedback data, Emotion feedback data
[0515] Output: Adjusted editing patterns, updated generative AI model
[0516] Step 10:
[0517] Finally, the server provides the user with a re-edited video based on the updated generative AI model, and the user can provide further feedback as needed, which the server then incorporates to refine the final video.
[0518] Input: Re-edited video with updated generative AI model
[0519] Output: The final video file
[0520] Through these processing steps, users can efficiently create and provide personalized videos that are likely to go viral, even without any technical knowledge.
[0521] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0522] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0523] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0524] [Second embodiment]
[0525] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0526] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0527] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0528] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0529] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0530] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0531] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0532] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0533] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0534] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0535] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0536] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0537] ---
[0538] This invention is a system that uses the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: video analysis, element extraction, automatic editing, provision, and updating the AI based on feedback.
[0539] Program processing and explanation
[0540] Uploading videos
[0541] User:
[0542] Users upload video files from their own devices using the system's dedicated web interface or application, and can also enter a brief description and keywords for the video.
[0543] Receiving and storing videos
[0544] server:
[0545] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[0546] Extracting buzzworthy elements
[0547] server:
[0548] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted.
[0549] Analysis and editing of user videos
[0550] server:
[0551] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music features through audio analysis. Then, based on the extracted viral elements, it automatically performs the following edits:
[0552] Automatically cut unnecessary scenes
[0553] Highlighting the highlights
[0554] Adding sound effects and music
[0555] Inserting text and titles
[0556] Automatic thumbnail generation
[0557] Providing edited results
[0558] server:
[0559] To provide users with the completed edited video, a dedicated preview link is generated. This link is then sent to the user, allowing them to view the edited video on the preview screen. Suggested thumbnails and titles are also displayed at the same time.
[0560] Receiving feedback and updating the AI
[0561] User:
[0562] Users can view the edited video via a preview link and provide feedback, including requested corrections and a satisfaction rating.
[0563] server:
[0564] The server receives user feedback and incorporates it into the generation AI as training data. This updates the AI algorithm so that subsequent edits are more suited to the user's preferences. This feedback loop gradually improves the system's accuracy and user satisfaction.
[0565] ---
[0566] Specific examples
[0567] Example 1: Travel video editing
[0568] User:
[0569] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[0570] server:
[0571] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[0572] User:
[0573] Check the edited results and provide feedback such as "I would like the text color in the thumbnail to be changed."
[0574] server:
[0575] The text color is changed based on the user's instructions, and the final video is provided to the user. The AI continues to learn based on the feedback.
[0576] Example 2: Editing a gadget review video
[0577] User:
[0578] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[0579] server:
[0580] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0581] User:
[0582] We will review the edited video and provide feedback if we are satisfied, and if there are any requests for re-editing, we will provide specific instructions on what needs to be corrected.
[0583] server:
[0584] Based on the feedback, we make any necessary corrections and provide the final video. We also update the AI based on the feedback information to improve accuracy.
[0585] ---
[0586] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[0587] The processing flow will be explained below.
[0588] ---
[0589] Step 1: User uploads a video
[0590] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[0591] Step 2: The server receives and stores the video
[0592] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[0593] Step 3: Extract elements that will make the server buzz
[0594] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[0595] Step 4: The server analyzes the user's video
[0596] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[0597] Step 5: The server will automatically edit the video
[0598] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[0599] Cutting out unnecessary scenes
[0600] Highlighting the highlights
[0601] Adding sound effects and music
[0602] Inserting text and titles
[0603] Automatic thumbnail generation
[0604] Step 6: The server serves the edits
[0605] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[0606] Step 7: User reviews the video and provides feedback
[0607] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[0608] Step 8: The server updates the AI based on the feedback
[0609] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[0610] ---
[0611] The above is a specific explanation of the processing steps of an automatic video editing system using generative AI.
[0612] Example 1
[0613] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0614] Conventional video editing systems require users to edit videos manually, which requires a great deal of time and effort. Furthermore, determining whether a video will go viral and determining the optimal editing method requires specialized knowledge, making it difficult for average users to use. Furthermore, the manual improvement process after receiving feedback is inefficient, making it difficult to improve user satisfaction.
[0615] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0616] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated AI and extracting viral elements, means for receiving and saving video files uploaded by users, means for automatically editing the received videos based on the extracted viral elements, means for generating a preview link of the edited results and providing it to the user and receiving feedback, and means for updating the generated AI based on the feedback. This automates the editing work that users would otherwise do manually, enabling optimal editing to increase the likelihood of a video going viral. Furthermore, by updating the AI based on user feedback, editing accuracy can be improved from the next time onwards, thereby continuously increasing user satisfaction.
[0617] 1. "Generated artificial intelligence" refers to algorithms that learn from a large amount of data and then use the results to automatically perform specific tasks.
[0618] 2. "Highly viewed video" means a video that has been viewed by a large number of viewers on an online platform.
[0619] 3. "Buzz elements" refer to the features and factors that increase the number of views of a video, including the video content, editing patterns, thumbnails, titles, keywords, etc.
[0620] 4. "User" refers to an individual or corporation that uses this system to upload videos and receive edited results.
[0621] 5. "Video File" means a digital file containing video and associated audio data.
[0622] 6. "Storage means" refers to a method or device for temporarily or permanently retaining video files in storage.
[0623] 7. "Editing means" means an algorithm or program that automatically processes video, such as cutting out unnecessary scenes, adding sound effects, or inserting text.
[0624] 8. "Preview Link" means a URL that allows users to view the edited video online.
[0625] 9. "Feedback" means any opinions or requests for improvements regarding edits provided by a User.
[0626] 10. "Updating means" means a method or device for correcting or improving the generated AI algorithm based on the feedback received.
[0627] This invention is a system that utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: uploading videos, receiving and saving them, extracting viral elements, analyzing and editing them, providing the editing results, receiving feedback, and updating the AI.
[0628] First, a user uploads a video file from their device through a dedicated web interface or application, and enters the video title, description, and keywords. For example, the title might be "Summer Hawaii Trip" and the keywords "beach, scenery."
[0629] The device sends the selected video file to the system server. The server receives the video file sent from the device and temporarily stores it in the system's storage. It analyzes the video's metadata (resolution, format, file size, etc.) and records that information in a database. For example, it may be analyzed that the video's resolution is 1920x1080 pixels.
[0630] Next, the generative AI model run by the server analyzes a large amount of video data with high past views. Specifically, it references a dataset including the video content, editing patterns, thumbnails, titles, and keywords to extract elements that will create buzz. For example, in a video containing the keywords "beach" and "scenery," it detects the points where many viewers played the video as scenes of waterfronts.
[0631] The server analyzes the videos uploaded by users, tags each scene in a timeline, and extracts important lines and background music through audio analysis.Then, based on the extracted viral elements, it performs the following editing:
[0632] Automatically cut unnecessary scenes
[0633] Highlighting the highlights
[0634] Adding sound effects and music
[0635] Inserting text and titles
[0636] Automatic thumbnail generation
[0637] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[0638] To provide the user with the completed edited video, the server generates a dedicated preview link, which is sent to the user via email or app notification. The preview screen allows the user to view the edited video, along with suggested thumbnails and titles.
[0639] Users can view the edited video through a preview link and provide feedback, including requests for corrections and a satisfaction rating. For example, a user could send feedback such as, "Please change the text color in the thumbnail to blue."
[0640] The server receives user feedback and feeds it into the generation AI as training data. This routine updates the AI algorithm and improves its accuracy so that the next edit will be more tailored to the user's preferences.
[0641] Specific examples
[0642] Example 1: Travel video editing
[0643] The user uploads a video they shot during their trip. They enter a brief description of their trip, "Summer trip to Hawaii," and the keywords "beach, scenery." The server analyzes the video and audio from the travel video and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generation AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails. The user reviews the edited results and provides feedback, such as "I'd like the text color in the thumbnail changed." The server changes the text color based on the instructions and provides the final video to the user. The AI continues to learn based on this feedback.
[0644] Example 2: Editing a gadget review video
[0645] A user uploads a review video of a new gadget. They enter "technology, review, new product" as keywords. The server analyzes the content of the gadget review video and extracts scenes that highlight the product's unique functions and features. The generation AI extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails. The user reviews the edited video and provides feedback if satisfied. If there is a request for re-editing, the user can specify specific corrections. The server makes the necessary corrections based on the feedback and provides the final video. The AI is updated based on the feedback to improve accuracy.
[0646] Prompt Sentence Examples
[0647] 1. "I'm uploading a video of a beach in Hawaii that I took during my trip. Please auto-edit it. The keywords are 'beach' and 'scenery.'"
[0648] 2. "I'm going to upload a video reviewing a new gadget. I need a catchy title and thumbnail. The keywords are 'technology,' 'review,' and 'new product.'"
[0649] This allows users to easily create high-quality videos that are likely to receive many views, reducing the burden on users and increasing the chances of a video's success.
[0650] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0651] Step 1: Upload your video
[0652] User:
[0653] A user uploads a video file from their own device through a dedicated web interface or application. As input, the user specifies the video file, title, description, and keywords. For example, the user enters the title "Summer Hawaii Trip" and the keywords "Beach, Scenery." As output, the specified video file is sent to the system server.
[0654] Step 2: Receive and save the video
[0655] server:
[0656] The server receives video files sent from the device. The inputs are the video files and their metadata (resolution, format, file size, etc.). This is analyzed, and the metadata is temporarily saved in the system's storage along with the video files. The metadata is recorded in a database. For example, the data that the video resolution is 1920x1080 pixels is recorded. The saved video files and metadata are obtained as output.
[0657] Step 3: Identifying viral elements
[0658] server:
[0659] The generative AI model run by the server analyzes large amounts of video data that have recorded high numbers of views in the past. The input is a dataset of videos with high numbers of views in the past. This analysis extracts buzzworthy elements from the dataset, which includes the video content, editing patterns, thumbnails, titles, and keywords. For example, in a video containing the keywords "beach" and "scenery," the points where many viewers played the video are detected as scenes of waterfronts. The output is a list of the extracted buzzworthy elements.
[0660] Step 4: Analyze and edit user videos
[0661] server:
[0662] The server analyzes videos uploaded by users. As input, it receives the uploaded video file and a list of extracted viral elements. It tags each scene in a timeline and extracts important dialogue and background music features through audio analysis. It then performs the following edits and generates an edited video file as output:
[0663] Automatically cut unnecessary scenes
[0664] Highlighting the highlights
[0665] Adding sound effects and music
[0666] Inserting text and titles
[0667] Automatic thumbnail generation
[0668] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[0669] Step 5: Submit your edits
[0670] server:
[0671] The server generates a dedicated preview link to provide the user with the edited video. The inputs are the edited video file and the thumbnail and title suggested by the generation AI. The output is a notification message containing the preview link sent to the user. For example, the email notifying the user includes the message "Check out the preview of your new video" and a link.
[0672] Step 6: Receive feedback and update the AI
[0673] User:
[0674] Users can view the edited video through a preview link and provide feedback. The input is feedback or requests for improvements. For example, a user might send feedback such as, "Please change the text color in the thumbnail to blue." The output is reflected in the system.
[0675] server:
[0676] The server receives feedback from the user and incorporates it into the generation AI as learning data. The input is the user's feedback information, and the AI algorithm is updated based on this. The output is an updated AI algorithm, which improves the accuracy of future edits. Specifically, the text color of the thumbnail is changed to blue and provided to the user again. This feedback information is also saved as reference data for the next edit.
[0677] (Application example 1)
[0678] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0679] In recent years, with the spread of video distribution services, there has been a demand for ways to improve the quality of video content created by individuals and companies and increase the number of views. In particular, it is not easy for individuals to easily upload videos from mobile devices such as smartphones and effectively edit them. Furthermore, there are not enough methods in place to reflect user feedback and use it in editing the next video. Therefore, providing an efficient system for automatically creating videos that are likely to go viral is a challenge.
[0680] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0681] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence to extract buzzworthy elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzzworthy elements, means for providing the edited videos to users as preview links and receiving feedback, and means for updating the generated artificial intelligence based on the feedback, thereby enabling users to efficiently and easily create videos that are likely to go viral and continuously improve the quality through feedback.
[0682] "Generated artificial intelligence" refers to algorithms or systems that are trained to analyze large amounts of video data and learn patterns to edit and extract elements from videos.
[0683] "Buzz elements" are specific visual, audio, editing patterns, and other characteristics that significantly increase the number of views and viewer response of a video.
[0684] "Receiving" is the process in which the system takes in the video file uploaded by the user and temporarily stores it for processing.
[0685] "Analysis" is the process of analyzing the content of a video and identifying its features and patterns, including elements such as video, audio, and scene composition.
[0686] "Editing" is the process of changing, adding, or deleting the content of a video, and includes emphasizing specific scenes, cutting unnecessary scenes, adding sound effects, music, etc.
[0687] The "preview link" is a temporary URL provided to the user so that the user can check the video after editing is complete.
[0688] "Feedback" refers to the opinions and ratings users provide on edited videos, which are used to improve the AI model.
[0689] "Updating" is the process by which the generated AI learns from new data and feedback to improve the accuracy of the next video analysis and editing.
[0690] The "server" is a computer system that receives videos from users, analyzes, edits, stores, and processes feedback.
[0691] This invention relates to a system that uses generated artificial intelligence to automatically edit videos to make them more likely to go viral. This system automatically performs a series of processes, including video analysis, element extraction, automatic editing, provision, and updating the artificial intelligence based on feedback.
[0692] Uploading and saving videos
[0693] Users upload videos using a dedicated application on their smartphones or other devices. The videos are sent to a server and stored immediately upon receipt. Users can also enter a brief description of the video and keywords.
[0694] Extracting and analyzing buzzworthy elements
[0695] The server uses the generated AI to analyze a large number of highly viewed videos and extract elements that will create buzz, including video content, editing patterns, music selection, thumbnails, titles, keywords, etc. Based on this, videos uploaded by users are analyzed.
[0696] Automatic editing function
[0697] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, it performs the following edits:
[0698] Cutting out unnecessary scenes
[0699] Highlighting the highlights
[0700] Adding sound effects and music
[0701] Inserting text and titles
[0702] Automatic thumbnail generation
[0703] Preview and Feedback
[0704] Once the video is edited, it will be provided to the user as a dedicated preview link. Users can use the preview link to check the edited results and provide feedback. For example, users can upload a video of their "summer trip to Hawaii" and edit it to incorporate elements that will create buzz (unique scenery, lively music, catchy title). Users can also upload a review video of a new gadget and edit it to highlight unique product features.
[0705] AI Updates
[0706] The feedback provided by users is stored on a server and then incorporated into the generated AI, which then updates the AI algorithm to ensure that future edits are more accurate and tailored to the user's preferences.
[0707] Details of the hardware and software you will be using
[0708] The system uses servers equipped with high-performance CPUs and GPUs. The main software used is Django (a Python framework), FFmpeg (video processing), and TensorFlow (AI model). Django is used to manage the reception, storage, analysis, editing, and preview link generation of uploaded videos. FFmpeg is used for video analysis and editing, while TensorFlow is used to train and update the generative AI model.
[0709] (Examples of specific examples and prompts)
[0710] Specific examples
[0711] Example 1: A user uploads a video of their "Summer Trip to Hawaii" that they shot during their trip and edits it to emphasize beautiful scenery and upbeat music.
[0712] Example 2: A user uploads a review video for a new gadget and edits it to highlight the product's unique features.
[0713] Prompt Sentence Examples
[0714] Please provide us with an automatically edited video of your "Summer Trip to Hawaii" with elements that will create buzz (unique scenery, lively music, catchy title).
[0715] In this way, users can easily create high-quality videos that are likely to go viral. This system reduces the burden on creators and increases the success rate of videos.
[0716] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0717] Step 1:
[0718] The user starts the application, selects a video file, and uploads it. The user can also enter a description and keywords for the video. This input data (video file and metadata) is sent to the server. The server receives the sent data and temporarily saves it in storage. This process saves the video file and metadata on the server.
[0719] Step 2:
[0720] The server analyzes the metadata of the video file and records it in a database. As a result of the analysis, the video format, resolution, file size, etc. are extracted. This information is stored in the database and used for subsequent processing.
[0721] Step 3:
[0722] The server uses the generative AI model to analyze a large number of highly viewed videos stored in a database and extract viral elements. This extraction process identifies features such as video content, editing patterns, music selection, thumbnails, titles, and keywords. The extracted viral elements are stored as internal data in the AI model.
[0723] Step 4:
[0724] The server uses a generative AI model to analyze videos uploaded by users. It tags each scene along the video's timeline and extracts important lines and background music through audio analysis. Based on the results of this analysis, the next step is automatic editing.
[0725] Step 5:
[0726] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, the following editing is performed:
[0727] Cut Unwanted Scenes: Automatically remove unwanted scenes from your video.
[0728] Highlighting the highlights: Highlight important scenes.
[0729] Add sound effects and music: Add sound effects and background music that suit your video.
[0730] Insert text or title: Insert catchy text or title.
[0731] Auto-generate thumbnails: Automatically generate visually appealing thumbnails.
[0732] Step 6:
[0733] A preview link is generated for the completed edited video and provided to the user. The user can access this preview link through the application and check the edited results. This process generates the preview link and notifies the user.
[0734] Step 7:
[0735] Users can review the edited video and provide feedback, including ratings and specific suggestions for correction, which is then sent to the server.
[0736] Step 8:
[0737] The server receives feedback from users and updates the generative AI model. Based on the feedback data, the AI model's algorithm is adjusted to improve the accuracy of video analysis and editing in future videos. This feedback loop continuously improves the performance of the entire system.
[0738] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0739] ---
[0740] This invention utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral, and also combines it with an emotion engine that recognizes the user's emotions. This system automatically performs a series of processes including video analysis, element extraction, automatic editing, and the collection and utilization of emotional feedback.
[0741] Program processing and explanation
[0742] Uploading videos
[0743] User:
[0744] Users upload video files from their own devices using the system's dedicated web interface or application, and can enter a brief description and keywords for the video.
[0745] Receiving and storing videos
[0746] server:
[0747] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[0748] Extracting buzzworthy elements
[0749] server:
[0750] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted and used to generate patterns.
[0751] Analysis and editing of user videos
[0752] server:
[0753] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music through audio analysis. It then automatically edits the video based on the extracted viral elements, as follows:
[0754] Cutting out unnecessary scenes
[0755] Highlighting the highlights
[0756] Adding sound effects and music
[0757] Inserting text and titles
[0758] Automatic thumbnail generation
[0759] Submitting edits and gathering feedback
[0760] server:
[0761] The edited video, along with the suggested thumbnail and title, is provided to the user via a preview link that is generated for the user. The user is then notified of this link, allowing them to view the edited video on the preview screen.
[0762] User:
[0763] Visit the preview link provided to see the edited video and provide feedback, including any correction requests and a satisfaction rating.
[0764] Use of emotion engine
[0765] server:
[0766] The server uses an emotion engine to recognize the user's emotions in real time while watching videos. The emotion engine analyzes the user's facial expressions, tone of voice, and reactions to collect emotional data such as joy, surprise, sadness, and excitement.
[0767] Utilizing Emotional Feedback and Updating AI
[0768] server:
[0769] Based on the emotional feedback collected by the emotion engine, the generative AI adjusts the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This allows for further personalization.
[0770] User:
[0771] Review the final edit and provide further feedback as needed, including requests for re-edits and additional emotional feedback.
[0772] server:
[0773] Based on the feedback, necessary corrections are made and the final video is provided to the user. The generation AI and emotion engine are updated based on the feedback information to improve accuracy.
[0774] Specific examples
[0775] Example 1: Travel video editing
[0776] User:
[0777] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[0778] server:
[0779] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[0780] User:
[0781] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[0782] server:
[0783] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[0784] Example 2: Editing a gadget review video
[0785] User:
[0786] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[0787] server:
[0788] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0789] User:
[0790] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[0791] server:
[0792] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[0793] ---
[0794] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[0795] The processing flow will be explained below.
[0796] ---
[0797] Step 1: User uploads a video
[0798] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[0799] Step 2: The server receives and stores the video
[0800] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[0801] Step 3: Extract elements that will make the server buzz
[0802] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[0803] Step 4: The server analyzes the user's video
[0804] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[0805] Step 5: The server will automatically edit the video
[0806] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[0807] Cutting out unnecessary scenes
[0808] Highlighting the highlights
[0809] Adding sound effects and music
[0810] Inserting text and titles
[0811] Automatic thumbnail generation
[0812] Step 6: The server serves the edits
[0813] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[0814] Step 7: User reviews the video and provides feedback
[0815] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[0816] Step 8: The server updates the AI based on the feedback
[0817] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[0818] Step 9: The server starts the emotion engine that recognizes the user's emotions.
[0819] Server: When a user watches a video, the emotion engine is activated and analyzes the user's facial expressions, tone of voice, and reactions in real time. Through the collected data, the server recognizes the user's emotions (happiness, surprise, sadness, excitement, etc.).
[0820] Step 10: The server collects and uses emotional feedback
[0821] Server: Adjusts editing patterns based on the emotional feedback collected by the emotion engine. For example, if a user expresses strong positive emotions in a particular scene, the server emphasizes that scene. Conversely, if the user expresses negative emotions, the server deletes or edits that scene.
[0822] Specific examples
[0823] Example 1: Travel video editing
[0824] User:
[0825] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[0826] server:
[0827] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[0828] User:
[0829] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[0830] server:
[0831] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[0832] Example 2: Editing a gadget review video
[0833] User:
[0834] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[0835] server:
[0836] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0837] User:
[0838] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[0839] server:
[0840] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[0841] ---
[0842] These are the processing steps of an automatic video editing system that uses generative AI combined with an emotion engine.
[0843] Example 2
[0844] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0845] On modern video distribution platforms, videos need to be properly edited and have elements that capture viewers' interest in order to be viewed by a large audience. However, video editing is time-consuming and requires specialized knowledge, placing a heavy burden on many content creators. Furthermore, it is difficult to reflect viewer emotions and feedback in real time, limiting the improvement of video quality. There is a need for a system that can solve these problems and automatically create videos with a high probability of going viral.
[0846] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing a large number of highly viewed videos using generated artificial intelligence and extracting buzzworthy elements; means for receiving and analyzing videos uploaded by users; means for automatically editing the received videos based on the extracted buzzworthy elements; means for providing the edited videos to users and receiving feedback; means for updating the generated artificial intelligence based on the feedback and user emotional data; means for recognizing users' emotions while watching videos in real time and collecting the emotional data; and means for adjusting the video editing pattern using the emotional data. This allows users to automatically create videos that are likely to go viral without requiring specialized knowledge and personalize them in real time based on viewer emotions.
[0847] "Generated artificial intelligence" is a data processing system that is trained on past data and is a model that performs advanced analysis and predictions on new data.
[0848] "Highly viewed videos" refer to video content that has recorded a large number of views on a video distribution platform.
[0849] "Buzz elements" refer to the editing points and content features necessary for a video to be well-received by viewers and widely distributed.
[0850] "User" refers to an individual or corporation that uses this system, uploads videos, and receives editing services from the system.
[0851] "Uploading" refers to the act of a user sending a video file from their device to a server via the Internet.
[0852] "Analysis" refers to the process of examining the content, structure, metadata, etc. of a video in detail and extracting important information.
[0853] "Automatic editing" refers to the process of cutting video, adding music, or other edits based on specific algorithms and rules without user input.
[0854] "Feedback" refers to the action of a user providing opinions or requests for improvement regarding a video edited by the system.
[0855] "Emotion data" is data that expresses, as numerical values or indices, emotions such as joy, sadness, and surprise that a user shows while watching a video.
[0856] "Real-time recognition" refers to the ability to instantly analyze and record the user's emotions at each moment while watching a video.
[0857] "Editing pattern" refers to a series of editing techniques or styles in video editing, such as cutting, rearranging scenes, and adding sound effects.
[0858] "Personalization" is a method of individually optimizing services and content based on individual user preferences and feedback.
[0859] MODE FOR CARRYING OUT THE INVENTION
[0860] Uploading videos
[0861] User:
[0862] Users upload video files from their own devices using a dedicated web interface or application. In this step, they enter a brief description of the video (e.g., "Summer trip to Hawaii") and keywords (e.g., "beach, scenery").
[0863] Receiving and storing videos
[0864] server:
[0865] The server receives video files uploaded by users and temporarily stores them in the system's storage. During this process, the video format and metadata (resolution, format, file size, etc.) are analyzed and recorded in a database. The server uses a high-performance computer system and database system to efficiently analyze and edit video data.
[0866] Extracting buzzworthy elements
[0867] server:
[0868] Using a generative AI model, information such as footage, thumbnails, titles, and keywords is collected and analyzed from a database of videos with high past views. This allows elements for buzz (e.g., emphasis points, editing patterns, use of sound effects, etc.) to be extracted. The generative AI model is trained using deep learning techniques based on a large dataset.
[0869] Analysis and editing of user videos
[0870] server:
[0871] The server analyzes the video uploaded by the user scene by scene and tags important scenes. Next, it performs audio analysis to extract important dialogue and background music characteristics. Based on this, it performs the following automatic editing operations:
[0872] Cutting out unnecessary scenes
[0873] Highlighting the highlights
[0874] Adding sound effects and music
[0875] Inserting text and titles
[0876] Automatic thumbnail generation
[0877] Submitting edits and gathering feedback
[0878] server:
[0879] The completed edited video, along with the suggested thumbnail and title, is generated as a preview link for the user and notified to the user.
[0880] User:
[0881] Users can access the preview link to view the edited video and provide feedback on their satisfaction and corrections, including comments on the editing process and specific requests for improvements.
[0882] Use of emotion engine
[0883] server:
[0884] While the user is watching the preview, the server uses an emotion engine to recognize the user's emotions in real time, and collects and analyzes the data. Emotion data includes happiness, surprise, sadness, excitement, etc., obtained through facial recognition and voice analysis.
[0885] Utilizing Emotional Feedback and Updating AI
[0886] server:
[0887] The server adjusts the editing patterns of the generative AI model based on the emotion data collected by the emotion engine and user feedback, enabling more personalized video editing by emphasizing scenes with strong positive emotions and editing or deleting scenes that show negative emotions.
[0888] User:
[0889] Users can review the final edits and provide additional feedback, which improves the quality of the resulting video.
[0890] Specific examples
[0891] Example 1: Travel video editing
[0892] A user uploads a video taken during a summer trip to Hawaii and enters "Summer trip to Hawaii" and the keywords "beach, scenery" as a brief description.
[0893] The server analyzes the video and audio of travel videos and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generative AI model emphasizes scenic scenes, adds lively music, and automatically uses beautiful beach images as thumbnails.
[0894] The user reviews the edited results and provides emotional feedback on the edits, such as requesting that a particular scene be emphasized if it moved them.
[0895] The server adjusts the edit based on the emotional feedback and delivers the final video, allowing the generative AI model and emotion engine to learn and improve over time.
[0896] Example 2: Editing a gadget review video
[0897] A user uploads a video review of a new gadget and enters the keywords "technology, review, new product."
[0898] The server analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the products. The generative AI model extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[0899] Users review the edited video and provide emotional feedback, highlighting scenes where they express surprise at a particular product feature.
[0900] The server adjusts the editing content based on the emotional feedback and provides the final video, thereby improving the accuracy of the generative AI model and emotion engine.
[0901] This system allows users to easily create personalized, high-quality videos that are likely to go viral, even without specialized editing skills, by reflecting viewers' real-time reactions.
[0902] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0903] Step 1: Select and upload your video
[0904] User:
[0905] The user selects a video file from their device and uploads it through the system's dedicated web interface or application. The information entered in this step is the video file (e.g., "Hawaii Trip.mp4"), a video description (e.g., "Summer Hawaii Trip"), and keywords (e.g., "Beach, Scenery"). The server receives this data and prepares it for the next step.
[0906] Step 2: Receive and save the video
[0907] server:
[0908] The server receives video files uploaded by users and temporarily stores them in the system's storage. The input is the video file uploaded by the user and its associated metadata. The server analyzes the metadata, such as the video format, resolution, format, and file size, and records it in a database as output. This step makes it easier to manage and access the videos.
[0909] Step 3: Identifying viral elements
[0910] server:
[0911] The server uses a generative AI model to analyze past videos with high view counts stored in a database. The input is data on past videos with high view counts, and the analysis targets video content, thumbnails, titles, keywords, etc. The server extracts buzz-generating elements from these and updates the generative AI model. The output is patterns of buzz-generating elements (e.g., timing of noteworthy scenes, use of sound effects).
[0912] Step 4: Analyzing user videos
[0913] server:
[0914] The server analyzes videos uploaded by users for each scene and assigns specific tags. The input is the user's video file, and the server divides the scenes using a timeline and performs audio analysis to extract important dialogue and musical features. The output is tagged scene information and extracted audio data.
[0915] Step 5: Auto Edit
[0916] server:
[0917] The server automatically edits the user's video based on the extracted viral elements. The input is the analyzed scene information and viral element patterns. This process involves cutting out unnecessary scenes, emphasizing highlights, adding sound effects and music, inserting text and titles, and automatically generating thumbnails. The output is an edited video file.
[0918] Step 6: Submit your edits and get feedback
[0919] server:
[0920] The server generates a preview link for the completed video, along with the suggested thumbnail and title, and notifies the user. The input is the edited video file, and the output is the preview link.
[0921] User:
[0922] The user accesses the provided preview link to view the edited video. The input is the preview link, and the user provides feedback on satisfaction and corrections. The output is the feedback information sent to the system.
[0923] Step 7: Real-time emotion recognition
[0924] server:
[0925] The server uses an emotion engine to recognize and collect data in real time about the emotions users express while watching videos. The input is the user's viewing data (facial expressions and voice), which the emotion engine analyzes. The output is emotional data such as joy, surprise, sadness, and excitement.
[0926] Step 8: Leverage emotional feedback and update the AI
[0927] server:
[0928] The server adjusts the editing patterns of the generative AI model based on the emotional data collected by the emotion engine and feedback from users. The input is emotional data and feedback, and the server uses this to consider whether to emphasize or delete specific scenes. The output is an updated editing pattern and model.
[0929] Step 9: Final edits and user confirmation
[0930] server:
[0931] The server then re-edits the video based on the revised editing pattern and generates the final video. The input is the updated editing pattern, and the output is the final edited video file.
[0932] User:
[0933] The user then reviews the final edit and provides further feedback if necessary. The input is the final video, and the output is the feedback sent back to the server.
[0934] In this way, by going through a series of processing steps, users can automatically generate videos that are likely to go viral and obtain personalized content based on emotional data.
[0935] (Application example 2)
[0936] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0937] When creating video content, users need to spend time and effort on editing in order to increase the number of views. However, finding effective editing techniques and buzzworthy elements is difficult, placing a significant burden on users without specialized knowledge. Furthermore, while there is a demand for personalized content that reflects the emotions of viewers, there is a lack of means to achieve this. A system that solves these issues and allows anyone to easily create high-quality, buzzworthy videos is needed.
[0938] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0939] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence and extracting buzz-generating elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzz-generating elements, means for providing the edited videos to users and receiving feedback, means for updating the generated artificial intelligence based on the feedback, means for recognizing user emotions in real time using an emotion engine and collecting emotional feedback, and means for adjusting editing patterns based on the emotional feedback. This enables users, even without technical knowledge, to efficiently create and provide videos that reflect viewer emotions and are likely to go viral.
[0940] "Generated artificial intelligence" is artificial intelligence that learns from data and is used to automate or optimize specific tasks.
[0941] A "highly viewed video" is video content that has been viewed by many users in a short period of time.
[0942] "Buzz elements" are characteristic features that make a video more likely to be shared and spread by many viewers.
[0943] "User-uploaded videos" are video files that users submit to the system from their devices.
[0944] "Means for analysis" refers to technologies or processes that provide the functionality to analyze the content of a video and extract specific elements or patterns.
[0945] "Automatic editing methods" refer to technologies and algorithms that use analyzed data to edit videos without human intervention.
[0946] "Means for receiving feedback" is a function for collecting ratings and comments from users.
[0947] "Means of updating" refers to the ability to improve artificial intelligence models and algorithms based on collected feedback.
[0948] The "emotion engine" is a technology that analyzes emotions from a user's facial expressions and voice in real time.
[0949] "Emotional feedback" is data based on the emotional reactions of users when they watch videos.
[0950] "Means for adjusting editing patterns" refers to a function that changes the method and content of video editing based on collected emotional data.
[0951] (Mode for carrying out the invention)
[0952] The system for implementing this invention is mainly composed of a server, a user's device, a generative AI model, and an emotion engine. Each step and its specific implementation method are described below.
[0953] The server receives video files uploaded by users and stores them in storage. The received video files are analyzed for their metadata (resolution, format, file size, etc.) and recorded in a database.
[0954] The server then analyzes a large number of highly viewed videos using a generative AI model. This analysis process extracts elements such as video content, editing patterns, thumbnails, titles, and keywords to identify buzzworthy elements. The generative AI model is trained using past video data and viewer feedback.
[0955] When the server analyzes a user's video, it tags each scene based on the timeline and extracts important lines and background music through audio analysis. Based on the extracted buzzworthy elements, it then cuts out unnecessary scenes, emphasizes highlights, adds sound effects and music, inserts text and titles, and automatically generates thumbnails.
[0956] Once the edits are complete, a preview link is generated and provided to the user, where the user can review the edited video and provide feedback, including requests for corrections and a satisfaction rating.
[0957] Furthermore, the emotion engine analyzes the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data such as joy, surprise, sadness, and excitement, thereby obtaining emotional feedback.
[0958] The server uses this emotional feedback to update the AI and adjust the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This feedback loop results in personalized videos and improved quality.
[0959] The main hardware used includes:
[0960] Video editing server (equipped with high-performance CPU / GPU)
[0961] Smartphone or PC (user interface)
[0962] The main software used includes:
[0963] moviepy (video editing library)
[0964] SentimentAnalyzer (custom model for sentiment analysis)
[0965] VideoEditorAI (generative AI model)
[0966] Specific examples
[0967] Example 1: Travel video editing
[0968] Users upload videos they have taken during their trip. For example, a "summer trip to Hawaii" includes scenes of the beach, scenery, and local culture. The server then selects highlights based on these scenes and adds appropriate music and sound effects. Further editing is performed based on user feedback and emotional feedback to create the final video.
[0969] Prompt Sentence Examples
[0970] "Analyze video files and generate optimal editing patterns based on elements of past viral videos."
[0971] "Improve your video editing process and create personalized edits based on user emotional feedback data."
[0972] This system enables users, even without technical knowledge, to efficiently create and distribute videos that reflect viewers' emotions and are likely to go viral.
[0973] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0974] Step 1:
[0975] The user selects a video file and uploads it through the system's dedicated web interface or application, and can enter a brief description and keywords for the video. The entered video file and metadata are then sent to the server.
[0976] Input: Video file, description, keywords
[0977] Output: Video files and metadata sent to the server
[0978] Step 2:
[0979] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video file's metadata (resolution, format, file size, etc.) and records it in a database.
[0980] Input: Received video files and metadata
[0981] Output: Metadata of video files recorded in a database
[0982] Step 3:
[0983] The server uses a generative AI model to analyze large amounts of video data that has recorded high numbers of views in the past. During this analysis process, elements such as the content of the video, editing patterns, thumbnails, titles, and keywords are extracted to identify buzzworthy elements.
[0984] Input: Data of past videos with high views
[0985] Output: Extracted buzzworthy elements
[0986] Step 4:
[0987] The server analyzes the video files uploaded by users, tagging each scene along the timeline and extracting important lines and background music characteristics through audio analysis.
[0988] Input: A video file uploaded by the user
[0989] Output: Scene-specific tagged data, audio analysis results
[0990] Step 5:
[0991] The server automatically edits the video by cutting out unnecessary scenes and emphasizing highlights based on the extracted buzzworthy elements, adding sound effects and music, inserting text and titles, and automatically generating thumbnails.
[0992] Input: Extracted buzz elements, tagged scene data, audio analysis results
[0993] Output: Edited video file
[0994] Step 6:
[0995] The server generates a preview link containing the edited video file and provides it to the user, allowing the user to view the edited video and provide feedback.
[0996] Input: Edited video file
[0997] Output: Preview link provided to the user
[0998] Step 7:
[0999] Users can access the provided preview link to view the edited video, and provide feedback by requesting corrections or rating their satisfaction on the feedback screen.
[1000] Input: Preview link, user feedback
[1001] Output: Feedback data sent to the server
[1002] Step 8:
[1003] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data, which is then stored as emotional feedback such as happiness, surprise, sadness, and excitement.
[1004] Input: User's facial expression, voice
[1005] Output: Collected emotional feedback data
[1006] Step 9:
[1007] The server updates the generative AI and adjusts the editing patterns based on the collected feedback data and emotional feedback, enhancing the quality of the video by highlighting scenes that show positive emotions and deleting or editing scenes that show negative emotions.
[1008] Input: Feedback data, Emotion feedback data
[1009] Output: Adjusted editing patterns, updated generative AI model
[1010] Step 10:
[1011] Finally, the server provides the user with a re-edited video based on the updated generative AI model, and the user can provide further feedback as needed, which the server then incorporates to refine the final video.
[1012] Input: Re-edited video with updated generative AI model
[1013] Output: The final video file
[1014] Through these processing steps, users can efficiently create and provide personalized videos that are likely to go viral, even without any technical knowledge.
[1015] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1016] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1017] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1018] [Third embodiment]
[1019] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1020] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1021] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1022] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1023] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1024] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1025] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1026] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1027] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1028] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1029] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1030] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1031] ---
[1032] This invention is a system that uses the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: video analysis, element extraction, automatic editing, provision, and updating the AI based on feedback.
[1033] Program processing and explanation
[1034] Uploading videos
[1035] User:
[1036] Users upload video files from their own devices using the system's dedicated web interface or application, and can also enter a brief description and keywords for the video.
[1037] Receiving and storing videos
[1038] server:
[1039] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[1040] Extracting buzzworthy elements
[1041] server:
[1042] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted.
[1043] Analysis and editing of user videos
[1044] server:
[1045] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music features through audio analysis. Then, based on the extracted viral elements, it automatically performs the following edits:
[1046] Automatically cut unnecessary scenes
[1047] Highlighting the highlights
[1048] Adding sound effects and music
[1049] Inserting text and titles
[1050] Automatic thumbnail generation
[1051] Providing edited results
[1052] server:
[1053] To provide users with the completed edited video, a dedicated preview link is generated. This link is then sent to the user, allowing them to view the edited video on the preview screen. Suggested thumbnails and titles are also displayed at the same time.
[1054] Receiving feedback and updating the AI
[1055] User:
[1056] Users can view the edited video via a preview link and provide feedback, including requested corrections and a satisfaction rating.
[1057] server:
[1058] The server receives user feedback and incorporates it into the generation AI as training data. This updates the AI algorithm so that subsequent edits are more suited to the user's preferences. This feedback loop gradually improves the system's accuracy and user satisfaction.
[1059] ---
[1060] Specific examples
[1061] Example 1: Travel video editing
[1062] User:
[1063] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[1064] server:
[1065] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[1066] User:
[1067] Check the edited results and provide feedback such as "I would like the text color in the thumbnail to be changed."
[1068] server:
[1069] The text color is changed based on the user's instructions, and the final video is provided to the user. The AI continues to learn based on the feedback.
[1070] Example 2: Editing a gadget review video
[1071] User:
[1072] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[1073] server:
[1074] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1075] User:
[1076] We will review the edited video and provide feedback if we are satisfied, and if there are any requests for re-editing, we will provide specific instructions on what needs to be corrected.
[1077] server:
[1078] Based on the feedback, we make any necessary corrections and provide the final video. We also update the AI based on the feedback information to improve accuracy.
[1079] ---
[1080] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[1081] The processing flow will be explained below.
[1082] ---
[1083] Step 1: User uploads a video
[1084] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[1085] Step 2: The server receives and stores the video
[1086] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[1087] Step 3: Extract elements that will make the server buzz
[1088] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[1089] Step 4: The server analyzes the user's video
[1090] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[1091] Step 5: The server will automatically edit the video
[1092] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[1093] Cutting out unnecessary scenes
[1094] Highlighting the highlights
[1095] Adding sound effects and music
[1096] Inserting text and titles
[1097] Automatic thumbnail generation
[1098] Step 6: The server serves the edits
[1099] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[1100] Step 7: User reviews the video and provides feedback
[1101] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[1102] Step 8: The server updates the AI based on the feedback
[1103] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[1104] ---
[1105] The above is a specific explanation of the processing steps of an automatic video editing system using generative AI.
[1106] Example 1
[1107] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1108] Conventional video editing systems require users to edit videos manually, which requires a great deal of time and effort. Furthermore, determining whether a video will go viral and determining the optimal editing method requires specialized knowledge, making it difficult for average users to use. Furthermore, the manual improvement process after receiving feedback is inefficient, making it difficult to improve user satisfaction.
[1109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1110] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated AI and extracting viral elements, means for receiving and saving video files uploaded by users, means for automatically editing the received videos based on the extracted viral elements, means for generating a preview link of the edited results and providing it to the user and receiving feedback, and means for updating the generated AI based on the feedback. This automates the editing work that users would otherwise do manually, enabling optimal editing to increase the likelihood of a video going viral. Furthermore, by updating the AI based on user feedback, editing accuracy can be improved from the next time onwards, thereby continuously increasing user satisfaction.
[1111] 1. "Generated artificial intelligence" refers to algorithms that learn from a large amount of data and then use the results to automatically perform specific tasks.
[1112] 2. "Highly viewed video" means a video that has been viewed by a large number of viewers on an online platform.
[1113] 3. "Buzz elements" refer to the features and factors that increase the number of views of a video, including the video content, editing patterns, thumbnails, titles, keywords, etc.
[1114] 4. "User" refers to an individual or corporation that uses this system to upload videos and receive edited results.
[1115] 5. "Video File" means a digital file containing video and associated audio data.
[1116] 6. "Storage means" refers to a method or device for temporarily or permanently retaining video files in storage.
[1117] 7. "Editing means" means an algorithm or program that automatically processes video, such as cutting out unnecessary scenes, adding sound effects, or inserting text.
[1118] 8. "Preview Link" means a URL that allows users to view the edited video online.
[1119] 9. "Feedback" means any opinions or requests for improvements regarding edits provided by a User.
[1120] 10. "Updating means" means a method or device for correcting or improving the generated AI algorithm based on the feedback received.
[1121] This invention is a system that utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: uploading videos, receiving and saving them, extracting viral elements, analyzing and editing them, providing the editing results, receiving feedback, and updating the AI.
[1122] First, a user uploads a video file from their device through a dedicated web interface or application, and enters the video title, description, and keywords. For example, the title might be "Summer Hawaii Trip" and the keywords "beach, scenery."
[1123] The device sends the selected video file to the system server. The server receives the video file sent from the device and temporarily stores it in the system's storage. It analyzes the video's metadata (resolution, format, file size, etc.) and records that information in a database. For example, it may be analyzed that the video's resolution is 1920x1080 pixels.
[1124] Next, the generative AI model run by the server analyzes a large amount of video data with high past views. Specifically, it references a dataset including the video content, editing patterns, thumbnails, titles, and keywords to extract elements that will create buzz. For example, in a video containing the keywords "beach" and "scenery," it detects the points where many viewers played the video as scenes of waterfronts.
[1125] The server analyzes the videos uploaded by users, tags each scene in a timeline, and extracts important lines and background music through audio analysis.Then, based on the extracted viral elements, it performs the following editing:
[1126] Automatically cut unnecessary scenes
[1127] Highlighting the highlights
[1128] Adding sound effects and music
[1129] Inserting text and titles
[1130] Automatic thumbnail generation
[1131] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[1132] To provide the user with the completed edited video, the server generates a dedicated preview link, which is sent to the user via email or app notification. The preview screen allows the user to view the edited video, along with suggested thumbnails and titles.
[1133] Users can view the edited video through a preview link and provide feedback, including requests for corrections and a satisfaction rating. For example, a user could send feedback such as, "Please change the text color in the thumbnail to blue."
[1134] The server receives user feedback and feeds it into the generation AI as training data. This routine updates the AI algorithm and improves its accuracy so that the next edit will be more tailored to the user's preferences.
[1135] Specific examples
[1136] Example 1: Travel video editing
[1137] The user uploads a video they shot during their trip. They enter a brief description of their trip, "Summer trip to Hawaii," and the keywords "beach, scenery." The server analyzes the video and audio from the travel video and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generation AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails. The user reviews the edited results and provides feedback, such as "I'd like the text color in the thumbnail changed." The server changes the text color based on the instructions and provides the final video to the user. The AI continues to learn based on this feedback.
[1138] Example 2: Editing a gadget review video
[1139] A user uploads a review video of a new gadget. They enter "technology, review, new product" as keywords. The server analyzes the content of the gadget review video and extracts scenes that highlight the product's unique functions and features. The generation AI extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails. The user reviews the edited video and provides feedback if satisfied. If there is a request for re-editing, the user can specify specific corrections. The server makes the necessary corrections based on the feedback and provides the final video. The AI is updated based on the feedback to improve accuracy.
[1140] Prompt Sentence Examples
[1141] 1. "I'm uploading a video of a beach in Hawaii that I took during my trip. Please auto-edit it. The keywords are 'beach' and 'scenery.'"
[1142] 2. "I'm going to upload a video reviewing a new gadget. I need a catchy title and thumbnail. The keywords are 'technology,' 'review,' and 'new product.'"
[1143] This allows users to easily create high-quality videos that are likely to receive many views, reducing the burden on users and increasing the chances of a video's success.
[1144] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1145] Step 1: Upload your video
[1146] User:
[1147] A user uploads a video file from their own device through a dedicated web interface or application. As input, the user specifies the video file, title, description, and keywords. For example, the user enters the title "Summer Hawaii Trip" and the keywords "Beach, Scenery." As output, the specified video file is sent to the system server.
[1148] Step 2: Receive and save the video
[1149] server:
[1150] The server receives video files sent from the device. The inputs are the video files and their metadata (resolution, format, file size, etc.). This is analyzed, and the metadata is temporarily saved in the system's storage along with the video files. The metadata is recorded in a database. For example, the data that the video resolution is 1920x1080 pixels is recorded. The saved video files and metadata are obtained as output.
[1151] Step 3: Identifying viral elements
[1152] server:
[1153] The generative AI model run by the server analyzes large amounts of video data that have recorded high numbers of views in the past. The input is a dataset of videos with high numbers of views in the past. This analysis extracts buzzworthy elements from the dataset, which includes the video content, editing patterns, thumbnails, titles, and keywords. For example, in a video containing the keywords "beach" and "scenery," the points where many viewers played the video are detected as scenes of waterfronts. The output is a list of the extracted buzzworthy elements.
[1154] Step 4: Analyze and edit user videos
[1155] server:
[1156] The server analyzes videos uploaded by users. As input, it receives the uploaded video file and a list of extracted viral elements. It tags each scene in a timeline and extracts important dialogue and background music features through audio analysis. It then performs the following edits and generates an edited video file as output:
[1157] Automatically cut unnecessary scenes
[1158] Highlighting the highlights
[1159] Adding sound effects and music
[1160] Inserting text and titles
[1161] Automatic thumbnail generation
[1162] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[1163] Step 5: Submit your edits
[1164] server:
[1165] The server generates a dedicated preview link to provide the user with the edited video. The inputs are the edited video file and the thumbnail and title suggested by the generation AI. The output is a notification message containing the preview link sent to the user. For example, the email notifying the user includes the message "Check out the preview of your new video" and a link.
[1166] Step 6: Receive feedback and update the AI
[1167] User:
[1168] Users can view the edited video through a preview link and provide feedback. The input is feedback or requests for improvements. For example, a user might send feedback such as, "Please change the text color in the thumbnail to blue." The output is reflected in the system.
[1169] server:
[1170] The server receives feedback from the user and incorporates it into the generation AI as learning data. The input is the user's feedback information, and the AI algorithm is updated based on this. The output is an updated AI algorithm, which improves the accuracy of future edits. Specifically, the text color of the thumbnail is changed to blue and provided to the user again. This feedback information is also saved as reference data for the next edit.
[1171] (Application example 1)
[1172] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1173] In recent years, with the spread of video distribution services, there has been a demand for ways to improve the quality of video content created by individuals and companies and increase the number of views. In particular, it is not easy for individuals to easily upload videos from mobile devices such as smartphones and effectively edit them. Furthermore, there are not enough methods in place to reflect user feedback and use it in editing the next video. Therefore, providing an efficient system for automatically creating videos that are likely to go viral is a challenge.
[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1175] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence to extract buzzworthy elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzzworthy elements, means for providing the edited videos to users as preview links and receiving feedback, and means for updating the generated artificial intelligence based on the feedback, thereby enabling users to efficiently and easily create videos that are likely to go viral and continuously improve the quality through feedback.
[1176] "Generated artificial intelligence" refers to algorithms or systems that are trained to analyze large amounts of video data and learn patterns to edit and extract elements from videos.
[1177] "Buzz elements" are specific visual, audio, editing patterns, and other characteristics that significantly increase the number of views and viewer response of a video.
[1178] "Receiving" is the process in which the system takes in the video file uploaded by the user and temporarily stores it for processing.
[1179] "Analysis" is the process of analyzing the content of a video and identifying its features and patterns, including elements such as video, audio, and scene composition.
[1180] "Editing" is the process of changing, adding, or deleting the content of a video, and includes emphasizing specific scenes, cutting unnecessary scenes, adding sound effects, music, etc.
[1181] The "preview link" is a temporary URL provided to the user so that the user can check the video after editing is complete.
[1182] "Feedback" refers to the opinions and ratings users provide on edited videos, which are used to improve the AI model.
[1183] "Updating" is the process by which the generated AI learns from new data and feedback to improve the accuracy of the next video analysis and editing.
[1184] The "server" is a computer system that receives videos from users, analyzes, edits, stores, and processes feedback.
[1185] This invention relates to a system that uses generated artificial intelligence to automatically edit videos to make them more likely to go viral. This system automatically performs a series of processes, including video analysis, element extraction, automatic editing, provision, and updating the artificial intelligence based on feedback.
[1186] Uploading and saving videos
[1187] Users upload videos using a dedicated application on their smartphones or other devices. The videos are sent to a server and stored immediately upon receipt. Users can also enter a brief description of the video and keywords.
[1188] Extracting and analyzing buzzworthy elements
[1189] The server uses the generated AI to analyze a large number of highly viewed videos and extract elements that will create buzz, including video content, editing patterns, music selection, thumbnails, titles, keywords, etc. Based on this, videos uploaded by users are analyzed.
[1190] Automatic editing function
[1191] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, it performs the following edits:
[1192] Cutting out unnecessary scenes
[1193] Highlighting the highlights
[1194] Adding sound effects and music
[1195] Inserting text and titles
[1196] Automatic thumbnail generation
[1197] Preview and Feedback
[1198] Once the video is edited, it will be provided to the user as a dedicated preview link. Users can use the preview link to check the edited results and provide feedback. For example, users can upload a video of their "summer trip to Hawaii" and edit it to incorporate elements that will create buzz (unique scenery, lively music, catchy title). Users can also upload a review video of a new gadget and edit it to highlight unique product features.
[1199] AI Updates
[1200] The feedback provided by users is stored on a server and then incorporated into the generated AI, which then updates the AI algorithm to ensure that future edits are more accurate and tailored to the user's preferences.
[1201] Details of the hardware and software you will be using
[1202] The system uses servers equipped with high-performance CPUs and GPUs. The main software used is Django (a Python framework), FFmpeg (video processing), and TensorFlow (AI model). Django is used to manage the reception, storage, analysis, editing, and preview link generation of uploaded videos. FFmpeg is used for video analysis and editing, while TensorFlow is used to train and update the generative AI model.
[1203] (Examples of specific examples and prompts)
[1204] Specific examples
[1205] Example 1: A user uploads a video of their "Summer Trip to Hawaii" that they shot during their trip and edits it to emphasize beautiful scenery and upbeat music.
[1206] Example 2: A user uploads a review video for a new gadget and edits it to highlight the product's unique features.
[1207] Prompt Sentence Examples
[1208] Please provide us with an automatically edited video of your "Summer Trip to Hawaii" with elements that will create buzz (unique scenery, lively music, catchy title).
[1209] In this way, users can easily create high-quality videos that are likely to go viral. This system reduces the burden on creators and increases the success rate of videos.
[1210] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1211] Step 1:
[1212] The user starts the application, selects a video file, and uploads it. The user can also enter a description and keywords for the video. This input data (video file and metadata) is sent to the server. The server receives the sent data and temporarily saves it in storage. This process saves the video file and metadata on the server.
[1213] Step 2:
[1214] The server analyzes the metadata of the video file and records it in a database. As a result of the analysis, the video format, resolution, file size, etc. are extracted. This information is stored in the database and used for subsequent processing.
[1215] Step 3:
[1216] The server uses the generative AI model to analyze a large number of highly viewed videos stored in a database and extract viral elements. This extraction process identifies features such as video content, editing patterns, music selection, thumbnails, titles, and keywords. The extracted viral elements are stored as internal data in the AI model.
[1217] Step 4:
[1218] The server uses a generative AI model to analyze videos uploaded by users. It tags each scene along the video's timeline and extracts important lines and background music through audio analysis. Based on the results of this analysis, the next step is automatic editing.
[1219] Step 5:
[1220] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, the following editing is performed:
[1221] Cut Unwanted Scenes: Automatically remove unwanted scenes from your video.
[1222] Highlighting the highlights: Highlight important scenes.
[1223] Add sound effects and music: Add sound effects and background music that suit your video.
[1224] Insert text or title: Insert catchy text or title.
[1225] Auto-generate thumbnails: Automatically generate visually appealing thumbnails.
[1226] Step 6:
[1227] A preview link is generated for the completed edited video and provided to the user. The user can access this preview link through the application and check the edited results. This process generates the preview link and notifies the user.
[1228] Step 7:
[1229] Users can review the edited video and provide feedback, including ratings and specific suggestions for correction, which is then sent to the server.
[1230] Step 8:
[1231] The server receives feedback from users and updates the generative AI model. Based on the feedback data, the AI model's algorithm is adjusted to improve the accuracy of video analysis and editing in future videos. This feedback loop continuously improves the performance of the entire system.
[1232] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1233] ---
[1234] This invention utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral, and also combines it with an emotion engine that recognizes the user's emotions. This system automatically performs a series of processes including video analysis, element extraction, automatic editing, and the collection and utilization of emotional feedback.
[1235] Program processing and explanation
[1236] Uploading videos
[1237] User:
[1238] Users upload video files from their own devices using the system's dedicated web interface or application, and can enter a brief description and keywords for the video.
[1239] Receiving and storing videos
[1240] server:
[1241] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[1242] Extracting buzzworthy elements
[1243] server:
[1244] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted and used to generate patterns.
[1245] Analysis and editing of user videos
[1246] server:
[1247] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music through audio analysis. It then automatically edits the video based on the extracted viral elements, as follows:
[1248] Cutting out unnecessary scenes
[1249] Highlighting the highlights
[1250] Adding sound effects and music
[1251] Inserting text and titles
[1252] Automatic thumbnail generation
[1253] Submitting edits and gathering feedback
[1254] server:
[1255] The edited video, along with the suggested thumbnail and title, is provided to the user via a preview link that is generated for the user. The user is then notified of this link, allowing them to view the edited video on the preview screen.
[1256] User:
[1257] Visit the preview link provided to see the edited video and provide feedback, including any correction requests and a satisfaction rating.
[1258] Use of emotion engine
[1259] server:
[1260] The server uses an emotion engine to recognize the user's emotions in real time while watching videos. The emotion engine analyzes the user's facial expressions, tone of voice, and reactions to collect emotional data such as joy, surprise, sadness, and excitement.
[1261] Utilizing Emotional Feedback and Updating AI
[1262] server:
[1263] Based on the emotional feedback collected by the emotion engine, the generative AI adjusts the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This allows for further personalization.
[1264] User:
[1265] Review the final edit and provide further feedback as needed, including requests for re-edits and additional emotional feedback.
[1266] server:
[1267] Based on the feedback, necessary corrections are made and the final video is provided to the user. The generation AI and emotion engine are updated based on the feedback information to improve accuracy.
[1268] Specific examples
[1269] Example 1: Travel video editing
[1270] User:
[1271] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[1272] server:
[1273] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[1274] User:
[1275] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[1276] server:
[1277] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[1278] Example 2: Editing a gadget review video
[1279] User:
[1280] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[1281] server:
[1282] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1283] User:
[1284] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[1285] server:
[1286] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[1287] ---
[1288] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[1289] The processing flow will be explained below.
[1290] ---
[1291] Step 1: User uploads a video
[1292] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[1293] Step 2: The server receives and stores the video
[1294] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[1295] Step 3: Extract elements that will make the server buzz
[1296] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[1297] Step 4: The server analyzes the user's video
[1298] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[1299] Step 5: The server will automatically edit the video
[1300] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[1301] Cutting out unnecessary scenes
[1302] Highlighting the highlights
[1303] Adding sound effects and music
[1304] Inserting text and titles
[1305] Automatic thumbnail generation
[1306] Step 6: The server serves the edits
[1307] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[1308] Step 7: User reviews the video and provides feedback
[1309] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[1310] Step 8: The server updates the AI based on the feedback
[1311] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[1312] Step 9: The server starts the emotion engine that recognizes the user's emotions.
[1313] Server: When a user watches a video, the emotion engine is activated and analyzes the user's facial expressions, tone of voice, and reactions in real time. Through the collected data, the server recognizes the user's emotions (happiness, surprise, sadness, excitement, etc.).
[1314] Step 10: The server collects and uses emotional feedback
[1315] Server: Adjusts editing patterns based on the emotional feedback collected by the emotion engine. For example, if a user expresses strong positive emotions in a particular scene, the server emphasizes that scene. Conversely, if the user expresses negative emotions, the server deletes or edits that scene.
[1316] Specific examples
[1317] Example 1: Travel video editing
[1318] User:
[1319] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[1320] server:
[1321] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[1322] User:
[1323] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[1324] server:
[1325] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[1326] Example 2: Editing a gadget review video
[1327] User:
[1328] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[1329] server:
[1330] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1331] User:
[1332] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[1333] server:
[1334] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[1335] ---
[1336] These are the processing steps of an automatic video editing system that uses generative AI combined with an emotion engine.
[1337] Example 2
[1338] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1339] On modern video distribution platforms, videos need to be properly edited and have elements that capture viewers' interest in order to be viewed by a large audience. However, video editing is time-consuming and requires specialized knowledge, placing a heavy burden on many content creators. Furthermore, it is difficult to reflect viewer emotions and feedback in real time, limiting the improvement of video quality. There is a need for a system that can solve these problems and automatically create videos with a high probability of going viral.
[1340] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing a large number of highly viewed videos using generated artificial intelligence and extracting buzzworthy elements; means for receiving and analyzing videos uploaded by users; means for automatically editing the received videos based on the extracted buzzworthy elements; means for providing the edited videos to users and receiving feedback; means for updating the generated artificial intelligence based on the feedback and user emotional data; means for recognizing users' emotions while watching videos in real time and collecting the emotional data; and means for adjusting the video editing pattern using the emotional data. This allows users to automatically create videos that are likely to go viral without requiring specialized knowledge and personalize them in real time based on viewer emotions.
[1341] "Generated artificial intelligence" is a data processing system that is trained on past data and is a model that performs advanced analysis and predictions on new data.
[1342] "Highly viewed videos" refer to video content that has recorded a large number of views on a video distribution platform.
[1343] "Buzz elements" refer to the editing points and content features necessary for a video to be well-received by viewers and widely distributed.
[1344] "User" refers to an individual or corporation that uses this system, uploads videos, and receives editing services from the system.
[1345] "Uploading" refers to the act of a user sending a video file from their device to a server via the Internet.
[1346] "Analysis" refers to the process of examining the content, structure, metadata, etc. of a video in detail and extracting important information.
[1347] "Automatic editing" refers to the process of cutting video, adding music, or other edits based on specific algorithms and rules without user input.
[1348] "Feedback" refers to the action of a user providing opinions or requests for improvement regarding a video edited by the system.
[1349] "Emotion data" is data that expresses, as numerical values or indices, emotions such as joy, sadness, and surprise that a user shows while watching a video.
[1350] "Real-time recognition" refers to the ability to instantly analyze and record the user's emotions at each moment while watching a video.
[1351] "Editing pattern" refers to a series of editing techniques or styles in video editing, such as cutting, rearranging scenes, and adding sound effects.
[1352] "Personalization" is a method of individually optimizing services and content based on individual user preferences and feedback.
[1353] MODE FOR CARRYING OUT THE INVENTION
[1354] Uploading videos
[1355] User:
[1356] Users upload video files from their own devices using a dedicated web interface or application. In this step, they enter a brief description of the video (e.g., "Summer trip to Hawaii") and keywords (e.g., "beach, scenery").
[1357] Receiving and storing videos
[1358] server:
[1359] The server receives video files uploaded by users and temporarily stores them in the system's storage. During this process, the video format and metadata (resolution, format, file size, etc.) are analyzed and recorded in a database. The server uses a high-performance computer system and database system to efficiently analyze and edit video data.
[1360] Extracting buzzworthy elements
[1361] server:
[1362] Using a generative AI model, information such as footage, thumbnails, titles, and keywords is collected and analyzed from a database of videos with high past views. This allows elements for buzz (e.g., emphasis points, editing patterns, use of sound effects, etc.) to be extracted. The generative AI model is trained using deep learning techniques based on a large dataset.
[1363] Analysis and editing of user videos
[1364] server:
[1365] The server analyzes the video uploaded by the user scene by scene and tags important scenes. Next, it performs audio analysis to extract important dialogue and background music characteristics. Based on this, it performs the following automatic editing operations:
[1366] Cutting out unnecessary scenes
[1367] Highlighting the highlights
[1368] Adding sound effects and music
[1369] Inserting text and titles
[1370] Automatic thumbnail generation
[1371] Submitting edits and gathering feedback
[1372] server:
[1373] The completed edited video, along with the suggested thumbnail and title, is generated as a preview link for the user and notified to the user.
[1374] User:
[1375] Users can access the preview link to view the edited video and provide feedback on their satisfaction and corrections, including comments on the editing process and specific requests for improvements.
[1376] Use of emotion engine
[1377] server:
[1378] While the user is watching the preview, the server uses an emotion engine to recognize the user's emotions in real time, and collects and analyzes the data. Emotion data includes happiness, surprise, sadness, excitement, etc., obtained through facial recognition and voice analysis.
[1379] Utilizing Emotional Feedback and Updating AI
[1380] server:
[1381] The server adjusts the editing patterns of the generative AI model based on the emotion data collected by the emotion engine and user feedback, enabling more personalized video editing by emphasizing scenes with strong positive emotions and editing or deleting scenes that show negative emotions.
[1382] User:
[1383] Users can review the final edits and provide additional feedback, which improves the quality of the resulting video.
[1384] Specific examples
[1385] Example 1: Travel video editing
[1386] A user uploads a video taken during a summer trip to Hawaii and enters "Summer trip to Hawaii" and the keywords "beach, scenery" as a brief description.
[1387] The server analyzes the video and audio of travel videos and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generative AI model emphasizes scenic scenes, adds lively music, and automatically uses beautiful beach images as thumbnails.
[1388] The user reviews the edited results and provides emotional feedback on the edits, such as requesting that a particular scene be emphasized if it moved them.
[1389] The server adjusts the edit based on the emotional feedback and delivers the final video, allowing the generative AI model and emotion engine to learn and improve over time.
[1390] Example 2: Editing a gadget review video
[1391] A user uploads a video review of a new gadget and enters the keywords "technology, review, new product."
[1392] The server analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the products. The generative AI model extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1393] Users review the edited video and provide emotional feedback, highlighting scenes where they express surprise at a particular product feature.
[1394] The server adjusts the editing content based on the emotional feedback and provides the final video, thereby improving the accuracy of the generative AI model and emotion engine.
[1395] This system allows users to easily create personalized, high-quality videos that are likely to go viral, even without specialized editing skills, by reflecting viewers' real-time reactions.
[1396] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1397] Step 1: Select and upload your video
[1398] User:
[1399] The user selects a video file from their device and uploads it through the system's dedicated web interface or application. The information entered in this step is the video file (e.g., "Hawaii Trip.mp4"), a video description (e.g., "Summer Hawaii Trip"), and keywords (e.g., "Beach, Scenery"). The server receives this data and prepares it for the next step.
[1400] Step 2: Receive and save the video
[1401] server:
[1402] The server receives video files uploaded by users and temporarily stores them in the system's storage. The input is the video file uploaded by the user and its associated metadata. The server analyzes the metadata, such as the video format, resolution, format, and file size, and records it in a database as output. This step makes it easier to manage and access the videos.
[1403] Step 3: Identifying viral elements
[1404] server:
[1405] The server uses a generative AI model to analyze past videos with high view counts stored in a database. The input is data on past videos with high view counts, and the analysis targets video content, thumbnails, titles, keywords, etc. The server extracts buzz-generating elements from these and updates the generative AI model. The output is patterns of buzz-generating elements (e.g., timing of noteworthy scenes, use of sound effects).
[1406] Step 4: Analyzing user videos
[1407] server:
[1408] The server analyzes videos uploaded by users for each scene and assigns specific tags. The input is the user's video file, and the server divides the scenes using a timeline and performs audio analysis to extract important dialogue and musical features. The output is tagged scene information and extracted audio data.
[1409] Step 5: Auto Edit
[1410] server:
[1411] The server automatically edits the user's video based on the extracted viral elements. The input is the analyzed scene information and viral element patterns. This process involves cutting out unnecessary scenes, emphasizing highlights, adding sound effects and music, inserting text and titles, and automatically generating thumbnails. The output is an edited video file.
[1412] Step 6: Submit your edits and get feedback
[1413] server:
[1414] The server generates a preview link for the completed video, along with the suggested thumbnail and title, and notifies the user. The input is the edited video file, and the output is the preview link.
[1415] User:
[1416] The user accesses the provided preview link to view the edited video. The input is the preview link, and the user provides feedback on satisfaction and corrections. The output is the feedback information sent to the system.
[1417] Step 7: Real-time emotion recognition
[1418] server:
[1419] The server uses an emotion engine to recognize and collect data in real time about the emotions users express while watching videos. The input is the user's viewing data (facial expressions and voice), which the emotion engine analyzes. The output is emotional data such as joy, surprise, sadness, and excitement.
[1420] Step 8: Leverage emotional feedback and update the AI
[1421] server:
[1422] The server adjusts the editing patterns of the generative AI model based on the emotional data collected by the emotion engine and feedback from users. The input is emotional data and feedback, and the server uses this to consider whether to emphasize or delete specific scenes. The output is an updated editing pattern and model.
[1423] Step 9: Final edits and user confirmation
[1424] server:
[1425] The server then re-edits the video based on the revised editing pattern and generates the final video. The input is the updated editing pattern, and the output is the final edited video file.
[1426] User:
[1427] The user then reviews the final edit and provides further feedback if necessary. The input is the final video, and the output is the feedback sent back to the server.
[1428] In this way, by going through a series of processing steps, users can automatically generate videos that are likely to go viral and obtain personalized content based on emotional data.
[1429] (Application example 2)
[1430] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1431] When creating video content, users need to spend time and effort on editing in order to increase the number of views. However, finding effective editing techniques and buzzworthy elements is difficult, placing a significant burden on users without specialized knowledge. Furthermore, while there is a demand for personalized content that reflects the emotions of viewers, there is a lack of means to achieve this. A system that solves these issues and allows anyone to easily create high-quality, buzzworthy videos is needed.
[1432] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1433] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence and extracting buzz-generating elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzz-generating elements, means for providing the edited videos to users and receiving feedback, means for updating the generated artificial intelligence based on the feedback, means for recognizing user emotions in real time using an emotion engine and collecting emotional feedback, and means for adjusting editing patterns based on the emotional feedback. This enables users, even without technical knowledge, to efficiently create and provide videos that reflect viewer emotions and are likely to go viral.
[1434] "Generated artificial intelligence" is artificial intelligence that learns from data and is used to automate or optimize specific tasks.
[1435] A "highly viewed video" is video content that has been viewed by many users in a short period of time.
[1436] "Buzz elements" are characteristic features that make a video more likely to be shared and spread by many viewers.
[1437] "User-uploaded videos" are video files that users submit to the system from their devices.
[1438] "Means for analysis" refers to technologies or processes that provide the functionality to analyze the content of a video and extract specific elements or patterns.
[1439] "Automatic editing methods" refer to technologies and algorithms that use analyzed data to edit videos without human intervention.
[1440] "Means for receiving feedback" is a function for collecting ratings and comments from users.
[1441] "Means of updating" refers to the ability to improve artificial intelligence models and algorithms based on collected feedback.
[1442] The "emotion engine" is a technology that analyzes emotions from a user's facial expressions and voice in real time.
[1443] "Emotional feedback" is data based on the emotional reactions of users when they watch videos.
[1444] "Means for adjusting editing patterns" refers to a function that changes the method and content of video editing based on collected emotional data.
[1445] (Mode for carrying out the invention)
[1446] The system for implementing this invention is mainly composed of a server, a user's device, a generative AI model, and an emotion engine. Each step and its specific implementation method are described below.
[1447] The server receives video files uploaded by users and stores them in storage. The received video files are analyzed for their metadata (resolution, format, file size, etc.) and recorded in a database.
[1448] The server then analyzes a large number of highly viewed videos using a generative AI model. This analysis process extracts elements such as video content, editing patterns, thumbnails, titles, and keywords to identify buzzworthy elements. The generative AI model is trained using past video data and viewer feedback.
[1449] When the server analyzes a user's video, it tags each scene based on the timeline and extracts important lines and background music through audio analysis. Based on the extracted buzzworthy elements, it then cuts out unnecessary scenes, emphasizes highlights, adds sound effects and music, inserts text and titles, and automatically generates thumbnails.
[1450] Once the edits are complete, a preview link is generated and provided to the user, where the user can review the edited video and provide feedback, including requests for corrections and a satisfaction rating.
[1451] Furthermore, the emotion engine analyzes the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data such as joy, surprise, sadness, and excitement, thereby obtaining emotional feedback.
[1452] The server uses this emotional feedback to update the AI and adjust the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This feedback loop results in personalized videos and improved quality.
[1453] The main hardware used includes:
[1454] Video editing server (equipped with high-performance CPU / GPU)
[1455] Smartphone or PC (user interface)
[1456] The main software used includes:
[1457] moviepy (video editing library)
[1458] SentimentAnalyzer (custom model for sentiment analysis)
[1459] VideoEditorAI (generative AI model)
[1460] Specific examples
[1461] Example 1: Travel video editing
[1462] Users upload videos they have taken during their trip. For example, a "summer trip to Hawaii" includes scenes of the beach, scenery, and local culture. The server then selects highlights based on these scenes and adds appropriate music and sound effects. Further editing is performed based on user feedback and emotional feedback to create the final video.
[1463] Prompt Sentence Examples
[1464] "Analyze video files and generate optimal editing patterns based on elements of past viral videos."
[1465] "Improve your video editing process and create personalized edits based on user emotional feedback data."
[1466] This system enables users, even without technical knowledge, to efficiently create and distribute videos that reflect viewers' emotions and are likely to go viral.
[1467] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1468] Step 1:
[1469] The user selects a video file and uploads it through the system's dedicated web interface or application, and can enter a brief description and keywords for the video. The entered video file and metadata are then sent to the server.
[1470] Input: Video file, description, keywords
[1471] Output: Video files and metadata sent to the server
[1472] Step 2:
[1473] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video file's metadata (resolution, format, file size, etc.) and records it in a database.
[1474] Input: Received video files and metadata
[1475] Output: Metadata of video files recorded in a database
[1476] Step 3:
[1477] The server uses a generative AI model to analyze large amounts of video data that has recorded high numbers of views in the past. During this analysis process, elements such as the content of the video, editing patterns, thumbnails, titles, and keywords are extracted to identify buzzworthy elements.
[1478] Input: Data of past videos with high views
[1479] Output: Extracted buzzworthy elements
[1480] Step 4:
[1481] The server analyzes the video files uploaded by users, tagging each scene along the timeline and extracting important lines and background music characteristics through audio analysis.
[1482] Input: A video file uploaded by the user
[1483] Output: Scene-specific tagged data, audio analysis results
[1484] Step 5:
[1485] The server automatically edits the video by cutting out unnecessary scenes and emphasizing highlights based on the extracted buzzworthy elements, adding sound effects and music, inserting text and titles, and automatically generating thumbnails.
[1486] Input: Extracted buzz elements, tagged scene data, audio analysis results
[1487] Output: Edited video file
[1488] Step 6:
[1489] The server generates a preview link containing the edited video file and provides it to the user, allowing the user to view the edited video and provide feedback.
[1490] Input: Edited video file
[1491] Output: Preview link provided to the user
[1492] Step 7:
[1493] Users can access the provided preview link to view the edited video, and provide feedback by requesting corrections or rating their satisfaction on the feedback screen.
[1494] Input: Preview link, user feedback
[1495] Output: Feedback data sent to the server
[1496] Step 8:
[1497] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data, which is then stored as emotional feedback such as happiness, surprise, sadness, and excitement.
[1498] Input: User's facial expression, voice
[1499] Output: Collected emotional feedback data
[1500] Step 9:
[1501] The server updates the generative AI and adjusts the editing patterns based on the collected feedback data and emotional feedback, enhancing the quality of the video by highlighting scenes that show positive emotions and deleting or editing scenes that show negative emotions.
[1502] Input: Feedback data, Emotion feedback data
[1503] Output: Adjusted editing patterns, updated generative AI model
[1504] Step 10:
[1505] Finally, the server provides the user with a re-edited video based on the updated generative AI model, and the user can provide further feedback as needed, which the server then incorporates to refine the final video.
[1506] Input: Re-edited video with updated generative AI model
[1507] Output: The final video file
[1508] Through these processing steps, users can efficiently create and provide personalized videos that are likely to go viral, even without any technical knowledge.
[1509] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1510] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1511] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1512] [Fourth embodiment]
[1513] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1514] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1515] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1516] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1517] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1518] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1519] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1520] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1521] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1522] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1523] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1524] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1525] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1526] ---
[1527] This invention is a system that uses the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: video analysis, element extraction, automatic editing, provision, and updating the AI based on feedback.
[1528] Program processing and explanation
[1529] Uploading videos
[1530] User:
[1531] Users upload video files from their own devices using the system's dedicated web interface or application, and can also enter a brief description and keywords for the video.
[1532] Receiving and storing videos
[1533] server:
[1534] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[1535] Extracting buzzworthy elements
[1536] server:
[1537] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted.
[1538] Analysis and editing of user videos
[1539] server:
[1540] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music features through audio analysis. Then, based on the extracted viral elements, it automatically performs the following edits:
[1541] Automatically cut unnecessary scenes
[1542] Highlighting the highlights
[1543] Adding sound effects and music
[1544] Inserting text and titles
[1545] Automatic thumbnail generation
[1546] Providing edited results
[1547] server:
[1548] To provide users with the completed edited video, a dedicated preview link is generated. This link is then sent to the user, allowing them to view the edited video on the preview screen. Suggested thumbnails and titles are also displayed at the same time.
[1549] Receiving feedback and updating the AI
[1550] User:
[1551] Users can view the edited video via a preview link and provide feedback, including requested corrections and a satisfaction rating.
[1552] server:
[1553] The server receives user feedback and incorporates it into the generation AI as training data. This updates the AI algorithm so that subsequent edits are more suited to the user's preferences. This feedback loop gradually improves the system's accuracy and user satisfaction.
[1554] ---
[1555] Specific examples
[1556] Example 1: Travel video editing
[1557] User:
[1558] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[1559] server:
[1560] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[1561] User:
[1562] Check the edited results and provide feedback such as "I would like the text color in the thumbnail to be changed."
[1563] server:
[1564] The text color is changed based on the user's instructions, and the final video is provided to the user. The AI continues to learn based on the feedback.
[1565] Example 2: Editing a gadget review video
[1566] User:
[1567] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[1568] server:
[1569] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1570] User:
[1571] We will review the edited video and provide feedback if we are satisfied, and if there are any requests for re-editing, we will provide specific instructions on what needs to be corrected.
[1572] server:
[1573] Based on the feedback, we make any necessary corrections and provide the final video. We also update the AI based on the feedback information to improve accuracy.
[1574] ---
[1575] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[1576] The processing flow will be explained below.
[1577] ---
[1578] Step 1: User uploads a video
[1579] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[1580] Step 2: The server receives and stores the video
[1581] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[1582] Step 3: Extract elements that will make the server buzz
[1583] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[1584] Step 4: The server analyzes the user's video
[1585] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[1586] Step 5: The server will automatically edit the video
[1587] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[1588] Cutting out unnecessary scenes
[1589] Highlighting the highlights
[1590] Adding sound effects and music
[1591] Inserting text and titles
[1592] Automatic thumbnail generation
[1593] Step 6: The server serves the edits
[1594] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[1595] Step 7: User reviews the video and provides feedback
[1596] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[1597] Step 8: The server updates the AI based on the feedback
[1598] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[1599] ---
[1600] The above is a specific explanation of the processing steps of an automatic video editing system using generative AI.
[1601] Example 1
[1602] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1603] Conventional video editing systems require users to edit videos manually, which requires a great deal of time and effort. Furthermore, determining whether a video will go viral and determining the optimal editing method requires specialized knowledge, making it difficult for average users to use. Furthermore, the manual improvement process after receiving feedback is inefficient, making it difficult to improve user satisfaction.
[1604] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1605] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated AI and extracting viral elements, means for receiving and saving video files uploaded by users, means for automatically editing the received videos based on the extracted viral elements, means for generating a preview link of the edited results and providing it to the user and receiving feedback, and means for updating the generated AI based on the feedback. This automates the editing work that users would otherwise do manually, enabling optimal editing to increase the likelihood of a video going viral. Furthermore, by updating the AI based on user feedback, editing accuracy can be improved from the next time onwards, thereby continuously increasing user satisfaction.
[1606] 1. "Generated artificial intelligence" refers to algorithms that learn from a large amount of data and then use the results to automatically perform specific tasks.
[1607] 2. "Highly viewed video" means a video that has been viewed by a large number of viewers on an online platform.
[1608] 3. "Buzz elements" refer to the features and factors that increase the number of views of a video, including the video content, editing patterns, thumbnails, titles, keywords, etc.
[1609] 4. "User" refers to an individual or corporation that uses this system to upload videos and receive edited results.
[1610] 5. "Video File" means a digital file containing video and associated audio data.
[1611] 6. "Storage means" refers to a method or device for temporarily or permanently retaining video files in storage.
[1612] 7. "Editing means" means an algorithm or program that automatically processes video, such as cutting out unnecessary scenes, adding sound effects, or inserting text.
[1613] 8. "Preview Link" means a URL that allows users to view the edited video online.
[1614] 9. "Feedback" means any opinions or requests for improvements regarding edits provided by a User.
[1615] 10. "Updating means" means a method or device for correcting or improving the generated AI algorithm based on the feedback received.
[1616] This invention is a system that utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral. This system automatically performs a series of processes: uploading videos, receiving and saving them, extracting viral elements, analyzing and editing them, providing the editing results, receiving feedback, and updating the AI.
[1617] First, a user uploads a video file from their device through a dedicated web interface or application, and enters the video title, description, and keywords. For example, the title might be "Summer Hawaii Trip" and the keywords "beach, scenery."
[1618] The device sends the selected video file to the system server. The server receives the video file sent from the device and temporarily stores it in the system's storage. It analyzes the video's metadata (resolution, format, file size, etc.) and records that information in a database. For example, it may be analyzed that the video's resolution is 1920x1080 pixels.
[1619] Next, the generative AI model run by the server analyzes a large amount of video data with high past views. Specifically, it references a dataset including the video content, editing patterns, thumbnails, titles, and keywords to extract elements that will create buzz. For example, in a video containing the keywords "beach" and "scenery," it detects the points where many viewers played the video as scenes of waterfronts.
[1620] The server analyzes the videos uploaded by users, tags each scene in a timeline, and extracts important lines and background music through audio analysis.Then, based on the extracted viral elements, it performs the following editing:
[1621] Automatically cut unnecessary scenes
[1622] Highlighting the highlights
[1623] Adding sound effects and music
[1624] Inserting text and titles
[1625] Automatic thumbnail generation
[1626] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[1627] To provide the user with the completed edited video, the server generates a dedicated preview link, which is sent to the user via email or app notification. The preview screen allows the user to view the edited video, along with suggested thumbnails and titles.
[1628] Users can view the edited video through a preview link and provide feedback, including requests for corrections and a satisfaction rating. For example, a user could send feedback such as, "Please change the text color in the thumbnail to blue."
[1629] The server receives user feedback and feeds it into the generation AI as training data. This routine updates the AI algorithm and improves its accuracy so that the next edit will be more tailored to the user's preferences.
[1630] Specific examples
[1631] Example 1: Travel video editing
[1632] The user uploads a video they shot during their trip. They enter a brief description of their trip, "Summer trip to Hawaii," and the keywords "beach, scenery." The server analyzes the video and audio from the travel video and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generation AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails. The user reviews the edited results and provides feedback, such as "I'd like the text color in the thumbnail changed." The server changes the text color based on the instructions and provides the final video to the user. The AI continues to learn based on this feedback.
[1633] Example 2: Editing a gadget review video
[1634] A user uploads a review video of a new gadget. They enter "technology, review, new product" as keywords. The server analyzes the content of the gadget review video and extracts scenes that highlight the product's unique functions and features. The generation AI extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails. The user reviews the edited video and provides feedback if satisfied. If there is a request for re-editing, the user can specify specific corrections. The server makes the necessary corrections based on the feedback and provides the final video. The AI is updated based on the feedback to improve accuracy.
[1635] Prompt Sentence Examples
[1636] 1. "I'm uploading a video of a beach in Hawaii that I took during my trip. Please auto-edit it. The keywords are 'beach' and 'scenery.'"
[1637] 2. "I'm going to upload a video reviewing a new gadget. I need a catchy title and thumbnail. The keywords are 'technology,' 'review,' and 'new product.'"
[1638] This allows users to easily create high-quality videos that are likely to receive many views, reducing the burden on users and increasing the chances of a video's success.
[1639] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1640] Step 1: Upload your video
[1641] User:
[1642] A user uploads a video file from their own device through a dedicated web interface or application. As input, the user specifies the video file, title, description, and keywords. For example, the user enters the title "Summer Hawaii Trip" and the keywords "Beach, Scenery." As output, the specified video file is sent to the system server.
[1643] Step 2: Receive and save the video
[1644] server:
[1645] The server receives video files sent from the device. The inputs are the video files and their metadata (resolution, format, file size, etc.). This is analyzed, and the metadata is temporarily saved in the system's storage along with the video files. The metadata is recorded in a database. For example, the data that the video resolution is 1920x1080 pixels is recorded. The saved video files and metadata are obtained as output.
[1646] Step 3: Identifying viral elements
[1647] server:
[1648] The generative AI model run by the server analyzes large amounts of video data that have recorded high numbers of views in the past. The input is a dataset of videos with high numbers of views in the past. This analysis extracts buzzworthy elements from the dataset, which includes the video content, editing patterns, thumbnails, titles, and keywords. For example, in a video containing the keywords "beach" and "scenery," the points where many viewers played the video are detected as scenes of waterfronts. The output is a list of the extracted buzzworthy elements.
[1649] Step 4: Analyze and edit user videos
[1650] server:
[1651] The server analyzes videos uploaded by users. As input, it receives the uploaded video file and a list of extracted viral elements. It tags each scene in a timeline and extracts important dialogue and background music features through audio analysis. It then performs the following edits and generates an edited video file as output:
[1652] Automatically cut unnecessary scenes
[1653] Highlighting the highlights
[1654] Adding sound effects and music
[1655] Inserting text and titles
[1656] Automatic thumbnail generation
[1657] For example, if a scene with waves breaking on the beach attracts a large number of viewers, the video can highlight that scene and add upbeat music.
[1658] Step 5: Submit your edits
[1659] server:
[1660] The server generates a dedicated preview link to provide the user with the edited video. The inputs are the edited video file and the thumbnail and title suggested by the generation AI. The output is a notification message containing the preview link sent to the user. For example, the email notifying the user includes the message "Check out the preview of your new video" and a link.
[1661] Step 6: Receive feedback and update the AI
[1662] User:
[1663] Users can view the edited video through a preview link and provide feedback. The input is feedback or requests for improvements. For example, a user might send feedback such as, "Please change the text color in the thumbnail to blue." The output is reflected in the system.
[1664] server:
[1665] The server receives feedback from the user and incorporates it into the generation AI as learning data. The input is the user's feedback information, and the AI algorithm is updated based on this. The output is an updated AI algorithm, which improves the accuracy of future edits. Specifically, the text color of the thumbnail is changed to blue and provided to the user again. This feedback information is also saved as reference data for the next edit.
[1666] (Application example 1)
[1667] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1668] In recent years, with the spread of video distribution services, there has been a demand for ways to improve the quality of video content created by individuals and companies and increase the number of views. In particular, it is not easy for individuals to easily upload videos from mobile devices such as smartphones and effectively edit them. Furthermore, there are not enough methods in place to reflect user feedback and use it in editing the next video. Therefore, providing an efficient system for automatically creating videos that are likely to go viral is a challenge.
[1669] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1670] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence to extract buzzworthy elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzzworthy elements, means for providing the edited videos to users as preview links and receiving feedback, and means for updating the generated artificial intelligence based on the feedback, thereby enabling users to efficiently and easily create videos that are likely to go viral and continuously improve the quality through feedback.
[1671] "Generated artificial intelligence" refers to algorithms or systems that are trained to analyze large amounts of video data and learn patterns to edit and extract elements from videos.
[1672] "Buzz elements" are specific visual, audio, editing patterns, and other characteristics that significantly increase the number of views and viewer response of a video.
[1673] "Receiving" is the process in which the system takes in the video file uploaded by the user and temporarily stores it for processing.
[1674] "Analysis" is the process of analyzing the content of a video and identifying its features and patterns, including elements such as video, audio, and scene composition.
[1675] "Editing" is the process of changing, adding, or deleting the content of a video, and includes emphasizing specific scenes, cutting unnecessary scenes, adding sound effects, music, etc.
[1676] The "preview link" is a temporary URL provided to the user so that the user can check the video after editing is complete.
[1677] "Feedback" refers to the opinions and ratings users provide on edited videos, which are used to improve the AI model.
[1678] "Updating" is the process by which the generated AI learns from new data and feedback to improve the accuracy of the next video analysis and editing.
[1679] The "server" is a computer system that receives videos from users, analyzes, edits, stores, and processes feedback.
[1680] This invention relates to a system that uses generated artificial intelligence to automatically edit videos to make them more likely to go viral. This system automatically performs a series of processes, including video analysis, element extraction, automatic editing, provision, and updating the artificial intelligence based on feedback.
[1681] Uploading and saving videos
[1682] Users upload videos using a dedicated application on their smartphones or other devices. The videos are sent to a server and stored immediately upon receipt. Users can also enter a brief description of the video and keywords.
[1683] Extracting and analyzing buzzworthy elements
[1684] The server uses the generated AI to analyze a large number of highly viewed videos and extract elements that will create buzz, including video content, editing patterns, music selection, thumbnails, titles, keywords, etc. Based on this, videos uploaded by users are analyzed.
[1685] Automatic editing function
[1686] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, it performs the following edits:
[1687] Cutting out unnecessary scenes
[1688] Highlighting the highlights
[1689] Adding sound effects and music
[1690] Inserting text and titles
[1691] Automatic thumbnail generation
[1692] Preview and Feedback
[1693] Once the video is edited, it will be provided to the user as a dedicated preview link. Users can use the preview link to check the edited results and provide feedback. For example, users can upload a video of their "summer trip to Hawaii" and edit it to incorporate elements that will create buzz (unique scenery, lively music, catchy title). Users can also upload a review video of a new gadget and edit it to highlight unique product features.
[1694] AI Updates
[1695] The feedback provided by users is stored on a server and then incorporated into the generated AI, which then updates the AI algorithm to ensure that future edits are more accurate and tailored to the user's preferences.
[1696] Details of the hardware and software you will be using
[1697] The system uses servers equipped with high-performance CPUs and GPUs. The main software used is Django (a Python framework), FFmpeg (video processing), and TensorFlow (AI model). Django is used to manage the reception, storage, analysis, editing, and preview link generation of uploaded videos. FFmpeg is used for video analysis and editing, while TensorFlow is used to train and update the generative AI model.
[1698] (Examples of specific examples and prompts)
[1699] Specific examples
[1700] Example 1: A user uploads a video of their "Summer Trip to Hawaii" that they shot during their trip and edits it to emphasize beautiful scenery and upbeat music.
[1701] Example 2: A user uploads a review video for a new gadget and edits it to highlight the product's unique features.
[1702] Prompt Sentence Examples
[1703] Please provide us with an automatically edited video of your "Summer Trip to Hawaii" with elements that will create buzz (unique scenery, lively music, catchy title).
[1704] In this way, users can easily create high-quality videos that are likely to go viral. This system reduces the burden on creators and increases the success rate of videos.
[1705] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1706] Step 1:
[1707] The user starts the application, selects a video file, and uploads it. The user can also enter a description and keywords for the video. This input data (video file and metadata) is sent to the server. The server receives the sent data and temporarily saves it in storage. This process saves the video file and metadata on the server.
[1708] Step 2:
[1709] The server analyzes the metadata of the video file and records it in a database. As a result of the analysis, the video format, resolution, file size, etc. are extracted. This information is stored in the database and used for subsequent processing.
[1710] Step 3:
[1711] The server uses the generative AI model to analyze a large number of highly viewed videos stored in a database and extract viral elements. This extraction process identifies features such as video content, editing patterns, music selection, thumbnails, titles, and keywords. The extracted viral elements are stored as internal data in the AI model.
[1712] Step 4:
[1713] The server uses a generative AI model to analyze videos uploaded by users. It tags each scene along the video's timeline and extracts important lines and background music through audio analysis. Based on the results of this analysis, the next step is automatic editing.
[1714] Step 5:
[1715] The server automatically edits the videos uploaded by users based on the extracted viral elements. Specifically, the following editing is performed:
[1716] Cut Unwanted Scenes: Automatically remove unwanted scenes from your video.
[1717] Highlighting the highlights: Highlight important scenes.
[1718] Add sound effects and music: Add sound effects and background music that suit your video.
[1719] Insert text or title: Insert catchy text or title.
[1720] Auto-generate thumbnails: Automatically generate visually appealing thumbnails.
[1721] Step 6:
[1722] A preview link is generated for the completed edited video and provided to the user. The user can access this preview link through the application and check the edited results. This process generates the preview link and notifies the user.
[1723] Step 7:
[1724] Users can review the edited video and provide feedback, including ratings and specific suggestions for correction, which is then sent to the server.
[1725] Step 8:
[1726] The server receives feedback from users and updates the generative AI model. Based on the feedback data, the AI model's algorithm is adjusted to improve the accuracy of video analysis and editing in future videos. This feedback loop continuously improves the performance of the entire system.
[1727] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1728] ---
[1729] This invention utilizes the generated AI to automatically edit videos uploaded by users to make them more likely to go viral, and also combines it with an emotion engine that recognizes the user's emotions. This system automatically performs a series of processes including video analysis, element extraction, automatic editing, and the collection and utilization of emotional feedback.
[1730] Program processing and explanation
[1731] Uploading videos
[1732] User:
[1733] Users upload video files from their own devices using the system's dedicated web interface or application, and can enter a brief description and keywords for the video.
[1734] Receiving and storing videos
[1735] server:
[1736] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video format and metadata (resolution, format, file size, etc.) and records this information in a database.
[1737] Extracting buzzworthy elements
[1738] server:
[1739] Generative AI is used to analyze large amounts of data from videos that have recorded high numbers of views in the past. This analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. From this data, buzz-generating elements are extracted and used to generate patterns.
[1740] Analysis and editing of user videos
[1741] server:
[1742] The server analyzes videos uploaded by users, tags each scene along the timeline, and extracts important lines and background music through audio analysis. It then automatically edits the video based on the extracted viral elements, as follows:
[1743] Cutting out unnecessary scenes
[1744] Highlighting the highlights
[1745] Adding sound effects and music
[1746] Inserting text and titles
[1747] Automatic thumbnail generation
[1748] Submitting edits and gathering feedback
[1749] server:
[1750] The edited video, along with the suggested thumbnail and title, is provided to the user via a preview link that is generated for the user. The user is then notified of this link, allowing them to view the edited video on the preview screen.
[1751] User:
[1752] Visit the preview link provided to see the edited video and provide feedback, including any correction requests and a satisfaction rating.
[1753] Use of emotion engine
[1754] server:
[1755] The server uses an emotion engine to recognize the user's emotions in real time while watching videos. The emotion engine analyzes the user's facial expressions, tone of voice, and reactions to collect emotional data such as joy, surprise, sadness, and excitement.
[1756] Utilizing Emotional Feedback and Updating AI
[1757] server:
[1758] Based on the emotional feedback collected by the emotion engine, the generative AI adjusts the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This allows for further personalization.
[1759] User:
[1760] Review the final edit and provide further feedback as needed, including requests for re-edits and additional emotional feedback.
[1761] server:
[1762] Based on the feedback, necessary corrections are made and the final video is provided to the user. The generation AI and emotion engine are updated based on the feedback information to improve accuracy.
[1763] Specific examples
[1764] Example 1: Travel video editing
[1765] User:
[1766] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[1767] server:
[1768] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[1769] User:
[1770] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[1771] server:
[1772] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[1773] Example 2: Editing a gadget review video
[1774] User:
[1775] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[1776] server:
[1777] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1778] User:
[1779] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[1780] server:
[1781] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[1782] ---
[1783] This allows users to effortlessly create high-quality videos that are likely to go viral, thereby reducing the burden on creators and improving the success rate of their videos.
[1784] The processing flow will be explained below.
[1785] ---
[1786] Step 1: User uploads a video
[1787] User: The user selects a video file using the system's dedicated web interface or application on their device, clicks the upload button, and enters a brief description and keywords related to the video.
[1788] Step 2: The server receives and stores the video
[1789] Server: Receives video files sent by users. Stores the received video files in a temporary storage location. Analyzes the video format and metadata (resolution, format, file size, etc.) and records them in a database.
[1790] Step 3: Extract elements that will make the server buzz
[1791] Server: Launches the generation AI and analyzes a large number of past videos with high view counts. Analysis includes video content, editing patterns, thumbnails, titles, keywords, etc. Buzzworthy elements are extracted and used to generate patterns.
[1792] Step 4: The server analyzes the user's video
[1793] Server: Performs detailed analysis of the content of videos uploaded by users, tagging each scene along the timeline, and extracts important lines and background music through audio analysis.
[1794] Step 5: The server will automatically edit the video
[1795] Server: Based on the extracted buzzworthy elements, the following edits are automatically made:
[1796] Cutting out unnecessary scenes
[1797] Highlighting the highlights
[1798] Adding sound effects and music
[1799] Inserting text and titles
[1800] Automatic thumbnail generation
[1801] Step 6: The server serves the edits
[1802] Server: Generates a preview link for the user with the edited video and suggested thumbnail and title. Notifies the user of the preview link.
[1803] Step 7: User reviews the video and provides feedback
[1804] Users: Visit the provided preview link to view the edited video and provide feedback, including requested corrections and a satisfaction rating.
[1805] Step 8: The server updates the AI based on the feedback
[1806] Server: Receives user feedback and initiates the re-editing process. The generative AI incorporates the feedback as training data and updates its algorithms and models. If the user is satisfied, the final edited result is saved in the user's account and a download link is provided.
[1807] Step 9: The server starts the emotion engine that recognizes the user's emotions.
[1808] Server: When a user watches a video, the emotion engine is activated and analyzes the user's facial expressions, tone of voice, and reactions in real time. Through the collected data, the server recognizes the user's emotions (happiness, surprise, sadness, excitement, etc.).
[1809] Step 10: The server collects and uses emotional feedback
[1810] Server: Adjusts editing patterns based on the emotional feedback collected by the emotion engine. For example, if a user expresses strong positive emotions in a particular scene, the server emphasizes that scene. Conversely, if the user expresses negative emotions, the server deletes or edits that scene.
[1811] Specific examples
[1812] Example 1: Travel video editing
[1813] User:
[1814] A user uploads a video taken during a trip. The user enters "Summer trip to Hawaii" as a brief description and the keywords "beach, scenery."
[1815] server:
[1816] The AI analyzes the video and audio of travel videos to extract scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the AI emphasizes scenic scenes, adds upbeat music, and automatically uses beautiful beach images as thumbnails.
[1817] User:
[1818] Review the edit and provide emotional feedback, for example, if you were very moved by a particular scene, request that it be emphasized.
[1819] server:
[1820] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine continue to learn based on the feedback.
[1821] Example 2: Editing a gadget review video
[1822] User:
[1823] A user uploads a video review of a new gadget. Enter the keywords "technology, review, new product."
[1824] server:
[1825] The AI analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the product. It also extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1826] User:
[1827] Review edited videos and provide emotional feedback, for example highlighting scenes where a user expresses surprise at a particular product feature.
[1828] server:
[1829] Based on the emotional feedback, the editing content is adjusted and the final video is delivered. The AI and emotion engine are updated based on the feedback information to improve accuracy.
[1830] ---
[1831] These are the processing steps of an automatic video editing system that uses generative AI combined with an emotion engine.
[1832] Example 2
[1833] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1834] On modern video distribution platforms, videos need to be properly edited and have elements that capture viewers' interest in order to be viewed by a large audience. However, video editing is time-consuming and requires specialized knowledge, placing a heavy burden on many content creators. Furthermore, it is difficult to reflect viewer emotions and feedback in real time, limiting the improvement of video quality. There is a need for a system that can solve these problems and automatically create videos with a high probability of going viral.
[1835] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes: means for analyzing a large number of highly viewed videos using generated artificial intelligence and extracting buzzworthy elements; means for receiving and analyzing videos uploaded by users; means for automatically editing the received videos based on the extracted buzzworthy elements; means for providing the edited videos to users and receiving feedback; means for updating the generated artificial intelligence based on the feedback and user emotional data; means for recognizing users' emotions while watching videos in real time and collecting the emotional data; and means for adjusting the video editing pattern using the emotional data. This allows users to automatically create videos that are likely to go viral without requiring specialized knowledge and personalize them in real time based on viewer emotions.
[1836] "Generated artificial intelligence" is a data processing system that is trained on past data and is a model that performs advanced analysis and predictions on new data.
[1837] "Highly viewed videos" refer to video content that has recorded a large number of views on a video distribution platform.
[1838] "Buzz elements" refer to the editing points and content features necessary for a video to be well-received by viewers and widely distributed.
[1839] "User" refers to an individual or corporation that uses this system, uploads videos, and receives editing services from the system.
[1840] "Uploading" refers to the act of a user sending a video file from their device to a server via the Internet.
[1841] "Analysis" refers to the process of examining the content, structure, metadata, etc. of a video in detail and extracting important information.
[1842] "Automatic editing" refers to the process of cutting video, adding music, or other edits based on specific algorithms and rules without user input.
[1843] "Feedback" refers to the action of a user providing opinions or requests for improvement regarding a video edited by the system.
[1844] "Emotion data" is data that expresses, as numerical values or indices, emotions such as joy, sadness, and surprise that a user shows while watching a video.
[1845] "Real-time recognition" refers to the ability to instantly analyze and record the user's emotions at each moment while watching a video.
[1846] "Editing pattern" refers to a series of editing techniques or styles in video editing, such as cutting, rearranging scenes, and adding sound effects.
[1847] "Personalization" is a method of individually optimizing services and content based on individual user preferences and feedback.
[1848] MODE FOR CARRYING OUT THE INVENTION
[1849] Uploading videos
[1850] User:
[1851] Users upload video files from their own devices using a dedicated web interface or application. In this step, they enter a brief description of the video (e.g., "Summer trip to Hawaii") and keywords (e.g., "beach, scenery").
[1852] Receiving and storing videos
[1853] server:
[1854] The server receives video files uploaded by users and temporarily stores them in the system's storage. During this process, the video format and metadata (resolution, format, file size, etc.) are analyzed and recorded in a database. The server uses a high-performance computer system and database system to efficiently analyze and edit video data.
[1855] Extracting buzzworthy elements
[1856] server:
[1857] Using a generative AI model, information such as footage, thumbnails, titles, and keywords is collected and analyzed from a database of videos with high past views. This allows elements for buzz (e.g., emphasis points, editing patterns, use of sound effects, etc.) to be extracted. The generative AI model is trained using deep learning techniques based on a large dataset.
[1858] Analysis and editing of user videos
[1859] server:
[1860] The server analyzes the video uploaded by the user scene by scene and tags important scenes. Next, it performs audio analysis to extract important dialogue and background music characteristics. Based on this, it performs the following automatic editing operations:
[1861] Cutting out unnecessary scenes
[1862] Highlighting the highlights
[1863] Adding sound effects and music
[1864] Inserting text and titles
[1865] Automatic thumbnail generation
[1866] Submitting edits and gathering feedback
[1867] server:
[1868] The completed edited video, along with the suggested thumbnail and title, is generated as a preview link for the user and notified to the user.
[1869] User:
[1870] Users can access the preview link to view the edited video and provide feedback on their satisfaction and corrections, including comments on the editing process and specific requests for improvements.
[1871] Use of emotion engine
[1872] server:
[1873] While the user is watching the preview, the server uses an emotion engine to recognize the user's emotions in real time, and collects and analyzes the data. Emotion data includes happiness, surprise, sadness, excitement, etc., obtained through facial recognition and voice analysis.
[1874] Utilizing Emotional Feedback and Updating AI
[1875] server:
[1876] The server adjusts the editing patterns of the generative AI model based on the emotion data collected by the emotion engine and user feedback, enabling more personalized video editing by emphasizing scenes with strong positive emotions and editing or deleting scenes that show negative emotions.
[1877] User:
[1878] Users can review the final edits and provide additional feedback, which improves the quality of the resulting video.
[1879] Specific examples
[1880] Example 1: Travel video editing
[1881] A user uploads a video taken during a summer trip to Hawaii and enters "Summer trip to Hawaii" and the keywords "beach, scenery" as a brief description.
[1882] The server analyzes the video and audio of travel videos and extracts scenes that include beautiful scenery and local culture. Based on the viral elements of past travel videos, the generative AI model emphasizes scenic scenes, adds lively music, and automatically uses beautiful beach images as thumbnails.
[1883] The user reviews the edited results and provides emotional feedback on the edits, such as requesting that a particular scene be emphasized if it moved them.
[1884] The server adjusts the edit based on the emotional feedback and delivers the final video, allowing the generative AI model and emotion engine to learn and improve over time.
[1885] Example 2: Editing a gadget review video
[1886] A user uploads a video review of a new gadget and enters the keywords "technology, review, new product."
[1887] The server analyzes the content of gadget review videos and extracts scenes that highlight the unique features and functions of the products. The generative AI model extracts buzzworthy elements from past gadget review videos, adds sound effects, generates catchy titles, and creates impactful thumbnails.
[1888] Users review the edited video and provide emotional feedback, highlighting scenes where they express surprise at a particular product feature.
[1889] The server adjusts the editing content based on the emotional feedback and provides the final video, thereby improving the accuracy of the generative AI model and emotion engine.
[1890] This system allows users to easily create personalized, high-quality videos that are likely to go viral, even without specialized editing skills, by reflecting viewers' real-time reactions.
[1891] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1892] Step 1: Select and upload your video
[1893] User:
[1894] The user selects a video file from their device and uploads it through the system's dedicated web interface or application. The information entered in this step is the video file (e.g., "Hawaii Trip.mp4"), a video description (e.g., "Summer Hawaii Trip"), and keywords (e.g., "Beach, Scenery"). The server receives this data and prepares it for the next step.
[1895] Step 2: Receive and save the video
[1896] server:
[1897] The server receives video files uploaded by users and temporarily stores them in the system's storage. The input is the video file uploaded by the user and its associated metadata. The server analyzes the metadata, such as the video format, resolution, format, and file size, and records it in a database as output. This step makes it easier to manage and access the videos.
[1898] Step 3: Identifying viral elements
[1899] server:
[1900] The server uses a generative AI model to analyze past videos with high view counts stored in a database. The input is data on past videos with high view counts, and the analysis targets video content, thumbnails, titles, keywords, etc. The server extracts buzz-generating elements from these and updates the generative AI model. The output is patterns of buzz-generating elements (e.g., timing of noteworthy scenes, use of sound effects).
[1901] Step 4: Analyzing user videos
[1902] server:
[1903] The server analyzes videos uploaded by users for each scene and assigns specific tags. The input is the user's video file, and the server divides the scenes using a timeline and performs audio analysis to extract important dialogue and musical features. The output is tagged scene information and extracted audio data.
[1904] Step 5: Auto Edit
[1905] server:
[1906] The server automatically edits the user's video based on the extracted viral elements. The input is the analyzed scene information and viral element patterns. This process involves cutting out unnecessary scenes, emphasizing highlights, adding sound effects and music, inserting text and titles, and automatically generating thumbnails. The output is an edited video file.
[1907] Step 6: Submit your edits and get feedback
[1908] server:
[1909] The server generates a preview link for the completed video, along with the suggested thumbnail and title, and notifies the user. The input is the edited video file, and the output is the preview link.
[1910] User:
[1911] The user accesses the provided preview link to view the edited video. The input is the preview link, and the user provides feedback on satisfaction and corrections. The output is the feedback information sent to the system.
[1912] Step 7: Real-time emotion recognition
[1913] server:
[1914] The server uses an emotion engine to recognize and collect data in real time about the emotions users express while watching videos. The input is the user's viewing data (facial expressions and voice), which the emotion engine analyzes. The output is emotional data such as joy, surprise, sadness, and excitement.
[1915] Step 8: Leverage emotional feedback and update the AI
[1916] server:
[1917] The server adjusts the editing patterns of the generative AI model based on the emotional data collected by the emotion engine and feedback from users. The input is emotional data and feedback, and the server uses this to consider whether to emphasize or delete specific scenes. The output is an updated editing pattern and model.
[1918] Step 9: Final edits and user confirmation
[1919] server:
[1920] The server then re-edits the video based on the revised editing pattern and generates the final video. The input is the updated editing pattern, and the output is the final edited video file.
[1921] User:
[1922] The user then reviews the final edit and provides further feedback if necessary. The input is the final video, and the output is the feedback sent back to the server.
[1923] In this way, by going through a series of processing steps, users can automatically generate videos that are likely to go viral and obtain personalized content based on emotional data.
[1924] (Application example 2)
[1925] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1926] When creating video content, users need to spend time and effort on editing in order to increase the number of views. However, finding effective editing techniques and buzzworthy elements is difficult, placing a significant burden on users without specialized knowledge. Furthermore, while there is a demand for personalized content that reflects the emotions of viewers, there is a lack of means to achieve this. A system that solves these issues and allows anyone to easily create high-quality, buzzworthy videos is needed.
[1927] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1928] In this invention, the server includes means for analyzing a large number of highly viewed videos using the generated artificial intelligence and extracting buzz-generating elements, means for receiving and analyzing videos uploaded by users, means for automatically editing the received videos based on the extracted buzz-generating elements, means for providing the edited videos to users and receiving feedback, means for updating the generated artificial intelligence based on the feedback, means for recognizing user emotions in real time using an emotion engine and collecting emotional feedback, and means for adjusting editing patterns based on the emotional feedback. This enables users, even without technical knowledge, to efficiently create and provide videos that reflect viewer emotions and are likely to go viral.
[1929] "Generated artificial intelligence" is artificial intelligence that learns from data and is used to automate or optimize specific tasks.
[1930] A "highly viewed video" is video content that has been viewed by many users in a short period of time.
[1931] "Buzz elements" are characteristic features that make a video more likely to be shared and spread by many viewers.
[1932] "User-uploaded videos" are video files that users submit to the system from their devices.
[1933] "Means for analysis" refers to technologies or processes that provide the functionality to analyze the content of a video and extract specific elements or patterns.
[1934] "Automatic editing methods" refer to technologies and algorithms that use analyzed data to edit videos without human intervention.
[1935] "Means for receiving feedback" is a function for collecting ratings and comments from users.
[1936] "Means of updating" refers to the ability to improve artificial intelligence models and algorithms based on collected feedback.
[1937] The "emotion engine" is a technology that analyzes emotions from a user's facial expressions and voice in real time.
[1938] "Emotional feedback" is data based on the emotional reactions of users when they watch videos.
[1939] "Means for adjusting editing patterns" refers to a function that changes the method and content of video editing based on collected emotional data.
[1940] (Mode for carrying out the invention)
[1941] The system for implementing this invention is mainly composed of a server, a user's device, a generative AI model, and an emotion engine. Each step and its specific implementation method are described below.
[1942] The server receives video files uploaded by users and stores them in storage. The received video files are analyzed for their metadata (resolution, format, file size, etc.) and recorded in a database.
[1943] The server then analyzes a large number of highly viewed videos using a generative AI model. This analysis process extracts elements such as video content, editing patterns, thumbnails, titles, and keywords to identify buzzworthy elements. The generative AI model is trained using past video data and viewer feedback.
[1944] When the server analyzes a user's video, it tags each scene based on the timeline and extracts important lines and background music through audio analysis. Based on the extracted buzzworthy elements, it then cuts out unnecessary scenes, emphasizes highlights, adds sound effects and music, inserts text and titles, and automatically generates thumbnails.
[1945] Once the edits are complete, a preview link is generated and provided to the user, where the user can review the edited video and provide feedback, including requests for corrections and a satisfaction rating.
[1946] Furthermore, the emotion engine analyzes the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data such as joy, surprise, sadness, and excitement, thereby obtaining emotional feedback.
[1947] The server uses this emotional feedback to update the AI and adjust the editing patterns. For example, if the user expresses strong positive emotions in a particular scene, it will emphasize that scene. Conversely, if the user expresses negative emotions, it will delete or edit that scene. This feedback loop results in personalized videos and improved quality.
[1948] The main hardware used includes:
[1949] Video editing server (equipped with high-performance CPU / GPU)
[1950] Smartphone or PC (user interface)
[1951] The main software used includes:
[1952] moviepy (video editing library)
[1953] SentimentAnalyzer (custom model for sentiment analysis)
[1954] VideoEditorAI (generative AI model)
[1955] Specific examples
[1956] Example 1: Travel video editing
[1957] Users upload videos they have taken during their trip. For example, a "summer trip to Hawaii" includes scenes of the beach, scenery, and local culture. The server then selects highlights based on these scenes and adds appropriate music and sound effects. Further editing is performed based on user feedback and emotional feedback to create the final video.
[1958] Prompt Sentence Examples
[1959] "Analyze video files and generate optimal editing patterns based on elements of past viral videos."
[1960] "Improve your video editing process and create personalized edits based on user emotional feedback data."
[1961] This system enables users, even without technical knowledge, to efficiently create and distribute videos that reflect viewers' emotions and are likely to go viral.
[1962] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1963] Step 1:
[1964] The user selects a video file and uploads it through the system's dedicated web interface or application, and can enter a brief description and keywords for the video. The entered video file and metadata are then sent to the server.
[1965] Input: Video file, description, keywords
[1966] Output: Video files and metadata sent to the server
[1967] Step 2:
[1968] The server receives the video file sent by the user and temporarily stores it in the system's storage. It analyzes the video file's metadata (resolution, format, file size, etc.) and records it in a database.
[1969] Input: Received video files and metadata
[1970] Output: Metadata of video files recorded in a database
[1971] Step 3:
[1972] The server uses a generative AI model to analyze large amounts of video data that has recorded high numbers of views in the past. During this analysis process, elements such as the content of the video, editing patterns, thumbnails, titles, and keywords are extracted to identify buzzworthy elements.
[1973] Input: Data of past videos with high views
[1974] Output: Extracted buzzworthy elements
[1975] Step 4:
[1976] The server analyzes the video files uploaded by users, tagging each scene along the timeline and extracting important lines and background music characteristics through audio analysis.
[1977] Input: A video file uploaded by the user
[1978] Output: Scene-specific tagged data, audio analysis results
[1979] Step 5:
[1980] The server automatically edits the video by cutting out unnecessary scenes and emphasizing highlights based on the extracted buzzworthy elements, adding sound effects and music, inserting text and titles, and automatically generating thumbnails.
[1981] Input: Extracted buzz elements, tagged scene data, audio analysis results
[1982] Output: Edited video file
[1983] Step 6:
[1984] The server generates a preview link containing the edited video file and provides it to the user, allowing the user to view the edited video and provide feedback.
[1985] Input: Edited video file
[1986] Output: Preview link provided to the user
[1987] Step 7:
[1988] Users can access the provided preview link to view the edited video, and provide feedback by requesting corrections or rating their satisfaction on the feedback screen.
[1989] Input: Preview link, user feedback
[1990] Output: Feedback data sent to the server
[1991] Step 8:
[1992] The server uses an emotion engine to analyze the user's facial expressions, tone of voice, and reactions while watching the video, and collects emotional data, which is then stored as emotional feedback such as happiness, surprise, sadness, and excitement.
[1993] Input: User's facial expression, voice
[1994] Output: Collected emotional feedback data
[1995] Step 9:
[1996] The server updates the generative AI and adjusts the editing patterns based on the collected feedback data and emotional feedback, enhancing the quality of the video by highlighting scenes that show positive emotions and deleting or editing scenes that show negative emotions.
[1997] Input: Feedback data, Emotion feedback data
[1998] Output: Adjusted editing patterns, updated generative AI model
[1999] Step 10:
[2000] Finally, the server provides the user with a re-edited video based on the updated generative AI model, and the user can provide further feedback as needed, which the server then incorporates to refine the final video.
[2001] Input: Re-edited video with updated generative AI model
[2002] Output: The final video file
[2003] Through these processing steps, users can efficiently create and provide personalized videos that are likely to go viral, even without any technical knowledge.
[2004] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2005] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2006] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2007] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2008] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2009] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2010] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2011] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2012] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2013] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2014] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2015] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2016] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2017] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2018] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2019] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2020] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2021] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2022] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2023] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2024] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2025] The following is further disclosed regarding the above embodiment.
[2026] (Claim 1)
[2027] The generated AI is used to analyze many highly viewed videos and extract elements that will create buzz.
[2028] means for receiving and analyzing videos uploaded by users;
[2029] A means for automatically editing the received video based on the extracted buzz...
Claims
1. The generated AI is used to analyze many highly viewed videos and extract elements that will create buzz. means for receiving and analyzing videos uploaded by users; A means for automatically editing the received video based on the extracted buzzworthy elements; a means for providing the edited video to users and receiving feedback; The system includes means for updating the generated artificial intelligence based on said feedback.
2. The system according to claim 1, wherein unnecessary scenes from the video are automatically cut out and highlight scenes are emphasized based on the extracted buzzworthy elements.
3. 10. The system of claim 1, wherein the system automatically adds sound effects and music, and inserts text and visual effects into the edited video.
4. The system according to claim 1, wherein the system tags each scene along the timeline of a video uploaded by a user and performs audio analysis.
5. The system according to claim 1, further comprising: generating a preview link for the edited video and notifying the user of the preview link;
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A