System
The system addresses the challenge of manually selecting video music by using AI to analyze video content and generate synchronized music, allowing users to effortlessly create harmonious video content with scene-specific music.
Patent Information
- Application Number
- JP2024123945
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional video editing requires significant time and effort to manually select and synchronize music suitable for video data, often necessitating specialized knowledge and skills, and there is a demand for a flexible system that can automatically generate and apply different music patterns for various genres and scenes.
A system that utilizes an image analysis module to analyze video content and a music generation module to automatically generate and synchronize appropriate music based on user preferences, allowing users to specify genres and scene-specific music patterns.
Enables users to quickly and easily add high-quality, harmonious music to their videos without manual effort, ensuring the music matches the video's atmosphere and genre, thus simplifying the video editing process.
Smart Images

Figure 2026022428000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In conventional video editing, the task of selecting and inserting music suitable for video data is extremely time-consuming and laborious, often requiring specialized knowledge and skills. This makes it difficult for video creators to quickly and easily add music suitable for video data. There has also been a demand for a flexible system that can automatically generate and apply different music patterns for different music genres and scenes. [Means for solving the problem]
[0005] The present invention provides a system that uses an image analysis module to analyze the content of video data uploaded by a user, and then uses a music generation module to automatically generate and synchronize appropriate music based on the analysis results. Specifically, the above-mentioned problems are solved by building a system that includes the following means.
[0006] 1. A means for users to upload video data;
[0007] 2. A means for storing the video data received by the server;
[0008] 3. A means for the server to analyze the content of the video data using an image analysis module;
[0009] 4. A means for the server to generate appropriate music using a music generation module based on the analysis results;
[0010] 5. A means for synchronizing the generated music with the video data;
[0011] 6. A means for providing the generated video data to the user.
[0012] Furthermore, by providing a means for the user to specify the genre of music and a means for generating different music patterns for each scene based on the analysis results, it becomes possible to add music more flexibly and according to the user's wishes.
[0013] A "user" is someone who creates video data and uploads it to the system.
[0014] A "terminal" is an electronic device such as a computer or smartphone used by a user, and is a device for uploading and downloading video data.
[0015] The "server" is a central processing unit that receives, stores, analyzes, and processes video data, and performs various processes for editing.
[0016] "Video data" refers to moving image data such as moving image files, movies, animations, etc. created by users.
[0017] "Uploading means" refers to a function or method for transmitting video data from a user's terminal to a server.
[0018] An "image analysis module" is a function or software that analyzes the content of each frame of video data and extracts the features of the scene.
[0019] The "analysis results" are the feature amounts and data for each scene of the video data extracted by the image analysis module.
[0020] The "music generation module" is a function or software that automatically generates optimal music for video based on the analysis results.
[0021] "Appropriate music" is music that matches the content of each scene of the video data and that matches the genre specified by the user.
[0022] "Means for synchronizing" refers to a function or method for matching the generated music with the timeline of the video data.
[0023] The "means for providing to the user" refers to a method or function for the user to download or view the generated final edited video data.
[0024] The "means for specifying a genre" is a function or interface that allows the user to specify the style or type of music that they desire.
[0025] "Different musical patterns for each scene" refers to different musical portions or patterns that correspond to the content of each scene. [Brief explanation of the drawings]
[0026] [Figure 1]1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0027] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0028] First, the terms used in the following description will be explained.
[0029] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0030] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0031] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0032] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0033] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0034] [First embodiment]
[0035] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0036] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0037] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0038] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0039] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0040] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0041] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0042] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0043] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0044] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0045] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0046] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0047] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data. Each step of this system is described in detail below.
[0048] First, the user uploads video data from their own device to the server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[0049] The server then stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[0050] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[0051] The results of the image analysis module are returned to the server, which then determines the appropriate music generation requirements for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0052] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[0053] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[0054] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0055] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[0056] Specific examples
[0057] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[0058] 1. User A uses a dedicated app to upload a travel video file to the server.
[0059] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0060] 3. The server passes the saved video file to the image analysis module, which extracts the features of each scene.
[0061] 4. User A selects a rock-style music genre in the application and sends it to the server.
[0062] 5. The server passes the analysis results and genre specification to the music generation module, which then generates rock-style music.
[0063] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[0064] 7. User A can download the completed video from the application and watch it.
[0065] In this way, the system of the present invention allows the user to automatically assign suitable music to video data simply and quickly.
[0066] The processing flow will be explained below.
[0067] Step 1:
[0068] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[0069] Step 2:
[0070] The server detects the received video data and stores it in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0071] Step 3:
[0072] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[0073] Step 4:
[0074] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[0075] Step 5:
[0076] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[0077] Step 6:
[0078] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates music appropriate for each scene based on the specified requirements. This music is then adjusted to create a harmonious overall composition.
[0079] Step 7:
[0080] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[0081] Step 8:
[0082] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[0083] Specific operation example
[0084] Step 1:
[0085] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[0086] Step 2:
[0087] The server receives the video data and saves it in the "temporary / uploads / " folder.
[0088] Step 3:
[0089] The server starts the image analysis module and analyzes the video data. The analysis module extracts the features of each frame.
[0090] Step 4:
[0091] The analysis module returns the extracted features to the server, which then sets the music generation requirements appropriate for each scene based on the features.
[0092] Step 5:
[0093] User A selects the music genre "rock" from the application and sends it to the server. The server receives this information.
[0094] Step 6:
[0095] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates rock-style music.
[0096] Step 7:
[0097] The server synchronizes the generated rock music with the timeline of the video data.
[0098] Step 8:
[0099] The server encodes the final edited video data and provides a downloadable link to User A. User A downloads the edited video from the link and watches it.
[0100] Example 1
[0101] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0102] Conventional video editing systems require users to manually select and edit music appropriate for the video, which requires a great deal of time and effort. It is also not easy to select music that matches the different atmospheres of each video scene. Furthermore, synchronizing music and video requires advanced editing skills, making it difficult for average users to create the high-quality video content desired quickly and easily.
[0103] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0104] In this invention, the server includes a means for a user to upload video data, a means for the server to store the received video data, a means for the server to analyze the content of the video data using an image analysis module, a means for the server to generate appropriate music using a music generation module based on the analysis results, a means for synchronizing the generated music with the video data, and a means for providing the generated video data to the user. This enables the user to quickly create high-quality video content using AI technology without any manual effort.
[0105] "User" refers to an entity that uses the system to upload video data and request music creation.
[0106] "Server" refers to a computer system that performs a series of processes including receiving, storing, analyzing, and generating music from video data.
[0107] "Video data" refers to video files uploaded by users and supports a variety of formats.
[0108] "Uploading means" refers to an application or interface that allows a user to send video data to the server.
[0109] The "storage means" refers to a storage area in which the server temporarily stores the video data received.
[0110] An "image analysis module" refers to a software component that analyzes the contents of video data and extracts features for each scene.
[0111] "Music generation module" refers to a software component that automatically generates music suitable for video based on analysis results.
[0112] "Synchronization means" refers to a process that performs operations to integrate the generated music with the video data.
[0113] "Providing means" refers to a method for providing the generated video data to the user.
[0114] "Metadata" refers to information contained in video data (e.g., resolution, frame rate, file format, etc.).
[0115] "Generative AI model" refers to the artificial intelligence algorithm used by the music generation module, enabling the automatic generation of music.
[0116] A "prompt" is an instruction given to a generative AI model that describes the requirements for generating specific music.
[0117] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data.
[0118] First, the user uploads video data from their device to the server using a dedicated application. Video data can be video files in various formats, ranging from a few minutes to several hours in length. To do this, the user clicks the "Upload" button in the application and selects the appropriate video file from the file selection dialog. The video is then uploaded to the server via cloud storage.
[0119] Next, the server stores the received video data in a temporary storage area (directory "temporary / uploads / "). At the same time, it obtains the video data's metadata (resolution, frame rate, file format, etc.) to prepare for subsequent analysis.
[0120] The server launches an image analysis module to analyze the stored video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). This analysis uses common libraries such as OpenCV and TensorFlow. The analyzed data is temporarily stored in intermediate data storage.
[0121] Based on the analysis results from the image analysis module, the server then defines the music generation requirements for each scene. For example, a quiet landscape scene requires calm music, while an active scene requires energetic music. These defined requirements are recorded in JSON format.
[0122] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through a dedicated application, selects the desired genre from a drop-down menu in the application, and sends the information to the server.
[0123] The server passes the analysis results and genre specification to the music generation module, which uses AI algorithms to automatically generate optimal music based on this information. This generation process can utilize OpenAI's MuseNet or other generative AI models.
[0124] The generated music is then synchronized with the video data by a server. During the synchronization process, the timing of the music is adjusted to match each scene in the video data, resulting in a harmonious overall music and video. Specifically, the music track is merged into the video file using libraries such as FFmpeg.
[0125] Finally, the server encodes the edited video data and generates a download link for the user, who can click the "Download" button in the application to download and watch the completed video.
[0126] Specific examples
[0127] For example, if user A wants to create a virtual travel video and add rock-style background music to it, he or she would follow the steps below.
[0128] 1. User A clicks the "Upload" button in the application, selects a travel video file, and uploads it to the server.
[0129] 2. The server receives and stores the video file and retrieves the metadata.
[0130] 3. The server uses OpenCV to analyze the video data and extract features for each scene.
[0131] 4. User A selects a rock-style music genre in the application and sends it to the server.
[0132] 5. The server passes the analysis results and genre specification to the music generation module, which generates rock-style music using MuseNet.
[0133] 6. The server uses FFmpeg to merge the generated music into the video file and generate the edited video.
[0134] 7. User A downloads the completed video from the application and watches it.
[0135] Prompt Sentence Examples
[0136] "Analyze the footage of the virtual trip and add calming music to the quiet scenes and energetic music to the dynamic scenes. Overall, keep the music taste rock-like."
[0137] As described above, the present invention enables a user to automatically assign suitable music to video data simply and quickly.
[0138] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0139] Step 1:
[0140] The user uploads the video data.
[0141] Specifically, the user launches the dedicated application, clicks the "Upload" button, and selects the video file they want to upload from the file selection dialog. The input is the user's video data, and the output is the video data sent to the server via cloud storage. The video data can be in a variety of video formats, ranging from a few minutes to a few hours.
[0142] Step 2:
[0143] The server stores the received video data.
[0144] Specifically, the server stores the video data in the "temporary / uploads / " directory. At this time, the server acquires the video data's metadata (resolution, frame rate, file format, etc.). The input is the uploaded video data, and the output is the saved data and the acquired metadata. This makes subsequent analysis easier.
[0145] Step 3:
[0146] The server starts the image analysis module and analyzes the video data.
[0147] Specifically, the server uses libraries such as OpenCV and TensorFlow to analyze each frame of video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). The input is the saved video data, and the output is the analysis results (features for each scene). These analysis results are temporarily stored in intermediate data storage.
[0148] Step 4:
[0149] The server defines the requirements for music generation based on the analysis results.
[0150] Specifically, based on the data obtained from the analysis module, the music specifications appropriate for each scene (e.g., gentle music for quiet scenes, energetic music for active scenes) are written in JSON format. The input is the analysis results, and the output is JSON data that records the requirements for music generation.
[0151] Step 5:
[0152] The user specifies the genre of music.
[0153] Specifically, the user selects the desired music genre (e.g., rock, jazz, etc.) from a drop-down menu in the application and sends that information to the server. The input is the user's genre selection, and the output is the genre information sent to the server.
[0154] Step 6:
[0155] The server passes the analysis results and genre designation to the music generation module.
[0156] Specifically, it uses a generative AI model such as OpenAI's MuseNet to generate music based on the analysis results and the user's genre selection. The input is the analysis results and genre information, and the output is the generated music data. Based on this information, the music generation module automatically generates music customized for each scene.
[0157] Step 7:
[0158] The server synchronizes the generated music with the video data.
[0159] Specifically, we use libraries such as FFmpeg to merge the generated music with the video data and adjust the timing. The input is the video data and the generated music, and the output is edited video data synchronized with the music. In this editing process, we optimize the timing of the music for each scene.
[0160] Step 8:
[0161] The server provides the edited video data to the user.
[0162] Specifically, the edited video data is encoded and a download link is generated. Users can click the "Download" button in the application to download and watch the completed video. The input is the edited video data, and the output is the download link.
[0163] (Application example 1)
[0164] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0165] In the past, adding appropriate music to video data created by users required manual selection of music and editing to match the video, which required a great deal of time and effort. Furthermore, the video and music often did not harmonize, making it difficult to create content that is visually and aurally unified. To solve this problem, there is a demand for a system that can automatically generate appropriate music based on the content of the video data and easily provide content in which the video and music are in harmony.
[0166] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0167] In this invention, the server includes means for a user to upload video data, means for the server to store the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user via a terminal such as a smartphone, thereby enabling the user to quickly and easily create video data with optimal music added to the video, and to provide content that is visually and aurally harmonious.
[0168] "User" means a person or legal entity that has the right to use the system to upload video data and generate music.
[0169] "Video data" refers to all video files uploaded by users, the contents of which are analyzed to generate music.
[0170] "Upload" refers to the act of a user sending video data from their own device to a server.
[0171] The "server" is a central computer system that receives and stores video data, and processes it for analysis and music generation.
[0172] An "image analysis module" is a software component that analyzes video data and understands its contents.
[0173] A "music generation module" is a software component for automatically generating music based on the analysis results.
[0174] "Generating" "music" refers to the process by which the system automatically creates new music.
[0175] "Synchronization" refers to the process of matching the generated music to the appropriate timing for each scene in the video data.
[0176] "Devices such as smartphones" refer to portable information and communication devices used by users, which are capable of uploading video data and downloading generated video data.
[0177] "Generated video data" refers to the final video file that has been analyzed and processed to generate music, and is synchronized with the music.
[0178] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data with a music generation module to generate original music according to the user's preferences and add it to the video data.
[0179] First, a user uploads video data from their smartphone or other device to a server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[0180] The server stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[0181] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[0182] The results of the image analysis module are returned to the server, which then determines the appropriate music generation requirements for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0183] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[0184] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[0185] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0186] Finally, the server encodes the edited video data and generates a download link for users, who can then download and watch the completed video on their smartphones or other devices.
[0187] For example, if a user creates a travel video and wants to add rock-style background music to it, they can use this system as follows:
[0188] Example prompt sentence:
[0189] Travel Videos
[0190] Genre: Rock
[0191] Description: Adventurous scenery, mountain climbing, river rafting
[0192] Based on these prompts, the system executes each step and provides the user with a final edited video that is visually and audibly consistent.
[0193] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0194] Step 1:
[0195] The user uploads video data from their smartphone or other device. The input is a video file selected by the user, and the output is that video file is sent to the server. At this time, the user also specifies the desired music genre.
[0196] Step 2:
[0197] The server saves the received video data. It stores the video file in a temporary storage area and obtains the file metadata (resolution, frame rate, file format, etc.). The input is the video data received in step 1, and the output is the video data and its metadata stored in the storage area.
[0198] Step 3:
[0199] The server uses an image analysis module to analyze the content of the stored video data. The input is the video data stored in the storage area, and the output is the extracted results of features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). Specifically, it analyzes each frame of the video data and recognizes specific patterns and objects.
[0200] Step 4:
[0201] Based on the analysis results, the server defines the requirements for music generation appropriate for each scene in the video data. The input is the analysis results obtained in step 3, and the output is the requirements necessary for music generation (for example, calm music for quiet scenes, energetic music for active scenes). Specifically, the server analyzes the analysis results and determines the style and atmosphere of the music that suits each scene.
[0202] Step 5:
[0203] The user specifies the desired music genre through the application. The input is the music genre information specified by the user, and the output is that information is sent to the server.
[0204] Step 6:
[0205] The server passes the analysis results and genre specification to the music generation module. The input is the analysis results and the user's genre specification information, and the output is the input data for the music generation module.
[0206] Step 7:
[0207] The music generation module automatically generates optimal music. The input is the data obtained in step 6, and the output is the generated music. Specifically, an AI model (using TensorFlow, for example) generates music data based on the analysis results and genre information.
[0208] Step 8:
[0209] The generated music is synchronized with the video data. The server adjusts the timing of the music to match each scene in the video data. The input is the generated music and video data, and the output is edited video data in which the music and video are synchronized. Specifically, the start and end points of the music are adjusted for each scene to create integrated data.
[0210] Step 9:
[0211] The server encodes the edited video data and generates a link that allows the user to download it. The input is the edited video data generated in step 8, and the output is a download link. Specifically, the server encodes the data and provides the link to the user.
[0212] Step 10:
[0213] The user downloads the completed video from a device such as a smartphone and watches it. The input is a download link provided by the server, and the output is the final video data with music added, stored on the user's device.
[0214] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0215] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data, a music generation module for generating music, and an emotion engine for recognizing the user's emotions to generate original music according to the user's preferences and add it to the video data. Each step of this system is described in detail below.
[0216] First, the user selects video data using a dedicated application on their own device and clicks the upload button. The device then sends the selected video data file to the server.
[0217] Next, the server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0218] After receiving and saving the video data, the server starts the image analysis module and passes the saved video data to it. The image analysis module analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.). This allows for a detailed understanding of the content of the video data.
[0219] The results of the image analysis module are returned to the server, which then determines the appropriate music generation requirements for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0220] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and adds it to the requirements for music generation.
[0221] The present invention also includes an emotion engine. The emotion engine analyzes the user's facial expressions, voice, and text input to evaluate the user's emotional state. The emotion engine analyzes the data collected by a camera and microphone, including facial expressions and tone of voice, while the user is manipulating video data. The results of this analysis (the user's emotional state, such as whether they are enjoying, surprised, or sad) are reflected in the parameters for generating music.
[0222] The server then passes the analysis results, the user's emotional state, and the genre designation to the music generation module. The music generation module automatically generates optimal music based on this information. The tempo, tone, and genre of the music are dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright, fast-paced music is generated, while if the user is calm, calm, slow music is generated.
[0223] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0224] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[0225] Specific examples
[0226] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[0227] 1. User A uses the dedicated app to select a travel video file and presses the upload button. The device sends the video data to the server storage.
[0228] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0229] 3. The server passes the video data to the image analysis module, which extracts the features of each scene.
[0230] 4. User A specified a rock-style music genre in the application, and the emotion he felt when operating the video was determined to be joy.
[0231] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[0232] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[0233] 7. User A can download the completed video from the application and watch it.
[0234] In this way, the system of the present invention allows users to easily and quickly automatically assign suitable music to video data, and can provide more personalized music depending on the user's emotional state.
[0235] The processing flow will be explained below.
[0236] Step 1:
[0237] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[0238] Step 2:
[0239] The server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0240] Step 3:
[0241] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[0242] Step 4:
[0243] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[0244] Step 5:
[0245] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[0246] Step 6:
[0247] When a user interacts with video data (e.g., through facial expressions or voice input), the emotion engine evaluates the user's emotional state. The user's facial expressions and voice are recorded via the device's camera and microphone, and then analyzed by the emotion engine.
[0248] Step 7:
[0249] The emotion engine sends the analysis results to the server, which then incorporates this emotion information (e.g., whether the user is happy or surprised) into the requirements for music generation.
[0250] Step 8:
[0251] The server passes the analysis results, the user's emotional state, and genre information to the music generation module. The music generation module automatically generates optimal music based on this information. If the user is happy, it generates bright, fast-paced music, and if the user is calm, it generates calm, slow music.
[0252] Step 9:
[0253] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[0254] Step 10:
[0255] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[0256] Specific operation example
[0257] Step 1:
[0258] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[0259] Step 2:
[0260] The server receives the video data and saves it in the "temporary / uploads / " directory.
[0261] Step 3:
[0262] The server starts an image analysis module, analyzes each frame of the video data, and extracts features.
[0263] Step 4:
[0264] The image analysis module returns the analysis results to the server, which then defines the necessary music generation requirements.
[0265] Step 5:
[0266] User A selects the music genre "rock" from the application and sends it to the server.
[0267] Step 6:
[0268] When User A is operating the video, the device sends data to the emotion engine based on User A's facial expressions and voice. The emotion engine then analyzes User A's emotions.
[0269] Step 7:
[0270] The emotion engine sends the emotional state of user A (e.g., joy) to the server.
[0271] Step 8:
[0272] The server passes the analysis results, emotional state, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[0273] Step 9:
[0274] The server synchronizes the generated music with the video timeline to generate an edited video.
[0275] Step 10:
[0276] The server provides the edited video in a downloadable format to User A. User A then downloads the completed video from the application and watches it.
[0277] Example 2
[0278] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0279] Conventional video editing systems require users to individually select appropriate music and manually synchronize it with the video data, which requires time and effort. Furthermore, even systems capable of automatically generating music for video data lack the functionality to personalize the music based on the user's emotions and preferences. This makes it difficult to automatically perform high-quality, personalized video editing that satisfies users.
[0280] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0281] In this invention, the server includes means for a user to upload video data, means for saving the received video data, means for analyzing the content of the video data using an image analysis module, means for defining requirements for music generation using an emotion engine that evaluates the analysis results and the user's emotional state, means for the user to specify a desired music genre, means for generating appropriate music based on the analysis results, the emotional state, and the specified genre using a music generation module, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user. This allows the user to automatically generate appropriate music for specific video data and personalize the music according to the user's emotions and preferences.
[0282] A "user" is an individual or corporation that operates the system and uploads video data or specifies music genres.
[0283] "Video data" refers to video files uploaded by users to the system, including metadata such as resolution, frame rate, and file format.
[0284] A "server" is a computer system that receives, stores, analyzes, and generates music from video data.
[0285] An "image analysis module" is a software component that analyzes each frame of video data and extracts features for each scene.
[0286] "Metadata" is attribute information about video data, and includes resolution, frame rate, file format, and the like.
[0287] An "emotion engine" is an algorithm or software component that analyzes a user's facial expressions and tone of voice to assess their emotional state.
[0288] The "music generation module" is a software component that automatically generates music based on the results of image analysis, emotional state, and specified music genre.
[0289] "Music genre" refers to the style or type of music designated by the user, examples of which include rock, jazz, classical, etc.
[0290] "Synchronization" refers to the process of adjusting the timing of the generated music to match each scene of the video data and integrating them together.
[0291] "Generated video data" refers to a moving image file in which the generated music is added to the original video data and editing is complete.
[0292] "Upload" refers to the operation by which a user sends video data from their own terminal to a server.
[0293] "Providing" refers to the server presenting the generated video data to the user in a downloadable form.
[0294] The present invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by a user. This system is composed of a combination of an image analysis module for analyzing the content of the video data, a music generation module for generating appropriate music, and an emotion engine for recognizing the user's emotions. Specific embodiments for implementing the present invention are described below.
[0295] System configuration
[0296] Hardware and Software
[0297] 1. The server is primarily responsible for the following:
[0298] Receiving and storing video data
[0299] Image analysis
[0300] Emotion analysis
[0301] Music Generation
[0302] Synchronization of video data and music
[0303] Provision of edited video data
[0304] Examples of use include AWS EC2 servers, Python programs, and the ffmpeg library.
[0305] 2. A terminal is a device (smartphone, tablet, PC, etc.) operated by a user that has the following functions:
[0306] Uploading video data
[0307] Specifying the music genre
[0308] Collecting Emotional Data
[0309] A concrete example would be a mobile application using React Native.
[0310] Operation flow
[0311] The system begins operation when the user selects video data using a dedicated application and clicks the upload button. The device then sends the selected video data to the server. The server then saves the received video data in the "temporary / uploads / " directory and uses the ffmpeg library to obtain the metadata of the video data.
[0312] After saving is complete, the server launches an image analysis module (e.g., OpenCV or TensorFlow library) to analyze each frame and extract features for each scene (motion detection, face recognition, type of scenery, etc.), thereby providing a detailed understanding of the contents of the video data.
[0313] Based on the analysis results, the system defines the music generation requirements for each scene. For example, calm music is required for quiet landscape scenes, while energetic music is required for active scenes. Next, the user selects the desired music genre. For example, rock, jazz, classical, etc. can be selected.
[0314] Furthermore, the emotion engine analyzes the user's facial expressions and tone of voice to assess their emotional state. For example, if the user is smiling, the emotional state is determined to be "joy." Analysis data is collected from the camera and microphone and processed using OpenVINO and Microsoft Azure facial recognition APIs.
[0315] The server passes the analysis results, the user's emotional state, and genre selection to a music generation module. The music generation module uses libraries such as Magenta to automatically generate appropriate music based on this data. The generated music is synchronized with the video data by the server, and the timeline is adjusted using the Python moviepy library to complete the video. Finally, the edited video data is encoded, and a download link is generated for the user.
[0316] Specific examples and prompts
[0317] Specific examples
[0318] If user A wants to create a virtual travel video and add rock-style background music to it, he can use the system by following the steps below:
[0319] 1. User A uses the dedicated app to select and upload a travel video file. The device sends the video data to the server storage.
[0320] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0321] 3. The server passes the video data to the image analysis module, which extracts scene features.
[0322] 4. User A selects a rock-style music genre in the application, and the emotion while operating the video is determined to be joy.
[0323] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which automatically generates rock-style music that is suitable for joy.
[0324] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[0325] 7. User A downloads the completed video from the application and can watch it.
[0326] Prompt Sentence Examples
[0327] "Automatically generate music that matches the video data based on the user's facial recognition data."
[0328] "When a user uploads a video containing a tranquil landscape scene, it generates calming music that matches the scene."
[0329] "Analyze the user's emotions and automatically generate and add rock-style music that matches the video data."
[0330] In this way, by utilizing the system of the present invention, users can easily and quickly automatically add suitable music to video data, thereby creating personalized video content.
[0331] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0332] Step 1:
[0333] The user selects the video data using a dedicated application and presses the upload button. This causes the device to send the video data (input: video data file) to the server (output: video data sent to server). A progress bar is displayed on the device to show the user the progress of the upload.
[0334] Step 2:
[0335] The server saves the received video data (input: video data file) in a temporary storage area ("temporary / uploads / " directory) (output: saved video data). Next, the server uses the ffmpeg library to obtain the video data's metadata (resolution, frame rate, file format) (output: metadata).
[0336] Step 3:
[0337] The server starts an image analysis module (e.g., OpenCV, TensorFlow) and passes the saved video data (input: video data file) to the image analysis module (output: analysis results). The image analysis module analyzes each frame of the video and extracts features for each scene (motion detection, face recognition, type of scenery) (output: feature data for each scene).
[0338] Step 4:
[0339] Based on the analysis results (input: feature data for each scene), the server defines the music generation requirements appropriate for each scene (output: music generation requirements). For example, it defines calm music for quiet scenes and energetic music for active scenes.
[0340] Step 5:
[0341] The user specifies the desired music genre (input: music genre specification) through the application (output: selected music genre). The user selects the genre using a pull-down menu or radio buttons.
[0342] Step 6:
[0343] The server launches an emotion engine (e.g., OpenVINO, facial recognition API) and collects the user's facial expressions and tone of voice using a camera or microphone (input: facial expression data, voice data). The emotion engine analyzes this data and evaluates the user's emotional state (output: emotional state data). For example, smiling data is recognized as "joy."
[0344] Step 7:
[0345] The server passes the analysis results, emotional state, and genre designation to the music generation module (input: feature data for each scene, emotional state data, and music genre designation).The music generation module uses the Magenta library and other tools to automatically generate optimal music based on this information (output: generated music file).
[0346] Step 8:
[0347] The generated music (input: music file) is synchronized with the video data by the server. Specifically, the timing of the music is adjusted to match each scene using the Python moviepy library, and the integrated edited video data (output: edited video data) is generated.
[0348] Step 9:
[0349] The server encodes the edited video data (input: edited video data) and generates a download link (output: download link) that the user can use to download the video. The link is stored in a cloud storage service (e.g., AWS S3) and displayed in the user's application.
[0350] Step 10:
[0351] Users can click the download link from the application to download the edited video data (input: download link) and watch it (output: downloaded video data). This allows users to easily enjoy personalized video content.
[0352] (Application example 2)
[0353] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0354] When automatically generating and adding appropriate music to video data created by a user, there is a demand for providing personalized music that reflects the user's preferences and emotional state. However, conventional systems have difficulty generating music that fully takes the user's emotional state into consideration, and the automatic generation and synchronization of music appropriate for video data is complex and time-consuming. To address these issues, the present invention aims to improve the user experience by providing optimal music for video data in a more intuitive and simple manner.
[0355] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0356] In this invention, the server includes means for a user to upload video data, means for storing the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, means for providing the generated video data to the user, means for specifying a desired music genre from the user's terminal, and means for recognizing facial expressions and voices and analyzing the emotional state when the user operates the video data. This makes it possible to provide a system that automatically and personalizedly generates music optimal for video data and is easy for users to use.
[0357] A "user" is an individual or organization who uploads video data to use this system and wishes to generate music.
[0358] "Video data" refers to digital data including video files that users have shot, edited, and saved, as well as their metadata.
[0359] A "server" is a computer system that stores video data received from users and performs the analysis and music generation process.
[0360] An "image analysis module" is a program or algorithm that analyzes each frame of video data and extracts scene features (e.g., motion detection, landscape type, facial recognition, etc.).
[0361] The "music generation module" is a program or algorithm for automatically generating appropriate music based on the results of image analysis and the user's wishes and emotional state.
[0362] "Emotional state" refers to the user's psychological and emotional state (e.g., joy, surprise, sadness, etc.) obtained by analyzing the user's facial expressions, voice, and text input.
[0363] "Synchronization" is the process of adjusting the timing of the generated music to match the scenes in the video data, so that the music and video are integrated.
[0364] "Genre" is a classification that refers to the style or type of music, and represents the kind of music a user desires (for example, rock, jazz, classical, etc.).
[0365] "Upload" is the process of sending video data from a user's terminal to a server.
[0366] "Recognition" is the process of collecting a user's facial expressions and voice using a camera and microphone, and analyzing them to assess their emotional state.
[0367] The system for implementing this invention is configured using the following hardware and software. The hardware requires a user terminal (such as a smartphone), a server, a camera, and a microphone. The software requires an image analysis module, a music generation module, an emotion engine, and an upload and synchronization management program.
[0368] First, the user uploads the video data they created using a device with a dedicated application installed. This device is equipped with a camera and microphone, which can collect the user's facial expressions and voice. The application also includes an interface for the user to specify the music genre they want.
[0369] When a user uploads video data from their device, the server stores the received video data in a storage directory (e.g., the "uploads" directory). The server then launches an image analysis module to analyze each frame of the video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). This allows the content of the video data to be understood in detail.
[0370] After obtaining the results of the image analysis, the server uses the emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice. For example, it analyzes their facial expressions and tone of voice while they are operating the video data to determine whether they are enjoying, surprised, or sad.
[0371] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification to automatically generate optimal music. The tempo and tone of the music can be dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright and fast-paced music is generated, and if the user is calm, calm and slow music is generated.
[0372] The server synchronizes the generated music with the timeline of the video data. Specifically, the timing of the generated music is adjusted to match each scene in the video data, thereby generating edited video data in which the music and video are integrated.
[0373] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data via the application.
[0374] Examples of concrete examples and prompts
[0375] Examples:
[0376] "A user used a smartphone app to upload a video taken at the beach during their summer vacation. The app specified 'bossa nova' as the user's desired music genre and recognized the emotion of joy from the user's facial expression while interacting with the video. The app then generated bossa nova music that matched the serene beach scene and added it to the video."
[0377] Example prompt sentence:
[0378] "I want to edit a video I shot at the beach during my summer vacation. I want to generate bossa nova music that matches the emotion of joy and add it to the video."
[0379] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0380] Step 1:
[0381] The user launches the dedicated application on the device, selects video data, and specifies the desired music genre.
[0382] Input: Video data file, desired music genre
[0383] Output: Video data file path, music genre information
[0384] How it works: A user selects a video file from their smartphone's gallery and clicks the "Upload" button through the app's interface, along with specifying the desired music genre, such as "Bossa Nova."
[0385] Step 2:
[0386] The terminal transmits the selected video data to the server, which stores the received video data in a storage directory.
[0387] Input: Video data file path
[0388] Output: Video data storage directory
[0389] Specific operation: The device uploads the video file to the server using a file transfer protocol, and the server stores the video data in the "uploads" directory.
[0390] Step 3:
[0391] The server launches an image analysis module, analyzes each frame of the video data, and extracts features for each scene.
[0392] Input: saved video data file, image analysis module
[0393] Output: Feature data for each scene
[0394] Specific operation: The server analyzes each frame of video data using Python scripts and extracts features such as scene motion detection, landscape type, and face recognition.
[0395] Step 4:
[0396] The server uses an emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice.
[0397] Input: facial expression data, voice data, music genre information
[0398] Output: Emotional state data
[0399] Specific operation: The server analyzes the camera and microphone data sent from the device, evaluates facial expressions and vocal tone, and uses an emotion engine to estimate the emotional state.
[0400] Step 5:
[0401] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification, and automatically generates optimal music.
[0402] Input: Scene feature data, emotional state data, music genre information
[0403] Output: Generated music file
[0404] Specific operation: Using a music generation AI model, music optimized for the characteristics of each scene and the user's emotions is generated and saved as a music file.
[0405] Step 6:
[0406] The server synchronizes the generated music with the timeline of the video data.
[0407] Input: Video data file, generated music file
[0408] Output: Edited video data file
[0409] What it does: The server uses a video editing library (such as MoviePy) to synchronize the music and video on the timeline and generate the final edited video file.
[0410] Step 7:
[0411] The server encodes the edited video data and generates a link that users can download.
[0412] Input: Edited video data file
[0413] Output: Download link
[0414] Specific operation: The server encodes the edited video data into a distributable format such as MP4 and provides the user with a download link, allowing them to watch the video.
[0415] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0416] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0417] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0418] [Second embodiment]
[0419] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0420] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0421] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0422] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0423] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0424] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0425] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0426] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0427] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0428] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0429] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0430] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0431] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data. Each step of this system is described in detail below.
[0432] First, the user uploads video data from their own device to the server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[0433] The server then stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[0434] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[0435] The results of the image analysis module are returned to the server, which then determines the appropriate music generation requirements for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0436] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[0437] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[0438] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0439] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[0440] Specific examples
[0441] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[0442] 1. User A uses a dedicated app to upload a travel video file to the server.
[0443] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0444] 3. The server passes the saved video file to the image analysis module, which extracts the features of each scene.
[0445] 4. User A selects a rock-style music genre in the application and sends it to the server.
[0446] 5. The server passes the analysis results and genre specification to the music generation module, which then generates rock-style music.
[0447] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[0448] 7. User A can download the completed video from the application and watch it.
[0449] In this way, the system of the present invention allows the user to automatically assign suitable music to video data simply and quickly.
[0450] The processing flow will be explained below.
[0451] Step 1:
[0452] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[0453] Step 2:
[0454] The server detects the received video data and stores it in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0455] Step 3:
[0456] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[0457] Step 4:
[0458] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[0459] Step 5:
[0460] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[0461] Step 6:
[0462] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates music appropriate for each scene based on the specified requirements. This music is then adjusted to create a harmonious overall composition.
[0463] Step 7:
[0464] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[0465] Step 8:
[0466] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[0467] Specific operation example
[0468] Step 1:
[0469] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[0470] Step 2:
[0471] The server receives the video data and saves it in the "temporary / uploads / " folder.
[0472] Step 3:
[0473] The server starts the image analysis module and analyzes the video data. The analysis module extracts the features of each frame.
[0474] Step 4:
[0475] The analysis module returns the extracted features to the server, which then sets the music generation requirements appropriate for each scene based on the features.
[0476] Step 5:
[0477] User A selects the music genre "rock" from the application and sends it to the server. The server receives this information.
[0478] Step 6:
[0479] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates rock-style music.
[0480] Step 7:
[0481] The server synchronizes the generated rock music with the timeline of the video data.
[0482] Step 8:
[0483] The server encodes the final edited video data and provides a downloadable link to User A. User A downloads the edited video from the link and watches it.
[0484] Example 1
[0485] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0486] Conventional video editing systems require users to manually select and edit music appropriate for the video, which requires a great deal of time and effort. It is also not easy to select music that matches the different atmospheres of each video scene. Furthermore, synchronizing music and video requires advanced editing skills, making it difficult for average users to create the high-quality video content desired quickly and easily.
[0487] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0488] In this invention, the server includes a means for a user to upload video data, a means for the server to store the received video data, a means for the server to analyze the content of the video data using an image analysis module, a means for the server to generate appropriate music using a music generation module based on the analysis results, a means for synchronizing the generated music with the video data, and a means for providing the generated video data to the user. This enables the user to quickly create high-quality video content using AI technology without any manual effort.
[0489] "User" refers to an entity that uses the system to upload video data and request music creation.
[0490] "Server" refers to a computer system that performs a series of processes including receiving, storing, analyzing, and generating music from video data.
[0491] "Video data" refers to video files uploaded by users and supports a variety of formats.
[0492] "Uploading means" refers to an application or interface that allows a user to send video data to the server.
[0493] The "storage means" refers to a storage area in which the server temporarily stores the video data received.
[0494] An "image analysis module" refers to a software component that analyzes the contents of video data and extracts features for each scene.
[0495] "Music generation module" refers to a software component that automatically generates music suitable for video based on analysis results.
[0496] "Synchronization means" refers to a process that performs operations to integrate the generated music with the video data.
[0497] "Providing means" refers to a method for providing the generated video data to the user.
[0498] "Metadata" refers to information contained in video data (e.g., resolution, frame rate, file format, etc.).
[0499] "Generative AI model" refers to the artificial intelligence algorithm used by the music generation module, enabling the automatic generation of music.
[0500] A "prompt" is an instruction given to a generative AI model that describes the requirements for generating specific music.
[0501] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data.
[0502] First, the user uploads video data from their device to the server using a dedicated application. Video data can be video files in various formats, ranging from a few minutes to several hours in length. To do this, the user clicks the "Upload" button in the application and selects the appropriate video file from the file selection dialog. The video is then uploaded to the server via cloud storage.
[0503] Next, the server stores the received video data in a temporary storage area (directory "temporary / uploads / "). At the same time, it obtains the video data's metadata (resolution, frame rate, file format, etc.) to prepare for subsequent analysis.
[0504] The server launches an image analysis module to analyze the stored video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). This analysis uses common libraries such as OpenCV and TensorFlow. The analyzed data is temporarily stored in intermediate data storage.
[0505] Based on the analysis results from the image analysis module, the server then defines the music generation requirements for each scene. For example, a quiet landscape scene requires calm music, while an active scene requires energetic music. These defined requirements are recorded in JSON format.
[0506] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through a dedicated application, selects the desired genre from a drop-down menu in the application, and sends the information to the server.
[0507] The server passes the analysis results and genre specification to the music generation module, which uses AI algorithms to automatically generate optimal music based on this information. This generation process can utilize OpenAI's MuseNet or other generative AI models.
[0508] The generated music is then synchronized with the video data by a server. During the synchronization process, the timing of the music is adjusted to match each scene in the video data, resulting in a harmonious overall music and video. Specifically, the music track is merged into the video file using libraries such as FFmpeg.
[0509] Finally, the server encodes the edited video data and generates a download link for the user, who can click the "Download" button in the application to download and watch the completed video.
[0510] Specific examples
[0511] For example, if user A wants to create a virtual travel video and add rock-style background music to it, he or she would follow the steps below.
[0512] 1. User A clicks the "Upload" button in the application, selects a travel video file, and uploads it to the server.
[0513] 2. The server receives and stores the video file and retrieves the metadata.
[0514] 3. The server uses OpenCV to analyze the video data and extract features for each scene.
[0515] 4. User A selects a rock-style music genre in the application and sends it to the server.
[0516] 5. The server passes the analysis results and genre specification to the music generation module, which generates rock-style music using MuseNet.
[0517] 6. The server uses FFmpeg to merge the generated music into the video file and generate the edited video.
[0518] 7. User A downloads the completed video from the application and watches it.
[0519] Prompt Sentence Examples
[0520] "Analyze the footage of the virtual trip and add calming music to the quiet scenes and energetic music to the dynamic scenes. Overall, keep the music taste rock-like."
[0521] As described above, the present invention enables a user to automatically assign suitable music to video data simply and quickly.
[0522] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0523] Step 1:
[0524] The user uploads the video data.
[0525] Specifically, the user launches the dedicated application, clicks the "Upload" button, and selects the video file they want to upload from the file selection dialog. The input is the user's video data, and the output is the video data sent to the server via cloud storage. The video data can be in a variety of video formats, ranging from a few minutes to a few hours.
[0526] Step 2:
[0527] The server stores the received video data.
[0528] Specifically, the server stores the video data in the "temporary / uploads / " directory. At this time, the server acquires the video data's metadata (resolution, frame rate, file format, etc.). The input is the uploaded video data, and the output is the saved data and the acquired metadata. This makes subsequent analysis easier.
[0529] Step 3:
[0530] The server starts the image analysis module and analyzes the video data.
[0531] Specifically, the server uses libraries such as OpenCV and TensorFlow to analyze each frame of video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). The input is the saved video data, and the output is the analysis results (features for each scene). These analysis results are temporarily stored in intermediate data storage.
[0532] Step 4:
[0533] The server defines the requirements for music generation based on the analysis results.
[0534] Specifically, based on the data obtained from the analysis module, the music specifications appropriate for each scene (e.g., gentle music for quiet scenes, energetic music for active scenes) are written in JSON format. The input is the analysis results, and the output is JSON data that records the requirements for music generation.
[0535] Step 5:
[0536] The user specifies the genre of music.
[0537] Specifically, the user selects the desired music genre (e.g., rock, jazz, etc.) from a drop-down menu in the application and sends that information to the server. The input is the user's genre selection, and the output is the genre information sent to the server.
[0538] Step 6:
[0539] The server passes the analysis results and genre designation to the music generation module.
[0540] Specifically, it uses a generative AI model such as OpenAI's MuseNet to generate music based on the analysis results and the user's genre selection. The input is the analysis results and genre information, and the output is the generated music data. Based on this information, the music generation module automatically generates music customized for each scene.
[0541] Step 7:
[0542] The server synchronizes the generated music with the video data.
[0543] Specifically, we use libraries such as FFmpeg to merge the generated music with the video data and adjust the timing. The input is the video data and the generated music, and the output is edited video data synchronized with the music. In this editing process, we optimize the timing of the music for each scene.
[0544] Step 8:
[0545] The server provides the edited video data to the user.
[0546] Specifically, the edited video data is encoded and a download link is generated. Users can click the "Download" button in the application to download and watch the completed video. The input is the edited video data, and the output is the download link.
[0547] (Application example 1)
[0548] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0549] In the past, adding appropriate music to video data created by users required manual selection of music and editing to match the video, which required a great deal of time and effort. Furthermore, the video and music often did not harmonize, making it difficult to create content that is visually and aurally unified. To solve this problem, there is a demand for a system that can automatically generate appropriate music based on the content of the video data and easily provide content in which the video and music are in harmony.
[0550] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0551] In this invention, the server includes means for a user to upload video data, means for the server to store the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user via a terminal such as a smartphone, thereby enabling the user to quickly and easily create video data with optimal music added to the video, and to provide content that is visually and aurally harmonious.
[0552] "User" means a person or legal entity that has the right to use the system to upload video data and generate music.
[0553] "Video data" refers to all video files uploaded by users, the contents of which are analyzed to generate music.
[0554] "Upload" refers to the act of a user sending video data from their own device to a server.
[0555] The "server" is a central computer system that receives and stores video data, and processes it for analysis and music generation.
[0556] An "image analysis module" is a software component that analyzes video data and understands its contents.
[0557] A "music generation module" is a software component for automatically generating music based on the analysis results.
[0558] "Generating" "music" refers to the process by which the system automatically creates new music.
[0559] "Synchronization" refers to the process of matching the generated music to the appropriate timing for each scene in the video data.
[0560] "Devices such as smartphones" refer to portable information and communication devices used by users, which are capable of uploading video data and downloading generated video data.
[0561] "Generated video data" refers to the final video file that has been analyzed and processed to generate music, and is synchronized with the music.
[0562] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data with a music generation module to generate original music according to the user's preferences and add it to the video data.
[0563] First, a user uploads video data from their smartphone or other device to a server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[0564] The server stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[0565] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[0566] The results of the image analysis module are returned to the server, which then determines the appropriate music generation requirements for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0567] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[0568] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[0569] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0570] Finally, the server encodes the edited video data and generates a download link for users, who can then download and watch the completed video on their smartphones or other devices.
[0571] For example, if a user creates a travel video and wants to add rock-style background music to it, they can use this system as follows:
[0572] Example prompt sentence:
[0573] Travel Videos
[0574] Genre: Rock
[0575] Description: Adventurous scenery, mountain climbing, river rafting
[0576] Based on these prompts, the system executes each step and provides the user with a final edited video that is visually and audibly consistent.
[0577] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0578] Step 1:
[0579] The user uploads video data from their smartphone or other device. The input is a video file selected by the user, and the output is that video file is sent to the server. At this time, the user also specifies the desired music genre.
[0580] Step 2:
[0581] The server saves the received video data. It stores the video file in a temporary storage area and obtains the file metadata (resolution, frame rate, file format, etc.). The input is the video data received in step 1, and the output is the video data and its metadata stored in the storage area.
[0582] Step 3:
[0583] The server uses an image analysis module to analyze the content of the stored video data. The input is the video data stored in the storage area, and the output is the extracted results of features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). Specifically, it analyzes each frame of the video data and recognizes specific patterns and objects.
[0584] Step 4:
[0585] Based on the analysis results, the server defines the requirements for music generation appropriate for each scene in the video data. The input is the analysis results obtained in step 3, and the output is the requirements necessary for music generation (for example, calm music for quiet scenes, energetic music for active scenes). Specifically, the server analyzes the analysis results and determines the style and atmosphere of the music that suits each scene.
[0586] Step 5:
[0587] The user specifies the desired music genre through the application. The input is the music genre information specified by the user, and the output is that information is sent to the server.
[0588] Step 6:
[0589] The server passes the analysis results and genre specification to the music generation module. The input is the analysis results and the user's genre specification information, and the output is the input data for the music generation module.
[0590] Step 7:
[0591] The music generation module automatically generates optimal music. The input is the data obtained in step 6, and the output is the generated music. Specifically, an AI model (using TensorFlow, for example) generates music data based on the analysis results and genre information.
[0592] Step 8:
[0593] The generated music is synchronized with the video data. The server adjusts the timing of the music to match each scene in the video data. The input is the generated music and video data, and the output is edited video data in which the music and video are synchronized. Specifically, the start and end points of the music are adjusted for each scene to create integrated data.
[0594] Step 9:
[0595] The server encodes the edited video data and generates a link that allows the user to download it. The input is the edited video data generated in step 8, and the output is a download link. Specifically, the server encodes the data and provides the link to the user.
[0596] Step 10:
[0597] The user downloads the completed video from a device such as a smartphone and watches it. The input is a download link provided by the server, and the output is the final video data with music added, stored on the user's device.
[0598] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0599] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data, a music generation module for generating music, and an emotion engine for recognizing the user's emotions to generate original music according to the user's preferences and add it to the video data. Each step of this system is described in detail below.
[0600] First, the user selects video data using a dedicated application on their own device and clicks the upload button. The device then sends the selected video data file to the server.
[0601] Next, the server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0602] After receiving and saving the video data, the server starts the image analysis module and passes the saved video data to it. The image analysis module analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.). This allows for a detailed understanding of the content of the video data.
[0603] The results of the image analysis module are returned to the server, which then determines the appropriate music generation requirements for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0604] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and adds it to the requirements for music generation.
[0605] The present invention also includes an emotion engine. The emotion engine analyzes the user's facial expressions, voice, and text input to evaluate the user's emotional state. The emotion engine analyzes the data collected by a camera and microphone, including facial expressions and tone of voice, while the user is manipulating video data. The results of this analysis (the user's emotional state, such as whether they are enjoying, surprised, or sad) are reflected in the parameters for generating music.
[0606] The server then passes the analysis results, the user's emotional state, and the genre designation to the music generation module. The music generation module automatically generates optimal music based on this information. The tempo, tone, and genre of the music are dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright, fast-paced music is generated, while if the user is calm, calm, slow music is generated.
[0607] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0608] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[0609] Specific examples
[0610] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[0611] 1. User A uses the dedicated app to select a travel video file and presses the upload button. The device sends the video data to the server storage.
[0612] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0613] 3. The server passes the video data to the image analysis module, which extracts the features of each scene.
[0614] 4. User A specified a rock-style music genre in the application, and the emotion he felt when operating the video was determined to be joy.
[0615] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[0616] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[0617] 7. User A can download the completed video from the application and watch it.
[0618] In this way, the system of the present invention allows users to easily and quickly automatically assign suitable music to video data, and can provide more personalized music depending on the user's emotional state.
[0619] The processing flow will be explained below.
[0620] Step 1:
[0621] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[0622] Step 2:
[0623] The server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0624] Step 3:
[0625] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[0626] Step 4:
[0627] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[0628] Step 5:
[0629] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[0630] Step 6:
[0631] When a user interacts with video data (e.g., through facial expressions or voice input), the emotion engine evaluates the user's emotional state. The user's facial expressions and voice are recorded via the device's camera and microphone, and then analyzed by the emotion engine.
[0632] Step 7:
[0633] The emotion engine sends the analysis results to the server, which then incorporates this emotion information (e.g., whether the user is happy or surprised) into the requirements for music generation.
[0634] Step 8:
[0635] The server passes the analysis results, the user's emotional state, and genre information to the music generation module. The music generation module automatically generates optimal music based on this information. If the user is happy, it generates bright, fast-paced music, and if the user is calm, it generates calm, slow music.
[0636] Step 9:
[0637] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[0638] Step 10:
[0639] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[0640] Specific operation example
[0641] Step 1:
[0642] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[0643] Step 2:
[0644] The server receives the video data and saves it in the "temporary / uploads / " directory.
[0645] Step 3:
[0646] The server starts an image analysis module, analyzes each frame of the video data, and extracts features.
[0647] Step 4:
[0648] The image analysis module returns the analysis results to the server, which then defines the necessary music generation requirements.
[0649] Step 5:
[0650] User A selects the music genre "rock" from the application and sends it to the server.
[0651] Step 6:
[0652] When User A is operating the video, the device sends data to the emotion engine based on User A's facial expressions and voice. The emotion engine then analyzes User A's emotions.
[0653] Step 7:
[0654] The emotion engine sends the emotional state of user A (e.g., joy) to the server.
[0655] Step 8:
[0656] The server passes the analysis results, emotional state, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[0657] Step 9:
[0658] The server synchronizes the generated music with the video timeline to generate an edited video.
[0659] Step 10:
[0660] The server provides the edited video in a downloadable format to User A. User A then downloads the completed video from the application and watches it.
[0661] Example 2
[0662] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0663] Conventional video editing systems require users to individually select appropriate music and manually synchronize it with the video data, which requires time and effort. Furthermore, even systems capable of automatically generating music for video data lack the functionality to personalize the music based on the user's emotions and preferences. This makes it difficult to automatically perform high-quality, personalized video editing that satisfies users.
[0664] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0665] In this invention, the server includes means for a user to upload video data, means for saving the received video data, means for analyzing the content of the video data using an image analysis module, means for defining requirements for music generation using an emotion engine that evaluates the analysis results and the user's emotional state, means for the user to specify a desired music genre, means for generating appropriate music based on the analysis results, the emotional state, and the specified genre using a music generation module, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user. This allows the user to automatically generate appropriate music for specific video data and personalize the music according to the user's emotions and preferences.
[0666] A "user" is an individual or corporation that operates the system and uploads video data or specifies music genres.
[0667] "Video data" refers to video files uploaded by users to the system, including metadata such as resolution, frame rate, and file format.
[0668] A "server" is a computer system that receives, stores, analyzes, and generates music from video data.
[0669] An "image analysis module" is a software component that analyzes each frame of video data and extracts features for each scene.
[0670] "Metadata" is attribute information about video data, and includes resolution, frame rate, file format, and the like.
[0671] An "emotion engine" is an algorithm or software component that analyzes a user's facial expressions and tone of voice to assess their emotional state.
[0672] The "music generation module" is a software component that automatically generates music based on the results of image analysis, emotional state, and specified music genre.
[0673] "Music genre" refers to the style or type of music designated by the user, examples of which include rock, jazz, classical, etc.
[0674] "Synchronization" refers to the process of adjusting the timing of the generated music to match each scene of the video data and integrating them together.
[0675] "Generated video data" refers to a moving image file in which the generated music is added to the original video data and editing is complete.
[0676] "Upload" refers to the operation by which a user sends video data from their own terminal to a server.
[0677] "Providing" refers to the server presenting the generated video data to the user in a downloadable form.
[0678] The present invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by a user. This system is composed of a combination of an image analysis module for analyzing the content of the video data, a music generation module for generating appropriate music, and an emotion engine for recognizing the user's emotions. Specific embodiments for implementing the present invention are described below.
[0679] System configuration
[0680] Hardware and Software
[0681] 1. The server is primarily responsible for the following:
[0682] Receiving and storing video data
[0683] Image analysis
[0684] Emotion analysis
[0685] Music Generation
[0686] Synchronization of video data and music
[0687] Provision of edited video data
[0688] Examples of use include AWS EC2 servers, Python programs, and the ffmpeg library.
[0689] 2. A terminal is a device (smartphone, tablet, PC, etc.) operated by a user that has the following functions:
[0690] Uploading video data
[0691] Specifying the music genre
[0692] Collecting Emotional Data
[0693] A concrete example would be a mobile application using React Native.
[0694] Operation flow
[0695] The system begins operation when the user selects video data using a dedicated application and clicks the upload button. The device then sends the selected video data to the server. The server then saves the received video data in the "temporary / uploads / " directory and uses the ffmpeg library to obtain the metadata of the video data.
[0696] After saving is complete, the server launches an image analysis module (e.g., OpenCV or TensorFlow library) to analyze each frame and extract features for each scene (motion detection, face recognition, type of scenery, etc.), thereby providing a detailed understanding of the contents of the video data.
[0697] Based on the analysis results, the system defines the music generation requirements for each scene. For example, calm music is required for quiet landscape scenes, while energetic music is required for active scenes. Next, the user selects the desired music genre. For example, rock, jazz, classical, etc. can be selected.
[0698] Furthermore, the emotion engine analyzes the user's facial expressions and tone of voice to assess their emotional state. For example, if the user is smiling, the emotional state is determined to be "joy." Analysis data is collected from the camera and microphone and processed using OpenVINO and Microsoft Azure facial recognition APIs.
[0699] The server passes the analysis results, the user's emotional state, and genre selection to a music generation module. The music generation module uses libraries such as Magenta to automatically generate appropriate music based on this data. The generated music is synchronized with the video data by the server, and the timeline is adjusted using the Python moviepy library to complete the video. Finally, the edited video data is encoded, and a download link is generated for the user.
[0700] Specific examples and prompts
[0701] Specific examples
[0702] If user A wants to create a virtual travel video and add rock-style background music to it, he can use the system by following the steps below:
[0703] 1. User A uses the dedicated app to select and upload a travel video file. The device sends the video data to the server storage.
[0704] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0705] 3. The server passes the video data to the image analysis module, which extracts scene features.
[0706] 4. User A selects a rock-style music genre in the application, and the emotion while operating the video is determined to be joy.
[0707] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which automatically generates rock-style music that is suitable for joy.
[0708] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[0709] 7. User A downloads the completed video from the application and can watch it.
[0710] Prompt Sentence Examples
[0711] "Automatically generate music that matches the video data based on the user's facial recognition data."
[0712] "When a user uploads a video containing a tranquil landscape scene, it generates calming music that matches the scene."
[0713] "Analyze the user's emotions and automatically generate and add rock-style music that matches the video data."
[0714] In this way, by utilizing the system of the present invention, users can easily and quickly automatically add suitable music to video data, thereby creating personalized video content.
[0715] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0716] Step 1:
[0717] The user selects the video data using a dedicated application and presses the upload button. This causes the device to send the video data (input: video data file) to the server (output: video data sent to server). A progress bar is displayed on the device to show the user the progress of the upload.
[0718] Step 2:
[0719] The server saves the received video data (input: video data file) in a temporary storage area ("temporary / uploads / " directory) (output: saved video data). Next, the server uses the ffmpeg library to obtain the video data's metadata (resolution, frame rate, file format) (output: metadata).
[0720] Step 3:
[0721] The server starts an image analysis module (e.g., OpenCV, TensorFlow) and passes the saved video data (input: video data file) to the image analysis module (output: analysis results). The image analysis module analyzes each frame of the video and extracts features for each scene (motion detection, face recognition, type of scenery) (output: feature data for each scene).
[0722] Step 4:
[0723] Based on the analysis results (input: feature data for each scene), the server defines the music generation requirements appropriate for each scene (output: music generation requirements). For example, it defines calm music for quiet scenes and energetic music for active scenes.
[0724] Step 5:
[0725] The user specifies the desired music genre (input: music genre specification) through the application (output: selected music genre). The user selects the genre using a pull-down menu or radio buttons.
[0726] Step 6:
[0727] The server launches an emotion engine (e.g., OpenVINO, facial recognition API) and collects the user's facial expressions and tone of voice using a camera or microphone (input: facial expression data, voice data). The emotion engine analyzes this data and evaluates the user's emotional state (output: emotional state data). For example, smiling data is recognized as "joy."
[0728] Step 7:
[0729] The server passes the analysis results, emotional state, and genre designation to the music generation module (input: feature data for each scene, emotional state data, and music genre designation).The music generation module uses the Magenta library and other tools to automatically generate optimal music based on this information (output: generated music file).
[0730] Step 8:
[0731] The generated music (input: music file) is synchronized with the video data by the server. Specifically, the timing of the music is adjusted to match each scene using the Python moviepy library, and the integrated edited video data (output: edited video data) is generated.
[0732] Step 9:
[0733] The server encodes the edited video data (input: edited video data) and generates a download link (output: download link) that the user can use to download the video. The link is stored in a cloud storage service (e.g., AWS S3) and displayed in the user's application.
[0734] Step 10:
[0735] Users can click the download link from the application to download the edited video data (input: download link) and watch it (output: downloaded video data). This allows users to easily enjoy personalized video content.
[0736] (Application example 2)
[0737] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0738] When automatically generating and adding appropriate music to video data created by a user, there is a demand for providing personalized music that reflects the user's preferences and emotional state. However, conventional systems have difficulty generating music that fully takes the user's emotional state into consideration, and the automatic generation and synchronization of music appropriate for video data is complex and time-consuming. To address these issues, the present invention aims to improve the user experience by providing optimal music for video data in a more intuitive and simple manner.
[0739] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0740] In this invention, the server includes means for a user to upload video data, means for storing the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, means for providing the generated video data to the user, means for specifying a desired music genre from the user's terminal, and means for recognizing facial expressions and voices and analyzing the emotional state when the user operates the video data. This makes it possible to provide a system that automatically and personalizedly generates music optimal for video data and is easy for users to use.
[0741] A "user" is an individual or organization who uploads video data to use this system and wishes to generate music.
[0742] "Video data" refers to digital data including video files that users have shot, edited, and saved, as well as their metadata.
[0743] A "server" is a computer system that stores video data received from users and performs the analysis and music generation process.
[0744] An "image analysis module" is a program or algorithm that analyzes each frame of video data and extracts scene features (e.g., motion detection, landscape type, facial recognition, etc.).
[0745] The "music generation module" is a program or algorithm for automatically generating appropriate music based on the results of image analysis and the user's wishes and emotional state.
[0746] "Emotional state" refers to the user's psychological and emotional state (e.g., joy, surprise, sadness, etc.) obtained by analyzing the user's facial expressions, voice, and text input.
[0747] "Synchronization" is the process of adjusting the timing of the generated music to match the scenes in the video data, so that the music and video are integrated.
[0748] "Genre" is a classification that refers to the style or type of music, and represents the kind of music a user desires (for example, rock, jazz, classical, etc.).
[0749] "Upload" is the process of sending video data from a user's terminal to a server.
[0750] "Recognition" is the process of collecting a user's facial expressions and voice using a camera and microphone, and analyzing them to assess their emotional state.
[0751] The system for implementing this invention is configured using the following hardware and software. The hardware requires a user terminal (such as a smartphone), a server, a camera, and a microphone. The software requires an image analysis module, a music generation module, an emotion engine, and an upload and synchronization management program.
[0752] First, the user uploads the video data they created using a device with a dedicated application installed. This device is equipped with a camera and microphone, which can collect the user's facial expressions and voice. The application also includes an interface for the user to specify the music genre they want.
[0753] When a user uploads video data from their device, the server stores the received video data in a storage directory (e.g., the "uploads" directory). The server then launches an image analysis module to analyze each frame of the video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). This allows the content of the video data to be understood in detail.
[0754] After obtaining the results of the image analysis, the server uses the emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice. For example, it analyzes their facial expressions and tone of voice while they are operating the video data to determine whether they are enjoying, surprised, or sad.
[0755] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification to automatically generate optimal music. The tempo and tone of the music can be dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright and fast-paced music is generated, and if the user is calm, calm and slow music is generated.
[0756] The server synchronizes the generated music with the timeline of the video data. Specifically, the timing of the generated music is adjusted to match each scene in the video data, thereby generating edited video data in which the music and video are integrated.
[0757] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data via the application.
[0758] Examples of concrete examples and prompts
[0759] Examples:
[0760] "A user used a smartphone app to upload a video taken at the beach during their summer vacation. The app specified 'bossa nova' as the user's desired music genre and recognized the emotion of joy from the user's facial expression while interacting with the video. The app then generated bossa nova music that matched the serene beach scene and added it to the video."
[0761] Example prompt sentence:
[0762] "I want to edit a video I shot at the beach during my summer vacation. I want to generate bossa nova music that matches the emotion of joy and add it to the video."
[0763] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0764] Step 1:
[0765] The user launches the dedicated application on the device, selects video data, and specifies the desired music genre.
[0766] Input: Video data file, desired music genre
[0767] Output: Video data file path, music genre information
[0768] How it works: A user selects a video file from their smartphone's gallery and clicks the "Upload" button through the app's interface, along with specifying the desired music genre, such as "Bossa Nova."
[0769] Step 2:
[0770] The terminal transmits the selected video data to the server, which stores the received video data in a storage directory.
[0771] Input: Video data file path
[0772] Output: Video data storage directory
[0773] Specific operation: The device uploads the video file to the server using a file transfer protocol, and the server stores the video data in the "uploads" directory.
[0774] Step 3:
[0775] The server launches an image analysis module, analyzes each frame of the video data, and extracts features for each scene.
[0776] Input: saved video data file, image analysis module
[0777] Output: Feature data for each scene
[0778] Specific operation: The server analyzes each frame of video data using Python scripts and extracts features such as scene motion detection, landscape type, and face recognition.
[0779] Step 4:
[0780] The server uses an emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice.
[0781] Input: facial expression data, voice data, music genre information
[0782] Output: Emotional state data
[0783] Specific operation: The server analyzes the camera and microphone data sent from the device, evaluates facial expressions and vocal tone, and uses an emotion engine to estimate the emotional state.
[0784] Step 5:
[0785] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification, and automatically generates optimal music.
[0786] Input: Scene feature data, emotional state data, music genre information
[0787] Output: Generated music file
[0788] Specific operation: Using a music generation AI model, music optimized for the characteristics of each scene and the user's emotions is generated and saved as a music file.
[0789] Step 6:
[0790] The server synchronizes the generated music with the timeline of the video data.
[0791] Input: Video data file, generated music file
[0792] Output: Edited video data file
[0793] What it does: The server uses a video editing library (such as MoviePy) to synchronize the music and video on the timeline and generate the final edited video file.
[0794] Step 7:
[0795] The server encodes the edited video data and generates a link that users can download.
[0796] Input: Edited video data file
[0797] Output: Download link
[0798] Specific operation: The server encodes the edited video data into a distributable format such as MP4 and provides the user with a download link, allowing them to watch the video.
[0799] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0800] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0801] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0802] [Third embodiment]
[0803] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0804] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0805] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0806] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0807] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0808] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0809] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0810] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0811] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0812] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0813] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0814] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0815] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data. Each step of this system is described in detail below.
[0816] First, the user uploads video data from their own device to the server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[0817] The server then stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[0818] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[0819] The analysis results from the image analysis module are returned to the server, which then uses the results to define the music generation requirements appropriate for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0820] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[0821] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[0822] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0823] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[0824] Specific examples
[0825] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[0826] 1. User A uses a dedicated app to upload a travel video file to the server.
[0827] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0828] 3. The server passes the saved video file to the image analysis module, which extracts the features of each scene.
[0829] 4. User A selects a rock-style music genre in the application and sends it to the server.
[0830] 5. The server passes the analysis results and genre specification to the music generation module, which then generates rock-style music.
[0831] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[0832] 7. User A can download the completed video from the application and watch it.
[0833] In this way, the system of the present invention allows the user to automatically assign suitable music to video data simply and quickly.
[0834] The processing flow will be explained below.
[0835] Step 1:
[0836] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[0837] Step 2:
[0838] The server detects the received video data and stores it in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0839] Step 3:
[0840] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[0841] Step 4:
[0842] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[0843] Step 5:
[0844] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[0845] Step 6:
[0846] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates music appropriate for each scene based on the specified requirements. This music is then adjusted to create a harmonious overall composition.
[0847] Step 7:
[0848] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[0849] Step 8:
[0850] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[0851] Specific operation example
[0852] Step 1:
[0853] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[0854] Step 2:
[0855] The server receives the video data and saves it in the "temporary / uploads / " folder.
[0856] Step 3:
[0857] The server starts the image analysis module and analyzes the video data. The analysis module extracts the features of each frame.
[0858] Step 4:
[0859] The analysis module returns the extracted features to the server, which then sets the music generation requirements appropriate for each scene based on the features.
[0860] Step 5:
[0861] User A selects the music genre "rock" from the application and sends it to the server. The server receives this information.
[0862] Step 6:
[0863] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates rock-style music.
[0864] Step 7:
[0865] The server synchronizes the generated rock music with the timeline of the video data.
[0866] Step 8:
[0867] The server encodes the final edited video data and provides a downloadable link to User A. User A downloads the edited video from the link and watches it.
[0868] Example 1
[0869] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0870] Conventional video editing systems require users to manually select and edit music appropriate for the video, which requires a great deal of time and effort. It is also not easy to select music that matches the different atmospheres of each video scene. Furthermore, synchronizing music and video requires advanced editing skills, making it difficult for average users to create the high-quality video content desired quickly and easily.
[0871] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0872] In this invention, the server includes a means for a user to upload video data, a means for the server to store the received video data, a means for the server to analyze the content of the video data using an image analysis module, a means for the server to generate appropriate music using a music generation module based on the analysis results, a means for synchronizing the generated music with the video data, and a means for providing the generated video data to the user. This enables the user to quickly create high-quality video content using AI technology without any manual effort.
[0873] "User" refers to an entity that uses the system to upload video data and request music creation.
[0874] "Server" refers to a computer system that performs a series of processes including receiving, storing, analyzing, and generating music from video data.
[0875] "Video data" refers to video files uploaded by users and supports a variety of formats.
[0876] "Uploading means" refers to an application or interface that allows a user to send video data to the server.
[0877] The "storage means" refers to a storage area in which the server temporarily stores the video data received.
[0878] An "image analysis module" refers to a software component that analyzes the contents of video data and extracts features for each scene.
[0879] "Music generation module" refers to a software component that automatically generates music suitable for video based on analysis results.
[0880] "Synchronization means" refers to a process that performs operations to integrate the generated music with the video data.
[0881] "Providing means" refers to a method for providing the generated video data to the user.
[0882] "Metadata" refers to information contained in video data (e.g., resolution, frame rate, file format, etc.).
[0883] "Generative AI model" refers to the artificial intelligence algorithm used by the music generation module, enabling the automatic generation of music.
[0884] A "prompt" is an instruction given to a generative AI model that describes the requirements for generating specific music.
[0885] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data.
[0886] First, the user uploads video data from their device to the server using a dedicated application. Video data can be video files in various formats, ranging from a few minutes to several hours in length. To do this, the user clicks the "Upload" button in the application and selects the appropriate video file from the file selection dialog. The video is then uploaded to the server via cloud storage.
[0887] Next, the server stores the received video data in a temporary storage area (directory "temporary / uploads / "). At the same time, it obtains the video data's metadata (resolution, frame rate, file format, etc.) to prepare for subsequent analysis.
[0888] The server launches an image analysis module to analyze the stored video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). This analysis uses common libraries such as OpenCV and TensorFlow. The analyzed data is temporarily stored in intermediate data storage.
[0889] Based on the analysis results from the image analysis module, the server then defines the music generation requirements for each scene. For example, a quiet landscape scene requires calm music, while an active scene requires energetic music. These defined requirements are recorded in JSON format.
[0890] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through a dedicated application, selects the desired genre from a drop-down menu in the application, and sends the information to the server.
[0891] The server passes the analysis results and genre specification to the music generation module, which uses this information to automatically generate optimal music using AI algorithms. This generation process can utilize OpenAI's MuseNet or other generative AI models.
[0892] The generated music is then synchronized with the video data by a server. During the synchronization process, the timing of the music is adjusted to match each scene in the video data, resulting in a harmonious overall music and video. Specifically, the music track is merged into the video file using libraries such as FFmpeg.
[0893] Finally, the server encodes the edited video data and generates a download link for the user, who can click the "Download" button in the application to download and watch the completed video.
[0894] Specific examples
[0895] For example, if user A wants to create a virtual travel video and add rock-style background music to it, he or she would follow the steps below.
[0896] 1. User A clicks the "Upload" button in the application, selects a travel video file, and uploads it to the server.
[0897] 2. The server receives and stores the video file and retrieves the metadata.
[0898] 3. The server uses OpenCV to analyze the video data and extract features for each scene.
[0899] 4. User A selects a rock-style music genre in the application and sends it to the server.
[0900] 5. The server passes the analysis results and genre specification to the music generation module, which generates rock-style music using MuseNet.
[0901] 6. The server uses FFmpeg to merge the generated music into the video file and generate the edited video.
[0902] 7. User A downloads the completed video from the application and watches it.
[0903] Prompt Sentence Examples
[0904] "Analyze the footage of the virtual trip and add calming music to the quiet scenes and energetic music to the dynamic scenes. Overall, keep the music taste rock-like."
[0905] As described above, the present invention enables a user to automatically assign suitable music to video data simply and quickly.
[0906] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0907] Step 1:
[0908] The user uploads the video data.
[0909] Specifically, the user launches the dedicated application, clicks the "Upload" button, and selects the video file they want to upload from the file selection dialog. The input is the user's video data, and the output is the video data sent to the server via cloud storage. The video data can be in a variety of video formats, ranging from a few minutes to a few hours.
[0910] Step 2:
[0911] The server stores the received video data.
[0912] Specifically, the server stores the video data in the "temporary / uploads / " directory. At this time, the server acquires the video data's metadata (resolution, frame rate, file format, etc.). The input is the uploaded video data, and the output is the saved data and the acquired metadata. This makes subsequent analysis easier.
[0913] Step 3:
[0914] The server starts the image analysis module and analyzes the video data.
[0915] Specifically, the server uses libraries such as OpenCV and TensorFlow to analyze each frame of video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). The input is the saved video data, and the output is the analysis results (features for each scene). These analysis results are temporarily stored in intermediate data storage.
[0916] Step 4:
[0917] The server defines the requirements for music generation based on the analysis results.
[0918] Specifically, based on the data obtained from the analysis module, the music specifications appropriate for each scene (e.g., gentle music for quiet scenes, energetic music for active scenes) are written in JSON format. The input is the analysis results, and the output is JSON data that records the requirements for music generation.
[0919] Step 5:
[0920] The user specifies the genre of music.
[0921] Specifically, the user selects the desired music genre (e.g., rock, jazz, etc.) from a drop-down menu in the application and sends that information to the server. The input is the user's genre selection, and the output is the genre information sent to the server.
[0922] Step 6:
[0923] The server passes the analysis results and genre designation to the music generation module.
[0924] Specifically, it uses a generative AI model such as OpenAI's MuseNet to generate music based on the analysis results and the user's genre selection. The input is the analysis results and genre information, and the output is the generated music data. Based on this information, the music generation module automatically generates music customized for each scene.
[0925] Step 7:
[0926] The server synchronizes the generated music with the video data.
[0927] Specifically, we use libraries such as FFmpeg to merge the generated music with the video data and adjust the timing. The input is the video data and the generated music, and the output is edited video data synchronized with the music. In this editing process, we optimize the timing of the music for each scene.
[0928] Step 8:
[0929] The server provides the edited video data to the user.
[0930] Specifically, the edited video data is encoded and a download link is generated. Users can click the "Download" button in the application to download and watch the completed video. The input is the edited video data, and the output is the download link.
[0931] (Application example 1)
[0932] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0933] In the past, adding appropriate music to video data created by users required manual selection of music and editing to match the video, which required a great deal of time and effort. Furthermore, the video and music often did not harmonize, making it difficult to create content that is visually and aurally unified. To solve this problem, there is a demand for a system that can automatically generate appropriate music based on the content of the video data and easily provide content in which the video and music are in harmony.
[0934] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0935] In this invention, the server includes means for a user to upload video data, means for the server to store the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user via a terminal such as a smartphone, thereby enabling the user to quickly and easily create video data with optimal music added to the video, and to provide content that is visually and aurally harmonious.
[0936] "User" means a person or legal entity that has the right to use the system to upload video data and generate music.
[0937] "Video data" refers to all video files uploaded by users, the contents of which are analyzed to generate music.
[0938] "Upload" refers to the act of a user sending video data from their own device to a server.
[0939] The "server" is a central computer system that receives and stores video data, and processes it for analysis and music generation.
[0940] An "image analysis module" is a software component that analyzes video data and understands its contents.
[0941] A "music generation module" is a software component for automatically generating music based on the analysis results.
[0942] "Generating" "music" refers to the process by which the system automatically creates new music.
[0943] "Synchronization" refers to the process of matching the generated music to the appropriate timing for each scene in the video data.
[0944] "Devices such as smartphones" refer to portable information and communication devices used by users, which are capable of uploading video data and downloading generated video data.
[0945] "Generated video data" refers to the final video file that has been analyzed and processed to generate music, and is synchronized with the music.
[0946] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data with a music generation module to generate original music according to the user's preferences and add it to the video data.
[0947] First, a user uploads video data from their smartphone or other device to a server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[0948] The server stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[0949] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[0950] The analysis results from the image analysis module are returned to the server, which then uses the results to define the music generation requirements appropriate for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0951] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[0952] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[0953] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0954] Finally, the server encodes the edited video data and generates a download link for users, who can then download and watch the completed video on their smartphones or other devices.
[0955] For example, if a user creates a travel video and wants to add rock-style background music to it, they can use this system as follows:
[0956] Example prompt sentence:
[0957] Travel Videos
[0958] Genre: Rock
[0959] Description: Adventurous scenery, mountain climbing, river rafting
[0960] Based on these prompts, the system executes each step and provides the user with a final edited video that is visually and audibly consistent.
[0961] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0962] Step 1:
[0963] The user uploads video data from their smartphone or other device. The input is a video file selected by the user, and the output is that video file is sent to the server. At this time, the user also specifies the desired music genre.
[0964] Step 2:
[0965] The server saves the received video data. It stores the video file in a temporary storage area and obtains the file metadata (resolution, frame rate, file format, etc.). The input is the video data received in step 1, and the output is the video data and its metadata stored in the storage area.
[0966] Step 3:
[0967] The server uses an image analysis module to analyze the content of the stored video data. The input is the video data stored in the storage area, and the output is the extracted results of features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). Specifically, it analyzes each frame of the video data and recognizes specific patterns and objects.
[0968] Step 4:
[0969] Based on the analysis results, the server defines the requirements for music generation appropriate for each scene in the video data. The input is the analysis results obtained in step 3, and the output is the requirements necessary for music generation (for example, calm music for quiet scenes, energetic music for active scenes). Specifically, the server analyzes the analysis results and determines the style and atmosphere of the music that suits each scene.
[0970] Step 5:
[0971] The user specifies the desired music genre through the application. The input is the music genre information specified by the user, and the output is that information is sent to the server.
[0972] Step 6:
[0973] The server passes the analysis results and genre specification to the music generation module. The input is the analysis results and the user's genre specification information, and the output is the input data for the music generation module.
[0974] Step 7:
[0975] The music generation module automatically generates optimal music. The input is the data obtained in step 6, and the output is the generated music. Specifically, an AI model (using TensorFlow, for example) generates music data based on the analysis results and genre information.
[0976] Step 8:
[0977] The generated music is synchronized with the video data. The server adjusts the timing of the music to match each scene in the video data. The input is the generated music and video data, and the output is edited video data in which the music and video are synchronized. Specifically, the start and end points of the music are adjusted for each scene to create integrated data.
[0978] Step 9:
[0979] The server encodes the edited video data and generates a link that allows the user to download it. The input is the edited video data generated in step 8, and the output is a download link. Specifically, the server encodes the data and provides the link to the user.
[0980] Step 10:
[0981] The user downloads the completed video from a device such as a smartphone and watches it. The input is a download link provided by the server, and the output is the final video data with music added, stored on the user's device.
[0982] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0983] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data, a music generation module for generating music, and an emotion engine for recognizing the user's emotions to generate original music according to the user's preferences and add it to the video data. Each step of this system is described in detail below.
[0984] First, the user selects the video data using a dedicated application on their own device and clicks the upload button. The device then sends the selected video data file to the server.
[0985] Next, the server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[0986] After receiving and saving the video data, the server starts the image analysis module and passes the saved video data to it. The image analysis module analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.). This allows for a detailed understanding of the content of the video data.
[0987] The analysis results from the image analysis module are returned to the server, which then uses the results to define the music generation requirements appropriate for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[0988] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and adds it to the requirements for music generation.
[0989] The present invention also includes an emotion engine. The emotion engine analyzes the user's facial expressions, voice, and text input to evaluate the user's emotional state. The emotion engine analyzes the data collected by a camera and microphone, including facial expressions and tone of voice, while the user is manipulating video data. The results of this analysis (the user's emotional state, such as whether they are enjoying, surprised, or sad) are reflected in the parameters for generating music.
[0990] The server then passes the analysis results, the user's emotional state, and the genre designation to the music generation module. The music generation module automatically generates optimal music based on this information. The tempo, tone, and genre of the music are dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright, fast-paced music is generated, while if the user is calm, calm, slow music is generated.
[0991] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[0992] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[0993] Specific examples
[0994] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[0995] 1. User A uses the dedicated app to select a travel video file and presses the upload button. The device sends the video data to the server storage.
[0996] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[0997] 3. The server passes the video data to the image analysis module, which extracts the features of each scene.
[0998] 4. User A specified a rock-style music genre in the application, and the emotion he felt when operating the video was determined to be joy.
[0999] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[1000] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[1001] 7. User A can download the completed video from the application and watch it.
[1002] In this way, the system of the present invention allows users to easily and quickly automatically assign suitable music to video data, and can provide more personalized music depending on the user's emotional state.
[1003] The processing flow will be explained below.
[1004] Step 1:
[1005] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[1006] Step 2:
[1007] The server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[1008] Step 3:
[1009] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[1010] Step 4:
[1011] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[1012] Step 5:
[1013] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[1014] Step 6:
[1015] When a user interacts with video data (e.g., through facial expressions or voice input), the emotion engine evaluates the user's emotional state. The user's facial expressions and voice are recorded via the device's camera and microphone, and then analyzed by the emotion engine.
[1016] Step 7:
[1017] The emotion engine sends the analysis results to the server, which then incorporates this emotion information (e.g., whether the user is happy or surprised) into the requirements for music generation.
[1018] Step 8:
[1019] The server passes the analysis results, the user's emotional state, and genre information to the music generation module. The music generation module automatically generates optimal music based on this information. If the user is happy, it generates bright, fast-paced music, and if the user is calm, it generates calm, slow music.
[1020] Step 9:
[1021] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[1022] Step 10:
[1023] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[1024] Specific operation example
[1025] Step 1:
[1026] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[1027] Step 2:
[1028] The server receives the video data and saves it in the "temporary / uploads / " directory.
[1029] Step 3:
[1030] The server starts an image analysis module, analyzes each frame of the video data, and extracts features.
[1031] Step 4:
[1032] The image analysis module returns the analysis results to the server, which then defines the necessary music generation requirements.
[1033] Step 5:
[1034] User A selects the music genre "rock" from the application and sends it to the server.
[1035] Step 6:
[1036] When User A is operating the video, the device sends data to the emotion engine based on User A's facial expressions and voice. The emotion engine then analyzes User A's emotions.
[1037] Step 7:
[1038] The emotion engine sends the emotional state of user A (e.g., joy) to the server.
[1039] Step 8:
[1040] The server passes the analysis results, emotional state, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[1041] Step 9:
[1042] The server synchronizes the generated music with the video timeline to generate an edited video.
[1043] Step 10:
[1044] The server provides the edited video in a downloadable format to User A. User A then downloads the completed video from the application and watches it.
[1045] Example 2
[1046] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1047] Conventional video editing systems require users to individually select appropriate music and manually synchronize it with the video data, which requires time and effort. Furthermore, even systems capable of automatically generating music for video data lack the functionality to personalize the music based on the user's emotions and preferences. This makes it difficult to automatically perform high-quality, personalized video editing that satisfies users.
[1048] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1049] In this invention, the server includes means for a user to upload video data, means for saving the received video data, means for analyzing the content of the video data using an image analysis module, means for defining requirements for music generation using an emotion engine that evaluates the analysis results and the user's emotional state, means for the user to specify a desired music genre, means for generating appropriate music based on the analysis results, the emotional state, and the specified genre using a music generation module, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user. This allows the user to automatically generate appropriate music for specific video data and personalize the music according to the user's emotions and preferences.
[1050] A "user" is an individual or corporation that operates the system and uploads video data or specifies music genres.
[1051] "Video data" refers to video files uploaded by users to the system, including metadata such as resolution, frame rate, and file format.
[1052] A "server" is a computer system that receives, stores, analyzes, and generates music from video data.
[1053] An "image analysis module" is a software component that analyzes each frame of video data and extracts features for each scene.
[1054] "Metadata" is attribute information about video data, and includes resolution, frame rate, file format, and the like.
[1055] An "emotion engine" is an algorithm or software component that analyzes a user's facial expressions and tone of voice to assess their emotional state.
[1056] The "music generation module" is a software component that automatically generates music based on the results of image analysis, emotional state, and specified music genre.
[1057] "Music genre" refers to the style or type of music designated by the user, examples of which include rock, jazz, classical, etc.
[1058] "Synchronization" refers to the process of adjusting the timing of the generated music to match each scene of the video data and integrating them together.
[1059] "Generated video data" refers to a video file in which the generated music is added to the original video data and editing is complete.
[1060] "Upload" refers to the operation by which a user sends video data from their own terminal to a server.
[1061] "Providing" refers to the server presenting the generated video data to the user in a downloadable form.
[1062] The present invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by a user. This system is composed of a combination of an image analysis module for analyzing the content of the video data, a music generation module for generating appropriate music, and an emotion engine for recognizing the user's emotions. Specific embodiments for implementing the present invention are described below.
[1063] System configuration
[1064] Hardware and Software
[1065] 1. The server is primarily responsible for the following:
[1066] Receiving and storing video data
[1067] Image analysis
[1068] Emotion analysis
[1069] Music Generation
[1070] Synchronization of video data and music
[1071] Provision of edited video data
[1072] Examples of use include AWS EC2 servers, Python programs, and the ffmpeg library.
[1073] 2. A terminal is a device (smartphone, tablet, PC, etc.) operated by a user that has the following functions:
[1074] Uploading video data
[1075] Specifying the music genre
[1076] Collecting Emotional Data
[1077] A concrete example would be a mobile application using React Native.
[1078] Operation flow
[1079] The system begins operation when the user selects video data using a dedicated application and clicks the upload button. The device then sends the selected video data to the server. The server then saves the received video data in the "temporary / uploads / " directory and uses the ffmpeg library to obtain the metadata of the video data.
[1080] After saving is complete, the server launches an image analysis module (e.g., OpenCV or TensorFlow library) to analyze each frame and extract features for each scene (motion detection, face recognition, type of scenery, etc.), thereby providing a detailed understanding of the contents of the video data.
[1081] Based on the analysis results, the system defines the music generation requirements for each scene. For example, calm music is required for quiet landscape scenes, while energetic music is required for active scenes. Next, the user selects the desired music genre. For example, rock, jazz, classical, etc. can be selected.
[1082] Furthermore, the emotion engine analyzes the user's facial expressions and tone of voice to assess their emotional state. For example, if the user is smiling, the emotional state is determined to be "joy." Analysis data is collected from the camera and microphone and processed using OpenVINO and Microsoft Azure facial recognition APIs.
[1083] The server passes the analysis results, the user's emotional state, and genre selection to a music generation module. The music generation module uses libraries such as Magenta to automatically generate appropriate music based on this data. The generated music is synchronized with the video data by the server, and the timeline is adjusted using the Python moviepy library to complete the video. Finally, the edited video data is encoded, and a download link is generated for the user.
[1084] Specific examples and prompts
[1085] Specific examples
[1086] If user A wants to create a virtual travel video and add rock-style background music to it, he can use the system by following the steps below:
[1087] 1. User A uses the dedicated app to select and upload a travel video file. The device sends the video data to the server storage.
[1088] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[1089] 3. The server passes the video data to the image analysis module, which extracts scene features.
[1090] 4. User A selects a rock-style music genre in the application, and the emotion while operating the video is determined to be joy.
[1091] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which automatically generates rock-style music that is suitable for joy.
[1092] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[1093] 7. User A downloads the completed video from the application and can watch it.
[1094] Prompt Sentence Examples
[1095] "Automatically generate music that matches the video data based on the user's facial recognition data."
[1096] "When a user uploads a video containing a tranquil landscape scene, it generates calming music that matches the scene."
[1097] "Analyze the user's emotions and automatically generate and add rock-style music that matches the video data."
[1098] In this way, by utilizing the system of the present invention, users can easily and quickly automatically add suitable music to video data, thereby creating personalized video content.
[1099] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1100] Step 1:
[1101] The user selects the video data using a dedicated application and presses the upload button. This causes the device to send the video data (input: video data file) to the server (output: video data sent to server). A progress bar is displayed on the device to show the user the progress of the upload.
[1102] Step 2:
[1103] The server saves the received video data (input: video data file) in a temporary storage area ("temporary / uploads / " directory) (output: saved video data). Next, the server uses the ffmpeg library to obtain the video data's metadata (resolution, frame rate, file format) (output: metadata).
[1104] Step 3:
[1105] The server starts an image analysis module (e.g., OpenCV, TensorFlow) and passes the saved video data (input: video data file) to the image analysis module (output: analysis results). The image analysis module analyzes each frame of the video and extracts features for each scene (motion detection, face recognition, type of scenery) (output: feature data for each scene).
[1106] Step 4:
[1107] Based on the analysis results (input: feature data for each scene), the server defines the music generation requirements appropriate for each scene (output: music generation requirements). For example, it defines calm music for quiet scenes and energetic music for active scenes.
[1108] Step 5:
[1109] The user specifies the desired music genre (input: music genre specification) through the application (output: selected music genre). The user selects the genre using a pull-down menu or radio buttons.
[1110] Step 6:
[1111] The server launches an emotion engine (e.g., OpenVINO, facial recognition API) and collects the user's facial expressions and tone of voice using a camera or microphone (input: facial expression data, voice data). The emotion engine analyzes this data and evaluates the user's emotional state (output: emotional state data). For example, smiling data is recognized as "joy."
[1112] Step 7:
[1113] The server passes the analysis results, emotional state, and genre designation to the music generation module (input: feature data for each scene, emotional state data, and music genre designation).The music generation module uses the Magenta library and other tools to automatically generate optimal music based on this information (output: generated music file).
[1114] Step 8:
[1115] The generated music (input: music file) is synchronized with the video data by the server. Specifically, the timing of the music is adjusted to match each scene using the Python moviepy library, and the integrated edited video data (output: edited video data) is generated.
[1116] Step 9:
[1117] The server encodes the edited video data (input: edited video data) and generates a download link (output: download link) that the user can use to download the video. The link is stored in a cloud storage service (e.g., AWS S3) and displayed in the user's application.
[1118] Step 10:
[1119] Users can click the download link from the application to download the edited video data (input: download link) and watch it (output: downloaded video data). This allows users to easily enjoy personalized video content.
[1120] (Application example 2)
[1121] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1122] When automatically generating and adding appropriate music to video data created by a user, there is a demand for providing personalized music that reflects the user's preferences and emotional state. However, conventional systems have difficulty generating music that fully takes the user's emotional state into consideration, and the automatic generation and synchronization of music appropriate for video data is complex and time-consuming. To address these issues, the present invention aims to improve the user experience by providing optimal music for video data in a more intuitive and simple manner.
[1123] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1124] In this invention, the server includes means for a user to upload video data, means for storing the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, means for providing the generated video data to the user, means for specifying a desired music genre from the user's terminal, and means for recognizing facial expressions and voices and analyzing the emotional state when the user operates the video data. This makes it possible to provide a system that automatically and personalizedly generates music optimal for video data and is easy for users to use.
[1125] A "user" is an individual or organization who uploads video data to use this system and wishes to generate music.
[1126] "Video data" refers to digital data including video files that users have shot, edited, and saved, as well as their metadata.
[1127] A "server" is a computer system that stores video data received from users and performs the analysis and music generation process.
[1128] An "image analysis module" is a program or algorithm that analyzes each frame of video data and extracts scene features (e.g., motion detection, landscape type, facial recognition, etc.).
[1129] The "music generation module" is a program or algorithm for automatically generating appropriate music based on the results of image analysis and the user's wishes and emotional state.
[1130] "Emotional state" refers to the user's psychological and emotional state (e.g., joy, surprise, sadness, etc.) obtained by analyzing the user's facial expressions, voice, and text input.
[1131] "Synchronization" is the process of adjusting the timing of the generated music to match the scenes in the video data, so that the music and video are integrated.
[1132] "Genre" is a classification that refers to the style or type of music, and represents the kind of music a user desires (for example, rock, jazz, classical, etc.).
[1133] "Upload" is the process of sending video data from a user's terminal to a server.
[1134] "Recognition" is the process of collecting a user's facial expressions and voice using a camera and microphone, and analyzing them to assess their emotional state.
[1135] The system for implementing this invention is configured using the following hardware and software. The hardware requires a user terminal (such as a smartphone), a server, a camera, and a microphone. The software requires an image analysis module, a music generation module, an emotion engine, and an upload and synchronization management program.
[1136] First, the user uploads the video data they created using a device with a dedicated application installed. This device is equipped with a camera and microphone, which can collect the user's facial expressions and voice. The application also includes an interface for the user to specify the music genre they want.
[1137] When a user uploads video data from their device, the server stores the received video data in a storage directory (e.g., the "uploads" directory). The server then launches an image analysis module to analyze each frame of the video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). This allows the content of the video data to be understood in detail.
[1138] After obtaining the results of the image analysis, the server uses the emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice. For example, it analyzes their facial expressions and tone of voice while they are operating the video data to determine whether they are enjoying, surprised, or sad.
[1139] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification to automatically generate optimal music. The tempo and tone of the music can be dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright and fast-paced music is generated, and if the user is calm, calm and slow music is generated.
[1140] The server synchronizes the generated music with the timeline of the video data. Specifically, the timing of the generated music is adjusted to match each scene in the video data, thereby generating edited video data in which the music and video are integrated.
[1141] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data via the application.
[1142] Examples of concrete examples and prompts
[1143] Examples:
[1144] "A user used a smartphone app to upload a video taken at the beach during their summer vacation. The app specified 'bossa nova' as the user's desired music genre and recognized the emotion of joy from the user's facial expression while interacting with the video. The app then generated bossa nova music that matched the serene beach scene and added it to the video."
[1145] Example prompt sentence:
[1146] "I want to edit a video I shot at the beach during my summer vacation. I want to generate bossa nova music that matches the emotion of joy and add it to the video."
[1147] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1148] Step 1:
[1149] The user launches the dedicated application on the device, selects video data, and specifies the desired music genre.
[1150] Input: Video data file, desired music genre
[1151] Output: Video data file path, music genre information
[1152] How it works: A user selects a video file from their smartphone's gallery and clicks the "Upload" button through the app's interface, along with specifying the desired music genre, such as "Bossa Nova."
[1153] Step 2:
[1154] The terminal transmits the selected video data to the server, which stores the received video data in a storage directory.
[1155] Input: Video data file path
[1156] Output: Video data storage directory
[1157] Specific operation: The device uploads the video file to the server using a file transfer protocol, and the server stores the video data in the "uploads" directory.
[1158] Step 3:
[1159] The server launches an image analysis module, analyzes each frame of the video data, and extracts features for each scene.
[1160] Input: saved video data file, image analysis module
[1161] Output: Feature data for each scene
[1162] Specific operation: The server analyzes each frame of video data using Python scripts and extracts features such as scene motion detection, landscape type, and face recognition.
[1163] Step 4:
[1164] The server uses an emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice.
[1165] Input: facial expression data, voice data, music genre information
[1166] Output: Emotional state data
[1167] Specific operation: The server analyzes the camera and microphone data sent from the device, evaluates facial expressions and vocal tone, and uses an emotion engine to estimate the emotional state.
[1168] Step 5:
[1169] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification, and automatically generates optimal music.
[1170] Input: Scene feature data, emotional state data, music genre information
[1171] Output: Generated music file
[1172] Specific operation: Using a music generation AI model, music optimized for the characteristics of each scene and the user's emotions is generated and saved as a music file.
[1173] Step 6:
[1174] The server synchronizes the generated music with the timeline of the video data.
[1175] Input: Video data file, generated music file
[1176] Output: Edited video data file
[1177] What it does: The server uses a video editing library (such as MoviePy) to synchronize the music and video on the timeline and generate the final edited video file.
[1178] Step 7:
[1179] The server encodes the edited video data and generates a link that users can download.
[1180] Input: Edited video data file
[1181] Output: Download link
[1182] Specific operation: The server encodes the edited video data into a distributable format such as MP4 and provides the user with a download link, allowing them to watch the video.
[1183] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1184] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1185] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1186] [Fourth embodiment]
[1187] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1188] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1189] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1190] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1191] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1192] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1193] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1194] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1195] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1196] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1197] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1198] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1199] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1200] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data. Each step of this system is described in detail below.
[1201] First, the user uploads video data from their own device to the server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[1202] The server then stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[1203] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[1204] The analysis results from the image analysis module are returned to the server, which then uses the results to define the music generation requirements appropriate for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[1205] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[1206] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[1207] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[1208] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[1209] Specific examples
[1210] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[1211] 1. User A uses a dedicated app to upload a travel video file to the server.
[1212] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[1213] 3. The server passes the saved video file to the image analysis module, which extracts the features of each scene.
[1214] 4. User A selects a rock-style music genre in the application and sends it to the server.
[1215] 5. The server passes the analysis results and genre specification to the music generation module, which then generates rock-style music.
[1216] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[1217] 7. User A can download the completed video from the application and watch it.
[1218] In this way, the system of the present invention allows the user to automatically assign suitable music to video data simply and quickly.
[1219] The processing flow will be explained below.
[1220] Step 1:
[1221] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[1222] Step 2:
[1223] The server detects the received video data and stores it in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[1224] Step 3:
[1225] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[1226] Step 4:
[1227] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[1228] Step 5:
[1229] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[1230] Step 6:
[1231] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates music appropriate for each scene based on the specified requirements. This music is then adjusted to create a harmonious overall composition.
[1232] Step 7:
[1233] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[1234] Step 8:
[1235] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[1236] Specific operation example
[1237] Step 1:
[1238] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[1239] Step 2:
[1240] The server receives the video data and saves it in the "temporary / uploads / " folder.
[1241] Step 3:
[1242] The server starts the image analysis module and analyzes the video data. The analysis module extracts the features of each frame.
[1243] Step 4:
[1244] The analysis module returns the extracted features to the server, which then sets the music generation requirements appropriate for each scene based on the features.
[1245] Step 5:
[1246] User A selects the music genre "rock" from the application and sends it to the server. The server receives this information.
[1247] Step 6:
[1248] The server passes the analysis results and user-specified genre information to the music generation module, which then automatically generates rock-style music.
[1249] Step 7:
[1250] The server synchronizes the generated rock music with the timeline of the video data.
[1251] Step 8:
[1252] The server encodes the final edited video data and provides a downloadable link to User A. User A downloads the edited video from the link and watches it.
[1253] Example 1
[1254] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1255] Conventional video editing systems require users to manually select and edit music appropriate for the video, which requires a great deal of time and effort. It is also not easy to select music that matches the different atmospheres of each video scene. Furthermore, synchronizing music and video requires advanced editing skills, making it difficult for average users to create the high-quality video content desired quickly and easily.
[1256] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1257] In this invention, the server includes a means for a user to upload video data, a means for the server to store the received video data, a means for the server to analyze the content of the video data using an image analysis module, a means for the server to generate appropriate music using a music generation module based on the analysis results, a means for synchronizing the generated music with the video data, and a means for providing the generated video data to the user. This enables the user to quickly create high-quality video content using AI technology without any manual effort.
[1258] "User" refers to an entity that uses the system to upload video data and request music creation.
[1259] "Server" refers to a computer system that performs a series of processes including receiving, storing, analyzing, and generating music from video data.
[1260] "Video data" refers to video files uploaded by users and supports a variety of formats.
[1261] "Uploading means" refers to an application or interface that allows a user to send video data to the server.
[1262] The "storage means" refers to a storage area in which the server temporarily stores the video data received.
[1263] An "image analysis module" refers to a software component that analyzes the contents of video data and extracts features for each scene.
[1264] "Music generation module" refers to a software component that automatically generates music suitable for video based on analysis results.
[1265] "Synchronization means" refers to a process that performs operations to integrate the generated music with the video data.
[1266] "Providing means" refers to a method for providing the generated video data to the user.
[1267] "Metadata" refers to information contained in video data (e.g., resolution, frame rate, file format, etc.).
[1268] "Generative AI model" refers to the artificial intelligence algorithm used by the music generation module, enabling the automatic generation of music.
[1269] A "prompt" is an instruction given to a generative AI model that describes the requirements for generating specific music.
[1270] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module that analyzes the content of the video data with a music generation module that generates music, generating original music according to the user's preferences and adding it to the video data.
[1271] First, the user uploads video data from their device to the server using a dedicated application. Video data can be video files in various formats, ranging from a few minutes to several hours in length. To do this, the user clicks the "Upload" button in the application and selects the appropriate video file from the file selection dialog. The video is then uploaded to the server via cloud storage.
[1272] Next, the server stores the received video data in a temporary storage area (directory "temporary / uploads / "). At the same time, it obtains the video data's metadata (resolution, frame rate, file format, etc.) to prepare for subsequent analysis.
[1273] The server launches an image analysis module to analyze the stored video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). This analysis uses common libraries such as OpenCV and TensorFlow. The analyzed data is temporarily stored in intermediate data storage.
[1274] Based on the analysis results from the image analysis module, the server then defines the music generation requirements for each scene. For example, a quiet landscape scene requires calm music, while an active scene requires energetic music. These defined requirements are recorded in JSON format.
[1275] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through a dedicated application, selects the desired genre from a drop-down menu in the application, and sends the information to the server.
[1276] The server passes the analysis results and genre specification to the music generation module, which uses this information to automatically generate optimal music using AI algorithms. This generation process can utilize OpenAI's MuseNet or other generative AI models.
[1277] The generated music is then synchronized with the video data by a server. During the synchronization process, the timing of the music is adjusted to match each scene in the video data, resulting in a harmonious overall music and video. Specifically, the music track is merged into the video file using libraries such as FFmpeg.
[1278] Finally, the server encodes the edited video data and generates a download link for the user, who can click the "Download" button in the application to download and watch the completed video.
[1279] Specific examples
[1280] For example, if user A wants to create a virtual travel video and add rock-style background music to it, he or she would follow the steps below.
[1281] 1. User A clicks the "Upload" button in the application, selects a travel video file, and uploads it to the server.
[1282] 2. The server receives and stores the video file and retrieves the metadata.
[1283] 3. The server uses OpenCV to analyze the video data and extract features for each scene.
[1284] 4. User A selects a rock-style music genre in the application and sends it to the server.
[1285] 5. The server passes the analysis results and genre specification to the music generation module, which generates rock-style music using MuseNet.
[1286] 6. The server uses FFmpeg to merge the generated music into the video file and generate the edited video.
[1287] 7. User A downloads the completed video from the application and watches it.
[1288] Prompt Sentence Examples
[1289] "Analyze the footage of the virtual trip and add calming music to the quiet scenes and energetic music to the dynamic scenes. Overall, keep the music taste rock-like."
[1290] As described above, the present invention enables a user to automatically assign suitable music to video data simply and quickly.
[1291] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1292] Step 1:
[1293] The user uploads the video data.
[1294] Specifically, the user launches the dedicated application, clicks the "Upload" button, and selects the video file they want to upload from the file selection dialog. The input is the user's video data, and the output is the video data sent to the server via cloud storage. The video data can be in a variety of video formats, ranging from a few minutes to a few hours.
[1295] Step 2:
[1296] The server stores the received video data.
[1297] Specifically, the server stores the video data in the "temporary / uploads / " directory. At this time, the server acquires the video data's metadata (resolution, frame rate, file format, etc.). The input is the uploaded video data, and the output is the saved data and the acquired metadata. This makes subsequent analysis easier.
[1298] Step 3:
[1299] The server starts the image analysis module and analyzes the video data.
[1300] Specifically, the server uses libraries such as OpenCV and TensorFlow to analyze each frame of video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). The input is the saved video data, and the output is the analysis results (features for each scene). These analysis results are temporarily stored in intermediate data storage.
[1301] Step 4:
[1302] The server defines the requirements for music generation based on the analysis results.
[1303] Specifically, based on the data obtained from the analysis module, the music specifications appropriate for each scene (e.g., gentle music for quiet scenes, energetic music for active scenes) are written in JSON format. The input is the analysis results, and the output is JSON data that records the requirements for music generation.
[1304] Step 5:
[1305] The user specifies the genre of music.
[1306] Specifically, the user selects the desired music genre (e.g., rock, jazz, etc.) from a drop-down menu in the application and sends that information to the server. The input is the user's genre selection, and the output is the genre information sent to the server.
[1307] Step 6:
[1308] The server passes the analysis results and genre designation to the music generation module.
[1309] Specifically, it uses a generative AI model such as OpenAI's MuseNet to generate music based on the analysis results and the user's genre selection. The input is the analysis results and genre information, and the output is the generated music data. Based on this information, the music generation module automatically generates music customized for each scene.
[1310] Step 7:
[1311] The server synchronizes the generated music with the video data.
[1312] Specifically, we use libraries such as FFmpeg to merge the generated music with the video data and adjust the timing. The input is the video data and the generated music, and the output is edited video data synchronized with the music. In this editing process, we optimize the timing of the music for each scene.
[1313] Step 8:
[1314] The server provides the edited video data to the user.
[1315] Specifically, the edited video data is encoded and a download link is generated. Users can click the "Download" button in the application to download and watch the completed video. The input is the edited video data, and the output is the download link.
[1316] (Application example 1)
[1317] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1318] In the past, adding appropriate music to video data created by users required manual selection of music and editing to match the video, which required a great deal of time and effort. Furthermore, the video and music often did not harmonize, making it difficult to create content that is visually and aurally unified. To solve this problem, there is a demand for a system that can automatically generate appropriate music based on the content of the video data and easily provide content in which the video and music are in harmony.
[1319] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1320] In this invention, the server includes means for a user to upload video data, means for the server to store the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user via a terminal such as a smartphone, thereby enabling the user to quickly and easily create video data with optimal music added to the video, and to provide content that is visually and aurally harmonious.
[1321] "User" means a person or legal entity that has the right to use the system to upload video data and generate music.
[1322] "Video data" refers to all video files uploaded by users, the contents of which are analyzed to generate music.
[1323] "Upload" refers to the act of a user sending video data from their own device to a server.
[1324] The "server" is a central computer system that receives and stores video data, and processes it for analysis and music generation.
[1325] An "image analysis module" is a software component that analyzes video data and understands its contents.
[1326] A "music generation module" is a software component for automatically generating music based on the analysis results.
[1327] "Generating" "music" refers to the process by which the system automatically creates new music.
[1328] "Synchronization" refers to the process of matching the generated music to the appropriate timing for each scene in the video data.
[1329] "Devices such as smartphones" refer to portable information and communication devices used by users, which are capable of uploading video data and downloading generated video data.
[1330] "Generated video data" refers to the final video file that has been analyzed and processed to generate music, and is synchronized with the music.
[1331] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data with a music generation module to generate original music according to the user's preferences and add it to the video data.
[1332] First, a user uploads video data from their smartphone or other device to a server using a dedicated application. The video data is a video file in various formats, ranging in length from a few minutes to a few hours.
[1333] The server stores the received video data in a temporary storage area, which retrieves the file metadata (resolution, frame rate, file format, etc.) and prepares it for subsequent analysis.
[1334] After receiving and saving the video data, the server launches the image analysis module, which analyzes the saved video data. The image analysis module analyzes each frame of the video data and extracts features for each scene (such as motion detection, facial recognition, and type of scenery). This allows for a detailed understanding of the content of the video data.
[1335] The analysis results from the image analysis module are returned to the server, which then uses the results to define the music generation requirements appropriate for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[1336] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and incorporates it into the requirements for music generation.
[1337] The server then passes the analysis results and genre specification to the music generation module, which then automatically generates optimal music based on this information. The generated music is designed to contain different elements for each scene, yet create a harmonious overall composition.
[1338] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[1339] Finally, the server encodes the edited video data and generates a download link for users, who can then download and watch the completed video on their smartphones or other devices.
[1340] For example, if a user creates a travel video and wants to add rock-style background music to it, they can use this system as follows:
[1341] Example prompt sentence:
[1342] Travel Videos
[1343] Genre: Rock
[1344] Description: Adventurous scenery, mountain climbing, river rafting
[1345] Based on these prompts, the system executes each step and provides the user with a final edited video that is visually and audibly consistent.
[1346] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1347] Step 1:
[1348] The user uploads video data from their smartphone or other device. The input is a video file selected by the user, and the output is that video file is sent to the server. At this time, the user also specifies the desired music genre.
[1349] Step 2:
[1350] The server saves the received video data. It stores the video file in a temporary storage area and obtains the file metadata (resolution, frame rate, file format, etc.). The input is the video data received in step 1, and the output is the video data and its metadata stored in the storage area.
[1351] Step 3:
[1352] The server uses an image analysis module to analyze the content of the stored video data. The input is the video data stored in the storage area, and the output is the extracted results of features for each scene (e.g., motion detection, face recognition, type of scenery, etc.). Specifically, it analyzes each frame of the video data and recognizes specific patterns and objects.
[1353] Step 4:
[1354] Based on the analysis results, the server defines the requirements for music generation appropriate for each scene in the video data. The input is the analysis results obtained in step 3, and the output is the requirements necessary for music generation (for example, calm music for quiet scenes, energetic music for active scenes). Specifically, the server analyzes the analysis results and determines the style and atmosphere of the music that suits each scene.
[1355] Step 5:
[1356] The user specifies the desired music genre through the application. The input is the music genre information specified by the user, and the output is that information is sent to the server.
[1357] Step 6:
[1358] The server passes the analysis results and genre specification to the music generation module. The input is the analysis results and the user's genre specification information, and the output is the input data for the music generation module.
[1359] Step 7:
[1360] The music generation module automatically generates optimal music. The input is the data obtained in step 6, and the output is the generated music. Specifically, an AI model (using TensorFlow, for example) generates music data based on the analysis results and genre information.
[1361] Step 8:
[1362] The generated music is synchronized with the video data. The server adjusts the timing of the music to match each scene in the video data. The input is the generated music and video data, and the output is edited video data in which the music and video are synchronized. Specifically, the start and end points of the music are adjusted for each scene to create integrated data.
[1363] Step 9:
[1364] The server encodes the edited video data and generates a link that allows the user to download it. The input is the edited video data generated in step 8, and the output is a download link. Specifically, the server encodes the data and provides the link to the user.
[1365] Step 10:
[1366] The user downloads the completed video from a device such as a smartphone and watches it. The input is a download link provided by the server, and the output is the final video data with music added, stored on the user's device.
[1367] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1368] This invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by users. This system combines an image analysis module for analyzing the content of the video data, a music generation module for generating music, and an emotion engine for recognizing the user's emotions to generate original music according to the user's preferences and add it to the video data. Each step of this system is described in detail below.
[1369] First, the user selects the video data using a dedicated application on their own device and clicks the upload button. The device then sends the selected video data file to the server.
[1370] Next, the server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[1371] After receiving and saving the video data, the server starts the image analysis module and passes the saved video data to it. The image analysis module analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.). This allows for a detailed understanding of the content of the video data.
[1372] The analysis results from the image analysis module are returned to the server, which then uses the results to define the music generation requirements appropriate for each scene in the video data. For example, calm music is required for quiet landscape scenes, and energetic music for active scenes.
[1373] Next, the user specifies the desired music genre (e.g., rock, jazz, etc.) through the application. The server receives the user's specified genre information and adds it to the requirements for music generation.
[1374] The present invention also includes an emotion engine. The emotion engine analyzes the user's facial expressions, voice, and text input to evaluate the user's emotional state. The emotion engine analyzes the data collected by a camera and microphone, including facial expressions and tone of voice, while the user is manipulating video data. The results of this analysis (the user's emotional state, such as whether they are enjoying, surprised, or sad) are reflected in the parameters for generating music.
[1375] The server then passes the analysis results, the user's emotional state, and the genre designation to the music generation module. The music generation module automatically generates optimal music based on this information. The tempo, tone, and genre of the music are dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright, fast-paced music is generated, while if the user is calm, calm, slow music is generated.
[1376] The generated music is synchronized with the video data by the server, which adjusts the timing of the music to match each scene in the video data, and then generates the final edited video data in which the music and video are integrated.
[1377] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data from the application.
[1378] Specific examples
[1379] If user A wants to create a virtual travel video and add rock-style background music to it, he can use this system as follows.
[1380] 1. User A uses the dedicated app to select a travel video file and presses the upload button. The device sends the video data to the server storage.
[1381] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[1382] 3. The server passes the video data to the image analysis module, which extracts the features of each scene.
[1383] 4. User A specified a rock-style music genre in the application, and the emotion he felt when operating the video was determined to be joy.
[1384] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[1385] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[1386] 7. User A can download the completed video from the application and watch it.
[1387] In this way, the system of the present invention allows users to easily and quickly automatically assign suitable music to video data, and can provide more personalized music depending on the user's emotional state.
[1388] The processing flow will be explained below.
[1389] Step 1:
[1390] The user selects the video data using a dedicated application on their device and clicks the upload button. The device then sends the selected video data file to the server.
[1391] Step 2:
[1392] The server stores the received video data in a temporary storage area (e.g., temporary / uploads / directory). The server then obtains the video data metadata (resolution, frame rate, file format, etc.) and prepares it for analysis.
[1393] Step 3:
[1394] The server starts the image analysis module and passes the saved video data to the image analysis module, which analyzes each frame of the video data and extracts features for each scene (motion detection, face recognition, type of scenery, etc.).
[1395] Step 4:
[1396] The image analysis module returns the extracted features to the server, which then analyzes them and defines the music generation requirements for each scene. For example, a calm melody is required for a quiet landscape scene, and a rhythmic melody is required for an active scene.
[1397] Step 5:
[1398] The user inputs the desired music genre (e.g., rock, jazz, etc.) through the application and sends it to the server. The server receives the desired genre information and adds it to the parameters for music generation.
[1399] Step 6:
[1400] When a user interacts with video data (e.g., through facial expressions or voice input), the emotion engine evaluates the user's emotional state. The user's facial expressions and voice are recorded via the device's camera and microphone, and then analyzed by the emotion engine.
[1401] Step 7:
[1402] The emotion engine sends the analysis results to the server, which then incorporates this emotion information (e.g., whether the user is happy or surprised) into the requirements for music generation.
[1403] Step 8:
[1404] The server passes the analysis results, the user's emotional state, and genre information to the music generation module. The music generation module automatically generates optimal music based on this information. If the user is happy, it generates bright, fast-paced music, and if the user is calm, it generates calm, slow music.
[1405] Step 9:
[1406] The server synchronizes the generated music files with the timeline of the video data, matching the timing of the music to each scene and generating edited video data in which the music and video are integrated.
[1407] Step 10:
[1408] The server encodes the edited video data and generates a download link for the user. The user can then download the edited video data from the application, view it, and save it.
[1409] Specific operation example
[1410] Step 1:
[1411] User A uses the dedicated app to select the video data of the trip and presses the "Upload" button. The device sends the video data to the server storage.
[1412] Step 2:
[1413] The server receives the video data and saves it in the "temporary / uploads / " directory.
[1414] Step 3:
[1415] The server starts an image analysis module, analyzes each frame of the video data, and extracts features.
[1416] Step 4:
[1417] The image analysis module returns the analysis results to the server, which then defines the necessary music generation requirements.
[1418] Step 5:
[1419] User A selects the music genre "rock" from the application and sends it to the server.
[1420] Step 6:
[1421] When User A is operating the video, the device sends data to the emotion engine based on User A's facial expressions and voice. The emotion engine then analyzes User A's emotions.
[1422] Step 7:
[1423] The emotion engine sends the emotional state of user A (e.g., joy) to the server.
[1424] Step 8:
[1425] The server passes the analysis results, emotional state, and genre designation to the music generation module, which then automatically generates rock-style music that matches the emotion of joy.
[1426] Step 9:
[1427] The server synchronizes the generated music with the video timeline to generate an edited video.
[1428] Step 10:
[1429] The server provides the edited video in a downloadable format to User A. User A then downloads the completed video from the application and watches it.
[1430] Example 2
[1431] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1432] Conventional video editing systems require users to individually select appropriate music and manually synchronize it with the video data, which requires time and effort. Furthermore, even systems capable of automatically generating music for video data lack the functionality to personalize the music based on the user's emotions and preferences. This makes it difficult to automatically perform high-quality, personalized video editing that satisfies users.
[1433] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1434] In this invention, the server includes means for a user to upload video data, means for saving the received video data, means for analyzing the content of the video data using an image analysis module, means for defining requirements for music generation using an emotion engine that evaluates the analysis results and the user's emotional state, means for the user to specify a desired music genre, means for generating appropriate music based on the analysis results, the emotional state, and the specified genre using a music generation module, means for synchronizing the generated music with the video data, and means for providing the generated video data to the user. This allows the user to automatically generate appropriate music for specific video data and personalize the music according to the user's emotions and preferences.
[1435] A "user" is an individual or corporation that operates the system and uploads video data or specifies music genres.
[1436] "Video data" refers to video files uploaded by users to the system, including metadata such as resolution, frame rate, and file format.
[1437] A "server" is a computer system that receives, stores, analyzes, and generates music from video data.
[1438] An "image analysis module" is a software component that analyzes each frame of video data and extracts features for each scene.
[1439] "Metadata" is attribute information about video data, and includes resolution, frame rate, file format, and the like.
[1440] An "emotion engine" is an algorithm or software component that analyzes a user's facial expressions and tone of voice to assess their emotional state.
[1441] The "music generation module" is a software component that automatically generates music based on the results of image analysis, emotional state, and specified music genre.
[1442] "Music genre" refers to the style or type of music designated by the user, examples of which include rock, jazz, classical, etc.
[1443] "Synchronization" refers to the process of adjusting the timing of the generated music to match each scene of the video data and integrating them together.
[1444] "Generated video data" refers to a video file in which the generated music is added to the original video data and editing is complete.
[1445] "Upload" refers to the operation by which a user sends video data from their own terminal to a server.
[1446] "Providing" refers to the server presenting the generated video data to the user in a downloadable form.
[1447] The present invention relates to a system that uses AI technology to automatically generate and add appropriate music to video data created by a user. This system is composed of a combination of an image analysis module for analyzing the content of the video data, a music generation module for generating appropriate music, and an emotion engine for recognizing the user's emotions. Specific embodiments for implementing the present invention are described below.
[1448] System configuration
[1449] Hardware and Software
[1450] 1. The server is primarily responsible for the following:
[1451] Receiving and storing video data
[1452] Image analysis
[1453] Emotion analysis
[1454] Music Generation
[1455] Synchronization of video data and music
[1456] Provision of edited video data
[1457] Examples of use include AWS EC2 servers, Python programs, and the ffmpeg library.
[1458] 2. A terminal is a device (smartphone, tablet, PC, etc.) operated by a user that has the following functions:
[1459] Uploading video data
[1460] Specifying the music genre
[1461] Collecting Emotional Data
[1462] A concrete example would be a mobile application using React Native.
[1463] Operation flow
[1464] The system begins operation when the user selects video data using a dedicated application and clicks the upload button. The device then sends the selected video data to the server. The server then saves the received video data in the "temporary / uploads / " directory and uses the ffmpeg library to obtain the metadata of the video data.
[1465] After saving is complete, the server launches an image analysis module (e.g., OpenCV or TensorFlow library) to analyze each frame and extract features for each scene (motion detection, face recognition, type of scenery, etc.), thereby providing a detailed understanding of the contents of the video data.
[1466] Based on the analysis results, the system defines the music generation requirements for each scene. For example, calm music is required for quiet landscape scenes, while energetic music is required for active scenes. Next, the user selects the desired music genre. For example, rock, jazz, classical, etc. can be selected.
[1467] Furthermore, the emotion engine analyzes the user's facial expressions and tone of voice to assess their emotional state. For example, if the user is smiling, the emotional state is determined to be "joy." Analysis data is collected from the camera and microphone and processed using OpenVINO and Microsoft Azure facial recognition APIs.
[1468] The server passes the analysis results, the user's emotional state, and genre selection to a music generation module. The music generation module uses libraries such as Magenta to automatically generate appropriate music based on this data. The generated music is synchronized with the video data by the server, and the timeline is adjusted using the Python moviepy library to complete the video. Finally, the edited video data is encoded, and a download link is generated for the user.
[1469] Specific examples and prompts
[1470] Specific examples
[1471] If user A wants to create a virtual travel video and add rock-style background music to it, he can use the system by following the steps below:
[1472] 1. User A uses the dedicated app to select and upload a travel video file. The device sends the video data to the server storage.
[1473] 2. The server receives the video file and saves it in the "temporary / uploads / " directory.
[1474] 3. The server passes the video data to the image analysis module, which extracts scene features.
[1475] 4. User A selects a rock-style music genre in the application, and the emotion while operating the video is determined to be joy.
[1476] 5. The server passes the analysis results, emotion engine results, and genre designation to the music generation module, which automatically generates rock-style music that is suitable for joy.
[1477] 6. The server synchronizes the generated music with the video timeline and generates the edited video.
[1478] 7. User A downloads the completed video from the application and can watch it.
[1479] Prompt Sentence Examples
[1480] "Automatically generate music that matches the video data based on the user's facial recognition data."
[1481] "When a user uploads a video containing a tranquil landscape scene, it generates calming music that matches the scene."
[1482] "Analyze the user's emotions and automatically generate and add rock-style music that matches the video data."
[1483] In this way, by utilizing the system of the present invention, users can easily and quickly automatically add suitable music to video data, thereby creating personalized video content.
[1484] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1485] Step 1:
[1486] The user selects the video data using a dedicated application and presses the upload button. This causes the device to send the video data (input: video data file) to the server (output: video data sent to server). A progress bar is displayed on the device to show the user the progress of the upload.
[1487] Step 2:
[1488] The server saves the received video data (input: video data file) in a temporary storage area ("temporary / uploads / " directory) (output: saved video data). Next, the server uses the ffmpeg library to obtain the video data's metadata (resolution, frame rate, file format) (output: metadata).
[1489] Step 3:
[1490] The server starts an image analysis module (e.g., OpenCV, TensorFlow) and passes the saved video data (input: video data file) to the image analysis module (output: analysis results). The image analysis module analyzes each frame of the video and extracts features for each scene (motion detection, face recognition, type of scenery) (output: feature data for each scene).
[1491] Step 4:
[1492] Based on the analysis results (input: feature data for each scene), the server defines the music generation requirements appropriate for each scene (output: music generation requirements). For example, it defines calm music for quiet scenes and energetic music for active scenes.
[1493] Step 5:
[1494] The user specifies the desired music genre (input: music genre specification) through the application (output: selected music genre). The user selects the genre using a pull-down menu or radio buttons.
[1495] Step 6:
[1496] The server launches an emotion engine (e.g., OpenVINO, facial recognition API) and collects the user's facial expressions and tone of voice using a camera or microphone (input: facial expression data, voice data). The emotion engine analyzes this data and evaluates the user's emotional state (output: emotional state data). For example, smiling data is recognized as "joy."
[1497] Step 7:
[1498] The server passes the analysis results, emotional state, and genre designation to the music generation module (input: feature data for each scene, emotional state data, and music genre designation).The music generation module uses the Magenta library and other tools to automatically generate optimal music based on this information (output: generated music file).
[1499] Step 8:
[1500] The generated music (input: music file) is synchronized with the video data by the server. Specifically, the timing of the music is adjusted to match each scene using the Python moviepy library, and the integrated edited video data (output: edited video data) is generated.
[1501] Step 9:
[1502] The server encodes the edited video data (input: edited video data) and generates a download link (output: download link) that the user can use to download the video. The link is stored in a cloud storage service (e.g., AWS S3) and displayed in the user's application.
[1503] Step 10:
[1504] Users can click the download link from the application to download the edited video data (input: download link) and watch it (output: downloaded video data). This allows users to easily enjoy personalized video content.
[1505] (Application example 2)
[1506] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1507] When automatically generating and adding appropriate music to video data created by a user, there is a demand for providing personalized music that reflects the user's preferences and emotional state. However, conventional systems have difficulty generating music that fully takes the user's emotional state into consideration, and the automatic generation and synchronization of music appropriate for video data is complex and time-consuming. To address these issues, the present invention aims to improve the user experience by providing optimal music for video data in a more intuitive and simple manner.
[1508] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1509] In this invention, the server includes means for a user to upload video data, means for storing the received video data, means for the server to analyze the content of the video data using an image analysis module, means for the server to generate appropriate music using a music generation module based on the analysis results, means for synchronizing the generated music with the video data, means for providing the generated video data to the user, means for specifying a desired music genre from the user's terminal, and means for recognizing facial expressions and voices and analyzing the emotional state when the user operates the video data. This makes it possible to provide a system that automatically and personalizedly generates music optimal for video data and is easy for users to use.
[1510] A "user" is an individual or organization who uploads video data to use this system and wishes to generate music.
[1511] "Video data" refers to digital data including video files that users have shot, edited, and saved, as well as their metadata.
[1512] A "server" is a computer system that stores video data received from users and performs the analysis and music generation process.
[1513] An "image analysis module" is a program or algorithm that analyzes each frame of video data and extracts scene features (e.g., motion detection, landscape type, facial recognition, etc.).
[1514] The "music generation module" is a program or algorithm for automatically generating appropriate music based on the results of image analysis and the user's wishes and emotional state.
[1515] "Emotional state" refers to the user's psychological and emotional state (e.g., joy, surprise, sadness, etc.) obtained by analyzing the user's facial expressions, voice, and text input.
[1516] "Synchronization" is the process of adjusting the timing of the generated music to match the scenes in the video data, so that the music and video are integrated.
[1517] "Genre" is a classification that refers to the style or type of music, and represents the kind of music a user desires (for example, rock, jazz, classical, etc.).
[1518] "Upload" is the process of sending video data from a user's terminal to a server.
[1519] "Recognition" is the process of collecting a user's facial expressions and voice using a camera and microphone, and analyzing them to assess their emotional state.
[1520] The system for implementing this invention is configured using the following hardware and software. The hardware requires a user terminal (such as a smartphone), a server, a camera, and a microphone. The software requires an image analysis module, a music generation module, an emotion engine, and an upload and synchronization management program.
[1521] First, the user uploads the video data they created using a device with a dedicated application installed. This device is equipped with a camera and microphone, which can collect the user's facial expressions and voice. The application also includes an interface for the user to specify the music genre they want.
[1522] When a user uploads video data from their device, the server stores the received video data in a storage directory (e.g., the "uploads" directory). The server then launches an image analysis module to analyze each frame of the video data and extract features for each scene (motion detection, face recognition, type of scenery, etc.). This allows the content of the video data to be understood in detail.
[1523] After obtaining the results of the image analysis, the server uses the emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice. For example, it analyzes their facial expressions and tone of voice while they are operating the video data to determine whether they are enjoying, surprised, or sad.
[1524] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification to automatically generate optimal music. The tempo and tone of the music can be dynamically adjusted according to the user's emotional state. For example, if the user is happy, bright and fast-paced music is generated, and if the user is calm, calm and slow music is generated.
[1525] The server synchronizes the generated music with the timeline of the video data. Specifically, the timing of the generated music is adjusted to match each scene in the video data, thereby generating edited video data in which the music and video are integrated.
[1526] Finally, the server encodes the edited video data and generates a download link for the user, who can then download and view the edited video data via the application.
[1527] Examples of concrete examples and prompts
[1528] Examples:
[1529] "A user used a smartphone app to upload a video taken at the beach during their summer vacation. The app specified 'bossa nova' as the user's desired music genre and recognized the emotion of joy from the user's facial expression while interacting with the video. The app then generated bossa nova music that matched the serene beach scene and added it to the video."
[1530] Example prompt sentence:
[1531] "I want to edit a video I shot at the beach during my summer vacation. I want to generate bossa nova music that matches the emotion of joy and add it to the video."
[1532] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1533] Step 1:
[1534] The user launches the dedicated application on the device, selects video data, and specifies the desired music genre.
[1535] Input: Video data file, desired music genre
[1536] Output: Video data file path, music genre information
[1537] How it works: A user selects a video file from their smartphone's gallery and clicks the "Upload" button through the app's interface, along with specifying the desired music genre, such as "Bossa Nova."
[1538] Step 2:
[1539] The terminal transmits the selected video data to the server, which stores the received video data in a storage directory.
[1540] Input: Video data file path
[1541] Output: Video data storage directory
[1542] Specific operation: The device uploads the video file to the server using a file transfer protocol, and the server stores the video data in the "uploads" directory.
[1543] Step 3:
[1544] The server launches an image analysis module, analyzes each frame of the video data, and extracts features for each scene.
[1545] Input: saved video data file, image analysis module
[1546] Output: Feature data for each scene
[1547] Specific operation: The server analyzes each frame of video data using Python scripts and extracts features such as scene motion detection, landscape type, and face recognition.
[1548] Step 4:
[1549] The server uses an emotion engine along with the music genre information specified by the user to analyze the user's emotional state from their facial expressions and voice.
[1550] Input: facial expression data, voice data, music genre information
[1551] Output: Emotional state data
[1552] Specific operation: The server analyzes the camera and microphone data sent from the device, evaluates facial expressions and vocal tone, and uses an emotion engine to estimate the emotional state.
[1553] Step 5:
[1554] The server activates a music generation module based on the image analysis results, the user's emotional state, and the genre specification, and automatically generates optimal music.
[1555] Input: Scene feature data, emotional state data, music genre information
[1556] Output: Generated music file
[1557] Specific operation: Using a music generation AI model, music optimized for the characteristics of each scene and the user's emotions is generated and saved as a music file.
[1558] Step 6:
[1559] The server synchronizes the generated music with the timeline of the video data.
[1560] Input: Video data file, generated music file
[1561] Output: Edited video data file
[1562] What it does: The server uses a video editing library (such as MoviePy) to synchronize the music and video on the timeline and generate the final edited video file.
[1563] Step 7:
[1564] The server encodes the edited video data and generates a link that users can download.
[1565] Input: Edited video data file
[1566] Output: Download link
[1567] Specific operation: The server encodes the edited video data into a distributable format such as MP4 and provides the user with a download link, allowing them to watch the video.
[1568] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1569] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1570] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1571] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1572] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1573] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1574] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1575] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1576] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1577] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1578] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1579] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1580] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1581] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1582] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1583] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1584] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1585] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1586] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1587] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1588] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1589] The following is further disclosed regarding the above embodiment.
[1590] (Claim 1)
[1591] A means for a user to upload video data;
[1592] A means for storing the video data received by the server;
[1593] A means for the server to analyze the content of the video data using an image analysis module;
[1594] A means for the server to generate appropriate music using a music generation module based on the analysis results;
[1595] means for synchronizing the generated music with the video data;
[1596] means for providing the generated video data to a user;
[1597] A system including:
[1598] (Claim 2)
[1599] 10. The system of claim 1, further comprising means for a user to specify a genre of music.
[1600] (Claim 3)
[1601] 2. The system according to claim 1, further comprising means for generating a different musical pattern for each scene based on the analysis result.
[1602] "Example 1"
[1603] (Claim 1)
[1604] A means for a user to upload video data;
[1605] A means for storing the video data received by the server;
[1606] A means for the server to analyze the content of the video data using an image analysis module;
[1607] A means for the server to generate appropriate music using a music generation module based on the analysis results;
[1608] means for synchronizing the generated music with the video data;
[1609] means for providing the generated video data to a user;
[1610] A system including:
[1611] (Claim 2)
[1612] 10. The system of claim 1, further comprising means for a user to specify a genre of music.
[1613] (Claim 3)
[1614] 2. The system according to claim 1, further comprising means for generating a different musical pattern for each scene based on the analysis result.
[1615] (Claim 4)
[1616] 10. The system according to claim 1, further comprising means for the server to obtain metadata of the video data and prepare for subsequent analysis processing.
[1617] (Claim 5)
[1618] 10. The system of claim 1, further comprising means for the server to record the analysis results in JSON format.
[1619] (Claim 6)
[1620] 10. The system of claim 1, further comprising means for the server to use the generative AI model to operate the music generation module.
[1621] (Claim 7)
[1622] 10. The system of claim 1, further comprising: means for merging the generated music with the video data using a video editing library.
[1623] "Application Example 1"
[1624] (Claim 1)
[1625] A means for a user to upload video data;
[1626] A means for storing the video data received by the server;
[1627] A means for the server to analyze the content of the video data using an image analysis module;
[1628] A means for the server to generate appropriate music using a music generation module based on the analysis results;
[1629] means for synchronizing the generated music with the video data;
[1630] A means for providing the generated video data to a user via a terminal such as a smartphone;
[1631] A system including:
[1632] (Claim 2)
[1633] 10. The system of claim 1, further comprising means for a user to specify a genre of music.
[1634] (Claim 3)
[1635] 2. The system according to claim 1, further comprising means for generating a different musical pattern for each scene based on the analysis result.
[1636] "Example 2: Combining Emotion Engines"
[1637] (Claim 1)
[1638] A means for a user to upload video data;
[1639] A means for storing the video data received by the server;
[1640] A means for the server to analyze the content of the video data using an image analysis module;
[1641] a means for the server to define requirements for music generation using an emotion engine that evaluates the analysis results and the user's emotional state;
[1642] A means for a user to specify a desired music genre;
[1643] A means for the server to generate appropriate music based on the analysis result, emotional state, and genre designation using a music generation module;
[1644] means for synchronizing the generated music with the video data;
[1645] means for providing the generated video data to a user;
[1646] A system including:
[1647] (Claim 2)
[1648] 10. The system of claim 1, further comprising means for a user to specify a genre of music.
[1649] (Claim 3)
[1650] 2. The system according to claim 1, further comprising means for generating a different musical pattern for each scene based on the analysis result.
[1651] "Application example 2 when combining emotion engines"
[1652] (Claim 1)
[1653] A means for a user to upload video data;
[1654] A means for storing the video data received by the server;
[1655] A means for the server to analyze the content of the video data using an image analysis module;
[1656] A means for the server to generate appropriate music using a music generation module based on the analysis results;
[1657] means for synchronizing the generated music with the video data;
[1658] means for providing the generated video data to a user;
[1659] A means for specifying a desired music genre from a user's terminal;
[1660] A means for recognizing facial expressions and voices and analyzing emotional states when a user operates video data;
[1661] A system including:
[1662] (Claim 2)
[1663] 10. The system of claim 1, further comprising means for dynamically adjusting the tempo or tone of the music based on the emotional state of the user as he or she interacts with the video data.
[1664] (Claim 3)
[1665] 2. The system according to claim 1, further comprising means for generating a different musical pattern for each scene based on the analysis result, and generating music that reflects the emotional state of the user. [Explanation of symbols]
[1666] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for a user to upload video data; A means for storing the video data received by the server; A means for the server to analyze the content of the video data using an image analysis module; A means for the server to generate appropriate music using a music generation module based on the analysis results; means for synchronizing the generated music with the video data; means for providing the generated video data to a user; A system including:
2. 10. The system of claim 1, further comprising means for a user to specify a genre of music.
3. 2. The system according to claim 1, further comprising means for generating a different musical pattern for each scene based on the analysis results.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A