System
The system efficiently generates high-quality music videos by analyzing audio files and user input, addressing the challenge of high costs and skill requirements in traditional music video production.
Patent Information
- Application Number
- JP2024123994
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Producing professional-quality music videos requires significant costs and specialized skills, making it difficult for artists with limited resources to easily create high-quality music videos.
A system that analyzes audio files for lyrics and beat information, receives user-provided character images and selection information, and generates music videos using a generative model, allowing users to create high-quality music videos efficiently without specialized skills or expensive equipment.
Enables artists to produce professional-quality music videos quickly and cost-effectively, reducing the need for specialized skills and expensive equipment.
Smart Images

Figure 2026022477000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, many artists have expressed a desire to create music videos to widely promote their work. However, producing professional-quality music videos requires significant costs and specialized skills. This makes it difficult for artists with limited resources to easily produce high-quality music videos. The present invention aims to solve these problems by providing a system that can efficiently and easily generate high-quality music videos. [Means for solving the problem]
[0005] The present invention provides a system that receives an audio file from a user and analyzes lyrics and beat information from the audio file. It also includes a means for receiving from the user photos of characters and selection information regarding the image and atmosphere of the music video, and generating a music video using a generative model based on this information. It also includes a means for providing the generated music video to the user. Specifically, it includes a means for generating video clips based on the analyzed lyrics and beat information and the selection information, and rendering the video clips as a sequence. The audio file analysis means also includes an algorithm for extracting beat maps and lyrics. This invention enables many artists to easily create professional-quality music videos, efficiently promoting their work.
[0006] "Audio File" means a digital file containing music or audio data that a User uploads to the System.
[0007] "Lyrics" are linguistic phrases or poems contained in an audio file, and are text data used to convey the content of a song.
[0008] "Beat information" refers to information about the rhythm and tempo of an audio file, and is data that indicates the temporal structure of a song.
[0009] "Photos of characters" refer to image data of people to be featured in the music video provided by the user.
[0010] "Image / Mood" refers to information that indicates the user's preference for the look, style, and visuals of the music video.
[0011] "Selection Information" refers to specific selections or instructions regarding the photos of characters and the image and atmosphere provided by the user.
[0012] A "generative model" is an algorithm or AI (artificial intelligence) that automatically generates a music video based on the lyrics and beat information of a received audio file, as well as selection information.
[0013] "Video clips" are the individual video segments or sequences that make up a music video.
[0014] A "sequence" refers to the order or structure in which video clips are played consecutively, and is an element that creates the overall flow of a music video.
[0015] "Rendering" is the process of combining the generated video clips and exporting them as the final music video.
[0016] An "algorithm" is a computational procedure or processing method for solving a specific problem, and is a mathematical method used to analyze audio files. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] The system of the present invention realizes a series of processes that allow users to automatically generate music videos (MVs) based on audio files. The specific operation of the system and the processing of the program are explained below in natural language.
[0039] Server-side processing
[0040] 1. Receiving audio files
[0041] The server receives audio files uploaded by users.
[0042] Example: "Store the file "song.mp3" uploaded by the user on the server."
[0043] 2. Metadata Extraction
[0044] The server extracts metadata such as lyrics and beat information from the received audio file and stores it in a database.
[0045] Example: "Analyze lyrics and beatmaps using audio analysis algorithms."
[0046] 3. Receiving User Selection Information
[0047] The server receives user-submitted photos of the characters and selection information regarding the image and atmosphere of the music video.
[0048] Example: "Receive the settings for the user-selected 'Vintage' style."
[0049] 4. Launching the Generative Model
[0050] Based on the parsed metadata and user selection information, the server activates the AI generative model and begins generating the music video.
[0051] For example: "The AI model generates a video sequence that corresponds to the lyrics and applies the selected mood."
[0052] 5. Rendering
[0053] The server combines the generated video clips into a sequence to render a high-quality music video.
[0054] For example: "Combine the sequences and render to produce the final video file."
[0055] 6. Providing results
[0056] The server provides the generated music video to the user and generates and sends a download link.
[0057] For example: "Provide users with a link to the finished video file so they can download it."
[0058] Terminal side processing
[0059] 1. Interface display
[0060] The terminal displays an interface for the user to upload an audio file.
[0061] Example: "Display a UI (user interface) with a button to upload an audio file."
[0062] 2. Upload audio file
[0063] The terminal uploads the audio file selected by the user to the server.
[0064] Example: "Send the selected file 'song.mp3' to the server."
[0065] 3. Select the image and atmosphere
[0066] The terminal receives photos of characters and selection information about the image and atmosphere to be created from the user and transmits them to the server.
[0067] Example: "Send user-selected image information to the server."
[0068] 4. Progress Display
[0069] The terminal displays the progress of the creation process in real time as received from the server.
[0070] Example: "Show status 'Generating video... 50% complete'."
[0071] 5. View and download the final result
[0072] The device displays a preview of the generated music video and provides a download link.
[0073] Example: "Display a screen to preview the generated video and provide a download link."
[0074] User processing
[0075] 1. Select an audio file
[0076] The user selects an audio file to upload through the interface.
[0077] For example: "Select 'song.mp3' from local storage."
[0078] 2. Select the image and atmosphere
[0079] Users select photos of the characters and the image and atmosphere of the music video.
[0080] For example: "Choose a profile picture and a 'vintage' style."
[0081] 3. File upload and information transmission
[0082] The user transmits the selected audio file and image information to the server through the terminal.
[0083] For example: "Click the upload button to submit your audio file and selection information."
[0084] 4. Check the generation progress
[0085] The user checks the progress of the creation displayed on the terminal.
[0086] For example: "Check that the progress bar reaches 50%."
[0087] 5. Download and share your finished product
[0088] Users can download the generated music video and share it on social media and other platforms.
[0089] For example: "Click the download link to save the video and upload it to YouTube to share."
[0090] The above is a specific embodiment of the present invention, which enables artists and content creators to produce high-quality music videos in a short amount of time without requiring special skills or expensive equipment.
[0091] The processing flow will be explained below.
[0092] Step 1:
[0093] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[0094] Step 2:
[0095] The terminal uploads the audio file selected by the user to the server, specifically, by transmitting the selected audio file to the server via an HTTP request.
[0096] Step 3:
[0097] The server receives the audio file sent from the terminal and saves it in a specified directory.
[0098] Step 4:
[0099] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio files, and the results of this analysis are stored in a database.
[0100] Step 5:
[0101] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[0102] Step 6:
[0103] The terminal transmits the character photos and image / atmosphere selection information received from the user to the server. The selection information includes image files and text data.
[0104] Step 7:
[0105] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model.
[0106] Step 8:
[0107] The server then activates the AI generation model based on the analyzed lyrics and beat information, as well as the user's selection information, and begins generating the music video. Specifically, the AI model processes the input data and generates a corresponding video clip.
[0108] Step 9:
[0109] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[0110] Step 10:
[0111] The server saves the completed music video as a file and generates a download link.
[0112] Step 11:
[0113] The server sends the generated download link to the terminal, allowing the user to access it.
[0114] Step 12:
[0115] The device displays a preview of the generated music video to the user and provides a download link, allowing the user to view, download, and save the generated music video.
[0116] Step 13:
[0117] Users can download the finished music video and share it on social media and other platforms, at which point they can also view the video and provide feedback.
[0118] Example 1
[0119] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0120] Traditional music video production requires expensive equipment and specialized skills, making it difficult for many artists and content creators. Manual editing is also time-consuming and inefficient. In response, there was a need for a method to automatically generate high-quality music videos from audio files.
[0121] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0122] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and atmosphere of the music video, means for generating a music video using a generative AI model based on the analyzed lyrics and beat information and the selection information, means for providing the generated music video to the user and generating and transmitting a download link, and means for displaying the progress of the music video generation to the user in real time, thereby enabling users to efficiently generate high-quality music videos without requiring special skills or expensive equipment.
[0123] "Audio files" are music or audio data files uploaded by users.
[0124] "Lyrics" refers to the words or sentences within an audio file, which may contain poetic expressions or narrative.
[0125] "Beat information" is data that indicates the rhythm and tempo within an audio file.
[0126] "Character Images" are photographs or illustrations of people that appear in the music video.
[0127] "Music video atmosphere" refers to the overall style or theme of the video selected by the user, such as "vintage" or "futuristic."
[0128] A "generative AI model" is an artificial intelligence algorithm that generates video sequences or video clips based on audio file data and user selections.
[0129] "Rendering" is the process of combining the generated video clips into a sequence to create the final video file.
[0130] "Download link" refers to a URL that allows a user to download the generated music video.
[0131] The "means for displaying the generation progress status in real time" is a function for displaying the progress of the server-side processing on the user terminal.
[0132] The present invention provides a system that allows users to automatically generate music videos (MVs) based on audio files. The specific operation and processing of the system are described below. The system mainly operates in cooperation with three parties: a server, a terminal, and a user.
[0133] Hardware and software used
[0134] Hardware:
[0135] Server: Cloud server or on-premise server.
[0136] Device: A PC, smartphone, tablet, or other device capable of operating a user interface.
[0137] software:
[0138] Speech analysis library: Librosa.
[0139] Video generation algorithms: Generative AI models (e.g., OpenAI's GPT-3, DALL-E).
[0140] Video editing tool: FFmpeg.
[0141] Web frameworks: Django, Flask, etc.
[0142] Example of system operation
[0143] Receive audio files:
[0144] The user selects an audio file to upload through the device interface. Example: "The user selects 'song.mp3' from local storage."
[0145] The device uploads this audio file to the server. Example: "Send the selected 'song.mp3' file to the server."
[0146] Metadata Extraction:
[0147] The server receives the audio file and parses it for lyrics and beat information using an audio analysis library such as Librosa. Example: "Using Librosa, parse 'uploads / song.mp3' for lyrics and beatmap data and store them in a database."
[0148] Receive user selection information:
[0149] Through the interface, users select the style and character images for the music video. For example, "User selects profile picture and 'vintage' style."
[0150] The terminal transmits this information to the server.
[0151] Launch the generative model:
[0152] The server then triggers a generative AI model based on the parsed metadata and user selections to generate a music video. For example, "The AI model takes lyrics and beatmaps as input and generates a video sequence."
[0153] rendering:
[0154] The server combines the generated video clips using a video editing tool such as FFmpeg and renders a high-quality music video. Example: "Combine the generated video clips using FFmpeg as 'final_video.mp4' and render."
[0155] Results provided:
[0156] The server generates a download link for the generated music video and provides it to the user. Example: "Email the user a download link for the generated 'final_video.mp4'."
[0157] The device has the ability to display the generation progress to the user in real time. For example, "Display the status 'Generating video... 50% complete' on the screen."
[0158] Examples of prompt statements
[0159] For example, the following prompt sentence is input to a generative AI model:
[0160] "Generate a vintage-style video sequence based on the lyrics of this audio file."
[0161] The video clips thus generated can be combined into a sequence to obtain the final video file.
[0162] This invention enables users to efficiently generate high-quality music videos without requiring special skills or expensive equipment, which significantly reduces time and costs compared to conventional methods.
[0163] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0164] Server-side processing
[0165] Step 1: Receive the audio file
[0166] The server receives the audio file uploaded by the user and stores it in a specified directory.
[0167] Input: An audio file (e.g. song.mp3) that the user uploads from their device.
[0168] Data processing: Send the audio file to the server as an HTTP request.
[0169] Output: An audio file saved in the specified directory on the server (e.g. uploads / song.mp3).
[0170] Specific behavior: "The server receives the HTTP request and saves the audio file to 'uploads / song.mp3'."
[0171] Step 2: Metadata extraction
[0172] The server extracts lyrics and beat information from the received audio file using an audio analysis library such as Librosa.
[0173] Input: A saved audio file (e.g. uploads / song.mp3).
[0174] Data processing: Analysis of audio files using Librosa.
[0175] Output: Extracted lyrics and beat information (stored in a database).
[0176] Specific operation: "Using Librosa, parse the lyrics and beatmap from the audio file 'uploads / song.mp3' and store them in a database."
[0177] Step 3: Receiving User Selection Information
[0178] The server receives user-submitted images of the characters and selection information regarding the image and atmosphere of the music video.
[0179] Input: Image file and selection information (e.g., vintage style) sent by the user from the device.
[0180] Data processing: The selected information is saved in the server database.
[0181] Output: User selection information and images stored on the server.
[0182] Specific operation: "The image information selected by the user and the character's image are saved in the 'user_preferences' directory on the server."
[0183] Step 4: Launching the generative model
[0184] The server then activates the generative AI model based on the parsed metadata and user selections to begin generating the music video.
[0185] Input: Parsed lyrics, beat information, user selection information and images.
[0186] Data processing: Input to a generative AI model to generate a video sequence according to the prompt.
[0187] Output: The generated video clip.
[0188] What it does: "The AI model takes lyrics and a beatmap as input and generates a video sequence based on a prompt."
[0189] Step 5: Rendering
[0190] The server then combines the generated video clips using a video editing tool such as FFmpeg to render a high-quality music video.
[0191] Input: The generated video clip.
[0192] Data processing: Video sequences were combined using FFmpeg.
[0193] Output: The final video file (e.g. final_video.mp4).
[0194] Specific operation: "Use FFmpeg to combine the generated video clips into 'final_video.mp4' and render it."
[0195] Step 6: Delivering results
[0196] The server generates and provides a download link for the generated music video to the user.
[0197] Input: The rendered music video file.
[0198] Data processing: Generate download links.
[0199] Output: The download link sent to the user.
[0200] Specific action: "Send the download link for the generated 'final_video.mp4' to the user via email."
[0201] Terminal side processing
[0202] Step 1: Interface display
[0203] The terminal displays an interface for the user to upload an audio file and enter selection information.
[0204] Input: None.
[0205] Data processing: GUI generation.
[0206] Output: The displayed interface.
[0207] Specific behavior: "When a user opens a browser, a button for uploading an audio file appears."
[0208] Step 2: Upload your audio file
[0209] The terminal uploads the audio file selected by the user to the server.
[0210] Input: User selected audio file (e.g. song.mp3).
[0211] Data processing: Send the audio file to the server.
[0212] Output: Audio file uploaded to the server.
[0213] Specific behavior: "When the user selects 'song.mp3' and clicks the upload button, this file will be sent to the server."
[0214] Step 3: Choose the image and mood
[0215] The terminal receives from the user the images of the characters and the selection information regarding the image and atmosphere of the music video, and transmits the information to the server.
[0216] Input: User-selected image file and image information (e.g., vintage style).
[0217] Data processing: Send the selected information to the server.
[0218] Output: Selection information and images sent to the server.
[0219] What it does: "When a user selects the 'Vintage' style and uploads a photo, this information is sent to a server."
[0220] Step 4: Viewing progress
[0221] The terminal displays the progress of the creation process in real time as received from the server.
[0222] Input: Progress data sent from the server.
[0223] Data processing: Display of progress data.
[0224] Output: Real-time updated progress indicator.
[0225] Specific behavior: "Displays the status 'Generating video... 50% complete' on the screen."
[0226] Step 5: View and download the final result
[0227] The device displays a preview of the generated music video and provides a download link.
[0228] Input: The download link sent by the server.
[0229] Data processing: Display links and preview screens.
[0230] Output: Preview and downloadable link.
[0231] What it does: "You'll be presented with a screen that previews the generated video and provides a download link below it."
[0232] User processing
[0233] Step 1: Select an audio file
[0234] The user selects an audio file to upload through the terminal interface.
[0235] Input: None.
[0236] Data processing: Using the file selection dialog.
[0237] Output: The selected audio file.
[0238] Specific behavior: "The user selects 'song.mp3' from local storage in the file selection dialog."
[0239] Step 2: Choose the image and mood
[0240] Users select images of the characters and the image and atmosphere of the music video.
[0241] Input: None.
[0242] Data manipulation: Use of drop-down menus and file upload functions.
[0243] Output: Selected image files and image information.
[0244] What it does: "User drags and drops a profile photo and selects the 'Vintage' style from the drop-down menu."
[0245] Step 3: Upload files and submit information
[0246] The user transmits the selected audio file and image information to the server through the terminal.
[0247] Input: Audio files and image information.
[0248] Data processing: Sending audio files and selected information.
[0249] Output: Audio file and selection information sent to the server.
[0250] Specific action: "Click the upload button to send the audio file and selected information to the server."
[0251] Step 4: Check the generation progress
[0252] The user can check the progress of the creation displayed on the terminal in real time.
[0253] Input: Progress data sent from the server.
[0254] Data processing: Checking progress data.
[0255] Output: Progress display on screen.
[0256] What it does: "A progress bar appears on the screen and updates in real time to show progress, such as 50% complete."
[0257] Step 5: Download and share your finished product
[0258] Users can download the generated music videos to their devices and share them on social media and other platforms.
[0259] Input: Download link.
[0260] Data processing: Use of download links and retrieval of files.
[0261] Output: Downloaded music video files.
[0262] Action: "Click the download link to save 'final_video.mp4' locally, then upload the video to YouTube to share it."
[0263] (Application example 1)
[0264] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0265] Traditional music video production required advanced expertise and expensive equipment, making it difficult for average users to easily create high-quality music videos. Furthermore, there was no way for users to check the progress of automatically generated music videos in real time, nor was there an easy way to directly post generated videos to content distribution services. This reduced user convenience and prevented content from being quickly shared.
[0266] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0267] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for generating a music video using a generative model based on the analyzed lyrics and beat information and the selection information, means for displaying the progress of the generated music video, and means for providing the generated music video to the user and enabling it to be directly posted to a content distribution service. This allows users to create high-quality music videos without specialized knowledge or expensive equipment, to check the progress in real time, and to quickly share the generated video with a content distribution service.
[0268] An "audio file" is a digital file containing music or sound that is uploaded by a user.
[0269] "Lyrics" refers to words and phrases in an audio file, and are text data that make up the sung portion of the music.
[0270] "Beat information" is data about rhythm and tempo obtained from an audio file, and indicates the rhythmic structure of the music.
[0271] "Character images" are photographs or illustrations of people to appear in the music video.
[0272] "Video image and atmosphere" refers to the overall style and visual theme of the music video, and is the visual atmosphere and taste specified by the user.
[0273] "Selection information" refers to information about the image and atmosphere of the characters and video set by the user.
[0274] A "generative model" is an algorithm or framework that uses AI technology to automatically generate music videos.
[0275] A "music video" is video content generated based on an audio file and includes visual storytelling set to music.
[0276] "Content distribution service" means an online service for distributing digital content, such as generated music videos, and includes, for example, a video sharing platform.
[0277] MODE FOR CARRYING OUT THE INVENTION
[0278] A detailed description of an embodiment of the present invention is provided below: The system allows users to automatically generate music videos from audio files, display the progress of the generation, and post the generated music videos directly to a content distribution service.
[0279] Server-side processing
[0280] 1. Receiving audio files: The server receives and stores audio files uploaded by users. The software used is Flask (Python).
[0281] 2. Metadata extraction: The server parses the lyrics and beat information from the received audio files and stores them in a database. This process uses the librosa and speech_recognition libraries.
[0282] 3. Receiving user selection information: The server receives the character images and music video image / atmosphere selection information sent by the user and stores them in a database using Flask and MySQL.
[0283] 4. Launching the generative model: The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selection information. The generative model uses TensorFlow or PyTorch.
[0284] 5. Rendering: The server combines the generated video clips into a sequence and renders the final music video using OpenCV software.
[0285] 6. Progress display: The server manages the progress of the generation process in real time and notifies the user.
[0286] 7. Result Serving: The server serves the generated music video to the user, generates and sends a download link, and also allows the user to submit the generated video directly to a content distribution service.
[0287] Terminal side processing
[0288] 1. Interface display: The terminal displays an interface for users to upload audio files. This interface is built using React Native.
[0289] 2. Audio file upload: The terminal uploads the audio file selected by the user to the server.
[0290] 3. Image and atmosphere selection: The terminal receives selection information from the user regarding the character images and the image and atmosphere of the music video, and transmits it to the server.
[0291] 4. Progress display: The terminal displays the progress of the generation process received from the server in real time.
[0292] 5. View, download, and share the final result: The device displays a preview of the generated music video, provides a download link, and also allows users to post it directly to content distribution services such as video sharing platforms.
[0293] User processing
[0294] 1. Audio file selection: The user selects an audio file to upload through the interface.
[0295] 2. Image / Mood Selection: Users select the image of the characters and the image / mood of the music video.
[0296] 3. File upload and information transmission: The user transmits the selected audio file and image information to the server through the terminal.
[0297] 4. Checking the progress of generation: The user checks the progress of generation displayed on the terminal.
[0298] 5. Download and share the finished product: Users can download the generated music video and share it on social media or other platforms, or post the generated video directly to a content distribution service.
[0299] Specific examples
[0300] As a concrete example, consider a situation where a user wants to create a music video using a new song by a popular Japanese band that can be shared on social media. Using the application, the user uploads an MP3 file of the new song and selects a particular music video theme (e.g., cyberpunk). An example prompt for this would be:
[0301] "Create a cyberpunk-style video that makes it easy to hear the lyrics. Use the attached image as a photo of the characters and a futuristic city in the background."
[0302] Based on this prompt, the generative AI model automatically generates a music video that reflects the style and content specified by the user.
[0303] In this way, users can create professional music videos in a short time without needing special skills or expensive equipment, and post them to content distribution services.
[0304] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0305] Step 1:
[0306] The server receives audio files from users. The audio files (e.g., MP3 format) uploaded by users are received via HTTP requests and stored in cloud storage such as an AWS S3 bucket. The input of this process is the audio file uploaded by the user, and the output is the path to the file stored in cloud storage.
[0307] Step 2:
[0308] The server analyzes the lyrics and beat information from the received audio file. It uses the librosa and speech_recognition libraries to extract rhythm, tempo (beat information), and lyrics from the audio file. The input for this process is the audio file read from cloud storage, and the output is the metadata (beat information and lyric data) resulting from the analysis.
[0309] Step 3:
[0310] The server receives the character images and music video image / mood selections sent by the user and stores them in a database using JSON data received via an HTTP POST request. The input to this process is the user-selected image files and image / mood selections, and the output is a record of the selections stored in the database.
[0311] Step 4:
[0312] The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selections. TensorFlow or PyTorch is used as the generative model, and a video sequence is generated by inputting prompts into the AI model. The inputs for this process are metadata and the user's selections, and the output is the generated video clip.
[0313] Step 5:
[0314] The server combines the generated video clips into a sequence and renders the final music video. OpenCV is used to combine multiple video clips into a sequence and encode them to generate the final video file. The input to this process is a list of video clips, and the output is the final rendered music video file.
[0315] Step 6:
[0316] The server manages the progress of the generation process in real time and notifies the device. It tracks the progress and notifies the device of progress updates sequentially using WebSocket or a real-time communication protocol. The input of this process is the current status of the generation process, and the output is progress update messages.
[0317] Step 7:
[0318] The server provides the generated music video to the user and generates a download link to send to the device. It also allows the user to post the generated video directly to a content distribution service by generating an AWS S3 download link and providing the endpoint URL to the user. The input of this process is the generated music video file, and the output is the download link and the posting endpoint.
[0319] Step 8:
[0320] The terminal displays an interface for the user to upload an audio file. React Native is used to build the user interface and display buttons and selection items for uploading audio files. The input is the design information for the user interface, and the output is the interface screen displayed to the user.
[0321] Step 9:
[0322] The terminal uploads the audio file selected by the user to the server. The audio file obtained from the file input is sent to the server via an HTTP request. The input of this process is the audio file selected by the user, and the output is the request sent to the server.
[0323] Step 10:
[0324] The device receives the user's selection information about the character images and the image and atmosphere of the music video, and sends it to the server. The device then obtains the selection information and sends it to the server in JSON format. The input of this process is the user-selected image file and image and atmosphere information, and the output is the request sent to the server.
[0325] Step 11:
[0326] The terminal displays the progress of the creation process in real time as it receives it from the server. It receives the progress via WebSocket or real-time communication and displays it in the interface as a progress bar or status messages. The input to this process is the progress update messages from the server, and the output is a progress screen that is displayed to the user.
[0327] Step 12:
[0328] The device displays a preview of the generated music video, provides a download link, and also provides the ability to directly post the video to content distribution services such as video sharing platforms. The input to this process is a download link from the server and the video file, and the output is a preview screen and a download link displayed to the user.
[0329] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0330] The system of the present invention not only includes a series of processes for users to automatically generate music videos (MVs) based on audio files, but also a function to recognize the user's emotions and customize the music video based on those emotions. Specific system operations and program processing are explained in natural language below.
[0331] Server-side processing
[0332] 1. Receiving audio files
[0333] The server receives audio files uploaded by users. For example, "Receive and store the file 'song.mp3' uploaded by the user."
[0334] 2. Metadata Extraction
[0335] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio file. The analysis results are stored in a database. For example, "Store the analyzed lyrics and beat information."
[0336] 3. Receiving User Selection Information
[0337] The server receives the user-submitted character photos and selection information about the image and atmosphere of the music video. For example, "receives user-selected 'vintage' style setting information."
[0338] 4. Activating the Emotional Engine
[0339] The server launches an emotion engine to analyze the user's voice file and real-time emotion data. The emotion engine identifies the user's emotional state and applies corresponding parameters to the generative model. For example, "Detect the emotion 'happiness' from the voice file."
[0340] 5. Launching the Generative Model
[0341] The server then activates the AI generative model based on the parsed metadata, user selection information, and emotional data obtained from the emotion engine, and begins generating the music video, for example, "applying 'bright color tones' to the scenes based on the emotional state."
[0342] 6. Rendering
[0343] The server then stitches together the resulting video clips into a sequence and renders it into a high-quality music video. This rendering process ensures smooth transitions between frames. For example, "Rendered for saving as a final video file."
[0344] 7. Providing results
[0345] The server provides the generated music video to the user and generates and sends a download link. For example, "Provide the user with a download link for the completed music video."
[0346] Terminal side processing
[0347] 1. Interface display
[0348] The terminal displays an interface for the user to upload an audio file. The user selects an audio file and clicks the upload button. For example, "Display a web page with an audio file upload button."
[0349] 2. Upload audio file
[0350] The terminal uploads the audio file selected by the user to the server. The selected audio file is sent to the server via an HTTP request. For example, "Send audio file to server."
[0351] 3. Select the image and atmosphere
[0352] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos and images and clicks the submit button. For example, "Display a UI showing photos of the characters and image options."
[0353] 4. Progress Display
[0354] The device will display real-time progress of the generation process as it receives it from the server, e.g., "Displays status 'Generating video... 50% complete'."
[0355] 5. View and download the final result
[0356] The device displays a preview of the generated music video and provides a download link. The user can check and download the generated music video. For example, "Display a preview screen of the generated video and a download link."
[0357] User processing
[0358] 1. Select an audio file
[0359] The user selects an audio file to upload through the interface, for example, "Select 'song.mp3' from local storage."
[0360] 2. Select the image and atmosphere
[0361] Users can choose the image and atmosphere of the character photos and music videos. For example, "Choose a profile picture and a 'vintage' style."
[0362] 3. File upload and information transmission
[0363] The user sends the selected audio file and image information to the server through the terminal. For example, "Click the upload button to send the audio file and selected information."
[0364] 4. Check the generation progress
[0365] The user checks the generation progress displayed on the device, for example, "Check that the progress bar reaches 50%."
[0366] 5. Download and share your finished product
[0367] Users can download the generated music video and share it on social media or other platforms. For example, "Click the download link to save the video, then upload it to YouTube to share."
[0368] The above is a specific example of how to implement a system incorporating the emotion engine of the present invention. This allows artists and content creators to easily create high-quality music videos customized to their own emotional state. Utilizing emotion data from the emotion engine enables more personalized visual expression, making a strong impact on viewers.
[0369] The processing flow will be explained below.
[0370] Step 1:
[0371] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[0372] Step 2:
[0373] The device sends the audio file selected by the user to the server via an HTTP request, specifically by using an input form or drag-and-drop function.
[0374] Step 3:
[0375] The server receives the audio file sent from the device and saves it in a specified directory, for example, a file named "song.mp3."
[0376] Step 4:
[0377] The server analyzes the received audio files and runs audio analysis algorithms to extract lyrics and beat information, and stores the analyzed metadata in a database.
[0378] Step 5:
[0379] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[0380] Step 6:
[0381] The device sends the character photos and image / atmosphere selection information received from the user to the server by sending the image files and text data via an HTTP request.
[0382] Step 7:
[0383] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model, which includes preprocessing the image data and analyzing the text data.
[0384] Step 8:
[0385] The server activates an emotion engine to analyze the user's voice file to identify their emotional state, for example, using a voice analysis algorithm to identify emotions such as "happiness" or "sadness."
[0386] Step 9:
[0387] The server adjusts the parameters of the generative model based on the identified emotional state, specifically dynamically changing the color grading and effects of the video based on the emotional data.
[0388] Step 10:
[0389] The server then activates an AI generation model based on the analyzed metadata, user selection information, and emotional data to begin generating the music video. The AI model processes the input data and generates a corresponding video clip.
[0390] Step 11:
[0391] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[0392] Step 12:
[0393] The server saves the completed music video as a file and generates a download link that can be accessed by the user.
[0394] Step 13:
[0395] The server sends the generated download link to the terminal, allowing the user to access it.
[0396] Step 14:
[0397] The terminal displays a preview of the generated music video to the user and provides a download link, allowing the user to view and download the generated music video.
[0398] Step 15:
[0399] Users can download the completed music video and share it on social media or other platforms, such as uploading it to YouTube to share with their audience.
[0400] Example 2
[0401] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0402] Conventional music video generation systems have struggled to generate personalized videos that take into account the user's emotional state. Furthermore, extracting metadata from audio files and generating videos based on user-selected images and moods are time-consuming, limiting the ability to improve the user experience. The present invention aims to solve these problems.
[0403] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0404] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and mood of a video of a visual work, means for analyzing the audio file and real-time emotional data to identify the user's emotional state, means for generating a visual work using a generative model based on the analyzed lyrics and beat information, the selection information, and the identified emotional state, and means for providing the generated visual work to the user, thereby enabling the automatic generation of personalized, high-quality music videos that take the user's emotions into consideration.
[0405] "Audio File" refers to audio data stored in digital format.
[0406] "Lyrics" refers to the words or sentences in a song within an audio file.
[0407] "Beat information" refers to data about rhythm and tempo contained in an audio file.
[0408] "Analysis" refers to the process of extracting useful information from input data.
[0409] "Character Images" refers to photographs or illustrations of people appearing in the music video.
[0410] "Visual work" refers to content that is visually displayed, such as a film or video clip.
[0411] "Image / Mood" refers to the style or theme expressed in a visual work.
[0412] "Emotional state" refers to the psychological state of the user analyzed from the voice data.
[0413] "Real-time emotional data" refers to data that indicates the current emotional state of the user.
[0414] "Generative model" refers to an algorithm that uses AI techniques to automatically generate visual works or other content.
[0415] "Generation" refers to the process of creating new data or content based on specific data.
[0416] "Rendering" refers to the process of converting digital data into a visually displayable form.
[0417] "Serving" refers to the process of transmitting or displaying the generated visual work in a form accessible to a user.
[0418] MODE FOR CARRYING OUT THE INVENTION
[0419] The system of the present invention is a system that includes a series of processes for a user to automatically generate a music video (MV) based on an audio file, recognize the user's emotions, and customize the MV based on those emotions. Specific operations of the system and program processing are described in detail below.
[0420] Server-side processing
[0421] 1. Receiving audio files
[0422] The server receives the audio file uploaded by the user. In this receiving process, the HTTP protocol is used to send the file "song.mp3" to the server and save it in a specified folder.
[0423] 2. Metadata Extraction
[0424] The server uses audio analysis algorithms (e.g., Python libraries librosa or PyDub) to extract lyrics and beat information from audio files. For example, it loads audio using the librosa.load() function and extracts beat information using the librosa.beat.beat_track() function. The analysis results are stored in a database.
[0425] 3. Receiving User Selection Information
[0426] The server receives and stores the character images and the selection information about the image and atmosphere of the visual work sent by the user via HTTP request, for example, "vintage style" or "character: profile.jpg" in the database.
[0427] 4. Activating the Emotional Engine
[0428] The server invokes an emotion engine (e.g., a general emotion analysis service) to analyze the audio file and real-time emotion data. The engine identifies the user's emotional state (e.g., happiness, sadness, etc.) and provides corresponding parameters to the generative model.
[0429] 5. Launching the Generative Model
[0430] The server generates a music video by activating an AI generation model based on the analyzed lyrics, beat information, user selection information, and emotional state. For example, the server inputs prompts to customize scene settings and color tones according to the user's emotional state into the generation AI model.
[0431] 6. Rendering
[0432] The server then combines the generated video clips into a sequence and renders the final video. This process uses a video processing tool such as FFmpeg. For example, to concatenate multiple clips, a command like ffmpeg -i input1.mp4 -i input2.mp4 -filter_complex "[0:v][1:v] concat=n=2:v=1[outv]" -map "[outv]" output.mp4 is used.
[0433] 7. Providing results
[0434] The server generates a download link for the completed music video and provides it to the user, returning a response JSON object containing the link URL.
[0435] Terminal side processing
[0436] 1. Interface display
[0437] The terminal displays an interface for the user to upload an audio file. The interface includes a file selection button and an upload button. For example, the terminal uses an HTML form or JavaScript to prompt the user to select an audio file.
[0438] 2. Upload audio file
[0439] The device uploads the audio file selected by the user to the server via an HTTP request, specifically an AJAX request, sending the file "song.mp3" to the server.
[0440] 3. Select the image and atmosphere
[0441] The device displays an interface that allows users to select the image of the characters and the image and atmosphere of the music video. For example, a selection screen built with HTML5 and CSS is displayed, allowing users to select "profile.jpg" or a "vintage" style.
[0442] 4. Progress Display
[0443] The terminal displays the progress of the generation process received from the server in real time. It uses the JavaScript setInterval function to query the server for the progress at regular intervals and displays the results.
[0444] 5. View and download the final result
[0445] The device will display a preview of the generated music video and provide a download link. <video>View by tag and download link Provided by tag.
[0446] User processing
[0447] 1. Select an audio file
[0448] The user selects the audio file to upload from the interface, specifically the file "song.mp3" from local storage.
[0449] 2. Select the image and atmosphere
[0450] The user selects the image and atmosphere of the characters and music video. They select "profile.jpg" from local storage and choose "vintage style" from the options provided in the UI.
[0451] 3. File upload and information transmission
[0452] The user sends the selected audio file and image information to the server through the terminal, and clicks the upload button to send the audio file and the selected information.
[0453] 4. Check the generation progress
[0454] The user checks the generation progress displayed on the device, specifically checking the progress message such as "Make sure the progress bar reaches 50%."
[0455] 5. Download and share your finished product
[0456] Users can download the generated music video and share it on social media or other platforms by clicking the download link to save the 'output.mp4' video, then uploading the video to YouTube or social media to share it.
[0457] Examples of concrete examples and prompts
[0458] Prompt Sentence Examples
[0459] Audio file: song.mp3
[0460] Character photo:profile.jpg
[0461] Music video vibe: Vintage
[0462] These procedures and system configurations allow users to easily generate personalized, high-quality music videos based on audio files and emotions.
[0463] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0464] Step 1:
[0465] The server receives an audio file from the user. The input is the audio file selected by the user (e.g., 'song.mp3'), and the output is the audio file stored on the server. Specific operations include requesting the audio file using the HTTP protocol and saving it in a specific folder on the server.
[0466] Step 2:
[0467] The server analyzes the lyrics and beat information from the received audio file. The input is the audio file saved in step 1, and the output is the lyrics and beat information. Specifically, it uses a Python library (e.g., librosa or PyDub) to load the audio with librosa.load() and extract the beat with librosa.beat.beat_track(). The analysis results are stored in a database.
[0468] Step 3:
[0469] The server receives from the user the image of the characters and the selection information regarding the image and atmosphere of the visual work video. The input is the image file and text information selected by the user, and the output is the selection information saved on the server. Specifically, the data sent via the HTTP request is saved in a database. For example, information such as "vintage style" or "character: profile.jpg" is recorded.
[0470] Step 4:
[0471] The server starts an emotion engine to analyze the audio file and real-time emotion data to identify the user's emotional state. The input is the lyrics and beat information analyzed in step 2 and the audio file, and the output is the identified emotional state. Specifically, the emotion analysis service detects the user's emotion (e.g., happiness) from the audio.
[0472] Step 5:
[0473] The server activates a generative AI model based on the analyzed metadata, user selection information, and emotional data to generate a visual work. The input is the data obtained in Steps 2 and 3 and the emotional state identified in Step 4, and the output is the generated visual work (video clip). Specifically, the server provides the generative AI model with prompts to customize the scene setting and color tone according to the emotional state, and generates a music video.
[0474] Step 6:
[0475] The server combines the generated video clips into a sequence and renders it as a final video. The input is the video clip generated in step 5, and the output is a high-quality final video file. Specifically, it uses a video processing tool such as FFmpeg to concatenate multiple clips and executes the ffmpeg command to save it as "output.mp4."
[0476] Step 7:
[0477] The server provides the completed music video to the user. The input is the final video file generated in step 6, and the output is a download link URL. The server then returns a response JSON object containing the URL of the generated video file to the user.
[0478] Step 8:
[0479] The terminal provides the user with an upload interface, a selection interface, a progress indicator, a final result indicator, and a download link, allowing the user to upload audio files, submit selection information, check the progress, and download the generated video.
[0480] (Application example 2)
[0481] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0482] Currently, many artists and content creators spend a great deal of time and effort creating music videos that reflect the emotions and themes of their work. Furthermore, many automated music video generation systems are not customized based on the user's emotional state, and often lack visual appeal. The present invention aims to solve this problem by providing a system that automatically generates high-quality music videos that reflect the user's emotional state.
[0483] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for analyzing the user's emotional state, means for generating video using a generative model based on the analyzed lyrics and beat information, the selection information, and the emotional state, and means for providing the generated video to the user. This enables the generation of a personalized music video based on the user's emotional state.
[0484] An "audio file" is a file in which audio data is stored in digital format.
[0485] "Lyrics" are words or text that correspond to a musical piece in an audio file.
[0486] "Beat information" refers to the rhythm and tempo information within an audio file.
[0487] "Character images" are image data of people appearing in the music video that are uploaded or selected by the user.
[0488] "Image and atmosphere of the video" refers to the visual elements and atmosphere that express the style and theme of the music video.
[0489] "Emotional state" is data that indicates the user's current mental state and emotions.
[0490] "Generative model" refers to an algorithm or AI model that automatically generates music videos based on input data.
[0491] "Video clips" are short video segments that make up a music video.
[0492] "Rendering as a sequence" refers to the process of combining a series of video clips into a continuous video and rendering it at high quality.
[0493] The system of the present invention provides a set of processes for users to automatically generate music videos (MVs) based on audio files, and further includes the ability to recognize the user's emotions and customize the MV based on those emotions.
[0494] The server first receives an audio file from the user. This audio file is audio data stored in a digital format, such as mp3 or wav. The server then uses an audio analysis algorithm to extract lyrics and beat information from the audio file. The extracted metadata is then stored in a database.
[0495] Receives from the user character images and video image and mood selection information that reflects the user's designated visual style or theme.
[0496] The emotion engine is activated to analyze the user's emotional state. The emotion engine identifies emotions from the text data and voice data entered by the user and obtains the emotional state as data.
[0497] The server uses a generative AI model to generate video based on the analyzed lyrics and beat information, user selections, and emotional state. The generative AI model receives these input data as prompts and generates a series of video clips based on them. These video clips are rendered as a sequence and combined into a high-quality video.
[0498] Finally, the generated video is provided to the user: the server saves the video file and generates a link for the user to download it.
[0499] The hardware and software used are as follows: The server uses the Django framework (Python) and an appropriate library (e.g., librosa) for speech analysis; an NLP library (e.g., NLTK, spaCy) or a speech analysis library is used for the emotion engine; and a machine learning framework (e.g., TensorFlow, PyTorch) is used for the generative AI model.
[0500] For example, if a user uploads an up-tempo song called "Happy.mp3," the system extracts lyrics and beat information from the audio file and detects the user's emotional state of happiness. If the user selects a vintage theme, the system applies visual effects that match the theme and generates a light-hearted video that expresses a sense of happiness. The generated music video is immediately available for preview and distribution.
[0501] An example of a prompt is:
[0502] "User ID: 001
[0503] Audio file URL: / media / uploads / song.mp3
[0504] Extracted Metadata: {Lyrics: 'Happy song lyrics', Beat: 'Uptempo'}
[0505] User Sentiment: Happiness
[0506] Selected theme: Vintage
[0507] The above is a specific embodiment of the system of the present invention, which allows artists and content creators to easily create high-quality music videos that are personalized to their emotional state.
[0508] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0509] Step 1:
[0510] The user selects and uploads an audio file.
[0511] Input: An audio file selected by the user from local storage (e.g. song.mp3).
[0512] Specific operation: The terminal creates and sends an HTTP request to send the audio file selected by the user through the interface to the server.
[0513] Output: The audio file is sent to the server.
[0514] Step 2:
[0515] The server receives and stores the audio files.
[0516] Input: User submitted audio file (e.g. song.mp3).
[0517] Specific operation: The server stores the received audio file in local storage or cloud storage.
[0518] Output: The path to the audio file in local or cloud storage.
[0519] Step 3:
[0520] The server analyzes the audio file and extracts the metadata.
[0521] Input: The path to the saved audio file.
[0522] What it does: The server uses audio analysis algorithms to extract lyrics and beat information from the audio file (e.g., using the librosa library).
[0523] Data processing or data manipulation: Spectral analysis of audio data to extract beat maps and lyrics in text format.
[0524] Output: Parsed lyrics and beat information metadata.
[0525] Step 4:
[0526] The server receives from the user the images of the characters and selection information regarding the image and atmosphere of the video.
[0527] Input: User-submitted image file and image selection information.
[0528] Specific operation: The server reads the image file and selection information received as an HTTP request and saves them in a database or storage.
[0529] Output: Path to the saved image file and selection information.
[0530] Step 5:
[0531] The server analyzes the user's emotional state.
[0532] Input: User text or voice data.
[0533] Specific operation: The server launches an emotion engine and analyzes the user's input data to identify their emotional state (e.g., using NLTK or spaCy).
[0534] Data processing or data calculation: Analyzing text or audio data with natural language processing algorithms to generate emotion labels (e.g., happy, sad).
[0535] Output: The user's emotional state.
[0536] Step 6:
[0537] The server generates the video using a generative AI model.
[0538] Input: Parsed lyrics and beat information metadata, user image files and selections, user emotional state.
[0539] Specific operation: The server inputs these input data as prompt sentences into a generative AI model to generate video clips (e.g., using TensorFlow or PyTorch).
[0540] Data processing or data computation: A generative AI model generates video clips based on prompts and renders them as a sequence.
[0541] Output: The generated video clip.
[0542] Step 7:
[0543] The server provides the generated video to the user.
[0544] Input: The generated video clip.
[0545] Specific operation: The server saves the video clip, generates a URL from which the user can download it, and creates and sends an HTTP response that returns the URL to the user.
[0546] Output: A URL to the video clip that users can download.
[0547] These are the processing steps of the system for realizing this application example. Through the specific operations performed at each step, users can create and download music videos that reflect their emotional state.
[0548] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0549] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0550] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0551] [Second embodiment]
[0552] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0553] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0554] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0555] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0556] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0557] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0558] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0559] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0560] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0561] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0562] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0563] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0564] The system of the present invention realizes a series of processes that allow users to automatically generate music videos (MVs) based on audio files. The specific operation of the system and the processing of the program are explained below in natural language.
[0565] Server-side processing
[0566] 1. Receiving audio files
[0567] The server receives audio files uploaded by users.
[0568] Example: "Store the file "song.mp3" uploaded by the user on the server."
[0569] 2. Metadata Extraction
[0570] The server extracts metadata such as lyrics and beat information from the received audio file and stores it in a database.
[0571] Example: "Analyze lyrics and beatmaps using audio analysis algorithms."
[0572] 3. Receiving User Selection Information
[0573] The server receives user-submitted photos of the characters and selection information regarding the image and atmosphere of the music video.
[0574] Example: "Receive the settings for the user-selected 'Vintage' style."
[0575] 4. Launching the Generative Model
[0576] Based on the parsed metadata and user selection information, the server activates the AI generative model and begins generating the music video.
[0577] For example: "The AI model generates a video sequence that corresponds to the lyrics and applies the selected mood."
[0578] 5. Rendering
[0579] The server combines the generated video clips into a sequence to render a high-quality music video.
[0580] For example: "Combine the sequences and render to produce the final video file."
[0581] 6. Providing results
[0582] The server provides the generated music video to the user and generates and sends a download link.
[0583] For example: "Provide users with a link to the finished video file so they can download it."
[0584] Terminal side processing
[0585] 1. Interface display
[0586] The terminal displays an interface for the user to upload an audio file.
[0587] Example: "Display a UI (user interface) with a button to upload an audio file."
[0588] 2. Upload audio file
[0589] The terminal uploads the audio file selected by the user to the server.
[0590] Example: "Send the selected file 'song.mp3' to the server."
[0591] 3. Select the image and atmosphere
[0592] The terminal receives photos of characters and selection information about the image and atmosphere to be created from the user and transmits them to the server.
[0593] Example: "Send user-selected image information to the server."
[0594] 4. Progress Display
[0595] The terminal displays the progress of the creation process in real time as received from the server.
[0596] Example: "Show status 'Generating video... 50% complete'."
[0597] 5. View and download the final result
[0598] The device displays a preview of the generated music video and provides a download link.
[0599] Example: "Display a screen to preview the generated video and provide a download link."
[0600] User processing
[0601] 1. Select an audio file
[0602] The user selects an audio file to upload through the interface.
[0603] For example: "Select 'song.mp3' from local storage."
[0604] 2. Select the image and atmosphere
[0605] Users select photos of the characters and the image and atmosphere of the music video.
[0606] For example: "Choose a profile picture and a 'vintage' style."
[0607] 3. File upload and information transmission
[0608] The user transmits the selected audio file and image information to the server through the terminal.
[0609] For example: "Click the upload button to submit your audio file and selection information."
[0610] 4. Check the generation progress
[0611] The user checks the progress of the creation displayed on the terminal.
[0612] For example: "Check that the progress bar reaches 50%."
[0613] 5. Download and share your finished product
[0614] Users can download the generated music video and share it on social media and other platforms.
[0615] For example: "Click the download link to save the video and upload it to YouTube to share."
[0616] The above is a specific embodiment of the present invention, which enables artists and content creators to produce high-quality music videos in a short amount of time without requiring special skills or expensive equipment.
[0617] The processing flow will be explained below.
[0618] Step 1:
[0619] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[0620] Step 2:
[0621] The terminal uploads the audio file selected by the user to the server, specifically, by transmitting the selected audio file to the server via an HTTP request.
[0622] Step 3:
[0623] The server receives the audio file sent from the terminal and saves it in a specified directory.
[0624] Step 4:
[0625] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio files, and the results of this analysis are stored in a database.
[0626] Step 5:
[0627] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[0628] Step 6:
[0629] The terminal transmits the character photos and image / atmosphere selection information received from the user to the server. The selection information includes image files and text data.
[0630] Step 7:
[0631] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model.
[0632] Step 8:
[0633] The server then activates the AI generation model based on the analyzed lyrics and beat information, as well as the user's selection information, and begins generating the music video. Specifically, the AI model processes the input data and generates a corresponding video clip.
[0634] Step 9:
[0635] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[0636] Step 10:
[0637] The server saves the completed music video as a file and generates a download link.
[0638] Step 11:
[0639] The server sends the generated download link to the terminal, allowing the user to access it.
[0640] Step 12:
[0641] The device displays a preview of the generated music video to the user and provides a download link, allowing the user to view, download, and save the generated music video.
[0642] Step 13:
[0643] Users can download the finished music video and share it on social media and other platforms, at which point they can also view the video and provide feedback.
[0644] Example 1
[0645] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0646] Traditional music video production requires expensive equipment and specialized skills, making it difficult for many artists and content creators. Manual editing is also time-consuming and inefficient. In response, there was a need for a method to automatically generate high-quality music videos from audio files.
[0647] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0648] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and atmosphere of the music video, means for generating a music video using a generative AI model based on the analyzed lyrics and beat information and the selection information, means for providing the generated music video to the user and generating and transmitting a download link, and means for displaying the progress of the music video generation to the user in real time, thereby enabling users to efficiently generate high-quality music videos without requiring special skills or expensive equipment.
[0649] "Audio files" are music or audio data files uploaded by users.
[0650] "Lyrics" refers to the words or sentences within an audio file, which may contain poetic expressions or narrative.
[0651] "Beat information" is data that indicates the rhythm and tempo within an audio file.
[0652] "Character Images" are photographs or illustrations of people that appear in the music video.
[0653] "Music video atmosphere" refers to the overall style or theme of the video selected by the user, such as "vintage" or "futuristic."
[0654] A "generative AI model" is an artificial intelligence algorithm that generates video sequences or video clips based on audio file data and user selections.
[0655] "Rendering" is the process of combining the generated video clips into a sequence to create the final video file.
[0656] "Download link" refers to a URL that allows a user to download the generated music video.
[0657] The "means for displaying the generation progress status in real time" is a function for displaying the progress of the server-side processing on the user terminal.
[0658] The present invention provides a system that allows users to automatically generate music videos (MVs) based on audio files. The specific operation and processing of the system are described below. The system mainly operates in cooperation with three parties: a server, a terminal, and a user.
[0659] Hardware and software used
[0660] Hardware:
[0661] Server: Cloud server or on-premise server.
[0662] Device: A PC, smartphone, tablet, or other device capable of operating a user interface.
[0663] software:
[0664] Speech analysis library: Librosa.
[0665] Video generation algorithms: Generative AI models (e.g., OpenAI's GPT-3, DALL-E).
[0666] Video editing tool: FFmpeg.
[0667] Web frameworks: Django, Flask, etc.
[0668] Example of system operation
[0669] Receive audio files:
[0670] The user selects an audio file to upload through the device interface. Example: "The user selects 'song.mp3' from local storage."
[0671] The device uploads this audio file to the server. Example: "Send the selected 'song.mp3' file to the server."
[0672] Metadata Extraction:
[0673] The server receives the audio file and parses it for lyrics and beat information using an audio analysis library such as Librosa. Example: "Using Librosa, parse 'uploads / song.mp3' for lyrics and beatmap data and store them in a database."
[0674] Receive user selection information:
[0675] Through the interface, users select the style and character images for the music video. For example, "User selects profile picture and 'vintage' style."
[0676] The terminal transmits this information to the server.
[0677] Launch the generative model:
[0678] The server then triggers a generative AI model based on the parsed metadata and user selections to generate a music video. For example, "The AI model takes lyrics and beatmaps as input and generates a video sequence."
[0679] rendering:
[0680] The server combines the generated video clips using a video editing tool such as FFmpeg and renders a high-quality music video. Example: "Combine the generated video clips using FFmpeg as 'final_video.mp4' and render."
[0681] Results provided:
[0682] The server generates a download link for the generated music video and provides it to the user. Example: "Email the user a download link for the generated 'final_video.mp4'."
[0683] The device has the ability to display the generation progress to the user in real time. For example, "Display the status 'Generating video... 50% complete' on the screen."
[0684] Examples of prompt statements
[0685] For example, the following prompt sentence is input to a generative AI model:
[0686] "Generate a vintage-style video sequence based on the lyrics of this audio file."
[0687] The video clips thus generated can be combined into a sequence to obtain the final video file.
[0688] This invention enables users to efficiently generate high-quality music videos without requiring special skills or expensive equipment, which significantly reduces time and costs compared to conventional methods.
[0689] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0690] Server-side processing
[0691] Step 1: Receive the audio file
[0692] The server receives the audio file uploaded by the user and stores it in a specified directory.
[0693] Input: An audio file (e.g. song.mp3) that the user uploads from their device.
[0694] Data processing: Send the audio file to the server as an HTTP request.
[0695] Output: An audio file saved in the specified directory on the server (e.g. uploads / song.mp3).
[0696] Specific behavior: "The server receives the HTTP request and saves the audio file to 'uploads / song.mp3'."
[0697] Step 2: Metadata extraction
[0698] The server extracts lyrics and beat information from the received audio file using an audio analysis library such as Librosa.
[0699] Input: A saved audio file (e.g. uploads / song.mp3).
[0700] Data processing: Analysis of audio files using Librosa.
[0701] Output: Extracted lyrics and beat information (stored in a database).
[0702] Specific operation: "Using Librosa, parse the lyrics and beatmap from the audio file 'uploads / song.mp3' and store them in a database."
[0703] Step 3: Receiving User Selection Information
[0704] The server receives user-submitted images of the characters and selection information regarding the image and atmosphere of the music video.
[0705] Input: Image file and selection information (e.g., vintage style) sent by the user from the device.
[0706] Data processing: The selected information is saved in the server database.
[0707] Output: User selection information and images stored on the server.
[0708] Specific operation: "The image information selected by the user and the character's image are saved in the 'user_preferences' directory on the server."
[0709] Step 4: Launching the generative model
[0710] The server then activates the generative AI model based on the parsed metadata and user selections to begin generating the music video.
[0711] Input: Parsed lyrics, beat information, user selection information and images.
[0712] Data processing: Input to a generative AI model to generate a video sequence according to the prompt.
[0713] Output: The generated video clip.
[0714] What it does: "The AI model takes lyrics and a beatmap as input and generates a video sequence based on a prompt."
[0715] Step 5: Rendering
[0716] The server then combines the generated video clips using a video editing tool such as FFmpeg to render a high-quality music video.
[0717] Input: The generated video clip.
[0718] Data processing: Video sequences were combined using FFmpeg.
[0719] Output: The final video file (e.g. final_video.mp4).
[0720] Specific operation: "Use FFmpeg to combine the generated video clips into 'final_video.mp4' and render it."
[0721] Step 6: Delivering results
[0722] The server generates and provides a download link for the generated music video to the user.
[0723] Input: The rendered music video file.
[0724] Data processing: Generate download links.
[0725] Output: The download link sent to the user.
[0726] Specific action: "Send the download link for the generated 'final_video.mp4' to the user via email."
[0727] Terminal side processing
[0728] Step 1: Interface display
[0729] The terminal displays an interface for the user to upload an audio file and enter selection information.
[0730] Input: None.
[0731] Data processing: GUI generation.
[0732] Output: The displayed interface.
[0733] Specific behavior: "When a user opens a browser, a button for uploading an audio file appears."
[0734] Step 2: Upload your audio file
[0735] The terminal uploads the audio file selected by the user to the server.
[0736] Input: User selected audio file (e.g. song.mp3).
[0737] Data processing: Send the audio file to the server.
[0738] Output: Audio file uploaded to the server.
[0739] Specific behavior: "When the user selects 'song.mp3' and clicks the upload button, this file will be sent to the server."
[0740] Step 3: Choose the image and mood
[0741] The terminal receives from the user the images of the characters and the selection information regarding the image and atmosphere of the music video, and transmits the information to the server.
[0742] Input: User-selected image file and image information (e.g., vintage style).
[0743] Data processing: Send the selected information to the server.
[0744] Output: Selection information and images sent to the server.
[0745] What it does: "When a user selects the 'Vintage' style and uploads a photo, this information is sent to a server."
[0746] Step 4: Viewing progress
[0747] The terminal displays the progress of the creation process in real time as received from the server.
[0748] Input: Progress data sent from the server.
[0749] Data processing: Display of progress data.
[0750] Output: Real-time updated progress indicator.
[0751] Specific behavior: "Displays the status 'Generating video... 50% complete' on the screen."
[0752] Step 5: View and download the final result
[0753] The device displays a preview of the generated music video and provides a download link.
[0754] Input: The download link sent by the server.
[0755] Data processing: Display links and preview screens.
[0756] Output: Preview and downloadable link.
[0757] What it does: "You'll be presented with a screen that previews the generated video and provides a download link below it."
[0758] User processing
[0759] Step 1: Select an audio file
[0760] The user selects an audio file to upload through the terminal interface.
[0761] Input: None.
[0762] Data processing: Using the file selection dialog.
[0763] Output: The selected audio file.
[0764] Specific behavior: "The user selects 'song.mp3' from local storage in the file selection dialog."
[0765] Step 2: Choose the image and mood
[0766] Users select images of the characters and the image and atmosphere of the music video.
[0767] Input: None.
[0768] Data manipulation: Use of drop-down menus and file upload functions.
[0769] Output: Selected image files and image information.
[0770] What it does: "User drags and drops a profile photo and selects the 'Vintage' style from the drop-down menu."
[0771] Step 3: Upload files and submit information
[0772] The user transmits the selected audio file and image information to the server through the terminal.
[0773] Input: Audio files and image information.
[0774] Data processing: Sending audio files and selected information.
[0775] Output: Audio file and selection information sent to the server.
[0776] Specific action: "Click the upload button to send the audio file and selected information to the server."
[0777] Step 4: Check the generation progress
[0778] The user can check the progress of the creation displayed on the terminal in real time.
[0779] Input: Progress data sent from the server.
[0780] Data processing: Checking progress data.
[0781] Output: Progress display on screen.
[0782] What it does: "A progress bar appears on the screen and updates in real time to show progress, such as 50% complete."
[0783] Step 5: Download and share your finished product
[0784] Users can download the generated music videos to their devices and share them on social media and other platforms.
[0785] Input: Download link.
[0786] Data processing: Use of download links and retrieval of files.
[0787] Output: Downloaded music video files.
[0788] Action: "Click the download link to save 'final_video.mp4' locally, then upload the video to YouTube to share it."
[0789] (Application example 1)
[0790] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0791] Traditional music video production required advanced expertise and expensive equipment, making it difficult for average users to easily create high-quality music videos. Furthermore, there was no way for users to check the progress of automatically generated music videos in real time, nor was there an easy way to directly post generated videos to content distribution services. This reduced user convenience and prevented content from being quickly shared.
[0792] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0793] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for generating a music video using a generative model based on the analyzed lyrics and beat information and the selection information, means for displaying the progress of the generated music video, and means for providing the generated music video to the user and enabling it to be directly posted to a content distribution service. This allows users to create high-quality music videos without specialized knowledge or expensive equipment, to check the progress in real time, and to quickly share the generated video with a content distribution service.
[0794] An "audio file" is a digital file containing music or sound that is uploaded by a user.
[0795] "Lyrics" refers to words and phrases in an audio file, and are text data that make up the sung portion of the music.
[0796] "Beat information" is data about rhythm and tempo obtained from an audio file, and indicates the rhythmic structure of the music.
[0797] "Character images" are photographs or illustrations of people to appear in the music video.
[0798] "Video image and atmosphere" refers to the overall style and visual theme of the music video, and is the visual atmosphere and taste specified by the user.
[0799] "Selection information" refers to information about the image and atmosphere of the characters and video set by the user.
[0800] A "generative model" is an algorithm or framework that uses AI technology to automatically generate music videos.
[0801] A "music video" is video content generated based on an audio file and includes visual storytelling set to music.
[0802] "Content distribution service" means an online service for distributing digital content, such as generated music videos, and includes, for example, a video sharing platform.
[0803] MODE FOR CARRYING OUT THE INVENTION
[0804] A detailed description of an embodiment of the present invention is provided below: The system allows users to automatically generate music videos from audio files, display the progress of the generation, and post the generated music videos directly to a content distribution service.
[0805] Server-side processing
[0806] 1. Receiving audio files: The server receives and stores audio files uploaded by users. The software used is Flask (Python).
[0807] 2. Metadata extraction: The server parses the lyrics and beat information from the received audio files and stores them in a database. This process uses the librosa and speech_recognition libraries.
[0808] 3. Receiving user selection information: The server receives the character images and music video image / atmosphere selection information sent by the user and stores them in a database using Flask and MySQL.
[0809] 4. Launching the generative model: The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selection information. The generative model uses TensorFlow or PyTorch.
[0810] 5. Rendering: The server combines the generated video clips into a sequence and renders the final music video using OpenCV software.
[0811] 6. Progress display: The server manages the progress of the generation process in real time and notifies the user.
[0812] 7. Result Serving: The server serves the generated music video to the user, generates and sends a download link, and also allows the user to submit the generated video directly to a content distribution service.
[0813] Terminal side processing
[0814] 1. Interface display: The terminal displays an interface for users to upload audio files. This interface is built using React Native.
[0815] 2. Audio file upload: The terminal uploads the audio file selected by the user to the server.
[0816] 3. Image and atmosphere selection: The terminal receives selection information from the user regarding the character images and the image and atmosphere of the music video, and transmits it to the server.
[0817] 4. Progress display: The terminal displays the progress of the generation process received from the server in real time.
[0818] 5. View, download, and share the final result: The device displays a preview of the generated music video, provides a download link, and also allows users to post it directly to content distribution services such as video sharing platforms.
[0819] User processing
[0820] 1. Audio file selection: The user selects an audio file to upload through the interface.
[0821] 2. Image / Mood Selection: Users select the image of the characters and the image / mood of the music video.
[0822] 3. File upload and information transmission: The user transmits the selected audio file and image information to the server through the terminal.
[0823] 4. Checking the progress of generation: The user checks the progress of generation displayed on the terminal.
[0824] 5. Download and share the finished product: Users can download the generated music video and share it on social media or other platforms, or post the generated video directly to a content distribution service.
[0825] Specific examples
[0826] As a concrete example, consider a situation where a user wants to create a music video using a new song by a popular Japanese band that can be shared on social media. Using the application, the user uploads an MP3 file of the new song and selects a particular music video theme (e.g., cyberpunk). An example prompt for this would be:
[0827] "Create a cyberpunk-style video that makes it easy to hear the lyrics. Use the attached image as a photo of the characters and a futuristic city in the background."
[0828] Based on this prompt, the generative AI model automatically generates a music video that reflects the style and content specified by the user.
[0829] In this way, users can create professional music videos in a short time without needing special skills or expensive equipment, and post them to content distribution services.
[0830] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0831] Step 1:
[0832] The server receives audio files from users. The audio files (e.g., MP3 format) uploaded by users are received via HTTP requests and stored in cloud storage such as an AWS S3 bucket. The input of this process is the audio file uploaded by the user, and the output is the path to the file stored in cloud storage.
[0833] Step 2:
[0834] The server analyzes the lyrics and beat information from the received audio file. It uses the librosa and speech_recognition libraries to extract rhythm, tempo (beat information), and lyrics from the audio file. The input for this process is the audio file read from cloud storage, and the output is the metadata (beat information and lyric data) resulting from the analysis.
[0835] Step 3:
[0836] The server receives the character images and music video image / mood selections sent by the user and stores them in a database using JSON data received via an HTTP POST request. The input to this process is the user-selected image files and image / mood selections, and the output is a record of the selections stored in the database.
[0837] Step 4:
[0838] The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selections. TensorFlow or PyTorch is used as the generative model, and a video sequence is generated by inputting prompts into the AI model. The inputs for this process are metadata and the user's selections, and the output is the generated video clip.
[0839] Step 5:
[0840] The server combines the generated video clips into a sequence and renders the final music video. OpenCV is used to combine multiple video clips into a sequence and encode them to generate the final video file. The input to this process is a list of video clips, and the output is the final rendered music video file.
[0841] Step 6:
[0842] The server manages the progress of the generation process in real time and notifies the device. It tracks the progress and notifies the device of progress updates sequentially using WebSocket or a real-time communication protocol. The input of this process is the current status of the generation process, and the output is progress update messages.
[0843] Step 7:
[0844] The server provides the generated music video to the user and generates a download link to send to the device. It also allows the user to post the generated video directly to a content distribution service by generating an AWS S3 download link and providing the endpoint URL to the user. The input of this process is the generated music video file, and the output is the download link and the posting endpoint.
[0845] Step 8:
[0846] The terminal displays an interface for the user to upload an audio file. React Native is used to build the user interface and display buttons and selection items for uploading audio files. The input is the design information for the user interface, and the output is the interface screen displayed to the user.
[0847] Step 9:
[0848] The terminal uploads the audio file selected by the user to the server. The audio file obtained from the file input is sent to the server via an HTTP request. The input of this process is the audio file selected by the user, and the output is the request sent to the server.
[0849] Step 10:
[0850] The device receives the user's selection information about the character images and the image and atmosphere of the music video, and sends it to the server. The device then obtains the selection information and sends it to the server in JSON format. The input of this process is the user-selected image file and image and atmosphere information, and the output is the request sent to the server.
[0851] Step 11:
[0852] The terminal displays the progress of the creation process in real time as it receives it from the server. It receives the progress via WebSocket or real-time communication and displays it in the interface as a progress bar or status messages. The input to this process is the progress update messages from the server, and the output is a progress screen that is displayed to the user.
[0853] Step 12:
[0854] The device displays a preview of the generated music video, provides a download link, and also provides the ability to directly post the video to content distribution services such as video sharing platforms. The input to this process is a download link from the server and the video file, and the output is a preview screen and a download link displayed to the user.
[0855] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0856] The system of the present invention not only includes a series of processes for users to automatically generate music videos (MVs) based on audio files, but also a function to recognize the user's emotions and customize the music video based on those emotions. Specific system operations and program processing are explained in natural language below.
[0857] Server-side processing
[0858] 1. Receiving audio files
[0859] The server receives audio files uploaded by users. For example, "Receive and store the file 'song.mp3' uploaded by the user."
[0860] 2. Metadata Extraction
[0861] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio file. The analysis results are stored in a database. For example, "Store the analyzed lyrics and beat information."
[0862] 3. Receiving User Selection Information
[0863] The server receives the user-submitted character photos and selection information about the image and atmosphere of the music video. For example, "receives user-selected 'vintage' style setting information."
[0864] 4. Activating the Emotional Engine
[0865] The server launches an emotion engine to analyze the user's voice file and real-time emotion data. The emotion engine identifies the user's emotional state and applies corresponding parameters to the generative model. For example, "Detect the emotion 'happiness' from the voice file."
[0866] 5. Launching the Generative Model
[0867] The server then activates the AI generative model based on the parsed metadata, user selection information, and emotional data obtained from the emotion engine, and begins generating the music video, for example, "applying 'bright color tones' to the scene based on the emotional state."
[0868] 6. Rendering
[0869] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video. This rendering process ensures smooth transitions between frames. For example, "Rendered for saving as a final video file."
[0870] 7. Providing results
[0871] The server provides the generated music video to the user and generates and sends a download link. For example, "Provide the user with a download link for the completed music video."
[0872] Terminal side processing
[0873] 1. Interface display
[0874] The terminal displays an interface for the user to upload an audio file. The user selects an audio file and clicks the upload button. For example, "Display a web page with an audio file upload button."
[0875] 2. Upload audio file
[0876] The terminal uploads the audio file selected by the user to the server. The selected audio file is sent to the server via an HTTP request. For example, "Send audio file to server."
[0877] 3. Select the image and atmosphere
[0878] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos and images and clicks the submit button. For example, "Display a UI showing photos of the characters and image options."
[0879] 4. Progress Display
[0880] The device will display real-time progress of the generation process as it receives it from the server, e.g., "Displays status 'Generating video... 50% complete'."
[0881] 5. View and download the final result
[0882] The device displays a preview of the generated music video and provides a download link. The user can check and download the generated music video. For example, "Display a preview screen of the generated video and a download link."
[0883] User processing
[0884] 1. Select an audio file
[0885] The user selects an audio file to upload through the interface, for example, "Select 'song.mp3' from local storage."
[0886] 2. Select the image and atmosphere
[0887] Users can choose the image and atmosphere of the character photos and music videos. For example, "Choose a profile picture and a 'vintage' style."
[0888] 3. File upload and information transmission
[0889] The user sends the selected audio file and image information to the server through the terminal. For example, "Click the upload button to send the audio file and selected information."
[0890] 4. Check the generation progress
[0891] The user checks the generation progress displayed on the device, for example, "Check that the progress bar reaches 50%."
[0892] 5. Download and share your finished product
[0893] Users can download the generated music video and share it on social media or other platforms. For example, "Click the download link to save the video, then upload it to YouTube to share."
[0894] The above is a specific example of how to implement a system incorporating the emotion engine of the present invention. This allows artists and content creators to easily create high-quality music videos customized to their own emotional state. Utilizing emotion data from the emotion engine enables more personalized visual expression, making a strong impact on viewers.
[0895] The processing flow will be explained below.
[0896] Step 1:
[0897] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[0898] Step 2:
[0899] The device sends the audio file selected by the user to the server via an HTTP request, specifically by using an input form or drag-and-drop function.
[0900] Step 3:
[0901] The server receives the audio file sent from the device and saves it in a specified directory, for example, a file named "song.mp3."
[0902] Step 4:
[0903] The server analyzes the received audio files and runs audio analysis algorithms to extract lyrics and beat information, and stores the analyzed metadata in a database.
[0904] Step 5:
[0905] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[0906] Step 6:
[0907] The device sends the photos of the characters and the image and atmosphere selection information received from the user to the server. Specifically, it sends the image files and text data via an HTTP request.
[0908] Step 7:
[0909] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model, which includes preprocessing the image data and analyzing the text data.
[0910] Step 8:
[0911] The server activates an emotion engine that analyzes the user's voice file to identify their emotional state, for example using a voice analysis algorithm to identify emotions such as "happiness" or "sadness."
[0912] Step 9:
[0913] The server adjusts the parameters of the generative model based on the identified emotional state, specifically dynamically changing the color grading and effects of the video based on the emotional data.
[0914] Step 10:
[0915] The server then activates an AI generation model based on the analyzed metadata, user selection information, and emotional data to begin generating the music video. The AI model processes the input data and generates a corresponding video clip.
[0916] Step 11:
[0917] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[0918] Step 12:
[0919] The server saves the completed music video as a file and generates a download link that can be accessed by the user.
[0920] Step 13:
[0921] The server sends the generated download link to the terminal, allowing the user to access it.
[0922] Step 14:
[0923] The terminal displays a preview of the generated music video to the user and provides a download link, allowing the user to view and download the generated music video.
[0924] Step 15:
[0925] Users can download the completed music video and share it on social media or other platforms, such as uploading it to YouTube to share with their audience.
[0926] Example 2
[0927] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0928] Conventional music video generation systems have struggled to generate personalized videos that take into account the user's emotional state. Furthermore, extracting metadata from audio files and generating videos based on user-selected images and moods are time-consuming, limiting the ability to improve the user experience. The present invention aims to solve these problems.
[0929] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0930] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and mood of a video of a visual work, means for analyzing the audio file and real-time emotional data to identify the user's emotional state, means for generating a visual work using a generative model based on the analyzed lyrics and beat information, the selection information, and the identified emotional state, and means for providing the generated visual work to the user, thereby enabling the automatic generation of personalized, high-quality music videos that take the user's emotions into consideration.
[0931] "Audio File" refers to audio data stored in digital format.
[0932] "Lyrics" refers to the words or sentences in a song within an audio file.
[0933] "Beat information" refers to data about rhythm and tempo contained in an audio file.
[0934] "Analysis" refers to the process of extracting useful information from input data.
[0935] "Character Images" refers to photographs or illustrations of people appearing in the music video.
[0936] "Visual work" refers to content that is visually displayed, such as a film or video clip.
[0937] "Image / Mood" refers to the style or theme expressed in a visual work.
[0938] "Emotional state" refers to the psychological state of the user analyzed from the voice data.
[0939] "Real-time emotional data" refers to data that indicates the current emotional state of the user.
[0940] "Generative model" refers to an algorithm that uses AI techniques to automatically generate visual works or other content.
[0941] "Generation" refers to the process of creating new data or content based on specific data.
[0942] "Rendering" refers to the process of converting digital data into a visually displayable form.
[0943] "Serving" refers to the process of transmitting or displaying the generated visual work in a form accessible to a user.
[0944] MODE FOR CARRYING OUT THE INVENTION
[0945] The system of the present invention is a system that includes a series of processes for a user to automatically generate a music video (MV) based on an audio file, recognize the user's emotions, and customize the MV based on those emotions. Specific operations of the system and program processing are described in detail below.
[0946] Server-side processing
[0947] 1. Receiving audio files
[0948] The server receives the audio file uploaded by the user. In this receiving process, the HTTP protocol is used to send the file "song.mp3" to the server and save it in a specified folder.
[0949] 2. Metadata Extraction
[0950] The server uses audio analysis algorithms (e.g., Python libraries librosa or PyDub) to extract lyrics and beat information from audio files. For example, it loads audio using the librosa.load() function and extracts beat information using the librosa.beat.beat_track() function. The analysis results are stored in a database.
[0951] 3. Receiving User Selection Information
[0952] The server receives and stores the character images and the selection information about the image and atmosphere of the visual work sent by the user via HTTP request, for example, "vintage style" or "character: profile.jpg" in the database.
[0953] 4. Activating the Emotional Engine
[0954] The server invokes an emotion engine (e.g., a general emotion analysis service) to analyze the audio file and real-time emotion data. The engine identifies the user's emotional state (e.g., happiness, sadness, etc.) and provides corresponding parameters to the generative model.
[0955] 5. Launching the Generative Model
[0956] The server generates a music video by activating an AI generation model based on the analyzed lyrics, beat information, user selection information, and emotional state. For example, the server inputs prompts to customize scene settings and color tones according to the user's emotional state into the generation AI model.
[0957] 6. Rendering
[0958] The server then combines the generated video clips into a sequence and renders the final video. This process uses a video processing tool such as FFmpeg. For example, to concatenate multiple clips, a command like ffmpeg -i input1.mp4 -i input2.mp4 -filter_complex "[0:v][1:v] concat=n=2:v=1[outv]" -map "[outv]" output.mp4 is used.
[0959] 7. Providing results
[0960] The server generates a download link for the completed music video and provides it to the user, returning a response JSON object containing the link URL.
[0961] Terminal side processing
[0962] 1. Interface display
[0963] The terminal displays an interface for the user to upload an audio file. The interface includes a file selection button and an upload button. For example, the terminal uses an HTML form or JavaScript to prompt the user to select an audio file.
[0964] 2. Upload audio file
[0965] The device uploads the audio file selected by the user to the server via an HTTP request, specifically, an AJAX request to send the file "song.mp3" to the server.
[0966] 3. Select the image and atmosphere
[0967] The device displays an interface that allows users to select images of the characters and the image and atmosphere of the music video. For example, a selection screen built with HTML5 and CSS is displayed, allowing users to select "profile.jpg" or a "vintage" style.
[0968] 4. Progress Display
[0969] The terminal displays the progress of the generation process received from the server in real time. It uses the JavaScript setInterval function to query the server for the progress at regular intervals and displays the results.
[0970] 5. View and download the final result
[0971] The device will display a preview of the generated music video and provide a download link. <video> View by tag and download link< / video> < / url:> Provided by tag.
[0972] User processing
[0973] 1. Select an audio file
[0974] The user selects the audio file to upload from the interface, specifically the file "song.mp3" from local storage.
[0975] 2. Select the image and atmosphere
[0976] The user selects the image and atmosphere of the characters and music video. They select "profile.jpg" from local storage and choose "vintage style" from the options provided in the UI.
[0977] 3. File upload and information transmission
[0978] The user sends the selected audio file and image information to the server through the terminal, and clicks the upload button to send the audio file and the selected information.
[0979] 4. Check the generation progress
[0980] The user checks the generation progress displayed on the device, specifically checking the progress message such as "Make sure the progress bar reaches 50%."
[0981] 5. Download and share your finished product
[0982] Users can download the generated music video and share it on social media or other platforms by clicking the download link to save the 'output.mp4' video, then uploading the video to YouTube or social media to share it.
[0983] Examples of concrete examples and prompts
[0984] Prompt Sentence Examples
[0985] Audio file: song.mp3
[0986] Character photo:profile.jpg
[0987] Music video vibe: Vintage
[0988] These procedures and system configurations allow users to easily generate personalized, high-quality music videos based on audio files and emotions.
[0989] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0990] Step 1:
[0991] The server receives an audio file from the user. The input is the audio file selected by the user (e.g., 'song.mp3'), and the output is the audio file stored on the server. Specific operations include requesting the audio file using the HTTP protocol and saving it in a specific folder on the server.
[0992] Step 2:
[0993] The server analyzes the lyrics and beat information from the received audio file. The input is the audio file saved in step 1, and the output is the lyrics and beat information. Specifically, it uses a Python library (e.g., librosa or PyDub) to load the audio with librosa.load() and extract the beat with librosa.beat.beat_track(). The analysis results are stored in a database.
[0994] Step 3:
[0995] The server receives from the user the image of the characters and the selection information regarding the image and atmosphere of the visual work video. The input is the image file and text information selected by the user, and the output is the selection information saved on the server. Specifically, the data sent via the HTTP request is saved in a database. For example, information such as "vintage style" or "character: profile.jpg" is recorded.
[0996] Step 4:
[0997] The server starts an emotion engine to analyze the audio file and real-time emotion data to identify the user's emotional state. The input is the lyrics and beat information analyzed in step 2 and the audio file, and the output is the identified emotional state. Specifically, the emotion analysis service detects the user's emotion (e.g., happiness) from the audio.
[0998] Step 5:
[0999] The server activates a generative AI model based on the analyzed metadata, user selection information, and emotional data to generate a visual work. The input is the data obtained in Steps 2 and 3 and the emotional state identified in Step 4, and the output is the generated visual work (video clip). Specifically, the server provides the generative AI model with prompts to customize the scene setting and color tone according to the emotional state, and generates a music video.
[1000] Step 6:
[1001] The server combines the generated video clips into a sequence and renders it as a final video. The input is the video clip generated in step 5, and the output is a high-quality final video file. Specifically, it uses a video processing tool such as FFmpeg to concatenate multiple clips and executes the ffmpeg command to save it as "output.mp4."
[1002] Step 7:
[1003] The server provides the completed music video to the user. The input is the final video file generated in step 6, and the output is a download link URL. The server then returns a response JSON object containing the URL of the generated video file to the user.
[1004] Step 8:
[1005] The terminal provides the user with an upload interface, a selection interface, a progress indicator, a final result indicator, and a download link, allowing the user to upload audio files, submit selection information, check the progress, and download the generated video.
[1006] (Application example 2)
[1007] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1008] Currently, many artists and content creators spend a great deal of time and effort creating music videos that reflect the emotions and themes of their work. Furthermore, many automated music video generation systems are not customized based on the user's emotional state, and often lack visual appeal. The present invention aims to solve this problem by providing a system that automatically generates high-quality music videos that reflect the user's emotional state.
[1009] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for analyzing the user's emotional state, means for generating video using a generative model based on the analyzed lyrics and beat information, the selection information, and the emotional state, and means for providing the generated video to the user. This enables the generation of a personalized music video based on the user's emotional state.
[1010] An "audio file" is a file in which audio data is stored in digital format.
[1011] "Lyrics" are words or text that correspond to a musical piece in an audio file.
[1012] "Beat information" refers to the rhythm and tempo information within an audio file.
[1013] "Character images" are image data of people appearing in the music video that are uploaded or selected by the user.
[1014] "Image and atmosphere of the video" refers to the visual elements and atmosphere that express the style and theme of the music video.
[1015] "Emotional state" is data that indicates the user's current mental state and emotions.
[1016] "Generative model" refers to an algorithm or AI model that automatically generates music videos based on input data.
[1017] "Video clips" are short video segments that make up a music video.
[1018] "Rendering as a sequence" refers to the process of combining a series of video clips into a continuous video and rendering it at high quality.
[1019] The system of the present invention provides a set of processes for users to automatically generate music videos (MVs) based on audio files, and further includes the ability to recognize the user's emotions and customize the MV based on those emotions.
[1020] The server first receives an audio file from the user. This audio file is audio data stored in a digital format, such as mp3 or wav. The server then uses an audio analysis algorithm to extract lyrics and beat information from the audio file. The extracted metadata is then stored in a database.
[1021] Receives from the user character images and video image and mood selection information that reflects the user's designated visual style or theme.
[1022] The emotion engine is activated to analyze the user's emotional state. The emotion engine identifies emotions from the text data and voice data entered by the user and obtains the emotional state as data.
[1023] The server uses a generative AI model to generate video based on the analyzed lyrics and beat information, user selections, and emotional state. The generative AI model receives these input data as prompts and generates a series of video clips based on them. These video clips are rendered as a sequence and combined into a high-quality video.
[1024] Finally, the generated video is provided to the user: the server saves the video file and generates a link for the user to download it.
[1025] The hardware and software used are as follows: The server uses the Django framework (Python) and an appropriate library (e.g., librosa) for speech analysis; an NLP library (e.g., NLTK, spaCy) or a speech analysis library is used for the emotion engine; and a machine learning framework (e.g., TensorFlow, PyTorch) is used for the generative AI model.
[1026] For example, if a user uploads an up-tempo song called "Happy.mp3," the system extracts lyrics and beat information from the audio file and detects the user's emotional state of happiness. If the user selects a vintage theme, the system applies visual effects that match the theme and generates a light-hearted video that expresses a sense of happiness. The generated music video is immediately available for preview and distribution.
[1027] An example of a prompt is:
[1028] "User ID: 001
[1029] Audio file URL: / media / uploads / song.mp3
[1030] Extracted Metadata: {Lyrics: 'Happy song lyrics', Beat: 'Uptempo'}
[1031] User Sentiment: Happiness
[1032] Selected theme: Vintage
[1033] The above is a specific embodiment of the system of the present invention, which allows artists and content creators to easily create high-quality music videos that are personalized to their emotional state.
[1034] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1035] Step 1:
[1036] The user selects and uploads an audio file.
[1037] Input: An audio file selected by the user from local storage (e.g. song.mp3).
[1038] Specific operation: The terminal creates and sends an HTTP request to send the audio file selected by the user through the interface to the server.
[1039] Output: The audio file is sent to the server.
[1040] Step 2:
[1041] The server receives and stores the audio files.
[1042] Input: User submitted audio file (e.g. song.mp3).
[1043] Specific operation: The server stores the received audio file in local storage or cloud storage.
[1044] Output: The path to the audio file in local or cloud storage.
[1045] Step 3:
[1046] The server analyzes the audio file and extracts the metadata.
[1047] Input: The path to the saved audio file.
[1048] What it does: The server uses audio analysis algorithms to extract lyrics and beat information from the audio file (e.g., using the librosa library).
[1049] Data processing or data manipulation: Spectral analysis of audio data to extract beat maps and lyrics in text format.
[1050] Output: Parsed lyrics and beat information metadata.
[1051] Step 4:
[1052] The server receives from the user the images of the characters and selection information regarding the image and atmosphere of the video.
[1053] Input: User-submitted image file and image selection information.
[1054] Specific operation: The server reads the image file and selection information received as an HTTP request and saves them in a database or storage.
[1055] Output: Path to the saved image file and selection information.
[1056] Step 5:
[1057] The server analyzes the user's emotional state.
[1058] Input: User text or voice data.
[1059] Specific operation: The server launches an emotion engine and analyzes the user's input data to identify their emotional state (e.g., using NLTK or spaCy).
[1060] Data processing or data calculation: Analyzing text or audio data with natural language processing algorithms to generate emotion labels (e.g., happy, sad).
[1061] Output: The user's emotional state.
[1062] Step 6:
[1063] The server generates the video using a generative AI model.
[1064] Input: Parsed lyrics and beat information metadata, user image files and selections, user emotional state.
[1065] Specific operation: The server inputs these input data as prompt sentences into a generative AI model to generate video clips (e.g., using TensorFlow or PyTorch).
[1066] Data processing or data computation: A generative AI model generates video clips based on prompts and renders them as a sequence.
[1067] Output: The generated video clip.
[1068] Step 7:
[1069] The server provides the generated video to the user.
[1070] Input: The generated video clip.
[1071] Specific operation: The server saves the video clip, generates a URL from which the user can download it, and creates and sends an HTTP response that returns the URL to the user.
[1072] Output: A URL to the video clip that users can download.
[1073] These are the processing steps of the system for realizing this application example. Through the specific operations performed at each step, users can create and download music videos that reflect their emotional state.
[1074] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1075] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1076] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1077] [Third embodiment]
[1078] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1079] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1080] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1081] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1082] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1083] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1084] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1085] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1086] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1087] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1088] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1089] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1090] The system of the present invention realizes a series of processes that allow users to automatically generate music videos (MVs) based on audio files. The specific operation of the system and the processing of the program are explained below in natural language.
[1091] Server-side processing
[1092] 1. Receiving audio files
[1093] The server receives audio files uploaded by users.
[1094] Example: "Store the file "song.mp3" uploaded by the user on the server."
[1095] 2. Metadata Extraction
[1096] The server extracts metadata such as lyrics and beat information from the received audio file and stores it in a database.
[1097] Example: "Analyze lyrics and beatmaps using audio analysis algorithms."
[1098] 3. Receiving User Selection Information
[1099] The server receives user-submitted photos of the characters and selection information regarding the image and atmosphere of the music video.
[1100] Example: "Receive the settings for the user-selected 'Vintage' style."
[1101] 4. Launching the Generative Model
[1102] Based on the parsed metadata and user selection information, the server activates the AI generative model and begins generating the music video.
[1103] For example: "The AI model generates a video sequence that corresponds to the lyrics and applies the selected mood."
[1104] 5. Rendering
[1105] The server combines the generated video clips into a sequence to render a high-quality music video.
[1106] For example: "Combine the sequences and render to produce the final video file."
[1107] 6. Providing results
[1108] The server provides the generated music video to the user and generates and sends a download link.
[1109] For example: "Provide users with a link to the finished video file so they can download it."
[1110] Terminal side processing
[1111] 1. Interface display
[1112] The terminal displays an interface for the user to upload an audio file.
[1113] Example: "Display a UI (user interface) with a button to upload an audio file."
[1114] 2. Upload audio file
[1115] The terminal uploads the audio file selected by the user to the server.
[1116] Example: "Send the selected file 'song.mp3' to the server."
[1117] 3. Select the image and atmosphere
[1118] The terminal receives photos of characters and selection information about the image and atmosphere to be created from the user and transmits them to the server.
[1119] Example: "Send user-selected image information to the server."
[1120] 4. Progress Display
[1121] The terminal displays the progress of the creation process in real time as received from the server.
[1122] Example: "Show status 'Generating video... 50% complete'."
[1123] 5. View and download the final result
[1124] The device displays a preview of the generated music video and provides a download link.
[1125] Example: "Display a screen to preview the generated video and provide a download link."
[1126] User processing
[1127] 1. Select an audio file
[1128] The user selects an audio file to upload through the interface.
[1129] For example: "Select 'song.mp3' from local storage."
[1130] 2. Select the image and atmosphere
[1131] Users select photos of the characters and the image and atmosphere of the music video.
[1132] For example: "Choose a profile picture and a 'vintage' style."
[1133] 3. File upload and information transmission
[1134] The user transmits the selected audio file and image information to the server through the terminal.
[1135] For example: "Click the upload button to submit your audio file and selection information."
[1136] 4. Check the generation progress
[1137] The user checks the progress of the creation displayed on the terminal.
[1138] For example: "Check that the progress bar reaches 50%."
[1139] 5. Download and share your finished product
[1140] Users can download the generated music video and share it on social media and other platforms.
[1141] For example: "Click the download link to save the video and upload it to YouTube to share."
[1142] The above is a specific embodiment of the present invention, which enables artists and content creators to produce high-quality music videos in a short amount of time without requiring special skills or expensive equipment.
[1143] The processing flow will be explained below.
[1144] Step 1:
[1145] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[1146] Step 2:
[1147] The terminal uploads the audio file selected by the user to the server, specifically, by transmitting the selected audio file to the server via an HTTP request.
[1148] Step 3:
[1149] The server receives the audio file sent from the terminal and saves it in a specified directory.
[1150] Step 4:
[1151] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio files, and the results of this analysis are stored in a database.
[1152] Step 5:
[1153] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[1154] Step 6:
[1155] The terminal transmits the character photos and image / atmosphere selection information received from the user to the server. The selection information includes image files and text data.
[1156] Step 7:
[1157] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model.
[1158] Step 8:
[1159] The server then activates the AI generation model based on the analyzed lyrics and beat information, as well as the user's selection information, and begins generating the music video. Specifically, the AI model processes the input data and generates a corresponding video clip.
[1160] Step 9:
[1161] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[1162] Step 10:
[1163] The server saves the completed music video as a file and generates a download link.
[1164] Step 11:
[1165] The server sends the generated download link to the terminal, allowing the user to access it.
[1166] Step 12:
[1167] The device displays a preview of the generated music video to the user and provides a download link, allowing the user to view, download, and save the generated music video.
[1168] Step 13:
[1169] Users can download the finished music video and share it on social media and other platforms, at which point they can also view the video and provide feedback.
[1170] Example 1
[1171] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1172] Traditional music video production requires expensive equipment and specialized skills, making it difficult for many artists and content creators. Manual editing is also time-consuming and inefficient. In response, there was a need for a method to automatically generate high-quality music videos from audio files.
[1173] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1174] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and atmosphere of the music video, means for generating a music video using a generative AI model based on the analyzed lyrics and beat information and the selection information, means for providing the generated music video to the user and generating and transmitting a download link, and means for displaying the progress of the music video generation to the user in real time, thereby enabling users to efficiently generate high-quality music videos without requiring special skills or expensive equipment.
[1175] "Audio files" are music or audio data files uploaded by users.
[1176] "Lyrics" refers to the words or sentences within an audio file, which may contain poetic expressions or narrative.
[1177] "Beat information" is data that indicates the rhythm and tempo within an audio file.
[1178] "Character Images" are photographs or illustrations of people that appear in the music video.
[1179] "Music video atmosphere" refers to the overall style or theme of the video selected by the user, such as "vintage" or "futuristic."
[1180] A "generative AI model" is an artificial intelligence algorithm that generates video sequences or video clips based on audio file data and user selections.
[1181] "Rendering" is the process of combining the generated video clips into a sequence to create the final video file.
[1182] "Download link" refers to a URL that allows a user to download the generated music video.
[1183] The "means for displaying the generation progress status in real time" is a function for displaying the progress of the server-side processing on the user terminal.
[1184] The present invention provides a system that allows users to automatically generate music videos (MVs) based on audio files. The specific operation and processing of the system are described below. The system mainly operates in cooperation with three parties: a server, a terminal, and a user.
[1185] Hardware and software used
[1186] Hardware:
[1187] Server: Cloud server or on-premise server.
[1188] Device: A PC, smartphone, tablet, or other device capable of operating a user interface.
[1189] software:
[1190] Speech analysis library: Librosa.
[1191] Video generation algorithms: Generative AI models (e.g., OpenAI's GPT-3, DALL-E).
[1192] Video editing tool: FFmpeg.
[1193] Web frameworks: Django, Flask, etc.
[1194] Example of system operation
[1195] Receive audio files:
[1196] The user selects an audio file to upload through the device interface. Example: "The user selects 'song.mp3' from local storage."
[1197] The device uploads this audio file to the server. Example: "Send the selected 'song.mp3' file to the server."
[1198] Metadata Extraction:
[1199] The server receives the audio file and parses it for lyrics and beat information using an audio analysis library such as Librosa. Example: "Using Librosa, parse 'uploads / song.mp3' for lyrics and beatmap data and store them in a database."
[1200] Receive user selection information:
[1201] Through the interface, users select the style and character images for the music video. For example, "User selects profile picture and 'vintage' style."
[1202] The terminal transmits this information to the server.
[1203] Launch the generative model:
[1204] The server then triggers a generative AI model based on the parsed metadata and user selections to generate a music video. For example, "The AI model takes lyrics and beatmaps as input and generates a video sequence."
[1205] rendering:
[1206] The server combines the generated video clips using a video editing tool such as FFmpeg and renders a high-quality music video. Example: "Combine the generated video clips using FFmpeg as 'final_video.mp4' and render."
[1207] Results provided:
[1208] The server generates a download link for the generated music video and provides it to the user. Example: "Email the user a download link for the generated 'final_video.mp4'."
[1209] The device has the ability to display the generation progress to the user in real time. For example, "Display the status 'Generating video... 50% complete' on the screen."
[1210] Examples of prompt statements
[1211] For example, the following prompt sentence is input to a generative AI model:
[1212] "Generate a vintage-style video sequence based on the lyrics of this audio file."
[1213] The video clips thus generated can be combined into a sequence to obtain the final video file.
[1214] This invention enables users to efficiently generate high-quality music videos without requiring special skills or expensive equipment, which significantly reduces time and costs compared to conventional methods.
[1215] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1216] Server-side processing
[1217] Step 1: Receive the audio file
[1218] The server receives the audio file uploaded by the user and stores it in a specified directory.
[1219] Input: An audio file (e.g. song.mp3) that the user uploads from their device.
[1220] Data processing: Send the audio file to the server as an HTTP request.
[1221] Output: An audio file saved in the specified directory on the server (e.g. uploads / song.mp3).
[1222] Specific behavior: "The server receives the HTTP request and saves the audio file to 'uploads / song.mp3'."
[1223] Step 2: Metadata extraction
[1224] The server extracts lyrics and beat information from the received audio file using an audio analysis library such as Librosa.
[1225] Input: A saved audio file (e.g. uploads / song.mp3).
[1226] Data processing: Analysis of audio files using Librosa.
[1227] Output: Extracted lyrics and beat information (stored in a database).
[1228] Specific operation: "Using Librosa, parse the lyrics and beatmap from the audio file 'uploads / song.mp3' and store them in a database."
[1229] Step 3: Receiving User Selection Information
[1230] The server receives user-submitted images of the characters and selection information regarding the image and atmosphere of the music video.
[1231] Input: Image file and selection information (e.g., vintage style) sent by the user from the device.
[1232] Data processing: The selected information is saved in the server database.
[1233] Output: User selection information and images stored on the server.
[1234] Specific operation: "The image information selected by the user and the character's image are saved in the 'user_preferences' directory on the server."
[1235] Step 4: Launching the generative model
[1236] The server then activates the generative AI model based on the parsed metadata and user selections to begin generating the music video.
[1237] Input: Parsed lyrics, beat information, user selection information and images.
[1238] Data processing: Input to a generative AI model to generate a video sequence according to the prompt.
[1239] Output: The generated video clip.
[1240] What it does: "The AI model takes lyrics and a beatmap as input and generates a video sequence based on a prompt."
[1241] Step 5: Rendering
[1242] The server then combines the generated video clips using a video editing tool such as FFmpeg to render a high-quality music video.
[1243] Input: The generated video clip.
[1244] Data processing: Video sequences were combined using FFmpeg.
[1245] Output: The final video file (e.g. final_video.mp4).
[1246] Specific operation: "Use FFmpeg to combine the generated video clips into 'final_video.mp4' and render it."
[1247] Step 6: Delivering results
[1248] The server generates and provides a download link for the generated music video to the user.
[1249] Input: The rendered music video file.
[1250] Data processing: Generate download links.
[1251] Output: The download link sent to the user.
[1252] Specific action: "Send the download link for the generated 'final_video.mp4' to the user via email."
[1253] Terminal side processing
[1254] Step 1: Interface display
[1255] The terminal displays an interface for the user to upload an audio file and enter selection information.
[1256] Input: None.
[1257] Data processing: GUI generation.
[1258] Output: The displayed interface.
[1259] Specific behavior: "When a user opens a browser, a button for uploading an audio file appears."
[1260] Step 2: Upload your audio file
[1261] The terminal uploads the audio file selected by the user to the server.
[1262] Input: User selected audio file (e.g. song.mp3).
[1263] Data processing: Send the audio file to the server.
[1264] Output: Audio file uploaded to the server.
[1265] Specific behavior: "When the user selects 'song.mp3' and clicks the upload button, this file will be sent to the server."
[1266] Step 3: Choose the image and mood
[1267] The terminal receives from the user the images of the characters and the selection information regarding the image and atmosphere of the music video, and transmits the information to the server.
[1268] Input: User-selected image file and image information (e.g., vintage style).
[1269] Data processing: Send the selected information to the server.
[1270] Output: Selection information and images sent to the server.
[1271] What it does: "When a user selects the 'Vintage' style and uploads a photo, this information is sent to a server."
[1272] Step 4: Viewing progress
[1273] The terminal displays the progress of the creation process in real time as received from the server.
[1274] Input: Progress data sent from the server.
[1275] Data processing: Display of progress data.
[1276] Output: Real-time updated progress indicator.
[1277] Specific behavior: "Displays the status 'Generating video... 50% complete' on the screen."
[1278] Step 5: View and download the final result
[1279] The device displays a preview of the generated music video and provides a download link.
[1280] Input: The download link sent by the server.
[1281] Data processing: Display links and preview screens.
[1282] Output: Preview and downloadable link.
[1283] What it does: "You'll be presented with a screen that previews the generated video and provides a download link below it."
[1284] User processing
[1285] Step 1: Select an audio file
[1286] The user selects an audio file to upload through the terminal interface.
[1287] Input: None.
[1288] Data processing: Using the file selection dialog.
[1289] Output: The selected audio file.
[1290] Specific behavior: "The user selects 'song.mp3' from local storage in the file selection dialog."
[1291] Step 2: Choose the image and mood
[1292] Users select images of the characters and the image and atmosphere of the music video.
[1293] Input: None.
[1294] Data manipulation: Use of drop-down menus and file upload functions.
[1295] Output: Selected image files and image information.
[1296] What it does: "User drags and drops a profile photo and selects the 'Vintage' style from the drop-down menu."
[1297] Step 3: Upload files and submit information
[1298] The user transmits the selected audio file and image information to the server through the terminal.
[1299] Input: Audio files and image information.
[1300] Data processing: Sending audio files and selected information.
[1301] Output: Audio file and selection information sent to the server.
[1302] Specific action: "Click the upload button to send the audio file and selected information to the server."
[1303] Step 4: Check the generation progress
[1304] The user can check the progress of the creation displayed on the terminal in real time.
[1305] Input: Progress data sent from the server.
[1306] Data processing: Checking progress data.
[1307] Output: Progress display on screen.
[1308] What it does: "A progress bar appears on the screen and updates in real time to show progress, such as 50% complete."
[1309] Step 5: Download and share your finished product
[1310] Users can download the generated music videos to their devices and share them on social media and other platforms.
[1311] Input: Download link.
[1312] Data processing: Use of download links and retrieval of files.
[1313] Output: Downloaded music video files.
[1314] Action: "Click the download link to save 'final_video.mp4' locally, then upload the video to YouTube to share it."
[1315] (Application example 1)
[1316] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1317] Traditional music video production required advanced expertise and expensive equipment, making it difficult for average users to easily create high-quality music videos. Furthermore, there was no way for users to check the progress of automatically generated music videos in real time, nor was there an easy way to directly post generated videos to content distribution services. This reduced user convenience and prevented content from being quickly shared.
[1318] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1319] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for generating a music video using a generative model based on the analyzed lyrics and beat information and the selection information, means for displaying the progress of the generated music video, and means for providing the generated music video to the user and enabling it to be directly posted to a content distribution service. This allows users to create high-quality music videos without specialized knowledge or expensive equipment, to check the progress in real time, and to quickly share the generated video with a content distribution service.
[1320] An "audio file" is a digital file containing music or sound that is uploaded by a user.
[1321] "Lyrics" refers to words and phrases in an audio file, and are text data that make up the sung portion of the music.
[1322] "Beat information" is data about rhythm and tempo obtained from an audio file, and indicates the rhythmic structure of the music.
[1323] "Character images" are photographs or illustrations of people to appear in the music video.
[1324] "Video image and atmosphere" refers to the overall style and visual theme of the music video, and is the visual atmosphere and taste specified by the user.
[1325] "Selection information" refers to information about the image and atmosphere of the characters and video set by the user.
[1326] A "generative model" is an algorithm or framework that uses AI technology to automatically generate music videos.
[1327] A "music video" is video content generated based on an audio file and includes visual storytelling set to music.
[1328] "Content distribution service" means an online service for distributing digital content, such as generated music videos, and includes, for example, a video sharing platform.
[1329] MODE FOR CARRYING OUT THE INVENTION
[1330] A detailed description of an embodiment of the present invention is provided below: The system allows users to automatically generate music videos from audio files, display the progress of the generation, and post the generated music videos directly to a content distribution service.
[1331] Server-side processing
[1332] 1. Receiving audio files: The server receives and stores audio files uploaded by users. The software used is Flask (Python).
[1333] 2. Metadata extraction: The server parses the lyrics and beat information from the received audio files and stores them in a database. This process uses the librosa and speech_recognition libraries.
[1334] 3. Receiving user selection information: The server receives the character images and music video image / atmosphere selection information sent by the user and stores them in a database using Flask and MySQL.
[1335] 4. Launching the generative model: The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selection information. The generative model uses TensorFlow or PyTorch.
[1336] 5. Rendering: The server combines the generated video clips into a sequence and renders the final music video using OpenCV software.
[1337] 6. Progress display: The server manages the progress of the generation process in real time and notifies the user.
[1338] 7. Result Serving: The server serves the generated music video to the user, generates and sends a download link, and also allows the user to submit the generated video directly to a content distribution service.
[1339] Terminal side processing
[1340] 1. Interface display: The terminal displays an interface for users to upload audio files. This interface is built using React Native.
[1341] 2. Audio file upload: The terminal uploads the audio file selected by the user to the server.
[1342] 3. Image and atmosphere selection: The terminal receives selection information from the user regarding the character images and the image and atmosphere of the music video, and transmits it to the server.
[1343] 4. Progress display: The terminal displays the progress of the generation process received from the server in real time.
[1344] 5. View, download, and share the final result: The device displays a preview of the generated music video, provides a download link, and also allows users to post it directly to content distribution services such as video sharing platforms.
[1345] User processing
[1346] 1. Audio file selection: The user selects an audio file to upload through the interface.
[1347] 2. Image / Mood Selection: Users select the image of the characters and the image / mood of the music video.
[1348] 3. File upload and information transmission: The user transmits the selected audio file and image information to the server through the terminal.
[1349] 4. Checking the progress of generation: The user checks the progress of generation displayed on the terminal.
[1350] 5. Download and share the finished product: Users can download the generated music video and share it on social media or other platforms, or post the generated video directly to a content distribution service.
[1351] Specific examples
[1352] As a concrete example, consider a situation where a user wants to create a music video using a new song by a popular Japanese band that can be shared on social media. Using the application, the user uploads an MP3 file of the new song and selects a particular music video theme (e.g., cyberpunk). An example prompt for this would be:
[1353] "Create a cyberpunk-style video that makes it easy to hear the lyrics. Use the attached image as a photo of the characters and a futuristic city in the background."
[1354] Based on this prompt, the generative AI model automatically generates a music video that reflects the style and content specified by the user.
[1355] In this way, users can create professional music videos in a short time without needing special skills or expensive equipment, and post them to content distribution services.
[1356] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1357] Step 1:
[1358] The server receives audio files from users. The audio files (e.g., MP3 format) uploaded by users are received via HTTP requests and stored in cloud storage such as an AWS S3 bucket. The input of this process is the audio file uploaded by the user, and the output is the path to the file stored in cloud storage.
[1359] Step 2:
[1360] The server analyzes the lyrics and beat information from the received audio file. It uses the librosa and speech_recognition libraries to extract rhythm, tempo (beat information), and lyrics from the audio file. The input for this process is the audio file read from cloud storage, and the output is the metadata (beat information and lyric data) resulting from the analysis.
[1361] Step 3:
[1362] The server receives the character images and music video image / mood selections sent by the user and stores them in a database using JSON data received via an HTTP POST request. The input to this process is the user-selected image files and image / mood selections, and the output is a record of the selections stored in the database.
[1363] Step 4:
[1364] The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selections. TensorFlow or PyTorch is used as the generative model, and a video sequence is generated by inputting prompts into the AI model. The inputs for this process are metadata and the user's selections, and the output is the generated video clip.
[1365] Step 5:
[1366] The server combines the generated video clips into a sequence and renders the final music video. OpenCV is used to combine multiple video clips into a sequence and encode them to generate the final video file. The input to this process is a list of video clips, and the output is the final rendered music video file.
[1367] Step 6:
[1368] The server manages the progress of the generation process in real time and notifies the device. It tracks the progress and notifies the device of progress updates sequentially using WebSocket or a real-time communication protocol. The input of this process is the current status of the generation process, and the output is progress update messages.
[1369] Step 7:
[1370] The server provides the generated music video to the user and generates a download link to send to the device. It also allows the user to post the generated video directly to a content distribution service by generating an AWS S3 download link and providing the endpoint URL to the user. The input of this process is the generated music video file, and the output is the download link and the posting endpoint.
[1371] Step 8:
[1372] The terminal displays an interface for the user to upload an audio file. React Native is used to build the user interface and display buttons and selection items for uploading audio files. The input is the design information for the user interface, and the output is the interface screen displayed to the user.
[1373] Step 9:
[1374] The terminal uploads the audio file selected by the user to the server. The audio file obtained from the file input is sent to the server via an HTTP request. The input of this process is the audio file selected by the user, and the output is the request sent to the server.
[1375] Step 10:
[1376] The device receives the user's selection information about the character images and the image and atmosphere of the music video, and sends it to the server. The device then obtains the selection information and sends it to the server in JSON format. The input of this process is the user-selected image file and image and atmosphere information, and the output is the request sent to the server.
[1377] Step 11:
[1378] The terminal displays the progress of the creation process in real time as it receives it from the server. It receives the progress via WebSocket or real-time communication and displays it in the interface as a progress bar or status messages. The input to this process is the progress update messages from the server, and the output is a progress screen that is displayed to the user.
[1379] Step 12:
[1380] The device displays a preview of the generated music video, provides a download link, and also provides the ability to directly post the video to content distribution services such as video sharing platforms. The input to this process is a download link from the server and the video file, and the output is a preview screen and a download link displayed to the user.
[1381] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1382] The system of the present invention not only includes a series of processes for users to automatically generate music videos (MVs) based on audio files, but also a function to recognize the user's emotions and customize the music video based on those emotions. Specific system operations and program processing are explained in natural language below.
[1383] Server-side processing
[1384] 1. Receiving audio files
[1385] The server receives audio files uploaded by users. For example, "Receive and store the file 'song.mp3' uploaded by the user."
[1386] 2. Metadata Extraction
[1387] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio file. The analysis results are stored in a database. For example, "Store the analyzed lyrics and beat information."
[1388] 3. Receiving User Selection Information
[1389] The server receives the user-submitted character photos and selection information about the image and atmosphere of the music video. For example, "receives user-selected 'vintage' style setting information."
[1390] 4. Activating the Emotional Engine
[1391] The server launches an emotion engine to analyze the user's voice file and real-time emotion data. The emotion engine identifies the user's emotional state and applies corresponding parameters to the generative model. For example, "Detect the emotion 'happiness' from the voice file."
[1392] 5. Launching the Generative Model
[1393] The server then activates the AI generative model based on the parsed metadata, user selection information, and emotional data obtained from the emotion engine, and begins generating the music video, for example, "applying 'bright color tones' to the scene based on the emotional state."
[1394] 6. Rendering
[1395] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video. This rendering process ensures smooth transitions between frames. For example, "Rendered for saving as a final video file."
[1396] 7. Providing results
[1397] The server provides the generated music video to the user and generates and sends a download link. For example, "Provide the user with a download link for the completed music video."
[1398] Terminal side processing
[1399] 1. Interface display
[1400] The terminal displays an interface for the user to upload an audio file. The user selects an audio file and clicks the upload button. For example, "Display a web page with an audio file upload button."
[1401] 2. Upload audio file
[1402] The terminal uploads the audio file selected by the user to the server. The selected audio file is sent to the server via an HTTP request. For example, "Send audio file to server."
[1403] 3. Select the image and atmosphere
[1404] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos and images and clicks the submit button. For example, "Display a UI showing photos of the characters and image options."
[1405] 4. Progress Display
[1406] The device will display real-time progress of the generation process as it receives it from the server, e.g., "Displays status 'Generating video... 50% complete'."
[1407] 5. View and download the final result
[1408] The device displays a preview of the generated music video and provides a download link. The user can check and download the generated music video. For example, "Display a preview screen of the generated video and a download link."
[1409] User processing
[1410] 1. Select an audio file
[1411] The user selects an audio file to upload through the interface, for example, "Select 'song.mp3' from local storage."
[1412] 2. Select the image and atmosphere
[1413] Users can choose the image and atmosphere of the character photos and music videos. For example, "Choose a profile picture and a 'vintage' style."
[1414] 3. File upload and information transmission
[1415] The user sends the selected audio file and image information to the server through the terminal. For example, "Click the upload button to send the audio file and selected information."
[1416] 4. Check the generation progress
[1417] The user checks the generation progress displayed on the device, for example, "Check that the progress bar reaches 50%."
[1418] 5. Download and share your finished product
[1419] Users can download the generated music video and share it on social media or other platforms. For example, "Click the download link to save the video, then upload it to YouTube to share."
[1420] The above is a specific example of how to implement a system incorporating the emotion engine of the present invention. This allows artists and content creators to easily create high-quality music videos customized to their own emotional state. Utilizing emotion data from the emotion engine enables more personalized visual expression, making a strong impact on viewers.
[1421] The processing flow will be explained below.
[1422] Step 1:
[1423] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[1424] Step 2:
[1425] The device sends the audio file selected by the user to the server via an HTTP request, specifically by using an input form or drag-and-drop function.
[1426] Step 3:
[1427] The server receives the audio file sent from the device and saves it in a specified directory, for example, a file named "song.mp3."
[1428] Step 4:
[1429] The server analyzes the received audio files and runs audio analysis algorithms to extract lyrics and beat information, and stores the analyzed metadata in a database.
[1430] Step 5:
[1431] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[1432] Step 6:
[1433] The device sends the photos of the characters and the image and atmosphere selection information received from the user to the server. Specifically, it sends the image files and text data via an HTTP request.
[1434] Step 7:
[1435] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model, which includes preprocessing the image data and analyzing the text data.
[1436] Step 8:
[1437] The server activates an emotion engine that analyzes the user's voice file to identify their emotional state, for example using a voice analysis algorithm to identify emotions such as "happiness" or "sadness."
[1438] Step 9:
[1439] The server adjusts the parameters of the generative model based on the identified emotional state, specifically dynamically changing the color grading and effects of the video based on the emotional data.
[1440] Step 10:
[1441] The server then activates an AI generation model based on the analyzed metadata, user selection information, and emotional data to begin generating the music video. The AI model processes the input data and generates a corresponding video clip.
[1442] Step 11:
[1443] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[1444] Step 12:
[1445] The server saves the completed music video as a file and generates a download link that can be accessed by the user.
[1446] Step 13:
[1447] The server sends the generated download link to the terminal, allowing the user to access it.
[1448] Step 14:
[1449] The terminal displays a preview of the generated music video to the user and provides a download link, allowing the user to view and download the generated music video.
[1450] Step 15:
[1451] Users can download the completed music video and share it on social media or other platforms, such as uploading it to YouTube to share with their audience.
[1452] Example 2
[1453] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1454] Conventional music video generation systems have struggled to generate personalized videos that take into account the user's emotional state. Furthermore, extracting metadata from audio files and generating videos based on user-selected images and moods are time-consuming, limiting the ability to improve the user experience. The present invention aims to solve these problems.
[1455] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1456] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and mood of a video of a visual work, means for analyzing the audio file and real-time emotional data to identify the user's emotional state, means for generating a visual work using a generative model based on the analyzed lyrics and beat information, the selection information, and the identified emotional state, and means for providing the generated visual work to the user, thereby enabling the automatic generation of personalized, high-quality music videos that take the user's emotions into consideration.
[1457] "Audio File" refers to audio data stored in digital format.
[1458] "Lyrics" refers to the words or sentences in a song within an audio file.
[1459] "Beat information" refers to data about rhythm and tempo contained in an audio file.
[1460] "Analysis" refers to the process of extracting useful information from input data.
[1461] "Character Images" refers to photographs or illustrations of people appearing in the music video.
[1462] "Visual work" refers to content that is visually displayed, such as a film or video clip.
[1463] "Image / Mood" refers to the style or theme expressed in a visual work.
[1464] "Emotional state" refers to the psychological state of the user analyzed from the voice data.
[1465] "Real-time emotional data" refers to data that indicates the current emotional state of the user.
[1466] "Generative model" refers to an algorithm that uses AI techniques to automatically generate visual works or other content.
[1467] "Generation" refers to the process of creating new data or content based on specific data.
[1468] "Rendering" refers to the process of converting digital data into a visually displayable form.
[1469] "Serving" refers to the process of transmitting or displaying the generated visual work in a form accessible to a user.
[1470] MODE FOR CARRYING OUT THE INVENTION
[1471] The system of the present invention is a system that includes a series of processes for a user to automatically generate a music video (MV) based on an audio file, recognize the user's emotions, and customize the MV based on those emotions. Specific operations of the system and program processing are described in detail below.
[1472] Server-side processing
[1473] 1. Receiving audio files
[1474] The server receives the audio file uploaded by the user. In this receiving process, the HTTP protocol is used to send the file "song.mp3" to the server and save it in a specified folder.
[1475] 2. Metadata Extraction
[1476] The server uses audio analysis algorithms (e.g., Python libraries librosa or PyDub) to extract lyrics and beat information from audio files. For example, it loads audio using the librosa.load() function and extracts beat information using the librosa.beat.beat_track() function. The analysis results are stored in a database.
[1477] 3. Receiving User Selection Information
[1478] The server receives and stores the character images and the selection information about the image and atmosphere of the visual work sent by the user via HTTP request, for example, "vintage style" or "character: profile.jpg" in the database.
[1479] 4. Activating the Emotional Engine
[1480] The server invokes an emotion engine (e.g., a general emotion analysis service) to analyze the audio file and real-time emotion data. The engine identifies the user's emotional state (e.g., happiness, sadness, etc.) and provides corresponding parameters to the generative model.
[1481] 5. Launching the Generative Model
[1482] The server generates a music video by activating an AI generation model based on the analyzed lyrics, beat information, user selection information, and emotional state. For example, the server inputs prompts to customize scene settings and color tones according to the user's emotional state into the generation AI model.
[1483] 6. Rendering
[1484] The server then combines the generated video clips into a sequence and renders the final video. This process uses a video processing tool such as FFmpeg. For example, to concatenate multiple clips, a command like ffmpeg -i input1.mp4 -i input2.mp4 -filter_complex "[0:v][1:v] concat=n=2:v=1[outv]" -map "[outv]" output.mp4 is used.
[1485] 7. Providing results
[1486] The server generates a download link for the completed music video and provides it to the user, returning a response JSON object containing the link URL.
[1487] Terminal side processing
[1488] 1. Interface display
[1489] The terminal displays an interface for the user to upload an audio file. The interface includes a file selection button and an upload button. For example, the terminal uses an HTML form or JavaScript to prompt the user to select an audio file.
[1490] 2. Upload audio file
[1491] The device uploads the audio file selected by the user to the server via an HTTP request, specifically, an AJAX request to send the file "song.mp3" to the server.
[1492] 3. Select the image and atmosphere
[1493] The device displays an interface that allows users to select images of the characters and the image and atmosphere of the music video. For example, a selection screen built with HTML5 and CSS is displayed, allowing users to select "profile.jpg" or a "vintage" style.
[1494] 4. Progress Display
[1495] The terminal displays the progress of the generation process received from the server in real time. It uses the JavaScript setInterval function to query the server for the progress at regular intervals and displays the results.
[1496] 5. View and download the final result
[1497] The device will display a preview of the generated music video and provide a download link. <video> View by tag and download link< / video> < / url:> Provided by tag.
[1498] User processing
[1499] 1. Select an audio file
[1500] The user selects the audio file to upload from the interface, specifically the file "song.mp3" from local storage.
[1501] 2. Select the image and atmosphere
[1502] The user selects the image and atmosphere of the characters and music video. They select "profile.jpg" from local storage and choose "vintage style" from the options provided in the UI.
[1503] 3. File upload and information transmission
[1504] The user sends the selected audio file and image information to the server through the terminal, and clicks the upload button to send the audio file and the selected information.
[1505] 4. Check the generation progress
[1506] The user checks the generation progress displayed on the device, specifically checking the progress message such as "Make sure the progress bar reaches 50%."
[1507] 5. Download and share your finished product
[1508] Users can download the generated music video and share it on social media or other platforms by clicking the download link to save the 'output.mp4' video, then uploading the video to YouTube or social media to share it.
[1509] Examples of concrete examples and prompts
[1510] Prompt Sentence Examples
[1511] Audio file: song.mp3
[1512] Character photo:profile.jpg
[1513] Music video vibe: Vintage
[1514] These procedures and system configurations allow users to easily generate personalized, high-quality music videos based on audio files and emotions.
[1515] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1516] Step 1:
[1517] The server receives an audio file from the user. The input is the audio file selected by the user (e.g., 'song.mp3'), and the output is the audio file stored on the server. Specific operations include requesting the audio file using the HTTP protocol and saving it in a specific folder on the server.
[1518] Step 2:
[1519] The server analyzes the lyrics and beat information from the received audio file. The input is the audio file saved in step 1, and the output is the lyrics and beat information. Specifically, it uses a Python library (e.g., librosa or PyDub) to load the audio with librosa.load() and extract the beat with librosa.beat.beat_track(). The analysis results are stored in a database.
[1520] Step 3:
[1521] The server receives from the user the image of the characters and the selection information regarding the image and atmosphere of the visual work video. The input is the image file and text information selected by the user, and the output is the selection information saved on the server. Specifically, the data sent via the HTTP request is saved in a database. For example, information such as "vintage style" or "character: profile.jpg" is recorded.
[1522] Step 4:
[1523] The server starts an emotion engine to analyze the audio file and real-time emotion data to identify the user's emotional state. The input is the lyrics and beat information analyzed in step 2 and the audio file, and the output is the identified emotional state. Specifically, the emotion analysis service detects the user's emotion (e.g., happiness) from the audio.
[1524] Step 5:
[1525] The server activates a generative AI model based on the analyzed metadata, user selection information, and emotional data to generate a visual work. The input is the data obtained in Steps 2 and 3 and the emotional state identified in Step 4, and the output is the generated visual work (video clip). Specifically, the server provides the generative AI model with prompts to customize the scene setting and color tone according to the emotional state, and generates a music video.
[1526] Step 6:
[1527] The server combines the generated video clips into a sequence and renders it as a final video. The input is the video clip generated in step 5, and the output is a high-quality final video file. Specifically, it uses a video processing tool such as FFmpeg to concatenate multiple clips and executes the ffmpeg command to save it as "output.mp4."
[1528] Step 7:
[1529] The server provides the completed music video to the user. The input is the final video file generated in step 6, and the output is a download link URL. The server then returns a response JSON object containing the URL of the generated video file to the user.
[1530] Step 8:
[1531] The terminal provides the user with an upload interface, a selection interface, a progress indicator, a final result indicator, and a download link, allowing the user to upload audio files, submit selection information, check the progress, and download the generated video.
[1532] (Application example 2)
[1533] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1534] Currently, many artists and content creators spend a great deal of time and effort creating music videos that reflect the emotions and themes of their work. Furthermore, many automated music video generation systems are not customized based on the user's emotional state, and often lack visual appeal. The present invention aims to solve this problem by providing a system that automatically generates high-quality music videos that reflect the user's emotional state.
[1535] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for analyzing the user's emotional state, means for generating video using a generative model based on the analyzed lyrics and beat information, the selection information, and the emotional state, and means for providing the generated video to the user. This enables the generation of a personalized music video based on the user's emotional state.
[1536] An "audio file" is a file in which audio data is stored in digital format.
[1537] "Lyrics" are words or text that correspond to a musical piece in an audio file.
[1538] "Beat information" refers to the rhythm and tempo information within an audio file.
[1539] "Character images" are image data of people appearing in the music video that are uploaded or selected by the user.
[1540] "Image and atmosphere of the video" refers to the visual elements and atmosphere that express the style and theme of the music video.
[1541] "Emotional state" is data that indicates the user's current mental state and emotions.
[1542] "Generative model" refers to an algorithm or AI model that automatically generates music videos based on input data.
[1543] "Video clips" are short video segments that make up a music video.
[1544] "Rendering as a sequence" refers to the process of combining a series of video clips into a continuous video and rendering it at high quality.
[1545] The system of the present invention provides a set of processes for users to automatically generate music videos (MVs) based on audio files, and further includes the ability to recognize the user's emotions and customize the MV based on those emotions.
[1546] The server first receives an audio file from the user. This audio file is audio data stored in a digital format, such as mp3 or wav. The server then uses an audio analysis algorithm to extract lyrics and beat information from the audio file. The extracted metadata is then stored in a database.
[1547] Receives from the user character images and video image and mood selection information that reflects the user's designated visual style or theme.
[1548] The emotion engine is activated to analyze the user's emotional state. The emotion engine identifies emotions from the text data and voice data entered by the user and obtains the emotional state as data.
[1549] The server uses a generative AI model to generate video based on the analyzed lyrics and beat information, user selections, and emotional state. The generative AI model receives these input data as prompts and generates a series of video clips based on them. These video clips are rendered as a sequence and combined into a high-quality video.
[1550] Finally, the generated video is provided to the user: the server saves the video file and generates a link for the user to download it.
[1551] The hardware and software used are as follows: The server uses the Django framework (Python) and an appropriate library (e.g., librosa) for speech analysis; an NLP library (e.g., NLTK, spaCy) or a speech analysis library is used for the emotion engine; and a machine learning framework (e.g., TensorFlow, PyTorch) is used for the generative AI model.
[1552] For example, if a user uploads an up-tempo song called "Happy.mp3," the system extracts lyrics and beat information from the audio file and detects the user's emotional state of happiness. If the user selects a vintage theme, the system applies visual effects that match the theme and generates a light-hearted video that expresses a sense of happiness. The generated music video is immediately available for preview and distribution.
[1553] An example of a prompt is:
[1554] "User ID: 001
[1555] Audio file URL: / media / uploads / song.mp3
[1556] Extracted Metadata: {Lyrics: 'Happy song lyrics', Beat: 'Uptempo'}
[1557] User Sentiment: Happiness
[1558] Selected theme: Vintage
[1559] The above is a specific embodiment of the system of the present invention, which allows artists and content creators to easily create high-quality music videos that are personalized to their emotional state.
[1560] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1561] Step 1:
[1562] The user selects and uploads an audio file.
[1563] Input: An audio file selected by the user from local storage (e.g. song.mp3).
[1564] Specific operation: The terminal creates and sends an HTTP request to send the audio file selected by the user through the interface to the server.
[1565] Output: The audio file is sent to the server.
[1566] Step 2:
[1567] The server receives and stores the audio files.
[1568] Input: User submitted audio file (e.g. song.mp3).
[1569] Specific operation: The server stores the received audio file in local storage or cloud storage.
[1570] Output: The path to the audio file in local or cloud storage.
[1571] Step 3:
[1572] The server analyzes the audio file and extracts the metadata.
[1573] Input: The path to the saved audio file.
[1574] What it does: The server uses audio analysis algorithms to extract lyrics and beat information from the audio file (e.g., using the librosa library).
[1575] Data processing or data manipulation: Spectral analysis of audio data to extract beat maps and lyrics in text format.
[1576] Output: Parsed lyrics and beat information metadata.
[1577] Step 4:
[1578] The server receives from the user the images of the characters and selection information regarding the image and atmosphere of the video.
[1579] Input: User-submitted image file and image selection information.
[1580] Specific operation: The server reads the image file and selection information received as an HTTP request and saves them in a database or storage.
[1581] Output: Path to the saved image file and selection information.
[1582] Step 5:
[1583] The server analyzes the user's emotional state.
[1584] Input: User text or voice data.
[1585] Specific operation: The server launches an emotion engine and analyzes the user's input data to identify their emotional state (e.g., using NLTK or spaCy).
[1586] Data processing or data calculation: Analyzing text or audio data with natural language processing algorithms to generate emotion labels (e.g., happy, sad).
[1587] Output: The user's emotional state.
[1588] Step 6:
[1589] The server generates the video using a generative AI model.
[1590] Input: Parsed lyrics and beat information metadata, user image files and selections, user emotional state.
[1591] Specific operation: The server inputs these input data as prompt sentences into a generative AI model to generate video clips (e.g., using TensorFlow or PyTorch).
[1592] Data processing or data computation: A generative AI model generates video clips based on prompts and renders them as a sequence.
[1593] Output: The generated video clip.
[1594] Step 7:
[1595] The server provides the generated video to the user.
[1596] Input: The generated video clip.
[1597] Specific operation: The server saves the video clip, generates a URL from which the user can download it, and creates and sends an HTTP response that returns the URL to the user.
[1598] Output: A URL to the video clip that users can download.
[1599] These are the processing steps of the system for realizing this application example. Through the specific operations performed at each step, users can create and download music videos that reflect their emotional state.
[1600] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1601] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1602] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1603] [Fourth embodiment]
[1604] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1605] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1606] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1607] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1608] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1609] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1610] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1611] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1612] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1613] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1614] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1615] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1616] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1617] The system of the present invention realizes a series of processes that allow users to automatically generate music videos (MVs) based on audio files. The specific operation of the system and the processing of the program are explained below in natural language.
[1618] Server-side processing
[1619] 1. Receiving audio files
[1620] The server receives audio files uploaded by users.
[1621] Example: "Store the file "song.mp3" uploaded by the user on the server."
[1622] 2. Metadata Extraction
[1623] The server extracts metadata such as lyrics and beat information from the received audio file and stores it in a database.
[1624] Example: "Analyze lyrics and beatmaps using audio analysis algorithms."
[1625] 3. Receiving User Selection Information
[1626] The server receives user-submitted photos of the characters and selection information regarding the image and atmosphere of the music video.
[1627] Example: "Receive the settings for the user-selected 'Vintage' style."
[1628] 4. Launching the Generative Model
[1629] Based on the parsed metadata and user selection information, the server activates the AI generative model and begins generating the music video.
[1630] For example: "The AI model generates a video sequence that corresponds to the lyrics and applies the selected mood."
[1631] 5. Rendering
[1632] The server combines the generated video clips into a sequence to render a high-quality music video.
[1633] For example: "Combine the sequences and render to produce the final video file."
[1634] 6. Providing results
[1635] The server provides the generated music video to the user and generates and sends a download link.
[1636] For example: "Provide users with a link to the finished video file so they can download it."
[1637] Terminal side processing
[1638] 1. Interface display
[1639] The terminal displays an interface for the user to upload an audio file.
[1640] Example: "Display a UI (user interface) with a button to upload an audio file."
[1641] 2. Upload audio file
[1642] The terminal uploads the audio file selected by the user to the server.
[1643] Example: "Send the selected file 'song.mp3' to the server."
[1644] 3. Select the image and atmosphere
[1645] The terminal receives photos of characters and selection information about the image and atmosphere to be created from the user and transmits them to the server.
[1646] Example: "Send user-selected image information to the server."
[1647] 4. Progress Display
[1648] The terminal displays the progress of the creation process in real time as received from the server.
[1649] Example: "Show status 'Generating video... 50% complete'."
[1650] 5. View and download the final result
[1651] The device displays a preview of the generated music video and provides a download link.
[1652] Example: "Display a screen to preview the generated video and provide a download link."
[1653] User processing
[1654] 1. Select an audio file
[1655] The user selects an audio file to upload through the interface.
[1656] For example: "Select 'song.mp3' from local storage."
[1657] 2. Select the image and atmosphere
[1658] Users select photos of the characters and the image and atmosphere of the music video.
[1659] For example: "Choose a profile picture and a 'vintage' style."
[1660] 3. File upload and information transmission
[1661] The user transmits the selected audio file and image information to the server through the terminal.
[1662] For example: "Click the upload button to submit your audio file and selection information."
[1663] 4. Check the generation progress
[1664] The user checks the progress of the creation displayed on the terminal.
[1665] For example: "Check that the progress bar reaches 50%."
[1666] 5. Download and share your finished product
[1667] Users can download the generated music video and share it on social media and other platforms.
[1668] For example: "Click the download link to save the video and upload it to YouTube to share."
[1669] The above is a specific embodiment of the present invention, which enables artists and content creators to produce high-quality music videos in a short amount of time without requiring special skills or expensive equipment.
[1670] The processing flow will be explained below.
[1671] Step 1:
[1672] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[1673] Step 2:
[1674] The terminal uploads the audio file selected by the user to the server, specifically, by transmitting the selected audio file to the server via an HTTP request.
[1675] Step 3:
[1676] The server receives the audio file sent from the terminal and saves it in a specified directory.
[1677] Step 4:
[1678] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio files, and the results of this analysis are stored in a database.
[1679] Step 5:
[1680] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[1681] Step 6:
[1682] The terminal transmits the character photos and image / atmosphere selection information received from the user to the server. The selection information includes image files and text data.
[1683] Step 7:
[1684] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model.
[1685] Step 8:
[1686] The server then activates the AI generation model based on the analyzed lyrics and beat information, as well as the user's selection information, and begins generating the music video. Specifically, the AI model processes the input data and generates a corresponding video clip.
[1687] Step 9:
[1688] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[1689] Step 10:
[1690] The server saves the completed music video as a file and generates a download link.
[1691] Step 11:
[1692] The server sends the generated download link to the terminal, allowing the user to access it.
[1693] Step 12:
[1694] The device displays a preview of the generated music video to the user and provides a download link, allowing the user to view, download, and save the generated music video.
[1695] Step 13:
[1696] Users can download the finished music video and share it on social media and other platforms, at which point they can also view the video and provide feedback.
[1697] Example 1
[1698] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1699] Traditional music video production requires expensive equipment and specialized skills, making it difficult for many artists and content creators. Manual editing is also time-consuming and inefficient. In response, there was a need for a method to automatically generate high-quality music videos from audio files.
[1700] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1701] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and atmosphere of the music video, means for generating a music video using a generative AI model based on the analyzed lyrics and beat information and the selection information, means for providing the generated music video to the user and generating and transmitting a download link, and means for displaying the progress of the music video generation to the user in real time, thereby enabling users to efficiently generate high-quality music videos without requiring special skills or expensive equipment.
[1702] "Audio files" are music or audio data files uploaded by users.
[1703] "Lyrics" refers to the words or sentences within an audio file, which may contain poetic expressions or narrative.
[1704] "Beat information" is data that indicates the rhythm and tempo within an audio file.
[1705] "Character Images" are photographs or illustrations of people that appear in the music video.
[1706] "Music video atmosphere" refers to the overall style or theme of the video selected by the user, such as "vintage" or "futuristic."
[1707] A "generative AI model" is an artificial intelligence algorithm that generates video sequences or video clips based on audio file data and user selections.
[1708] "Rendering" is the process of combining the generated video clips into a sequence to create the final video file.
[1709] "Download link" refers to a URL that allows a user to download the generated music video.
[1710] The "means for displaying the generation progress status in real time" is a function for displaying the progress of the server-side processing on the user terminal.
[1711] The present invention provides a system that allows users to automatically generate music videos (MVs) based on audio files. The specific operation and processing of the system are described below. The system mainly operates in cooperation with three parties: a server, a terminal, and a user.
[1712] Hardware and software used
[1713] Hardware:
[1714] Server: Cloud server or on-premise server.
[1715] Device: A PC, smartphone, tablet, or other device capable of operating a user interface.
[1716] software:
[1717] Speech analysis library: Librosa.
[1718] Video generation algorithms: Generative AI models (e.g., OpenAI's GPT-3, DALL-E).
[1719] Video editing tool: FFmpeg.
[1720] Web frameworks: Django, Flask, etc.
[1721] Example of system operation
[1722] Receive audio files:
[1723] The user selects an audio file to upload through the device interface. Example: "The user selects 'song.mp3' from local storage."
[1724] The device uploads this audio file to the server. Example: "Send the selected 'song.mp3' file to the server."
[1725] Metadata Extraction:
[1726] The server receives the audio file and parses it for lyrics and beat information using an audio analysis library such as Librosa. Example: "Using Librosa, parse 'uploads / song.mp3' for lyrics and beatmap data and store them in a database."
[1727] Receive user selection information:
[1728] Through the interface, users select the style and character images for the music video. For example, "User selects profile picture and 'vintage' style."
[1729] The terminal transmits this information to the server.
[1730] Launch the generative model:
[1731] The server then triggers a generative AI model based on the parsed metadata and user selections to generate a music video. For example, "The AI model takes lyrics and beatmaps as input and generates a video sequence."
[1732] rendering:
[1733] The server combines the generated video clips using a video editing tool such as FFmpeg and renders a high-quality music video. Example: "Combine the generated video clips using FFmpeg as 'final_video.mp4' and render."
[1734] Results provided:
[1735] The server generates a download link for the generated music video and provides it to the user. Example: "Email the user a download link for the generated 'final_video.mp4'."
[1736] The device has the ability to display the generation progress to the user in real time. For example, "Display the status 'Generating video... 50% complete' on the screen."
[1737] Examples of prompt statements
[1738] For example, the following prompt sentence is input to a generative AI model:
[1739] "Generate a vintage-style video sequence based on the lyrics of this audio file."
[1740] The video clips thus generated can be combined into a sequence to obtain the final video file.
[1741] This invention enables users to efficiently generate high-quality music videos without requiring special skills or expensive equipment, which significantly reduces time and costs compared to conventional methods.
[1742] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1743] Server-side processing
[1744] Step 1: Receive the audio file
[1745] The server receives the audio file uploaded by the user and stores it in a specified directory.
[1746] Input: An audio file (e.g. song.mp3) that the user uploads from their device.
[1747] Data processing: Send the audio file to the server as an HTTP request.
[1748] Output: An audio file saved in the specified directory on the server (e.g. uploads / song.mp3).
[1749] Specific behavior: "The server receives the HTTP request and saves the audio file to 'uploads / song.mp3'."
[1750] Step 2: Metadata extraction
[1751] The server extracts lyrics and beat information from the received audio file using an audio analysis library such as Librosa.
[1752] Input: A saved audio file (e.g. uploads / song.mp3).
[1753] Data processing: Analysis of audio files using Librosa.
[1754] Output: Extracted lyrics and beat information (stored in a database).
[1755] Specific operation: "Using Librosa, parse the lyrics and beatmap from the audio file 'uploads / song.mp3' and store them in a database."
[1756] Step 3: Receiving User Selection Information
[1757] The server receives user-submitted images of the characters and selection information regarding the image and atmosphere of the music video.
[1758] Input: Image file and selection information (e.g., vintage style) sent by the user from the device.
[1759] Data processing: The selected information is saved in the server database.
[1760] Output: User selection information and images stored on the server.
[1761] Specific operation: "The image information selected by the user and the character's image are saved in the 'user_preferences' directory on the server."
[1762] Step 4: Launching the generative model
[1763] The server then activates the generative AI model based on the parsed metadata and user selections to begin generating the music video.
[1764] Input: Parsed lyrics, beat information, user selection information and images.
[1765] Data processing: Input to a generative AI model to generate a video sequence according to the prompt.
[1766] Output: The generated video clip.
[1767] What it does: "The AI model takes lyrics and a beatmap as input and generates a video sequence based on a prompt."
[1768] Step 5: Rendering
[1769] The server then combines the generated video clips using a video editing tool such as FFmpeg to render a high-quality music video.
[1770] Input: The generated video clip.
[1771] Data processing: Video sequences were combined using FFmpeg.
[1772] Output: The final video file (e.g. final_video.mp4).
[1773] Specific operation: "Use FFmpeg to combine the generated video clips into 'final_video.mp4' and render it."
[1774] Step 6: Delivering results
[1775] The server generates and provides a download link for the generated music video to the user.
[1776] Input: The rendered music video file.
[1777] Data processing: Generate download links.
[1778] Output: The download link sent to the user.
[1779] Specific action: "Send the download link for the generated 'final_video.mp4' to the user via email."
[1780] Terminal side processing
[1781] Step 1: Interface display
[1782] The terminal displays an interface for the user to upload an audio file and enter selection information.
[1783] Input: None.
[1784] Data processing: GUI generation.
[1785] Output: The displayed interface.
[1786] Specific behavior: "When a user opens a browser, a button for uploading an audio file appears."
[1787] Step 2: Upload your audio file
[1788] The terminal uploads the audio file selected by the user to the server.
[1789] Input: User selected audio file (e.g. song.mp3).
[1790] Data processing: Send the audio file to the server.
[1791] Output: Audio file uploaded to the server.
[1792] Specific behavior: "When the user selects 'song.mp3' and clicks the upload button, this file will be sent to the server."
[1793] Step 3: Choose the image and mood
[1794] The terminal receives from the user the images of the characters and the selection information regarding the image and atmosphere of the music video, and transmits the information to the server.
[1795] Input: User-selected image file and image information (e.g., vintage style).
[1796] Data processing: Send the selected information to the server.
[1797] Output: Selection information and images sent to the server.
[1798] What it does: "When a user selects the 'Vintage' style and uploads a photo, this information is sent to a server."
[1799] Step 4: Viewing progress
[1800] The terminal displays the progress of the creation process in real time as received from the server.
[1801] Input: Progress data sent from the server.
[1802] Data processing: Display of progress data.
[1803] Output: Real-time updated progress indicator.
[1804] Specific behavior: "Displays the status 'Generating video... 50% complete' on the screen."
[1805] Step 5: View and download the final result
[1806] The device displays a preview of the generated music video and provides a download link.
[1807] Input: The download link sent by the server.
[1808] Data processing: Display links and preview screens.
[1809] Output: Preview and downloadable link.
[1810] What it does: "You'll be presented with a screen that previews the generated video and provides a download link below it."
[1811] User processing
[1812] Step 1: Select an audio file
[1813] The user selects an audio file to upload through the terminal interface.
[1814] Input: None.
[1815] Data processing: Using the file selection dialog.
[1816] Output: The selected audio file.
[1817] Specific behavior: "The user selects 'song.mp3' from local storage in the file selection dialog."
[1818] Step 2: Choose the image and mood
[1819] Users select images of the characters and the image and atmosphere of the music video.
[1820] Input: None.
[1821] Data manipulation: Use of drop-down menus and file upload functions.
[1822] Output: Selected image files and image information.
[1823] What it does: "User drags and drops a profile photo and selects the 'Vintage' style from the drop-down menu."
[1824] Step 3: Upload files and submit information
[1825] The user transmits the selected audio file and image information to the server through the terminal.
[1826] Input: Audio files and image information.
[1827] Data processing: Sending audio files and selected information.
[1828] Output: Audio file and selection information sent to the server.
[1829] Specific action: "Click the upload button to send the audio file and selected information to the server."
[1830] Step 4: Check the generation progress
[1831] The user can check the progress of the creation displayed on the terminal in real time.
[1832] Input: Progress data sent from the server.
[1833] Data processing: Checking progress data.
[1834] Output: Progress display on screen.
[1835] What it does: "A progress bar appears on the screen and updates in real time to show progress, such as 50% complete."
[1836] Step 5: Download and share your finished product
[1837] Users can download the generated music videos to their devices and share them on social media and other platforms.
[1838] Input: Download link.
[1839] Data processing: Use of download links and retrieval of files.
[1840] Output: Downloaded music video files.
[1841] Action: "Click the download link to save 'final_video.mp4' locally, then upload the video to YouTube to share it."
[1842] (Application example 1)
[1843] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1844] Traditional music video production required advanced expertise and expensive equipment, making it difficult for average users to easily create high-quality music videos. Furthermore, there was no way for users to check the progress of automatically generated music videos in real time, nor was there an easy way to directly post generated videos to content distribution services. This reduced user convenience and prevented content from being quickly shared.
[1845] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1846] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for generating a music video using a generative model based on the analyzed lyrics and beat information and the selection information, means for displaying the progress of the generated music video, and means for providing the generated music video to the user and enabling it to be directly posted to a content distribution service. This allows users to create high-quality music videos without specialized knowledge or expensive equipment, to check the progress in real time, and to quickly share the generated video with a content distribution service.
[1847] An "audio file" is a digital file containing music or sound that is uploaded by a user.
[1848] "Lyrics" refers to words and phrases in an audio file, and are text data that make up the sung portion of the music.
[1849] "Beat information" is data about rhythm and tempo obtained from an audio file, and indicates the rhythmic structure of the music.
[1850] "Character images" are photographs or illustrations of people to appear in the music video.
[1851] "Video image and atmosphere" refers to the overall style and visual theme of the music video, and is the visual atmosphere and taste specified by the user.
[1852] "Selection information" refers to information about the image and atmosphere of the characters and video set by the user.
[1853] A "generative model" is an algorithm or framework that uses AI technology to automatically generate music videos.
[1854] A "music video" is video content generated based on an audio file and includes visual storytelling set to music.
[1855] "Content distribution service" means an online service for distributing digital content, such as generated music videos, and includes, for example, a video sharing platform.
[1856] MODE FOR CARRYING OUT THE INVENTION
[1857] A detailed description of an embodiment of the present invention is provided below: The system allows users to automatically generate music videos from audio files, display the progress of the generation, and post the generated music videos directly to a content distribution service.
[1858] Server-side processing
[1859] 1. Receiving audio files: The server receives and stores audio files uploaded by users. The software used is Flask (Python).
[1860] 2. Metadata extraction: The server parses the lyrics and beat information from the received audio files and stores them in a database. This process uses the librosa and speech_recognition libraries.
[1861] 3. Receiving user selection information: The server receives the character images and music video image / atmosphere selection information sent by the user and stores them in a database using Flask and MySQL.
[1862] 4. Launching the generative model: The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selection information. The generative model uses TensorFlow or PyTorch.
[1863] 5. Rendering: The server combines the generated video clips into a sequence and renders the final music video using OpenCV software.
[1864] 6. Progress display: The server manages the progress of the generation process in real time and notifies the user.
[1865] 7. Result Serving: The server serves the generated music video to the user, generates and sends a download link, and also allows the user to submit the generated video directly to a content distribution service.
[1866] Terminal side processing
[1867] 1. Interface display: The terminal displays an interface for users to upload audio files. This interface is built using React Native.
[1868] 2. Audio file upload: The terminal uploads the audio file selected by the user to the server.
[1869] 3. Image and atmosphere selection: The terminal receives selection information from the user regarding the character images and the image and atmosphere of the music video, and transmits it to the server.
[1870] 4. Progress display: The terminal displays the progress of the generation process received from the server in real time.
[1871] 5. View, download, and share the final result: The device displays a preview of the generated music video, provides a download link, and also allows users to post it directly to content distribution services such as video sharing platforms.
[1872] User processing
[1873] 1. Audio file selection: The user selects an audio file to upload through the interface.
[1874] 2. Image / Mood Selection: Users select the image of the characters and the image / mood of the music video.
[1875] 3. File upload and information transmission: The user transmits the selected audio file and image information to the server through the terminal.
[1876] 4. Checking the progress of generation: The user checks the progress of generation displayed on the terminal.
[1877] 5. Download and share the finished product: Users can download the generated music video and share it on social media or other platforms, or post the generated video directly to a content distribution service.
[1878] Specific examples
[1879] As a concrete example, consider a situation where a user wants to create a music video using a new song by a popular Japanese band that can be shared on social media. Using the application, the user uploads an MP3 file of the new song and selects a particular music video theme (e.g., cyberpunk). An example prompt for this would be:
[1880] "Create a cyberpunk-style video that makes it easy to hear the lyrics. Use the attached image as a photo of the characters and a futuristic city in the background."
[1881] Based on this prompt, the generative AI model automatically generates a music video that reflects the style and content specified by the user.
[1882] In this way, users can create professional music videos in a short time without needing special skills or expensive equipment, and post them to content distribution services.
[1883] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1884] Step 1:
[1885] The server receives audio files from users. The audio files (e.g., MP3 format) uploaded by users are received via HTTP requests and stored in cloud storage such as an AWS S3 bucket. The input of this process is the audio file uploaded by the user, and the output is the path to the file stored in cloud storage.
[1886] Step 2:
[1887] The server analyzes the lyrics and beat information from the received audio file. It uses the librosa and speech_recognition libraries to extract rhythm, tempo (beat information), and lyrics from the audio file. The input for this process is the audio file read from cloud storage, and the output is the metadata (beat information and lyric data) resulting from the analysis.
[1888] Step 3:
[1889] The server receives the character images and music video image / mood selections sent by the user and stores them in a database using JSON data received via an HTTP POST request. The input to this process is the user-selected image files and image / mood selections, and the output is a record of the selections stored in the database.
[1890] Step 4:
[1891] The server generates a music video using a generative model based on the analyzed lyrics and beat information, as well as the user's selections. TensorFlow or PyTorch is used as the generative model, and a video sequence is generated by inputting prompts into the AI model. The inputs for this process are metadata and the user's selections, and the output is the generated video clip.
[1892] Step 5:
[1893] The server combines the generated video clips into a sequence and renders the final music video. OpenCV is used to combine multiple video clips into a sequence and encode them to generate the final video file. The input to this process is a list of video clips, and the output is the final rendered music video file.
[1894] Step 6:
[1895] The server manages the progress of the generation process in real time and notifies the device. It tracks the progress and notifies the device of progress updates sequentially using WebSocket or a real-time communication protocol. The input of this process is the current status of the generation process, and the output is progress update messages.
[1896] Step 7:
[1897] The server provides the generated music video to the user and generates a download link to send to the device. It also allows the user to post the generated video directly to a content distribution service by generating an AWS S3 download link and providing the endpoint URL to the user. The input of this process is the generated music video file, and the output is the download link and the posting endpoint.
[1898] Step 8:
[1899] The terminal displays an interface for the user to upload an audio file. React Native is used to build the user interface and display buttons and selection items for uploading audio files. The input is the design information for the user interface, and the output is the interface screen displayed to the user.
[1900] Step 9:
[1901] The terminal uploads the audio file selected by the user to the server. The audio file obtained from the file input is sent to the server via an HTTP request. The input of this process is the audio file selected by the user, and the output is the request sent to the server.
[1902] Step 10:
[1903] The device receives the user's selection information about the character images and the image and atmosphere of the music video, and sends it to the server. The device then obtains the selection information and sends it to the server in JSON format. The input of this process is the user-selected image file and image and atmosphere information, and the output is the request sent to the server.
[1904] Step 11:
[1905] The terminal displays the progress of the creation process in real time as it receives it from the server. It receives the progress via WebSocket or real-time communication and displays it in the interface as a progress bar or status messages. The input to this process is the progress update messages from the server, and the output is a progress screen that is displayed to the user.
[1906] Step 12:
[1907] The device displays a preview of the generated music video, provides a download link, and also provides the ability to directly post the video to content distribution services such as video sharing platforms. The input to this process is a download link from the server and the video file, and the output is a preview screen and a download link displayed to the user.
[1908] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1909] The system of the present invention not only includes a series of processes for users to automatically generate music videos (MVs) based on audio files, but also a function to recognize the user's emotions and customize the music video based on those emotions. Specific system operations and program processing are explained in natural language below.
[1910] Server-side processing
[1911] 1. Receiving audio files
[1912] The server receives audio files uploaded by users. For example, "Receive and store the file 'song.mp3' uploaded by the user."
[1913] 2. Metadata Extraction
[1914] The server runs an audio analysis algorithm to extract lyrics and beat information from the received audio file. The analysis results are stored in a database. For example, "Store the analyzed lyrics and beat information."
[1915] 3. Receiving User Selection Information
[1916] The server receives the user-submitted character photos and selection information about the image and atmosphere of the music video. For example, "receives user-selected 'vintage' style setting information."
[1917] 4. Activating the Emotional Engine
[1918] The server launches an emotion engine to analyze the user's voice file and real-time emotion data. The emotion engine identifies the user's emotional state and applies corresponding parameters to the generative model. For example, "Detect the emotion 'happiness' from the voice file."
[1919] 5. Launching the Generative Model
[1920] The server then activates the AI generative model based on the parsed metadata, user selection information, and emotional data obtained from the emotion engine, and begins generating the music video, for example, "applying 'bright color tones' to the scene based on the emotional state."
[1921] 6. Rendering
[1922] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video. This rendering process ensures smooth transitions between frames. For example, "Rendered for saving as a final video file."
[1923] 7. Providing results
[1924] The server provides the generated music video to the user and generates and sends a download link. For example, "Provide the user with a download link for the completed music video."
[1925] Terminal side processing
[1926] 1. Interface display
[1927] The terminal displays an interface for the user to upload an audio file. The user selects an audio file and clicks the upload button. For example, "Display a web page with an audio file upload button."
[1928] 2. Upload audio file
[1929] The terminal uploads the audio file selected by the user to the server. The selected audio file is sent to the server via an HTTP request. For example, "Send audio file to server."
[1930] 3. Select the image and atmosphere
[1931] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos and images and clicks the submit button. For example, "Display a UI showing photos of the characters and image options."
[1932] 4. Progress Display
[1933] The device will display real-time progress of the generation process as it receives it from the server, e.g., "Displays status 'Generating video... 50% complete'."
[1934] 5. View and download the final result
[1935] The device displays a preview of the generated music video and provides a download link. The user can check and download the generated music video. For example, "Display a preview screen of the generated video and a download link."
[1936] User processing
[1937] 1. Select an audio file
[1938] The user selects an audio file to upload through the interface, for example, "Select 'song.mp3' from local storage."
[1939] 2. Select the image and atmosphere
[1940] Users can choose the image and atmosphere of the character photos and music videos. For example, "Choose a profile picture and a 'vintage' style."
[1941] 3. File upload and information transmission
[1942] The user sends the selected audio file and image information to the server through the terminal. For example, "Click the upload button to send the audio file and selected information."
[1943] 4. Check the generation progress
[1944] The user checks the generation progress displayed on the device, for example, "Check that the progress bar reaches 50%."
[1945] 5. Download and share your finished product
[1946] Users can download the generated music video and share it on social media or other platforms. For example, "Click the download link to save the video, then upload it to YouTube to share."
[1947] The above is a specific example of how to implement a system incorporating the emotion engine of the present invention. This allows artists and content creators to easily create high-quality music videos customized to their own emotional state. Utilizing emotion data from the emotion engine enables more personalized visual expression, making a strong impact on viewers.
[1948] The processing flow will be explained below.
[1949] Step 1:
[1950] The terminal displays an interface for the user to upload an audio file, and the user selects the audio file and clicks the upload button.
[1951] Step 2:
[1952] The device sends the audio file selected by the user to the server via an HTTP request, specifically by using an input form or drag-and-drop function.
[1953] Step 3:
[1954] The server receives the audio file sent from the device and saves it in a specified directory, for example, a file named "song.mp3."
[1955] Step 4:
[1956] The server analyzes the received audio files and runs audio analysis algorithms to extract lyrics and beat information, and stores the analyzed metadata in a database.
[1957] Step 5:
[1958] The device displays an interface for the user to select photos of the characters and the image and atmosphere of the music video. The user selects the required photos, image and atmosphere and clicks the send button.
[1959] Step 6:
[1960] The device sends the photos of the characters and the image and atmosphere selection information received from the user to the server. Specifically, it sends the image files and text data via an HTTP request.
[1961] Step 7:
[1962] The server receives the photos and image / atmosphere selection information sent from the device and prepares them for input into the generative model, which includes preprocessing the image data and analyzing the text data.
[1963] Step 8:
[1964] The server activates an emotion engine that analyzes the user's voice file to identify their emotional state, for example using a voice analysis algorithm to identify emotions such as "happiness" or "sadness."
[1965] Step 9:
[1966] The server adjusts the parameters of the generative model based on the identified emotional state, specifically dynamically changing the color grading and effects of the video based on the emotional data.
[1967] Step 10:
[1968] The server then activates an AI generation model based on the analyzed metadata, user selection information, and emotional data to begin generating the music video. The AI model processes the input data and generates a corresponding video clip.
[1969] Step 11:
[1970] The server then stitches the resulting video clips together into a sequence and renders it into a high-quality music video, ensuring smooth transitions between frames during the rendering process.
[1971] Step 12:
[1972] The server saves the completed music video as a file and generates a download link that can be accessed by the user.
[1973] Step 13:
[1974] The server sends the generated download link to the terminal, allowing the user to access it.
[1975] Step 14:
[1976] The terminal displays a preview of the generated music video to the user and provides a download link, allowing the user to view and download the generated music video.
[1977] Step 15:
[1978] Users can download the completed music video and share it on social media or other platforms, such as uploading it to YouTube to share with their audience.
[1979] Example 2
[1980] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1981] Conventional music video generation systems have struggled to generate personalized videos that take into account the user's emotional state. Furthermore, extracting metadata from audio files and generating videos based on user-selected images and moods are time-consuming, limiting the ability to improve the user experience. The present invention aims to solve these problems.
[1982] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1983] In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving from the user selection information regarding images of characters and the image and mood of a video of a visual work, means for analyzing the audio file and real-time emotional data to identify the user's emotional state, means for generating a visual work using a generative model based on the analyzed lyrics and beat information, the selection information, and the identified emotional state, and means for providing the generated visual work to the user, thereby enabling the automatic generation of personalized, high-quality music videos that take the user's emotions into consideration.
[1984] "Audio File" refers to audio data stored in digital format.
[1985] "Lyrics" refers to the words or sentences in a song within an audio file.
[1986] "Beat information" refers to data about rhythm and tempo contained in an audio file.
[1987] "Analysis" refers to the process of extracting useful information from input data.
[1988] "Character Images" refers to photographs or illustrations of people appearing in the music video.
[1989] "Visual work" refers to content that is visually displayed, such as a film or video clip.
[1990] "Image / Mood" refers to the style or theme expressed in a visual work.
[1991] "Emotional state" refers to the psychological state of the user analyzed from the voice data.
[1992] "Real-time emotional data" refers to data that indicates the current emotional state of the user.
[1993] "Generative model" refers to an algorithm that uses AI techniques to automatically generate visual works or other content.
[1994] "Generation" refers to the process of creating new data or content based on specific data.
[1995] "Rendering" refers to the process of converting digital data into a visually displayable form.
[1996] "Serving" refers to the process of transmitting or displaying the generated visual work in a form accessible to a user.
[1997] MODE FOR CARRYING OUT THE INVENTION
[1998] The system of the present invention is a system that includes a series of processes for a user to automatically generate a music video (MV) based on an audio file, recognize the user's emotions, and customize the MV based on those emotions. Specific operations of the system and program processing are described in detail below.
[1999] Server-side processing
[2000] 1. Receiving audio files
[2001] The server receives the audio file uploaded by the user. In this receiving process, the HTTP protocol is used to send the file "song.mp3" to the server and save it in a specified folder.
[2002] 2. Metadata Extraction
[2003] The server uses audio analysis algorithms (e.g., Python libraries librosa or PyDub) to extract lyrics and beat information from audio files. For example, it loads audio using the librosa.load() function and extracts beat information using the librosa.beat.beat_track() function. The analysis results are stored in a database.
[2004] 3. Receiving User Selection Information
[2005] The server receives and stores the character images and the selection information about the image and atmosphere of the visual work sent by the user via HTTP request, for example, "vintage style" or "character: profile.jpg" in the database.
[2006] 4. Activating the Emotional Engine
[2007] The server invokes an emotion engine (e.g., a general emotion analysis service) to analyze the audio file and real-time emotion data. The engine identifies the user's emotional state (e.g., happiness, sadness, etc.) and provides corresponding parameters to the generative model.
[2008] 5. Launching the Generative Model
[2009] The server generates a music video by activating an AI generation model based on the analyzed lyrics, beat information, user selection information, and emotional state. For example, the server inputs prompts to customize scene settings and color tones according to the user's emotional state into the generation AI model.
[2010] 6. Rendering
[2011] The server then combines the generated video clips into a sequence and renders the final video. This process uses a video processing tool such as FFmpeg. For example, to concatenate multiple clips, a command like ffmpeg -i input1.mp4 -i input2.mp4 -filter_complex "[0:v][1:v] concat=n=2:v=1[outv]" -map "[outv]" output.mp4 is used.
[2012] 7. Providing results
[2013] The server generates a download link for the completed music video and provides it to the user, returning a response JSON object containing the link URL.
[2014] Terminal side processing
[2015] 1. Interface display
[2016] The terminal displays an interface for the user to upload an audio file. The interface includes a file selection button and an upload button. For example, the terminal uses an HTML form or JavaScript to prompt the user to select an audio file.
[2017] 2. Upload audio file
[2018] The device uploads the audio file selected by the user to the server via an HTTP request, specifically, an AJAX request to send the file "song.mp3" to the server.
[2019] 3. Select the image and atmosphere
[2020] The device displays an interface that allows users to select images of the characters and the image and atmosphere of the music video. For example, a selection screen built with HTML5 and CSS is displayed, allowing users to select "profile.jpg" or a "vintage" style.
[2021] 4. Progress Display
[2022] The terminal displays the progress of the generation process received from the server in real time. It uses the JavaScript setInterval function to query the server for the progress at regular intervals and displays the results.
[2023] 5. View and download the final result
[2024] The device will display a preview of the generated music video and provide a download link. <video> View by tag and download link< / video> < / url:> Provided by tag.
[2025] User processing
[2026] 1. Select an audio file
[2027] The user selects the audio file to upload from the interface, specifically the file "song.mp3" from local storage.
[2028] 2. Select the image and atmosphere
[2029] The user selects the image and atmosphere of the characters and music video. They select "profile.jpg" from local storage and choose "vintage style" from the options provided in the UI.
[2030] 3. File upload and information transmission
[2031] The user sends the selected audio file and image information to the server through the terminal, and clicks the upload button to send the audio file and the selected information.
[2032] 4. Check the generation progress
[2033] The user checks the generation progress displayed on the device, specifically checking the progress message such as "Make sure the progress bar reaches 50%."
[2034] 5. Download and share your finished product
[2035] Users can download the generated music video and share it on social media or other platforms by clicking the download link to save the 'output.mp4' video, then uploading the video to YouTube or social media to share it.
[2036] Examples of concrete examples and prompts
[2037] Prompt Sentence Examples
[2038] Audio file: song.mp3
[2039] Character photo:profile.jpg
[2040] Music video vibe: Vintage
[2041] These procedures and system configurations allow users to easily generate personalized, high-quality music videos based on audio files and emotions.
[2042] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2043] Step 1:
[2044] The server receives an audio file from the user. The input is the audio file selected by the user (e.g., 'song.mp3'), and the output is the audio file stored on the server. Specific operations include requesting the audio file using the HTTP protocol and saving it in a specific folder on the server.
[2045] Step 2:
[2046] The server analyzes the lyrics and beat information from the received audio file. The input is the audio file saved in step 1, and the output is the lyrics and beat information. Specifically, it uses a Python library (e.g., librosa or PyDub) to load the audio with librosa.load() and extract the beat with librosa.beat.beat_track(). The analysis results are stored in a database.
[2047] Step 3:
[2048] The server receives from the user the image of the characters and the selection information regarding the image and atmosphere of the visual work video. The input is the image file and text information selected by the user, and the output is the selection information saved on the server. Specifically, the data sent via the HTTP request is saved in a database. For example, information such as "vintage style" or "character: profile.jpg" is recorded.
[2049] Step 4:
[2050] The server starts an emotion engine to analyze the audio file and real-time emotion data to identify the user's emotional state. The input is the lyrics and beat information analyzed in step 2 and the audio file, and the output is the identified emotional state. Specifically, the emotion analysis service detects the user's emotion (e.g., happiness) from the audio.
[2051] Step 5:
[2052] The server activates a generative AI model based on the analyzed metadata, user selection information, and emotional data to generate a visual work. The input is the data obtained in Steps 2 and 3 and the emotional state identified in Step 4, and the output is the generated visual work (video clip). Specifically, the server provides the generative AI model with prompts to customize the scene setting and color tone according to the emotional state, and generates a music video.
[2053] Step 6:
[2054] The server combines the generated video clips into a sequence and renders it as a final video. The input is the video clip generated in step 5, and the output is a high-quality final video file. Specifically, it uses a video processing tool such as FFmpeg to concatenate multiple clips and executes the ffmpeg command to save it as "output.mp4."
[2055] Step 7:
[2056] The server provides the completed music video to the user. The input is the final video file generated in step 6, and the output is a download link URL. The server then returns a response JSON object containing the URL of the generated video file to the user.
[2057] Step 8:
[2058] The terminal provides the user with an upload interface, a selection interface, a progress indicator, a final result indicator, and a download link, allowing the user to upload audio files, submit selection information, check the progress, and download the generated video.
[2059] (Application example 2)
[2060] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2061] Currently, many artists and content creators spend a great deal of time and effort creating music videos that reflect the emotions and themes of their work. Furthermore, many automated music video generation systems are not customized based on the user's emotional state, and often lack visual appeal. The present invention aims to solve this problem by providing a system that automatically generates high-quality music videos that reflect the user's emotional state.
[2062] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving an audio file from a user, means for analyzing lyrics and beat information from the audio file, means for receiving selection information from the user regarding images of characters and the image and atmosphere of the video, means for analyzing the user's emotional state, means for generating video using a generative model based on the analyzed lyrics and beat information, the selection information, and the emotional state, and means for providing the generated video to the user. This enables the generation of a personalized music video based on the user's emotional state.
[2063] An "audio file" is a file in which audio data is stored in digital format.
[2064] "Lyrics" are words or text that correspond to a musical piece in an audio file.
[2065] "Beat information" refers to the rhythm and tempo information within an audio file.
[2066] "Character images" are image data of people appearing in the music video that are uploaded or selected by the user.
[2067] "Image and atmosphere of the video" refers to the visual elements and atmosphere that express the style and theme of the music video.
[2068] "Emotional state" is data that indicates the user's current mental state and emotions.
[2069] "Generative model" refers to an algorithm or AI model that automatically generates music videos based on input data.
[2070] "Video clips" are short video segments that make up a music video.
[2071] "Rendering as a sequence" refers to the process of combining a series of video clips into a continuous video and rendering it at high quality.
[2072] The system of the present invention provides a set of processes for users to automatically generate music videos (MVs) based on audio files, and further includes the ability to recognize the user's emotions and customize the MV based on those emotions.
[2073] The server first receives an audio file from the user. This audio file is audio data stored in a digital format, such as mp3 or wav. The server then uses an audio analysis algorithm to extract lyrics and beat information from the audio file. The extracted metadata is then stored in a database.
[2074] Receives from the user character images and video image and mood selection information that reflects the user's designated visual style or theme.
[2075] The emotion engine is activated to analyze the user's emotional state. The emotion engine identifies emotions from the text data and voice data entered by the user and obtains the emotional state as data.
[2076] The server uses a generative AI model to generate video based on the analyzed lyrics and beat information, user selections, and emotional state. The generative AI model receives these input data as prompts and generates a series of video clips based on them. These video clips are rendered as a sequence and combined into a high-quality video.
[2077] Finally, the generated video is provided to the user: the server saves the video file and generates a link for the user to download it.
[2078] The hardware and software used are as follows: The server uses the Django framework (Python) and an appropriate library (e.g., librosa) for speech analysis; an NLP library (e.g., NLTK, spaCy) or a speech analysis library is used for the emotion engine; and a machine learning framework (e.g., TensorFlow, PyTorch) is used for the generative AI model.
[2079] For example, if a user uploads an up-tempo song called "Happy.mp3," the system extracts lyrics and beat information from the audio file and detects the user's emotional state of happiness. If the user selects a vintage theme, the system applies visual effects that match the theme and generates a light-hearted video that expresses a sense of happiness. The generated music video is immediately available for preview and distribution.
[2080] An example of a prompt is:
[2081] "User ID: 001
[2082] Audio file URL: / media / uploads / song.mp3
[2083] Extracted Metadata: {Lyrics: 'Happy song lyrics', Beat: 'Uptempo'}
[2084] User Sentiment: Happiness
[2085] Selected theme: Vintage
[2086] The above is a specific embodiment of the system of the present invention, which allows artists and content creators to easily create high-quality music videos that are personalized to their emotional state.
[2087] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2088] Step 1:
[2089] The user selects and uploads an audio file.
[2090] Input: An audio file selected by the user from local storage (e.g. song.mp3).
[2091] Specific operation: The terminal creates and sends an HTTP request to send the audio file selected by the user through the interface to the server.
[2092] Output: The audio file is sent to the server.
[2093] Step 2:
[2094] The server receives and stores the audio files.
[2095] Input: User submitted audio file (e.g. song.mp3).
[2096] Specific operation: The server stores the received audio file in local storage or cloud storage.
[2097] Output: The path to the audio file in local or cloud storage.
[2098] Step 3:
[2099] The server analyzes the audio file and extracts the metadata.
[2100] Input: The path to the saved audio file.
[2101] What it does: The server uses audio analysis algorithms to extract lyrics and beat information from the audio file (e.g., using the librosa library).
[2102] Data processing or data manipulation: Spectral analysis of audio data to extract beat maps and lyrics in text format.
[2103] Output: Parsed lyrics and beat information metadata.
[2104] Step 4:
[2105] The server receives from the user the images of the characters and selection information regarding the image and atmosphere of the video.
[2106] Input: User-submitted image file and image selection information.
[2107] Specific operation: The server reads the image file and selection information received as an HTTP request and saves them in a database or storage.
[2108] Output: Path to the saved image file and selection information.
[2109] Step 5:
[2110] The server analyzes the user's emotional state.
[2111] Input: User text or voice data.
[2112] Specific operation: The server launches an emotion engine and analyzes the user's input data to identify their emotional state (e.g., using NLTK or spaCy).
[2113] Data processing or data calculation: Analyzing text or audio data with natural language processing algorithms to generate emotion labels (e.g., happy, sad).
[2114] Output: The user's emotional state.
[2115] Step 6:
[2116] The server generates the video using a generative AI model.
[2117] Input: Parsed lyrics and beat information metadata, user image files and selections, user emotional state.
[2118] Specific operation: The server inputs these input data as prompt sentences into a generative AI model to generate video clips (e.g., using TensorFlow or PyTorch).
[2119] Data processing or data computation: A generative AI model generates video clips based on prompts and renders them as a sequence.
[2120] Output: The generated video clip.
[2121] Step 7:
[2122] The server provides the generated video to the user.
[2123] Input: The generated video clip.
[2124] Specific operation: The server saves the video clip, generates a URL from which the user can download it, and creates and sends an HTTP response that returns the URL to the user.
[2125] Output: A URL to the video clip that users can download.
[2126] These are the processing steps of the system for realizing this application example. Through the specific operations performed at each step, users can create and download music videos that reflect their emotional state.
[2127] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[2128] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2129] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[2130] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2131] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[2132] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[2133] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[2134] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[2135] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[2136] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[2137] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[2138] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[2139] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[2140] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[2141] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[2142] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[2143] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[2144] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[2145] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[2146] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[2147] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[2148] The following is further disclosed regarding the above embodiment.
[2149] (Claim 1)
[2150] means for receiving an audio file from a user;
[2151] means for analyzing lyrics and beat information from the audio file;
[2152] means for receiving from a user selection information regarding photos of characters and the image and atmosphere of the music video;
[2153] means for generating a music video using a generative model based on the analyzed lyrics and beat information and the selection information;
[2154] means for providing the generated music video to a user;
[2155] A system including:
[2156] (Claim 2)
[2157] 2. The system of claim 1, wherein the generative model generates video clips based on user-selected photos of characters and the image and atmosphere of the music video, and renders the video clips as a sequence.
[2158] (Claim 3)
[2159] 2. The system of claim 1, wherein the means for analyzing the audio file comprises an algorithm for extracting beatmaps and lyrics from the audio file.
[2160] "Example 1"
[2161] (Claim 1)
[2162] means for receiving an audio file from a user;
[2163] means for analyzing lyrics and beat information from the audio file;
[2164] means for receiving from a user selection information regarding the image of the characters and the image and atmosphere of the music video;
[2165] means for generating a music video using a generative AI model based on the analyzed lyrics and beat information and the selection information;
[2166] means for providing the generated music video to a user and generating and transmitting a download link;
[2167] means for displaying the progress of the music video generation to a user in real time;
[2168] A system including:
[2169] (Claim 2)
[2170] 2. The system of claim 1, wherein the generative AI model generates video clips based on user-selected character images and the image and atmosphere of the music video, and renders the video clips as a sequence.
[2171] (Claim 3)
[2172] 2. The system of claim 1, wherein the means for analyzing the audio file comprises an audio analysis algorithm for extracting beat maps and lyrics from the audio file.
[2173] "Application Example 1"
[2174] (Claim 1)
[2175] means for receiving an audio file from a user;
[2176] means for analyzing lyrics and beat information from the audio file;
[2177] means for receiving selection information from a user regarding the image of the characters and the image and atmosphere of the video;
[2178] means for generating a music video using a generative model based on the analyzed lyrics and beat information and the selection information;
[2179] means for displaying the progress of the generated music video;
[2180] means for providing the generated music video to a user so that the user can post the music video directly to a content distribution service;
[2181] A system including:
[2182] (Claim 2)
[2183] The system of claim 1, wherein the generative model generates video clips based on user-selected character images and video image / ambience, and renders the video clips as a sequence.
[2184] (Claim 3)
[2185] 2. The system of claim 1, wherein the means for analyzing the audio file comprises an algorithm for extracting beatmaps and lyrics from the audio file.
[2186] "Example 2: Combining Emotion Engines"
[2187] (Claim 1)
[2188] means for receiving an audio file from a user;
[2189] means for analyzing lyrics and beat information from the audio file;
[2190] means for receiving from a user a selection of images of characters and image / atmosphere information for the video of the visual work;
[2191] means for analyzing the audio file and real-time emotional data to identify an emotional state of the user;
[2192] means for generating a visual composition using a generative model based on the analyzed lyric and beat information, the selection information, and the identified emotional state;
[2193] means for providing the generated visual work to a user;
[2194] A system including:
[2195] (Claim 2)
[2196] 2. The s...
Claims
1. means for receiving an audio file from a user; means for analyzing lyrics and beat information from the audio file; means for receiving from a user photos of characters and selection information regarding the image and atmosphere of the music video; means for generating a music video using a generative model based on the analyzed lyrics and beat information and the selection information; means for providing the generated music video to a user; A system including:
2. The system of claim 1 , wherein the generative model generates video clips based on user-selected photos of characters and the image and atmosphere of the music video, and renders the video clips as a sequence.
3. 2. The system of claim 1, wherein the means for analyzing the audio file comprises an algorithm for extracting beat maps and lyrics from the audio file.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A