Video content generation device providing ai-based music arrangement function and method for generating video content thereof
Patent Information
- Application Number
- KR1020250063622
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2045-05-16
Smart Images

Figure 112025054781815-PAT00001_ABST
Abstract
Description
Technology Field
[0001] Embodiments of the present invention relate to a video content generation device and a video content generation method that provide an artificial intelligence-based music arrangement function, and to a video content generation device and a video content generation method that automatically arrange music to match the emotional flow of the video. Background Technology
[0003] Recently, automated video content generation technologies that integrate various modalities such as text, images, video, and audio using multimodal generative AI are being actively developed. In particular, there is an increasing number of attempts to automatically insert background music into videos by utilizing AI-based music generation technology.
[0004] However, since most of these music generation technologies are based on pre-trained, general music generation models, music is frequently generated or inserted that does not match the emotional flow or atmosphere of the video, or does not match the style desired by the user. This results in issues such as a lack of emotional consistency between the video and the music, reduced immersion, and lowered content quality.
[0005] In addition, for a user to arrange or adjust music data to suit the video, professional music arrangement knowledge is required, and separate music arrangement software must be used, which imposes limitations on music arrangement. Prior art literature
[0007] Korean Registered Patent Publication No. 10-2612572, "Audio Data Generation Using Artificial Intelligence" The problem to be solved
[0008] The embodiments of the present invention are intended to provide an AI-based video content generation device and a video content generation method that automatically arrange music to match the emotional flow of a video by analyzing music data input by a user to generate music composition information and adjusting the music composition information by referring to an emotional state graph of the video. means of solving the problem
[0010] A video content generation device providing an AI-based music arrangement function according to an embodiment of the present invention for achieving the above objective comprises: an input unit that receives at least one music data from a user; a music information generation unit that analyzes the music data and generates music composition information along a time axis; a video analysis unit that analyzes visual and auditory elements within a video to generate an emotional state graph that quantitatively expresses the emotional state of the video along a time axis, and dynamically divides the video into a plurality of segments based on the point in time when the emotional state changes in the emotional state graph; a music arrangement unit that matches the music composition information to correspond to the emotional state graph and arranges the music data by adjusting the music composition information for each of the plurality of segments; a music rendering unit that renders the arranged music data and outputs it as an audio file; and a video-music synthesis unit that aligns the audio file with the timeline of the video to generate integrated video content.
[0011] According to an embodiment, the music information generating unit can analyze the music data and generate composition information of the music arranged along a time axis, at least one of genre, melody, rhythm, tempo, structural information, chord progression, and instrument arrangement.
[0012] According to an embodiment, the image analysis unit analyzes the visual element using at least one of camera movement, image brightness, and image color within the image, and analyzes the voice element using at least one of emotion keywords, voice pitch, voice speed, and voice intensity from dialogue voice or narration voice within the image, calculates the emotion type and emotion intensity based on the measurement values for each of the analyzed visual element and voice element, and can generate the emotion state graph by arranging the calculated emotion type and emotion intensity along a time axis.
[0013] According to an embodiment, the image analysis unit determines a continuous section in the emotion state graph where the same emotion type is maintained at a constant emotion intensity as one emotion section, recognizes a point where the emotion type or emotion intensity changes as an emotion turning point, and can dynamically divide the image into multiple emotion sections based on the emotion turning point.
[0014] According to an embodiment, the music arrangement unit can automatically arrange the music data by re-dividing the composition information of the music to correspond to the plurality of sections dynamically divided in the emotional state graph, and adjusting at least one of the genre, melody, beat, tempo, structural information, chord progression, and instrument arrangement for each re-divided section of the composition information of the music.
[0015] According to an embodiment, when the playback time of the video is shorter than the playback time of the music data, the music arrangement unit may select one or more music segments within the music data that emotionally match the emotional state of the video based on the emotional state graph, and arrange the music data by reconstructing the selected music segments to match the playback time of the video.
[0016] According to an embodiment, if the playback time of the video is longer than the music data, the music arrangement unit may repeat a portion of the music data or automatically generate music through a pre-trained music generation model to expand and arrange the music to fit the playback time of the video.
[0017] Meanwhile, a video content generation method providing an AI-based music arrangement function according to an embodiment of the present invention comprises the steps of: receiving at least one music data input from a user; analyzing the music data to generate composition information of the music along a time axis; analyzing visual and auditory elements within the video to generate an emotional state graph that quantitatively expresses the emotional state of the video along a time axis; dynamically dividing the video into a plurality of segments based on the point in time when the emotional state changes in the emotional state graph; arranging the music data by matching the composition information of the music to correspond to the emotional state graph and adjusting the composition information of the music for each of the plurality of segments; rendering the arranged music data to output it as an audio file; and aligning the audio file with the timeline of the video to generate integrated video content. Effects of the invention
[0019] According to embodiments of the present invention having the above-described configuration, music data input by a user can be automatically arranged to match the emotional flow of the video, thereby improving emotional consistency and immersion between the video content and the music.
[0020] In addition, video-related creators, such as video producers, content creators, advertising planners, and educational video producers, can generate customized background music for videos using their desired music data, even if they lack professional knowledge of music arrangement or cannot use music arrangement software. Brief explanation of the drawing
[0022] FIG. 1 is a block diagram showing the configuration of a video content generation device that provides an artificial intelligence-based music arrangement function according to an embodiment of the present invention. FIG. 2 is a block diagram showing the configuration of a video content generation device that provides an artificial intelligence-based music arrangement function according to another embodiment of the present invention. Figure 3 is a drawing showing an emotional state graph according to an embodiment of the present invention. FIG. 4 is a flowchart illustrating a method for generating video content that provides an artificial intelligence-based music arrangement function according to an embodiment of the present invention. Specific details for implementing the invention
[0023] Specific structural or functional descriptions of embodiments according to the concept of the present invention disclosed herein are provided merely for the purpose of explaining embodiments according to the concept of the present invention, and embodiments according to the concept of the present invention may be implemented in various forms and are not limited to the embodiments described herein.
[0024] Embodiments according to the concept of the present invention may be subject to various modifications and may take various forms; therefore, embodiments are illustrated in the drawings and described in detail in this specification. However, this is not intended to limit the embodiments according to the concept of the present invention to specific disclosed forms, and includes modifications, equivalents, or substitutions that fall within the spirit and scope of the present invention.
[0025] Terms such as "first" or "second" may be used to describe various components, but said components should not be limited by said terms. For the sole purpose of distinguishing one component from another, for example, without departing from the scope of rights according to the concept of the present invention, the first component may be named the second component, and similarly, the second component may be named the first component.
[0026] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. Conversely, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between. Expressions describing the relationships between components, such as "between," "exactly between," or "directly adjacent to," should be interpreted in the same way.
[0027] The terms used herein are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as “comprising” or “having” are intended to specify the existence of the described features, numbers, steps, actions, components, parts, or combinations thereof, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0028] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the present invention pertains. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this specification.
[0029] Hereinafter, embodiments will be described in detail with reference to the attached drawings. However, the scope of the patent application is not limited or restricted by these embodiments. Identical reference numerals in each drawing indicate identical components.
[0031] FIG. 1 is a block diagram showing the configuration of a video content generation device that provides an artificial intelligence-based music arrangement function according to an embodiment of the present invention. The video content generation device (100) according to FIG. 1 includes an input unit (110), a music information generation unit (120), a video analysis unit (130), a music arrangement unit (140), a music rendering unit (150), and a video-music synthesis unit (160).
[0032] As used in this specification, "video" refers to original video data that includes visual and audio elements such as dialogue, narration, and scene transitions, but without background music or sound effects.
[0033] In addition, "video content" refers to a completed video product provided to the user, which is the result of aligning, synthesizing, and synthesizing audio files generated through the music arrangement technology according to the present invention with a video.
[0034] The input unit (110) receives at least one piece of music data from the user.
[0035] Music data may be musical score data such as MIDI or MusicXML, or may include audio files such as WAV or MP3. When a user wishes to insert music into a video, they may input audio sources of music they have performed or musical score data they have composed, or they may input audio sources or musical score data they have personally selected.
[0036] The music information generation unit (120) analyzes music data and generates compositional information of music along the time axis. Specifically, before executing music editing, the music information generation unit (120) analyzes sheet music data and generates compositional information of music including genre, melody, beat, tempo, structural information, chord progression, instrument arrangement, etc.
[0037] The video analysis unit (130) analyzes visual and audio elements within the video and generates an emotional state graph that quantitatively expresses the emotional state of the video along the time axis. Here, the 'video' is generated through a video generation unit (not shown) within the video content generation device (100), and may be a scene video (scene video) generated by a prompt, but is not limited thereto and may be a sequence video composed of multiple scene videos.
[0038] The video analysis unit (130) dynamically divides the video into multiple segments based on the point in time when an emotional change occurs in the emotional state graph. For example, in a 1-minute video, if an emotional change occurs in the segments of 0 to 10 seconds, 10 to 25 seconds, 25 to 55 seconds, and 55 to 60 seconds, the video can be dynamically divided into 4 segments.
[0039] The music arrangement unit (140) matches music components to the video and arranges music data by adjusting music composition information for multiple sections. A detailed explanation of the music data arrangement will be provided later in FIG. 2.
[0040] The music rendering unit (160) interprets and plays arranged music data through a virtual instrument or sound sampler, renders it, converts it into an actual listenable audio file (WAV, MP3, etc.), and outputs it.
[0041] The video-music synthesis unit (170) aligns and synchronizes the rendered audio file and video data along the time axis and integrates them to synthesize them into a single video content. At this time, the playback timing of the audio is precisely aligned to match the scene progression and emotional flow of the video, and finally, an integrated video content is created in which the video and music are harmoniously combined.
[0042] According to the video content generating device (100) of Fig. 1, music data input by a user can be arranged into music suitable for the video atmosphere without the need to use separate music arrangement software.
[0044] FIG. 2 is a block diagram showing the configuration of a video content generation device that provides an artificial intelligence-based music arrangement function according to another embodiment of the present invention. The video content generation device (200) according to FIG. 2 includes an input unit (210), a video generation unit (220), a music information generation unit (230), a video analysis unit (240), a music arrangement unit (250), a music rendering unit (260), a video-music synthesis unit (270), and a control unit (280). Here, the control unit (280) can control the overall operation of the video content generation device (200).
[0045] The input unit (210) provides a user interface function for receiving input data required for video generation and music arrangement from a user, and includes a first input unit (211) and a second input unit (212).
[0046] The first input unit (211) is configured to receive a text prompt or command for video generation, and the second input unit (212) is configured to receive music data such as sheet music data (MIDI, MusicXML, etc.) or sound source files (WAV, MP3, etc.), and the input music data is subsequently used as basic data for music information analysis and arrangement processing.
[0047] The control unit (280) can control the operation of the image generation unit (220) or the music information generation unit (230) according to commands, data, etc. input from the first input unit (211) or the second input unit (212).
[0048] The video generation unit (220) generates a video by automatically configuring audiovisual elements required for each scene according to a prompt input through the first input unit (211). Specifically, the video generation unit (220) analyzes the prompt input in natural language to identify the components of the scene, and generates a video for each scene (scene video) by automatically generating and combining various data such as images, video clips, dialogue, and narration (TTS: Text-to-Speech) suitable for the scene through an AI-based multimodal generation model. In this process, the control unit (280) receives the video generation results for each scene in real time to monitor the video generation status, and can control the video generation unit (220) to regenerate the video if an abnormality occurs, or control it to provide a user notification through a display.
[0049] Although not illustrated in the drawings, the video content generation device (200) may further include a music generation unit (not shown) that additionally generates music appropriate for each scene of the video. In the embodiments, the generation of music based on music data input by a user through the second input unit (212) is processed first, and thereafter, when a separate music generation command is transmitted from the control unit (280), the music generation unit may additionally proceed with the music generation operation.
[0050] The music information generation unit (230) analyzes the music data input through the second input unit (212) and generates composition information of the music along the time axis. Here, the music data may be input as sheet music data (MIDI, MusicXML, etc.) or sound source files (WAV, MP3, etc.), and it is also possible to input multiple music data simultaneously.
[0051] The music information generation unit (230) may include a music data conversion unit (not shown) for data preprocessing depending on the form of the music data. When sheet music data is input, the music information generation unit (230) can directly analyze the sheet music data, and when an audio file is input, the audio file can be converted into sheet music data through the music data conversion unit.
[0052] The music information generation unit (230) can generate music composition information by analyzing music data and arranging at least one of genre, melody, beat, tempo, structural information (intro, verse, chorus, etc.), chord progression, and instrument arrangement along a time axis. At this time, the music data can be divided into multiple time segments based on a point in time where the music composition information (melody, beat, tempo, structural information, chord progression, instrument arrangement, etc.) changes significantly along the time axis.
[0053] The music information generation unit (230) classifies the genre of the music data by analyzing the combination pattern of each characteristic based on the melody, beat, tempo, chord progression, and instrument arrangement information extracted from the sheet music data. At this time, the genre classification can be performed by comparing with a predefined genre-specific feature database or through a machine learning-based genre classification model.
[0054] The music information generation unit (230) can extract major notes (e.g., C4, E4, G4, etc.) from the musical score data and analyze beats based on measure division and timing information from the musical score data.
[0055] The music information generation unit (230) can extract the tempo by first applying the value when the score data includes a metronome marking (e.g., ♩=72), and when there is no metronome marking, it can extract the tempo by calculating the beats per minute (BPM) based on the time signature, the time length of the notes included in each measure, and the beats.
[0056] The music information generation unit (230) can be divided into a development structure such as an introduction (intro), verse, climax, chorus, bridge, and outro based on a melody pattern, chord progression, and rhythm composition that are repeated in each time interval.
[0057] The music information generation unit (230) can generate an entire code progression sequence by analyzing the combination of notes included in each time interval to extract a major code and arranging them in chronological order.
[0058] The music information generation unit (230) can extract instrument arrangement information used in each time interval based on instrument-specific track or part information (e.g., Piano, Strings, Drum, etc.) included in the score data. When the score data is MusicXML <score-part>You can identify the type of instrument by checking the tag or the program number for each MIDI channel.
[0059] In addition, the items included in the composition information of the music are not limited to genre, melody, rhythm, tempo (speed), structural information, chords, and instrumentation, and may also include dynamics (volume change), note density, etc., and if the music data includes lyrics, lyric information may also be included.
[0060] Table 1 shows an example of composition information of music generated by analyzing predetermined music data.
[0061] division 1st time interval (0~30 seconds) 2nd time interval (30~60 seconds) 3rd time interval (60~90 seconds) Genre ballade ballade ballade melody C4-E4-G4-F4-D4… G4-F4-E4-D4-C4… A4-C5-B4-G4-F4… beat 4 / 4 4 / 4 4 / 4 Tempo 70bpm 72bpm 75bpm Structural Information Intro Verse Chorus Code progression CG-Am-F Dm-GCF Am-FCG Instrumentation piano Piano, string pads Piano, strings, drums
[0062] According to Table 1, the music information generation unit (230) divided the sheet music data for 1 minute and 30 seconds into three time intervals along the time axis and extracted composition information of the music for each time interval. For convenience of explanation, the music data was divided into 30-second intervals, but in reality, it may be divided into irregular time intervals.
[0063] The video analysis unit (240) analyzes visual and audio elements within the video to generate an emotional state graph that quantifies the emotional state of the video along the time axis. Specifically, the video analysis unit (240) analyzes visual elements such as camera movement, video color, and video brightness within the video, and can analyze audio elements by extracting emotional keywords, voice pitch, voice speed, and voice intensity from dialogue voice or narration voice within the video.
[0064] By using the visual and auditory elements analyzed in this way, it is possible to identify the sections in the video where an emotional state (e.g., sadness, joy, anger, etc.) is maintained, and to generate an emotional state graph that quantifies the emotional state of the corresponding section into a value between 0.1 and 1.0.
[0065] Table 2 shows an example of items and reference values for the video analysis unit (250) to analyze visual and audio elements within the video.
[0066] division item Unit of measurement Reference value Visual elements Camera movement (Distance of main objects moved between frames (px / frame) Slow : 0~10px / frame Normal : 11~15px / frame Fast : 16px / frame< Video colors (Hue, 0˚~360˚) 30˚ or less: Red~Orange 30˚~90˚: Orange~Light Green 90˚~150˚: Light Green~Teal 150˚~180˚: Teal~Blue 180˚~300˚: Blue~Purple 300˚~330˚: Magenta~Pink 330˚~360˚: Magenta~Red Video brightness (Value, 0~255) - 200 or more (Bright) - 80 or less (Dark) Voice elements Emotion keywords 1. Joy (Joyful, funny, happy, fun, good, etc.) 2. Sadness (Sad, empty, lonely, etc.) 3. Anger (Annoyed, pissed off, furious, etc.) 4. Fear (Help me, scared, run away, etc.) 5. Surprise (Really?, That's impossible!, I can't believe it!, etc.) Sentiment Keyword Extraction Voice pitch (Pitch, Hz) Low: <150Hz Medium: 150Hz~200Hz High: 200Hz< Voice speed (Speaking speed, WPM) Slow: <100WPM Normal: 100~180WPM Fast: 180WPM< Voice intensity (Volume, dB) Standard: 70dB (higher indicates higher emotional intensity)
[0067] The items shown in Table 2 are examples, and additional items for analyzing visual and audio elements within the video may be added or changed.
[0068] The video analysis unit (240) can analyze visual and audio elements within the video to generate an emotional state graph that quantitatively expresses the emotional state for each time interval.
[0069] First, the video analysis unit (240) analyzes visual elements and audio elements within the video in units of a certain time interval or video frames. Visual elements include camera movement, video color (Hue), video brightness (Value), etc., and audio elements include voice pitch (F0), voice rate (WPM), voice intensity (dB), emotion keywords (based on an emotion dictionary), etc.
[0070] The video analysis unit (240) can determine the emotion type by comparing the measured values of each visual element and voice element with an emotion standard table. The emotion standard table consists of reference data that specifies the corresponding emotion type when the measured value of each element exceeds a specific threshold or falls within a specific range.
[0071] For example, it can be configured such that if the camera movement speed value is 20px / frame or higher, or the voice pitch (F0) is 250Hz or higher, it is classified as 'anger', and if the hue is 30˚ or lower and the voice speed is 180WPM or higher, it is classified as 'hope' or 'excitement'.
[0072] Once the emotion type is determined through the measurements of each visual and auditory element, the most frequently derived emotion type can be designated as the representative emotion type for that section or video frame. However, this is not limited to this, and the emotion type associated with a specific element (e.g., emotion keyword) may also be designated as the representative emotion type.
[0073] When the video analysis unit (240) determines the type of emotion, it can normalize and quantify the emotional intensity for that type of emotion.
[0074] Emotional intensity can be calculated for each element by applying weights to the normalized values of the individual visual and auditory measurements, and the emotional intensity for a specific video frame or segment can be calculated by summing the emotional intensities for each element. Note that weights can be set differently for each element.
[0075] For example, if the speed value of the camera movement is 12px / frame and the reference value (maximum value) of the camera movement is 20px / frame, the normalized value for the camera movement element is calculated as "12 / 20" and yields 0.6, and if a weight of "0.15" is applied to the camera movement element, it is calculated as "0.6X0.15" and yields a value of 0.09. In this way, by calculating the emotional intensity for all other visual and audio elements and summing them up, the emotional intensity for the corresponding video frame or corresponding video segment can be calculated.
[0076] As a result, the video analysis unit (240) generates an emotional state graph containing emotional types (e.g., calm, sad, hope, etc.) and corresponding emotional intensities (e.g., 0.2, 0.8, etc.) for each time interval, which can then be provided as data for interval division and music arrangement.
[0077] The emotional state graph visually represents the flow of emotions throughout the video along the time axis, and the video analysis unit (240) determines a continuous section in which the same emotional type and emotional intensity are maintained as a single emotional section, recognizes the point at which the emotional type or emotional intensity changes as an emotional turning point, and can dynamically divide the entire video into multiple emotional sections based on this. The emotional sections divided in this way are subsequently used as reference information for the music arrangement unit to perform music arrangements corresponding to the emotional state of each section.
[0078] FIG. 3 is a diagram showing an emotional state graph according to an embodiment of the present invention. Referring to FIG. 3, the emotional state graph is a graph that quantitatively expresses the type of emotion and the intensity of emotion for each segment along the time axis for a video of 1 minute and 30 seconds. Segments where the type of emotion and the intensity of emotion remain the same or similar are displayed as a flat graph, while points where a sudden change in the intensity of emotion or a transition in the type of emotion occurs are displayed as a stepped structure of the graph.
[0079] In Figure 3, for the convenience of explanation, the emotional state is depicted as being maintained constant in each section, but if the emotional intensity differs even for the same type of emotion, the emotional state graph may be displayed as a straight line or curve with a slope.
[0080] In the first section (0–22 seconds), the "calm" emotion type is maintained at an intensity of 0.3, and in the second section (22–45 seconds), the "tension" emotion type is maintained at an intensity of 0.6. Additionally, in the third section (45–70 seconds), the "anger" emotion type is maintained at an intensity of 0.9, and in the fourth section (70–90 seconds), the "anxiety" emotion type is maintained at an intensity of 0.7.
[0081] The video analysis unit (240) determines the point at which a change in emotional intensity or emotional type is detected in the emotional state graph as an emotional turning point, and dynamically divides the video into multiple sections based on this turning point. For a video having an emotional state graph like that of FIG. 3, since a total of 4 emotional turning points are detected, the video can be divided into a total of 4 emotional sections from the first section to the fourth section, and each section is utilized so that a music element suitable for the emotional change corresponds to the music arrangement.
[0082] The music arrangement unit (250) divides the composition information of the music generated by the music information generation unit (230) into time intervals based on the emotional state graph received from the video analysis unit (240), and automatically arranges the music data to correspond to the emotional state of each interval.
[0083] More specifically, the music arrangement unit (250) re-divides the composition information of the music based on the emotional sections (e.g., sections 1 through 4) separated in the emotional state graph, and adjusts at least one of the genre, tempo, rhythm, chord progression, and instrument arrangement of the music to correspond to the type of emotion and emotional intensity in each section.
[0084] Table 3 shows the composition information of the music in the arranged state, which was divided into four time intervals based on the emotional state graph of Figure 3 and the composition information of the music shown in Table 1.
[0085] division First time interval (0~22 seconds) 2nd time interval (22~45 seconds) 3rd time interval (45~70 seconds) 4th hour interval (70~90 seconds) Genre ballade ballade rock ballad Modern ballad melody C4-E4-G4… F4-D4-G4… A4-C5-B4… G4-F4-E4… beat 4 / 4 4 / 4 4 / 4 4 / 4 Tempo 70bpm 76bpm 85bpm 78bpm Structural Information Intro Verse (initial) Verse (second half) + introduction of climax Climax latter half + Outro Code progression CG-Am-F Dm-GCF Am-FCG G-Em-Am-D Instrumentation Piano, string pads Piano, string pads Piano, strings, drums, brass Piano, strings
[0086] In the section classified as 'anger' with high emotional intensity on the emotional state graph (e.g., section 3, 45–70 seconds), the tempo is increased from the existing 72 bpm to 85 bpm, dynamic instruments such as drums and brass are added to the instrumentation to provide a strong rhythmic sense, and the chord progression is changed to harmonize major and minor keys suitable for emotional heightening to induce the maximization of emotion.
[0087] On the other hand, in the 'calm' state (e.g., Section 1, 0–22 seconds), the tempo is maintained, and the music is arranged with emotional instruments such as a piano or string pads to induce emotional alignment with the atmosphere of the video.
[0088] At this time, the music arrangement department (250) uses an AI-based automatic arrangement model, such as a Transformer or VAE-based sequence-based model, to reconstruct elements such as genre, beat, tempo, structural information, chord progression, and instrument arrangement in combination while maintaining the main pattern of the melody for each section.
[0089] In addition, when the emotional type changes rapidly between sections or when the tempo and genre differ, transition music can be generated and inserted to maintain a sense of musical continuity and minimize the sense of musical discontinuity. This transition music can be automatically generated by mixing the musical attributes (tempo, chords, instrumentation, etc.) of both sections.
[0090] In Figure 2, for the convenience of explanation, an example was given in which the music data and video input by the user have the same playback time (1 minute 30 seconds), but in an actual environment, the playback times of the music data and video content may differ from each other.
[0091] Accordingly, the music arrangement unit (250) can ensure consistency between the music and the video by comparing the playback time of the video and the playback time of the music data, and performing an appropriate arrangement and length adjustment process when they differ.
[0092] First, if the playback time of the video is shorter than the playback time of the music data, the music arrangement unit (250) can select sections within the music data that emotionally match the emotional flow of the video based on an emotional state graph input from the video analysis unit (240). The selected music sections are filtered according to the type of emotion and the intensity of the emotion, and the sections can be clipped to match the video playback time or reconstructed by sequencing to arrange the music data to suit the length of the video.
[0093] Conversely, if the playback time of the video is longer than the playback time of the music data, the music arrangement unit (250) may expand the music data in one or more of the following ways to maintain the emotional consistency and emotional flow of the music.
[0094] First, missing time segments can be supplemented by repeatedly playing segments within the music data that have emotional attributes identical or similar to the emotional state graph, or second, by using a pre-trained AI-based music generation model to automatically generate new music segments that match the tempo, chord progression, rhythm, and instrument composition of the existing music.
[0095] Accordingly, the music arrangement unit (250) supports the creation of video content that provides emotional consistency and immersion by flexibly rearranging the input music data to match the playback time and emotional flow of the video content.
[0096] The music rendering unit (260) can render the music data arranged through the music arrangement unit (250) using a virtual instrument or a sampler-based sound synthesizer and output it as an audio file in WAV, MP3, or MIDI format.
[0097] The music rendering unit (260) performs the role of converting the instrument arrangement, chord progression, tempo, and other musical elements configured for each section into actual sound sources based on the music data arranged by the music arrangement unit (250).
[0098] Specifically, the music rendering unit (260) utilizes a virtual instrument (VSTi: Virtual Studio Technology Instrument) or a sampler-based sound synthesis engine to render input music data into an audio signal and can generate a rendering result according to the output format (WAV, MP3, MIDI) set by the user. Here, WAV and MP3 are output in the form of completed audio files that can be immediately inserted into a video, while MIDI is provided in the form of sequence information for subsequent editing.
[0099] Additionally, the music rendering unit (260) is controlled by the control unit (280) for the rendering target section, output format, instrument library used, and whether rendering is complete, and can monitor the rendering status in real time and transmit it to the control unit (280).
[0100] Meanwhile, the video-music synthesis unit (270) can synthesize the video generated by the video generation unit (220) and the audio file output from the music rendering unit (260) into integrated video content by aligning them along the time axis. At this time, the playback timing of the two data is precisely matched based on the timecode of the video and the timeline sequence of the audio file, and synchronization between the emotional segments and the music segments is ensured to enable consistent emotional expression.
[0101] The video-music synthesis unit (270) is also controlled by the control unit (280), and the control unit (280) can automatically detect and correct whether there is a sync error between the video and the music, and can control the entire synthesis process including setting the format (MP4, MOV, etc.) of the final output video and generating a preview video.
[0102] According to the video content generating device (100) of FIG. 2, music optimized for the emotional state of the video is automatically arranged based on music data provided by the user, and integrated video content can be automatically generated by precisely synchronizing the music with the emotional flow of the video.
[0103] In particular, through segment division based on emotional state graphs and adjustment of music composition information, customized music arrangements that match the emotional flow of the video are possible, thereby maximizing viewer emotional immersion and the completeness of the content.
[0104] In addition, since the video content generation device (100) automatically performs the music arrangement process based on artificial intelligence, video producers, creators, advertising planners, educational video producers, etc., can automatically generate background music suitable for the mood and emotion of the video by utilizing their own music data without needing specialized knowledge of music arrangement or separate arrangement software.
[0105] FIG. 4 is a flowchart illustrating a method for generating video content that provides an artificial intelligence-based music arrangement function according to an embodiment of the present invention. The present method may be executed by a video content generating device illustrated in FIG. 1 or FIG. 2, or implemented as a program executable on a computer and stored on a computer-readable recording medium.
[0106] First, the video content generation method receives at least one piece of music data in the form of a musical score (MIDI, MusicXML, etc.) or a sound source (WAV, MP3, etc.) from the user (S410).
[0107] Afterwards, input music data is analyzed to generate composition information of music along the time axis (S420). Here, the composition information of music may include genre, melody, beat, tempo (speed), structural information, chord progression, instrument arrangement, etc.
[0108] A method for generating video content analyzes visual and audio elements within a video to generate an emotional state graph that quantitatively expresses the emotional state of the video along a time axis (S430). Here, the visual elements are analyzed through camera movement, video color, video brightness, etc. within the video, and the audio elements can be analyzed by extracting emotional keywords, voice pitch, voice speed, and voice intensity from dialogue or narration voice within the video.
[0109] When an emotional state graph is generated, the video is dynamically divided into multiple segments based on the point in time when the emotional state changes in the video (S440). At this time, if the same emotion persists for a certain period of time or longer, it is defined as a single segment, and the point in time when the emotion changes is determined as an emotional turning point.
[0110] Then, the video is dynamically divided into multiple segments based on the point in time when the emotional state changes in the emotional state graph (S440).
[0111] A method for generating video content involves matching music composition information to correspond to an emotional state graph and arranging music data by adjusting the music composition information for multiple sections (S450). For each section, at least one element such as melody, genre, beat, tempo, structural information, chord progression, and instrument arrangement can be adjusted in the music composition information to arrange the music data to suit the emotional state of the corresponding section.
[0112] The arranged music data is rendered into an audio file (WAV, MP3, etc.) using a virtual instrument or a sampler-based sound synthesizer (S460).
[0113] Finally, the rendered audio file is precisely aligned with the video timeline to create integrated video content that matches the emotional flow (S470).
[0114] According to this method, music optimized for the emotional flow can be automatically generated and arranged based on music data provided by the user, making it possible to produce high-quality content that matches the mood and emotional flow of the video without professional knowledge of music arrangement.
[0116] The steps of the method or algorithm described in connection with embodiments of the present invention may be implemented directly in hardware, implemented as a software module executed by hardware, or implemented by a combination thereof. The software module may reside in RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), Flash Memory, a hard disk, a removable disk, a CD-ROM, or any form of computer-readable recording medium well known in the art to which the present invention belongs.
[0117] Although embodiments of the present invention have been described above with reference to the attached drawings, those skilled in the art will understand that the present invention may be implemented in other specific forms without altering its technical concept or essential features. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. Explanation of the symbols
[0119] 110 : Input section 120 : Music Information Generation Unit 130 : Video Analysis Department 140 : Music Arrangement Department
Claims
Claim 1 An AI-based video content generation device comprising: an input unit that receives at least one piece of music data from a user; a music information generation unit that analyzes the music data and generates composition information of music along a time axis; a video analysis unit that analyzes visual and auditory elements within a video to generate an emotional state graph that quantitatively expresses the emotional state of the video along a time axis, and dynamically divides the video into multiple segments based on the point in time when the emotional state changes in the emotional state graph; a music arrangement unit that matches the composition information of the music to correspond to the emotional state graph and arranges the music data by adjusting the composition information of the music for each of the multiple segments; a music rendering unit that renders the arranged music data and outputs it as an audio file; and a video-music synthesis unit that aligns the audio file with the timeline of the video to generate integrated video content, wherein the music arrangement unit, when the playback time of the video is shorter than the playback time of the music data, selects one or more music segments within the music data that emotionally match the emotional state of the video based on the emotional state graph, and arranges the music data by reconstructing the selected music segments to match the playback time of the video. Claim 2 A video content generating device that provides an artificial intelligence-based music arrangement function, wherein the music information generating unit analyzes the music data to generate composition information of the music arranged along a time axis, at least one of genre, melody, rhythm, tempo, structural information, chord progression, and instrument arrangement. Claim 3 A video content generation device providing an AI-based music arrangement function according to claim 1, wherein the video analysis unit analyzes the visual element using at least one of camera movement, video brightness, and video color within the video, analyzes the voice element using at least one of emotion keywords, voice pitch, voice speed, and voice intensity from dialogue voice or narration voice within the video, calculates the emotion type and emotion intensity based on the measurement values for each of the analyzed visual element and voice element, and generates the emotion state graph by arranging the calculated emotion type and emotion intensity along a time axis. Claim 4 A video content generation device providing an AI-based music arrangement function, wherein, in paragraph 3, the video analysis unit determines a continuous section in which the same emotion type is maintained at a constant emotion intensity in the emotion state graph as a single emotion section, recognizes a point in time when the emotion type or emotion intensity changes as an emotion turning point, and dynamically divides the video into multiple emotion sections based on the emotion turning point. Claim 5 A video content generation device according to claim 1, wherein the music arrangement unit re-divides the composition information of the music to correspond to the plurality of sections dynamically divided in the emotional state graph, and automatically arranges the music data by adjusting at least one of genre, melody, beat, tempo, structural information, chord progression, and instrument arrangement for each re-divided section of the composition information of the music. Claim 6 delete Claim 7 A video content generation device according to claim 1, wherein the music arrangement unit provides an artificial intelligence-based music arrangement function that, when the playback time of the video is longer than the music data, repeats a portion of the music data or automatically generates music through a pre-trained music generation model and expands and arranges it to fit the playback time of the video. Claim 8 A method for generating video content that provides an AI-based music arrangement function, comprising: receiving at least one music data input from a user; analyzing the music data to generate composition information of the music along a time axis; analyzing visual and auditory elements within a video to generate an emotional state graph that quantitatively expresses the emotional state of the video along a time axis; dynamically dividing the video into multiple segments based on the point in time when the emotional state changes in the emotional state graph; arranging the music data by matching the composition information of the music to correspond to the emotional state graph and adjusting the composition information of the music for each of the multiple segments; rendering the arranged music data to output it as an audio file; and aligning the audio file with the timeline of the video to generate integrated video content, wherein the step of arranging the music data includes, when the playback time of the video is shorter than the playback time of the music data, selecting one or more music segments within the music data that emotionally match the emotional state of the video based on the emotional state graph, and reconstructing the selected music segments to match the playback time of the video to arrange the music data.
Citation Information
Patent Citations
Intelligent system for matching audio with video
US20230015498A1