Vlog generation methods and related devices
By acquiring video footage and emotional information, and using neural network processing to generate personalized background music, the problem of low music-video matching in vlogs is solved. This achieves matching between music and video emotions and transition points, thus improving the user experience.
Patent Information
- Application Number
- CN202210562029.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-05-23
AI Technical Summary
In existing technologies, it is difficult for users to personalize their background music selection based on the video content when shooting vlogs, resulting in low matching between music and video and failing to meet users' personalized needs.
By acquiring video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals, a neural network is used to generate personalized background music that matches the video transition points and emotions. This includes track splitting and transcription processing to obtain MIDI scores and phrase/section segmentation points, and adjusting the music to meet user needs.
It enables personalized matching of background music and video content, enhances the vlog creation experience, ensures that the mood of the music is consistent with the mood of the video, and that transition points correspond, thus solving the problem of insufficient music recommendation in existing technologies.
Smart Images

Figure CN117156173B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing, and in particular to a method for generating Vlogs and related apparatus. Background Technology
[0002] Creating vlogs to document their travels and lives has become an important lifestyle for many young people. When making a vlog, users often rack their brains trying to choose music that matches their current mood and the content of the video. Ordinary users often lack sufficient music knowledge; they may know what type of songs they like or what songs they want to listen to, but they don't know how to match those songs with the video.
[0003] Short video apps primarily recommend existing music, directly suggesting trending songs based on user-selected tags and song popularity. Music recommended using this approach often has a low match with the user's personalized video content, providing a generic experience and failing to create a personalized experience. It also fails to meet users' personalized needs to use specific songs in their vlogs. Summary of the Invention
[0004] This application provides a Vlog generation method and related apparatus. Using this application, personalized background music and Vlogs can be generated based on user needs, and the transition points of the background music match the transition points of the Vlog.
[0005] In a first aspect, embodiments of this application provide a Vlog generation method, including:
[0006] The process involves: acquiring video footage, the target Vlog's duration, reference transition points, video mood information, and reference audio signals; determining the edited video footage, target Vlog's transition points, and reference background music information based on these elements; performing track splitting and transcription on the reference audio signal to obtain its MIDI score and phrase / section segmentation points; processing the reference audio signal based on its transition points, reference background music information, MIDI score, and phrase / section segmentation points to obtain the target Vlog's background music, ensuring that the background music information matches the reference background music information and that the transition points of the background music match the transition points of the target Vlog; and finally, obtaining the target Vlog based on the edited video footage and its background music.
[0007] Among these, the video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals are selected by the user.
[0008] Background music for the target Vlog is generated based on the user's selected video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals. This ensures that the generated background music meets the user's personalized needs and satisfies the user's personalized desire to use specific music in the Vlog.
[0009] In one possible embodiment, the reference background music information includes at least one of the following: music meta information, musical form, harmonic progression, musical mood, and phrase / section syncopation points; the background music information of the target Vlog is matched with the reference background music information, including:
[0010] The background music meta information of the target vlog matches the background music meta information of the referenced background music, and / or,
[0011] The musical form of the background music in the target vlog matches the musical form included in the reference background music information, and / or...
[0012] The harmonic progression of the background music in the target Vlog matches the harmonic progression included in the reference background music information, and / or...
[0013] The background music in the target vlog matches the emotional tone of the music in the reference background music information, and / or...
[0014] The phrase and segment breakpoints of the background music in the target Vlog match the phrase and segment breakpoints included in the reference background music information.
[0015] In one possible embodiment, the reference audio signal is processed based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music for the target Vlog, including:
[0016] Based on the MIDI score of the reference audio signal, determine the music meta information, form, harmonic progression, and musical mood of the reference audio signal; determine whether the music meta information, form, harmonic progression, musical mood, and phrase / section breakpoints of the reference audio signal match the reference background music information, including music meta information, form, harmonic progression, musical mood, and phrase / section breakpoints.
[0017] When there are parts in the reference audio signal that do not match the reference background music information, including its meta information, form, harmonic progression, emotional tone, and phrase / section breakpoints, the reference audio signal is modified to obtain the background music for the target Vlog.
[0018] The reference audio signal is processed in terms of music meta information, musical form, harmonic progression, and musical emotional direction, so that the background music of the target Vlog obtained based on the reference audio signal meets the user's personalized needs.
[0019] In one possible embodiment, the edited video footage, the transition points of the target Vlog, and the reference background music information are determined based on the video footage, the duration of the target Vlog, reference transition points, and video mood information, including:
[0020] The video footage, the duration of the target Vlog, reference transition points, and video emotional information are input into a pre-trained neural network for processing to obtain the edited video footage, the transition points of the target Vlog, and reference background music information.
[0021] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video emotional information, the method of this application further includes:
[0022] After detecting the user's first modification command, the transition points of the target Vlog are modified to obtain the modified transition points of the target Vlog.
[0023] The reference audio signal is processed based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music for the target Vlog, including:
[0024] The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the modified target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0025] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video emotional information, the method of this application further includes:
[0026] After detecting the user's second modification instruction, the reference background music information is modified to obtain the modified reference background music information;
[0027] The reference audio signal is processed based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music for the target Vlog, including:
[0028] The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0029] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video emotional information, the method of this application further includes:
[0030] After detecting the user's first and second modification commands, the transition click reference background music information of the target Vlog is modified to obtain the modified transition points and modified reference background music information of the target Vlog.
[0031] The reference audio signal is processed based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music for the target Vlog, including:
[0032] The background music for the target Vlog is obtained by processing the reference audio signal based on the modified transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0033] To avoid generating background music for the target Vlog that is not what the user wants, after obtaining the transition points and reference background music information of the target Vlog, it checks whether the user needs to modify them. If it is detected that the user needs to modify the transition points and / or reference background music information of the target Vlog, the transition points and / or reference background music information of the target Vlog are modified to obtain the modified transition points and / or reference background music information of the target Vlog. Then, based on the modified transition points and / or reference background music information of the target Vlog and other information, the background music of the target Vlog is obtained, further ensuring that the background music of the target Vlog meets the user's personalized needs.
[0034] In one possible embodiment, the reference audio signal is split into tracks and transcribed to obtain the MIDI score of the reference audio signal and the phrase / section segmentation points of the reference audio signal, including:
[0035] The reference audio signal is input into a pre-trained track-splitting neural network for processing to obtain a multi-track audio signal; the multi-track audio signal is then input into a pre-trained transcription neural network for processing to obtain the MIDI score and phrase / section segmentation points corresponding to the multi-track audio signal.
[0036] The MIDI score of the reference audio signal includes the MIDI score corresponding to the multi-track audio signal, and the phrase / section segmentation points of the reference audio signal include the phrase / section segmentation points corresponding to the multi-track audio signal.
[0037] Secondly, embodiments of this application provide a video generation apparatus, comprising:
[0038] The acquisition unit is used to acquire video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals;
[0039] The determination unit is used to determine the edited video footage, the target Vlog's transition points, and reference background music information based on the video footage, the target Vlog's duration, reference transition points, and video emotional information.
[0040] The track-splitting transcription unit is used to split and transcribe the reference audio signal to obtain the MIDI score of the reference audio signal and the phrase / section segmentation points of the reference audio signal.
[0041] The processing unit is used to process the reference audio signal according to the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase segmentation points of the reference audio signal to obtain the background music of the target Vlog. The background music information of the target Vlog matches the reference background music information, and the transition points of the background music of the target Vlog match the transition points of the target Vlog. The target Vlog is obtained based on the edited video footage and the background music of the target Vlog.
[0042] In one possible embodiment, the reference background music information includes at least one of the following: music meta information, musical form, harmonic progression, musical mood, and phrase / section syncopation points; the background music information of the target Vlog is matched with the reference background music information, including:
[0043] The background music meta information of the target vlog matches the background music meta information of the referenced background music, and / or,
[0044] The musical form of the background music in the target vlog matches the musical form included in the reference background music information, and / or...
[0045] The harmonic progression of the background music in the target Vlog matches the harmonic progression included in the reference background music information, and / or...
[0046] The background music in the target vlog matches the emotional tone of the music in the reference background music information, and / or...
[0047] The phrase and segment breakpoints of the background music in the target Vlog match the phrase and segment breakpoints included in the reference background music information.
[0048] In one possible embodiment, in processing the reference audio signal based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit is specifically used for:
[0049] Based on the MIDI score of the reference audio signal, determine the music meta information, form, harmonic progression, and musical mood of the reference audio signal; determine whether the music meta information, form, harmonic progression, musical mood, and phrase / section breakpoints of the reference audio signal match the reference background music information, including music meta information, form, harmonic progression, musical mood, and phrase / section breakpoints.
[0050] When there are parts in the reference audio signal that do not match the reference background music information, including its meta information, form, harmonic progression, emotional tone, and phrase / section breakpoints, the reference audio signal is modified to obtain the background music for the target Vlog.
[0051] In one possible embodiment, the determining unit is specifically used for:
[0052] The video footage, the duration of the target Vlog, reference transition points, and video emotional information are input into a pre-trained neural network for processing to obtain the edited video footage, the transition points of the target Vlog, and reference background music information.
[0053] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video mood information, the video generation device further includes:
[0054] The modification unit is used to modify the transition points of the target Vlog after detecting the user's first modification command, so as to obtain the modified transition points of the target Vlog.
[0055] In processing the reference audio signal based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit is specifically used for:
[0056] The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the modified target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0057] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video mood information, the video generation device further includes:
[0058] The modification unit is used to modify the reference background music information after detecting the user's second modification instruction, so as to obtain the modified reference background music information;
[0059] In processing the reference audio signal to obtain the background music for the target Vlog, based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal, the processing unit is used for:
[0060] The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0061] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video mood information, the video generation device further includes:
[0062] The modification unit is used to modify the transition click reference background music information of the target Vlog after detecting the user's first modification command and second modification command, so as to obtain the modified transition points and modified reference background music information of the target Vlog.
[0063] In processing the reference audio signal based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit is specifically used for:
[0064] The background music for the target Vlog is obtained by processing the reference audio signal based on the modified transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0065] In one possible embodiment, the track-splitting transcription unit is specifically used for:
[0066] The reference audio signal is input into a pre-trained track-splitting neural network for processing to obtain a multi-track audio signal; the multi-track audio signal is then input into a pre-trained transcription neural network for processing to obtain the MIDI score and phrase / section segmentation points corresponding to the multi-track audio signal.
[0067] The MIDI score of the reference audio signal includes the MIDI score corresponding to the multi-track audio signal, and the phrase / section segmentation points of the reference audio signal include the phrase / section segmentation points corresponding to the multi-track audio signal.
[0068] Thirdly, embodiments of this application also provide an electronic device, including a processor and a memory, wherein the processor and the memory are connected, wherein the memory is used to store program code, and the processor is used to call the program code to execute part or all of the method described in the first aspect.
[0069] Fourthly, embodiments of this application also provide a chip system applied to an electronic device; the chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected via lines; the interface circuits are used to receive signals from the memory of the electronic device and send the signals to the processor, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs part or all of the method described in the first aspect.
[0070] Fifthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that is executed by a processor to implement part or all of the method described in the first aspect.
[0071] Sixthly, embodiments of this application also provide a computer program that is executed to implement part or all of the methods described in the first aspect.
[0072] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0073] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0074] Figure 1 A system schematic diagram is provided for an embodiment of this application;
[0075] Figure 2 This application provides a flowchart illustrating a method for generating Vlogs.
[0076] Figure 3a It illustrates the transition points, the duration of the Vlog, and the emotional trajectory of the video.
[0077] Figure 3b This is a schematic diagram of a user input interface;
[0078] Figure 4 This is a schematic diagram illustrating the relationship between audio transition points and accents.
[0079] Figure 5 This is a schematic diagram of a display result provided in this embodiment;
[0080] Figure 6 This is a schematic diagram illustrating another display result provided in this embodiment;
[0081] Figure 7 This is a schematic diagram of the structure of a video generation device provided in an embodiment of this application;
[0082] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0083] The following sections will provide detailed explanations.
[0084] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.
[0085] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0086] "Multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0087] The following is an explanation of the terms used in this application.
[0088] A transition refers to moving from one scene to another. A transition point is the point in time within a video where the scene changes from one location to another.
[0089] Musical form: the formal structure of music, such as the general-specific-general structure, the general-general-specific structure, etc.
[0090] Harmonic progressions are used to characterize the relationship between one chord and the next.
[0091] The musical mood progression is used to characterize changes in the emotional tone of music. For example, suppose a piece of music is divided into three parts: the first part is calm, the second part is exciting, and the third part is gentle.
[0092] Musical phrase / section syncopation points include phrase syncopation points and section syncopation points. Phrase syncopation points indicate the end of a musical phrase or the boundary between two phrases. Section syncopation points indicate the end of a musical section or the boundary between two sections.
[0093] The embodiments of this application will now be described with reference to the accompanying drawings.
[0094] See Figure 1, Figure 1 This is a schematic diagram of a system provided for an embodiment of this application. For example... Figure 1 As shown, the system includes a terminal device 101 and a server 102.
[0095] Terminal device 101 is a device capable of data processing and graphics rendering. Common terminal devices include: mobile phones, tablets, laptops, PDAs, mobile internet devices (MIDs), IoT devices, and wearable devices (such as smartwatches, smart bracelets, and pedometers).
[0096] Server 102 is a device that can be used for data storage, processing, and transmission. For example, it can be a cloud server, distributed server, integrated server, rack server, cabinet server, blade server, etc.
[0097] In one example, terminal device 101 sends a Vlog retrieval request to server 102. This request carries video footage, the duration of the target Vlog, reference transition points, emotional information, and a reference audio signal. Server 102 processes the video footage, the duration of the target Vlog, the reference transition points, the emotional information, and the reference audio signal to obtain edited video footage, the transition points of the target Vlog, and reference background music information. Server 102 performs track splitting and transcription processing on the reference audio signal to obtain the MIDI score and the phrase / section segmentation points of the reference audio signal. Server 102 processes the reference audio signal according to the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog. The server obtains the target Vlog based on the edited video footage and the background music of the target Vlog. The background music information of the target Vlog matches the reference background music information, and the transition points of the background music of the target Vlog match the transition points of the target Vlog. Server 102 sends the target Vlog to the terminal device in response to the Vlog retrieval request.
[0098] In another example, because terminal device 101 has powerful computing capabilities, it can obtain the target Vlog itself based on the above information without relying on server 102. The specific implementation process of terminal device 101 obtaining the target Vlog can be found in the relevant description of server 102, and will not be repeated here.
[0099] As can be seen, the solution in this application, by style-changing the music uploaded by users, allows users to hear their favorite songs in their vlogs, and the mood of the music matches the mood of the video, as well as the transition points of the music and video, thus achieving a better vlog creation experience. This solves the problem of monotonous video background music and avoids the current situation of weak correlation between audio and video content, thereby enabling users to create personalized, high-quality vlogs.
[0100] See Figure 2 , Figure 2 This is a flowchart illustrating a Vlog generation method provided in an embodiment of this application. Figure 2 As shown, the method includes:
[0101] S201, The video generation device acquires video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals.
[0102] The video material includes one or more video clips. The target vlog's duration can be a precise length, such as 15s, 30s, or 45s; or it can be a range, such as 0-30s, 30s-1min, or 2min-3min. Video sentiment information is used to indicate the emotional direction of the target vlog. For example, the target vlog's duration is 2 minutes. Figure 3a Figure a in the diagram illustrates the reference transition points. These reference transition points include transition point 1 and transition point 2. On the timeline, transition point 1 and transition point 2 correspond to times of 45 seconds and 75 seconds, respectively. Video emotional information can be represented using a video emotion curve. Figure 3a Graph b in the image illustrates the emotional trajectory of the target vlog. For example... Figure 3a As shown in Figure a, the emotional tone of the target Vlog is low from 0-45 seconds; stable from 45-75 seconds; and excited from 75-120 seconds.
[0103] It should be understood that video footage can be stored in the video generating device or obtained by the video generating device from other devices. The duration of the target vlog can be a default value or entered by the user according to their needs. Reference transition points can be default or entered by the user. Video emotional information can be default or entered by the user.
[0104] Figure 3b This illustrates a user input interface. For example... Figure 3bAs shown, the user input interface includes a desired duration input window, a video timeline where key points can be selected, and an interactive window for drawing emotion curves. Users can enter the desired duration of the vlog in the desired duration input window (if it exceeds the total duration of all uploaded video footage, the user will be prompted to reselect). After selecting the desired duration, the total desired duration and the corresponding timeline will be displayed at the bottom of the interactive interface. Users can add, remove, or move transition points on the timeline. After selecting transition points, users can draw the emotion curve of the entire vlog in the emotion curve drawing window of the interactive interface.
[0105] Optionally, the video generation device can be a terminal device 101 or a server 102.
[0106] S202, The video generation device processes the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information to obtain the edited video footage, the transition points of the target Vlog, and the reference background music information.
[0107] The background music reference information includes music meta information, musical form, harmonic progression, musical mood, and phrasing / section divisions. Music meta information includes, but is not limited to, beats per minute (BPM), music style, and time signature. Music style refers to the genre, such as pop, classical, and rock. The musical mood of the reference background music matches the emotional tone of the video.
[0108] In one example, the video generation device processes video footage, the duration of the target Vlog, reference transition points, and video mood information to obtain edited video footage, transition points of the target Vlog, and reference background music information, based on a trained neural network. Specifically, the video generation device inputs the video footage, the duration of the target Vlog, reference transition points, and video mood information into the trained neural network for processing to obtain edited video footage, transition points of the target Vlog, and reference background music information.
[0109] It should be noted that the transition point of the target Vlog can be the same as the reference transition point, but it can also be different. The reason is that the transition point of the target Vlog matches the time point corresponding to the accented beat of the background music, but the time point corresponding to the accented beat of the background music may not be the same as the reference transition point. Therefore, the video generation device will determine the time point corresponding to the accented beat that is closest to the reference transition point in time as the transition point of the target Vlog.
[0110] like Figure 4As shown, assuming the target Vlog is 2 minutes long, the reference transition points include transition point 1 (45s) and transition point 2 (75s). The five accented notes in the background music correspond to the following time points: t1, t2, t3, t4, and t5. Looking at the timeline, time point t2 is closest to transition point 1, and time point t5 is closest to transition point 2. Therefore, the video generation device determines time points t2 and t5 as the transition points for the target Vlog.
[0111] In one possible embodiment, to avoid the user's desired transition points and reference background music information for the target Vlog, determined using the above method, being undesirable, the video generation device displays the edited video footage, the target Vlog's transition points, and the reference background music information for the user to view and determine if these are what they want. When the target Vlog's transition points and reference background music information are undesirable, the user can modify the corresponding information. Upon detecting the user's first modification instruction, the video generation device modifies the target Vlog's transition points to obtain the modified transition points. Upon detecting the user's second modification instruction, the video generation device modifies the reference background music information accordingly to obtain the modified background music information. The first or second modification instruction includes, but is not limited to, touch commands, voice commands, and gesture commands.
[0112] The reference background music information includes musical meta information, musical form, harmonic progression, musical mood, and phrase / section syncopation points. The modified background music information can be obtained by modifying some or all of the musical meta information, musical form, harmonic progression, musical mood, and phrase / section syncopation points included in the reference background music information.
[0113] In one possible embodiment, the video generation device includes a display screen. The device displays edited video footage, transition points of the target Vlog, and reference background music information on the display screen's interface, allowing the user to view and determine if the transition points and reference background music information are as desired. In another possible embodiment, the video generation device does not have a display screen. Instead, it sends the edited video footage, transition points of the target Vlog, and reference background music information to another device with a display screen, such as the user's terminal device. The other device displays the edited video footage, transition points of the target Vlog, and reference background music information, allowing the user to view and determine if the information is as desired. If the transition points and reference background music information are not as desired, the user can modify the corresponding information. Upon detecting the user's modification instruction, the other device modifies the transition points and reference background music information accordingly, obtaining modified transition points and / or modified background music information. The other device then sends the modified information to the video generation device.
[0114] Figure 5 This is a schematic diagram illustrating a display result provided in this embodiment. For example... Figure 5 As shown, the display interface shows the playback order of the video footage, the transition points of the target Vlog, and background music information. The background music information includes the music information for the verse, interlude, and chorus. When the user is satisfied with the displayed information, they click the "Apply" icon on the display interface. Upon detecting the user's action on the "Apply" icon, the video generation device executes subsequent processes based on the transition points and background music information of the target Vlog. If the user is not satisfied with the displayed information, they can modify the corresponding information. In one example, the user can drag and drop the transition point icon to modify the transition point. The user can also click the verse icon, interlude icon, or chorus icon, and then modify the corresponding music information on the pop-up display interface, including but not limited to music meta information, musical form, harmonic progression, musical mood, and phrase / section syncopation. After modification, if the user is satisfied with the modified information, they click the "Apply" icon on the display interface. Upon detecting the user's action on the "Apply" icon, the video generation device executes subsequent processes based on the modified information. If the user is not satisfied with the modified information, the user can click the "Regenerate" icon. When the video generation device detects the user's action on the "Regenerate" icon, the video generation device re-executes the relevant content of S201-S202 to re-acquire the edited video footage, the transition points of the target Vlog, and the reference background music information.
[0115] S203 The video generation device performs track splitting and transcription processing on the reference audio signal to obtain the MIDI score of the reference audio signal and the phrase / section segmentation points of the reference audio signal.
[0116] To convert the reference audio signal into an audio signal that matches the edited video footage, the video generation device needs to determine relevant information about the reference audio signal, such as its MIDI score and phrase / section breakpoints. Specifically, the video generation device performs track splitting on the reference audio signal to obtain multi-track audio signals. A piece of music is typically derived from audio produced by multiple instruments, or in other words, a piece of music consists of multi-track audio signals, with each instrument's audio signal corresponding to a specific instrument. By splitting the reference audio signal into tracks, the multi-track audio signals that constitute a piece of music can be obtained. For example, music A is produced by drums, bass, guitar, and piano. By splitting the audio signal corresponding to music A into tracks, four audio tracks can be obtained, corresponding to drums, bass, guitar, and piano, respectively.
[0117] In one example, the video generation device performs track splitting on the reference audio signal based on a track splitting neural network. Specifically, the video generation device inputs the reference audio signal into a trained track splitting neural network for processing to obtain multi-track audio signals.
[0118] After obtaining the multitrack audio signal, the video generation device performs transcription processing on the multitrack audio signal to obtain the MIDI score and phrase / section segmentation points corresponding to the multitrack audio signal.
[0119] In one example, the video generation device transcribes multi-track audio signals to obtain multi-track audio signals based on a transcriptional neural network. Specifically, the video generation device inputs the multi-track audio signals into a trained transcriptional neural network for processing to obtain the corresponding MIDI sheet music and phrase / section segmentation points. Optionally, the video generation device can simultaneously input the multi-track audio signals into the transcriptional neural network for processing, or it can input one audio signal into the transcriptional neural network to obtain the corresponding MIDI sheet music and phrase / section segmentation points for that track, and then input the next audio signal into the transcriptional neural network for processing. That is, the processing of multi-track audio signals into the transcriptional neural network can be parallel or serial.
[0120] S204. The video generation device processes the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal to obtain the background music of the target Vlog.
[0121] Among them, the background music information of the target Vlog matches the background music information of the reference Vlog, and the transition points of the background music of the target Vlog match the transition points of the target Vlog.
[0122] The background information of the target Vlog includes at least one of the following: musical meta information, musical form, harmonic progression, musical mood, and phrasing / section breakpoints. The background music information of the target Vlog matches the reference background music information, including:
[0123] The background music meta information of the target Vlog matches the background music meta information included in the reference background music information, and / or, the musical form of the target Vlog matches the musical form of the reference background music information, and / or, the harmonic progression of the target Vlog matches the harmonic progression of the reference background music information, and / or, the musical mood of the target Vlog matches the musical mood of the reference background music information, and / or, the phrase / section syncopation points of the target Vlog's background music match the phrase / section syncopation points of the reference background music information.
[0124] The music meta information includes, but is not limited to, BPM, musical style, and time signature. In one example, matching the music meta information of the background music in the target Vlog with the music meta information in the reference background music information means that at least one of the BPM, musical style, and time signature in the background music meta information of the target Vlog is the same as the BPM, musical style, and time signature in the reference background music information. In one example, matching the musical form of the background music in the target Vlog with the musical form of the reference background music information specifically means that the musical form of the background music in the target Vlog is the same as the musical form of the reference background music information. In one example, matching the harmonic progression of the background music in the target Vlog with the harmonic progression in the reference background music information specifically means that the harmonic progression of the background music in the target Vlog is the same as the harmonic progression in the reference background music information. In one example, matching the musical mood of the background music in the target Vlog with the musical mood of the reference background music information specifically means that the musical mood of the background music in the target Vlog is the same as the musical mood of the reference background music information. In one example, matching the phrase / section breakpoints of the background music in the target Vlog with the phrase / section breakpoints included in the reference background music information specifically means that the phrase / section breakpoints of the background music in the target Vlog are the same as those included in the reference background music information.
[0125] In one example, matching the background music information of the target Vlog with the background music information of the reference Vlog specifically means that the music meta information, musical form, harmonic progression, musical mood, and phrase / section segmentation points of the target Vlog are the same as those of the reference background music information.
[0126] Matching the transition points of the background music in the target Vlog with the transition points of the target Vlog's background music means that the time point corresponding to the transition point in the target Vlog's background music is the same as the time point corresponding to the transition point in the target Vlog's background music, or the difference between the two is less than a preset threshold. Generally, the duration of the target Vlog and the duration of its background music are the same, therefore, the time point corresponding to the transition point in the target Vlog and the time point corresponding to the transition point in the target Vlog's background music are the same.
[0127] To ensure the reference audio signal matches the edited video footage, the video generation device determines the musical meta information, formal structure, harmonic progression, and musical mood of the reference audio signal based on the multi-track MIDI score corresponding to the multi-track audio signal. The video generation device then checks whether the musical meta information, formal structure, harmonic progression, and musical mood of each of the multi-track audio signals are identical to those included in the reference background music information. If any discrepancies exist, the video generation device modifies the reference audio signal to ensure that the musical meta information, formal structure, harmonic progression, and musical mood of each of the reference audio signals are identical to those included in the reference background music information.
[0128] For example, the musical structure of the reference audio signal is "general-part-general", and the musical structure of the reference background music is "general-general-part". The video generation device divides the reference audio signal into three parts according to the general-part-general musical structure, and then reassembles the three parts of the reference audio signal based on the general-general-part musical structure to obtain the modified audio signal. The musical structure of the modified audio signal is always "general-general-part".
[0129] For example, if the reference audio signal is in the classical music style and the background music information includes a pop music style, the video generation device performs rhythmic and adjustment processing on the main melody of the reference audio signal to obtain a modified reference audio signal with a pop style. Optionally, the video generation device can also rewrite the accompaniment part of the reference audio signal and / or modify the orchestration of the reference audio signal.
[0130] It should be noted that the orchestration of the reference audio signal refers to the instruments used to produce the reference audio signal. Modifying the orchestration of the reference audio signal means adding one or more audio tracks to the reference audio signal, and / or deleting one or more audio tracks from the multi-track audio signals corresponding to the reference audio signal. For example, if the reference audio signal is classical music played by a piano and violin, the reference audio signal includes two audio tracks, corresponding to the piano and violin respectively. If the reference background music information includes music of the pop music style, the video generation device can delete the audio track corresponding to the violin from the reference audio signal and add three audio tracks corresponding to the guitar, bass, and drum kit.
[0131] It should be noted that, before executing S204, if the video generation device detects the first modification instruction, it processes the reference audio signal according to the modified transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / segment breakpoints of the reference audio signal to obtain the background music of the target Vlog; if the video generation device detects the second modification instruction, it processes the reference audio signal according to the modified transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / segment breakpoints of the reference audio signal to obtain the background music of the target Vlog; if the video generation device detects both the first and second modification instructions, it processes the reference audio signal according to the modified transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / segment breakpoints of the reference audio signal to obtain the background music of the target Vlog. The specific implementation process can be found in the relevant description of S204, and will not be repeated here.
[0132] S205. The video generation device generates the target Vlog based on the edited video footage and the background music of the target Vlog.
[0133] After obtaining the edited video footage and the background music for the target Vlog, the video generation device combines the edited video footage and the background music for the target Vlog to create the target Vlog.
[0134] In one possible embodiment, after obtaining the target Vlog, the video generation device displays the target Vlog on a display interface to facilitate user viewing of the generated Vlog's quality. For example, Figure 6 As shown, the display interface includes a video display area, an audio display area, and a transport display area. The video display area displays the edited video footage; the audio display area displays the accompaniment tracks and MIDI notes corresponding to each instrument in the background music of the target Vlog. Users can modify and select the accompaniment tracks and MIDI notes corresponding to the instruments using touch controls in the audio display area to modify the background music of the target Vlog. The transport display area displays a progress bar. Users can drag the progress bar to view the overall matching effect between the video and audio.
[0135] As can be seen, the solution in this application allows users to hear their favorite songs in their vlogs by style-changing the music uploaded by the user. The music's mood matches the video's mood, and the music's transition points match the video's transition points, resulting in a better vlog creation experience. By generating background music for the target vlog based on the user's selected video footage, the target vlog's length, reference transition points, video mood information, and reference audio signals, the generated background music meets the user's personalized needs. This solves the problem of monotonous video background music and avoids the current situation of weak correlation between audio and video content, thereby enabling users to create personalized, high-quality vlogs.
[0136] See Figure 7 , Figure 7 This is a schematic diagram of a video generation device provided in an embodiment of this application. Figure 7 As shown, the video generation apparatus 700 includes:
[0137] The acquisition unit 701 is used to acquire video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals;
[0138] The determination unit 702 is used to determine the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information.
[0139] Track splitting and transcription unit 703 is used to split and transcribe the reference audio signal to obtain the MIDI score of the reference audio signal and the phrase / section segmentation points of the reference audio signal.
[0140] The processing unit 704 is used to process the reference audio signal according to the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase segmentation points of the reference audio signal to obtain the background music of the target Vlog. The background music information of the target Vlog matches the reference background music information, and the transition points of the background music of the target Vlog match the transition points of the target Vlog. The target Vlog is obtained based on the edited video material and the background music of the target Vlog.
[0141] In one possible embodiment, the reference background music information includes at least one of the following: music meta information, musical form, harmonic progression, musical mood, and phrase / section syncopation points; the background music information of the target Vlog is matched with the reference background music information, including:
[0142] The background music meta information of the target vlog matches the background music meta information of the referenced background music, and / or,
[0143] The musical form of the background music in the target vlog matches the musical form included in the reference background music information, and / or...
[0144] The harmonic progression of the background music in the target Vlog matches the harmonic progression included in the reference background music information, and / or...
[0145] The background music in the target vlog matches the emotional tone of the music in the reference background music information, and / or...
[0146] The phrase and segment breakpoints of the background music in the target Vlog match the phrase and segment breakpoints included in the reference background music information.
[0147] In one possible embodiment, in processing the reference audio signal based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit 704 is specifically used for:
[0148] Based on the MIDI score of the reference audio signal, determine the music meta information, form, harmonic progression, and musical mood of the reference audio signal; determine whether the music meta information, form, harmonic progression, musical mood, and phrase / section breakpoints of the reference audio signal match the reference background music information, including music meta information, form, harmonic progression, musical mood, and phrase / section breakpoints.
[0149] When there are parts in the reference audio signal that do not match the reference background music information, including its meta information, form, harmonic progression, emotional tone, and phrase / section breakpoints, the reference audio signal is modified to obtain the background music for the target Vlog.
[0150] In one possible embodiment, the determining unit 702 is specifically used for:
[0151] The video footage, the duration of the target Vlog, reference transition points, and video emotional information are input into a pre-trained neural network for processing to obtain the edited video footage, the transition points of the target Vlog, and reference background music information.
[0152] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video mood information, the video generation device 700 further includes:
[0153] Modification unit 705 is used to modify the transition points of the target Vlog after detecting the user's first modification command, so as to obtain the modified transition points of the target Vlog.
[0154] In processing the reference audio signal based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit 704 is specifically used for:
[0155] The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the modified target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0156] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video mood information, the video generation device 700 further includes:
[0157] The modification unit 705 is used to modify the reference background music information after detecting the user's second modification instruction, so as to obtain the modified reference background music information.
[0158] In processing the reference audio signal based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit 704 is used for:
[0159] The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0160] In one possible embodiment, after determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, reference transition points, and video mood information, the video generation device 700 further includes:
[0161] The modification unit 705 is used to modify the transition click reference background music information of the target Vlog after detecting the user's first modification instruction and second modification instruction, so as to obtain the modified transition point and modified reference background music information of the target Vlog.
[0162] In processing the reference audio signal based on the transition points of the target Vlog, reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit 704 is specifically used for:
[0163] The background music for the target Vlog is obtained by processing the reference audio signal based on the modified transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal.
[0164] In one possible embodiment, the track-splitting transcription unit 703 is specifically used for:
[0165] The reference audio signal is input into a pre-trained track-splitting neural network for processing to obtain a multi-track audio signal; the multi-track audio signal is then input into a pre-trained transcription neural network for processing to obtain the MIDI score and phrase / section segmentation points corresponding to the multi-track audio signal.
[0166] The MIDI score of the reference audio signal includes the MIDI score corresponding to the multi-track audio signal, and the phrase / section segmentation points of the reference audio signal include the phrase / section segmentation points corresponding to the multi-track audio signal.
[0167] It should be noted that the aforementioned units (acquisition unit 701, determination unit 702, track-by-track transcription unit 703, processing unit 704, and modification unit 705) are used to perform the relevant steps of the above method. Specifically, acquisition unit 701 is used to implement the relevant content of S201, determination unit 702 is used to implement the relevant content of S202, track-by-track transcription unit 703 is used to implement the relevant content of S203, and processing unit 704 and modification unit 705 are used to implement the relevant content of S204 and S205.
[0168] In this embodiment, the video generation apparatus 700 is presented in the form of a unit. Here, "unit" can refer to an application-specific integrated circuit (ASIC), a processor and memory executing one or more software or firmware programs, integrated logic circuits, and / or other devices that can provide the aforementioned functions. Furthermore, the acquisition unit 701, the determination unit 702, the track-splitting transcription unit 703, the processing unit 704, and the modification unit 705 can be... Figure 8 The processor 801 of the electronic device shown is used to implement this.
[0169] like Figure 8 The electronic device 800 shown can be used as Figure 8 The electronic device 800 is implemented using the structure described above. It includes at least one processor 801, at least one memory 802, and at least one communication interface 803. The processor 801, the memory 802, and the communication interface 803 are connected through the communication bus and communicate with each other. Optionally, the electronic device 800 also includes a display screen 804.
[0170] The processor 801 may be a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of programs in the above schemes; the processor 801 may also include a GPU.
[0171] The communication interface 803 is used to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Networks (WLAN), etc.
[0172] The memory 802 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory may exist independently and be connected to the processor via a bus. The memory may also be integrated with the processor.
[0173] The memory 802 stores the application code for executing the above scheme, and its execution is controlled by the processor 801. The processor 801 executes the application code stored in the memory 802.
[0174] The code stored in memory 802 can execute any of the Vlog generation methods provided above, such as:
[0175] The process involves: acquiring video footage, the target Vlog's duration, reference transition points, video mood information, and reference audio signals; determining the edited video footage, target Vlog's transition points, and reference background music information based on these elements; performing track splitting and transcription on the reference audio signal to obtain its MIDI score and phrase / section segmentation points; processing the reference audio signal based on its transition points, reference background music information, MIDI score, and phrase / section segmentation points to obtain the target Vlog's background music, ensuring that the background music information matches the reference background music information and that the transition points of the background music match the transition points of the target Vlog; and finally, obtaining the target Vlog based on the edited video footage and its background music.
[0176] Processor 801 executes relevant code to control the display of the target Vlog and reference background music information on the display screen 804, allowing the user to view and determine if the transition points and reference background music information of the target Vlog are as desired. If the transition points and reference background music information of the target Vlog are not as desired, the user can modify the corresponding information. Upon detecting the user's first modification instruction, processor 801 modifies the transition points of the target Vlog to obtain the modified transition points. Upon detecting the user's second modification instruction, processor 801 modifies the reference background music information accordingly to obtain the modified background music information.
[0177] This application also provides a computer storage medium, wherein the computer storage medium may store a program, which, when executed, includes some or all of the steps of any of the Vlog generation methods described in the above method embodiments.
[0178] This application also provides a computer program that is executed to implement some or all of the steps of any of the Vlog generation methods described in the above method embodiments.
[0179] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0180] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0181] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.
[0182] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0183] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0184] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, ROM, RAM, portable hard drives, magnetic disks, or optical disks.
[0185] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include a flash drive, ROM, RAM, disk, or optical disk, etc.
[0186] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for generating Vlogs, characterized in that, include: Acquire video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals; The edited video footage, the target Vlog's duration, the reference transition points, and the video's emotional information are determined based on the video footage, the target Vlog's transition points, and the reference background music information. The reference audio signal is processed by track splitting and transcription to obtain the MIDI score of the reference audio signal and the phrase / section segmentation points of the reference audio signal; The reference audio signal is processed based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal to obtain the background music of the target Vlog. The background music information of the target Vlog matches the reference background music information, and the transition points of the background music of the target Vlog match the transition points of the target Vlog. The target Vlog is obtained based on the edited video footage and the background music of the target Vlog.
2. The method according to claim 1, characterized in that, The reference background music information includes at least one of the following: music meta information, musical form, harmonic progression, musical mood, and phrase / section segmentation points; The background music information of the target Vlog matches the reference background music information, including: The music meta information of the background music of the target Vlog matches the music meta information included in the reference background music information, and / or, The musical structure of the background music in the target Vlog matches the musical structure included in the reference background music information, and / or, The harmonic progression of the background music in the target Vlog matches the harmonic progression included in the reference background music information, and / or, The emotional tone of the background music in the target Vlog matches the emotional tone of the background music included in the reference background music information, and / or, The phrase / segment breakpoints of the background music in the target Vlog match the phrase / segment breakpoints included in the reference background music information.
3. The method according to claim 2, characterized in that, The process of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog includes: The musical meta information, formal structure, harmonic progression, and musical emotional direction of the reference audio signal are determined based on the MIDI score of the reference audio signal. Determine whether the music meta information, musical form, harmonic progression, musical mood, and phrase / section segmentation points of the reference audio signal match the reference background music information, including music meta information, musical form, harmonic progression, musical mood, and phrase / section segmentation points. When there are parts in the music meta information, musical form, harmonic progression, musical mood, and phrase / section breakpoints of the reference audio signal that do not match the reference background music information, including music meta information, musical form, harmonic progression, musical mood, and phrase / section breakpoints, the reference audio signal is modified to obtain the background music for the target Vlog.
4. The method according to any one of claims 1-3, characterized in that, The step of determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information includes: The video footage, the duration of the target Vlog, the reference transition points, and the video emotion information are input into a trained neural network for processing to obtain the edited video footage, the transition points of the target Vlog, and the reference background music information.
5. The method according to any one of claims 1-4, characterized in that, After determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information, the method further includes: After detecting the user's first modification command, the transition points of the target Vlog are modified to obtain the modified transition points of the target Vlog. The process of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog includes: The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the modified target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal.
6. The method according to any one of claims 1-4, characterized in that, After determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information, the method further includes: After detecting the user's second modification instruction, the reference background music information is modified to obtain the modified reference background music information; The process of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog includes: The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal.
7. The method according to any one of claims 1-4, characterized in that, After determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information, the method further includes: After detecting the user's first modification command and second modification command, the transition click of the target Vlog is modified according to the reference background music information to obtain the modified transition point of the target Vlog and the modified reference background music information. The process of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog includes: The background music for the target Vlog is obtained by processing the reference audio signal based on the modified transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal.
8. The method according to any one of claims 1-7, characterized in that, The step of performing track-splitting and transcription processing on the reference audio signal to obtain the MIDI score of the reference audio signal and the phrase / section segmentation points of the reference audio signal includes: The reference audio signal is input into a trained track-splitting neural network for processing to obtain a multi-track audio signal. The multitrack audio signal is input into a trained transcription neural network for processing to obtain the MIDI score and phrase / section segmentation points corresponding to the multitrack audio signal. The MIDI score of the reference audio signal includes the MIDI score corresponding to the multi-track audio signal, and the phrase / segment segmentation points of the reference audio signal include the phrase / segment segmentation points corresponding to the multi-track audio signal.
9. A video generation apparatus, characterized in that, The device includes: The acquisition unit is used to acquire video footage, the duration of the target Vlog, reference transition points, video emotional information, and reference audio signals; The determining unit is used to determine the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information. The track-splitting transcription unit is used to split and transcribe the reference audio signal to obtain the MIDI score of the reference audio signal and the phrase / section segmentation points of the reference audio signal. The processing unit is configured to process the reference audio signal according to the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal to obtain the background music of the target Vlog. The background music information of the target Vlog matches the reference background music information, and the transition points of the background music of the target Vlog match the transition points of the target Vlog. The target Vlog is obtained based on the edited video footage and the background music of the target Vlog.
10. The apparatus according to claim 9, characterized in that, The reference background music information includes at least one of the following: music meta information, musical form, harmonic progression, musical mood, and phrase / section segmentation points; The background music information of the target Vlog is matched with the reference background music information, including: The music meta information of the background music of the target Vlog matches the music meta information included in the reference background music information, and / or, The musical structure of the background music in the target Vlog matches the musical structure included in the reference background music information, and / or, The harmonic progression of the background music in the target Vlog matches the harmonic progression included in the reference background music information, and / or, The emotional tone of the background music in the target Vlog matches the emotional tone of the background music included in the reference background music information, and / or, The phrase / segment breakpoints of the background music in the target Vlog match the phrase / segment breakpoints included in the reference background music information.
11. The apparatus according to claim 10, characterized in that, In the aspect of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit is specifically used for: The musical meta information, formal structure, harmonic progression, and musical emotional direction of the reference audio signal are determined based on the MIDI score of the reference audio signal. Determine whether the music meta information, musical form, harmonic progression, musical mood, and phrase / section segmentation points of the reference audio signal match the reference background music information, including music meta information, musical form, harmonic progression, musical mood, and phrase / section segmentation points. When there are parts in the music meta information, musical form, harmonic progression, musical mood, and phrase / section breakpoints of the reference audio signal that do not match the reference background music information, including music meta information, musical form, harmonic progression, musical mood, and phrase / section breakpoints, the reference audio signal is modified to obtain the background music for the target Vlog.
12. The apparatus according to any one of claims 9-11, characterized in that, The determining unit is specifically used for: The video footage, the duration of the target Vlog, the reference transition points, and the video emotion information are input into a trained neural network for processing to obtain the edited video footage, the transition points of the target Vlog, and the reference background music information.
13. The apparatus according to any one of claims 9-12, characterized in that, After determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information, the device further includes: The modification unit is used to modify the transition points of the target Vlog after detecting the user's first modification instruction, so as to obtain the modified transition points of the target Vlog. In the aspect of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit is specifically used for: The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the modified target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal.
14. The apparatus according to any one of claims 9-12, characterized in that, After determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information, the device further includes: The modification unit is used to modify the reference background music information after detecting the user's second modification instruction, so as to obtain the modified reference background music information; In the aspect of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit is specifically used for: The background music for the target Vlog is obtained by processing the reference audio signal based on the transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal.
15. The apparatus according to any one of claims 9-12, characterized in that, After determining the edited video footage, the transition points of the target Vlog, and the reference background music information based on the video footage, the duration of the target Vlog, the reference transition points, and the video emotional information, the device further includes: The modification unit is used to modify the reference background music information of the transition points of the target Vlog after detecting the user's first modification instruction and second modification instruction, so as to obtain the modified transition points of the target Vlog and the modified reference background music information. In the aspect of processing the reference audio signal based on the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog, the processing unit is specifically used for: The background music for the target Vlog is obtained by processing the reference audio signal based on the modified transition points of the target Vlog, the modified reference background music information, the MIDI score of the reference audio signal, and the phrase / segment segmentation points of the reference audio signal.
16. The apparatus according to any one of claims 9-15, characterized in that, The track-splitting transcription unit is specifically used for: The reference audio signal is input into a trained track-splitting neural network for processing to obtain a multi-track audio signal. The multitrack audio signal is input into a trained transcription neural network for processing to obtain the MIDI score and phrase / section segmentation points corresponding to the multitrack audio signal. The MIDI score of the reference audio signal includes the MIDI score corresponding to the multi-track audio signal, and the phrase / segment segmentation points of the reference audio signal include the phrase / segment segmentation points corresponding to the multi-track audio signal.
17. An electronic device comprising a processor and a memory, wherein, The processor is connected to a memory, wherein the memory is used to store program code, and the processor is used to call the program code to implement the method as described in any one of claims 1-8.
18. A computer storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Video music adaptation method and system based on AI video understanding
CN111970579A
Background music recommendation method and device based on short video key frame
CN113190709A