Method, device, equipment and product for generating background music of video
By analyzing video transition points and music features and using existing audio materials to generate video background music, the problems of dataset dependence and training complexity in existing technologies are solved, and efficient, logical and high-quality background music generation is achieved, thereby improving the user experience.
Patent Information
- Application Number
- CN202410295613.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies require large-scale, high-quality video music datasets when generating video background music, which makes training complex and costly. The generated music has insufficient sound quality and logic, and does not fully utilize the background music and template information in the video, affecting the user experience.
By analyzing the transition points of the video and the musical characteristics of the background music, and using existing audio materials, we can directly create new background music that matches the video content, and combine it with the audio track group collection pool to generate high-quality and coherent background music.
It improves the efficiency and sound quality of generating video background music, enhances user experience and satisfaction, and ensures the logic and coordination between the new background music and the video content.
Smart Images

Figure CN120658907A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computer technology, and more particularly, to a method, apparatus, electronic device, and program product for generating background music for a video. Background Art
[0002] Background music in a video is music that serves as background or accompaniment to a video. It often blends with the video content to create a specific atmosphere, mood, and rhythm. While not the primary focus of the video, background music can enhance the viewer's experience, adding depth and richness to the video.
[0003] In videos, background music often features a soft, understated presence that complements the content. It can range from a gentle melody to a high-octane rhythm, depending on the video's theme and style. When choosing background music, consider factors such as the tempo, melody, and volume. The tempo should match the rhythm of the video, and the melody should align with the content and emotion of the video. The volume should also be moderate, neither too loud nor too soft, as this can disrupt or diminish the overall effect of the video. Summary of the Invention
[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device, and program product for generating background music for a video.
[0005] According to a first disclosed aspect, a method for generating background music for a video is provided. The method includes determining a transition point in the video, where the transition point represents a transition between two scenes. The method also includes determining a track set for a second background music based on musical characteristics of the first background music in the video. The method also includes generating a second background music based on the transition point and the track set for the second background music.
[0006] In a second aspect of the disclosure, a device for generating background music for a video is provided. The device includes a transition point determination module configured to determine a transition point in a video, where a transition point represents a transition between two scenes. The device also includes a track group set determination module configured to determine a track group set for a second background music based on musical features of a first background music in the video. Furthermore, the device also includes a background music generation module configured to generate the second background music based on the transition point and the track group set for the second background music.
[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising a processor and a memory coupled to the processor, wherein the memory has instructions stored therein, and when the instructions are executed by the processor, the electronic device executes the method according to the first aspect.
[0008] In a fourth aspect of the present disclosure, a computer program product is provided, wherein a computer-readable storage medium stores computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to the first aspect.
[0009] This summary is intended to introduce a selection of concepts in a simplified form that are further described below in the detailed description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] Figure 1 A schematic diagram illustrating an example environment in which some embodiments of the present disclosure may be implemented;
[0012] Figure 2 A flowchart illustrating a method for generating background music for a video according to some embodiments of the present disclosure is shown;
[0013] Figure 3 A schematic diagram illustrating a process for generating background music for a video according to some embodiments of the present disclosure is provided;
[0014] Figure 4 A schematic diagram illustrating a method for determining a transition point of a video according to some embodiments of the present disclosure is shown;
[0015] Figure 5 A schematic diagram illustrating a module for generating background music for a video according to some embodiments of the present disclosure is shown;
[0016] Figure 6 A block diagram illustrating an apparatus for generating background music for a video according to some embodiments of the present disclosure; and
[0017] Figure 7 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0018] Throughout the drawings, the same or similar reference numbers denote the same or similar elements. DETAILED DESCRIPTION
[0019] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0020] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0021] For example, upon receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0022] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0023] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0024] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0025] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. can refer to different or the same objects, unless explicitly stated otherwise. Other explicit and implicit definitions may also be included below.
[0026] Typically, generating background music for videos relies primarily on multimodal training, which aims to generate background music by analyzing the correspondence between videos and music. However, this approach requires large-scale, high-quality video-music datasets to train the model, which is both difficult and expensive to acquire. Secondly, the training process requires a large amount of computing resources, increasing the complexity of the application. The background music generated by this method may have problems with sound quality and music logic, which directly affects the user's viewing experience. In addition, this method also ignores the background music that may already exist in the video, resulting in a lack of coordination and integration between the newly generated music and the original music. For videos, in addition to the background music of the video, additional information such as the video template is not fully considered and utilized, which limits the accuracy and adaptability of music generation.
[0027] In the embodiments of the present disclosure, by analyzing the transition points of the video and the musical characteristics of the background music of the video, and directly using the existing audio materials related to the background music of the video, new background music that matches the video content and the background music of the video is created. The new background music generated in this way not only maintains the coherence and logic of the new background music, but also ensures that the generated background music is high-quality music. This lightweight method improves the efficiency of generating background music for the video, thereby improving the user experience and satisfaction.
[0028] Figure 1 1 is a schematic diagram of an example environment 100 in which some embodiments of the present disclosure may be implemented. The schematic diagram of the example environment 100 is for illustrative purposes only and is not intended to limit the present invention. Figure 1 As shown, new background music can be configured for videos. These videos can be short or long, and they can contain template events set by the user or automatically. These template events can be template events such as stickers, filters, special effects, or sound effects. These template events can overlap or be separated individually. For example, there can be a special effect for a few seconds, and a "tick-tick" sound effect within those same few seconds. These template times can last for a certain period of time within the entire video. For example, a "tick-tick" sound effect can last for 0.5 seconds within the video. These videos may already have original background music configured. At 102, a video is input into the system. At 104, the input video can be pre-processed, such as extracting background music from the input video. This music can be pure music or other types of music. At 108, event template information for the video can be extracted, such as extracting the time sequence of template events in the video. At 106, once the background music for the video has been extracted, it is input into the system. It should be understood that the systems in the embodiments of the present disclosure can be the same system or different systems.
[0029] Continue to refer Figure 1At step 130 , a template-music analysis is performed on the input video music to obtain the video transition points. Transition points are the switching points between two different scenes or sections. These points can be obvious switching, such as switching from one scene to another, or more subtle transitions, such as gradually shifting from one topic to another.
[0030] Continue to refer Figure 1 At 120, the background music of the extracted video is analyzed, so that the music features of the background music of the video can be obtained at 122. These music features can be music features such as music structure, beat, rhythm, chords, timbre, harmony, emotion, etc. After the music features of the original background music of the video are obtained, at 124, a set of track groups that match the music features of the original background music of the video is found. These track group sets that match the input background music are used to generate new background music for the video. These track group sets that match the input background music are obtained from a pre-constructed track group set pool. Multiple track group sets can constitute a track group set pool, and these track group sets are sets of track groups. The track groups in the track group set pool are audio materials with similar chords, that is, they are materials with the same style and are harmonious and unified when combined. The track group set pool can be obtained by various methods, such as obtaining it from a music producer, or it can be obtained through sound source separation technology (MSS). The tracks, track groups, and track group sets in the track group are all audio materials of music. Continue to refer to Figure 1 When the transition point 132 of the video, the music feature 122 of the original background music of the video, and the track group set 124 that matches the input background music are determined, the new background music of the video can be output at 140.
[0031] In the disclosed embodiments, by analyzing the video's transition points and the musical characteristics of the original background music, and combining it with relevant audio material, new background music is created that is both coherent and logical, and highly compatible with the video content. This lightweight approach not only ensures excellent sound quality for the generated new background music, but also improves the efficiency of generating video background music without requiring video content, thereby enhancing the user experience.
[0032] It should be understood that the architecture and functions in the example environment 100 are described for exemplary purposes only and do not imply any limitation on the scope of the present disclosure. The embodiments of the present disclosure may also be applied to other environments with different structures and / or functions.
[0033] The following will be combined Figures 2 to 7The process according to the embodiment of the present disclosure is described in detail. For ease of understanding, the specific data mentioned in the following description are exemplary and are not intended to limit the scope of protection of the present disclosure. It is understood that the embodiments described below may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this respect.
[0034] Figure 2 A flowchart of a method 200 for generating background music for a video according to some embodiments of the present disclosure is shown. In block 202, transition points in a video are determined, where a transition point represents a transition between two scenes. These videos can be long or short. In some embodiments, the transition points in the video are determined based on template events in the video. These template events can be, for example, template events such as stickers, filters, special effects, or sound effects. In some embodiments, the template events 108 for the video can be pre-determined. For example, the transition points in the video can be determined based on changes in the density of temporal template events in the video. For example, the transition points in the video can also be determined based on changes in the offset intervals of template events in the video. For another example, the transition points in the video can also be determined based on the continuous changes in the video template events in the video. In some embodiments, the transition points in the video can be determined based on the original background music of the video, for example, based on energy changes in the original video music. In some embodiments, when multiple transition points are determined using multiple methods, a relatively suitable number of transition points can be selected to ultimately serve as the transition points in the video.
[0035] In box 204, based on the music features of the first background music of the video, determine the track group set for the second background music. The first background music is the original background music of the video, and the second background music is the new background music that is about to be generated. In some embodiments, the original background music of the input video can be analyzed to obtain the music features of the input original background music by multiple music information retrieval (MIR) models or a unified multi-task basic MIR model. These features can be such as genre, structure, beat and instrumentation timbre or other music features such as beat, speed, theme or emotion, harmony or instrumentation timbre. For example, the track group set for the new background music can be determined based on the genre of the original background music. Alternatively, the music features of the original background music can also be obtained by predetermined rules.
[0036] In block 206, based on the transition point and the track group set for the second background music, the second background music is generated. In some embodiments, the music features 122 of the original background music obtained according to the music analysis 120 are used to screen out a track group set that is related to or matches the original background music in its music features from a pre-made track group set or a mixed track pool. For example, the track group set of the new background music is similar to the original background music in genre or in instrumentation or in emotion and harmony. In some embodiments, the music structure planning of the newly generated background music is determined based on the transition point of the video via 130 and the music structure of the input original background music. In some embodiments, the newly generated video background music 140 is obtained based on the transition point of the video obtained through template-music analysis 130 and the original and determined track group set that matches the input original background music obtained through music analysis 120.
[0037] In the embodiments of the present disclosure, by analyzing the transition points of the video and the musical features of the background music of the video and directly using the existing audio materials related to the background music of the video to create new background music that is consistent with the video content and the background music of the video, the new background music generated in this way not only maintains the coherence and logic of the new background music, but also ensures that the generated background music is high-quality music. In addition, this lightweight method improves the efficiency of generating background music for the video, thereby improving the user experience and satisfaction.
[0038] Figure 3 FIG2 is a schematic diagram showing a process 300 for generating background music for a video according to some embodiments of the present disclosure. Figure 3 , at 301, the original background music is input. In some embodiments, the input original background music may be pure music. At 302, the input original background music is analyzed to obtain the musical features of the input original background music. In some embodiments, these musical features may be information such as structure, the time point when the sound starts, the beat position, emotion or harmony. In some embodiments, the musical features of the original background music may also be analyzed using an MIR model. MIR is a technology for automatically analyzing and extracting various information and features from music. By deeply analyzing the input original background music and extracting various embedded information using MIR technology, the understanding of the original background music can be enhanced.
[0039] Continue to refer Figure 3At 304, event template information in the video can be obtained. In some embodiments, the time series information of the event templates in the video can be obtained through pre-processing. At 305, the music input into the video and the event templates in the video are analyzed to obtain transition points at 306. In some embodiments, the transition points in the video can be determined based on the density changes of the event templates in the video. For example, the system selects a reference time unit, such as a musical beat. It then analyzes the template events in the video and calculates the number of template events occurring within each reference time unit, i.e., 1 beat. This generates a statistical graph of the template event density in the video, which shows the number of events occurring near each time point. Next, a sliding window can be used to analyze this statistical graph. The size and step size of this sliding window are pre-set, for example, it may cover a time range of 4 beats with a step size of 1 beat. At each time point, the system calculates the sum of the event density within the window. This generates a short-term density curve that shows the trend of event density changes near each time point. Next, this short-term density curve is analyzed to identify peaks that are significantly higher than the surrounding values. These peaks represent sudden increases in event density and are typically associated with transition points in the video. Next, a threshold can be set to filter out less significant peaks, resulting in the most likely transition points. By following these steps, we can automatically detect transition points in the video.
[0040] Continue to refer Figure 3At 305, the music input into the video and the event templates in the video are analyzed to obtain a transition point at 306. In some embodiments, the video transition point 404 can be determined based on the changes in the offset intervals of the template events in the video. For example, each template event in the video and the time intervals between them are first analyzed. These time intervals can be in seconds, milliseconds, or any other suitable time unit. Next, the longest of these time intervals is found. The start time of the event after this longest interval is then checked. If the start time of the event after the longest interval falls in the second half of the video, this time point is directly considered a candidate transition point. This is because a long interval may indicate that the video creator has intentionally made discontinuous edits, which is often a sign of a transition point. If the start time of the event after the longest interval falls in the first half of the video, further analysis is required. Specifically, the offset interval of the next event is first checked to see if it is not less than a certain experimental factor (e.g., 0.1) of the longest interval. If this condition is met, the start time of the next event is also checked to see if it is close enough to the downbeat (accent in the music). If both conditions are met, the start time of the event is also considered a candidate transition point. In this method, if the start time of an event is close enough to the strong beat, it may mean that there is a broken rhythm synchronization, and this time point may be a transition point. Through this method, events that may represent transition points can be identified.
[0041] Continue to refer Figure 3At 305, the music input into the video and the event templates in the video are analyzed to obtain a transition point at 306. In some embodiments, the transition point 406 of the video can be determined based on the continuous changes in the template events in the video. For example, first, all events in the video can be traversed and the duration of each event recorded. For example, one template event may last for 5 seconds, while the next template event lasts for 10 seconds. Assume that the 10-second template event is the template event with the longest duration. Next, possible transition points are determined based on the position of the event with the longest duration. If the 10-second event is the beginning event of the video, the start time of the next event will be considered as a candidate transition point. This is because there is usually a noticeable switch or change after a long beginning event. If the 10-second event is the last event in the video, the start time of the last event will also be considered as a candidate transition point. If the 10-second event occurs in the middle of the video, the start time and end time of this event will be compared to see which time point is closer to a strong beat (accent in the music) in the video. The time point closer to the strong beat will be selected as the candidate transition point. In this way, the system can identify potential transition points related to long-term events, which can help better understand the structure and content of the video without knowing the video content.
[0042] Continue to refer Figure 3In some embodiments, the transition point 408 of the video can be determined based on the energy changes of the original background music in the video. First, the background music in the video can be extracted and its frame-level root mean square (RMS) value can be calculated under certain settings. For example, the jump size can be set to 512 samples, the window length can be set to 2048 samples, and the operation can be performed at a sampling rate of 44100. This is done to capture the energy changes of the music signal. Next, the BPM (beats per minute) of the original background music will be identified. By understanding the rhythm of the original music, we can better understand the relationship between the template event in the video and the rhythm of the original background music. Then, the BPM of the music is used to set the jump size and window length for each beat. For example, if the BPM of the original background music is 120, then the duration of each beat is 0.5 seconds. The average RMS value is calculated within each beat to obtain a beat-level RMS statistical curve. Then, based on the beat-level RMS curve, significant peaks are determined. These peaks represent sudden increases in music energy and usually correspond to transition points in the video. Next, in order to determine whether the detected peak actually represents a transition point, two confidence values can be calculated: a pre-confidence and a post-confidence. The pre-confidence measures the stability of the music energy before the peak, while the post-confidence measures the stability of the music energy after the peak. If the pre-peak energy is low and the post-peak energy is stably high, the detected peak has a high confidence level. Next, based on these calculated confidence values, it is determined whether the detected peak is considered a transition point. Only when the confidence level is high enough is it considered that the peak corresponds to an actual transition point in the video. In this way, the system can use the energy changes of the background music to identify the transition points in the video, thereby providing useful information for the subsequent determination of the structure of the new background music output. In some embodiments, after the transition points of the video are determined by multiple methods, the best several transition points can be selected as the transition points of the video. Alternatively, the transition points of the video can also be determined by a data-driven method.
[0043] Continue to refer Figure 3 In order to generate new background music, the most suitable track group set can be selected from the track set pool or the mixed track pool. The mixed track pool 309 can be obtained in the track group set pool 307. In some embodiments, the track groups in these track group sets can be arranged in a certain order in the track group set according to the timbre of the instrument or the properties of the instrument or other rules. In some embodiments, the track groups in the track group set can be divided into track group segments of different lengths based on the music structure information of the track group extracted by the MIR model.
[0044] Continue to refer Figure 3, at 308, the track groups in the track group set are randomly mixed to obtain a mixed track. A mixed track refers to an independent audio, and a mixed track can be formed by mixing two or more track groups. In some embodiments, the track groups in the track group set pool are screened track groups with obvious audio content. In some embodiments, track groups can be randomly selected based on the beat tracking result, i.e., the beat position, and the track groups can be mixed using beat synchronization. By combining different track groups, richer and more complex music effects can be created, enhancing the expressiveness and appeal of the music. Continue to refer to Figure 3 At 309 , a mixed track pool is determined. In some embodiments, the mixed track pool includes a large number of mixed tracks, which can be arranged according to the attributes of the audio track group set.
[0045] Continue to refer Figure 3 At 303, a track group set is selected. After the track group set selection process, a selected track group combination can be obtained at 310. In some embodiments, a mixed track pool is called based on the musical characteristics of the original background music, and a track group set corresponding to the newly generated background music is selected from it. For example, this can be based on mood, theme, or timbre. By considering various musical characteristics such as genre, mood, theme, emotion, or harmony, more personalized music can be generated for the user.
[0046] In some embodiments, the track group combination corresponding to the new background music can be determined based on the genre of the original background music. For example, a K-nearest neighbor search (KNN) is performed in the emotional feature space to select a suitable track group combination for the original background music, the original background music is divided into multiple music segments according to the music structure, and the emotional feature vectors of the multiple music segments are extracted. At the same time, the emotional feature vectors of the mixed tracks in the mixed track pool are also extracted. KNN is used to recall multiple combinations in the random mixed track pool. For example, if there are M music segments, M combinations are recalled, and each combination has K mixed tracks. The mixed tracks in these combinations are all related to the music segments, and these mixed tracks are different from each other. Using KNN to search in the genre feature vector space can more accurately find the track group set related to the original background music, thereby improving the matching accuracy.
[0047] Continue to refer Figure 3The M combinations are voted and ranked, and the track set with the most mixed tracks is selected as the track set corresponding to the new background music. For example, if track set A recalls 10 mixed tracks related to a music segment, and track set B recalls 9 mixed tracks related to a music segment, and track set A is ranked higher than track set B, then track set A is selected as the track set related to the original background music. Through the voting and ranking process, the track set that best meets the requirements can be selected as candidate material, thereby optimizing generation efficiency. In some embodiments, the tempo and starting point information of each music segment can also be considered during the ranking of the M combinations. For example, different tempo (or BPM) multiple combinations are compared and searched for the optimal starting point density match for each music segment. During this matching stage, feedback coefficients for different time stretch ratios are simultaneously calculated to penalize the selection of large time stretch ratios. This feedback coefficient is applied to the ranking process; a larger feedback coefficient indicates a lower ranking. This is because selecting a large time stretch ratio can cause the music to sound unnatural or distorted. Time stretching is a technique that changes the speed of music without changing its pitch. For example, if a song has a BPM of 110, time stretching can make it sound like it has a BPM of 100, but the pitch of each note remains unchanged. By taking into account musical characteristics such as tonality and rhythm, new, more harmonious background music can be generated.
[0048] Continue to refer Figure 3 , at 311, the structure of the new background music is planned. After the track group set is selected at 310, the structure of the new background music is determined at 311. In some embodiments, the musical characteristics of the input original background music are first analyzed, including rhythm, chords, melody, emotion or harmony, etc. Based on the rhythm of the input original background music and the characteristics of the selected track group set, it can be determined how many bars to generate for each part of the newly generated background music. For example, if the input original background music is a fast-paced song, it may be necessary to generate more bars for each part of the newly generated background music to maintain synchronization.
[0049] In some embodiments, if the optimal time stretching ratio has been calculated in the track group set selection processing 303 stage, this ratio can be used to adjust the duration of the selected track group set to match the rhythm of the input original background music. If the optimal time stretching ratio is not calculated, the rhythm ratio between the track group set and the input original background music can be used directly. Next, in combination with the determined transition point 306, for each part containing a transition point, it can be divided into two sub-parts. For example, if there is an obvious transition point in the middle of a part, then this part will be divided into two parts. For each sub-part, the number of bars can be allocated from the corresponding undivided part according to the duration ratio. In this way, the parts before and after the transition point will have a corresponding number of bars to maintain synchronization with the input original background music. In this process, the start timestamp of the input original background music, as well as the duration of each audio content and the total audio duration, are also recorded. This information is very important for subsequent processing.
[0050] Continue to refer Figure 3 At 312, a subset of the mixed track pool is selected, the mixed tracks in this subset consisting of the mixed tracks within the selected track group set 303. In some embodiments, this subset is a subset of the mixed track pool, and the mixed tracks in this subset are a set of mixed tracks corresponding to the selected track group set 310.
[0051] Continue to refer Figure 3 , at 313, a random mixed track is selected from the selected track group set for the music arrangement, that is, the mixed track that is most suitable for the music arrangement 314 is selected again from the selected subset pool. In some embodiments, the mixed track that is most suitable for the music arrangement can be selected based on the musical characteristics of the original background music. Music arrangement refers to the various parts of the new background music that are conceived and output (such as the introduction, main song, chorus, etc.). In some embodiments, for each part, a KNN algorithm can be used to find the mixed track (the best track group set) that best matches the original background music in the embedding space of, for example, the instrumentation timbre. When selecting the track group set, some constraints can be set, such as a certain part requires a specific instrument or the track length of a certain part must match.
[0052] Continue to refer Figure 3In some embodiments, during the KNN search and recall process, the length of the mixed tracks and the number of bars required for each part can also be compared to ensure that they are time-matched. In some embodiments, the harmony or melody of the mixed tracks and the original background music can also be compared, and mixed tracks with similar harmony or melody can be selected. This ensures that the output new background music is also harmonically or melodically coordinated with the input original background music. In some embodiments, if the recalled mixed tracks do not meet the above constraints, a feedback coefficient is calculated and used in the ranking process to determine the ranking of the mixed track pool. For example, if the selection of a mixed track results in a shorter mixed audio content length, a higher feedback coefficient can be given to reduce such deviations from the constraints in subsequent selection processes. Similarly, if the selected mixed track does not match the musical characteristics, a corresponding feedback coefficient can also be given. These constraints will more accurately find a set of audio track groups that meet the requirements.
[0053] Continue to refer Figure 3 In some embodiments, the distribution of instrumental sounds in the newly output music can be adjusted based on the relative position of transition points. For example, if a transition point is located at the beginning of a particular section, and the first subsection of this section (i.e., the section before the transition point) originally contained a drum sound, then the drum sound in this section can be deleted to match the effect of the transition point. A drum sound can be added to the subsection after the transition point. If the section after the transition point originally did not have a drum sound, it can be added to enhance the rhythm or emotional effect. When the duration of a subsection exceeds a threshold value, another recalled track group set in the ranking can be selected. That is, if a track group set has already been selected for a subsection, but its duration is shorter than the subsection, then another track group set can be selected from the previously recalled track group sets to fill the subsection. If the duration of the newly selected track group set is still insufficient to fill the entire subsection, or if its duration exceeds the subsection, the excess portion will be replaced by the newly selected track group set. In some embodiments, the selected track group sets can be further optimized based on the musical characteristics of various original background music (such as timbre matching, rhythm synchronization, emotional similarity or harmonic similarity, etc.) to ensure that they are coordinated with the overall style and structure of the song.
[0054] Continue to refer Figure 3At 315, audio is generated and post-processed. After the music arrangement 314, the corresponding track groups are combined according to the arrangement structure of each part, that is, the number of bars required for each part, such as the chorus, introduction, etc., is taken into consideration. When the audio length of a certain part of the new background music is not enough to meet the required number of bars, the audio of this part can be looped to make the audio length meet the requirements. In some embodiments, the audio can be guaranteed to have a certain length by copying a certain audio segment and appending it to the end of the audio.
[0055] In some embodiments, after the audio of each part of each new background music is generated and expanded, the audio can be post-processed to ensure the continuity of the output audio. In some embodiments, the audio can be blank-filled. In some embodiments, the audio can also be time-stretched to change the playback speed of the audio without changing the pitch of the audio, so as to ensure that the output new background music matches the rhythm of the input original background music. In some embodiments, the length of the audio can also be trimmed. If the length of the generated audio exceeds the required length, it can be trimmed to ensure that it meets the requirements. In some embodiments, the audio can also be dynamically compressed, that is, the difference between the maximum and minimum volume in the audio is reduced. Continue to refer to Figure 3 At 316, new background music for the video that is related to the original background music is output. The newly generated background music can replace the original background music of the video. In this way, the user can obtain brand new background music that is related to the original background music and matches the rhythm content of the video.
[0056] Figure 4 FIG. 4 is a schematic diagram showing a method for determining a transition point 400 of a video according to some embodiments of the present disclosure. Figure 4The video's transition points can be determined based on changes in the density of event templates in the video. Specifically, at 402, for example, a benchmark time unit—the musical beat—can be introduced as a reference for quantifying the density of video events. Analysis of the video template events can count the number of template events occurring within each beat, and based on this, a histogram of the template event density can be constructed. This histogram intuitively displays the frequency distribution of events occurring around various time points in the video. A sliding window analysis method is then used to determine the dynamic changes in the template event density. The sliding window size and step size are pre-set, for example, covering a time range of five beats, with a step size of one beat. At each time point, the sum of the event density within the window is calculated, generating a short-term density curve that accurately reflects the changing trend of event density. The short-term density curve is then analyzed, identifying peaks that are significantly higher than the surrounding values. These peaks are closely associated with transition points in the video. To improve recognition accuracy, a threshold can be set to filter out less significant peaks, thereby accurately determining the most likely transition points. This approach is based on automated processing, which reduces the need for human intervention.
[0057] refer to Figure 4 The transition point of a video can be determined based on the changes in the offset intervals of template events in the video. Specifically, at 404, each template event in the video and the time intervals between them are first analyzed. These time intervals can be in seconds, milliseconds, or any other suitable time unit. Next, the longest of these time intervals is found. The start time of the event after this longest interval is then checked. If the start time of the event after the longest interval falls in the second half of the video, then this time point is directly considered a candidate transition point. This is because a long interval may indicate that the video editor has made intentional, discontinuous cuts there, which is often a sign of a transition point. If the start time of the event after the longest interval falls in the first half of the video, further analysis is required. Specifically, the system first checks whether the offset interval of the next event is not less than a fixed proportion (e.g., 10%) of the longest interval. If this condition is met, the system also checks whether the start time of the next event is sufficiently close to the downbeat (the accent in music). If both conditions are met, then the start time of the event is also considered a candidate transition point. In this method, if the start time of an event is close enough to the downbeat, it may indicate a break in rhythmic synchronization, and thus a possible transition point. This method can identify events that may represent transition points. By combining time interval analysis, event offset comparison, and musical rhythm considerations, events that may represent transition points can be more accurately identified.
[0058] Continue to refer Figure 4 The transition point of the video can be determined based on the continuous changes in the template events in the video. Specifically, at 406, for example, all template events in the video can be comprehensively traversed and the duration of each event can be recorded in detail. For example, a template event may last for 6 seconds, while the next template event may last for 12 seconds. In this scenario, it is assumed that the 12-second event is the template event with the longest duration. Subsequently, potential transition points can be preliminarily determined based on the location of this longest-duration event. If the 12-second event is located at the beginning of the video, the starting time of the next event will be considered as a possible candidate transition point, because long opening events are often followed by obvious scene or content changes. If the 12-second event is located at the end of the video, the starting time of this event will also be considered as a candidate transition point. If the template event is located in the middle of the video, the system will compare its start and end times to determine which time point is closer to the downbeat in the video (a musical term referring to the accent in the musical rhythm). The time point closer to the downbeat will be selected as a potential transition point. This strategy can identify potential transition points closely related to long-term events without knowing the specific content of the video.
[0059] refer to Figure 4, the transition point of the video can be determined based on the energy change of the original background music in the video. Specifically, in 408, for example, first, the background music in the video can be extracted and its frame-level root mean square (RMS) value can be calculated under certain settings. For example, the jump size can be set to 512 samples, the window length can be set to 2048 samples, and the operation can be performed at a sampling rate of 44100. This is done to capture the energy changes of the music signal. Next, the rhythm BPM (beats per minute) of the original background music will be identified. By understanding the rhythm of the original music, you can better understand the relationship between the template event in the video and the rhythm of the original background music. Then, use the BPM of the music to set the jump size and window length of each beat. For example, if the BPM of the original background music is 120, then the duration of each beat is 0.5 seconds. The average value of the RMS will be calculated within each beat to obtain the beat-level RMS statistics. Then, based on the beat-level RMS curve, significant peaks are determined. These peaks represent sudden increases in music energy, which usually correspond to transition points in the video. Next, in order to determine whether the detected peak actually represents a transition point, two confidence values can be calculated: the pre-confidence and the post-confidence. The pre-confidence measures the stability of the music energy before the peak, while the post-confidence measures the stability of the music energy after the peak. If the pre-peak energy is low and the post-peak energy is stably high, the detected peak has a high confidence. Next, based on these calculated confidence values, it is determined whether the detected peak is considered a transition point. Only when the confidence is high enough will the peak be considered to correspond to an actual transition point in the video. In this way, the system can use the energy changes of the background music to identify the transition points in the video, thereby providing useful information for the subsequent determination of the structure of the new background music output.
[0060] Figure 5 FIG2 shows a schematic diagram of a module for generating background music for a video according to some embodiments of the present disclosure. Figure 5 , pre-processing 510, music analysis 520, determining video transition points 530, calling the mixed track pool 540, and music generation 550 are five different functional parts for generating new background music for a video. Among them, the understanding module for generating new background music and the generation module for generating new background music are two different modules that are decoupled from each other. Figure 5 As shown, preprocessing 510 mainly involves preprocessing the video, for example, inputting the video at 512 and analyzing the input to extract the original background music of the video at 514. For example, time information of the template event sequence of the video can also be extracted.
[0061] Continue to refer Figure 5At 510, the input original background music is analyzed. Musical features of the input original background music can be extracted through music analysis, such as the genre, structure, harmony, emotion, and other musical features. In some embodiments, the input original background music can be divided into multiple music segments based on the structure of the original background music. Each music segment can then be input into a music information retrieval model to obtain the musical features of each segment.
[0062] like Figure 5 As shown, at 530, a video transition point is determined. In some embodiments, a video transition point 402 can be determined based on changes in the density of event templates in the video. In some embodiments, a video transition point 404 can be determined based on changes in the offset interval of template events in the video. In some embodiments, a video transition point 406 can be determined based on the continuous change of template events in the video. For example, all events in the video can be traversed first, and the duration of each event can be recorded. For example, one template event may last for 5 seconds, while the next template event lasts for 10 seconds. Assume that the 10-second template event is the template event with the longest duration. Next, possible transition points are determined based on the position of the event with the longest duration. If the 10-second event is the first event in the video, the start time of the next event will be considered as a candidate transition point. This is because a long first event is usually followed by a noticeable switch or change. If the 10-second event is the last event in the video, the start time of the last event will also be considered as a candidate transition point. If this 10-second event occurs in the middle of the video, the start time and end time of the event will be compared to see which time point is closer to a strong beat (accent in the music) in the video. The time point closer to the strong beat will be selected as the candidate transition point. In this way, the system can identify potential transition points related to long-term events, which can help to better understand the structure and content of the video without knowing the content of the video. In some embodiments, the transition point of the video can be determined based on the energy change of the original background music in the video. In some embodiments, after the transition point of the video is determined by a variety of methods, the best several transition points can be selected as the transition point of the video. Alternatively, the transition point of the video can also be determined by a data-driven method.
[0063] See you next time Figure 5At 540, a mixed track pool can be called. In some embodiments, track groups within a track group set can be randomly mixed to generate randomly mixed tracks. In some embodiments, the track groups in the track group set pool are screened track groups containing audio content. In some embodiments, track groups can be randomly selected based on beat tracking results, i.e., beat positions, and mixed using beat synchronization. By combining different track groups, richer and more complex musical effects can be created, enhancing the expressiveness and appeal of the music. At 540, by calling the mixed track pool and referencing the musical characteristics of the original background music obtained through music analysis 520, a set of track groups related to or corresponding to the output music can be obtained. For example, a set of relevant track groups can be first determined from the constructed randomly mixed track pool based on the genre of the original background music. Subsequently, a set of randomly mixed tracks most suitable for music arrangement can be determined from the determined track group set based on the timbre characteristics of the original background music.
[0064] refer to Figure 5 At 550, background music for the video may be generated. In 550, the process mainly involves planning the music structure of the new background music 552 and arranging the new background music 554 based on the secondary selected mixed track. Figure 3 As shown, after selecting the track group set at 310, music structure planning 532 will be performed on the output music at 311. Music structure planning can ensure that the new background music generated in the end is consistent with the original background music in structure and rhythm.
[0065] In some embodiments, the structure of the new background music can be determined based on the music features obtained according to the MIR model and the selected track group set, that is, how many bars each part of the new background music and each part corresponding to the music segment should be divided into. In some embodiments, if the optimal time stretch ratio has been calculated in the track group set selection process 303, the number of bars can be determined based on the optimal time stretch ratio. If the optimal time stretch ratio has not been calculated, the speed ratio between the selected track group set and the original background music can be directly used to calculate how many bars each part of the new background music needs to generate. This can facilitate adjusting the speed of the new background music to match the rhythm of the original background music, thereby outputting new background music with the same music duration as the original background music. In some embodiments, in this process, the starting time point of the new background music and the duration of the new background music audio content can also be calculated, thereby paving the way for the subsequent generation of new background music. By preparing information such as the start timestamp, the audio content duration, and the audio duration, necessary inputs can be provided for the subsequent audio processing process, improving overall efficiency.
[0066] Continue to refer Figure 5 At 554, a mixed track is selected from the selected track group set for use in the music arrangement. Figure 3 As described above, after the track group set selection process 303, a selected track group combination can be obtained at 310. Next, a second random mix of tracks is selected from the selected track group set for use in the music arrangement, i.e., the most suitable mixed track for the music arrangement is again selected from the selected subset 534. In some embodiments, the most suitable mixed track for the music arrangement can be selected based on the musical characteristics of the original background music. For example, the timbre of the instrument can be used to find the track group set that best matches each short segment of the original background music for the new background music in the instrument timbre feature space. For example, a KNN search can be performed in the timbre feature space to find the mixed track that best matches the original background music. In some embodiments, during the KNN search, the length of the mixed track and the number of bars required for each music segment can be compared to ensure that they are time-matched. In some embodiments, the harmony or melody of the mixed track pool is compared with the original background music, and the mixed track pool with similar harmony or melody is selected. This ensures that the output new background music is harmonically or melodically consistent with the input original background music.
[0067] Decoupling the music generation module and the music analysis module can improve the flexibility of the system. This design can enable the system to better generate a variety of personalized music and ensure the quality and efficiency of the generated background music for the video.
[0068] Figure 6 FIG. 6 is a block diagram of an apparatus 600 for generating background music for a video according to some embodiments of the present disclosure. Figure 6 As shown, apparatus 600 includes a transition point determination module 602 configured to determine a transition point of a video, wherein a transition point represents a transition between two scenes. Apparatus 600 also includes a track group set determination module 604 configured to determine a track group set for a second background music based on musical features of the first background music of the video. Apparatus 600 also includes a background music generation module 606 configured to generate a second background music based on the transition point and the track group set for the second background music.
[0069] Figure 7 FIG2 shows a block diagram of an electronic device 700 according to some embodiments of the present disclosure. The device 700 may be a device or apparatus described in the embodiments of the present disclosure. Figure 7As shown, the device 700 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 702 or computer program instructions loaded from a storage unit 708 into a random access memory (RAM) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The CPU / GPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 707. An input / output (I / O) interface 705 is also connected to the bus 704. Although not shown in FIG. Figure 7 As shown in FIG, device 700 may further include a co-processor.
[0070] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0071] The various methods or processes described above may be performed by the CPU / GPU 701. For example, in some embodiments, the methods may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the CPU / GPU 701, one or more steps or actions in the methods or processes described above may be performed.
[0072] In some embodiments, the methods and processes described above may be implemented as a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure.
[0073] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.
[0074] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0075] The computer program instructions for performing the disclosed operation can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data or source code or the object code written in any combination of one or more programming languages, programming languages include object-oriented programming languages, and conventional procedural programming languages.Computer-readable program instructions can be performed completely on a user's computer, partially on a user's computer, performed as an independent software package, partly on a user's computer and partly on a remote computer, or performed completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer by any type of network-including local area network (LAN) or wide area network (WAN), or can be connected to an external computer (such as utilizing an internet service provider to connect by the internet). In certain embodiments, by utilizing the state information of computer-readable program instructions to carry out personalized customization electronic circuits, such as programmable logic circuits, field programmable gate arrays (FPGAs) or programmable logic arrays (PLA), this electronic circuit can perform computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0076] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0077] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0078] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented by a special hardware-based system that performs the prescribed function or action, or can be implemented by a combination of special hardware and computer instructions.
[0079] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements to existing technologies, or to enable other persons skilled in the art to understand the embodiments disclosed herein.
[0080] Some example implementations of the present disclosure are listed below.
[0081] Example 1. A method for generating background music for a video, comprising:
[0082] Determining a transition point of the video, the transition point representing a transition between two scenes;
[0083] Determining a track group set for second background music based on music features of the first background music of the video; and
[0084] The second background music is generated based on the transition point and the track group set for the second background music.
[0085] Example 2. The method of Example 1, wherein determining a transition point of the video comprises:
[0086] determining, based on time information of the video template events, a plurality of density values of the plurality of video template events within a plurality of time intervals of the video; and
[0087] The transition point of the video is determined based on the multiple density values and the multiple time intervals.
[0088] Example 3. The method of any one of Examples 1-2, wherein determining a transition point of the video comprises:
[0089] Determining multiple offset intervals between multiple video template events based on time information of the video template event;
[0090] determining an offset interval satisfying an offset condition among the plurality of offset intervals; and
[0091] The transition point of the video is determined based on the offset interval that satisfies the offset condition.
[0092] Example 4. The method of any one of Examples 1-3, wherein determining a transition point of the video comprises:
[0093] Determining, based on time information of the video template event, a video template event that meets a persistence condition among the plurality of video template events; and
[0094] Based on the video template event that meets the persistence condition, the end time of the previous video template event or the start time of the next video template event of the video template event is determined as the transition point of the video.
[0095] Example 5. The method of any one of Examples 1-4, wherein determining a transition point of the video comprises:
[0096] determining, based on the first background music, a plurality of energy characteristic values of the first background music within a plurality of time intervals of the video;
[0097] Determining a curve fitting relationship between the plurality of energy characteristic values and the characteristic values of the plurality of time intervals;
[0098] Determining a corresponding peak point based on the characteristic value curve fitting relationship; and
[0099] In response to a first ratio of multiple times below the characteristic value threshold before the peak point and a second ratio of multiple times above the characteristic value threshold after the peak point satisfying a ratio condition, the peak point is determined to be the transition point of the video.
[0100] Example 6. The method of any one of Examples 1-5, wherein determining a track group set for the second background music based on music features of the first background music of the video comprises:
[0101] Invoke the hybrid track pool; and
[0102] Based on the mixed track pool, the track group set for the second background music is determined.
[0103] Example 7. The method of any one of Examples 1-6, wherein generating the second background music comprises:
[0104] Based on the determined track group set for the second background music, determining a mixed track corresponding to the determined track group set for the second background music from the mixed track pool;
[0105] generating a subset of the hybrid trajectory pool based on the corresponding hybrid trajectory; and
[0106] Based on the subset, a combination of mixed tracks for the second background music is determined.
[0107] Example 8. The method of any one of Examples 1-7, wherein generating the second background music further comprises:
[0108] Based on the determined track group set for the second background music, the music features of the first background music, and the transition points, the music structure for the second background music is determined.
[0109] Example 9. The method of any one of Examples 1-8, wherein generating the second background music further comprises:
[0110] The second background music is generated based on the music structure and the determined combination of the mixed tracks for the second background music.
[0111] Example 10. A device for generating background music for a video, comprising:
[0112] a transition point determination module, configured to determine a transition point of the video, wherein the transition point represents a transition between two scenes;
[0113] a track group set determining module configured to determine a track group set for the second background music based on the music features of the first background music of the video; and
[0114] The background music generation module is configured to generate the second background music based on the transition point and the track group set for the second background music.
[0115] Example 11. The apparatus of any one of Example 10, wherein the transition point determination module comprises:
[0116] a density value determining module configured to determine a plurality of density values of the plurality of video template events within a plurality of time intervals of the video based on time information of the video template events; and
[0117] The first transition point determination module is configured to determine the transition point of the video based on the multiple density values and the multiple time intervals.
[0118] Example 12. The apparatus of any one of Examples 10-11, wherein the transition point determination module comprises:
[0119] an offset interval determining module, configured to determine a plurality of offset intervals between a plurality of video template events based on time information of the video template events;
[0120] an offset interval determining module satisfying an offset condition, configured to determine an offset interval satisfying an offset condition from among the plurality of offset intervals; and
[0121] The second transition point determination module is configured to determine the transition point of the video based on the offset interval that meets the offset condition.
[0122] Example 13. The apparatus of any one of Examples 10-12, wherein the transition point determination module comprises:
[0123] a video template event determination module configured to determine a video template event that meets a persistence condition among a plurality of video template events based on time information of the video template event; and
[0124] The third transition point determination module is configured to determine the end time of the previous video template event or the start time of the next video template event of the video template event as the transition point of the video based on the video template event that meets the persistence condition.
[0125] Example 14. The apparatus of any one of Examples 10-13, wherein the transition point determination module comprises:
[0126] an energy characteristic value determination module, configured to determine, based on the first background music, a plurality of energy characteristic values of the first background music within a plurality of time intervals of the video;
[0127] a curve fitting relationship determination module, configured to determine a curve fitting relationship between the plurality of energy characteristic values and the characteristic values of the plurality of time intervals;
[0128] a peak point determination module configured to determine a corresponding peak point based on the characteristic value curve fitting relationship; and
[0129] The fourth transition point determination module is configured to determine that the peak point is the transition point of the video in response to a first ratio of multiple times below the characteristic value threshold before the peak point and a second ratio of multiple times above the characteristic value threshold after the peak point satisfying a ratio condition.
[0130] Example 15. The apparatus of any of Examples 10-14, wherein the module for determining a set of audio track groups comprises:
[0131] A hybrid track pool calling module configured to call the hybrid track pool; and
[0132] The first track group set determination module is configured to determine the track group set for the second background music based on the mixed track pool.
[0133] Example 16. The apparatus of any one of Examples 10-15, wherein the background music generation module comprises:
[0134] a mixed track determination module configured to determine, based on the determined track group set for the second background music, a mixed track corresponding to the determined track group set for the second background music from the mixed track pool;
[0135] a subset generation module, configured to generate a subset of the hybrid track pool based on the corresponding hybrid track; and
[0136] The combination determination module is configured to determine the combination of the mixed tracks for the second background music based on the subset.
[0137] Example 17. The apparatus of any one of Examples 10-16, wherein the background music generation module further comprises:
[0138] The music structure determination module is configured to determine the music structure of the second background music based on the determined track group set for the second background music, the music features of the first background music, and the transition point.
[0139] Example 18. The apparatus of any one of Examples 10-17, wherein the background music generation module further comprises:
[0140] The second background music generating module is configured to generate the second background music based on the music structure and the determined combination of the mixed tracks for the second background music.
[0141] Example 19. An electronic device comprising:
[0142] processor; and
[0143] A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device performs actions, the actions comprising:
[0144] Determining a transition point of the video, the transition point representing a transition between two scenes;
[0145] Determining a track group set for second background music based on music features of the first background music of the video; and
[0146] The second background music is generated based on the transition point and the track group set for the second background music.
[0147] Example 20. The electronic device of Example 19, wherein determining a transition point of the video comprises:
[0148] determining, based on time information of the video template events, a plurality of density values of the plurality of video template events within a plurality of time intervals of the video; and
[0149] The transition point of the video is determined based on the multiple density values and the multiple time intervals.
[0150] Example 21. The electronic device of any of Examples 19-20, wherein determining a transition point of the video comprises:
[0151] Determining multiple offset intervals between multiple video template events based on time information of the video template event;
[0152] determining an offset interval satisfying an offset condition among the plurality of offset intervals; and
[0153] The transition point of the video is determined based on the offset interval that satisfies the offset condition.
[0154] Example 22. The electronic device of any of Examples 19-21, wherein determining a transition point of the video comprises:
[0155] Determining, based on time information of the video template event, a video template event that meets a persistence condition among the plurality of video template events; and
[0156] Based on the video template event that meets the persistence condition, the end time of the previous video template event or the start time of the next video template event of the video template event is determined as the transition point of the video.
[0157] Example 23. The electronic device of any of Examples 19-22, wherein determining a transition point of the video comprises:
[0158] determining, based on the first background music, a plurality of energy characteristic values of the first background music within a plurality of time intervals of the video;
[0159] Determining a curve fitting relationship between the plurality of energy characteristic values and the characteristic values of the plurality of time intervals;
[0160] Determining a corresponding peak point based on the characteristic value curve fitting relationship; and
[0161] In response to a first ratio of multiple times below the characteristic value threshold before the peak point and a second ratio of multiple times above the characteristic value threshold after the peak point satisfying a ratio condition, the peak point is determined to be the transition point of the video.
[0162] Example 24. The electronic device of any of Examples 19-23, wherein determining a set of audio track groups for the second background music based on music features of the first background music of the video comprises:
[0163] Invoke the hybrid track pool; and
[0164] Based on the mixed track pool, the track group set for the second background music is determined.
[0165] Example 25. The electronic device of any one of Examples 19-24, wherein generating the second background music comprises:
[0166] Based on the determined track group set for the second background music, determining a mixed track corresponding to the determined track group set for the second background music from the mixed track pool;
[0167] generating a subset of the hybrid trajectory pool based on the corresponding hybrid trajectory; and
[0168] Based on the subset, a combination of mixed tracks for the second background music is determined.
[0169] Example 26. The electronic device of any one of Examples 19-25, wherein generating the second background music further comprises:
[0170] Based on the determined track group set for the second background music, the music features of the first background music, and the transition points, the music structure for the second background music is determined.
[0171] Example 27. The electronic device of any one of Examples 19-26, wherein generating the second background music further comprises:
[0172] Based on the music structure and the determined combination of mixed tracks for the second background music, the second background music is generated.
[0173] Example 28. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of Examples 1 to 9.
[0174] Example 29. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method of any one of Examples 1 to 9.
[0175] Although the present disclosure has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for generating background music for a video, comprising: Determining a transition point of the video, the transition point representing a transition between two scenes; Determining a track group set for second background music based on music features of the first background music of the video; as well as The second background music is generated based on the transition point and the track group set for the second background music.
2. The method according to claim 1, wherein determining the transition point of the video comprises: determining, based on time information of the video template events, a plurality of density values of the plurality of video template events within a plurality of time intervals of the video; as well as The transition point of the video is determined based on the multiple density values and the multiple time intervals.
3. The method according to claim 1, wherein determining the transition point of the video comprises: Determining multiple offset intervals between multiple video template events based on time information of the video template event; determining an offset interval satisfying an offset condition among the plurality of offset intervals; as well as The transition point of the video is determined based on the offset interval that satisfies the offset condition.
4. The method according to claim 1, wherein determining the transition point of the video comprises: Determining, based on time information of the video template event, a video template event that meets a persistence condition among the plurality of video template events; as well as Based on the video template event that meets the persistence condition, the end time of the previous video template event or the start time of the next video template event of the video template event is determined as the transition point of the video.
5. The method according to claim 1, wherein determining the transition point of the video comprises: determining, based on the first background music, a plurality of energy characteristic values of the first background music within a plurality of time intervals of the video; Determining a curve fitting relationship between the plurality of energy characteristic values and the characteristic values of the plurality of time intervals; Determining a corresponding peak point based on the characteristic value curve fitting relationship; as well as In response to a first ratio of multiple times below the characteristic value threshold before the peak point and a second ratio of multiple times above the characteristic value threshold after the peak point satisfying a ratio condition, the peak point is determined to be the transition point of the video.
6. The method according to claim 1, wherein determining a track group set for the second background music based on the music features of the first background music of the video comprises: Calling the hybrid track pool; as well as Based on the mixed track pool, the track group set for the second background music is determined.
7. The method according to claim 6, wherein generating the second background music comprises: Based on the determined track group set for the second background music, determining a mixed track corresponding to the determined track group set for the second background music from the mixed track pool; generating a subset of the hybrid track pool based on the corresponding hybrid track; as well as Based on the subset, a combination of mixed tracks for the second background music is determined.
8. The method according to claim 7, wherein generating the second background music further comprises: Based on the determined track group set for the second background music, the music features of the first background music, and the transition points, the music structure for the second background music is determined.
9. The method according to claim 8, wherein generating the second background music further comprises: The second background music is generated based on the music structure and the determined combination of the mixed tracks for the second background music.
10. A device for generating background music for a video, comprising: a transition point determination module, configured to determine a transition point of the video, wherein the transition point represents a transition between two scenes; a track group set determining module configured to determine a track group set for the second background music based on the music features of the first background music of the video; as well as The background music generation module is configured to generate the second background music based on the transition point and the track group set for the second background music.
11. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, wherein when the instructions are executed by the processor, the electronic device executes the method according to any one of claims 1 to 9.
12. A computer program product comprising computer executable instructions, wherein the computer executable instructions are executed by a processor to implement the method according to any one of claims 1 to 9.