A public relations film intelligent music matching system fusing music elements
Patent Information
- Application Number
- CN202610991829.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-07-06
AI Technical Summary
[0004]为解决上述技术问题,提供一种融合音乐元素的企宣片智能配乐匹配系统,本技术方案解决了上述人工制作中极易出现视频画面情绪起伏与音乐情感变化割裂的情况,严重影响企宣片的整体观感与传播力的问题
Smart Images

Figure CN122507904B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia information processing technology, specifically to an intelligent music matching system for corporate promotional videos that integrates musical elements. Background Technology
[0002] Corporate promotional videos serve as a core visual medium for brand display, product promotion, and cultural dissemination. Leveraging their intuitive and vivid communication advantages, they are widely used in various scenarios such as brand promotion, investment attraction, trade show presentations, and new media communication. High-quality background music is a key element in enhancing the quality of a corporate promotional video, creating a conducive atmosphere, conveying corporate emotions, and strengthening the audience's visual experience. The compatibility of the background music with the rhythm, atmosphere, and emotional tone of the video directly determines the video's communicative effectiveness and the quality of its brand presentation. Therefore, precise and efficient background music matching is an indispensable core step in the corporate promotional video production process.
[0003] Currently, most of the music composition for corporate promotional videos in China is done manually. The mainstream production process involves editors manually selecting music materials based on their personal aesthetic sense and production experience, and then manually comparing the music with the video shots, scenes, rhythm, and style to complete the music matching and editing. This traditional work model has several inherent flaws: First, manual music composition relies heavily on the professional experience and aesthetic sense of the personnel involved. Different producers may choose different music, easily leading to problems such as music style not matching the brand tone of the promotional video, music emotion being disconnected from the video scene, and music rhythm not matching shot transitions. The accuracy of music matching is inconsistent, making it difficult to guarantee the production quality of the promotional video. Second, the process of manually selecting massive amounts of music materials and adapting them segment by segment to video shots is cumbersome, time-consuming, and labor-intensive, resulting in extremely low production efficiency and failing to meet the current industry demands for rapid iteration and mass production of promotional videos. Third, manual music composition struggles to accurately capture the emotional trajectory of the entire video, achieving only a rough match of the overall style. It cannot dynamically adapt to the scene's emotional and rhythmic changes in different shot segments, easily resulting in a disconnect between the emotional fluctuations of the video footage and the emotional changes of the music, severely impacting the overall viewing experience and communicative power of the promotional video. To address these issues, we propose an intelligent music matching system for promotional videos that integrates musical elements. Summary of the Invention
[0004] To address the aforementioned technical issues, an intelligent music matching system for corporate promotional videos that integrates musical elements is provided. This technical solution resolves the problem that easily occurs in manual production, where the emotional fluctuations in the video footage are disconnected from the emotional changes in the music, severely impacting the overall viewing experience and dissemination power of the promotional video.
[0005] To achieve the above objectives, the technical solution adopted by this invention is: an intelligent music matching system for corporate promotional videos that integrates musical elements, comprising: The video import module is used to import the original video of the corporate promotional video that needs to be set to music, perform shot segmentation and scene division on the video, extract the emotional tags of each shot segment, and generate a video emotional trajectory sequence. The music material module is used to store background music materials labeled with musical emotional elements. Each piece of background music material includes a curve of musical emotional change. The data processing module receives the video emotion trajectory sequence output by the video import module and the music emotion change curve output by the music material module, performs dynamic time warping and alignment on multiple granular time scales, calculates the matching degree between the video emotion trajectory and each candidate music emotion change curve, and selects the optimal matching background music. The music output module is used to synthesize and output the best matching music selected by the data processing module with the original video.
[0006] Preferably, the specific steps for the video import module to extract emotional tags from each shot segment and generate a video emotional trajectory sequence are as follows: Perform shot boundary detection on the original corporate promotional video and segment the video into multiple shot segments; Visual features are extracted from each shot segment, including tone distribution, amplitude of image movement, shot duration, and image composition features; The extracted visual features are input into a pre-trained emotion classification model based on convolutional neural networks and attention mechanisms, and the output is the emotion label and its confidence level corresponding to each shot segment. The emotion label includes at least one of excitement, tension, relaxation and calmness. The emotional tags of each shot are arranged in chronological order to generate a video emotional trajectory sequence.
[0007] Preferably, the music emotion change curve in the music material module is constructed as follows: Each piece of background music is divided into sections or phrases, and the musical characteristics of each section are extracted, including tonality, rhythm density, dynamic range and timbre brightness. Based on the musical features of each segment, a pre-trained music emotion mapping model based on multilayer perceptron is used to map the musical features into emotion tags, and the emotion tags of each segment are labeled. The emotion tags and the emotion tags on the video end use the same tag set. The emotional tags of each paragraph are arranged in chronological order to construct a musical emotional change curve, with time as the horizontal axis and emotional tags as the vertical axis.
[0008] Preferably, the multi-granularity time scale includes three levels: shot level, scene level, and full-film level; Shot-level alignment uses individual shot segments as matching units, matching the emotional labels of each shot in the video emotional trajectory sequence with the emotional labels within the corresponding time window of the music; Scene-level alignment uses scenes as the matching unit, matching the emotional trend of a scene composed of multiple consecutive shot clips with the emotional direction of the corresponding musical passages; Full-video alignment uses the entire video as the matching unit to calculate the similarity between the overall fluctuation pattern of the video's emotional trajectory and the overall trend of the music's emotional change curve.
[0009] Preferably, the dynamic time warping alignment calculation process is as follows: Each emotional label is mapped to coordinates in the valence-arousal two-dimensional emotional space, where excitement is mapped to a high arousal positive valence coordinate, tension is mapped to a high arousal negative valence coordinate, relaxation is mapped to a low arousal positive valence coordinate, and calmness is mapped to a low arousal negative valence coordinate. At the shot-level time scale, the emotional label sequence of each shot segment in the video emotional trajectory sequence is taken as one end, and the emotional label sequence of each segment in the corresponding time window in the candidate music emotional change curve is taken as the other end. Based on the valence-arousal two-dimensional coordinates, the Euclidean distance between the emotional labels at both ends is calculated, and a distance matrix is constructed. A dynamic programming algorithm is used to search for the optimal alignment path in the distance matrix, which minimizes the cumulative Euclidean distance between sentiment tags on the alignment path; At both the scene-level and full-film-level time scales, the same dynamic time warping calculations are performed on the scene-level emotional trend sequence and the full-film emotional fluctuation pattern to obtain the alignment path and cumulative distance value at each granularity.
[0010] Preferably, the specific method for calculating the matching degree between the video emotion trajectory and the emotion change curves of each candidate music is as follows: Based on the alignment path and cumulative distance value at each granularity, the lens-level matching degree, scene-level matching degree and full-film-level matching degree are calculated respectively; The matching degree at each granularity is weighted and fused, with the weight of the whole film level being higher than that of the scene level, and the weight of the scene level being higher than that of the shot level. The weighted fusion result is used as the comprehensive matching degree score between the video and the candidate soundtrack. All candidate soundtracks are sorted in descending order based on their overall matching score, and the candidate soundtrack with the highest overall matching score is selected as the optimal matching soundtrack.
[0011] Preferably, the music output module specifically includes: Based on the alignment path output by the data processing module, the start and end times and playback speed of the background music material in each video segment are determined; the playback speed adjustment adopts a time-domain stretching algorithm to achieve variable speed without changing pitch. At the video clip switching points, an audio fade-in / fade-out transition is set. The transition duration is dynamically determined based on the Euclidean distance between the emotional tags of adjacent clips in the valence-arousal space. When the Euclidean distance is less than 0.3, the transition duration is 0.5 seconds. When the Euclidean distance is between 0.3 and 0.8, the transition duration increases linearly to 2 seconds. When the Euclidean distance is greater than 0.8, the transition duration is 2 seconds. The background music audio is mixed and rendered with the original video audio track to output a composite video file.
[0012] Preferably, the video import module further extracts a visual beat feature sequence based on the lens boundary detection results. The visual beat features include the lens switching frequency, the peak value of the image motion amplitude, and the duration of key image dwell time. After the data processing module completes the dynamic time warping alignment, it anchors and verifies the beat markers of the candidate background music with the visual beat feature sequence, calculates the time deviation between the music downbeat time point and the most recent visual beat peak time point, and when the time deviation exceeds 200 milliseconds, it uses the deviation value as the time offset compensation amount to apply a time shift in the corresponding direction to the alignment path, and re-verifies the anchoring deviation until it is less than 200 milliseconds or the compensation iteration count reaches 3 times.
[0013] Preferably, before selecting the optimal matching background music, the data processing module also performs brand audio DNA constraint screening: The brand audio fingerprint is extracted from the company's existing brand materials. The brand audio fingerprint is extracted by extracting spectral envelope features through fast Fourier transform, extracting main frequency band preference through spectral peak tracking, and extracting rhythmic templates through start point detection algorithm. Using the brand's audio fingerprint as a matching constraint, the spectral similarity and rhythmic matching degree of each candidate background music with the brand's audio fingerprint are calculated. Candidate background music with both spectral similarity and rhythmic matching degree lower than the preset brand consistency threshold is eliminated, and the remaining candidate background music enters the subsequent dynamic time normalization alignment and matching degree calculation.
[0014] Preferably, each piece of background music in the music material module also includes layered audio track data, which includes independent audio tracks for melody, rhythm, harmony, and low frequency. During synthesis, the music output module performs adaptive gain adjustment on each independent audio track layer according to the emotional tags of the video segments: in segments with a soothing or calming emotional tag, the gain of the rhythm layer and low frequency layer is reduced by 3 to 6 dB, while the gain of the melody layer and harmony layer remains unchanged; in segments with an exciting or tense emotional tag, all audio track layers are activated and the dynamic range is increased by 6 to 12 dB, thus presenting differentiated audio track forms for the same music in different video segments.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention effectively solves the problems of traditional corporate promotional video music selection, such as reliance on manual work, low efficiency, poor matching accuracy, and insufficient dynamic adaptation of existing intelligent solutions, through shot segmentation, emotional trajectory extraction, music emotional curve annotation, and multi-granularity dynamic matching. The system can perform fine-grained scene and shot division of videos, construct a continuous and complete video emotional trajectory sequence, accurately capture changes in the emotions of the scene, and avoid the music from being out of sync with the scene's emotions. At the same time, it uses dynamic emotional curve annotation of music materials to replace traditional static label matching, achieving a high degree of consistency between the emotional fluctuations of the music and the rhythm and transitions of the video. Through multi-granularity time scale dynamic regularization and alignment, it can adapt to complex shot rhythms, greatly improve matching accuracy, and reduce quality instability caused by human subjective differences. The system automates the entire process of video analysis, music selection, and audio-visual synthesis, greatly simplifying the process, shortening the production cycle, reducing reliance on professional personnel, meeting the needs of batch and high-efficiency production, and significantly improving the audio-visual effects and production efficiency of corporate promotional videos. It has strong practicality and promotional value. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the system framework of the present invention. Detailed Implementation
[0017] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0018] Example 1: System Overall Architecture: The intelligent music matching system for corporate promotional videos that integrates musical elements, as described in this invention, is as follows: Figure 1 As shown, it includes four core functional modules: video import module, music material module, data processing module, and background music output module. The modules interact with each other through a standard data interface.
[0019] The video import module is responsible for receiving the original video files of the promotional video to be set to music, supporting decoding input of mainstream video encoding formats. This module integrates a shot segmentation unit and a scene division unit, automatically identifying shot boundaries in the video and aggregating consecutive shots into semantically coherent scene units. During shot segmentation, a dual-threshold comparison method based on inter-frame difference measurement is used to determine shot transition positions, where the inter-frame difference is calculated by weighting the pixel domain and feature domain. The scene division unit, based on the shot segmentation results, clusters adjacent shots according to visual feature similarity and semantic coherence, organizing the video stream into a multi-level scene structure.
[0020] The music resource module constructs a database of background music materials, storing professionally annotated background music files and their associated emotional element metadata. The data structure of each piece of background music includes a file identifier, music duration, original sampling rate, emotional annotation tag set, and pre-extracted music feature vectors. The emotional annotation of the background music materials is completed by professional music editors, following a predefined emotional classification system to ensure the consistency and reliability of the annotation results. This module supports a fast search function based on metadata, enabling the filtering of candidate background music based on emotional tags, rhythm type, duration range, and other criteria.
[0021] The data processing module, as the computational core of the system, receives video emotion trajectory sequences from the video import module and candidate background music emotion data from the music material module, and performs dynamic alignment and matching degree calculations at multiple granular time scales. This module is internally divided into three sub-modules: a feature extraction unit, an alignment calculation unit, and a matching decision unit. The feature extraction unit is responsible for converting the emotional representations of the video and music into numerical vectors in a unified coordinate space; the alignment calculation unit implements a dynamic time warping algorithm to construct the optimal time correspondence; and the matching decision unit performs weighted fusion of the multi-granularity matching results and outputs a sorted list of background music candidates.
[0022] The music output module receives the optimal music matching scheme selected by the data processing module, performs time alignment processing, fade-in / fade-out transition processing, and audio / video mixing rendering, and finally outputs a ready-to-use music video file. This module supports multiple output formats and resolutions to meet the technical requirements of different distribution channels.
[0023] The data flow between modules follows the following temporal logic: After the video import module completes video parsing, the extracted video emotional trajectory sequence is transmitted to the data processing module in the form of a vector sequence with timestamp index; the candidate background music library and its emotional change curve data output by the music material module according to the pre-screening conditions are synchronously transmitted to the data processing module; after the data processing module completes the matching decision, it transmits the time alignment parameters and gain adjustment parameters of the optimal background music to the background music output module to complete the final synthesis output.
[0024] Example 2: Video Emotion Trajectory Extraction The emotion trajectory extraction process in the video import module involves the cascading collaboration of multiple processing stages. This embodiment details the complete conversion process from the original video frame sequence to the emotion tag sequence.
[0025] The shot boundary detection stage employs an adaptive thresholding algorithm based on the content change rate. Let the video frame sequence be F = {f1, f2, ..., f...}. n The inter-frame difference D(i, i+1) is calculated as follows: First, for two adjacent frames fi and f {i+1} Extract the color histogram feature vector H respectively i and H {i+1} Then, the intersection of the normalized color histograms is calculated as a measure of visual difference: ; Simultaneously, the optical flow field feature vector is calculated, and the average optical flow amplitude between adjacent frames is used as a measure of motion difference D. motion (i, i+1). The overall difference is defined as the weighted sum of the two: ; The weighting coefficient α is adaptively adjusted according to the video content type. When D(i, i+1) exceeds the preset combination of two thresholds (low threshold Tlow and high threshold Thigh), a shot change is determined to occur; when D(i, i+1) is between the two thresholds and continues for a certain number of frames, it is determined to be a gradual switch.
[0026] In the visual feature extraction stage, for each detected shot segment, the following four-dimensional feature vectors are calculated: the hue distribution feature uses histogram statistics in the LAB color space to quantify the overall hue bias and saturation level of the image; the image motion amplitude feature is based on the statistics of dense optical flow field, calculating the mean, variance, and peak ratio of optical flow amplitude; the shot duration feature is directly taken as the ratio of the number of frames to the frame rate of the segment; the image composition feature is extracted using an aesthetic scoring model, including the deviation of the rule of thirds, the detection results of the gaze guide line, and the foreground-background contrast.
[0027] A pre-trained sentiment classification model based on CNN and attention mechanism receives the aforementioned visual feature vectors and outputs sentiment labels and confidence scores. The model architecture uses ResNet50 as the backbone network, fine-tuned for sentiment classification tasks based on ImageNet pre-trained weights. The set of sentiment labels output by the model is defined as L = {excited, tense, relaxed, calm}, consistent with the music sentiment label set for subsequent cross-modal alignment calculations.
[0028] For each frame of image, the model outputs the probability distribution P = {p1, p2, p3, p4} of which the frame belongs to each category, where p j Indicates belonging to label l j The probability of the probability is calculated. The label corresponding to the maximum probability is taken as the sentiment label of the frame, and the probability value is output as the confidence level of the sentiment label. To ensure the temporal smoothness of the sentiment labels, a Markov chain-like spatiotemporal constraint is introduced between adjacent frames, and the Viterbi algorithm is used to decode the globally optimal label sequence.
[0029] The formal expression for the video emotion trajectory sequence is: Tvideo = {t1, t2, ..., t} m}, where m is the number of shot segments, t i = (start i end i emotion i valence i arousal i ) represents the start time, end time, sentiment tag, and valence-arousal coordinates used for alignment calculation of the i-th shot.
[0030] Example 3: Construction of Musical Emotional Change Curve: The process of constructing the emotional change curve of the background music in the music material module includes three core steps: segmentation processing, feature extraction, and emotion mapping. After processing, each piece of background music generates a music emotion change curve representation that is compatible with the video emotion trajectory format.
[0031] In the segmentation stage, a music structure analysis algorithm is used to divide the background music material into semantically coherent musical units. The algorithm first performs beat detection, using a beat tracking method based on spectrum analysis to extract the beat time sequence. Then, it divides the music into measures based on the beat position information, calculating the feature differences between adjacent measure boundaries to detect musical phrase boundaries. The segmentation results are organized by measure or musical phrase, with each segment unit including a start timestamp, an end timestamp, and a segment boundary confidence score.
[0032] In the music feature extraction stage, the following four-dimensional feature vectors are calculated for each segment unit: Tonal pattern feature: the tonality, modulation and tonality ambiguity regions are identified through harmony analysis algorithm, and the tonal stability index is output; Rhythm density feature: the average and variance of beats per minute (BPM) within the segment are statistically analyzed to quantify the temporal density of the music; Dynamic range feature: the ratio of the peak value to the root mean square value of the audio signal amplitude within the segment is calculated to reflect the amplitude of the music's intensity variation; Timbre brightness feature: obtained through spectral centroid calculation, the higher the proportion of high-frequency energy, the brighter the timbre.
[0033] This MLP-based pre-trained music sentiment mapping model maps extracted music feature vectors to sentiment labels. The model structure employs a three-layer fully connected network. The input layer dimension matches the dimension of the music feature vector, the hidden layer contains 256 neurons using the ReLU activation function, and the output layer contains 4 neurons corresponding to four sentiment label categories. The model is trained under supervision on an annotated music dataset containing 5000 professionally annotated music tracks of various types.
[0034] The construction of the music emotion change curve follows the same data structure specifications as the video emotion trajectory. For the background music material M, after segmentation, n music segments are obtained, and the emotion representation of each segment is ej = (emotion... j valence j ,arousal j These segments are organized into a sequence in chronological order: ; This sequence shares the same coordinate representation system as the video emotion trajectory sequence. The valence dimension describes the pleasantness attribute of the emotion, with a value range of [-1, +1], while the arousal dimension describes the activation level of the emotion, also with a value range of [-1, +1]. The specific two-dimensional coordinate mapping relationship for the four emotion tags is as follows: Exhilarating: ; nervous: ; Soothing: ; Silence: ; In actual output, the sentiment mapping model may output continuous values between the typical coordinates mentioned above, reflecting subtle differences and gradual transitions in sentiment.
[0035] Example 4: Multi-granularity DTW alignment and matching degree calculation: The core algorithm in the data processing module is Dynamic Time Warping (DTW). This algorithm performs alignment calculations between the video emotion trajectory and the emotion change curves of the candidate background music at three time granularity levels. This embodiment details the algorithm principle and matching degree fusion strategy.
[0036] The core objective of the DTW algorithm is to find the optimal temporal correspondence between two sequences of unequal length. For a video sentiment trajectory sequence Tvideo = {t1, t2, ..., t...} m The sequence of emotional changes in the background music is Tmusic = {s1, s2,..., s}. n First, the emotional tags are mapped to numerical vectors in the valence-arousal two-dimensional coordinate space.
[0037] For any element t in the sequence i Its coordinates are represented as (v i , a i ), where v i For effectiveness, a i This is the wake-up rate value. Similarly, s j The coordinates are represented as (u j , bj The similarity between two elements is measured by Euclidean distance: ; Dynamic programming search is performed based on the distance matrix D = {d(i, j)}. Assume the cumulative distance matrix C satisfies the recurrence relation: ; The boundary condition is C(1, 1) = d(t1, s1), and the recursion endpoint is C(m, n). The minimum cumulative distance D obtained by dynamic programming search is... dtw This is the optimal alignment cost between the two sequences: ; Shot-level alignment calculations perform Time-to-Write (DTW) at the smallest temporal granularity, establishing a point-to-point correspondence between each video shot and its corresponding music segment. The shot-level processing unit traverses the candidate music library, performs a complete DTW calculation for each piece of music, and outputs a shot-level matching score for initial screening.
[0038] Scene-level alignment calculations are performed on scene units composed of consecutive shots. Let a scene Sk contain k consecutive shots, and its emotional trajectory be represented as Tscene = {t {p} , t {p+1} , ..., t {q}}, where q-p+1=k. Scene-level sentiment trends are represented using a sliding weighted average vector: ; Where the weight w i Using an exponential decay method, shots closer to the center of the scene receive higher weight: ; The parameter λ controls the weight decay rate, and was determined to be λ = 0.5 through validation set experiments. Scene-level DTW calculation aligns the trend vector of each scene with the trend vector of the corresponding time interval of the candidate soundtrack.
[0039] Full-video alignment calculation is based on the analysis of the emotional fluctuation patterns of the entire video. The full-video emotional fluctuation pattern feature vector is represented using statistical features: ; The components are defined as follows: mean valence, standard deviation of valence, mean arousal, standard deviation of arousal, correlation coefficient between valence and arousal, valence skewness, arousal skewness, valence kurtosis, and arousal kurtosis. The full-video DTW calculation aligns the overall statistical feature vector of the video with the overall statistical feature vector of the background music.
[0040] After the DTW calculations at each granularity level are completed, the cumulative distance values need to be converted into standardized matching scores. The conversion uses an exponential decay function: ; Where dnorm is the normalized distance value and α is the decay coefficient. The normalization benchmark is determined by a predefined distance threshold: the maximum possible distance between the centers of two labels in the sentiment space is defined as dmax = 2√2 (corresponding to sentiment labels of opposite polarity), then the normalized distance is dnorm = d / dmax.
[0041] The multi-granularity matching degree fusion adopts a weighted summation strategy, and the weight allocation reflects the contribution of different granularities to the overall matching decision: ; Where Mshot represents shot-level matching degree, Mscene represents scene-level matching degree, and Mglobal represents overall film-level matching degree. The weight parameters are set to w1 = 0.2, w2 = 0.3, and w3 = 0.5. This allocation scheme has been experimentally verified to highlight the consistency of the overall emotional trend while preserving local detail matching.
[0042] The data processing module performs the aforementioned multi-granularity DTW calculation and matching degree fusion on all candidate background music in the candidate music library, and outputs a list of candidate background music arranged in descending order of matching degree score. The matching decision unit selects the background music with the highest score as the optimal matching solution. If there are multiple candidates with similar scores (score difference less than 0.05), a manual review process is triggered for the user to select and confirm.
[0043] Example 5: Visual beat anchoring and brand DNA constraint screening: Before final synthesis, the music output module needs to perform visual beat anchoring verification to ensure that the time alignment accuracy between the musical beats and key visual rhythm points meets the requirements. Visual beat extraction involves the calculation of three types of features.
[0044] Shot transition frequency characteristics are statistically analyzed to determine the number of times shot boundaries appear per unit time, reflecting the basic rhythm density of the video. Let the video duration be T seconds, and the sequence of shot boundary time points be B = {b1, b2, ..., b...} k If fshot = k / T, then the shot switching frequency is fshot = k / T.
[0045] Motion amplitude peak feature detection of time points in regions of intense motion within a video frame sequence. The optical flow field amplitude is calculated between consecutive frames, and peak detection is performed on the amplitude sequence to output a set of peak time points P = {p1, p2, ..., p...}. q}
[0046] The key frame dwell time feature identification identifies the duration threshold of static or slowly moving frames in a video, which is used to determine the pause points in visual rhythm.
[0047] After DTW alignment is completed, the beat time sequence B in the background music is extracted based on the time correspondence determined by the alignment path. music = {bm1, bm2, ..., bm r Beat anchoring verification calculates the time deviation between the visual beat point and the nearest musical beat point: ,in ; in Let be the time of the i-th visual beat. When the maximum deviation Δtmax exceeds the 200-millisecond threshold, time shift compensation is triggered. ; Compensation amount Based on the direction of the deviation, shift in the direction that reduces the deviation. After updating the alignment path, recalculate the deviation value, and iteratively execute the above verification process until the accuracy requirement is met or the maximum number of iterations (3) is reached.
[0048] Before performing DTW calculations, the data processing module pre-screens candidate background music based on brand DNA constraints, retaining only those that match the brand's audio style requirements for subsequent matching. Brand DNA extraction comprises two independent dimensions: spectral envelope analysis and rhythmic template extraction.
[0049] Spectral envelope analysis uses Fast Fourier Transform (FFT) to extract the spectral information of the background music material. Let the audio signal be x(n), and the FFT result of length N points be X(k). The spectral envelope is extracted using the spectral peak tracking method. ; Where Δf is the bandwidth parameter. The peak energy of each main frequency band is extracted to form the spectral envelope vector S = [s1, s2, ..., s...]. K ], where K is the number of main frequency bands. Spectral similarity is calculated using cosine similarity: ; Spectral similarity threshold T spectrum The threshold is set to 0.65; candidate background music below this threshold will be directly eliminated.
[0050] Rhythm pattern extraction identifies rhythmic patterns in the background music using a start-point detection algorithm. This algorithm, based on spectral derivative analysis, identifies the starting points of energy abrupt changes. The start-point time interval sequence is extracted, and the average BPM value and BPM stability index are calculated to construct the rhythm pattern template R = {bpm, σbpm}. The rhythm pattern matching degree is calculated using the reciprocal of the BPM deviation. ; Rhythm matching threshold T rhythm The value is set to 0.70, and candidate background music is retained for the DTW calculation process only when the matching degree of both dimensions simultaneously meets their respective threshold requirements.
[0051] The brand DNA constraint screening employs an AND logic combination of two independent thresholds; rejection is triggered if either dimension falls below the threshold. This design allows for the opposite scenario where spectral features match but rhythmic patterns do not, validating the necessity of independent thresholds.
[0052] Example 6: Music Output Synthesis and Layered Track Adaptive Gain: The music output module performs time alignment processing on the music footage based on the time correspondence determined by the DTW alignment path. When there is a deviation between the video shot boundaries and the music beat positions, time-domain stretching technology is used for fine-tuning. ; Where β is the stretching coefficient, satisfying β = Tvideosegment / Tmusicsegment. The stretching process maintains the pitch unchanged, only adjusting the playback speed.
[0053] The fade-in / fade-out transition is dynamically adjusted based on the emotional distance between the video and the background music. Let d be the emotional distance at the edge of the shot. boundary Gradual transition duration Determine using the following formula: ; in The unit is seconds. When the emotional distance d is less than 0.3, a short transition duration of 0.5 seconds is used; when the emotional distance is between 0.3 and 0.8, the transition duration is linearly interpolated between 0.5 and 2 seconds; when the emotional distance is greater than 0.8, the maximum transition duration of 2 seconds is used. The transition effect is achieved through an exponential gain curve: ; The background music material has been layered and parsed during the storage phase, extracted into four independent audio tracks: melody, rhythm, harmony, and low-frequency. The layered audio track adaptive gain control dynamically adjusts the gain coefficient of each audio track based on the emotional type of the current video clip.
[0054] For video clips with a soothing or calming emotional label, the gain adjustment strategy is to reduce the energy of the rhythm and low-frequency layers while maintaining the original levels of the melody and harmony layers. Let the original gain be G0 = 0 dB, and the adjusted gain be: ; ; ; ; The attenuation ΔG reduce The attenuation is dynamically determined based on the emotional intensity of the video; the stronger the emotional intensity, the greater the attenuation.
[0055] For video clips labeled as emotionally intense or tense, the gain adjustment strategy is to activate all layers and enhance dynamic range: ; Where ΔG boost The same method is applied to all four audio tracks to achieve full-frequency energy enhancement.
[0056] The formula for adjusting the gain of layered audio tracks can be uniformly expressed as: ; Where ΔG is the gain adjustment amount jointly determined by the emotion type (evideo) and the audio track layer type (llayer). The functional relationship between emotion intensity and gain adjustment amount was determined through experimental optimization.
[0057] The music output module's mixing and rendering unit receives time alignment parameters, transition processing parameters, and layer gain parameters, and performs audio mixing and video encoding operations.
[0058] The audio mixing process first applies corresponding gain coefficients to the four layered audio tracks of the background music material, and then superimposes them to synthesize a stereo mixed signal. The mixed signal is then mixed with the audio track of the original video (if synchronous sound material exists) at a preset ratio of background music: synchronous sound = 0.85:0.15, which can be adjusted according to user needs.
[0059] The video encoding uses H.264 / AVC or H.265 / HEVC codecs, and the output resolution and bitrate are automatically adapted according to the target publishing platform. The mixed audio and video streams are encapsulated in MP4 container format for output. The output file includes background music metadata tags, recording the source of the background music, emotional annotation information, and matching score, facilitating subsequent traceability and retrieval.
[0060] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A smart music matching system for corporate promotional videos that integrates musical elements, characterized in that, include: The video import module is used to import the original video of the corporate promotional video that needs to be set to music, perform shot segmentation and scene division on the video, extract the emotional tags of each shot segment, and generate a video emotional trajectory sequence. The music material module is used to store background music materials labeled with musical emotional elements. Each piece of background music material includes a curve of musical emotional change. The data processing module receives the video emotion trajectory sequence output by the video import module and the music emotion change curve output by the music material module, performs dynamic time warping and alignment on multiple granular time scales, calculates the matching degree between the video emotion trajectory and each candidate music emotion change curve, and selects the optimal matching background music. The music output module is used to synthesize and output the best matching music selected by the data processing module with the original video. The multi-granularity timescale includes three levels: shot-level, scene-level, and full-film-level. Shot-level alignment uses individual shot segments as matching units, matching the emotional labels of each shot in the video emotional trajectory sequence with the emotional labels within the corresponding time window of the music; Scene-level alignment uses scenes as the matching unit, matching the emotional trend of a scene composed of multiple consecutive shot clips with the emotional direction of the corresponding musical passages; Full-video alignment uses the entire video as the matching unit to calculate the similarity between the overall fluctuation pattern of the video's emotional trajectory and the overall trend of the music's emotional change curve; The dynamic time warping alignment calculation process is as follows: Each emotional label is mapped to coordinates in the valence-arousal two-dimensional emotional space, where excitement is mapped to a high arousal positive valence coordinate, tension is mapped to a high arousal negative valence coordinate, relaxation is mapped to a low arousal positive valence coordinate, and calmness is mapped to a low arousal negative valence coordinate. At the shot-level time scale, the emotional label sequence of each shot segment in the video emotional trajectory sequence is taken as one end, and the emotional label sequence of each segment in the corresponding time window in the candidate music emotional change curve is taken as the other end. Based on the coordinates in the valence-arousal two-dimensional emotional space, the Euclidean distance between the two emotional labels is calculated, and a distance matrix is constructed. A dynamic programming algorithm is used to search for the optimal alignment path in the distance matrix, which minimizes the cumulative Euclidean distance between sentiment tags on the alignment path; At both the scene-level and full-film-level time scales, the same dynamic time warping calculations are performed on the scene-level emotional trend sequence and the full-film emotional fluctuation pattern to obtain the alignment path and cumulative distance value at each granularity.
2. The intelligent music matching system for corporate promotional videos integrating musical elements as described in claim 1, characterized in that, The specific steps for the video import module to extract emotional tags from each shot segment and generate a video emotional trajectory sequence are as follows: Perform shot boundary detection on the original corporate promotional video and segment the video into multiple shot segments; Visual features are extracted from each shot segment, including tone distribution, amplitude of image movement, shot duration, and image composition features; The extracted visual features are input into a pre-trained emotion classification model based on convolutional neural networks and attention mechanisms, and the output is the emotion label and its confidence level corresponding to each shot segment. The emotion label includes at least one of excitement, tension, relaxation and calmness. The emotional tags of each shot are arranged in chronological order to generate a video emotional trajectory sequence.
3. The intelligent music matching system for corporate promotional videos integrating musical elements as described in claim 1, characterized in that, The music emotion change curve in the music material module is constructed as follows: Each piece of background music is divided into sections or phrases, and the musical characteristics of each section are extracted, including tonality, rhythm density, dynamic range and timbre brightness. Based on the musical features of each segment, a pre-trained music emotion mapping model based on multilayer perceptron is used to map the musical features into emotion tags, and the emotion tags of each segment are labeled. The emotion tags and the emotion tags on the video end use the same tag set. The emotional tags of each paragraph are arranged in chronological order to construct a musical emotional change curve, with time as the horizontal axis and emotional tags as the vertical axis.
4. The intelligent music matching system for corporate promotional videos integrating musical elements as described in claim 1, characterized in that, The specific steps for calculating the matching degree between the video emotion trajectory and the emotion change curves of each candidate music are as follows: Based on the alignment path and cumulative distance value at each granularity, the lens-level matching degree, scene-level matching degree and full-film-level matching degree are calculated respectively; The matching degree at each granularity is weighted and fused, with the weight of the whole film level being higher than that of the scene level, and the weight of the scene level being higher than that of the shot level. The weighted fusion result is used as the comprehensive matching degree score between the video and the candidate soundtrack. All candidate soundtracks are sorted in descending order based on their overall matching score, and the candidate soundtrack with the highest overall matching score is selected as the optimal matching soundtrack.
5. The intelligent music matching system for corporate promotional videos integrating musical elements as described in claim 1, characterized in that, The music output module specifically includes: Based on the alignment path output by the data processing module, the start and end times and playback speed of the background music material in each video segment are determined; the playback speed adjustment adopts a time-domain stretching algorithm to achieve variable speed without changing pitch. At the video clip switching points, an audio fade-in / fade-out transition is set. The transition duration is dynamically determined based on the Euclidean distance between the emotional tags of adjacent clips in the valence-arousal space. When the Euclidean distance is less than 0.3, the transition duration is 0.5 seconds. When the Euclidean distance is between 0.3 and 0.8, the transition duration increases linearly to 2 seconds. When the Euclidean distance is greater than 0.8, the transition duration is 2 seconds. The background music audio is mixed and rendered with the original video audio track to output a composite video file.
6. The intelligent music matching system for corporate promotional videos that integrates musical elements as described in claim 1, characterized in that: The video import module also extracts visual beat feature sequences based on the lens boundary detection results. The visual beat features include lens switching frequency, peak value of image motion amplitude, and duration of key image dwell time. After the data processing module completes the dynamic time warping alignment, it anchors and verifies the beat markers of the candidate background music with the visual beat feature sequence, calculates the time deviation between the music downbeat time point and the most recent visual beat peak time point, and when the time deviation exceeds 200 milliseconds, it uses the deviation value as the time offset compensation amount to apply a time shift in the corresponding direction to the alignment path, and re-verifies the anchoring deviation until it is less than 200 milliseconds or the compensation iteration count reaches 3 times.
7. The intelligent music matching system for corporate promotional videos integrating musical elements as described in claim 1, characterized in that: Before selecting the optimal matching background music, the data processing module also performs brand audio DNA constraint screening: The brand audio fingerprint is extracted from the company's existing brand materials. The brand audio fingerprint is extracted by extracting spectral envelope features through fast Fourier transform, extracting main frequency band preference through spectral peak tracking, and extracting rhythmic templates through start point detection algorithm. Using the brand's audio fingerprint as a matching constraint, the spectral similarity and rhythmic matching degree of each candidate background music with the brand's audio fingerprint are calculated. Candidate background music with both spectral similarity and rhythmic matching degree lower than the preset brand consistency threshold is eliminated, and the remaining candidate background music enters the subsequent dynamic time normalization alignment and matching degree calculation.
8. The intelligent music matching system for corporate promotional videos that integrates musical elements as described in claim 1, characterized in that: Each piece of background music in the music material module also contains layered audio track data, which includes independent audio tracks for melody layer, rhythm layer, harmony layer and low frequency layer. During synthesis, the music output module performs adaptive gain adjustment on each independent audio track layer according to the emotional tags of the video segments: in segments with a soothing or calming emotional tag, the gain of the rhythm layer and low frequency layer is reduced by 3 to 6 dB, while the gain of the melody layer and harmony layer remains unchanged; in segments with an exciting or tense emotional tag, all audio track layers are activated and the dynamic range is increased by 6 to 12 dB, thus presenting differentiated audio track forms for the same music in different video segments.
Citation Information
Patent Citations
Video generation method and system based on music rhythm and television on-demand method
CN118301382A