Advertising video rhythm analysis system
Through audio-visual modal separation and comprehensive analysis, the best sense of rhythm of advertising video is obtained, and the problem of ignoring rhythm in the existing technology is solved, and the automation optimization and effect improvement of advertising videos is achieved.
Patent Information
- Application Number
- CN202411534042.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-10-31
AI Technical Summary
The prior art ignores the importance of video rhythm in advertising analysis and lacks effective analytical methods to evaluate and optimize the rhythm of advertising videos.
The audio-visual modal separation module separates the advertising video into visual image information and auditory audio information, and performs image analysis and auditory analysis respectively to form a normalized significant frame difference sequence and pitch difference sequence, and obtains the best frame difference threshold and pitch difference threshold through a comprehensive analysis function to achieve the best rhythm coordination between video and audio.
It realizes automation, intelligent evaluation and optimization of the rhythmic sense of advertising videos, avoids manual intervention to the greatest extent, improves the attractiveness of advertising and information transmission efficiency, and enhances the brand image.
Smart Images

Figure CN119383371B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of advertising video analysis, and in particular to an advertising video rhythm analysis system. Background Art
[0002] In modern advertising, video ads, as a crucial medium, are crucial for capturing audience attention and enhancing message delivery. Traditional advertising analysis often focuses on video content, image quality, and creativity, while overlooking the importance of rhythm in advertising videos.
[0003] However, rhythm in advertising videos isn't just a technical detail; it's a core element of effective communication. Through appropriate rhythmic arrangement, ads can better engage viewers, enhance message delivery, strengthen brand image, and ultimately boost sales and increase brand loyalty. Therefore, creatives must prioritize the grasp and application of rhythm when designing their videos.
[0004] In order to improve the effectiveness of advertising, there is an urgent need for an efficient analysis method that can objectively analyze and evaluate the rhythm of videos. Summary of the Invention
[0005] The present invention provides a method to effectively solve the above-mentioned problems in the prior art.
[0006] Specifically, the present invention provides an advertising video rhythm analysis system for analyzing an advertising video comprising an N-frame video image sequence, where N is an integer and N>3. The system comprises an audiovisual modality separation module, an image analysis module, an auditory analysis module, and a comprehensive analysis module. The audiovisual modality separation module separates the advertising video into visual image information and auditory audio information. The visual image information is input into the image analysis module, which forms a normalized significant frame difference sequence Δp'1, Δp'2, ..., Δp' based on the N-frame video image sequence and a set frame difference threshold VP. M The auditory audio information is input into the auditory analysis module. In the auditory analysis module, the duration of the auditory audio information is evenly divided into N time periods, and a normalized significant pitch difference sequence Δh'1, Δh'2, ..., Δh' is formed based on the N time periods according to the set pitch difference threshold VH. m , normalized significant frame difference sequence Δp'1, Δp'2, ..., Δp' M and the normalized audio significant pitch difference sequence Δh'1, Δh'2, ..., Δh' mAll are input into the comprehensive analysis module, and the comprehensive analysis function is f(Δp, Δh, VP, VH). Here, Δp represents the normalized video significant frame difference sequence, Δh represents the normalized audio significant pitch difference sequence, VP represents the frame difference threshold as an independent variable, and VH represents the pitch difference threshold as an independent variable. Obtain the extreme points VP0 and VH0 of the comprehensive analysis function f(Δp, Δh, VP, VH). Input the frame difference threshold VP0 into the image analysis module to obtain the best normalized video significant frame difference sequence Δp0. Input the obtained pitch difference threshold VH0 into the auditory analysis module to obtain the best normalized audio significant pitch difference sequence Δh0.
[0007] Preferably, the image analysis module calculates the normalized significant frame difference sequences Δp’1, Δp’2, …, Δp’ through the following steps M : The image analysis module sets the image information of any k-th frame in the video image sequence from the 1st frame to the N-th frame as P k , where k is an integer and 1 ≤ k ≤ N - 1, and calculates the frame difference ΔP between the image information of the (k + 1)-th frame and the k-th frame k :
[0008] ΔP k =||P k+1 -P k ||
[0009] As the value of k ranges from 1 to N - 1, N - 1 frame differences ΔP1, ΔP2, …, ΔP are formed N-1 , extract all the frame differences greater than the frame difference threshold VP among the N - 1 frame differences to form M significant frame differences, that is, ΔP’1, ΔP’2, …, ΔP’ M , where M < N - 1. Divide each term of the M significant frame differences by the frame difference threshold VP to form the normalized significant frame difference sequences Δp’1, Δp’2, …, Δp’ M .
[0010] Preferably, in the auditory analysis module, the duration of the auditory audio information is evenly divided into N time periods to form a pitch sequence containing N pitches. The pitch of any j-th time period from the 1st time period to the N-th time period is H j , where j is an integer and 1 ≤ j ≤ N - 1, and calculate the pitch difference ΔH between the pitch H j of the (j + 1)-th time period and the pitch H j+1 of the j-th time period j : ΔH j =||H j+1 -H j ||. As the value of j ranges from 1 to N - 1, N - 1 pitch differences ΔH1, ΔH2, …, ΔH are formed N-1 , extract the N - 1 pitch differences ΔH1, ΔH2, …, ΔHN-1 All pitch differences greater than the pitch difference threshold VH in [the relevant content] form m significant pitch differences ΔH’1, ΔH’2, …, ΔH’ m , where m < N - 1. Divide each of the m significant pitch differences ΔH’1, ΔH’2, …, ΔH’ m by the pitch difference threshold VH to form a normalized significant pitch difference sequence Δh’1, Δh’2, …, Δh’ m .
[0011] Preferably, the comprehensive analysis function f(Δp, Δh, VP, VH) forms a two - dimensional surface in the x - y - z three - dimensional space. In the x - y - z three - dimensional space, the x - axis represents the value of the frame difference threshold VP, the y - axis represents the value of the pitch difference threshold VH, and the z - axis represents the function value of the comprehensive analysis function.
[0012] Preferably, the comprehensive analysis function f(Δp, Δh, VP, VH) is defined as:
[0013]
[0014] In the above formula, i is the number of each item in the normalized video significant frame difference sequence, i ranges from 1 to M, and each item in the normalized video significant frame difference sequence is expressed as Δp’ i ; n is the number of each item in the normalized audio significant pitch difference sequence, n ranges from 1 to m, and each item in the normalized audio significant pitch difference sequence is expressed as Δh’ n . Additionally, the s function in the summation formula above is defined as:
[0015] [[]]s(Δp′ i ,VP)=(Δp′ i - VP)×sigmoid(Δp′ i - VP);
[0016] s(Δh′ n ,VH)=(Δh' n - VH)×sigmoid(Δh' n - VH).
[0017] Preferably, the comprehensive analysis function f(Δp, Δh, VP, VH) is defined as:
[0018]
[0019] In the above formula, i is the number of each item in the normalized video significant frame difference sequence, i ranges from 1 to M, and each item in the normalized video significant frame difference sequence is expressed as Δp’ i; n is the number of each item in the normalized audio significant pitch difference sequence, n ranges from 1 to m, and each item in the normalized audio significant pitch difference sequence is expressed as Δh' n .
[0020] Preferably, the comprehensive analysis function f(Δp, Δh, VP, VH) is defined as:
[0021]
[0022] In the above formula, i is the number of each item in the normalized video significant frame difference sequence, i ranges from 1 to M, and each item in the normalized video significant frame difference sequence is expressed as Δp' i ; n is the number of each item in the normalized audio significant pitch difference sequence, n ranges from 1 to m, and each item in the normalized audio significant pitch difference sequence is expressed as Δh' n .
[0023] In summary, the present invention provides a system for analyzing the rhythm of advertising videos. Initially, a fixed advertisement video segment is split into visual image information and auditory audio information. An image analysis module then identifies clipping points on the visual image information, thereby forming a normalized sequence of significant video frame differences. In parallel, an audio analysis module performs time-segment pitch clipping on the auditory audio information, thereby forming a normalized sequence of significant pitch differences. Both sequences are input into a comprehensive analysis module, which uses a comprehensive analysis function to generate a two-dimensional surface. Based on this two-dimensional surface feedback, optimal frame and pitch difference thresholds are set to maximize the rhythmic coordination between the video and audio images. In the system provided by the present invention, audio and video information are processed separately and then parallelized to achieve the same end result, effectively achieving optimal rhythmic acquisition for video files. This minimizes manual intervention while maintaining automated and intelligent control of the rhythm of the advertising video. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be discussed below. Obviously, the technical solutions described in conjunction with the drawings are only some embodiments of the present invention. For ordinary technicians in this field, other embodiments and their drawings can be obtained based on the embodiments shown in these drawings without paying any creative work.
[0025] Figure 1 The system operation flow chart of the advertising video rhythm analysis system provided by the present invention is shown.
[0026] Figure 2 A two-dimensional surface example of a function run by the comprehensive analysis module in the advertising video rhythm analysis system provided by the present invention is shown. DETAILED DESCRIPTION
[0027] The following will clearly and completely describe the technical solutions of various embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments described in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] The specific structure and operation of the distributed crawler scheduling system based on server performance monitoring provided by the present invention will be described in detail below with reference to the accompanying drawings.
[0029] Figure 1 The system operation flow chart of the advertising video rhythm analysis system provided by the present invention is shown.
[0030] like Figure 1 As shown, a video advertisement is first introduced into the audiovisual modality separation module for audiovisual separation, thereby separating the video advertisement into visual image information and auditory audio information. Subsequent system modules perform both independent analysis of the visual image information and the auditory audio information, as well as comprehensive adjustments to these two types of information.
[0031] The visual image information is input into the image analysis module for clipping point recognition. Clipping point recognition can be achieved using an inter-frame difference method. Specifically, the inter-frame recognition method is used to perform a difference operation on two adjacent frames in the video image sequence formed by the visual graphic information.
[0032] For example, a video image sequence can be divided into N frames (N>3, in practice, N is often a large integer value, such as N≥10000, or even N≥100000 is normal), and the image information of any k-th frame in the video image sequence from the 1st frame to the Nth frame is summarized as P k , where k is an integer and 1≤k≤N-1.
[0033] The frame difference ΔP between the k+1th frame image information and the kth frame image information can be calculated. k :
[0034] ΔP k =||P k+1 -P k ||
[0035] As the value of k increases from 1 to N-1, N-1 frame differences ΔP1, ΔP2, ..., ΔP are formed. N-1 . Set the frame difference threshold VP>0.
[0036] Extract N-1 frame differences ΔP1, ΔP2, ..., ΔPN-1 All frame differences greater than the frame difference threshold VP form M significant frame differences, namely, ΔP'1, ΔP'2, ..., ΔP' M , where M <N-1。
[0037] This process aims to screen the frame differences, thereby selecting significant frame differences with larger values from N-1 frame differences using a frame difference threshold as a boundary.
[0038] The M significant frame differences ΔP'1, ΔP'2, ..., ΔP' M Each item in is divided by the frame difference threshold VP to form a normalized significant frame difference sequence Δp'1, Δp'2, ..., Δp' M , where Δp'1=ΔP'1 / VP, Δp'2=ΔP'2 / VP, and so on, until Δp' M =ΔP' M / VP.
[0039] Therefore, the image analysis module outputs the normalized video significant frame difference sequence Δp'1, Δp'2, ..., Δp' M .
[0040] In parallel, the auditory audio information is input into the auditory analysis module for rhythm point recognition. The auditory audio information forms a pitch sequence. In other words, the auditory audio information forms different degrees of pitch at each time point, and these pitches are arranged into a pitch sequence in time order.
[0041] The duration of the auditory audio information can also be divided into N time periods, each of which has a pitch, thus forming a pitch sequence containing N pitches. The pitch of any j-th time period from the 1st time period to the Nth time period is H j , where j is an integer and 1≤j≤N-1.
[0042] The pitch H of the j+1th period is calculated from this j and H in period j j+1 The pitch difference ΔH j :
[0043] ΔH j =||H j+1 -H j ||
[0044] As the value of j changes from 1 to N-1, N-1 pitch differences ΔH1, ΔH2, ..., ΔH are formed. N-1 . Set the pitch difference threshold VH>0.
[0045] Extract N-1 pitch differences ΔH1, ΔH2, ..., ΔH N-1All pitch differences greater than the pitch difference threshold VH form m significant pitch differences, i.e., ΔH'1, ΔH'2, ..., ΔH' m , where m <N-1。
[0046] This process aims to screen the pitch differences, thereby selecting significant pitch differences with larger values from N-1 pitch differences based on the pitch difference threshold.
[0047] The m significant pitch differences ΔH'1, ΔH'2, ..., ΔH' m Each term in is divided by the pitch difference threshold VH to form a normalized significant pitch difference sequence Δh'1, Δh'2, ..., Δh' m , where Δp'1=ΔH'1 / VH, Δh'2=ΔH'2 / HP, and so on, until Δh' m =ΔH' m / VH.
[0048] Thus, the audio analysis module outputs a normalized audio significant pitch difference sequence Δh'1, Δh'2, ..., Δh' m .
[0049] Next, normalize the video significant frame difference sequence Δp'1, Δp'2, ..., Δp' M and the normalized audio significant pitch difference sequence Δh'1, Δh'2, ..., Δh' m All are input into the comprehensive analysis module.
[0050] A comprehensive analysis function f(Δp, Δh, VP, VH) is set in the comprehensive analysis module, where Δp represents the normalized video significant frame difference sequence, Δh represents the normalized audio significant pitch difference sequence, and VP and VH have been mentioned above. VP represents the frame difference threshold, and VH represents the pitch difference threshold.
[0051] As mentioned above, while the original input advertisement video remains unchanged, the normalized video significant frame difference sequence will change with different values of the frame difference threshold VP. Conversely, the normalized audio significant pitch difference sequence will also change with different values of the pitch difference threshold VH. Therefore, in the comprehensive analysis function f(Δp, Δh, VP, VH), the frame difference threshold VP and the pitch difference threshold VH constitute the two independent variables of the function.
[0052] Therefore, as the values of these two independent variables are different, the comprehensive analysis function f(Δp, Δh, VP, VH) forms a two-dimensional surface in the xyz three-dimensional space, where the x-axis represents the value of the frame difference threshold VP, the y-axis represents the value of the pitch difference threshold VH, and the z-axis represents the function value of the comprehensive analysis function.
[0053] In this two-dimensional surface, there are extreme points VP0 and VH0 of the function values of the frame difference threshold VP and the pitch difference threshold VH. The audio sequence and image sequence formed at the extreme points indicate that the image sequence and the pitch sequence have achieved the best fit. At this time, the frame difference threshold VP is VP0, and the pitch difference threshold VH is VH0. By inputting the frame difference threshold VP0 into the image analysis module, the best normalized video significant frame difference sequence Δp0 can be obtained. By inputting the pitch difference threshold VH0 into the auditory analysis module, the best normalized audio significant pitch difference sequence Δh0 can be obtained.
[0054] By observing the optimal normalized video significant frame difference sequence Δp0 and the optimal normalized audio significant pitch difference sequence Δh0, users can obtain the optimal rhythm of the video image in the audio environment, and can also obtain the optimal rhythm of the auditory audio in the current video image environment.
[0055] The comprehensive analysis function can be defined according to specific circumstances. For example, the comprehensive analysis function f(Δp, Δh, VP, VH) can be defined as follows:
[0056]
[0057] In the above formula, i is the number of each item in the normalized video significant frame difference sequence, i ranges from 1 to M, and each item in the normalized video significant frame difference sequence is expressed as Δp' i ; n is the number of each item in the normalized audio significant pitch difference sequence, n ranges from 1 to m, and each item in the normalized audio significant pitch difference sequence is expressed as Δh' n .
[0058] In the above formula, each term in the previous and next summation is an s function, where:
[0059] s(Δp′ i ,VP)=(Δp′ i -VP)×sigmoid(Δp′ i -VP);
[0060] s(Δh' n ,VH)=(Δh' n -VH)×sigmoid(Δh′ n -VH).
[0061] The above function forms a two-dimensional surface in the xyz three-dimensional space with respect to VP and VH, where the X-axis represents the value of the frame difference threshold VP, the Y-axis represents the value of the pitch difference threshold VH, and the Z-axis represents the function value of the comprehensive analysis function. The shape of the two-dimensional surface is as follows: Figure 2 shown.
[0062] As mentioned above, the comprehensive analysis function is not limited to the above form. For example, it can also be defined as follows:
[0063]
[0064] Compared with the previous formula, this formula adds absolute value symbols on both sides of each summation term. The definitions of specific parameters are exactly the same, but the shape of the two-dimensional surface in the three-dimensional space is changed.
[0065] Alternatively, you can also consider transforming the above comprehensive analysis function into the following form:
[0066]
[0067] Compared with the previous formula, this formula is to take the square of each summation term. The definitions of specific parameters are exactly the same, but the shape of the two-dimensional surface in the three-dimensional space is changed.
[0068] This concludes the basic introduction to the technical solution of the present invention. In summary, the present invention provides a system for analyzing the rhythm of advertising videos. Initially, a fixed advertising video segment is split into visual image information and auditory audio information. An image analysis module then performs clipping point identification on the visual image information, thereby forming a normalized sequence of significant video frame differences. In parallel, an audio analysis module performs time-segment pitch clipping on the auditory audio information, thereby forming a normalized sequence of significant pitch differences. Both sequences are input into a comprehensive analysis module, which uses a comprehensive analysis function to generate a two-dimensional surface. Based on this two-dimensional surface feedback, optimal frame difference and pitch difference thresholds are set, thereby achieving the maximum rhythmic coordination between the video and audio images. In the system provided by the present invention, audio and video information are processed separately and then parallelized to achieve the same end result, effectively achieving optimal rhythmic acquisition for video files. This minimizes manual intervention while maintaining automated and intelligent control of the rhythm of advertising videos.
[0069] The above description is merely an exemplary embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A system for analyzing the rhythm of an advertising video, for analyzing an advertising video comprising a sequence of N video frames, where N is an integer and N>3, characterized in that: The system includes an audio-visual modality separation module, an image analysis module, an auditory analysis module, and a comprehensive analysis module, among which: The audio-visual modality separation module separates the advertisement video into visual image information and auditory audio information. The visual image information is input into the image analysis module, which forms a normalized significant frame difference sequence Δp'1, Δp'2, ..., Δp' based on the N-frame video image sequence according to the set frame difference threshold VP. M , The auditory audio information is input into the auditory analysis module. In the auditory analysis module, the duration of the auditory audio information is evenly divided into N time periods, and a normalized significant pitch difference sequence Δh'1, Δh'2, ..., Δh' is formed based on the N time periods according to the set pitch difference threshold VH. m , Normalized significant frame difference sequence Δp'1, Δp'2, ..., Δp' M and the normalized audio significant pitch difference sequence Δh'1, Δh'2, ..., Δh' m are all input into the comprehensive analysis module, in which a comprehensive analysis function f(Δp, Δh, VP, VH) is set, wherein Δp represents a normalized video significant frame difference sequence, Δh represents a normalized audio significant pitch difference sequence, VP represents a frame difference threshold as an independent variable, and VH represents a pitch difference threshold as an independent variable. The frame difference threshold extreme point VP0 and the pitch difference threshold extreme point VH0 in the comprehensive analysis function f(Δp, Δh, VP, VH) are taken. The frame difference threshold extreme point VP0 is input into the image analysis module to obtain the optimal normalized video significant frame difference sequence Δp0, and the pitch difference threshold extreme point VH0 is input into the auditory analysis module to obtain the optimal normalized audio significant pitch difference sequence Δh0.
2. The system according to claim 1, wherein: The image analysis module calculates the normalized significant frame difference sequence Δp'1, Δp'2, ..., Δp' through the following steps: M : The image analysis module sets the image information of any kth frame in the video image sequence from the 1st frame to the Nth frame as P k , where k is an integer and 1≤k≤N-1, calculate the frame difference ΔP between the k+1th frame image information and the kth frame image information k : ΔP k =‖P k+1 -P k ‖ As the value of k increases from 1 to N-1, N-1 frame differences ΔP1, ΔP2, ..., ΔP are formed. N-1 , extract all the frame differences greater than the frame difference threshold VP from the N-1 frame differences to form M significant frame differences, i.e., ΔP'1, ΔP'2, ..., ΔP' M , where M <N-1, Each of the M significant frame differences is divided by the frame difference threshold VP to form a normalized significant frame difference sequence Δp'1, Δp'2, ..., Δp' M .
3. The system according to claim 1, wherein: In the auditory analysis module, the duration of the auditory audio information is evenly divided into N time periods, thereby forming a pitch sequence containing N pitches. The pitch of any j-th time period from the 1st time period to the Nth time period is H j , where j is an integer and 1≤j≤N-1, calculate the pitch H of the j+1th time period j+1 and H in period j j The pitch difference ΔH j :ΔH j =||H j+1 -H j ||, as the value of j changes from 1 to N-1, N-1 pitch differences ΔH1, ΔH2, ..., ΔH are formed N-1 , Extract N-1 pitch differences ΔH1, ΔH2, ..., ΔH N-1 All pitch differences greater than the pitch difference threshold VH form m significant pitch differences ΔH'1, ΔH'2, ..., ΔH' m , where m <N-1, The m significant pitch differences ΔH'1, ΔH'2, ..., ΔH' m Each term in is divided by the pitch difference threshold VH to form a normalized significant pitch difference sequence Δh'1, Δh'2, ..., Δh' m .
4. The system according to claim 1, wherein: The comprehensive analysis function f(Δp, Δh, VP, VH) forms a two-dimensional surface in the xyz three-dimensional space, where the x-axis represents the value of the frame difference threshold VP, the y-axis represents the value of the pitch difference threshold VH, and the z-axis represents the function value of the comprehensive analysis function.
5. The system according to claim 4, characterized in that The comprehensive analysis function f(Δp, Δh, VP, VH) is defined as: In the above formula, i is the number of each item in the normalized video significant frame difference sequence, i ranges from 1 to M, and each item in the normalized video significant frame difference sequence is expressed as Δp' i ; n is the number of each item in the normalized audio significant pitch difference sequence, n ranges from 1 to m, and each item in the normalized audio significant pitch difference sequence is expressed as Δh' n , In addition, the s function in the summation of the above formula is defined as: s(Δp′ i ,VP)=(Δp′ i -VP)×sigmoid(Δp′ i -VP); s(Δh′ n ,VH)=(Δh′ n -VH)×sigmoid(Δh′ n -VH)。
Citation Information
Patent Citations
Video key position determination method and device
CN107222746A
Main melody automatic extraction algorithm based on probability model
CN112735365A