Traditional media advertisement separation and identification method based on multi-modal large model

Through the multimodal large model method, the traditional media advertising content is automatically identified and extracted, which solves the problem of time-consuming and manual dependence in traditional media advertising monitoring, and achieves efficient and accurate advertising monitoring.

CN120340545APending Publication Date: 2025-07-18江苏省广告监测中心
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510483820.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the traditional media advertising monitoring, advertising extraction takes a long time and relies on manual operations, resulting in inefficient efficiency and easy human error, making it difficult to meet the needs of fast and accurate monitoring.

Method used

Using a multimodal large model-based method, through multimodal preprocessing, timing feature enhancement, large language model retrieval and boundary optimization, the advertising content and boundary are automatically extracted, combined with neural networks and attention mechanisms, the automatic recognition and high-precision extraction of advertisements are realized.

Benefits of technology

It significantly improves the efficiency and accuracy of advertising monitoring, reduces manual intervention and errors, and can quickly process a large number of TV and radio program content, providing support for advertising supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340545A_ABST
    Figure CN120340545A_ABST
Patent Text Reader

Abstract

The invention discloses a traditional media advertisement separation and identification method based on a multi-modal large model, and the method comprises the steps: S1, extracting advertisement audio and video signals in a television and a broadcast, and carrying out the preprocessing of the advertisement audio and video signals; s2, extracting audio features of the television and the broadcast from the preprocessed audio and video signals, finding out a subset of the longest same fragment from continuous videos in a feature extraction and feature retrieval mode, and recording the head and the tail of the time point; s3, building a traditional media advertisement knowledge base by using previous artificial experience data, judging whether the fragment is an advertisement or not by applying a large language model and an enhanced retrieval mode, and extracting a text boundary of the advertisement; and S4, revising the beginning and ending boundary of the advertisement by using a VAD mute detection and speaker synchronization mode. According to the method, in the television and broadcast feature extraction process, the neural network and the attention mechanism are applied, the problem that the traditional media audio time span is long is effectively solved, the time sequence features of the audio can be better extracted, and the advertisement matching accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of advertising monitoring, and particularly relates to a method for separating and identifying traditional media advertisements based on a multimodal large model, which is used to automatically and efficiently separate and identify advertisement content from traditional media programs such as television and radio. Background Technique

[0002] Advertisements, as an important indicator of economic development, their publicity orientation is of crucial importance. The monitoring of traditional media advertisements, especially television and radio advertisements, is an important link in advertisement supervision. In the process of traditional media advertisement monitoring, extracting advertisement content segments from continuously played television and radio programs is the most time-consuming and crucial task.

[0003] Currently, it mainly relies on manual viewing of programs and manually marking the start and end times of advertisements, and then cropping out single advertisement versions. Although the repetition of advertisement placement is also utilized, and audio feature search is used to identify repeated advertisements, the first production and marking of new advertisements still rely on manual operations. This not only consumes a large amount of manpower and time but also is prone to human errors.

[0004] In terms of advertisement extraction in the prior art, although there are methods that utilize audio features (such as audio fingerprints, MFCC, etc.), visual features (such as specific colors, layouts in images), and specific neural network encoders, decoders, or Transformers for feature extraction and matching, they still cannot get rid of the dependence on manual processing for the first marking of new advertisements. These technologies are inefficient when faced with a large amount of television and radio program content and are difficult to meet the requirements of rapid and accurate advertisement monitoring. Summary of the Invention

[0005] The present invention aims to provide a method for separating and identifying traditional media advertisements based on a multimodal large model. By integrating various modal information and leveraging the powerful retrieval and judgment capabilities of the large model, it automatically extracts advertisement content and advertisement boundaries from continuous television and radio, greatly improving the efficiency of traditional media advertisement separation, reducing labor costs, and enhancing the accuracy and timeliness of advertisement monitoring.

[0006] To solve the above technical problems, the present invention is implemented as follows:

[0007] A method for separating and identifying traditional media advertisements based on a multimodal large model includes the following steps:

[0008] Step S1: Multimodal preprocessing process: Extract advertisement audio-visual signals from television and radio and perform preprocessing on them;

[0009] Step S2: Temporal Feature Enhancement: Extract the audio features of TV and radio from the preprocessed audio-visual signals, find the subset of the longest identical segments from the continuous video through feature extraction and feature retrieval, and record the start and end of this time point;

[0010] Step S3: Large Model Retrieval and Judgment: Build a traditional media advertising knowledge base using past artificial experience data, and use large language models and enhanced retrieval methods to determine whether the segment is an advertisement and extract the text boundaries of the advertisement;

[0011] Step S4: Boundary Optimization: Revise the start and end boundaries of the advertisement using VAD silent detection and speaker synchronization.

[0012] For further optimization, step S1 specifically includes:

[0013] Step S1.1: Use the FFMPEG tool to extract the advertising audio-visual signals in TV and radio;

[0014] Step S1.2: Use the Wiener filtering algorithm to denoise the audio signal:

[0015] Step S1.3: Use the VAD algorithm to detect the speech activity and silent periods in the audio; use the ASR model to transcribe the speech part into text, record the start and end times of the text, identify the speech segments of the same speaker, and slice the audio-visual file at equal-length intervals to generate text segments with timestamps.

[0016] For further optimization, step S2 specifically includes:

[0017] Step S2.1: Perform a short-time Fourier transform on the preprocessed audio signal to obtain the time-frequency representation X(t,f), and convert the time-frequency signal into the Mel spectrum M(t,f);

[0018] Step S2.2: Convert the Mel spectrum signal into the input format of a convolutional neural network (CNN), and use the CNN to process the audio features to obtain the feature matrix F cnn . Specifically: First, normalize the Mel spectrum so that its mean is 0 and its standard deviation is 1 to accelerate the training speed of the model. Second, build a CNN model and define the structures and parameters of the convolutional layer, pooling layer, and fully connected layer. Then, use the training data to train the CNN model, and then input the Mel spectrum to be processed into the trained model to obtain the feature matrix of the audio.

[0019] Step S2.3: Add the positional encoding PE of the Transformer to the feature matrix to obtain the audio features with positional encoding, and then perform multi-head self-attention calculation. After calculation by the feed-forward neural network, residual, and layer normalization, obtain the audio features F with temporal dependence trans 。

[0020] The multi-head attention calculation is specifically as follows: First, calculate the attention scores: For the input audio features with positional encoding, then multiply them by three different weight matrices respectively to obtain three matrices of query, key, and value. Then, calculate the dot product of the query matrix and the key matrix to obtain the attention score matrix. Second, apply the Softmax function: Apply the Softmax function to the attention score matrix to convert the scores into a probability distribution to obtain the attention weight matrix. Third, calculate the weighted sum: Multiply the attention weight matrix by the value matrix to obtain the weighted sum matrix, which is the output of the multi-head attention. Fourth, multi-head attention merging: Concatenate the outputs of multiple attention heads, and then perform a transformation through a linear layer to obtain the final multi-head attention output

[0021] The feed-forward neural network calculation is specifically as follows: Input the output of the multi-head attention into the feed-forward neural network, and after processing by two fully connected layers and a non-linear activation function (such as ReLU), obtain the output of the feed-forward neural network

[0022] Then add the output of the multi-head attention and the output of the feed-forward neural network to obtain the output of the residual connection. Then perform layer normalization on the output of the residual connection to obtain the final temporal feature matrix

[0023] Step S2.4: Perform a nearest neighbor search on the audio features in the vector retrieval engine, calculate the similarity using the Euclidean distance, obtain continuous similar continuous subsets through dynamic programming, and determine the time range [t n ,t n+m 。

[0024] Performing a nearest neighbor search on the audio features in the vector retrieval engine specifically includes:

[0025] 1). Vector storage: Store the audio feature vectors in Milvus, and each vector corresponds to an audio segment

[0026] 2). Query vector: Input the audio feature vector to be queried into Milvus

[0027] 3). Nearest neighbor search: Milvus uses the approximate nearest neighbor search algorithm (ANN) to quickly find the vector most similar to the query vector in the high-dimensional vector space, that is, the similar segment

[0028] 4) Similarity calculation: Calculate the Euclidean distance between the query vector and the found similar segment vector as a measure of similarity.

[0029] Obtain continuous similar continuous subsets through dynamic programming and determine the time start of the similar segments, specifically including:

[0030] 1) Define states: Define a state array to record the maximum length of the continuous similar subset ending with each similar segment.

[0031] 2) State transition equation: For each similar segment, traverse all the previous similar segments. If the similarity between the current segment and a previous segment meets the threshold requirement, update the maximum length of the current segment in the state array.

[0032] 3) Backtracking: According to the state array, find the continuous similar subset with the maximum length and determine its time start point.

[0033] For further optimization, in step S2.2, Resnet–50 is selected in the improved CNN to process the audio features of television broadcasts. The input of ResNet-50 is a single-channel Mel spectrogram.

[0034] For further optimization, step S3 large model retrieval and judgment specifically includes:

[0035] Step S3.1: Intercept the transcribed text in the preprocessing process according to the time range [t n , t n+m , and splice the feature T n , t n+m in the time range; i ;

[0036] Step S3.2: Create an advertising feature sample library A, calculate the similarity S(T i , A j ) between the feature T i and the advertising sample A j in the advertising database. Use cosine similarity for nearest neighbor search. If the similarity exceeds the set threshold, preliminarily determine that this segment is an advertisement;

[0037] Step S3.3: Use the large language model to design a Prompt, extract the advertising content in the transcribed text, and merge adjacent advertising segments if they have continuous high confidence.

[0038] For further optimization, the specific steps for creating the advertising feature sample library in step S3.2 include:

[0039] Step S3.2.1: Data collection: Collect historical monitoring data of traditional media advertisements in the early stage. These data include manually annotated advertisement audio, video clips, as well as corresponding text scripts, advertisement types, and placement time information. For example, advertisement materials from television and radio in the past few years can be collected, including advertisement videos, subtitle texts, records of advertisement placement time periods, etc. In addition to its own historical monitoring data, supplementary data can also be obtained from channels such as industry reports, advertisement databases, and relevant research institutions to enrich the content of the knowledge base.

[0040] Step S3.2.2: Data preprocessing:

[0041] Clean the collected data to remove noise, duplicate data, and misannotations. For example, check whether there are problems such as garbled characters and typos in the text script and correct them.

[0042] Perform detailed annotation on the advertisement clips. The annotation content includes the start and end times of the advertisement, the product or service type of the advertisement, and the theme of the advertisement, providing accurate labels for subsequent model training and retrieval.

[0043] Structurally process the cleaned and annotated data and store it in a database or file system. A relational database (such as MySQL) or a non-relational database (such as MongoDB) can be used to store the data for easy management and query.

[0044] Step S3.2.3: Knowledge base construction:

[0045] Extract features from the advertisement data, including audio features (such as MFCC, spectral centroid, etc.), video features (such as color histogram, motion vector, etc.), and text features (such as keywords, word frequency, etc.). These features will be used for subsequent retrieval and judgment.

[0046] Build an index for the data in the knowledge base for fast retrieval. Techniques such as inverted index and hash index can be used to improve the retrieval efficiency.

[0047] For further optimization, in step S3.3, adjacent advertisement clips with consecutive high confidence are merged, specifically including:

[0048] Step S3.3.1: Set the confidence definition and threshold:

[0049] When the large language model determines that a clip is an advertisement, output the corresponding confidence score (such as in the range of 0 - 1, the higher the value, the more likely it is an advertisement). For example, set the threshold to 0.8, and only retain the clips with a confidence level ≥ 0.8.

[0050] Check the chronological order and confidence of adjacent segments. If segment A (time t1 - t2, confidence 0.85) and segment B (time t2 - t3, confidence 0.88) are temporally continuous and both exceed the threshold, they are determined to be consecutive high-confidence advertisement segments.

[0051] Step S3.3.2: Merge the transcribed texts of consecutive segments. The new boundary is the start time of the first segment and the end time of the last segment, and the text content is spliced into the complete advertisement text.

[0052] If segment features (such as semantic vectors, keyword features) have been extracted previously, weighted average (such as by confidence weight) or splicing is used to fuse the features during merging to form the comprehensive features of the merged advertisement.

[0053] Result output: Output the merged advertisement text, time boundary, and high-confidence flag.

[0054] Merging improves the integrity of advertisement recognition, avoids fragmentation of advertisement segments due to audio segmentation and recognition errors, and restores the true boundaries of advertisements. At the same time, it reduces the number of fragmented advertisement segments, provides a more regular input for downstream tasks such as advertisement statistics and compliance review, and optimizes the efficiency of subsequent processing.

[0055] For further optimization, extract the text boundaries of the advertisement, specifically including:

[0056] 1). Boundary localization based on timestamps:

[0057] Time alignment: If the audio segment has been speech-to-text transcribed and each word or sentence has a corresponding timestamp, locate the start and end positions of the advertisement text based on the start and end times of the audio segment determined to be an advertisement.

[0058] Boundary adjustment: When locating the text boundary, some adjustments are needed, such as removing some irrelevant words or noises before and after the advertisement to ensure the accuracy of the boundary.

[0059] 2). Boundary optimization based on semantics:

[0060] Semantic analysis: Perform semantic analysis on the located advertisement text to determine whether the beginning and end of the text conform to the semantic characteristics of the advertisement. For example, advertisements usually contain promotional information about products or services, promotional activities, etc. If the beginning or end of the text does not contain such information, the boundary can be further adjusted.

[0061] Context understanding: Combine the context information of the advertisement text, such as the previous and subsequent sentences, paragraphs, etc., to optimize the text boundary. For example, if the end of the advertisement text is an incomplete sentence, the next sentence can be referred to determine the correct boundary.

[0062] Through the above steps, the large language model and enhanced retrieval can be used to determine whether an audio segment is an advertisement and extract the text boundaries of the advertisement.

[0063] For further optimization, in step S4 boundary optimization, it specifically includes: during the revision process, by fine-tuning the start and end times, the nearest silent time point is selected to obtain the final advertisement time.

[0064] Step S4.1: Use the VAD algorithm to detect the audio frame by frame, mark the silent intervals, and use the front and rear endpoints of the silent intervals as candidate positions for the advertisement boundaries. For example, if the detected silent interval is [t1, t2], the advertisement may start at t1 or end at t2.

[0065] Step S4.2: Use the speaker segmentation algorithm (such as pyannote.audio) to divide the audio into segments of different speakers and generate speaker change points. For example, the advertisement may be broadcast by a specific announcer, which is different from the voices of the front and back program hosts.

[0066] If the endpoint of the silent interval coincides with or is close to (such as the time difference < 0.5 seconds) the speaker change point, then this endpoint is more likely to be the advertisement boundary;

[0067] If the speaker changes before and after the silence (such as changing from the program host to the advertisement announcer), then the start point of the silence is the start of the advertisement; conversely, the end point of the silence is the end of the advertisement;

[0068] Step S4.3: Integrate the VAD silent intervals and speaker change points to determine the final boundary. For example: if the start point t1 of the silent interval [t1, t2] is accompanied by a speaker change, then the start boundary of the advertisement is revised to t1. If the speaker changes back to the program content after the end point t2 of the silent interval, then the end boundary of the advertisement is revised to t2.

[0069] Compared with the prior art, the beneficial effects of the present invention are:

[0070] 1. In the process of extracting features of television and radio, the method of the present invention uses a neural network and an attention mechanism, effectively solving the problem of the long time span of traditional media audio, being able to better extract the temporal features of the audio, and improving the accuracy of advertisement matching.

[0071] 2. The method of the present invention fully considers the characteristics of traditional media advertisement broadcasting, comprehensively uses techniques such as nearest neighbor search, dynamic programming, knowledge base retrieval, large language model judgment, and silence detection, realizes a higher-precision identification of the advertisement extraction time boundary, and greatly reduces manual intervention and errors.

[0072] 3. The method of the present invention realizes the automation of advertisement separation and recognition, significantly improves the efficiency of traditional media advertisement monitoring, can quickly process a large amount of television and radio program content, timely discovers and monitors advertisements, and provides strong support for advertisement supervision. Description of the Drawings

[0073] Figure 1 It is a schematic diagram of a method for separating and recognizing traditional media advertisements based on a multimodal large model. Detailed Embodiments

[0074] The present invention will be further described below in conjunction with embodiments, but it should not be understood that the above-mentioned subject matter scope of the present invention is limited to the following embodiments.

[0075] Embodiment 1:

[0076] As Figure 1 shown, a method for separating and recognizing traditional media advertisements based on a multimodal large model includes the following steps:

[0077] Step S1: Multimodal preprocessing process: Extract advertisement audio-visual signals in television and radio and perform preprocessing on them. Specifically, it includes:

[0078] Step S1.1: Use the FFMPEG tool to extract advertisement audio-visual signals in television and radio;

[0079] Step S1.2: Use the Wiener filtering algorithm to perform noise reduction processing on the audio signal:

[0080]

[0081] Among them, represents the estimated value of the audio signal at the r-th frequency component in the frequency domain, |H(r)| represents the amplitude response of the signal at the r-th frequency component, Y(r) represents the noisy signal received at the r-th frequency component, and λ represents the power spectrum of the noise.

[0082] Step S1.3: Use the VAD algorithm to detect speech activities and silent periods in the audio; use the ASR model to transcribe the speech part into text, record the start and end times of the text, identify speech segments of the same speaker, and slice the audio-visual file at equal-length intervals to generate text segments with timestamps.

[0083] Step S2: Temporal feature enhancement: Extract audio features of television and radio from the preprocessed audio-visual signals, find the subset of the longest identical segments from continuous videos through feature extraction and feature retrieval, and record the start and end of this time point.

[0084] Specifically, it includes:

[0085] Step S2.1: Perform a short-time Fourier transform on the preprocessed audio signal to obtain the time-frequency representation X(t, f),

[0086]

[0087] and convert the time-frequency signal to the Mel spectrogram M(t, f).

[0088]

[0089] Where x(n) represents the value of the preprocessed audio signal at the nth sampling point, representing the discrete sampling values of the audio signal. w(n - t) represents the window function (such as the Hann window, Hamming window), which is used to frame the audio signal. t represents the analysis time position, and different time segments of the audio are intercepted by sliding the window function. H g (f) represents the frequency response function of the gth filter in the Mel filter bank, which is used to filter the time-frequency signal X(t, f) and extract the frequency components that conform to the Mel scale. G represents the total number of filters in the Mel filter bank, which determines the dimension of the final Mel spectrogram (i.e., the number of features).

[0090] Step S2.2: Convert the Mel spectrogram signal into the input format of the convolutional neural network CNN, and use CNN for audio feature processing to obtain the feature matrix F cnn . Specifically: First, normalize the Mel spectrogram so that its mean is 0 and the standard deviation is 1 to accelerate the training speed of the model. Secondly, construct a CNN model and define the structures and parameters of the convolutional layer, pooling layer, and fully connected layer. Then, use the training data to train the CNN model, and then input the Mel spectrogram to be processed into the trained model to obtain the feature matrix F of the audio. cnn .

[0091] In this embodiment, CNN is used for audio feature processing of television broadcasts, and Resnet-50 is selected here. Since the ResNet pre-trained model expects RGB images (224×224×3), but the Mel spectrogram is single-channel (T×F×1). Duplicate the Mel spectrogram 3 times to become (T×F×3) to adapt to ResNet.

[0092] Input the Mel spectrogram M(t, f), and after passing through ResNet, output 512-dimensional high-dimensional features, including spectral patterns and background sound features.

[0093] F cnn = ResNet(M′);

[0094] C l+1 (t, f) = ReLU(W * C l (t, f)+b);

[0095] Among them, M' represents the Mel spectrogram data input to the ResNet network; C l (t, f) represents the feature map of the l-th layer in the convolutional neural network, with dimensions of time t and frequency f; W represents the convolutional kernel weight matrix; B represents the bias term, adding an offset to the result of the convolutional operation to increase the expressive power of the model; ReLU represents the rectified linear activation function (Rectified Linear Unit), performing a non-linear transformation on the convolutional result W * C l (t, f) + b, and the formula is ReLU(x) = max(0, x), which is used to introduce non-linearity and improve the network's ability to learn complex patterns. C l+1 (t, f) represents the feature map output to the (l + 1)-th layer after the l-th layer of convolution and activation operations, which integrates new spectral features for further processing by subsequent network layers.

[0096] Step S2.3: Add the position encoding PE of the Transformer to the feature matrix to obtain the audio feature with position encoding, then perform multi-head self-attention calculation, and through feed-forward neural network calculation, residual, and layer normalization calculation, obtain the audio feature F with temporal dependence trans .

[0097] Among them, the addition of the position encoding of the Transformer is as follows:

[0098] PE(t, 2i) = sin(t / 10000 2i / D )

[0099] PE(t, 2i + 1) = cos(t / 10000 2i / D )

[0100] Among them, T is the time step, D is the feature dimension, and i is the dimension index value.

[0101] Finally, add the position encoding to the audio feature matrix.

[0102] F input = F cnn + PE.

[0103] The specific multi-head attention calculation is as follows: First, calculate the attention scores: For the input audio feature with position encoding, first multiply it by three different weight matrices respectively to obtain three matrices of query (Query), key (Key), and value (Value).

[0104] Q = W Q F input , K = W K F input , V = W V Finput

[0105] Among them, Q, K, and V respectively represent the query matrix, the key matrix, and the value matrix; W Q 、W K 、W V respectively represent trainable parameter matrices;

[0106] Then, calculate the dot product of the query matrix and the key matrix to obtain the attention score matrix.

[0107]

[0108] Among them, d k is the scaling factor, and B represents the attention score matrix.

[0109] Calculate the attention weights, use Softmax normalization so that the sum of the weights is 1. Multiply V by the attention scores to obtain the aggregated information.

[0110] Attention(Q, K, V) = softmax(A)V

[0111] Finally, perform the multi-head attention MHSA(F input ) calculation.

[0112] MHSA(F input ) = Concat(head1,..., head h )W O ;

[0113] There is a feed-forward neural network after each Transformer layer to improve the model's expressive ability.

[0114] FFN(x) = max(0, xW1 + b1)W2 + b2;

[0115] Among them, head h represents the output result of the h-th attention head. W ODenote the projection matrix, which performs a linear transformation on the concatenated multi-head attention output (Concat(head1, …, headh)), integrates multi-head information, and outputs the final feature representation. x represents the input to the feed-forward neural network; W1 and b1 represent the weight matrix and bias term of the first-layer linear transformation, which perform a linear transformation on the input x (xW1 + b1), expanding the feature dimension or extracting preliminary abstract features. max(0, ·) represents the ReLU activation function, which introduces non-linearity, filters out negative values in the result of the linear transformation, and enhances the model's ability to express complex patterns. W2 and b2 represent the weight matrix and bias term of the second-layer linear transformation, which further transform the activated features (max(0, xW1 + b1)W2 + b2), compressing the feature dimension or generating the final feature representation, and enhancing the model's expressive ability.

[0116] Use residual connections and layer normalization. Use LayerNorm to normalize features to make training more stable. And use residual connections to prevent the vanishing gradient during the solution process.

[0117] F trans = LayerNorm(MHSA(F input ) + F input )

[0118] F trans = LayerNorm(FFN(F input ) + F input )。

[0119] Step S2.4: Perform a nearest neighbor search ANN on the audio features in the vector retrieval engine, calculate the similarity using the Euclidean distance, and obtain the continuous similar continuous subset through dynamic programming to determine the time range [t n , t n+m . Specifically:

[0120]

[0121] The similarity calculation uses the Euclidean distance d(F i , F j ):

[0122] d(F i , F j ) = ||F i - F j ||

[0123] F i , F j represent the feature vectors of audio segments, where i and j are segment indices, and NN(F i ) represents the nearest neighbor function, which is used to find the feature F i that is most similar to the feature Fj , output the F that maximizes the similarity sim(F i , F j ) j 's index j.

[0124] Set the similarity threshold θ. If sim(F i , F j ) > θ, then it is considered that a i and a j may belong to the same advertisement. Let S be the longest advertisement subset S = {a k , a k+1 ,..., a k+m};

[0125] where a i and a j represent audio segments, where i, j are segment indices; k is the index of the starting segment in the advertisement subset S; m represents the number offset of segments in the subset.

[0126] Obtain the target where M i,i+1 = 1, and calculate using dynamic programming. M i,i+1 = 1 represents the element of the indicator matrix.

[0127]

[0128] After calculation, the longest advertisement segment is max(dp[i]), and the advertisement time is [t k , t k+m .

[0129] Step S3: Retrieval and judgment by large model: Build a traditional media advertisement knowledge base using past artificial experience data, and use a large language model and enhanced retrieval method to judge whether the segment is an advertisement and extract the text boundary of the advertisement. Specifically, it includes:

[0130] Step S3.1: Intercept the transcribed text in the preprocessing process according to the time range [t n , t n+m , and splice the feature T n , t n+m in the range; i ;

[0131] Step S3.2: Make an advertisement feature sample library A, and calculate the similarity S(T i , A j ) between the audio feature T i and the advertisement sample A j in the advertisement database.

[0132]

[0133] Among them, F Ti , F Aj are the feature vectors of the feature T to be detected i and the advertisement sample A j respectively.

[0134] Perform a nearest neighbor search using cosine similarity,

[0135]

[0136] If S(T i , A * ) > θ (threshold), then T i may be an advertisement.

[0137] Step S3.3: Use the large language model to design a Prompt to extract the advertisement content in the transcribed text. If adjacent advertisement segments have continuous high confidence, merge them.

[0138] If adjacent advertisement segments (T i , T i+1 ,... T i+k ) have continuous high confidence, merge them.

[0139]

[0140] T final represents the final advertisement segment after merging. P ad (T i ) represents the confidence that the advertisement segment T i belongs to an advertisement. τ is the confidence threshold.

[0141] Step S4: Boundary optimization: Revise the beginning and ending boundaries of the advertisement using VAD silence detection and speaker synchronization.

[0142] Boundary revision process. Further fine-tune the start and end times through VAD silence detection:

[0143] (t start , t end ) = VAD(T final );

[0144] Finally, within the time range of the fine-tuned (t start , t end ), it is the final advertisement content obtained.

[0145] Application example:

[0146] Separate a TV advertisement of a certain TV drama.

[0147] Step S1, Input Data: A TV drama is 60 minutes long and contains 3 inserted advertisements; the resolution is 1920×1080, and the audio sampling rate is 44.1kHz.

[0148] Step S2, Preprocessing Results: An audio file is separated, with a size of 210MB; 360 10-second slices are generated, and 42 silent intervals are detected by VAD.

[0149] Step S3, Feature Enhancement Output: Mel spectrum matrix: 360×40×1; ResNet output features: 360×512; Transformer temporal features: 360×512.

[0150] Step S4, Large Model Retrieval: The longest matching segment is found, with a similarity of 0.92; LLM verification is performed, with a confidence of 0.96, and it is determined to be an advertisement.

[0151] Step S5, Boundary Optimization: VAD fine-tuning, the start time changes from 00:15:30→00:15:30.5, and the final output advertisement timestamp is: [00:15:30.5 - 00:16:15.5].

[0152] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. A traditional media advertisement separation and recognition method based on a multimodal large model, characterized in that It includes the following steps: Step S1: Multimodal preprocessing process: Extract the advertising audio-visual signals from TV and radio, and perform preprocessing on them; Step S2: Temporal feature enhancement: Extract the audio features of TV and radio from the preprocessed audio-visual signals, find the subset of the longest identical segments from continuous videos through feature extraction and feature retrieval methods, and record the start and end of this time point; Step S3: Large model retrieval and judgment: Build a traditional media advertising knowledge base using past artificial experience data, and use large language models and enhanced retrieval methods to judge whether the segment is an advertisement and extract the text boundaries of the advertisement; Step S4: Boundary optimization: Use VAD silence detection and speaker synchronization methods to revise the start and end boundaries of the advertisement.

2. The method for separating and identifying traditional media advertisements based on a multimodal large model according to claim 1, wherein The specific steps of step S1 include: Step S1.1: Use the FFMPEG tool to extract the advertising audio-visual signals from TV and radio; Step S1.2: Use the Wiener filtering algorithm to denoise the audio signal: Step S1.3: Use the VAD algorithm to detect the speech activity and silent periods in the audio; Use the ASR model to transcribe the speech part into text, record the start and end times of the text, and identify the speech segments of the same speaker. Slice the audio-visual file at equal-length intervals to generate text segments with time stamps.

3. The method for separating and identifying traditional media advertisements based on a multi-modal large model according to claim 2, wherein The specific steps of step S2 include: Step S2.1: Perform short-time Fourier transform on the preprocessed audio signal to obtain the time-frequency representation X(t,f), and convert the time-frequency signal to the Mel spectrum M(t,f); Step S2.2: Convert the Mel spectrum signal into the input format of a convolutional neural network (CNN), and use the CNN to process audio features to obtain a feature matrix F cnn ; Step S2.3: Add the positional encoding PE of the Transformer to the feature matrix to obtain the audio features with positional encoding, and then perform multi-head self-attention calculation. After calculation by the feed-forward neural network, residual, and layer normalization calculation, obtain the audio features F with temporal dependence trans ; Step S2.4: Perform a nearest neighbor search on the audio features in the vector retrieval engine, calculate the similarity using the Euclidean distance, obtain consecutive similar consecutive subsets through dynamic programming, and determine the starting time [t n , t n+m .

4. The method for separating and identifying traditional media advertisements based on a multi-modal large model according to claim 3, wherein, In step S2.2, Resnet–50 is selected in the improved CNN to process the audio features of TV and radio. The input of ResNet-50 is a single-channel Mel spectrum.

5. The method for separating and identifying traditional media advertisements based on a multimodal large model according to claim 4, wherein The large model retrieval and judgment in step S3 specifically includes: Step S3.1: According to the time range [t n ,t n+m ] intercept the transcribed text in the preprocessing process and concatenate the time [t n ,t n+m ] Features T in the range i ; Step S3.2: Create an advertising feature sample library A and calculate feature T i and the advertising sample A in the advertising database j to calculate the similarity S(T i , A j ). Use cosine similarity for nearest neighbor search. If the similarity exceeds the set threshold, preliminarily determine that this segment is an advertisement; Step S3.3: Use the large language model to design a Prompt to extract the advertising content in the transcribed text. If adjacent advertising segments have continuous high confidence, they are merged.

6. The method for separating and identifying traditional media advertisements based on a multi-modal large model according to claim 5, wherein The specific steps of making the advertising feature sample library in step S3.2 include: Step S3.2.1: Data collection: Collect historical data of traditional media advertisements in the early stage. These data include artificially labeled advertising audio, video segments, and corresponding text scripts, advertising types, and placement time information; Step S3.2.2: Data preprocessing: Clean the collected data to remove noise, duplicate data, and incorrect labels; Perform detailed annotation on the advertising segments. The annotation content includes the start and end times of the advertisement, the product or service type of the advertisement, and the theme of the advertisement, providing accurate labels for subsequent model training and retrieval; Structurally process the cleaned and annotated data and store it in a database or file system. Step S3.2.3: Knowledge base construction: Extract features from the advertising data, including audio features, video features, and text features; Build an index for the data in the knowledge base for quick retrieval.

7. The method for separating and identifying traditional media advertisements based on a multimodal large model according to claim 6, wherein In step S3.3, the merging of adjacent advertising segments with continuous high confidence specifically includes: Step S3.3.1: Set the confidence definition and threshold: When the large language model determines that a segment is an advertisement, output the corresponding confidence score; Check the chronological order and confidence of adjacent segments. If different segments are consecutive in time and their confidence levels both exceed the threshold, they are determined to be consecutive high-confidence advertisement segments; Step S3.3.2: Merge the transcribed texts of the consecutive segments. The new boundary is the start time of the first segment and the end time of the last segment, and the text content is spliced into the complete advertisement text; If segment features have been extracted previously, use weighted average or splicing methods to fuse the features during merging to form the comprehensive features of the merged advertisement; Output the merged advertisement text, time boundary, and high-confidence identifier.

8. The method for separating and identifying traditional media advertisements based on a multimodal large model according to claim 7, wherein The boundary optimization in step S4 specifically includes: during the revision process, by fine-tuning the start and end times, select the nearest silent time point to obtain the final advertisement time. Step S4.1: Use the VAD algorithm to detect the audio frame by frame, mark the silent intervals, and use the front and end points of the silent intervals as candidate positions for the advertisement boundary; Step S4.2: Use the speaker segmentation algorithm to divide the audio into segments of different speakers and generate speaker change points; If the end point of the silent interval coincides with or is adjacent to the speaker change point, then this end point is more likely to be the advertisement boundary; If the speaker changes before and after the silence, then the start point of the silence is the start of the advertisement; otherwise, the end point of the silence is the end of the advertisement; Step S4.3: Integrate the VAD silent intervals and speaker change points to determine the final boundary.

Citation Information

Cited By

  • Advertisement content automatic monitoring method and system

    CN121660749A

  • An automatic advertisement content monitoring method and system

    CN121660749B