Advertisement content automatic monitoring method and system

By using multimodal feature matrix fusion and dynamic template library similarity matching, the accuracy problem of advertising identification in streaming media data is solved, realizing an end-to-end advertising monitoring system with high accuracy and adaptability, adapting to new advertising variants and policy differences.

CN121660749AActive Publication Date: 2026-03-13BEIJING HIZHI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-03
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify advertising content within massive amounts of streaming media data, especially advertising segments with multimodal heterogeneous information, resulting in insufficient accuracy and adaptability of advertising monitoring systems.

Method used

Employing multimodal feature matrix fusion technology, this method extracts visual feature vectors from video streams, audio feature vectors from audio streams, and text feature vectors from auxiliary text. It then combines these with a dynamic advertising template library for multimodal similarity matching to generate a set of spatiotemporal coordinates for advertising segments. Based on audio and visual signals, it performs coarse-grained segmentation and optimizes boundaries using text clustering to generate a list of independent advertising items. Finally, it compares the list with a prohibited sample library for multimodal similarity and outputs a report of violations.

Benefits of technology

It achieves end-to-end automated supervision of streaming media advertising content, improves the accuracy of ad identification and the rationality of violation determination, can adapt to new variations and cross-regional policy differences, and has strong robustness and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121660749A_ABST
    Figure CN121660749A_ABST
Patent Text Reader

Abstract

The invention relates to an advertisement content automatic monitoring method and system, and belongs to the technical field of digital media content analysis, and the monitoring method comprises the steps: obtaining streaming media data; extracting a visual feature vector from the video stream, extracting an audio feature vector from the audio stream, extracting a text feature vector from the auxiliary text, and fusing to obtain a multi-modal feature matrix; performing multi-modal similarity matching on the multi-modal feature matrix and an advertisement template library to generate an advertisement paragraph space-time coordinate set; positioning a corresponding subset of the multi-modal feature matrix, performing coarse-grained segmentation based on an energy abrupt change point of an audio feature vector and a lens switching point of a visual feature vector, performing boundary optimization through a clustering result of a text feature vector, generating an independent advertisement item object list, performing multi-modal similarity comparison with a preset broadcast forbidding sample library, and performing broadcast forbidding on the independent advertisement item object list. And analyzing the semantic relevance of the context, outputting a violation item report and confidence, and triggering an alarm instruction in a grading manner. The accuracy of the advertisement monitoring system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital media content analysis technology, and in particular to a method and system for automatically monitoring advertising content. Background Technology

[0002] With the explosive growth of internet video traffic, traditional advertising supervision methods relying on manual inspection or rule-based matching are no longer sufficient to cope with the current complex and ever-changing content ecosystem. There is an urgent need for an intelligent monitoring system that can achieve end-to-end automation, high precision, and strong adaptability. Against this backdrop, accurately identifying advertising content from massive amounts of streaming media data and further determining its compliance has become one of the core challenges facing the industry.

[0003] Currently, advertising formats are becoming increasingly diversified, with significantly enhanced disguise and concealment. Many ads appear as soft placements, news broadcasts, or product reviews, deliberately avoiding the "advertisement" label, making it difficult to effectively identify them using only metadata or keyword matching methods. Furthermore, due to the highly multimodal nature of media content, visual, audio, and textual information often presents heterogeneous, asynchronous, and even contradictory information, making single-modal analysis prone to misjudgment. In particular, ad segments containing multiple independent brand carousels or product promotion units cannot be adequately labeled with only the overall start and end times, failing to support refined content management and accountability. This results in a lack of granular support for subsequent review and enforcement, reducing the accuracy of advertising monitoring systems. Summary of the Invention

[0004] To improve the accuracy of advertising monitoring systems, this application provides a method and system for automatic monitoring of advertising content.

[0005] Firstly, this application provides a method for automatically monitoring advertising content, employing the following technical solution: An automatic monitoring method for advertising content, the monitoring method comprising: Acquire streaming media data, including video streams, audio streams, and supplementary text; Visual feature vectors are extracted from the video stream, audio feature vectors are extracted from the audio stream, and text feature vectors are extracted from the auxiliary text. These are then fused to obtain a multimodal feature matrix. The multimodal feature matrix is ​​matched with a dynamically updated ad template library using multimodal similarity to generate a spatiotemporal coordinate set of ad segments; Based on the spatiotemporal coordinate set of the advertisement segment, locate the corresponding subset of the multimodal feature matrix, perform coarse-grained segmentation based on the energy mutation point of the audio feature vector and the shot switching point of the visual feature vector, and then perform boundary optimization through the clustering result of the text feature vector to generate a list of independent advertisement item objects. The list of independent advertising items is compared with a pre-set prohibited sample library using multimodal similarity analysis. The contextual semantic relevance is analyzed in conjunction with a regional pre-set policy database, and a report of violations and confidence level are output. An alarm command is triggered based on the confidence level of the reported violation item.

[0006] By adopting the above technical solution, end-to-end automated supervision of streaming media advertising content has been achieved. This method overcomes the limitations of traditional single-modal detection, utilizing a multimodal feature matrix to achieve collaborative analysis of audiovisual text. Combined with a dynamic template library and contextual semantic understanding mechanism, it significantly improves the accuracy of ad recognition and the rationality of violation determination. Its core innovation lies in advancing ad detection from "static comparison" to a new stage of "dynamic evolution + context awareness." It can not only efficiently identify known violation patterns but also adapt to new variations and cross-regional policy differences, possessing strong robustness and scalability.

[0007] Secondly, this application provides an automatic advertising content monitoring system, which adopts the following technical solution: An automatic advertising content monitoring system, the system comprising: The acquisition module is used to acquire streaming media data, including video streams, audio streams, and auxiliary text; The feature fusion module is used to extract visual feature vectors from the video stream, audio feature vectors from the audio stream, and text feature vectors from the auxiliary text, and fuse them to obtain a multimodal feature matrix; The multimodal similarity matching module is used to perform multimodal similarity matching between the multimodal feature matrix and the dynamically updated advertising template library to generate a set of spatiotemporal coordinates for advertising segments; The ad entry generation module is used to locate the corresponding subset of the multimodal feature matrix based on the spatiotemporal coordinate set of the ad segment, perform coarse-grained segmentation based on the energy mutation points of the audio feature vector and the shot switching points of the visual feature vector, and then perform boundary optimization through the clustering results of the text feature vector to generate a list of independent ad entry objects. The violation analysis module is used to perform multimodal similarity comparison between the list of independent advertising items and the pre-set prohibited sample library, analyze the contextual semantic relevance in conjunction with the regional preset policy database, and output a violation item report and confidence level. The graded alarm module is used to trigger alarm commands according to the confidence level of the reported violation items.

[0008] Thirdly, this application provides a computer device, which adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to perform the steps of the method as described in the first aspect.

[0009] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description

[0010] Figure 1 This is a first flowchart illustrating an automatic advertising content monitoring method according to one embodiment of this application.

[0011] Figure 2 This is a second flowchart illustrating an automatic advertising content monitoring method according to one embodiment of this application.

[0012] Figure 3 This is a schematic diagram of the third process of an automatic advertising content monitoring method according to one embodiment of this application.

[0013] Figure 4 This is a schematic diagram of the fourth process of an automatic advertising content monitoring method according to one embodiment of this application.

[0014] Figure 5 This is a schematic diagram of the fifth process of an automatic advertising content monitoring method according to one embodiment of this application.

[0015] Figure 6 This is a schematic diagram of the sixth process of an automatic advertising content monitoring method according to one embodiment of this application.

[0016] Figure 7 This is a schematic diagram of the seventh process of an automatic advertising content monitoring method according to one embodiment of this application. Detailed Implementation

[0017] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figure 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.

[0018] This application discloses a method for automatically monitoring advertising content.

[0019] Reference Figure 1 An automatic monitoring method for advertising content, the monitoring method includes: Step S101: Acquire streaming media data, including video stream, audio stream, and auxiliary text; Streaming media data typically originates from broadcast signal decoders, CDN distribution nodes, live streaming interfaces, or on-demand platform storage systems, forming a complete media content carrier. Video streams present dynamic visuals as frame sequences, audio streams record background music, narration, and sound effects, while supplementary text includes non-visual text information such as subtitles, closed-circuit captions (CC), metadata tags (e.g., program name, broadcast time), or real-time bullet comments. These three types of data are synchronized on the timeline, but their semantic expressions differ: video reflects visual presentation intentions, audio conveys auditory emotions and rhythm, and text provides explicit semantic cues.

[0020] By simultaneously collecting data from these three modalities, the system can build a comprehensive understanding of advertising content, avoiding information blind spots caused by the absence of a single modality (such as silent playback or obscured subtitles). For example, in a soft advertisement disguised as a news broadcast, the visuals may not explicitly label it as an "advertisement," but its audio tone has obvious sales-oriented characteristics, and brand keywords frequently appear in the subtitles. Multimodal collaborative analysis can effectively reveal its true nature.

[0021] Step S102: Extract visual feature vectors from the video stream, extract audio feature vectors from the audio stream, extract text feature vectors from the auxiliary text, and fuse them to obtain a multimodal feature matrix; For video streams, after frame sampling, the system uses pre-trained object detection models (such as YOLOv8 or Faster R-CNN) to identify key elements in the scene, such as facial expressions, product displays, logos, and price tags, and encodes these visual semantic information into high-dimensional visual feature vectors. At the same time, it calculates the HSV color histogram differences between adjacent frames for subsequent shot switching detection.

[0022] For audio streams, the system extracts Mel frequency cepstral coefficients (MFCCs), which can effectively simulate the human ear's perception of sound frequencies, capture the timbre and intonation changes of speech, and extract energy envelope features to reflect volume fluctuations and rhythmic intensity, and then combine them into an audio feature vector.

[0023] For auxiliary text, the system uses semantic embedding technology (such as BERT or Sentence-BERT) to transform it into a dense vector representation, and enhances its semantic sensitivity by matching it with a preset keyword library (such as sensitive words like "limited-time offer", "lowest price on the entire network", "special treatment" etc.) to generate text feature vectors.

[0024] Ultimately, these three types of features are aligned and concatenated along the time dimension to form a multimodal feature matrix that evolves over time. Each time segment corresponds to a set of joint feature vectors, fully characterizing the audiovisual semantic state of the media content at that moment. This structured representation not only preserves the independent characteristics of each modality but also provides a unified data foundation for subsequent cross-modal comparison and fusion analysis.

[0025] Step S103: Perform multimodal similarity matching between the multimodal feature matrix and the dynamically updated advertising template library to generate a spatiotemporal coordinate set of advertising segments; The spatiotemporal coordinate set of the advertisement segment includes the start time, end time, and confidence level; In this embodiment, the advertising template library is not a static resource pool, but a continuously evolving knowledge base that stores typical feature patterns of historically confirmed legal and illegal advertisements. The matching process employs a weighted cosine similarity or deep metric learning model to calculate the similarity scores between the current feature matrix and various advertising templates in the template library across visual, audio, and text dimensions, respectively. A fusion strategy is then used to derive a comprehensive matching confidence score. When the matching score within a certain time period exceeds a preset threshold, the system determines that time period as an advertising content interval and records its start time, end time, and matching confidence score, forming a spatiotemporal coordinate set of the advertising segment.

[0026] For example, on a certain television channel, if the system identifies a 30-second segment that is highly similar to a known health product advertisement in terms of screen layout, background music rhythm, and subtitle language, it can automatically label it as a "suspected advertisement" and output its time range and matching probability. This mechanism can not only identify standard advertisements but also detect slightly modified variant advertisements, significantly improving the robustness of detection.

[0027] Step S104: Locate the corresponding subset of the multimodal feature matrix based on the spatiotemporal coordinate set of the advertisement segment, perform coarse-grained segmentation based on the energy mutation point of the audio feature vector and the shot switching point of the visual feature vector, and then perform boundary optimization through the clustering result of the text feature vector to generate a list of independent advertisement item objects. Since an ad segment may contain multiple independent ads (such as multiple brand carousels), simple start and end time divisions are insufficient to support individualized management, so further segmentation is required.

[0028] First, the system performs coarse-grained segmentation based on energy abrupt changes in audio feature vectors and shot transition points in visual feature vectors. Energy abrupt changes are identified by calculating the second derivative of the MFCC sequence to detect abrupt changes in sound intensity (such as sudden music stops or vocal intrusions), often corresponding to advertising transitions. Shot transition points are determined by significant changes in the HSV histograms of adjacent frames, suitable for identifying transitions such as hard cuts and fade-ins. These physical signal changes constitute the initial segmentation boundaries. However, relying solely on signal abrupt changes is susceptible to noise interference (such as fluctuations in sound effects within the program or rapid editing). Therefore, the system further incorporates the clustering results of text feature vectors for boundary optimization.

[0029] Specifically, the system clusters the text embedding vectors within ad segments according to time series, identifies semantic topic transition nodes (such as switching from "skincare products" to "home appliance promotions"), and uses a Conditional Random Field (CRF) model to integrate visual switching frequency and textual semantic coherence, minimizing the variance of features within each entry to determine the optimal segmentation point. The final output list of independent ad entry objects contains a structured data unit for each entry, including an independent time interval, a subset of multimodal features, and semantic topic tags, laying the foundation for subsequent refined review.

[0030] Step S105: Compare the list of independent advertising items with the pre-set prohibited sample library for multimodal similarity, analyze the contextual semantic relevance in conjunction with the regional preset policy database, and output a report of violations and confidence level; The banned sample database stores the characteristics of advertising content that has been identified as illegal, non-compliant, or sensitive, covering categories such as false advertising, medical misrepresentation, politically sensitive content, and vulgar content. The system calculates the cross-modal similarity between the current advertising item and the banned samples using deep hashing or twin network structures to assess its likelihood of violation.

[0031] More importantly, the system does not judge individual advertisement content in isolation, but introduces a contextual semantic association analysis mechanism: it loads regional policy rules that match the time and place of the advertisement broadcast (such as prohibiting tobacco advertisements from being broadcast before or after youth programs in a certain area), and models the semantic relationship between the advertisement item and the content of the programs before and after it through a long short-term memory network (LSTM) to identify whether there is any inducement link (such as playing a drug advertisement immediately after a health lecture).

[0032] Understandably, this context-aware capability enables the system to not only identify explicit violations but also uncover implicit violation logic, significantly improving the depth of supervision. The final violation report not only lists suspected violations but also includes detailed information such as confidence scores, matching sample numbers, and violated policy clauses, providing sufficient evidence for manual review.

[0033] Step S106: Trigger an alarm command based on the confidence level of the violation report.

[0034] The alert strategy is not a "one-size-fits-all" approach, but rather a dynamic response based on risk level: high-confidence violations (such as the explicit appearance of prohibited drug names and efficacy promises) trigger a Level 1 alert, immediately notifying the regulatory platform and automatically cutting off the signal source; medium-confidence entries trigger a Level 2 alert, which is pushed to the reviewers for confirmation; and low-confidence entries are only logged for subsequent traceability.

[0035] Meanwhile, the system monitors its own resource usage status (such as GPU memory and CPU load) and dynamically adjusts feature processing strategies in high-concurrency scenarios. For example, it performs principal component analysis (PCA) to reduce the dimensionality of the multimodal feature matrix to alleviate computational pressure, or dynamically scales up parallel processing nodes through a container orchestration platform (such as Kubernetes) to ensure that the system can still run stably under high traffic.

[0036] The above implementation achieves end-to-end automated supervision of streaming media advertising content. This method overcomes the limitations of traditional single-modal detection, utilizing a multimodal feature matrix to achieve audiovisual and text collaborative analysis. Combined with a dynamic template library and contextual semantic understanding mechanism, it significantly improves the accuracy of ad recognition and the rationality of violation determination. Its core innovation lies in advancing ad detection from "static comparison" to a new stage of "dynamic evolution + context awareness." This not only efficiently identifies known violation patterns but also adapts to new variations and cross-regional policy differences, exhibiting strong robustness and scalability.

[0037] Reference Figure 2 As one implementation of step S102, the steps of extracting visual feature vectors from the video stream, extracting audio feature vectors from the audio stream, extracting text feature vectors from the auxiliary text, and fusing them to obtain a multimodal feature matrix include: Step S201: Obtain the video stream, audio stream, and auxiliary text; The video stream typically originates from a camera or other image recording device and contains continuously changing visual images; the audio stream corresponds to sound signals recorded synchronously with the video, such as human voices, ambient sounds, or alarm sounds; and the auxiliary text may be unstructured or semi-structured language information such as subtitles, speech-to-text results, user input descriptions, or metadata tags.

[0038] Step S202: Perform frame sampling processing on the video stream and extract visual feature vectors through the object detection model; Since raw videos typically have a high frame rate (e.g., 30fps), there is significant temporal redundancy between consecutive frames. Performing in-depth processing on every frame would greatly increase the system load. Therefore, employing uniform sampling or motion-based adaptive sampling strategies, and selecting representative frames for analysis on the timeline, can both preserve key nodes in the motion evolution and improve overall processing efficiency.

[0039] Subsequently, a pre-trained object detection model is used to identify and locate objects in the sampled video frames. This object detection model (such as YOLO, Faster R-CNN, etc.) can automatically identify the category of interest in the image (such as people, vehicles, animals, etc.) and output the bounding box coordinates (x, y, w, h) and classification confidence score of each detected object. This location and confidence information reflects the spatial distribution and salience of the object in the image, constituting an important part of visual semantics.

[0040] Building upon this, we further extract the low-level visual features of the regions within the bounding box, including the HSV color histogram and the Histogram of Oriented Gradients (HOG). Compared to RGB, the HSV color space better reflects the color attributes perceived by human vision. Its histogram statistically analyzes the distribution of hue, saturation, and value, making it suitable for expressing color consistency under varying lighting conditions. HOG features, on the other hand, capture edge and texture structure information by calculating the directional distribution of gradients in local image regions, providing excellent descriptive capabilities for shape and contour.

[0041] Finally, the bounding box coordinates, detection confidence, HSV histogram, and HOG features are concatenated to form a fixed-dimensional visual feature vector. This multi-level feature combination includes both high-level semantics (detection results) and mid-to-low-level visual attributes, enhancing the expressive power and robustness of the features, making it particularly suitable for target recognition and behavior understanding tasks in complex scenarios.

[0042] Step S203: Calculate the Mel frequency cepstral coefficients and energy envelope parameters for the audio stream to generate an audio feature vector; First, a sliding time window is used to segment the continuous audio signal into short time segments (e.g., 25ms window length, 10ms step size). This is because speech and ambient sounds can be approximated as stationary signals within a short time frame, facilitating spectral analysis. For each time window, its Mel-frequency cepstral coefficients (MFCCs) are calculated. This is a classic acoustic feature widely used in speech recognition and audio classification. The generation process of MFCCs simulates the nonlinear frequency response characteristics of the human auditory system: after the original audio is Fourier transformed to obtain the spectrum, it is weighted by a set of triangular filters distributed according to the Mel scale, emphasizing low-frequency information and compressing high-frequency resolution. Then, the logarithmic energy of the filter output is taken, and finally, a discrete cosine transform is performed to obtain the cepstral coefficients. The first 12-13 orders of MFCCs can effectively characterize the spectral envelope of the sound, reflecting the shape and resonance characteristics of the vocal organs, and are highly sensitive to information such as speech content and speaker identity.

[0043] At the same time, the system also extracts the energy envelope parameter of the audio, that is, the short-time energy value of each frame of signal, to characterize the intensity trend of the sound. Based on this, it identifies abrupt changes on the energy curve that exceed a preset threshold (such as sudden rises or falls). These abrupt changes often correspond to the occurrence of event-related sounds such as the start / end of speech, knocking sounds, switching actions, or emotional fluctuations, and have important contextual clue significance.

[0044] Understandably, encoding the MFCC sequence together with the mutation point timestamp into an audio feature vector not only preserves the essential spectral features of the sound, but also introduces time-sensitive dynamic event information, enabling the audio modality to not only express "what was said", but also reflect "when a significant sound occurred", thereby enhancing its discriminative ability in multimodal collaborative analysis.

[0045] Step S204: Perform semantic embedding processing on the auxiliary text and match it with a pre-stored keyword library to generate a text feature vector; First, the original text is transformed into dense semantic vectors using pre-trained language models (such as BERT, RoBERTa, or Sentence-BERT). These models, trained on large-scale corpora, can capture contextual relationships between words, syntactic structures, and deep semantic meanings. The generated vectors exhibit good clustering and computability in the semantic space; for example, the vector distance between "doctor" and "hospital" is relatively close, while it is relatively far from "car." This semantic vector, as a representation of the overall meaning of the text, possesses strong generalization ability.

[0046] However, relying solely on semantic embedding may overlook the importance of keywords in certain domains, especially in applications such as surveillance, security, or industrial control, where some keywords (e.g., "fire," "intrusion," "stop") have clear triggering meanings. To address this, the system further matches the text content against a pre-stored keyword library, generating a Boolean feature vector where each bit corresponds to the presence or absence of a keyword (1 for presence, 0 for absence). This explicit keyword tagging mechanism enhances the model's sensitivity to key events and compensates for potential ambiguity in semantic vectors.

[0047] Finally, the semantic vector and the Boolean feature vector are concatenated to form a comprehensive text feature vector. This dual-track design balances the depth of semantic understanding with the precision of keyword response, enabling the text modality to both understand the rich expressions of natural language and quickly respond to predefined key instructions or alarm signals, thus improving the interpretability and controllability of the system.

[0048] Step S205: Align the visual feature vector, audio feature vector, and text feature vector according to the timestamp, and generate a multimodal feature matrix through a feature fusion algorithm.

[0049] Because video, audio, and text acquisition devices may have clock asynchrony, transmission delays, or encoding differences, their timestamps are not naturally aligned. For example, there may be millisecond-level offsets between the camera and microphone, and the generation of subtitle text may lag behind the actual sound output. If the images are directly spliced ​​together based on nominal timestamps, the "seen images" will not match the "heard sounds" or "read text," severely affecting the fusion effect.

[0050] To address this, the system introduces a Dynamic Time Warping (DTW) algorithm to calibrate the timelines of audio and visual features. DTW is a non-linear time alignment technique that, while allowing for time scaling, finds the optimal matching path between two time series, minimizing the overall mismatch cost. By aligning the temporal relationship between audio energy abrupt changes and action start points (such as lip movements or object movement) in video frames, DTW can automatically correct temporal drift between modalities, achieving sub-second or even millisecond-level precise synchronization. Based on this, all modal features are remapped to a unified time reference, ensuring that visual, audio, and text features at the same timestamp truly correspond to multimodal observations at the same moment.

[0051] Subsequently, for each aligned time point, the three feature vectors are standardized (e.g., Z-score normalization) to eliminate the influence of differences in dimensions, scale, and distribution among the modalities. Next, a learnable fully connected neural network layer is used to weight and concatenate the three types of features; this layer essentially allocates importance among modalities. For example, higher weights are given to audio in speech-dominated scenarios, and image features are highlighted in visually salient scenarios.

[0052] The final output is a unified fusion feature vector, arranged in chronological order as a matrix. Each row represents the joint multimodal representation at a time step, and the columns correspond to the feature dimensions. This multimodal feature matrix not only preserves the original information of each modality but also mines the complementary and synergistic relationships between cross-modalities through nonlinear interactions, forming a more complete and discriminative high-level semantic representation than a single modality.

[0053] In the above implementation, based on the modality-based refined feature extraction strategy, the spatial-temporal structure of video, the spectral-energy dynamics of audio, and the semantic-keyword information of text are fully explored to ensure the quality and representativeness of each modality feature. Secondly, the introduction of a dynamic time warping mechanism effectively solves the asynchronous problem commonly found in multimodal systems, ensuring strict correspondence of cross-modal information in the temporal dimension and avoiding misjudgments caused by temporal misalignment. Thirdly, based on a standardized and learnable weighted feature fusion algorithm, adaptive integration between modalities is achieved, so that the fused features can reflect the unique contributions of each modality and also reflect the overall semantic consistency. Finally, the generated multimodal feature matrix, as a structured, high-dimensional, and temporally continuous representation, can directly serve downstream tasks such as anomaly detection, emotion recognition, and event classification, significantly improving the perception accuracy and decision-making ability of intelligent systems.

[0054] Reference Figure 3 As one implementation of step S103, the step of performing multimodal similarity matching between the multimodal feature matrix and the dynamically updated advertising template library to generate a spatiotemporal coordinate set of advertising segments includes: Step S301: Obtain the multimodal feature matrix and the dynamically updated advertising template library; Specifically, the system maintains a dynamically updated advertising template library, which stores standardized multimodal feature templates for known advertising types. Each template represents a typical piece of advertising content (such as a brand promotional video, promotional voiceover, or specific screen layout), and includes its modal feature vector and metadata. This template library is not a static resource, but rather continuously absorbs newly identified advertising patterns as the system runs, thus enabling it to adapt to emerging advertising formats.

[0055] Step S302: For each timestamp unit in the multimodal feature matrix, extract the visual feature vector, audio feature vector, and text feature vector; Since the multimodal feature matrix is ​​a fusion result existing as a whole, directly using the fused vector for template comparison will make it difficult to distinguish the specific contribution of each modality in the matching process, and will also prevent flexible adjustment of the importance weights of different modalities. Therefore, it is necessary to reconstruct the original feature vectors of each modality in order to perform subsequent modality-specific similarity calculations.

[0056] For example, at a certain point in time, a brand logo appears in the image (visually prominent), iconic background music plays (audio prominent), and the subtitle displays "Limited-Time Offer" (text prominent). Only by extracting these three dimensions of features independently can their correspondence with the template be evaluated separately. This "decoupling-re-fusion" strategy ensures the transparency and controllability of the matching process, avoiding information masking problems that may occur during feature fusion. Furthermore, this operation supports an asynchronous update mechanism: when a modal feature changes (such as changing the advertising slogan), it is not necessary to regenerate the entire fusion vector; only the corresponding component needs to be updated, improving the system's flexibility and maintainability.

[0057] Step S303: Perform multimodal weighted similarity calculation between the feature vector of the timestamp unit and each template in the advertising template library to generate an initial matching result set; Due to the varying data types and distribution characteristics of different modalities, a unified distance metric cannot be used. Visual features are typically high-dimensional dense vectors, making cosine similarity suitable for measuring directional consistency and reflecting the semantic closeness of image content. Audio features often exhibit temporal dynamics, especially in cases of rhythmic variations or intonation fluctuations. Simple Euclidean distance cannot capture their temporal patterns, making Dynamic Time Warping (DTW) distance more appropriate. It can flexibly align the timelines of two audio sequences, effectively addressing issues such as varying speech rates or recording delays, and is converted into similarity values ​​through normalization. Text features, often existing as keyword sets or sparse vectors, are effectively measured by Jaccard similarity, which is particularly suitable for comparing short texts and tagged content.

[0058] Building upon this foundation, an adjustable weighted fusion mechanism is introduced. The weight coefficients for visual (α), audio (β), and text (γ) are set according to different application scenarios or ad types, enabling the system to flexibly adapt to ad formats with different dominant modalities. For example, for music ads, the audio weight can be increased; for image-text ads, the emphasis is placed on visual and text. This multimodal weighted similarity calculation not only improves matching accuracy but also enhances the system's adaptability and robustness, ensuring reasonable judgments can still be made even in the event of missing modalities or noise interference.

[0059] Step S304: Perform temporal clustering analysis on the initial matching results of continuous timestamp units, merge matching units that meet the preset continuity threshold, and obtain merged units; Due to the continuity of media content and the sequential nature of advertising playback, the matching results at a single point in time often exhibit fluctuations or fragmentation. For example, a complete advertisement may experience a brief decrease in similarity in a few frames due to camera cuts or silence, leading to matching interruptions and the formation of multiple isolated matching fragments, thus affecting the completeness of the final recognition.

[0060] To address this, the system introduces a temporal clustering analysis mechanism. By performing contextual correlation analysis on the matching results of consecutive timestamps, it identifies candidate regions for ad segments with temporal consistency. Specifically, the system detects whether adjacent time units match the same ad template and determines whether the number of consecutive matches reaches a preset threshold (e.g., no less than 3 consecutive units). Simultaneously, it requires that the time interval between these units does not exceed a specified upper limit (e.g., 2 seconds) to eliminate false positives caused by random matching or repeated playback. Once the conditions are met, these discrete matching points are merged into a coherent candidate ad segment unit.

[0061] Understandably, this process is essentially a rule-based time-series clustering, similar in logic to the "contextual awareness" ability humans have when watching videos: they don't assume an ad ends just because there's a brief black screen, but rather judge its continuity based on the consistency of the preceding and following content. Through this mechanism, the system effectively overcomes the "fragmented recognition" problem caused by traditional frame-by-frame matching methods, improving the accuracy and completeness of ad segment boundaries.

[0062] Step S305: Generate a set of spatiotemporal coordinates for the ad segment based on the start and end timestamps of the merged unit and the highest matching confidence.

[0063] The matching confidence score is not simply the maximum similarity value, but is calculated using a difference-normalized confidence function. Its core idea is to measure the significance of the current best match relative to the second-best match. If the similarity to template A at a certain time point is 0.95, while the highest similarity to other templates is only 0.6, then the matching result is highly exclusive, and the confidence score should be close to 1. Conversely, if the similarities between two templates are 0.8 and 0.78 respectively, the system has difficulty determining the attribution, and the confidence score should be low.

[0064] Understandably, this mechanism simulates the "comparative judgment" process in human decision-making, avoiding false alarms that may result from relying solely on absolute thresholds (such as background music accidentally resembling the tune of an advertisement). The resulting spatiotemporal coordinate set of the advertisement segment not only indicates the specific time period in which the advertisement appears (time coordinates), but can also be supplemented with confidence levels for subsequent review priority ranking or selection of automated processing strategies, achieving a cognitive leap from "whether it matches" to "how confident it is of matching".

[0065] In addition, when the matching similarity is lower than the preset similarity threshold, the feature vector of the unmatched timestamp unit is stored as a new template in the advertising template library.

[0066] Specifically, when the highest matching similarity in a certain time unit is lower than a preset threshold, it indicates that it is not covered by any advertisement in the existing template library, and may represent a new type of advertisement or variant. At this time, the system does not simply ignore this feature, but automatically stores it in the advertisement template library as a potential new advertisement template for reuse in subsequent recognition tasks.

[0067] Understandably, this dynamic update mechanism empowers the system with the ability to continuously learn and self-evolve, enabling it to adapt to the rapidly iterating reality of advertising content. For example, when an e-commerce platform launches new promotional ads during holidays, traditional systems require manual labeling and model retraining, while this solution can capture their features and include them in the template library the first time they appear, allowing for accurate identification the next time they are played.

[0068] To further improve the quality of the template library, the system also integrates a manual review channel, regularly imports confirmed features of illegal advertisements, and uses the Locality Sensitive Hash (LSH) algorithm for approximate deduplication to prevent highly similar templates from being repeatedly added to the library, maintaining the simplicity and efficiency of the data within the library. LSH maps similar vectors to the same or neighboring buckets through a hash function, achieving fast deduplication without requiring a full comparison, greatly improving the management efficiency of large-scale vector libraries.

[0069] The above implementation achieves high accuracy, robustness, and adaptability in ad segment recognition. First, by using multimodal weighted similarity calculation, complementary information from visual, audio, and text is fully integrated, significantly improving recognition accuracy in complex scenarios. Second, the introduction of a temporal clustering mechanism effectively solves the problem of ad segment fragmentation, ensuring the continuity and integrity of recognition results. Finally, a confidence quantification method based on relative differences enhances the reliability of the system's judgment, providing a credible basis for subsequent automated processing.

[0070] Reference Figure 4As one implementation of step S104, the steps of locating the corresponding subset of the multimodal feature matrix based on the spatiotemporal coordinate set of the advertisement segment, performing coarse-grained segmentation based on the energy mutation points of the audio feature vector and the shot switching points of the visual feature vector, and then performing boundary optimization through the clustering results of the text feature vector to generate a list of independent advertisement item objects include: Step S401: Obtain the spatiotemporal coordinate set and multimodal feature matrix of the advertisement segment; The spatiotemporal coordinate set includes the start and end timestamps of the advertisement segments, and the feature matrix includes visual feature vectors, audio feature vectors, and text feature vectors. Specifically, the spatiotemporal coordinate set of the ad segments contains several confirmed ad segments, each marked with a clear start and end timestamp, representing the playback interval of the ad within the entire video stream. The multimodal feature matrix is ​​structured data formed by feature extraction, alignment, and fusion of the original video, audio, and text data. Each row corresponds to a timestamp, and each column represents the feature dimension under a certain modality, forming a high-dimensional vector sequence that is temporally continuous and modally aligned.

[0071] Step S402: Extract the feature subset of the corresponding time interval from the multimodal feature matrix based on the start and end timestamps of the spatiotemporal coordinate set of the advertisement segment; By extracting visual, audio, and text feature vector sequences from specific ad segments, the system can focus on the internal dynamic changes of the ad itself, eliminating interference from other irrelevant content. For example, in a 15-second brand advertisement, there may be multiple sub-events such as product displays, slogan broadcasts, and background music changes, all of which are implicit in the temporal evolution of the feature subset.

[0072] Step S403: Based on the audio feature vectors in the feature subset, calculate the abrupt change point of the energy envelope as the coarse-grained segmentation boundary. The audio energy envelope reflects the trend of sound intensity changes over time, typically exhibiting significant rhythmic fluctuations in advertising: moments such as the start of an advertising tagline, the introduction of background music, the burst of sound effects, or the transition to silence all create distinct rising or falling edges on the energy curve. By calculating the first derivative of this energy envelope and taking its absolute value sequence, these points of dramatic change can be effectively captured. When the derivative value exceeds a preset threshold, it is identified as an energy abrupt change point. These points often correspond to important transitions in the advertising content, such as a shift from brand introduction to promotional information, or a switch from the main advertisement to a disclaimer.

[0073] Understandably, using these mutation points as initial segmentation candidate boundaries can quickly locate the possible switching positions of functional units within the advertisement, forming a preliminary division of the advertisement's structure. This acoustically dynamic segmentation strategy is particularly suitable for voice-driven advertisements, providing reliable segmentation cues without relying on visual or textual information.

[0074] Step S404: Based on the visual feature vectors in the feature subset, detect the shot switching point as a supplementary segmentation boundary; Specifically, the system extracts the HSV color histogram corresponding to each timestamp. This histogram statistically analyzes the pixel distribution of the image across the three channels of hue, saturation, and value, effectively reflecting the overall color composition of the image. By calculating the Bhattacharyya distance between the HSV histograms of adjacent frames, the degree of difference in color distribution between the two images can be quantified. As a measure of probabilistic similarity, a larger Bhattacharyya distance indicates a greater dissimilarity in the color structures of the two frames. When this distance exceeds a set threshold, a scene change can be identified.

[0075] Understandably, this color statistics-based method is robust to changes in lighting and has high computational efficiency, making it suitable for real-time processing. The introduction of camera switching points compensates for the blind spots in audio modalities when there is silence or constant background noise, enabling the system to identify visually driven changes in the advertising structure, such as key nodes like product carousels, scene changes, or logo flashing.

[0076] Step S405: Divide the advertisement segment into initial sub-segments based on the coarse-grained segmentation boundary and the supplementary segmentation boundary; The system employs a unified sorting and interval segmentation along a timeline to arrange all detected boundary points chronologically, using these points as dividing points to segment the original ad into a series of continuous short time segments. Each segment represents a potential functional unit within the ad, such as "brand display," "price explanation," or "purchase guidance." Thanks to the dual-modal collaborative detection mechanism, both abrupt changes in sound and visuals are taken into account, significantly improving the completeness and coverage of the segmentation.

[0077] Step S406: Perform cluster analysis on the text feature vectors within each initial sub-segment, merge semantically similar adjacent sub-segments, and obtain the merged segment; Specifically, for example, an advertisement might first show the product's appearance (visual change), then announce the model name (audio abrupt change), and finally display the parameter list (text update). Although these three actions trigger different boundary detection mechanisms, they should be considered as a single "product introduction" entry from a semantic perspective.

[0078] To address this semantic fragmentation issue, the system performs cluster analysis on the text feature vectors within each initial sub-segment, identifying and merging semantically similar adjacent sub-segments. Specifically, density-based clustering algorithms (such as DBSCAN) can be employed. This algorithm automatically discovers high-density data clusters based on the cosine similarity between vectors, without requiring pre-specified cluster numbers, and is highly adaptable. When the cosine similarity between the text feature vectors of two adjacent sub-segments exceeds a set threshold (e.g., 0.8), it indicates that their semantic content is similar. The system then groups them into the same semantic cluster and merges them into a larger ad entry. This semantically driven merging mechanism effectively improves the semantic integrity of ad entries, avoids information fragmentation caused by technical segmentation, and makes the final output ad entry more consistent with human cognitive habits.

[0079] Step S407: The start and end timestamps of the merged segments are calibrated using a boundary optimization model to generate a list of independent advertising item objects.

[0080] The boundary optimization model uses a Conditional Random Field (CRF) architecture as its core, treating ad item segmentation as a sequence labeling problem: each time point is assigned a label indicating it "belongs to an item" or "is a boundary," with the goal of finding the optimal label sequence that minimizes the overall segmentation loss. The model uses coarse-grained segmentation boundaries as initial nodes, comprehensively considering both visual feature differences and text feature similarities within segments as node feature inputs. Visual feature differences are measured by calculating the variance of visual vectors within segments; a smaller variance indicates more stable content and a higher likelihood of belonging to the same visual unit. Text feature similarities reflect the consistency of language expression within segments; a higher average similarity indicates more semantic coherence. By solving the optimal path of the CRF model using the Viterbi algorithm, the system can dynamically adjust the initial boundaries, eliminating minor offsets caused by detection errors or noise, and generating smoother, more reasonable, and logically consistent final time boundaries.

[0081] Understandably, this optimization approach based on probabilistic graphical models not only integrates multimodal information but also introduces global contextual constraints to ensure that the segmentation results are accurate locally while maintaining overall consistency.

[0082] The above implementation achieves refined reconstruction from macro-level advertising segments to micro-level advertising items. First, the bimodal boundary detection mechanism fully utilizes the complementarity of audio energy mutations and visual camera transitions, solving the problem of incomplete single-modal segmentation. Second, text semantic-based density clustering effectively integrates semantically coherent but technically fragmented sub-segments, improving the semantic integrity of advertising items. Third, the conditional random field-driven boundary optimization model achieves sub-second refinement of time boundaries under the joint constraints of multimodal features, significantly improving segmentation accuracy. Finally, the output list of independent advertising item objects not only includes precise start and end timestamps but can also be extended with additional metadata such as the segment ID and feature fingerprint hash, facilitating subsequent storage, retrieval, and analysis.

[0083] Reference Figure 5 As one implementation of step S105, the steps of performing multimodal similarity comparison between the list of independent advertising items and a pre-set prohibited sample library, analyzing the contextual semantic relevance in conjunction with a regional preset policy database, and outputting a violation item report and confidence level include: Step S501: Obtain the list of independent advertising items, the preset prohibited sample library, and the regional preset policy database; The individual ad item object contains visual feature vectors, audio feature vectors, text feature vectors, and spatiotemporal metadata. Specifically, each entry in the list of independent advertising entries encapsulates complete multimodal feature information, including visual feature vectors, audio feature vectors, and text feature vectors. These vectors respectively represent the deep semantics of the advertisement in terms of image content, sound signals, and language expression. At the same time, each entry also includes spatiotemporal metadata, such as its precise start and end timestamps, broadcast channels, and geographical coverage, which constitute the temporal and spatial anchors for subsequent policy matching and contextual analysis.

[0084] The banned sample library is a dynamically maintained knowledge base of illegal content, which stores advertising templates that have been confirmed as illegal or inappropriate for dissemination. Its content exists in the form of multimodal feature vectors, covering typical illegal images (such as vulgar images), sensitive audio (such as suggestive voice), prohibited statements (such as false advertising), etc., and is used as a comparison benchmark to identify known risk patterns.

[0085] The regional pre-set policy database is a structured collection of laws and regulations, organized according to different administrative regions, industry categories, broadcast times, and other dimensions. It includes specific clauses such as "prohibiting alcohol advertising during periods of protection for minors" and "restricting medical advertising in specific regions."

[0086] Step S502: Extract the visual feature vector, audio feature vector, and text feature vector for each individual advertisement item; Step S503: Calculate the similarity between the visual feature vector and the visual templates in the pre-set banned sample library to generate a visual similarity value; This process typically employs metrics such as cosine similarity or Euclidean distance to measure the semantic similarity between the current advertisement and known violating templates in a high-dimensional feature space. For example, if an advertisement contains a character's pose or background composition that is highly similar to a banned vulgar advertisement, even if the brand logo is different, the visual similarity may still exceed a threshold, triggering an alert. This feature space-based comparison method surpasses traditional coarse-grained recognition based on keywords or image hashes, capturing deep-seated visual semantic imitation behavior and effectively preventing "skin-swapping violations."

[0087] Step S504: Perform dynamic time warping matching between the audio feature vector and the audio templates in the pre-set banned sample library to generate audio similarity values; The system performs Dynamic Time Warping (DTW) matching on audio feature vectors to address common issues in advertising audio such as rhythm changes, speech rate adjustments, or background noise interference. The DTW algorithm allows two audio sequences to be non-linearly stretched or compressed along the time axis to find the optimal alignment path, thus more accurately assessing their content consistency. For example, even if the speech of a violating advertisement is deliberately slowed down or has echo processing added, as long as its tonal structure and intonation pattern remain consistent, DTW can still identify its high degree of matching with the prohibited template and generate a corresponding audio similarity value.

[0088] Step S505: Calculate the spatial distance between the text feature vector and the semantic vector of the pre-set banned sample library to generate a text similarity value; The system calculates the spatial distance between the text feature vector and the semantic vectors in the prohibited sample library. Common methods include cosine distance or Euclidean distance, which are used to measure whether the advertising copy is semantically close to known violations. For example, absolute terms such as "completely cure" and "never relapse" may have been recorded as high-risk semantic vectors. When the semantic vector of a new advertising text is too close to its high-risk semantic vector, it is judged as a potential violation.

[0089] Step S506: Combine visual similarity values, audio similarity values, and text similarity values ​​to generate a comprehensive violation probability; This fusion process is not a simple averaging; rather, it flexibly sets the weight coefficients for visual (α), audio (β), and text (γ) based on the actual application scenario, ensuring that different types of advertisements are reasonably evaluated. For example, in pharmaceutical advertisements, text information often carries key efficacy claims and should be given higher weight; while in music advertisements, the risk of infringement for audio melodies is higher, and the audio weight can be increased accordingly. This configurable weighting mechanism gives the system good adaptability, enabling it to maintain consistent judgment logic across different industries and media environments. The overall violation probability serves as a unified risk score, quantifying the overall violation tendency of the advertisement at the content level, becoming the core basis for subsequent decision-making.

[0090] Step S507: Based on the regional preset policy database, locate the associated regional policy rules according to the spatiotemporal metadata of the independent advertising item object; Specifically, the system analyzes the geographic codes (such as province and city) and broadcast timestamps (accurate to the minute) in the entries, retrieves policy rule sets that match the spatiotemporal conditions from the regional preset policy database, and further filters out clauses marked "effective immediately," excluding expired or not yet implemented regulations. For example, if an educational advertisement is legal nationwide but temporarily banned from prime time in a pilot city due to the "double reduction" policy, the system will automatically identify this restriction based on the city's real-time policy database. This dynamic matching mechanism based on spatiotemporal dimensions enables precise implementation of regulations, avoiding misjudgments or omissions caused by "one-size-fits-all" reviews, and reflects a significant advancement in intelligent supervision.

[0091] Step S508: Analyze the semantic conflict between independent advertising items and adjacent program content using a contextual semantic association model; Specifically, for example, if an alcohol advertisement appears after a children's cartoon, even if the content itself is legal, it still constitutes a hidden violation from a communication ethics perspective. The system loads program text features for a preset duration (e.g., 5 minutes before and after) before and after the advertisement airs, and uses a pre-trained language model (such as BERT) to calculate the semantic association vector between the advertisement and adjacent programs, capturing the strength of their association in terms of theme, emotion, and audience targeting.

[0092] Subsequently, by combining a pre-defined keyword database for prohibited scenarios (such as "children's programs" and "alcohol," "health lectures" and "exaggerated claims about health products"), the system detects the existence of high-risk semantic combinations. Furthermore, the system can construct a knowledge graph of prohibited scenarios, using prohibited entities (such as "minors" and "prescription drugs") as nodes and prohibited relationships between entities as edges. By calculating the weight of the connection paths between the advertising semantic vector and the graph nodes, it determines whether a high-risk association is triggered. When the weight exceeds a preset threshold, it is marked as a semantic conflict, generating a semantic conflict degree index. This context-aware capability enables the system to identify those "legal but inappropriate" gray-area advertisements, improving the depth of review and social adaptability.

[0093] Step S509: Generate a violation report and confidence level based on the comprehensive violation probability, regional policy rules, and semantic conflict degree.

[0094] The confidence level is not a simple threshold judgment, but is calculated through a piecewise function and a dynamic weighting mechanism. The system first calculates the difference (ΔP) between the overall probability of violation and the high-risk threshold, reflecting the significance of the violation; at the same time, it weights and sums the semantic conflict degree and the policy violation level to reflect the severity of external constraints.

[0095] Based on this, a piecewise function is used to map the final confidence value: when ΔP is significantly higher than the threshold, it is directly judged as a high-confidence violation; when it is in the critical interval, the output is adjusted by combining contextual risk weighting. The weight coefficients can be dynamically adjusted according to the timeliness of the policy, for example, automatically increasing the weight of relevant policy clauses during major holidays or sensitive periods to enhance the system's responsiveness.

[0096] The final violation report not only includes a clear violation determination, but also a detailed chain of evidence (such as matching templates, violation of policy provisions, and descriptions of contextual conflicts) and a confidence level, supporting various downstream processing strategies such as manual review, automatic interception, or tiered alarms.

[0097] In the above implementation, firstly, the multimodal fusion comparison mechanism significantly improves the identification coverage of known violation patterns, effectively preventing behaviors that circumvent review; secondly, dynamic policy matching based on spatiotemporal metadata enables precise and real-time enforcement of regulations, adapting to the complex and ever-changing regulatory environment; thirdly, contextual semantic conflict analysis breaks through the limitations of isolated content judgment, enhancing the system's ability to perceive hidden violations; and finally, through a confidence quantification model combining piecewise functions and dynamic weights, it generates report outputs that are interpretable and provide operational guidance, providing strong support for regulatory decision-making.

[0098] Reference Figure 6As one implementation of step S106, the step of triggering an alarm instruction based on the confidence level of the violation report includes: Step S601: Parse the violation report to extract the confidence parameter value, and map the confidence parameter value to a predefined confidence level threshold range; The confidence parameter is a comprehensive score generated by the preceding violation detection module through complex calculations such as multimodal similarity fusion, contextual semantic conflict analysis, and regional policy matching. It is usually normalized to a real number in the range [0,1] to quantify the credibility of the advertisement being judged as a violation.

[0099] For example, an entry with a confidence level of 0.97 may mean that it is highly consistent with the banned sample in the visual, audio, and text modalities and violates a clear regional policy; while an entry with a confidence level of 0.45 may only have ambiguity at the text level and lack supporting evidence from other modalities.

[0100] In addition, to ensure data reliability, the system also has built-in verification logic: if the confidence value is found to be outside the reasonable range (such as negative or greater than 1), it will be automatically reset to the preset default value (such as 0.5) and the abnormal state will be marked in the alarm instruction metadata to prevent erroneous actions caused by upstream module abnormalities, which reflects the robust design of the system in terms of abnormality handling.

[0101] Step S602: Generate the corresponding confidence level identifier based on the mapping result; The confidence level is categorized into high confidence, medium confidence, and low confidence. Specifically, the grading mechanism is not fixed, but adopts a configurable dynamic threshold model. Typically, [0, α) is divided into low confidence level, [α, β) into medium confidence level, and [β, 1] into high confidence level. α and β are adjustable parameters that can be flexibly set according to different application scenarios.

[0102] For example, in live broadcast television scenarios, to ensure broadcast security, β can be set to 0.8, meaning that only when the confidence level exceeds 80% is it considered high risk, thus avoiding the accidental interruption of normal programs. In online platform content review, to improve the capture rate of sensitive content, α can be set to 0.3, expanding the medium-to-high confidence range and increasing the coverage of early warnings. This parametric design allows the system to adapt to the needs of different regulatory intensities, demonstrating its flexibility in engineering deployment.

[0103] Step S603: Match the pre-configured alarm action rule set according to the confidence level identifier, generate differentiated alarm instructions and send them to the target alarm execution terminal.

[0104] This rule set is the core carrier of the system's behavior strategy. It defines the response actions corresponding to different levels and supports dynamic updates through an external configuration management interface, ensuring that policy adjustments take effect without restarting the system.

[0105] In some embodiments, when the confidence level is high, the system determines that the violation is highly certain and strong intervention measures should be taken immediately. Therefore, it generates an instruction that includes real-time audio-visual alarms and automatic interception operations. For example, in a live broadcast, this instruction can trigger the front-end device to cut off video output, pop up a red warning box, and simultaneously push a text message to the on-duty personnel's mobile phone, achieving "second-level circuit breaker" and effectively preventing the spread of illegal content.

[0106] When the level is medium confidence, the system considers there to be a high probability but still requires manual confirmation. Therefore, it generates a delayed alarm instruction that requires manual review and pushes the item to the review workbench, along with a screenshot of the matching sample, policy basis, and a summary of semantic conflict analysis, for regulatory personnel to make a quick judgment.

[0107] For low-confidence entries, the system treats them as potential clues rather than explicit threats, generating only silent processing instructions and writing relevant information into the audit log database for subsequent big data analysis or model training feedback, avoiding disruption to normal business processes. This tiered response mechanism achieves optimal resource allocation: high-risk events are prioritized, medium-risk events are recorded for future investigation, and low-risk events are silently archived, significantly improving regulatory efficiency and system availability.

[0108] The above implementation achieves refined management of alerting behavior. Through a confidence level grading mechanism, it transforms ambiguous risk assessments into clear decision-making paths; through configurable rule sets, it empowers the system to flexibly adapt to different regulatory scenarios; and through differentiated instruction design, it balances automated intervention with manual review, thereby avoiding extreme situations such as excessive alerts or delayed responses. This technical solution effectively improves the practicality and reliability of the advertising supervision system, truly translating technical judgments into management actions, and providing crucial support for building a digital content governance system.

[0109] Reference Figure 7 As a further implementation of the automatic advertising content monitoring method, after the step of outputting violation reports and confidence levels, the method further includes: Step S701: Send violation entries with a confidence level lower than a preset threshold from the violation entry report to the manual review terminal; These low-confidence entries typically indicate significant uncertainty in the system's judgment process. For example, an advertisement might use boundary words like "ultimate experience" in its text, but the visuals do not contain any obvious violations, and no clear prohibited samples are matched, resulting in a comprehensive confidence level of only 0.45, which is insufficient to trigger a high-level alarm.

[0110] Step S702: Receive the correction data for misjudged violations from the manual review terminal; The misclassification correction data includes the misclassified entry identifier and the corrected category label; Specifically, after the reviewers complete the annotations on the terminal interface, the system receives the feedback data on the correction of misjudgments. This data includes the unique identifier (such as UUID) of the misjudged item and the corrected category label (such as "legitimate promotion", "soft implantation", or "new type of health product advertisement"). This data structure ensures the accurate traceability of the feedback information and avoids erroneous updates due to confusing labeling.

[0111] Step S703: Based on the misjudged item identifier, extract the multimodal feature vector of the corresponding advertising item from the historical processing record, and generate candidate template features through a clustering algorithm; These multimodal feature vectors are high-dimensional semantic representations obtained through deep encoding of video frames, audio signals, and text content during the ad detection phase. They fully preserve the joint feature state of the ad in terms of visual composition, sound rhythm, and language expression. For example, a skincare product ad misjudged as "false advertising" might have visual features including close-ups of the model's face and product packaging displays, audio features reflecting soft background music and a gentle narration tone, and text features containing non-absolute statements such as "improves skin texture" and "gentle formula." These vectors collectively constitute the essential characterization of the ad content and serve as the foundational material for subsequent template generation.

[0112] Building upon this foundation, the system performs incremental learning on the extracted multimodal feature vectors. First, it aggregates them into candidate template feature vectors using clustering algorithms (such as K-means or DBSCAN). This process aims to eliminate noise interference in individual samples and extract more representative common features. For example, if multiple entries confirmed as "legitimate daily chemical advertisements" show a clustering trend in the feature space, the system will calculate their cluster centers as candidate templates, thus forming a standardized representation of this type of advertisement. This clustering operation not only improves the generalization ability of the templates but also avoids knowledge base pollution caused by a single misjudged sample, ensuring that newly added or updated templates possess statistical stability and representativeness.

[0113] Step S704: Calculate the cosine similarity matrix between the candidate template features and the existing template features in the advertising template library; Cosine similarity, as a metric for vector orientation consistency, effectively reflects the semantic proximity of two templates in a high-dimensional feature space. By constructing a similarity matrix, the system can comprehensively evaluate the matching relationship between candidate templates and existing templates in the library.

[0114] Step S705: Determine whether there is a similarity value in the cosine similarity matrix that exceeds the preset merging threshold; if yes, proceed to step S706; if no, proceed to step S707. Step S706: The candidate template features are weighted and fused with the corresponding existing template features to generate a fused template, and the fused existing template features are deleted. If a similarity value in the cosine similarity matrix exceeds a preset merging threshold (e.g., 0.85), it indicates that the candidate template is highly similar to an existing template, potentially representing different instances or slight variations of the same type of advertisement. In this case, the system does not insert it as an independent new template, but instead initiates a weighted fusion mechanism: linearly combining the candidate template features with the existing template features according to certain weights (e.g., weighted by historical data volume or time decay factor) to generate a new fused template feature vector, which then replaces the original template.

[0115] Understandably, this operation enables dynamic optimization of the template library, preserving existing knowledge while incorporating the feature evolution of new samples. This prevents the template library from expanding due to repeated inputs and enhances the templates' inclusiveness towards advertising variations. For example, if a brand's advertisement changes its spokesperson and the visual style changes slightly, the system can update the template to adapt to the new version without having to recreate the entry.

[0116] Step S707: Add the candidate template features as new template entries to the advertising template library.

[0117] If all similarity values ​​are below the merging threshold, it indicates that the candidate template represents a previously unseen type of advertisement or a significantly different expression pattern, and the system formally includes it as a new template entry in the advertisement template library. This new mechanism empowers the system with the ability to autonomously recognize new types of illegal or legal advertisements, enabling it to gradually expand its knowledge boundaries.

[0118] For example, a new e-commerce platform adopted a completely new advertising narrative structure, which was not initially recognized. However, after manual verification, its characteristics were modeled into a new template, allowing for accurate categorization during subsequent playback. After all update operations are completed, the system generates an updated advertising template library with a version identifier. This version number is used to identify the timestamp of this update, the source of the operation, and the change summary, ensuring that every evolution of the template library is traceable and rollbackable.

[0119] In the above implementation, this incremental learning mechanism based on human feedback not only significantly reduces the system's repeated misjudgment rate but also enhances its adaptability to the evolution of the advertising ecosystem, making it particularly suitable for dealing with rapidly iterating online marketing strategies and cross-regional cultural differences. Human-machine collaboration achieves a knowledge loop, improving the system's long-term accuracy and robustness; clustering and fusion mechanisms ensure the template library's conciseness and efficiency, avoiding knowledge redundancy; and version control ensures system maintainability and compliance auditing capabilities. This design meets the core requirements of intelligent content supervision systems for self-learning capabilities.

[0120] This application also discloses an automatic advertising content monitoring system.

[0121] An automatic advertising content monitoring system, the system comprising: The acquisition module is used to acquire streaming media data, including video streams, audio streams, and auxiliary text; The feature fusion module is used to extract visual feature vectors from the video stream, audio feature vectors from the audio stream, and text feature vectors from the auxiliary text, and then fuse them to obtain a multimodal feature matrix. The multimodal similarity matching module is used to perform multimodal similarity matching between the multimodal feature matrix and the dynamically updated advertising template library to generate a set of spatiotemporal coordinates for advertising segments; The ad entry generation module is used to locate the corresponding subset of the multimodal feature matrix based on the spatiotemporal coordinate set of ad segments, perform coarse-grained segmentation based on the energy mutation points of audio feature vectors and the shot switching points of visual feature vectors, and then perform boundary optimization through the clustering results of text feature vectors to generate a list of independent ad entry objects. The violation analysis module is used to compare the list of independent advertising items with the pre-set prohibited sample library in a multimodal similarity, analyze the contextual semantic relevance in combination with the regional preset policy database, and output violation item reports and confidence scores. The tiered alarm module is used to trigger alarm commands based on the confidence level of the reported violation items.

[0122] The automatic advertising content monitoring system of this application embodiment can implement any of the above-described automatic advertising content monitoring methods, and the specific working process of each module in the automatic advertising content monitoring system can be referred to the corresponding process in the above-described method embodiments.

[0123] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0124] This application also discloses a computer device.

[0125] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the automatic monitoring method for advertising content as described above.

[0126] This application also discloses a computer-readable storage medium.

[0127] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the above-described methods for automatically detecting advertising content.

[0128] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0129] In this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0130] Although this application has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings, disclosure, and appended claims, will understand and implement other variations of the disclosed embodiments in carrying out the claimed application. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0131] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A method for automatically monitoring advertising content, characterized in that, The monitoring method includes: Acquire streaming media data, including video streams, audio streams, and supplementary text; Visual feature vectors are extracted from the video stream, audio feature vectors are extracted from the audio stream, and text feature vectors are extracted from the auxiliary text. These are then fused to obtain a multimodal feature matrix. The multimodal feature matrix is ​​matched with a dynamically updated ad template library using multimodal similarity to generate a spatiotemporal coordinate set of ad segments; Based on the spatiotemporal coordinate set of the advertisement segment, locate the corresponding subset of the multimodal feature matrix, perform coarse-grained segmentation based on the energy mutation point of the audio feature vector and the shot switching point of the visual feature vector, and then perform boundary optimization through the clustering result of the text feature vector to generate a list of independent advertisement item objects. The list of independent advertising items is compared with a pre-set prohibited sample library using multimodal similarity analysis. The contextual semantic relevance is analyzed in conjunction with a regional pre-set policy database, and a report of violations and confidence level are output. An alarm command is triggered based on the confidence level of the reported violation item.

2. The method for automatically monitoring advertising content according to claim 1, characterized in that, The steps of extracting visual feature vectors from the video stream, extracting audio feature vectors from the audio stream, extracting text feature vectors from the auxiliary text, and fusing them to obtain a multimodal feature matrix include: Acquire video streams, audio streams, and auxiliary text; Frame sampling processing is performed on the video stream, and visual feature vectors are extracted using an object detection model; The Mel frequency cepstral coefficients and energy envelope parameters are calculated for the audio stream to generate an audio feature vector; The auxiliary text is semantically embedded and matched with a pre-stored keyword library to generate a text feature vector; The visual feature vector, audio feature vector, and text feature vector are aligned by timestamps, and a multimodal feature matrix is ​​generated by a feature fusion algorithm.

3. The method for automatically monitoring advertising content according to claim 1, characterized in that, The steps of performing multimodal similarity matching between the multimodal feature matrix and the dynamically updated ad template library to generate a spatiotemporal coordinate set of ad segments include: Obtain multimodal feature matrices and a dynamically updated ad template library; For each timestamp unit in the multimodal feature matrix, extract the visual feature vector, audio feature vector, and text feature vector; The feature vector of the timestamp unit is compared with the templates in the advertising template library using a multimodal weighted similarity calculation to generate an initial matching result set; Temporal clustering analysis is performed on the initial matching results of continuous timestamp units, and matching units that meet the preset continuity threshold are merged to obtain merged units; Based on the start and end timestamps of the merged unit and the highest matching confidence, a set of spatiotemporal coordinates for the ad segment is generated.

4. The method for automatically monitoring advertising content according to claim 3, characterized in that, The steps of locating the corresponding subset of the multimodal feature matrix based on the spatiotemporal coordinate set of the advertisement segments, performing coarse-grained segmentation based on the energy mutation points of the audio feature vectors and the shot transition points of the visual feature vectors, and then performing boundary optimization based on the clustering results of the text feature vectors to generate a list of independent advertisement item objects include: Obtain the spatiotemporal coordinate set and multimodal feature matrix of the advertisement segment; wherein, the spatiotemporal coordinate set includes the start and end timestamps of the advertisement segment, and the feature matrix includes visual feature vectors, audio feature vectors and text feature vectors; Based on the start and end timestamps of the spatiotemporal coordinate set of the advertising segments, extract the feature subsets corresponding to the time intervals from the multimodal feature matrix; Based on the audio feature vectors in the feature subset, the abrupt change points of the energy envelope are calculated as coarse-grained segmentation boundaries. Based on the visual feature vectors in the feature subset, the camera switching points are detected as supplementary segmentation boundaries. Based on the coarse-grained segmentation boundary and the supplementary segmentation boundary, the advertisement segment is divided into initial sub-segments; Cluster analysis is performed on the text feature vectors within each initial sub-segment, and semantically similar adjacent sub-segments are merged to obtain the merged segment; The start and end timestamps of the merged segments are calibrated using a boundary optimization model to generate a list of independent ad item objects.

5. The method for automatically monitoring advertising content according to claim 1, characterized in that, The steps of comparing the list of independent advertising items with a pre-set prohibited sample library using multimodal similarity analysis, combining the analysis of contextual semantic relevance with a regional preset policy database, and outputting a violation report and confidence level include: Obtain a list of independent advertising item objects, a pre-set prohibited sample library, and a regional preset policy database; wherein, the independent advertising item objects include visual feature vectors, audio feature vectors, text feature vectors, and spatiotemporal metadata; Extract the visual feature vector, audio feature vector, and text feature vector for each individual ad item; The visual feature vector is compared with the visual templates in the pre-set banned sample library to generate a visual similarity value. The audio feature vector is dynamically time-normalized and matched with the audio templates in the pre-set banned sample library to generate an audio similarity value; The text feature vector is compared with the semantic vector of the pre-set banned sample library to calculate the spatial distance and generate a text similarity value. By combining the visual similarity value, audio similarity value, and text similarity value, a comprehensive violation probability is generated; Based on the regional preset policy database, the associated regional policy rules are located according to the spatiotemporal metadata of the independent advertising item object; The semantic conflict between the independent advertisement item and the adjacent program content is analyzed using a contextual semantic association model. Based on the comprehensive violation probability, regional policy rules, and semantic conflict degree, a violation entry report and confidence level are generated.

6. The method for automatically monitoring advertising content according to claim 5, characterized in that, The steps for triggering an alarm instruction based on the confidence level of the violation report include: The violation report is parsed to extract the confidence parameter value, and the confidence parameter value is mapped to a predefined confidence level threshold range; The corresponding confidence level identifier is generated based on the mapping result; wherein, the confidence level identifier includes high confidence level, medium confidence level and low confidence level; Based on the confidence level identifier, a pre-configured alarm action rule set is matched to generate a differentiated alarm instruction and send it to the target alarm execution terminal.

7. A method for automatically monitoring advertising content according to any one of claims 1 to 6, characterized in that, After the steps of outputting the violation report and confidence level, the following are also included: Send violation entries with a confidence level lower than a preset threshold from the violation report to the manual review terminal; The system receives erroneous judgment correction data for violation items from the manual review terminal; wherein the erroneous judgment correction data includes the erroneous item identifier and the corrected category label. Based on the misjudged entry identifier, the multimodal feature vector of the corresponding advertising entry is extracted from the historical processing records, and candidate template features are generated through a clustering algorithm; Calculate the cosine similarity matrix between the candidate template features and the existing template features in the advertising template library; Determine whether there are similarity values ​​in the cosine similarity matrix that exceed a preset merging threshold; If so, the candidate template features are weighted and fused with the corresponding existing template features to generate a fused template, and the fused existing template features are deleted. If not, the candidate template features will be added as new template entries to the advertising template library.

8. An automatic advertising content monitoring system, characterized in that, The system includes: The acquisition module is used to acquire streaming media data, including video streams, audio streams, and auxiliary text; The feature fusion module is used to extract visual feature vectors from the video stream, audio feature vectors from the audio stream, and text feature vectors from the auxiliary text, and fuse them to obtain a multimodal feature matrix; The multimodal similarity matching module is used to perform multimodal similarity matching between the multimodal feature matrix and the dynamically updated advertising template library to generate a set of spatiotemporal coordinates for advertising segments; The ad entry generation module is used to locate the corresponding subset of the multimodal feature matrix based on the spatiotemporal coordinate set of the ad segment, perform coarse-grained segmentation based on the energy mutation points of the audio feature vector and the shot switching points of the visual feature vector, and then perform boundary optimization through the clustering results of the text feature vector to generate a list of independent ad entry objects. The violation analysis module is used to perform multimodal similarity comparison between the list of independent advertising items and the pre-set prohibited sample library, analyze the contextual semantic relevance in conjunction with the regional preset policy database, and output a violation item report and confidence level. The graded alarm module is used to trigger alarm commands according to the confidence level of the reported violation items.

9. A computer device, characterized in that: The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Traditional media advertisement separation and identification method based on multi-modal large model

    CN120340545A

  • Video content understanding method of multi-mode advertisement inventory intelligent matching system

    CN120388324A

  • Advertisement risk early warning method based on multi-modal feature fusion

    CN120450774A

  • Advertisement video badness-oriented fine-grained detection method based on knowledge graph and large language model

    CN120451872A

  • In-stream advertisement matching method implementing selection of advertisement for smooth transition from stream of the digital contents

    KR102804642B1