A method and system for detecting compliance of app advertising content
By constructing a compliance ontology and evidence index, cross-modal semantic alignment and evidence association are achieved, solving the problem of cross-modal semantic alignment and evidence association in advertising compliance detection, and ensuring the consistency and readability of the display of advertising statements and disclaimers.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING HONGTU XINDA TECH CO LTD
- Filing Date
- 2025-10-24
- Publication Date
- 2026-04-17
AI Technical Summary
Existing advertising compliance detection methods cannot effectively perform cross-modal semantic alignment and evidence association, and cannot simultaneously verify the legality of advertising claims and the consistency of display. They are prone to omissions or errors in recognition, especially in dynamic or multimedia scenarios.
Construct a compliance ontology and evidence index, generate a partitioned execution plan, extract statement fragments and disclaimer fragments through cross-modal anchor point alignment, perform minimum evidence pairing and verify the consistency of the advertisement display, and generate a preliminary statement judgment list and a disposal list.
It achieves cross-modal semantic alignment and evidence association of advertising content, ensuring the readability and consistency of the display of statements and disclaimers, and providing unified semantic basis and interpretable compliance ontology support.
Smart Images

Figure CN121235759B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of advertising compliance detection technology, and in particular to a method and system for detecting the compliance of APP advertising content. Background Technology
[0002] With the diversification of internet advertising methods, the compliant management of advertising content on mobile apps has gradually become an important direction for regulation and platform governance. Current advertising compliance detection mainly relies on text comparison and keyword recognition methods, achieving initial screening by matching advertising copy with sensitive word databases. However, advertising materials typically contain multimodal information such as text, images, and audio. There are complex temporal and spatial relationships between advertising statements, disclaimers, and display behaviors. Relying solely on static detection of a single modality is insufficient to accurately identify illegal content, especially in dynamic or multimedia scenarios, where omissions or misjudgments are prone to occur.
[0003] The following shortcomings still exist in advertising compliance testing: First, a semantic mapping mechanism from legal provisions to advertising expression has not been established, and there is a lack of interpretable compliance ontology support, which makes it difficult to match rules and trace evidence; Second, it is impossible to perform multimodal consistency verification on the statements and disclaimers in advertisements in terms of timeline, screen coordinates and voice synchronization, which results in some statements being legal in form but not in presentation that meets regulatory requirements. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a method for detecting the compliance of APP advertising content, which solves the problems of existing advertising content detection methods that lack cross-modal semantic alignment and evidence association mechanisms and cannot simultaneously verify the legality of advertising claims and the consistency of display.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0007] This invention provides a method for detecting the compliance of APP advertising content, which includes,
[0008] Construct a compliance ontology and evidence index, and generate a partitioned execution plan based on the timeline and screen coordinates of the advertising material, and output a list of task slices;
[0009] Using the trigger information carried in the task slice list, the declaration fragment and the corresponding disclaimer fragment are extracted from the text of the advertising material, and cross-modal anchor point alignment is completed by combining the video frame coordinates, subtitle timestamps and audio fragments to generate a declaration card set and a cross-modal anchor point table;
[0010] Based on the subject name, product information and time elements recorded in the declaration card set and cross-modal anchor table, evidence retrieval is performed and minimum evidence matching is conducted to output a preliminary declaration judgment list;
[0011] Based on the conclusions in the preliminary judgment list and the cross-modal anchor table, the readability and presentation consistency of the advertisement display are verified, and contextual legality detection is performed. The detection results are aggregated with the conclusions in the preliminary judgment list to form disposal items, and a disposal list and evidence presentation report are output.
[0012] As a preferred embodiment of the method for detecting the compliance of APP advertising content according to the present invention, the compliance ontology is constructed in the following specific steps.
[0013] The text of advertising display specifications and industry standards is parsed to extract constraints, limitations, and conditions.
[0014] Extract constraints, limitations, and conditions and convert them into logical nodes;
[0015] Set corresponding trigger conditions and exemption conditions for each logical node and store them as a rule node table;
[0016] Generate a trigger word list based on the trigger conditions in the rule node table;
[0017] The entries in the rule node table are mapped to the trigger vocabulary to generate a compliance ontology.
[0018] As a preferred embodiment of the method for detecting the compliance of APP advertising content according to the present invention, the evidence index is constructed in the following specific steps.
[0019] The categories of entities to be verified are determined based on the rule nodes in the compliance ontology, and the entity name, product characteristics, and time elements are extracted from the advertising materials.
[0020] The extracted subject name, product characteristics, and time elements are organized into a three-layer field structure: subject layer, product layer, and time layer.
[0021] Establish field mapping relationships for external data and generate unique index codes, bind each layer of fields to the corresponding fields in the external data, and generate an evidence index library.
[0022] As a preferred embodiment of the method for detecting the compliance of APP advertising content according to the present invention, the step of generating a partitioned execution plan based on the time axis and screen coordinates of the advertising material and outputting a task slice list includes the following steps:
[0023] Read rule nodes and index fields from the compliance ontology and evidence index;
[0024] The text is segmented according to the timeline of the advertising material, and image grid blocks and audio segments are divided according to the screen coordinates to generate material slices;
[0025] Each material slice is matched with the trigger word list to identify potential trigger words and entries to be verified, and the time range, spatial range and trigger word number of the slice are recorded to form a task slice list.
[0026] As a preferred embodiment of the method for detecting the compliance of APP advertising content described in this invention, the generation of the declaration card set and the cross-modal anchor table includes the following steps:
[0027] Based on the task slice list, identify the claim content, object and qualifier in the corresponding text range, generate the claim fragment, and identify the disclaimer statement in the same interval to generate the disclaimer fragment;
[0028] Match the text position with the video frame coordinates and register the subtitle timestamp with the audio time segment to establish a synchronous association between text, image and audio;
[0029] The synchronization results are recorded as time anchors and spatial anchors, and a declaration card is generated based on the correspondence between the declaration fragment and the disclaimer fragment;
[0030] The declaration cards and their anchor information are organized into a declaration card set and a cross-modal anchor table.
[0031] As a preferred embodiment of the method for detecting the compliance of APP advertising content described in this invention, the step of performing evidence retrieval and minimum evidence matching to output a preliminary judgment list includes the following steps:
[0032] Read the subject name, product information, and time element from the declaration card set and cross-modal anchor table;
[0033] Use the evidence index to locate external data records and filter out the minimum record fragments that can cover the key points of the statement;
[0034] Establish a one-to-one correspondence between the recorded fragments and the statement cards to generate evidence pointers;
[0035] Based on the evidence pointers, conclusions of support, refutation, and insufficient evidence are generated for each statement and compiled into a preliminary list of statements.
[0036] As a preferred embodiment of the method for detecting the compliance of APP advertising content according to the present invention, the verification of the readability and presentation consistency of the advertising screen includes the following steps:
[0037] Based on the preliminary list of statements and the cross-modal anchor point table, determine the range in which each statement and its disclaimer appear in the image;
[0038] Check whether the disclaimer is displayed at the same time as the statement, and verify whether the display duration, font height, color contrast and occlusion ratio meet the display threshold requirements;
[0039] Verify the semantic and temporal synchronization consistency between the subtitle text and the audio content, and summarize the verification results.
[0040] As a preferred embodiment of the method for detecting the compliance of APP advertising content according to the present invention, the execution context legalization detection includes the following steps:
[0041] Read the validation results, locate the context window of each statement, identify the occurrence of extreme words and sensitive words, and determine whether the corresponding disclaimer fragment exists and whether the validation result is qualified.
[0042] Determine whether the corresponding evidence pointers in the preliminary judgment list exist;
[0043] The document is marked as unqualified when a disclaimer is missing, the presentation is inadequate, or one of the evidentiary points is missing; it is marked as valid when a disclaimer is present, the presentation is qualified, and the evidentiary points are present.
[0044] The invalid and valid tags are combined into a contextual validation result.
[0045] As a preferred embodiment of the method for detecting the compliance of APP advertising content described in this invention, the step of aggregating the detection results with the conclusions of the preliminary judgment list to form a disposal item includes the following steps:
[0046] The contextual legalization results are merged with the presentation verification results, and a correspondence is established with the conclusions in the initial judgment list of the declarations. The disposal type of each declaration is determined to be one of blocking, review, and release, and the corresponding disposal description is generated.
[0047] Organize the disposal instructions, time anchors, spatial anchors, and evidence pointers into disposal items;
[0048] The disposal items are compiled into a disposal list, and an evidence presentation report is generated using the disposal list as an index.
[0049] Secondly, this invention provides a system for detecting the compliance of APP advertising content, including,
[0050] The rule partitioning module constructs a compliance ontology and evidence index, and generates a partition execution plan based on the timeline and screen coordinates of the advertising material, outputting a list of task slices;
[0051] The text alignment module uses the trigger information carried in the task slice list to extract the declaration fragment and the corresponding disclaimer fragment from the advertisement text, and combines the video frame coordinates, subtitle timestamps and audio fragments to complete the cross-modal anchor point alignment, generating a declaration card set and a cross-modal anchor point table;
[0052] The evidence matching module performs evidence retrieval and minimum evidence matching based on the subject name, product information and time elements recorded in the declaration card set and cross-modal anchor table, and outputs a preliminary declaration judgment list.
[0053] The results output module verifies the readability and presentation consistency of the advertisement display based on the conclusions in the declaration preliminary judgment list and the cross-modal anchor point table, performs contextual legality detection, aggregates the detection results with the conclusions in the declaration preliminary judgment list to form disposal items, and outputs a disposal list and evidence presentation report.
[0054] The beneficial effects of this invention are as follows: by constructing a compliance ontology, a structured mapping is achieved between regulatory clauses, advertising standards, and detection logic, so that the judgment of advertising content has a unified semantic basis; by establishing field-level binding between the advertising subject, product characteristics, and time elements and authoritative external materials through evidence indexing, the minimum pairing verification of the statement content and factual evidence is achieved; and by synchronously associating text, images, and audio in the time and space dimensions through cross-modal anchor table, the readability and consistency of the display of statements and disclaimers are ensured. Attached Figure Description
[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0056] Figure 1 This is a flowchart of a method for detecting the compliance of APP advertising content.
[0057] Figure 2 This is a flowchart for text alignment.
[0058] Figure 3 This is a flowchart of the evidence matching process.
[0059] Figure 4 Output a flowchart of the results. Detailed Implementation
[0060] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0061] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0062] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0063] Reference Figures 1-4 As an embodiment of the present invention, this embodiment provides a method for detecting the compliance of APP advertising content, including the following steps:
[0064] S1. Construct a compliance ontology and evidence index, and generate a partitioned execution plan based on the timeline and screen coordinates of the advertising material, and output a list of task slices.
[0065] S1.1. Perform natural language processing analysis on the advertising display specification documents and industry standard description texts. Identify the constraints, limitations, and conditions that describe the restrictions on advertising display through syntactic structure decomposition and dependency relation analysis. Perform semantic parsing on the identified constraints, limitations, and conditions to extract the action words, limiting words, and modifying relationships, and transform them into logical nodes.
[0066] It should be noted that identifying the constraints, limitations, and conditions describing advertising display restrictions through syntactic structure decomposition and dependency relation analysis involves using part-of-speech tagging and dependency syntax tree analysis in natural language processing to determine the main predicate and modifiers of the sentence, and extracting words and phrases that are restrictive, limiting, and conditional. Keyword lists (containing words such as "must not," "prohibited," "should," "limited," "must," and "if") are matched with syntactic templates such as "[subject] + must not + [verb]" or "[conditional clause] + may + [verb]" to mark them as constraints, limitations, or conditions.
[0067] Among them, the constraint item refers to the component that describes the prohibition or restriction of advertising display behavior, such as sentences containing restrictive verbs or adverbs such as "must not", "prohibited", "strictly prohibited", and "should be avoided";
[0068] Limitations refer to elements that restrict or condition the scope of permitted advertising content, such as phrases like "limited to," "must be," "must not exceed," or "shall comply with."
[0069] Conditional items refer to components that describe the conditions under which an advertising display can be established, such as "when authorized", "after approval", "when the content is complete", etc.
[0070] Semantic parsing of the identified constraints, limitations, and conditions involves determining the agent, object, and constraint target of the action in the sentence through semantic role labeling; and identifying the relationship between modifiers (such as adverbs and prepositional phrases) and core predicates based on the dependency tree structure, thus binding the scope of constraints to the action.
[0071] S1.2. For each logical node, set trigger conditions and exemption conditions through logical rule matching. Group and organize the trigger conditions and exemption conditions of the logical nodes according to semantic consistency and store them as a rule node table. Based on the trigger conditions recorded in the rule node table, extract high-frequency trigger predicates using keyword extraction and template matching methods to generate a trigger word list. Map the entries in the rule node table to the trigger word list one by one, and generate a compliance ontology based on the mapping relationship. The compliance ontology contains a hierarchical structure of constraint type, trigger predicate, exemption conditions and display threshold.
[0072] It should be noted that the steps for setting trigger conditions and exemption conditions through logical rule matching are as follows:
[0073] Read the subject, behavior, and constraint triples from the logical nodes; match the behavior verbs with preset trigger predicate templates, such as "claim," "guarantee," "promise," "unique," "first," etc.; when the behavior in the logical node matches the trigger predicate template, the appearance of the predicate is defined as a trigger condition, meaning that if the corresponding expression appears in the advertising text, further evidence verification or disclaimer comparison is required; at the same time, extract exempt phrases from the condition items of the logical nodes, such as if the data source has been indicated, if a disclaimer appears in the image, or if the word "advertisement" is indicated, these phrases are exempt conditions.
[0074] Extracting high-frequency triggering predicates using keyword extraction and template matching involves calculating the frequency of each verb or phrase in all triggering condition statements using word frequency statistics and the TF-IDF algorithm; retaining verb phrases whose frequency exceeds a preset frequency threshold; comparing the retained results with predefined trigger templates, which include formats such as "claim + result", "guarantee + effect", "use + ready", "first-of-its-kind + technology", and "unique + product"; when a verb phrase completely matches or is semantically similar to a template, the corresponding phrase is identified as a high-frequency triggering predicate.
[0075] The preset frequency threshold is set based on the statistical results of the advertising corpus, and is set to 3~8. This value can achieve a balance between word frequency distribution and detection performance. When the frequency threshold is lower than 3, a large number of low-frequency noise phrases will be introduced. When the frequency threshold is higher than 8, some industry-specific trigger predicates will be missed.
[0076] The predefined trigger templates were obtained through statistical analysis and semantic summarization of the most frequent and typical violation risk expressions in advertising review cases, industry standards, and advertising corpora over the years.
[0077] S1.3. Based on the rule nodes in the compliance ontology, determine the subject categories that need to be verified in the advertising content, and extract the subject name, product features and time elements from the text, image and audio content of the advertising material; through named entity recognition and time phrase recognition, extract and classify the proper nouns and time expressions appearing in the text of the advertising material, and organize the extracted subject name, product features and time elements into subject layer, product layer and time layer fields respectively, forming a three-layer field structure;
[0078] Establish field mapping relationships for external materials that can be referenced in advertising materials, generate unique index codes for each field, bind the fields of the subject layer, product layer and time layer with the corresponding fields in the external materials, and build an evidence index library.
[0079] External information refers to relevant information related to the advertising entity, product, and claims, such as product testing reports and historical price databases.
[0080] S1.4. Read rule nodes from the compliance ontology, read index fields from the evidence index, and extract trigger vocabulary as the matching basis. Segment the advertising text according to the timeline of the advertising material to form text segments with time start and end markers. Divide the image content according to the screen coordinates of the advertising material to generate image grid blocks containing coordinate boundaries. Segment the audio signal according to the audio timeline of the advertising material to generate audio short segments with timestamps. Use timestamp alignment and feature frame matching to combine text segments, image grid blocks and audio short segments to form material slices.
[0081] The material slices are matched with the trigger word list. Keyword matching and context semantic comparison are performed on each material slice to identify material slices containing potential trigger words and determine their correspondence with rule nodes in the compliance ontology. For each material slice that matches a trigger word, the time range, spatial range, and trigger word number of the material slice are recorded, and the corresponding subject layer, product layer, and time layer field indexes are labeled. By integrating the labeling information of all material slices, a task slice list is generated. Each task slice carries potential trigger words, associated rule node identifiers, and fields of evidence to be verified.
[0082] S2. Using the trigger information carried in the task slice list, extract the declaration fragment and the corresponding disclaimer fragment from the text of the advertising material, and combine it with video frame coordinates, subtitle timestamps, and audio fragments to complete cross-modal anchor point alignment, generating a declaration card set and a cross-modal anchor point table. See details... Figure 2 .
[0083] S2.1. Read the time range, spatial range, trigger word number, and evidence field to be verified for each task slice in the task slice list. Perform word segmentation, part-of-speech tagging, and dependency analysis within the text range corresponding to the task slice. Perform pattern matching based on the trigger word list and syntactic template (trigger predicate + object + qualifier). Mark the sentence segments containing the trigger predicate as candidate declaration segments. Record the start and end positions of the characters of the candidate declaration segments, the task slice number to which they belong, the trigger word number, and the subtitle timestamp range to form a candidate declaration list.
[0084] S2.2. Identify the claim content, object, numerical components, and qualifiers in the sentence segments corresponding to the candidate claim list. Use subject-predicate, verb-object, modifier-head, and adverbial relationships in dependency relations to bind the object and qualifiers to the trigger predicate. Extract the referents of the subject name, product information, and time element and label them accordingly with the field indexes of the subject layer, product layer, and time layer. For sentence segments containing numerical or range expressions, supplement the record with numerical units and comparison words. Output a list of claim fragments. Each claim fragment contains text content, trigger word number, object and qualifier elements, subject layer / product layer / time layer field indexes, and subtitle timestamp range.
[0085] S2.3. Within the text window located in the same task slice as the disclaimer fragment, keyword matching and negative constraint recognition methods based on natural language processing are used to retrieve disclaimer keywords and fixed expressions (such as "for reference only", "subject to actual product", "subject to terms and conditions" type expressions). Small windows overlapping or adjacent to the disclaimer fragment are set up on the subtitle timeline for supplementary retrieval, and the start and end positions of the characters of the disclaimer fragment and the subtitle timestamp interval are recorded. In accordance with the order of priority of time overlap, priority of the same sentence, and priority of proximity of the same segment, one or more candidate associations of disclaimer fragments are determined for each disclaimer fragment, forming a disclaimer-disclaimer candidate correspondence table.
[0086] S2.4. Select the corresponding video frame within the subtitle timestamp interval, use optical character recognition to identify the text box and its coordinates within the video frame, perform string matching or edit distance matching between the text box content and the declaration and disclaimer segments to obtain the correspondence between the text box and the segment and generate the video frame coordinates; locate the audio segment based on the audio timeline within the same interval and confirm whether the spoken content covers the declaration or disclaimer segment through timestamp alignment; use the subtitle timestamp interval as the time anchor point and the video frame coordinates as the spatial anchor point, complete the time anchor point and spatial anchor point fields for each declaration-disclaimer candidate association, and output the aligned declaration-disclaimer association list.
[0087] S2.5. Based on the overlap ratio of time anchor points and the intra-frame positional relationship of spatial anchor points, adjudicate each group of claim-disclaimer candidate associations, prioritize the associations that are completely overlapping in time and adjacent to the screen area as the main association, and register other associations that meet the time overlap but cross spatial regions as candidate associations; record the text height, foreground and background contrast, and the proportion of possible occlusion area and other readability-related elements on the video frame of the main association to form a visibility record bound to the main association.
[0088] S2.6. Based on each declaration fragment, summarize the trigger word number, object and qualifier elements, subject layer / product layer / time layer field index, pointer to the main associated disclaimer fragment, time anchor point, video frame coordinates and visibility record to generate a declaration card; combine all declaration cards into a declaration card set, and establish a cross-modal anchor point table with time anchor point and video frame coordinates as the core fields. In the cross-modal anchor point table, store the subtitle timestamp range, video frame coordinates and corresponding audio segment identifier for each declaration card.
[0089] S3. Based on the subject name, product information, and time elements recorded in the declaration card set and the cross-modal anchor table, perform evidence retrieval and minimum evidence matching, and output a preliminary declaration judgment list. See details below. Figure 3 .
[0090] S3.1. Read the trigger word number, subject name, product information, time element, qualifier, and declaration object fields recorded in the declaration card set, and extract the corresponding time anchor points and video frame coordinates from the cross-modal anchor point table; generate an evidence retrieval plan based on the subject layer, product layer, and time layer field indexes in the declaration card. The evidence retrieval plan uses the subject name, product information, and time element as the main search keys, and combines the qualifier and trigger word number to determine the search scope, clearly defining the query target as authoritative data sources in external materials, including regulatory approvals, product licenses, price records, promotional announcements, and authorization documents, etc. The generated evidence retrieval plan uses the declaration card number as the index item, and each search item carries the corresponding query field, time range, and evidence type requirements, outputting the evidence retrieval plan table.
[0091] S3.2. Based on the subject, product, and time layers in the evidence retrieval plan table, perform field mapping matching in the evidence index library; quickly locate the set of external data records corresponding to the declaration card using unique index codes; for cases where multiple data versions exist for the same subject or product, prioritize the record closest to the declaration time anchor point based on the time element field to limit the record time interval; if the external data is in text format, extract the document title, publication date, and content summary; if the external data is in a structured database, extract the corresponding field values, such as license number, price, approval time, and validity period, establish a correspondence between the retrieved set of external data records and the declaration card number, and output the evidence matching list.
[0092] S3.3. In the evidence matching list, based on the key points of the statement recorded in the statement card (including objects, action verbs, numerical expressions and qualifiers), perform field-level comparison of the external data record set; use text comparison and keyword matching to filter out the smallest evidence fragments that can directly cover the key points of the statement; when a statement has multiple sources of evidence, prioritize the combination that meets the requirements of coverage and integrity and has the simplest information to ensure that the number of evidence fragments is the smallest but can fully support the conclusion; record the filtering results as the minimum necessary evidence fragments, and establish evidence pointers to identify the storage location of external data records in the evidence index library, and output the minimum evidence fragment table.
[0093] To further explain, using the declaration card as an index, the system extracts the minimum amount of external data records that can cover the key points of the declaration from the evidence index library; establishes a correspondence between declaration and evidence pointers, generates conclusions that support, refute, or indicate insufficient evidence, significantly reduces the amount of comparison required for manual review, achieves automated evidence summarization, ensures the legal traceability and objectivity of the conclusions, and adaptively tailors the evidence complexity of different advertising content, thereby improving the system's computational efficiency and versatility.
[0094] S3.4. Based on the minimum evidence fragment table, perform a consistency analysis on the key points of each statement card and the content of the evidence. If the factual description in the evidence fragment is completely consistent with the content of the statement, or if the evidence fields match the time, subject, and product information involved in the statement, it is determined to be supported. If the factual description in the evidence fragment contradicts or contradicts the content of the statement, it is determined to be refuted. If no valid evidence corresponding to the statement can be found in the evidence index or external materials, it is determined to be insufficient evidence. Add a conclusion mark to the judgment result of each statement, and record the corresponding evidence pointer and the explanation of the uncovered key points to form a statement judgment record table.
[0095] S3.5. Summarize all declaration judgment record sheets according to declaration card number, integrate the declaration content, evidence pointers, judgment conclusions and explanations of unmet evidence gaps to form a preliminary declaration judgment list; each record in the preliminary declaration judgment list includes the declaration text, time anchor point, video frame coordinates, judgment conclusion (support, refute or insufficient evidence), minimum necessary evidence fragments and corresponding external data source information.
[0096] S4. Based on the conclusions in the preliminary judgment list and the cross-modal anchor table, verify the readability and presentation consistency of the advertisement display, perform contextual legality checks, aggregate the check results with the conclusions in the preliminary judgment list to form disposal items, and output a disposal list and evidence presentation report. See details in [link to relevant documentation]. Figure 4 .
[0097] S4.1. Read the statement text, judgment conclusion, evidence pointer, and minimum necessary evidence fragment from the initial judgment list, and obtain the corresponding time anchor and video frame coordinates from the cross-modal anchor table; at the same time, retrieve the visibility record and disclaimer fragment pointer bound to the main association of statement and disclaimer to form a set of items to be verified for presentation.
[0098] The system obtains the duration of the statement and disclaimer text on the screen based on the time anchor point, obtains the font height based on the video frame coordinates and the pixel height of the text box obtained by optical character recognition, and determines the foreground and background color contrast and the proportion of possible occlusion area based on the visibility record. The system compares the duration of the statement, font height, color contrast and occlusion proportion with the display threshold item by item, records the compliance status and numerical details of each indicator, and outputs the readability verification results.
[0099] It should be noted that determining the foreground and background color contrast based on visibility records refers to judging the degree of visual distinction between text and background by reading the average brightness and hue differences between the text box area and the background area in the video frame. When the brightness difference between the text area and the background area is significant, and the text edges are clearly distinguishable in grayscale or chromaticity distribution, the color contrast is considered to meet the readability requirements.
[0100] The occlusion ratio is determined by comparing the pixel range of the declaration text box in the video frame with the pixel range of the overlapping part of the image element (such as icon, button or animation) appearing in the same frame. When the main character shape of the text area is fully displayed and the overlapping area is small, the occlusion ratio is considered to meet the visibility requirements.
[0101] The display threshold is set based on the advertising review experience and visual research of mainstream video platforms. The range of values takes into account the differences in human eye recognition ability and screen resolution environment. It is usually set as follows: the font height is not less than 2% of the screen height; the difference in brightness between the foreground and background is not less than 40%; the dwell time is not less than 1.0 second; and the occlusion area ratio does not exceed 20% of the total area of the text area.
[0102] When the font height is less than 2%, the disclaimer text will be difficult to read on small screen devices; when the color difference is less than 40%, the text will visually confuse with the background, making it impossible for users to read the content accurately; when the dwell time is less than 1 second, the human eye cannot complete the reading; when the obscured area exceeds 20%, the disclaimer content will be partially missing, affecting the integrity of the information.
[0103] S4.2. Determine whether the disclaimer is displayed in the same time period as the statement based on the time anchor point, determine whether the statement and the disclaimer satisfy the spatial proximity relationship in the same frame based on the video frame coordinates, and check the semantic and temporal synchronization consistency between the subtitle text and the spoken content based on the subtitle timestamp and the audio timeline; combine the judgment results of time overlap, spatial proximity relationship and voice-subtitle synchronization with the presentation readability verification results to form the presentation verification result.
[0104] It should be noted that spatial proximity refers to the closeness of the positions of the declaration text box and the disclaimer text box within the same video frame. By comparing the intra-frame coordinate regions of the two text boxes, a spatial proximity relationship is determined when they are within the same area of the screen.
[0105] S4.3. Based on the presentation verification results, locate the context window of each statement, identify the occurrence of extreme words and sensitive words, and verify whether the disclaimer fragment pointer exists and whether the presentation verification result is qualified. At the same time, verify whether the evidence pointer in the statement preliminary judgment list exists. When the disclaimer is missing, or the presentation verification result is unqualified, or the evidence pointer is missing, mark it as unqualified. When the disclaimer exists, the presentation verification result is qualified, and the evidence pointer exists, mark it as legal. Summarize the unqualified mark and the legalization mark into the context legalization result.
[0106] It should be noted that the criteria for passing the verification are as follows: the disclaimer and the statement are displayed in the same time period and the duration of each is not less than the display threshold; the disclaimer and the statement are in the same screen area and the distance between their center points is within a perceptible range; the font height, color contrast, and occlusion ratio all meet the display threshold requirements; and the subtitle text and the audio content are semantically consistent and time-synchronized, without any deviation or omission.
[0107] S4.4. Establish a correspondence between the contextual legalization result, the presentation verification result, and the judgment conclusion in the declaration preliminary judgment list: when the presentation verification result is unqualified or the judgment conclusion is refuted, the disposal type is blocked; when the judgment conclusion is insufficient evidence and the presentation verification result is qualified, the disposal type is reviewed; when the judgment conclusion is supportive, the presentation verification result is qualified, and the contextual legalization result is legal, the disposal type is released; organize the disposal type, time anchor point, video frame coordinates, declaration text, disclaimer fragment pointer, and evidence pointer into disposal items, and summarize them into a disposal list according to the declaration card number; generate an evidence presentation report using the disposal list as an index.
[0108] This embodiment also provides a system for detecting the compliance of APP advertising content, including:
[0109] The rule partitioning module constructs a compliance ontology and evidence index, and generates a partition execution plan based on the timeline and screen coordinates of the advertising material, outputting a list of task slices;
[0110] The text alignment module uses the trigger information carried in the task slice list to extract the declaration fragment and the corresponding disclaimer fragment from the advertisement text, and combines the video frame coordinates, subtitle timestamps and audio fragments to complete the cross-modal anchor point alignment, generating a declaration card set and a cross-modal anchor point table;
[0111] The evidence matching module performs evidence retrieval and minimum evidence matching based on the subject name, product information and time elements recorded in the declaration card set and cross-modal anchor table, and outputs a preliminary declaration judgment list.
[0112] The results output module verifies the readability and presentation consistency of the advertisement display based on the conclusions in the declaration preliminary judgment list and the cross-modal anchor point table, performs contextual legality detection, aggregates the detection results with the conclusions in the declaration preliminary judgment list to form disposal items, and outputs a disposal list and evidence presentation report.
[0113] This embodiment also provides a computer device applicable to the method for detecting the compliance of APP advertising content, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method for detecting the compliance of APP advertising content as proposed in the above embodiment.
[0114] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0115] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the method for detecting the compliance of APP advertising content as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0116] In summary, this invention achieves a structured mapping between regulatory provisions, advertising standards, and detection logic by constructing a compliance ontology, thus providing a unified semantic basis for judging advertising content; it establishes field-level binding between the advertising subject, product characteristics, and time elements and authoritative external materials through evidence indexing, achieving minimum pairing verification between the statement content and factual evidence; and it ensures the readability and consistency of the display of statements and disclaimers by synchronously associating text, images, and audio in the time and space dimensions through a cross-modal anchor table.
[0117] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for detecting compliance of APP advertising content, characterized in that: include, Construct a compliance ontology and evidence index, and generate a partitioned execution plan based on the timeline and screen coordinates of the advertising material, and output a list of task slices; Using the trigger information carried in the task slice list, the declaration fragment and the corresponding disclaimer fragment are extracted from the text of the advertising material, and cross-modal anchor point alignment is completed by combining the video frame coordinates, subtitle timestamps and audio fragments to generate a declaration card set and a cross-modal anchor point table; Based on the subject name, product information and time elements recorded in the declaration card set and cross-modal anchor table, evidence retrieval is performed and minimum evidence matching is performed to output a preliminary declaration judgment list; Based on the conclusions in the preliminary judgment list and the cross-modal anchor table, the readability and presentation consistency of the advertisement display are verified, and contextual legality detection is performed. The detection results are aggregated with the conclusions in the preliminary judgment list to form disposal items, and a disposal list and evidence presentation report are output.
2. The method of claim 1, wherein: The specific steps for constructing the compliance ontology are as follows. The text of advertising display specifications and industry standards is parsed to extract constraints, limitations, and conditions. Extract constraints, limitations, and conditions and convert them into logical nodes; Set corresponding trigger conditions and exemption conditions for each logical node and store them as a rule node table; Generate a trigger word list based on the trigger conditions in the rule node table; The entries in the rule node table are mapped to the trigger vocabulary to generate a compliance ontology.
3. The method of claim 2, wherein: The evidence index is constructed using the following specific steps. The categories of entities to be verified are determined based on the rule nodes in the compliance ontology, and the entity name, product characteristics, and time elements are extracted from the advertising materials. The extracted subject name, product characteristics, and time elements are organized into a three-layer field structure: subject layer, product layer, and time layer. Establish field mapping relationships for external data and generate unique index codes, bind each layer of fields to the corresponding fields in the external data, and generate an evidence index library.
4. The method for detecting the compliance of APP advertising content as described in claim 3, characterized in that: The process of generating a partitioned execution plan based on the timeline and screen coordinates of the advertising creative, and outputting a task slice list, includes the following steps. Read rule nodes and index fields from the compliance ontology and evidence index; The text is segmented according to the timeline of the advertising material, and image grid blocks and audio segments are divided according to the screen coordinates to generate material slices; Each material slice is matched with the trigger word list to identify potential trigger words and entries to be verified, and the time range, spatial range and trigger word number of the slice are recorded to form a task slice list.
5. The method of claim 4, wherein: The generation of the declaration card set and cross-modal anchor table includes the following steps: Based on the task slice list, identify the claim content, object and qualifier in the corresponding text range, generate the claim fragment, and identify the disclaimer statement in the same interval to generate the disclaimer fragment; Match the text position with the video frame coordinates and register the subtitle timestamp with the audio time segment to establish a synchronous association between text, image and audio; The synchronization results are recorded as time anchors and spatial anchors, and a declaration card is generated based on the correspondence between the declaration fragment and the disclaimer fragment; The declaration cards and their anchor information are organized into a declaration card set and a cross-modal anchor table.
6. The method for detecting the compliance of APP advertising content as described in claim 5, characterized in that: The process involves performing evidence retrieval and minimum evidence matching, and outputting a preliminary judgment list. Includes the following steps, Read the subject name, product information, and time element from the declaration card set and cross-modal anchor table; Use the evidence index to locate external data records and filter out the minimum record fragments that can cover the key points of the statement; Establish a one-to-one correspondence between the recorded fragments and the statement cards to generate evidence pointers; Based on the evidence pointers, conclusions of support, refutation, and insufficient evidence are generated for each statement and compiled into a preliminary list of statements.
7. The method of claim 6, wherein: The verification of the readability and consistency of the advertisement display includes the following steps: Based on the preliminary list of statements and the cross-modal anchor point table, determine the range in which each statement and its disclaimer appear in the image; Check whether the disclaimer is displayed at the same time as the statement, and verify whether the display duration, font height, color contrast and occlusion ratio meet the display threshold requirements; Verify the semantic and temporal synchronization consistency between the subtitle text and the audio content, and summarize the verification results.
8. The method of claim 7, wherein: The execution context legitimacy detection includes the following steps: Read the validation results, locate the context window of each statement, identify the occurrence of extreme words and sensitive words, and determine whether the corresponding disclaimer fragment exists and whether the validation result is qualified. Determine whether the corresponding evidence pointers in the preliminary judgment list exist; The document is marked as unqualified when a disclaimer is missing, the presentation is inadequate, or one of the evidentiary points is missing; it is marked as valid when a disclaimer is present, the presentation is qualified, and the evidentiary points are present. The invalid and valid tags are combined into a contextual validation result.
9. The method of claim 8, wherein: The process involves aggregating the test results with the conclusions of the preliminary judgment list to form a disposal item. Includes the following steps, The contextual legalization results are merged with the presentation verification results, and a correspondence is established with the conclusions in the initial judgment list of the declarations. The disposal type of each declaration is determined to be one of blocking, review, and release, and the corresponding disposal description is generated. Organize the disposal instructions, time anchors, spatial anchors, and evidence pointers into disposal items; The disposal items are compiled into a disposal list, and an evidence presentation report is generated using the disposal list as an index.
10. A system for detecting compliance of APP advertising content, based on the method for detecting compliance of APP advertising content according to any one of claims 1-9, characterized in that: include, The rule partitioning module constructs a compliance ontology and evidence index, and generates a partition execution plan based on the timeline and screen coordinates of the advertising material, outputting a list of task slices; The text alignment module uses the trigger information carried in the task slice list to extract the declaration fragment and the corresponding disclaimer fragment from the advertisement text, and combines the video frame coordinates, subtitle timestamps and audio fragments to complete cross-modal anchor point alignment, generating a declaration card set and a cross-modal anchor point table; The evidence matching module performs evidence retrieval and minimum evidence matching based on the subject name, product information and time elements recorded in the declaration card set and cross-modal anchor table, and outputs a preliminary declaration judgment list. The results output module verifies the readability and presentation consistency of the advertisement display based on the conclusions in the declaration preliminary judgment list and the cross-modal anchor point table, performs contextual legality detection, aggregates the detection results with the conclusions in the declaration preliminary judgment list to form disposal items, and outputs a disposal list and evidence presentation report.
Citation Information
Patent Citations
Advertisement management and distribution system and method based on Internet of Things
CN118608209A
Advertisement video badness-oriented fine-grained detection method based on knowledge graph and large language model
CN120451872A