Intelligent advertisement language auditing system
By designing an audit system for intelligent advertising language, the problem of difficulty in accurately judging advertising emotional tendencies and misleading publicity in the existing technology is solved, and multi-dimensional review of advertising content and efficient and accurate review results are achieved.
Patent Information
- Application Number
- CN202510324794.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-13
AI Technical Summary
The existing advertising review system is difficult to accurately judge the emotional tendencies of the advertisement, whether there are misleading publicity or false promises, and the handling of sensitive topics is inappropriate, resulting in inconsistent and inefficient review results.
An audit system for intelligent advertising language is designed, including data reception module, preprocessing module, language sentiment analysis module, violation vocabulary detection module, semantic analysis module, manual audit interface module and report generation module. Through the collaborative work of these modules, the system can automatically receive, pre-process, sentiment analysis, violation vocabulary detection, semantic analysis and report generation of advertising content.
It realizes multi-dimensional review of advertising content, and can more accurately judge the emotional tendencies of the advertisement and whether there are misleading publicity or false commitments, improves the review efficiency and accuracy, and ensures the consistency and quality of the review results.
Smart Images

Figure CN120146927A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of advertising technology, and in particular to an intelligent advertising language review system. Background Art
[0002] With the rapid development of Internet and mobile communication technologies, the advertising industry has witnessed unprecedented prosperity. Whether it is social media, search engines or online video platforms, advertising has become a bridge of communication between businesses and consumers and an essential part of brand promotion and product marketing. However, with the sharp increase in the number of advertisements, how to ensure the legality, compliance and morality of advertising content and maintain a healthy online environment has become a huge challenge.
[0003] Traditional advertising review methods highly rely on manual labor. Reviewers need to read each advertisement content item by item and make judgments according to relevant laws and regulations, platform rules and social moral standards. However, this method is not only time-consuming and laborious, but also easily affected by personal subjective factors, resulting in inconsistent and inefficient review results. With the rise of big data and artificial intelligence technologies, advertising review has started to develop towards automation and intelligence.
[0004] In the prior art, some advertising review systems have begun to attempt to use automated means to improve review efficiency. For example, by setting a list of keywords or phrases, simple text matching is performed on the advertising content to identify potential violations. However, this method has obvious limitations. On the one hand, it can only identify advertisements containing specific words or phrases and is powerless against violation advertisements that avoid keyword detection by changing the wording, using metaphors or rhetoric, etc. On the other hand, existing automated review systems often lack in-depth language analysis and semantic understanding capabilities and are difficult to accurately judge complex issues such as the emotional tendency of advertisements, whether there is misleading publicity or false promises. In addition, for advertisements involving sensitive topics such as politics, religion, race, etc., existing systems often cannot handle them appropriately, easily leading to disputes and misunderstandings. Summary of the Invention
[0005] By providing an intelligent advertising language review system, the present application solves the technical problem in the prior art that it can only perform simple text matching on advertising content and is difficult to accurately judge the emotional tendency of advertisements, whether there is misleading publicity or false promises; and achieves the technical effect of not only performing text matching on advertising content, but also being able to judge the emotional tendency of advertisements, whether there is misleading publicity or false promises.
[0006] The present application provides a review system for intelligent advertising language, including a data receiving module, a preprocessing module, a language sentiment analysis module, a violation vocabulary detection module, a semantic analysis module, a manual review interface module, and a report generation module. The data receiving module is docked with an advertising release platform or an advertising management system to receive advertising content data to be reviewed. The preprocessing module preprocesses the received advertising data. The language sentiment analysis module analyzes the sentiment tendency in the advertising text. The violation vocabulary detection module detects whether the advertising text contains violation vocabulary. The semantic analysis module deeply analyzes the semantic content of the advertising text to determine whether there is misleading publicity or false promise. The manual review interface module provides an interface for manual reviewers for manual review. The report generation module generates a final review report.
[0007] Further, the data receiving module includes an interface component, a data parsing component, a data verification component, and a data storage component. The interface component is responsible for docking with an external advertising release platform or an advertising management system. The data parsing component parses the received data and converts it into a format recognizable within the system. The data verification component verifies the parsed data to ensure the integrity and accuracy of the data. The data storage component stores the verified data in the internal database of the system for subsequent processing.
[0008] Further, the preprocessing module includes a text cleaning component, a word segmentation component, a stop word filtering component, and a text standardization component. The text cleaning component is used to remove irrelevant characters, special symbols, HTML tags, and extra spaces in the text. The word segmentation component is used to cut continuous text strings into meaningful words, and this component includes a dictionary and a word segmentation algorithm. The stop word filtering component includes a list of stop words. After word segmentation, the stop word filtering component traverses each word in the text. If the word exists in the stop word list, it is removed from the text to reduce data sparsity and improve the efficiency of subsequent analysis. The text standardization component is used to convert the text into a unified format or standard.
[0009] Further, the language sentiment analysis module includes a feature extraction component and a machine learning component. The feature extraction component extracts features from the preprocessed text and converts them into numerical vectors. The machine learning component uses machine learning algorithms to train a sentiment classifier. Among them, the model is trained through sentiment text data, and the model learns the mapping relationship from text features to sentiment labels. In the prediction stage, the model receives the preprocessed text feature vector and outputs a sentiment classification result.
[0010] Further, the machine learning component uses machine learning algorithms to train a sentiment classifier, and the machine learning algorithm is at least any one of a support vector machine, a naive Bayes, and a deep learning model.
[0011] Further, the violation vocabulary detection module includes a violation vocabulary library, a vocabulary matching algorithm component, and a whitelist management component; the violation vocabulary library stores a list of violation vocabularies internally, and these violation vocabularies include at least any one of insulting language, sensitive political vocabulary, and illegal information keywords; the vocabulary matching algorithm component is used to achieve efficient matching between the text and the violation vocabulary library to identify the violation vocabularies contained in the text; the whitelist management component stores a whitelist list internally, and for the vocabularies matched to the whitelist, even if they appear in the violation vocabulary library, they will be regarded as compliant content.
[0012] Further, the semantic analysis module includes a rule library and a logical judgment component; the rule library contains a series of rules for specific problems of advertising texts; the logical judgment component makes logical judgments on the vector representation or extracted features output by the semantic understanding component according to the rules in the rule library to determine whether the text conforms to the problems defined by the rules.
[0013] Further, the manual review interface module is respectively connected to the language emotion analysis module, the violation vocabulary detection module, and the semantic analysis module. When the triggering conditions for manual review are met, relevant personnel conduct manual review through the manual review interface module.
[0014] One or more technical solutions provided in this application have at least the following technical effects or advantages:
[0015] The data reception module docks with the advertising platform, stores it after parsing and verification; the preprocessing module cleans the text, performs word segmentation, removes stop words, and standardizes it. The emotion analysis and violation vocabulary detection modules respectively identify the emotion tendency and violation content through machine learning algorithms and matching with the violation vocabulary library. The semantic analysis module judges misleading publicity, etc. according to the rule library. The manual review interface module triggers a manual re-review under specific conditions to ensure the review quality; effectively solves the technical problem in the prior art that only simple text matching can be performed on the advertising content, and it is difficult to accurately judge the emotion tendency of the advertisement, whether there is misleading publicity or false promise; and then realizes the technical effect of not only performing text matching on the advertising content, but also being able to judge the emotion tendency of the advertisement, whether there is misleading publicity or false promise. Brief Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the overall structure of the intelligent advertisement language review system of the present invention;
[0017] Figure 2 It is a schematic diagram of the structure of the data reception module of the intelligent advertisement language review system of the present invention;
[0018] Figure 3 It is a schematic diagram of the structure of the preprocessing module of the intelligent advertisement language review system of the present invention;
[0019] Figure 4 It is a schematic structural diagram of the illegal vocabulary detection module of the intelligent advertising language review system of the present invention. Specific embodiments
[0020] To facilitate the understanding of the present invention, the present application will be described more comprehensively with reference to the relevant accompanying drawings; the preferred embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein; on the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0021] It should be noted that the terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only embodiments.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs; the terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention; the term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0023] Example: As Figures 1 to 4 shown, the intelligent advertising language review system of the present application includes a data receiving module, a preprocessing module, a language sentiment analysis module, an illegal vocabulary detection module, a semantic analysis module, an artificial review interface module and a report generation module.
[0024] Among them, the data receiving module is docked with the advertising publishing platform or the advertising management system to receive the advertising content data to be reviewed.
[0025] The preprocessing module preprocesses the received advertising data, such as text tokenization, stop word removal, etc.
[0026] The language sentiment analysis module analyzes the sentiment tendency in the advertising text, such as positive, neutral, negative.
[0027] The illegal vocabulary detection module detects whether the advertising text contains illegal vocabulary, such as porn, violence, fraud, etc.
[0028] The semantic analysis module deeply analyzes the semantic content of the advertising text to determine whether there is misleading publicity, false promise, etc.
[0029] The artificial review interface module provides an interface for artificial reviewers for manual review when necessary.
[0030] The report generation module generates a final review report according to the results of automatic review and manual review.
[0031] Furthermore, the data receiving module includes an interface component, a data parsing component, a data verification component, and a data storage component.
[0032] The interface component is responsible for docking with an external advertising platform or advertising management system; specifically, the interface component can adopt the form of an API (Application Programming Interface) or an SDK (Software Development Kit) to receive data in a standardized manner.
[0033] The data parsing component parses the received data and converts it into a format recognizable within the system; specifically, according to the format of the data (such as JSON, XML, etc.), the corresponding parser is used for parsing.
[0034] The data verification component verifies the parsed data to ensure the integrity and accuracy of the data; specifically, by setting data verification rules (such as required fields, data types, data ranges, etc.), the received data is verified one by one.
[0035] The data storage component stores the verified data in the internal database of the system for subsequent processing; specifically, a relational database (such as MySQL, PostgreSQL, etc.) or a non-relational database (such as MongoDB, Redis, etc.) is used for data storage.
[0036] It can be understood that when the advertising platform or advertising management system has advertising content data that needs to be reviewed, the data is sent to the data receiving module through a preset interface. The interface component of the data receiving module is responsible for receiving this data and passing it to the data parsing component; the data parsing component parses the received data and converts it from an external format to a format recognizable within the system. During the parsing process, the data parsing component will select an appropriate parser according to the format of the data for parsing; the data verification component verifies the parsed data to ensure the integrity and accuracy of the data; during the verification process, the data verification component will verify the received data one by one according to the preset verification rules; if the data does not conform to the verification rules, the data receiving module will return an error message to the advertising platform or advertising management system and request to resend the data; the data that has passed the verification will be passed to the data storage component; the data storage component will store the data in the internal database of the system for subsequent processing; during the storage process, the data storage component will select an appropriate storage method according to the type and characteristics of the data for storage.
[0037] Furthermore, the preprocessing module includes a text cleaning component, a word segmentation component, a stop word filtering component, and a text standardization component.
[0038] The text cleaning component is used to remove irrelevant characters, special symbols, HTML tags, extra spaces, etc. from the text to ensure the purity of the text content; specifically, the text cleaning component uses regular expression matching and replacement technology, or specialized libraries (such as BeautifulSoup in Python for HTML tag removal) to identify and remove impurities in the text.
[0039] The word segmentation component is used to cut continuous text strings into meaningful words, where the component includes a dictionary and a word segmentation algorithm; specifically, rule-based word segmentation (such as maximum positive match, minimum segmentation, etc.) or statistical-based word segmentation methods (such as hidden Markov model, conditional random field, etc.), combined with a predefined dictionary, cuts the text into words or phrases.
[0040] It should be noted that for English, word segmentation is not necessary, but word form restoration (such as restoring "running" to "run") and stemming (extracting the basic form of a word) may be involved.
[0041] The stop word filter component includes a stop word list; after word segmentation, the stop word filter component traverses each word in the text, and if the word exists in the stop word list, it is removed from the text to reduce data sparsity and improve the efficiency of subsequent analysis.
[0042] It should be noted that stop words are words that appear frequently in a language but do not contribute much to the meaning of the text, such as "的", "了", "in", "the", etc.
[0043] The text standardization component is used to convert text into a unified format or standard, such as converting all letters to lowercase, unifying the number format, processing abbreviations, etc.; specifically, through string operation functions or regular expressions, the text is formatted in a unified manner to ensure that the text has a consistent representation in different contexts, which is convenient for comparison and analysis in subsequent modules.
[0044] It can be understood that when data is passed from the data receiving module to the preprocessing module, it is firstly removed from impurities by the text cleaning component, followed by word segmentation, and then the stop word filtering component is used to remove meaningless words, and the text standardization component ensures consistent format.
[0045] Furthermore, the language sentiment analysis module includes a feature extraction component and a machine learning component.
[0046] The feature extraction component extracts features from the preprocessed text and converts them into numerical vectors, such as the bag-of-words model, TF-IDF, word embedding (Word2Vec, BERT, etc.).
[0047] Train an emotion classifier using machine learning algorithms (such as support vector machines, Naive Bayes, deep learning models, etc.); specifically, train the model through a large number of labeled emotion text data, and the model learns the mapping relationship from text features to emotion labels. In the prediction stage, the model receives the preprocessed text feature vectors and outputs the emotion classification results.
[0048] Furthermore, the illegal vocabulary detection module includes an illegal vocabulary library, a vocabulary matching algorithm component, and a whitelist management component.
[0049] The illegal vocabulary library stores a list of illegal vocabulary internally, and these illegal vocabulary may include insulting language, sensitive political vocabulary, keywords of illegal information, etc.
[0050] It should be noted that the illegal vocabulary library can be updated dynamically, adding newly discovered illegal vocabulary through manual review, or keeping up-to-date by synchronizing with a third-party database.
[0051] The vocabulary matching algorithm component is used to achieve efficient matching between the text and the illegal vocabulary library, and identify the illegal vocabulary contained in the text; specifically, various strategies such as exact matching, fuzzy matching (such as the Levenshtein distance algorithm), regular expression matching, etc. can be adopted; at the same time, in order to improve efficiency, data structures such as hash tables and Trie trees can be used to optimize the search speed.
[0052] The whitelist management component stores a whitelist list internally. For the vocabulary that matches the whitelist, even if they appear in the illegal vocabulary library, they will be regarded as compliant content.
[0053] Furthermore, the semantic analysis module includes a rule library and a logical judgment component.
[0054] The rule library contains a series of rules for specific problems of advertising texts, such as identification rules for misleading publicity, detection rules for false promises, etc., and these rules are defined based on keywords, phrase patterns, and semantic similarity thresholds.
[0055] The logical judgment component makes logical judgments on the vector representation or extracted features output by the semantic understanding component according to the rules in the rule library to determine whether the text conforms to the problems defined by the rules.
[0056] Specifically, a series of judgment criteria are defined through the rule library, and these criteria can be hard-coded logical rules or threshold judgments based on the output of machine learning models; the logical judgment engine matches the extracted features or vector representations with the rules in the rule library and judges whether there are problems with the text according to the matching results.
[0057] Furthermore, the manual review interface module is respectively connected to the language sentiment analysis module, the illegal vocabulary detection module, and the semantic analysis module. When the triggering conditions for manual review are met, relevant personnel conduct manual review through the manual review interface module.
[0058] The triggering conditions for manual review may include:
[0059] a. The language sentiment analysis module is unable to determine the sentiment tendency of the text, or the sentiment tendency is in a critical state.
[0060] b. The illegal vocabulary detection module detects ambiguously matched words and is unable to determine whether they are truly illegal.
[0061] c. The semantic analysis module is unable to accurately determine whether there is misleading publicity or false promise in the text.
[0062] d. The text involves sensitive topics such as politics, religion, race, etc.
[0063] e. The advertiser raises objections to the automatic review result and requests a manual re-review.
[0064] f. The automatic review system is unable to work properly during the upgrade or maintenance period.
[0065] It should be noted that, in order to ensure the review quality, the system randomly selects a certain proportion of advertisement content for manual review regularly.
[0066] Specifically, during the process of manual review, the system automatically marks the advertisement content that needs to be re-reviewed according to the preset triggering conditions and pushes it to the manual review queue; the system automatically assigns review tasks to the corresponding manual reviewers according to the qualifications, experience, and current workload of the reviewers; the reviewers log in to the manual review interface module to view the list of review tasks assigned to themselves; the reviewers click on the tasks to view the detailed information of the advertisement content, including the results and suggestions of the automatic review; the reviewers make judgments based on the advertisement content, the automatic review results, and their own professional knowledge; the reviewers can mark the advertisement content as "passed", "not passed", or "needs further investigation"; the reviewers need to record their review opinions and reasons in the system; after the reviewers complete the judgment, they submit the review results to the system.
[0067] The technical solutions in the embodiments of the present application at least have the following technical effects or advantages:
[0068] 1. Through the collaborative work of multiple modules, the automatic reception, preprocessing, sentiment analysis, illegal vocabulary detection, semantic analysis, and report generation of advertisement content are realized, greatly improving the review efficiency and reducing the labor cost.
[0069] 2. The use of machine learning algorithms for sentiment analysis and semantic understanding improves the accuracy of analysis;
[0070] 3. Preprocessing steps such as text cleaning, tokenization, stop word filtering, and text standardization effectively improve the quality of text data;
[0071] 4. The system audits from multiple dimensions such as sentiment tendency, violation words, and semantic content, and can more comprehensively evaluate the compliance of advertising content;
[0072] 5. Based on automatic auditing, a manual auditing interface is provided, enabling manual review in critical or uncertain situations to ensure the accuracy of the auditing results;
[0073] 6. The use of efficient data structures and algorithms for violation word matching improves the detection speed and reduces the false alarm rate;
[0074] 7. The dynamically updated violation word library and whitelist management enable the system to adapt to changing auditing requirements;
[0075] 8. The setting of the rule library enables the system to perform customized analysis according to specific advertising auditing requirements, improving the pertinence of auditing.
[0076] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent advertising language review system, characterized in that: It includes data receiving module, preprocessing module, language sentiment analysis module, illegal vocabulary detection module, semantic analysis module, manual review interface module and report generation module; The data receiving module is connected to the advertisement publishing platform or the advertisement management system to receive the advertisement content data to be reviewed; The preprocessing module preprocesses the received advertisement data; The language sentiment analysis module analyzes the emotional tendency in the advertising text; The illegal words detection module detects whether the advertisement text contains illegal words; The semantic analysis module deeply analyzes the semantic content of the advertisement text to determine whether there is misleading propaganda or false promises; The manual review interface module provides an interface for manual reviewers to conduct manual review; The report generation module generates the final audit report.
2. The intelligent advertising language review system according to claim 1, characterized in that: The data receiving module includes an interface component, a data parsing component, a data verification component and a data storage component; The interface component is responsible for connecting with the external advertising publishing platform or advertising management system; The data parsing component parses the received data and converts it into a format that can be recognized within the system; The data verification component verifies the parsed data to ensure the integrity and accuracy of the data; The data storage component stores the verified data in the system's internal database for subsequent processing.
3. The intelligent advertising language review system according to claim 1, characterized in that: The preprocessing module includes a text cleaning component, a word segmentation component, a stop word filtering component and a text standardization component; The text cleaning component is used to remove irrelevant characters, special symbols, HTML tags, and extra spaces in the text; The word segmentation component is used to cut continuous text strings into meaningful words, where the component contains a dictionary and a word segmentation algorithm; the stop word filter component includes a stop word list; after word segmentation, the stop word filter component traverses each word in the text, and if the word exists in the stop word list, it is removed from the text to reduce data sparsity and improve the efficiency of subsequent analysis; The text normalization component is used to convert text into a uniform format or standard.
4. The intelligent advertising language review system according to claim 1, characterized in that: The language sentiment analysis module includes a feature extraction component and a machine learning component; The feature extraction component extracts features from the preprocessed text and converts them into numerical vectors; The machine learning component trains the sentiment classifier using a machine learning algorithm; Specifically, the model is trained with sentiment text data, and the model learns the mapping relationship from text features to sentiment labels. In the prediction stage, the model receives the preprocessed text feature vector and outputs the sentiment classification result.
5. The intelligent advertising language review system according to claim 4, characterized in that: The machine learning component trains the sentiment classifier using a machine learning algorithm, where the machine learning algorithm is at least any one of a vector machine, a naive Bayes, and a deep learning model.
6. The intelligent advertising language review system according to claim 1, characterized in that: The illegal vocabulary detection module includes an illegal vocabulary library, a vocabulary matching algorithm component and a whitelist management component; The illegal word library stores a list of illegal words, which include at least one of insulting language, sensitive political words, and illegal information keywords; The vocabulary matching algorithm component is used to achieve efficient matching between text and illegal vocabulary and identify illegal vocabulary contained in the text; the whitelist management component stores a whitelist list internally. For words matched to the whitelist, even if they appear in the illegal vocabulary, they will be regarded as compliant content.
7. The intelligent advertising language review system according to claim 1, characterized in that: The semantic analysis module includes a rule base and a logic judgment component; The rule base contains a set of rules for specific issues in advertising texts; The logic judgment component performs logic judgment on the vector representation or extracted features output by the semantic understanding component according to the rules in the rule base to determine whether the text meets the requirements defined by the rules.
8. The intelligent advertising language review system according to claim 1, characterized in that: The manual review interface module is respectively connected to the language sentiment analysis module, the illegal vocabulary detection module and the semantic analysis module. When the triggering conditions for manual review are met, relevant personnel conduct manual review through the manual review interface module.