Internet-oriented illegal advertisement identification method, device and system
By eliminating and evaluating special symbols in Internet advertisements, combining image data and delivery time to identify illegal advertisements, the problem of bypassing keyword detection is solved and more accurate identification of illegal advertisements is achieved.
Patent Information
- Application Number
- CN202510803531.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The prior art is difficult to effectively identify illegal advertisements in Internet advertisements that bypass keyword detection through special symbols, resulting in failure of identification and affecting consumer rights.
By eliminating special symbols in the advertising text, analyzing local text data before and after the symbol processing, combining image data and ad delivery time, evaluating the advertisement's illegal scores and identifying illegal advertisements.
Improve the accuracy and comprehensiveness of illegal advertisement identification to ensure consumer rights protection.
Smart Images

Figure CN120338884A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and particularly to a method, device and system for identifying illegal advertisements for the Internet. Background Art
[0002] With the rapid development of the Internet ecosystem, the resources invested by merchants, enterprises, etc. on the Internet are constantly increasing, resulting in an increasing number of advertisements on the Internet. Some unethical merchants and enterprises take the opportunity to invest a large number of illegal advertisements on the Internet, mixing them with normal advertisement information for false propaganda, vulgar content guidance, etc., seriously damaging the trust of consumers. For these illegal advertisement information, automatic identification is required for processing such as blocking and intercepting, so as to protect the rights and interests of consumers.
[0003] In the application of illegal advertisement identification technology in the Internet, some unethical advertisement merchants, when placing advertisements on the Internet in order to bypass illegal identification, will adopt ways such as text distortion, character occlusion or splitting sensitive words to bypass keyword detection, resulting in the failure of illegal advertisement identification, and thus flowing into the Internet, causing an adverse impact on the Internet user group. Summary of the Invention
[0004] This application provides a method, device and system for identifying illegal advertisements for the Internet. By eliminating the special symbols that may be carried in the advertisement text, this method solves the problem that advertisement manufacturers deliberately insert special symbols in advertisements to bypass keyword detection and interfere with the identification of illegal advertisements, making the identification of illegal information in advertisements more accurate.
[0005] The first aspect of this application provides a method for identifying illegal advertisements for the Internet, including: Collect advertisements and perform preprocessing to obtain the image data and text data of each advertisement; determine each special character in the text data; Perform word segmentation on the local text data before and after symbol processing for each special character. According to the part-of-speech and word frequency degree of each word segment, obtain the key degree of each word segment in the local text data before and after symbol processing respectively; according to the high distribution approximation situation of the word segment key degrees in the local text data before and after symbol processing for each special character, obtain the processing text consistency of each special character; According to the connection correlation degree and consistent distribution degree between each special character and the surrounding word segments, obtain the connection association degree of each special character; analyze the similarity situation between the local text data before and after symbol processing for each special character, and combine the processing text consistency and connection association degree to obtain the processing necessity of each special character; perform elimination processing based on the size of the processing necessity of special characters in the text data to obtain the processed text data; Analyze the degree of violation for the processed text data and image data of each advertisement, and combine it with the abnormality degree of the advertisement placement time to obtain the violation score for each advertisement; conduct violation identification based on the violation score of the advertisement.
[0006] Furthermore, the method for obtaining the key degree includes: For any special character, perform word segmentation on the local text data before symbol processing and the local text data after symbol processing of the character respectively to obtain word segments; determine the key weight of each word segment according to the part of speech of each word segment. For any word segment, take the frequency of occurrence of the word segment in the text data of the advertisement where the word segment is located as the word frequency of the word segment; take the product of the value obtained by negatively correlating the word frequency of the word segment and the key weight as the key degree of the word segment.
[0007] Furthermore, the method for obtaining the text consistency includes: For any special character, arrange all the word segments in the local text data before symbol processing of the character in descending order of key degree to obtain a key sequence; take the first preset number of word segments in the key sequence as the pre-key word segments before symbol processing of the character; similarly, obtain the post-key word segments after symbol processing of the character based on the local text data after symbol processing of the character. Count the data of the same word segments in the pre-key word segments and the post-key word segments of the character and perform normalization processing to obtain the word segment distribution approximation degree of the character. After calculating the difference in key degree between the pre-key word segments and the post-key word segments with the same serial number in the key sequence, sum all the differences and perform negative correlation mapping to obtain the high-key word segment similarity of the character. Combine the word segment distribution approximation degree and the high-key word segment similarity of the character to obtain the text consistency of the character.
[0008] Furthermore, the method for obtaining the connection correlation degree includes: For any special character, take the previous word segment adjacent to the character in the local text data as the pre-word segment of the character, and take the next word segment adjacent to the character in the local text data as the post-word segment of the character. Respectively obtain the sentiment scores of the pre-word segment and the post-word segment, and perform negative correlation mapping on the difference in sentiment scores between the pre-word segment and the post-word segment to obtain the connection sentiment correlation of the character. Count the frequency of occurrence of the pre-word segment and the character in the text data of the advertisement where they are located, and the frequency of occurrence of the character and the post-word segment in the text data of the advertisement where they are located. Then, take the sum of the two frequencies of occurrence as the phrase consistency of the character. Combine the connection sentiment correlation and the phrase consistency of the character to obtain the connection correlation degree of the character.
[0009] Further, the method for obtaining the processing necessity includes: For any special character, calculate the similarity between the local text data before symbol processing and the local text data after symbol processing of the character, and obtain the local text similarity index of the character; Take the product of the value obtained by negatively correlating the connection relevance of the character with the local text similarity index and the processing text consistency as the processing necessity of the character.
[0010] Further, the method for obtaining the processed text data includes: In the text data of the advertisement, eliminate the special characters whose processing necessity is greater than the preset processing threshold, otherwise perform symbol processing to obtain the processed text data of the advertisement.
[0011] Further, the method for obtaining the advertisement violation score includes: For any advertisement, input the image data of the advertisement into the trained image violation scoring model to output the image violation score of the advertisement; based on the preset violation word library, take the proportion of the number of violation words in the processed text data of the advertisement as the text violation score of the advertisement; Obtain the abnormal degree of the advertising time period of the advertisement according to the approximation degree between the advertising time of the advertisement and the preset abnormal time; take the advertising frequency of the same type of advertisement in the preset local time period of each advertisement placement as the abnormal degree of the advertisement type frequency; combine the abnormal degree of the advertising time period and the abnormal degree of the advertisement type frequency of the advertisement to obtain the advertising time abnormal index of the advertisement; Perform weighted summation on the image violation score, text violation score and advertising time abnormal index of the advertisement to obtain the advertisement violation score of the advertisement.
[0012] Further, the violation identification based on the violation score of the advertisement includes: When the advertisement violation score is greater than the preset violation threshold, mark the corresponding advertisement as a violated advertisement; When the advertisement violation score is less than the preset non-violation threshold, mark the corresponding advertisement as a non-violated advertisement; Otherwise, conduct manual review and violation identification on the advertisement; the preset non-violation threshold is less than the preset violation threshold.
[0013] In a second aspect, the present application provides an Internet-oriented violated advertisement identification system, and the system includes: An advertisement collection unit, configured to collect advertisements and perform preprocessing to obtain the image data and text data of each advertisement; determine each special character in the text data; An advertisement text processing unit is used to perform word segmentation on the local text data before and after symbol processing for each special character, and obtain the key degree of each word segment in the local text data before and after symbol processing according to the part of speech and word frequency degree of each word segment; according to the high distribution approximation of the word segment key degrees in the local text data before and after symbol processing for each special character, obtain the processing text consistency of each special character; Obtain the connection correlation degree of each special character according to the connection correlation degree and consistent distribution degree between each special character and the surrounding word segments; analyze the similarity between the local text data before and after symbol processing for each special character, and combine the processing text consistency and connection correlation degree to obtain the processing necessity of each special character; perform elimination processing based on the size of the processing necessity of special characters in the text data to obtain the processed text data; An advertisement violation recognition unit is used to analyze the violation degree of the processed text data and image data of each advertisement, and combine the abnormal degree of the advertisement placement time to obtain the violation score of each advertisement; perform violation recognition based on the violation score of the advertisement.
[0014] In a third aspect, the present application provides an Internet-oriented illegal advertisement recognition device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method according to the first aspect or any embodiment of the first aspect of the present application is implemented.
[0015] In a fourth aspect, the present application provides a computer program product, which includes computer program code. When the computer program code is executed, the method according to the first aspect or any embodiment of the first aspect of the present application is executed.
[0016] In a fifth aspect, the present application provides a computer-readable storage medium, which stores computer program code. When the computer program code is executed, the method according to the first aspect or any embodiment of the first aspect of the present application is executed.
[0017] The present invention has the following beneficial effects: The present invention analyzes the text information data carried by advertisements in detail, compares and analyzes the consistency of the high distribution of the part-of-speech and word frequency of each word segment in the local text before and after the processing of special symbols, as well as the similarity of the overall text, combines the connection and association of the words before and after the special symbols, analyzes the consistency of the key subjects represented by the text before and after the processing of special symbols, analyzes the degree of interference caused by the special symbols, and the degree to which the special symbols damage the connection semantics of the word segments on both sides, so as to obtain the necessary processing degree of the special symbols, and provides more accurate text information for the subsequent identification of illegal information. By combining the processed text data with the anomalies of the image and the release time, the illegal score of the advertisement is evaluated for identification, and a more accurate evaluation result is obtained. The present invention processes the interfering special symbols carried in the advertisement text by comparing the semantic changes represented by the coherent text content before and after the processing of special symbols, improves the accuracy of the identification of advertisement illegal information, and makes the identification of illegal advertisements more comprehensive and reliable. Brief Description of the Drawings
[0018] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is a flowchart of a method for identifying illegal advertisements on the Internet provided by an embodiment of the present invention; Figure 2 It is a structural diagram of a system for identifying illegal advertisements on the Internet provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a device for identifying illegal advertisements on the Internet provided by an embodiment of the present invention. Detailed Embodiments
[0020] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following, in combination with the drawings and preferred embodiments, details the specific embodiments, structures, features and effects of a method, device and system for identifying illegal advertisements on the Internet proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0022] The following specifically describes the specific solutions of a method, device, and system for identifying illegal advertisements for the Internet provided by the present invention in conjunction with the accompanying drawings.
[0023] An embodiment of the present application provides a method for identifying illegal advertisements for the Internet. Please refer to Figure 1 , which shows a flowchart of a method for identifying illegal advertisements for the Internet provided by an embodiment of the present invention. The method includes the following steps: S1: Collect advertisements and perform preprocessing to obtain the image data and text data of each advertisement; determine each special character in the text data.
[0024] The identification of illegal advertisements is mainly based on the specific information of the advertisements existing in the Internet to analyze and identify whether there are illegal information in the advertisements. Since the advertisements in the Internet are distributed in a decentralized manner, including static web pages, dynamically rendered pages, in-App advertisements, etc., it is impossible to directly apply the identification of illegal advertisements to all advertisement data. Therefore, it is necessary to actively collect the advertisement content in multiple channels to identify illegal advertisements.
[0025] In the embodiment of the present invention, advertisement information in multiple channels is crawled by using a crawling tool, and relevant illegal advertisement information is also obtained by users marking illegal situations. Among them, the crawling tool can use tools such as Scrapy and Selenium. The advertisement information in multiple channels includes the advertisement content of mainstream websites, forums or social media, and in-App advertisements can also be crawled through an Android emulator, and the advertisement SDK interface can be extracted by parsing the APK file to obtain advertisements, etc. It should be noted that when using the crawler tool to crawl the advertisement information in multiple channels, it should ensure compliance with laws and the Robots protocol, avoid crawling personal privacy data, and the process does not violate relevant laws and regulations and does not violate public order and good customs. The specific advertisement collection method is a well-known technical means for those skilled in the art and will not be elaborated here.
[0026] In the process of identifying illegal advertisements, an emotion dictionary, etc. is usually used to analyze the semantic emotion contained in the text. However, in the process of identifying illegal advertisements, there will be special symbols in the text, and these special symbols are not included in the Chinese emotion dictionary. The special symbols include symbols with expressive meanings and meaningless symbols. For example, "*" or "-" are direct interference words, and the interference generated is like "low * vulgar", and at this time, it can be directly eliminated for sensitive information analysis.
[0027] However, there are also symbols containing distorted interference information, such as " "Such special characters have replacement meanings. After replacement processing, it becomes "dim sum and private message to receive benefits". At this time, if the characters are directly eliminated, it will become "point and private receive benefits", which will have a greater blurring effect on sensitive information and violate the analysis. Therefore, it is necessary to further analyze the impact on semantic understanding of the characters that can be replaced, and retain the characters with information impact.
[0028] To make the processing of special characters more accurate, in the embodiments of the present invention, first, the advertisement is split multimodally, that is, the image data and text data of each advertisement are separated, and the advertisement text data is preprocessed. The preprocessing includes simplified and traditional Chinese processing and translation processing to unify the text data into simplified Chinese content for the accuracy of subsequent analysis. Then, the special characters in the preprocessed text data are determined to facilitate the subsequent analysis of the impact of special characters in the text. It should be noted that preprocessing and special character determination are technical means well-known to those skilled in the art. For example, special characters can be determined by matching symbol dictionaries through regular expressions, etc., and will not be elaborated and limited here.
[0029] S2: Perform word segmentation on the local text data before and after symbol processing of each special character, and obtain the key degree of each word segment in the local text data before and after symbol processing according to the part of speech and word frequency degree of each word segment; according to the high distribution approximation of the word segment key degrees in the local text data before and after symbol processing of each special character, obtain the text consistency of the processing of each special character.
[0030] First, based on the symbol classification dictionary, different symbol processing methods are defined. For example, characters without literal meanings such as " " can be directly deleted, but for those with literal replacement meanings such as " ", the special characters can be symbolically replaced through the dictionary to replace them with literal meanings. For example, " " is replaced by "heart", " " is replaced by "star", " " is replaced by "8", etc. Therefore, the symbol processing of each special character is based on the symbol classification dictionary.
[0031] By analyzing the changes in the semantic representation of the text before and after symbol processing, it is determined whether special characters affect the expression of text data. Considering that the meaning expression in text data is conveyed by words of various parts of speech, word segmentation analysis is performed on the local text of special text. Through the word segmentation information in the local text data segment where each special sentence is located, the key degree of the expression information of each word segment is reflected. To ensure the analyzability of local text data, in the embodiments of the present invention, the local text data of special characters refers to the data with coherent text before and after the special characters, that is, the text data composed of the sentence where the special character is located, the previous sentence and the subsequent sentence of the sentence where the special character is located is used as the local text data of each special character to represent the meaning of a local text.
[0032] Preferably, in the embodiments of the present invention, the method for obtaining the key degree includes: For any special character, the local text data before symbol processing and the local text data after symbol processing of the character are respectively subjected to word segmentation to obtain word segments, and the key weight of each word segment is determined according to the part of speech of each word segment. In the embodiments of the present invention, the BERT model can be used to perform word segmentation and part-of-speech annotation on text data. The word segmentation and part-of-speech analysis method is a well-known technical means for those skilled in the art and will not be elaborated here. By assigning different key weights to different parts of speech, it can be understood that after word segmentation of the local text data before and after symbol processing, the word segmentation results may be the same or different.
[0033] In the embodiments of the present invention, since key information generally appears as object > predicate > subject > others, noun > verb > preposition > others, the corresponding key weight values are set as: 0.8 > 0.5 > 0.3 > 0.1, 0.8 > 0.5 > 0.3 > 0.1. For example, "Add WeChat to get free courses" highlights WeChat and courses, and "Scan the code and grab a 50% discount immediately" highlights scanning the code and a 50% discount. The specific implementation of assigning key weights can be adjusted by the implementer according to the specific implementation scenario and is not limited here.
[0034] For any word segment, the frequency of occurrence of the word segment in the text data of the advertisement where the word segment is located is used as the word frequency of the word segment, which reflects the high-frequency degree of the corresponding word segment in the advertisement. The higher the frequency of the word segment, the less likely it is to carry key information. For example, for high-frequency words that often appear in the text, such as "get", "of", etc., these words generally do not serve as the main predicate of a noun, so they carry less key information and are given a lower key degree. On the contrary, for low-frequency words such as "WeChat" and "courses", they generally serve as the main predicate of a noun, so they carry more key information related to the advertisement and are given a higher key degree.
[0035] Therefore, the product of the value obtained by negatively correlating the word frequency of the word segmentation and the key weight is used as the key degree of the word segmentation. When the word frequency is lower and the key weight is higher, it indicates that the information carried in the advertisement text may be more critical. It should be noted that negative correlation mapping is a technical means well-known to those skilled in the art, such as in the form of inverse proportion or negative exponential power, and will not be limited or elaborated here.
[0036] Through the key degree, by analyzing the changes in the word segmentations with higher key degrees before and after symbol processing, it can be reflected whether the special symbol has a critical impact on the expression of the text segment where it is located. The word segmentations with higher key degrees represent the expression subjects of the corresponding text segments. When the expression subjects before and after symbol processing are more consistent, it indicates that the impact of the special symbol is small.
[0037] Preferably, in the embodiments of the present invention, by analyzing the high distribution of the key degrees of the word segmentations before and after processing, the consistency of the processed text is obtained. The method for obtaining the consistency of the processed text includes: For any special character, all the word segmentations in the local text data before symbol processing of this character are arranged in descending order of key degree to obtain a key sequence. The first preset number of word segmentations in the key sequence are used as the pre-key word segmentations before symbol processing of this character. By sorting, several word segmentations with higher key degrees are determined for subsequent approximate analysis. In the embodiments of the present invention, the preset number is taken as 3, that is, the first 3 word segmentations in the key sequence are used as the pre-key word segmentations.
[0038] Similarly, the post-key word segmentations after symbol processing of this character are obtained based on the local text data after symbol processing of this character, that is, all the word segmentations in the local text data after symbol processing of this character are arranged in descending order of key degree to obtain a key sequence. The first preset number of word segmentations in the key sequence are used as the post-key word segmentations after symbol processing of this character. To ensure the consistency of the selection of the preset number, the same number of word segmentations are used for approximate analysis.
[0039] Furthermore, the data of the same word segmentations in the pre-key word segmentations and the post-key word segmentations of this character are statistically counted and normalized to obtain the approximate degree of the word segmentation distribution of this character. When the content of the word segmentations with high key degrees before and after symbol processing is more approximately consistent, that is, the more the same word segmentations in the pre-key word segmentations and the post-key word segmentations, it indicates that the impact of special symbol processing is smaller.
[0040] It should be noted that normalization is a technical means well-known to those skilled in the art. The choice of normalization can be linear normalization or standard normalization, etc. The specific normalization method is not limited here.
[0041] Furthermore, after calculating the difference in key degrees between the pre-key participles and post-key participles with the same serial numbers in the key sequence, the sum of all differences is obtained and negatively correlated and mapped to obtain the high key similarity of word segmentation for this character. In the order of the key sequence, the difference in key degrees between the participles with the same serial numbers in the key sequence is calculated successively. In the embodiment of the present invention, the difference in key degrees between the first participles, the difference in key degrees between the second participles, and the difference in key degrees between the third participles of the two key sequences are calculated respectively, and then the sum of the three differences is calculated to reflect the deviation between the high-distribution key degrees. The smaller the deviation, the more consistent the distribution. Therefore, through negative correlation mapping, the high key similarity of word segmentation is obtained.
[0042] Finally, by combining the word segmentation distribution approximation degree and the high key similarity of word segmentation of this character, the text consistency of the processed character is obtained. In the embodiment of the present invention, the product of the word segmentation distribution approximation degree and the high key similarity of word segmentation is used as the text consistency of the processed character. When both the word segmentation distribution approximation degree and the high key similarity of word segmentation are higher, it indicates that the expression of the key content in the text before and after processing is more consistent, and the interference impact of this special character is not high.
[0043] S3: Obtain the connection correlation degree of each special character according to the connection correlation degree and the consistent distribution degree between each special character and the pre- and post-word segmentation; analyze the similarity between the local text data before and after the symbol processing of each special character, and combine the text consistency of the processing and the connection correlation degree to obtain the necessity of processing for each special character; perform elimination processing based on the size of the necessity of processing of the special characters in the text data to obtain the processed text data.
[0044] When the consistency of the text expression before and after processing is high, there will also be reasons that some special symbols act as modifying components for the connected words. Even though the influence of the special symbols is less when participating in the text expression at this time, removing them will also affect the analysis of sentence coherence, such as some digital representation symbols. Therefore, through further analysis of the association between each special symbol and the surrounding words, the influence analysis result of the special symbol is improved from the aspect of emotional semantics.
[0045] Preferably, in the embodiment of the present invention, the connection correlation degree is obtained through the connection correlation degree and the consistent distribution degree between the special character and the pre- and post-word segmentation. The method for obtaining the connection correlation degree includes: For any special character, the previous participle adjacent to the character in the local text data is used as the pre-participle of the character, and the next participle adjacent to the character in the local text data is used as the post-participle of the character. The participles connected before and after each special character are respectively marked, and the connection situation is analyzed.
[0046] Obtain the sentiment scores of the pre-tokenization and post-tokenization respectively, perform a negative correlation mapping on the difference in sentiment scores between the pre-tokenization and post-tokenization to obtain the connection sentiment correlation of this character. Through the connection situation of word sentiments, reflect the degree of emotional semantic association of the text composition. For example, the emotional semantic of "free" is positive, and the sentiment score is +0.8, while the emotional semantic of "fraud" is negative, and the sentiment score is -0.9. When the difference in sentiment scores is larger, the semantic coherence is less relevant. Therefore, through the negative correlation of the difference in situation scores, reflect the correlation of the connection sentiment. In the embodiments of the present invention, the sentiment scores are obtained using the textblob library, and can also be obtained through a pre-trained neural network model. The obtaining method is not elaborated and limited here.
[0047] Further, after counting the occurrence frequencies of the pre-tokenization and this character in the text data of the advertisement where they are located, and the occurrence frequencies of this character and the post-tokenization in the advertisement text data, take the sum value of the two occurrence frequencies as the phrase consistency of this character. When a special symbol often appears together with a certain continuous tokenization, it reflects a relatively high correlation between the special symbol and the tokenization. Removing it may have an impact. Therefore, count the synchronization frequencies of the special character and the pre-tokenization, and the synchronization frequencies of the special character and the post-tokenization respectively. The sum of the frequencies gives the possible phrase consistency of the special character.
[0048] Finally, combine the connection sentiment correlation and phrase consistency of this character to obtain the connection association degree of this character. In the embodiments of the present invention, take the product of the connection sentiment correlation and phrase consistency of this character as the connection association degree. When the connection sentiment correlation and phrase consistency are larger, it indicates that the special character has a higher degree of correlation with the words before and after, and the greater the possible elimination of the impact.
[0049] Therefore, finally, from a global analysis, through the overall similarity of the text semantics before and after symbol processing, combined with the processed text consistency and connection association degree, obtain the necessity of processing for each special character. In the embodiments of the present invention, the obtaining method of the processing necessity includes: For any special character, calculate the similarity between the local text data before symbol processing and the local text data after symbol processing of this character to obtain the local text similarity index of this character. For the local text data before symbol processing and the local text data after symbol processing of this character, perform text vector conversion respectively. After obtaining two text vectors, calculate the cosine similarity between the text vectors to obtain the local text similarity index. It should be noted that the obtaining of text vectors and the calculation of cosine similarity are both well-known technical means to those skilled in the art, such as using a word embedding model for vector conversion, etc., which are not elaborated and limited here.
[0050] Furthermore, the product of the value obtained by negatively correlating the connection relevance of the character and the local text similarity index, as well as the product for processing text consistency, is used as the necessity for processing the character. When the local text similarity index and the processing text consistency of the local text data of the special character are higher, it indicates that the influence degree of the text data before and after processing the special character from the global and local analysis is not high, and the possibility of eliminating the processing is higher. When the connection relevance is smaller, it indicates that the connection association influence between the special character and the previous and subsequent word segmentation is not high, and the influence after elimination is smaller. Therefore, the necessity for processing is greater.
[0051] Therefore, elimination processing can be performed based on the magnitude of the necessity for processing special characters in the text data. In the embodiment of the present invention, the method for obtaining the processed text data includes: in the text data of the advertisement, eliminating the special characters whose values after normalizing the necessity for processing are greater than the preset processing threshold, and otherwise performing symbol processing to obtain the processed text data of the advertisement. The preset processing threshold can be set to 0.9. When it is greater than the processing threshold, it indicates that the influence of the elimination result is not significant, and direct elimination processing can be performed. The rest can be processed according to the processing method of the symbol classification dictionary, that is, retaining after replacing the symbol with text to ensure the integrity of the semantics, which is used to improve the accuracy of subsequent violation recognition.
[0052] S4: Analyze the violation degree of the processed text data and image data of each advertisement, and combine the abnormality degree of the advertisement placement time to obtain the violation score of each advertisement; perform violation recognition based on the violation score of the advertisement.
[0053] After processing the text data of the advertisement, the violation situations of the text data and the image data can be respectively identified. Considering the characteristic that violation advertisements usually are placed in the early morning, a more comprehensive evaluation of the final violation score is carried out by combining the abnormality degree of the advertisement placement. Preferably, in the embodiment of the present invention, the method for obtaining the advertisement violation score includes: For any advertisement, input the image data of the advertisement into the trained image violation score model to output the image violation score of the advertisement, and identify the violation elements in the image through the model to output the score result. Further, based on the preset violation word library, the proportion of the number of violation words in the processed text data of the advertisement is used as the text violation score of the advertisement. By comparing and identifying the processed text data of the advertisement with the preset violation word library, the number of violation words is determined, and the ratio of the number of violation words to the total number of words in the processed text data of the advertisement is used as the text violation score.
[0054] Further analyze the time anomaly situation, and obtain the anomaly degree of the advertising time period of the advertisement according to the approximation degree between the advertisement delivery time and the preset anomaly time. In the embodiment of the present invention, calculate the time difference between the advertisement delivery time and the preset anomaly time for negative correlation mapping as the time period anomaly degree. The smaller the time difference, the closer it is to the abnormal delivery time, and the higher the probability of advertising violation anomaly. Among them, the preset anomaly time can be set to zero o'clock every day, that is, the closer the delivery time is to the early morning, the higher the probability of anomaly. The implementer of the preset anomaly time can adjust it by himself.
[0055] Furthermore, take the advertising frequency of the same type of advertisement in the preset local time period of each advertisement delivery as the anomaly degree of the advertisement type frequency. The more frequent the same type of delivery, the more significant the violation performance is reflected. In the embodiment of the present invention, within one hour centered on each advertisement delivery, count the number of deliveries of the same type of advertisement as the type frequency anomaly degree. The preset local time period is one hour centered on each delivery time, and the specific value can be adjusted by the implementer himself.
[0056] Combine the anomaly degree of the advertising time period and the anomaly degree of the type frequency of the advertisement to obtain the advertising time anomaly index of the advertisement. In the embodiment of the present invention, take the product of the anomaly degree of the advertising time period and the anomaly degree of the type frequency as the advertising time anomaly index of the advertisement, which reflects the possible degree of anomaly in the advertising time.
[0057] Finally, perform weighted summation on the image violation score, text violation score, and advertising time anomaly index of the advertisement to obtain the advertisement violation score of the advertisement. In the embodiment of the present invention, set the weight values of the image violation score, text violation score, and advertising time anomaly index to 0.3, 0.6, and 0.1 respectively, and the implementer can adjust the specific weight size by himself.
[0058] As an example, the expression of the advertisement violation score is: ; In the formula, represents the advertisement violation score, represents the image violation score, represents the text violation score, represents the advertising time anomaly index. Through comprehensive multi-faceted violation evaluations, a more comprehensive advertisement violation evaluation result is obtained. The higher the advertisement violation score, the greater the presence of violation factors in the advertisement, and the more likely the advertisement is a violation advertisement.
[0059] Based on the violation score of the advertisement, further perform final violation identification. In the embodiments of the present invention, when the advertisement violation score is greater than the preset violation threshold, the advertisement is considered a violating advertisement, and the corresponding advertisement is marked as a violating advertisement. When the advertisement violation score is less than the preset non-violating threshold, the advertisement is considered not to be a violating advertisement, and the corresponding advertisement is marked as a non-violating advertisement. Otherwise, that is, when the advertisement violation score is less than or equal to the preset violation threshold and greater than or equal to the preset non-violating threshold, it is impossible to directly determine whether the advertisement is a violation, and it is necessary to make a judgment through human review. Therefore, the advertisement is subject to human review for violation identification. In the embodiments of the present invention, the higher the advertisement violation score, the more likely the advertisement is a violating advertisement. Therefore, the preset non-violating threshold is less than the preset violation threshold. The preset non-violating threshold is set to 0.4, and the preset violation threshold is set to 0.6. The specific values can be adjusted by the implementer himself.
[0060] In summary, the present invention specifically analyzes the text information data carried by the advertisement, compares and analyzes the consistency of the high distribution of the part-of-speech and word frequency of each word segment in the local text before and after the special symbol processing, as well as the similarity of the overall text, combines the connection and association of the words before and after the special symbol, analyzes the consistency of the key subjects represented by the text before and after the special symbol processing, analyzes the degree of interference caused by the special symbol, and the degree of damage to the connection semantics of the word segments on both sides by the special symbol, to obtain the necessary processing degree of the special symbol, and provides more accurate text information for the subsequent identification of violation information. By combining the processed text data with the abnormality possibility of the image and the release time, evaluate the violation score of the advertisement for identification, and obtain a more accurate evaluation result. The present invention processes the interfering special symbols carried in the advertisement text by comparing the semantic changes represented by the coherent text content before and after the special symbol processing, improves the accuracy of the advertisement violation information identification, and makes the identification of violating advertisements more comprehensive and reliable.
[0061] This application also provides an Internet-oriented violating advertisement identification system. Please refer to Figure 2 , which shows the structural diagram of an Internet-oriented violating advertisement identification system provided by an embodiment of the present invention. The system includes: an advertisement collection unit 201, an advertisement text processing unit 202, and an advertisement violation identification unit 203.
[0062] The advertisement collection unit 201 is used to collect advertisements and perform preprocessing to obtain the image data and text data of each advertisement; determine each special character in the text data; The advertisement text processing unit 202 is used to perform word segmentation on the local text data before and after the symbol processing of each special character, and respectively obtain the key degree of each word segment in the local text data before and after the symbol processing according to the part-of-speech and word frequency degree of each word segment; obtain the processing text consistency of each special character according to the approximate high distribution of the word segment key degrees in the local text data before and after the symbol processing of each special character; According to the connection relevance and uniform distribution degree between each special character and the surrounding word segmentation, obtain the connection correlation degree of each special character; analyze the similarity between the local text data before and after the symbol processing of each special character, and combine the processing text consistency and connection correlation degree to obtain the necessity of processing for each special character; perform elimination processing based on the magnitude of the necessity of processing special characters in the text data to obtain the processed text data. The advertisement violation recognition unit 203 is used to analyze the violation degree of the processed text data and image data of each advertisement, and combine the abnormality degree of the advertisement placement time to obtain the violation score of each advertisement; perform violation recognition based on the violation score of the advertisement.
[0063] It should be noted that the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the embodiments of an Internet-oriented illegal advertisement recognition system and an Internet-oriented illegal advertisement recognition method provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.
[0064] The embodiment of the present application further provides an Internet-oriented illegal advertisement recognition device. Please refer to Figure 3 , which shows a schematic structural diagram of an Internet-oriented illegal advertisement recognition device provided by an embodiment of the present invention. The computer device includes a memory 301, a processor 302, and a computer program 303 stored in the memory 301 and running on the processor 302. Among them, when the processor 302 executes the computer program 303, the computer device can execute any one of the Internet-oriented illegal advertisement recognition methods introduced above.
[0065] The embodiment of the present application further provides a computer program product. When the computer program product runs on a computer device, the computer device can execute any one of the Internet-oriented illegal advertisement recognition methods introduced above.
[0066] The embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores computer program code. When the computer program code runs on a computer device, the computer device can execute any one of the Internet-oriented illegal advertisement recognition methods introduced above.
[0067] In the embodiments provided in the present application, it should be understood that the provided computer device, computer program product, and computer-readable storage medium are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the methods provided above, and will not be elaborated here.
[0068] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0069] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized.
Claims
1. An illegal advertisement recognition method for the Internet, characterized in that The method includes: Collect advertisements and perform preprocessing to obtain the image data and text data of each advertisement; determine each special character in the text data; Perform word segmentation on the local text data before and after symbol processing for each special character. According to the part-of-speech and word frequency degree of each word segment, obtain the key degree of each word segment in the local text data before and after symbol processing respectively; according to the high distribution approximation of the word segment key degrees in the local text data before and after symbol processing for each special character, obtain the processing text consistency of each special character; According to the connection correlation degree and the consistent distribution degree between each special character and the adjacent word segments before and after, obtain the connection association degree of each special character; analyze the similarity between the local text data before and after symbol processing for each special character, and combine the processing text consistency and the connection association degree to obtain the processing necessity of each special character; perform elimination processing based on the size of the processing necessity of the special characters in the text data to obtain the processed text data; Analyze the violation degree of the processed text data and image data of each advertisement, and combine the abnormal degree of the advertisement placement time to obtain the violation score of each advertisement; perform violation identification based on the violation score of the advertisement.
2. The method for identifying illegal advertisements for the Internet according to claim 1, wherein The method for obtaining the key degree includes: For any special character, perform word segmentation on the local text data before symbol processing and the local text data after symbol processing of this character respectively to obtain word segments; determine the key weight of each word segment according to the part-of-speech of each word segment; For any word segment, take the occurrence frequency of this word segment in the text data of the advertisement where this word segment is located as the word frequency of this word segment; take the product of the value obtained by performing negative correlation mapping on the word frequency of this word segment and the key weight as the key degree of this word segment.
3. The method for identifying illegal advertisements for the Internet according to claim 1, characterized in that, The method for obtaining the processing text consistency includes: For any special character, sort all the word segments in the local text data before symbol processing of this character in descending order of key degree to obtain a key sequence; take the first preset number of word segments in the key sequence as the front key word segments before symbol processing of this character; similarly, obtain the back key word segments after symbol processing of this character based on the local text data after symbol processing of this character; Count the data of the same word segments in the front key word segments and the back key word segments of this character and perform normalization processing to obtain the word segment distribution approximation of this character; After calculating the difference in key degrees between the front key word segments and the back key word segments with the same serial number in the key sequence, sum all the differences and perform negative correlation mapping to obtain the high key similarity of the word segments of this character; Combine the word segment distribution approximation and the high key similarity of the word segments of this character to obtain the processing text consistency of this character.
4. The method for identifying illegal advertisements for the Internet according to claim 1, characterized in that, The method for obtaining the connection association degree includes: For any special character, take the previous adjacent word segment of this character in the local text data as the front word segment of this character, and take the next adjacent word segment of this character in the local text data as the back word segment of this character; Respectively obtain the sentiment scores of the front word segment and the back word segment, and perform negative correlation mapping on the difference in sentiment scores between the front word segment and the back word segment to obtain the connection sentiment correlation of this character; After counting the occurrence frequency of the pre-tokenization and the character in the text data of the advertisement where it is located, as well as the occurrence frequency of the character and the post-tokenization in the advertisement text data, the sum value of the two occurrence frequencies is used as the phrase consistency of the character; Combining the connection emotional relevance and phrase consistency of the character, the connection association degree of the character is obtained.
5. The method for identifying illegal advertisements for the Internet according to claim 1, characterized in that, The method for obtaining the processing necessity includes: For any special character, calculate the similarity between the local text data before symbol processing and the local text data after symbol processing of the character, and obtain the local text similarity index of the character; The product of the value obtained by negatively correlating the connection relevance of the character, the local text similarity index, and the processing text consistency is used as the processing necessity of the character.
6. The method for identifying illegal advertisements for the Internet according to claim 1, characterized in that The method for obtaining the processed text data includes: In the text data of the advertisement, eliminate the special characters with a processing necessity greater than the preset processing threshold, otherwise perform symbol processing to obtain the processed text data of the advertisement.
7. The method for identifying illegal advertisements for the Internet according to claim 1, characterized in that The method for obtaining the advertisement violation score includes: For any advertisement, input the image data of the advertisement into the trained image violation scoring model to output the image violation score of the advertisement; based on the preset violation word library, use the proportion of the number of violation words in the processed text data of the advertisement as the text violation score of the advertisement; According to the approximation degree between the advertisement placement time and the preset abnormal time, obtain the placement period abnormality degree of the advertisement; use the placement frequency of the same type of advertisement in the preset local period of each placement of the advertisement as the type frequency abnormality degree of the advertisement; combine the placement period abnormality degree and the type frequency abnormality degree of the advertisement to obtain the placement time abnormality index of the advertisement; Perform a weighted sum of the image violation score, text violation score, and placement time abnormality index of the advertisement to obtain the advertisement violation score of the advertisement.
8. The method for identifying illegal advertisements for the Internet according to claim 1, characterized in that, The violation identification based on the violation score of the advertisement includes: When the advertisement violation score is greater than the preset violation threshold, mark the corresponding advertisement as a violated advertisement; When the advertisement violation score is less than the preset non-violation threshold, mark the corresponding advertisement as a non-violated advertisement; Otherwise, conduct a human review for violation identification of the advertisement; the preset non-violation threshold is less than the preset violation threshold.
9. An illegal advertisement recognition system for the Internet, characterized in that, The system includes: An advertisement collection unit for collecting advertisements and performing preprocessing to obtain the image data and text data of each advertisement; determining each special character in the text data; An advertisement text processing unit for tokenizing in the local text data before and after symbol processing of each special character, and respectively obtaining the key degree of each token in the local text data before and after symbol processing according to the part of speech and word frequency degree of each token; obtaining the processing text consistency of each special character according to the high distribution approximation of the token key degrees in the local text data before and after symbol processing of each special character; According to the connection relevance and uniform distribution degree between each special character and the surrounding word segments, obtain the connection correlation degree of each special character; analyze the similarity between the local text data before and after the symbol processing of each special character, and combine the processing text consistency and connection correlation degree to obtain the necessity of processing for each special character; based on the size of the necessity of processing special characters in the text data, perform elimination processing to obtain the processed text data. An advertisement violation recognition unit is used to analyze the violation degree of the processed text data and image data of each advertisement, and combine the abnormality degree of the advertisement placement time to obtain the violation score of each advertisement; based on the violation score of the advertisement, perform violation recognition.
10. An illegal advertisement recognition device for the Internet, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that When the processor executes the computer program, it implements an internet-oriented illegal advertisement recognition method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Chinese variation text matching recognition method
CN101976253A
Text categorization method and device
CN103514174A
Sensitive text recognition method, device, medium and computer equipment
CN110472234A
Information detection method, electronic equipment and computer storage medium
CN112559672A
Illegal advertisement identification method combining character visual features and character content features
CN114155529A
Cited By
Advertisement content generation system based on big data analysis
CN121329512A