Advertisement click flow analysis and anomaly recognition system and method

Through real-time data comparison and cross-domain correlation analysis, abnormal behavior of advertising clicks can be identified, solving the problem of misjudgment in existing technologies, achieving accurate anomaly identification and effective advertising strategy adjustment, and improving the efficiency and economic benefits of advertising.

CN120634640AActive Publication Date: 2025-09-12BEIJING HILONG TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510791643.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-12
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

Existing technologies cannot accurately determine whether there are abnormalities in the number of ad clicks, and are prone to misjudgment during the identification process, resulting in waste of advertising resources and disrupted market order.

Method used

By obtaining ad click data in real time and comparing it with historical data, data with fluctuations exceeding a threshold are marked as abnormal data, and correlation analysis is performed to extract co-occurrence feature combinations. Graph neural networks are used to build cross-domain correlation maps to identify abnormal data.

Benefits of technology

It improves the accuracy of anomaly identification, reduces the cost of invalid advertising, helps advertisers adjust their advertising strategies, improves advertising effectiveness and return on investment, and ensures a healthy and orderly advertising ecosystem.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634640A_ABST
    Figure CN120634640A_ABST
Patent Text Reader

Abstract

The invention discloses an advertisement click flow analysis and anomaly recognition system and method, and relates to the technical field of internet advertisements, and the method comprises the steps: obtaining daily advertisement click rate data in a preset time period in real time; determining the daily advertisement click rate data in the current preset time period and comparing the daily advertisement click rate data with the daily advertisement click rate data in the same time period in the historical data; the advertisement click data volume with the fluctuation amplitude exceeding a preset threshold value is marked as abnormal data based on the comparison result, the abnormal data comprises determined abnormal data and to-be-determined abnormal data, the data meeting all abnormal feature judgment standards are determined abnormal data, the data meeting only part of abnormal feature judgment standards are to-be-determined abnormal data, and the data meeting only part of abnormal feature judgment standards are to-be-determined abnormal data. According to the method, through real-time acquisition and accurate comparison with historical data, the abnormal fluctuation of the click rate can be quickly found, the determined abnormal data and the undetermined abnormal data can be strictly distinguished, and co-occurrence feature combinations are mined by means of correlation analysis, so that misjudgment and missed judgment are avoided, and the accuracy of abnormal recognition is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of Internet advertising, and specifically relates to a system and method for analyzing and identifying anomalies of advertisement click traffic. Background Art

[0002] In today's digital marketing landscape, click-through advertising plays a crucial role, serving as an essential and crucial interactive link in the entire marketing ecosystem. When users engage with various ad creatives—whether they're colorful and creative image ads, concise and engaging text links that directly address key concerns, or vivid and engaging video ads—each active click embodies complex psychological and behavioral logic, carrying immense commercial value and market significance.

[0003] Essentially, clicking on an ad is a quantitative reflection of the degree of precise match between user interests and ad content. When an ad can accurately grasp a user's needs, preferences, and potential purchasing intentions, and is presented in an appropriate manner, users are more likely to be attracted and click on it. This click-through behavior acts as a signal to the advertiser that the ad has successfully attracted the target audience's attention to a certain extent, building a bridge for subsequent conversions and sales, and injecting momentum into the continued advancement of marketing activities.

[0004] However, in actual application scenarios, this marketing model, which is supposed to be based on real user interest and positive interaction, faces a severe challenge: click manipulation. To gain illicit profits, some lawless individuals or unscrupulous businesses use various technical means, such as writing malicious scripts, using robots to simulate user clicks, and forming professional click manipulation teams, to create large amounts of false click data. This artificial manipulation seriously distorts the actual performance of ads and market feedback.

[0005] On the one hand, fake click-through rates can create a false impression, leading advertisers to mistakenly believe their advertising strategies are effective and their creatives are highly appealing and popular. Based on this misconception, advertisers may further increase their advertising budgets and scale up their campaigns, even though these investments fail to reach real users with actual demand and potential to purchase, resulting in a significant waste of resources.

[0006] On the other hand, for advertisers, every click can be associated with a significant cost. Whether using a pay-per-click (CPC) advertising model or other click-based billing methods, click manipulation directly leads to unnecessary consumption of the advertiser's advertising budget. The large number of fraudulent clicks drowns out genuine, valid clicks in a sea of ​​data, significantly reducing the input-output ratio of advertising, eroding a company's marketing effectiveness, and bringing heavy financial losses and market competition pressure to advertisers.

[0007] More seriously, click-through manipulation disrupts the order and ecosystem of the entire digital marketing market. It undermines the advertising delivery mechanism based on real data and fair competition, putting advertisers who truly prioritize ad quality and user experience at a disadvantage while allowing unscrupulous individuals to profit from fraudulent practices. If this unhealthy phenomenon is not effectively curbed, it will severely impact the credibility and sustainable development of the digital marketing industry, hindering innovation and progress across the industry.

[0008] Therefore, how to accurately determine whether there are abnormalities in the number of ad clicks and effectively avoid misjudgments during the identification process. Summary of the Invention

[0009] The purpose of the present invention is to provide a system and method for analyzing and identifying anomalies in advertising click traffic, which solves the technical problem in the prior art that it is impossible to accurately determine whether there are anomalies in the advertising click volume and effectively avoids misjudgment during the identification process.

[0010] A method for analyzing and identifying anomalies in advertisement click traffic, comprising:

[0011] S1. Obtain daily ad click data within a preset time period in real time;

[0012] S2. Determine the daily ad click volume data for the current preset time period and compare it with the daily ad click volume data for the same time period in historical data;

[0013] S3. Based on the comparison results, the amount of ad click data whose fluctuation exceeds a preset threshold is marked as abnormal data. The abnormal data includes confirmed abnormal data and pending abnormal data. Among them, the abnormal data that meets all the abnormal feature judgment criteria is confirmed abnormal data, and the abnormal data that only meets some of the abnormal feature judgment criteria is pending abnormal data.

[0014] S4, performing correlation analysis on the determined abnormal data and the abnormal data to be determined, extracting co-occurring feature combinations between the two, analyzing the abnormal data to be determined based on the co-occurring feature combinations, and determining whether the abnormal data to be determined is determined abnormal data;

[0015] S5. Combining all confirmed abnormal data, a final determination result of abnormal ad clicks is obtained.

[0016] As a further solution of the present invention: the abnormal feature judgment standard includes at least one of high-frequency clicks, false device IDs, and abnormal click time periods.

[0017] As a further solution of the present invention: the S4 step is specifically as follows:

[0018] Extracting at least two features from the determined abnormal data to form a feature combination;

[0019] Obtaining abnormal data to be determined and extracting features of the abnormal data to be determined;

[0020] The features of the abnormal data to be determined are matched with the feature combination of the determined abnormal data. If the features of the abnormal data to be determined include all the features in the feature combination of the determined abnormal data, it is determined that the abnormal data to be determined is associated with the determined abnormal data, and the abnormal data to be determined is determined as abnormal data; if the features of the abnormal data to be determined only include some of the features in the feature combination of the determined abnormal data, the matching status of the features of the abnormal data to be determined and other feature combinations of other determined abnormal data is further determined until the association judgment is completed.

[0021] As a further solution of the present invention: the characteristics include at least two of the click time, click device identification, click IP address, and click frequency.

[0022] As a further solution of the present invention: also include:

[0023] Acquire determined abnormal data, perform feature co-occurrence analysis on the determined abnormal data, and generate a feature correlation matrix, wherein element values ​​in the matrix represent the degree of correlation between features;

[0024] Acquire the abnormal data to be determined, extract its features and perform weighted matching with the feature correlation matrix;

[0025] Calculate a weighted matching score of the features of the to-be-determined abnormal data and the features of the determined abnormal data. If the score reaches a preset matching threshold, determine that the to-be-determined abnormal data is associated with the determined abnormal data, and determine the to-be-determined abnormal data as abnormal data.

[0026] If the score does not reach the threshold, a list of abnormal suspicions is generated based on the missing features for subsequent manual review or automatic completion verification.

[0027] As a further solution of the present invention: the calculation of the weighted matching score includes:

[0028] Determine the corresponding weight of each feature of the abnormal data to be determined in the feature correlation matrix;

[0029] Calculate the matching score of each feature based on the feature matching situation and corresponding weight;

[0030] The matching scores of all features are accumulated to obtain the weighted total matching score.

[0031] As a further solution of the present invention: also include:

[0032] Perform time series pattern mining on the click time series in the identified abnormal data to generate a time series abnormal pattern library;

[0033] Build user behavior profiles based on historical user click behavior data and extract normal behavior pattern features;

[0034] Obtain the abnormal data to be determined, determine whether its click time series matches any pattern in the time series abnormal pattern library, and at the same time, compare the characteristics of the abnormal data to be determined with the normal pattern characteristics of the user behavior profile, and identify the feature items that deviate from the normal pattern; if the abnormal data to be determined meets any of the following conditions: the click time series matches the time series abnormal pattern; the number of feature items that deviate from the normal pattern reaches a preset threshold; there are specific key features that deviate from the normal pattern, then the abnormal data to be determined is determined to be abnormal data; otherwise, the unmatched features and deviation analysis results are used as optimization feedback data to update the time series abnormal pattern library and user behavior profile.

[0035] As a further solution of the present invention: also include:

[0036] Integrate cross-domain data from advertising delivery platforms, including but not limited to ad creative types, delivery channel data, and user conversion behavior data;

[0037] Build a cross-domain correlation graph based on graph neural networks to analyze the potential correlation paths between the abnormal data to be determined and the cross-domain data;

[0038] If it is found that the abnormal data to be determined overlaps with the known abnormal pattern in a cross-domain correlation path, or its correlation path causes an abnormal interruption in the advertising conversion chain, it will be determined as abnormal data;

[0039] Otherwise, the cross-domain association analysis results are included in the optimization feedback data to update the association weights and abnormal path rules of the cross-domain association graph.

[0040] As a further solution of the present invention: the construction of the cross-domain association graph includes:

[0041] Extract entity data of ad click data, material data, and channel data;

[0042] Mining relationships between entities through association rules and assigning dynamic weights to relationship edges;

[0043] Graph embedding technology is used to convert the graph into a feature vector for rapid retrieval of abnormal association paths.

[0044] On the other hand, the present invention also proposes an advertisement click traffic analysis and anomaly identification system, which is applicable to the above-mentioned advertisement click traffic analysis and anomaly identification method, and the method comprises the following steps:

[0045] Data acquisition module, which obtains daily advertising click data within a preset time period in real time;

[0046] The data comparison module determines the daily ad click data within the current preset time period and compares it with the daily ad click data of the same time period in the historical data;

[0047] A judgment module, based on the comparison result, marks the amount of advertising click data whose fluctuation exceeds a preset threshold as abnormal data, wherein the abnormal data includes confirmed abnormal data and pending abnormal data. Among them, the abnormal data that meets all the abnormal feature judgment criteria is confirmed abnormal data, and the abnormal data that meets only some of the abnormal feature judgment criteria is pending abnormal data;

[0048] An association analysis module performs association analysis on the confirmed abnormal data and the abnormal data to be determined, extracts a co-occurring feature combination between the two, analyzes the abnormal data to be determined based on the co-occurring feature combination, and determines whether the abnormal data to be determined is confirmed abnormal data;

[0049] The result output module integrates all the confirmed abnormal data to obtain the final judgment result of abnormal ad clicks.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] By accurately comparing real-time collection with historical data, the present invention can quickly discover abnormal fluctuations in click volume, strictly distinguish between confirmed and pending abnormal data, and use association analysis to mine co-occurrence feature combinations, which is conducive to avoiding misjudgments and missed judgments, and is conducive to significantly improving the accuracy of abnormality identification; the final comprehensive judgment output results provide advertisers with intuitive early warnings and optimization suggestions, which is conducive to effectively reducing the cost of invalid advertising, helping advertisers to adjust their advertising strategies in a timely manner, improving the effectiveness and return on investment of advertising, and ensuring a healthy and orderly advertising ecosystem. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 Schematic diagram of the framework structure of the method of the present invention. DETAILED DESCRIPTION

[0053] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0054] First aspect: please refer to Figure 1 , this application provides a method for analyzing and identifying anomalies in advertising click traffic, including:

[0055] S1. Obtain daily ad click data within a preset time period in real time;

[0056] S2. Determine the daily ad click volume data for the current preset time period and compare it with the daily ad click volume data for the same time period in historical data;

[0057] S3. Based on the comparison results, the amount of ad click data whose fluctuation exceeds a preset threshold is marked as abnormal data. The abnormal data includes confirmed abnormal data and pending abnormal data. Among them, the abnormal data that meets all the abnormal feature judgment criteria is confirmed abnormal data, and the abnormal data that only meets some of the abnormal feature judgment criteria is pending abnormal data.

[0058] S4. Perform correlation analysis on the confirmed abnormal data and the abnormal data to be determined, extract the co-occurring feature combination between the two, analyze the abnormal data to be determined based on the co-occurring feature combination, and determine whether the abnormal data to be determined is confirmed abnormal data;

[0059] S5. Combining all confirmed abnormal data, a final determination result of abnormal ad clicks is obtained.

[0060] As an optional embodiment, the abnormal feature judgment standard includes at least one of high-frequency clicks, false device IDs, and abnormal click time periods.

[0061] It should be understood that by connecting with advertising platforms and traffic statistics tools (such as Google Analytics and Baidu Statistics) through data interfaces, ad click data can be pulled at a frequency of minutes (e.g., every 10 minutes). For example, if an e-commerce platform places a product advertisement on Douyin, the system can obtain real-time information such as the number of times users click on the ad to jump to the product details page, the click timestamp, and the user's IP address.

[0062] The collected data is stored in a time-series database (such as InfluxDB) and indexed by dimensions such as date, ad placement ID, and ad creative type. For example, daily ad click data from the National Day promotion period from October 1 to January 7, 2024, can be stored by category for different product ads, facilitating quick retrieval of historical data from the same time period.

[0063] The system automatically matches historical data for the same period based on the current date. For example, if the current date is June 2025, the daily ad click volume for June 2024 and June 2023 will be automatically retrieved. Furthermore, to account for changes in ad delivery strategies, only historical data with the same delivery channels and creatives will be used as a comparison benchmark.

[0064] The calculation method uses the percentage fluctuation formula to calculate the difference, specifically: Come to.

[0065] Furthermore, abnormal data that meets the following conditions simultaneously (taking e-commerce ads as an example): fluctuations exceeding 45%; click times concentrated between midnight and 6 a.m. (a non-active period); more than 10 clicks from the same IP address within a short period of time (e.g., 10 minutes); and data that meets only one or two of the above conditions, such as a fluctuation of 50% but a normal click time distribution, is considered abnormal data to be determined.

[0066] Furthermore, the preset threshold can be dynamically adjusted according to the type of advertisement. For example, for new product promotion advertisements, the threshold is set to 35% (due to large fluctuations in initial traffic); for mature brand advertisements, the threshold is set to 20% (traffic is relatively stable). It also supports manual fine-tuning of the threshold based on historical data.

[0067] Furthermore, more than 10 features are extracted from the click data, including: time features: click period, click interval; user features: IP address, device type, geographical distribution; behavioral features: whether a purchase is generated after the click and the length of stay.

[0068] By using association rule mining algorithms (such as the Apriori algorithm), we can identify high-frequency co-occurring feature combinations in confirmed abnormal data. For example, we found that 80% of confirmed abnormal data have the combined features of "clicks at midnight + multiple clicks from the same IP + no purchase behavior"; we apply this feature combination to the data to be determined: if a piece of data to be determined contains more than 80% of these co-occurring features, it is determined to be confirmed abnormal data. For example, if a piece of data to be determined fluctuates by 60% and meets the features of "clicks at midnight + multiple clicks from the same IP", it is determined to be abnormal data.

[0069] For the evaluation of the degree of abnormality, the number and distribution range of the abnormal data are determined and calculated. The weight coefficient is adjusted based on the advertising budget, with higher budgets resulting in larger coefficients. The final judgment results are presented in a visual report format, including information such as the percentage of abnormal data, key abnormal characteristics, and affected ad placements. The system also automatically sends alert emails to advertisers and recommends adjusting their advertising strategies (such as suspending abnormal placements or switching traffic channels). For example, if an ad's abnormality index reaches 30%, the system recommends temporarily shutting down its delivery on a low-quality channel.

[0070] In summary, through real-time collection and precise comparison with historical data, we can quickly discover abnormal fluctuations in click volume, strictly distinguish between confirmed and pending abnormal data, and use association analysis to mine co-occurring feature combinations, which is conducive to avoiding misjudgments and missed judgments, and significantly improve the accuracy of anomaly identification; the final comprehensive judgment output results provide advertisers with intuitive warnings and optimization suggestions, which is conducive to effectively reducing the cost of invalid advertising, helping advertisers to adjust their advertising strategies in a timely manner, improving the effectiveness and return on investment of advertising, and ensuring a healthy and orderly advertising ecosystem.

[0071] As an optional embodiment, step S4 is specifically as follows:

[0072] Extracting at least two features from the abnormal data to form a feature combination;

[0073] Obtaining abnormal data to be determined and extracting features of the abnormal data to be determined;

[0074] The features of the abnormal data to be determined are matched with the feature combination of the determined abnormal data. If the features of the abnormal data to be determined include all the features in the feature combination of the determined abnormal data, it is determined that the abnormal data to be determined is associated with the determined abnormal data, and the abnormal data to be determined is determined as abnormal data; if the features of the abnormal data to be determined only include some of the features in the feature combination of the determined abnormal data, the matching status of the features of the abnormal data to be determined and other feature combinations of other determined abnormal data is further determined until the association judgment is completed.

[0075] As an optional embodiment, the features include at least two of click time, click device identification, click IP address, and click frequency.

[0076] Specifically, we can obtain the distribution of click periods (such as high-frequency clicks between 0:00 and 6:00 in the morning), click time intervals (such as repeated clicks within every 10 minutes); the number of clicks from the same IP address, the jump depth after clicking (only browsing the homepage without entering the details page), and the length of stay (less than 10 seconds); abnormally concentrated delivery channels (such as a low-quality alliance website), and device type distribution (only low-end Android models have concentrated clicks).

[0077] Among them, a combination optimization algorithm is used to automatically generate feature combinations. 2-3 features are randomly selected from the above feature library to generate basic combinations, such as "clicks in the early morning + more than 10 clicks from the same IP + no product purchase behavior". The frequency of occurrence of each combination in determining abnormal data is calculated, and combinations with an occurrence frequency of less than 30% are eliminated, retaining high-frequency co-occurrence combinations as the core judgment basis.

[0078] After identifying abnormal data, the system initiates a special extraction step, parsing corresponding features from the original click log using regular expressions and data mapping rules. For example, from the log "2025-06-15 03:12:00, IP:192.168.1.10, clicked on Ad A, no redirect," features such as the click time (midnight) and IP address are extracted. These extracted features are then converted to a unified format, such as mapping geographic information to provincial administrative region codes, to ensure the accuracy of subsequent matching.

[0079] Furthermore, the features of the abnormal data to be determined are compared one by one with the feature combination of the confirmed abnormal data. If the data to be determined completely contains all the features of a certain feature combination (such as "clicks at midnight + multiple clicks from the same IP + no purchase behavior"), it is immediately determined to be abnormal data and the associated confirmed abnormal data number is marked. If it only contains some features (such as only "clicks at midnight + no purchase behavior"), it enters the recursive matching stage.

[0080] Furthermore, for partially matched pending data, the following operations are performed to traverse other feature combinations that determine abnormal data and continue matching attempts. For example, if the pending data that was not successfully matched initially contains the "midnight click" feature, other combinations containing this feature will be matched first. When 80% of the features in the feature combination are matched (configurable threshold), it is determined to be abnormal data. If the standard is still not met after traversing all combinations, it is marked as normal data.

[0081] It should be understood that a systematic feature extraction and matching mechanism significantly improves the accuracy and reliability of abnormal ad click detection. On the one hand, the combined extraction of multi-dimensional features can capture the underlying patterns of abnormal clicks, helping to avoid misjudgments or omissions caused by single-feature judgments. On the other hand, recursive matching logic deeply mines data that meets the characteristics, catching no potential anomalies and effectively reducing the risk of misjudgment of ambiguous data. Furthermore, this step is closely linked to the preceding and following steps, automatically optimizing the judgment strategy and forming a closed detection loop. This helps advertisers accurately identify invalid clicks, adjust their advertising strategies in a timely manner, and reduce wasted advertising budgets.

[0082] As an optional embodiment, the present invention further includes:

[0083] Obtaining confirmed abnormal data, performing feature co-occurrence analysis on the confirmed abnormal data, and generating a feature correlation matrix, wherein the element values ​​in the matrix represent the degree of correlation between the features;

[0084] Obtain the abnormal data to be determined, extract its features and perform weighted matching with the feature correlation matrix;

[0085] Calculate a weighted matching score of the features of the to-be-determined abnormal data and the features of the determined abnormal data. If the score reaches a preset matching threshold, determine that the to-be-determined abnormal data is associated with the determined abnormal data, and determine the to-be-determined abnormal data as abnormal data.

[0086] If the score does not reach the threshold, a list of abnormal suspicions is generated based on the missing features for subsequent manual review or automatic completion verification.

[0087] As an optional embodiment, the calculation of the weighted matching score includes:

[0088] Determine the corresponding weight of each feature of the abnormal data to be determined in the feature correlation matrix;

[0089] Calculate the matching score of each feature based on the feature matching situation and corresponding weight;

[0090] The matching scores of all features are accumulated to obtain the weighted total matching score.

[0091] Specifically, by performing a full feature scan on the identified anomaly data output in step S3, the frequency of any two features appearing together in the data set is counted. For example, among 100 identified anomaly data, "clicks at midnight" and "clicks from the same IP address more than 10 times" appear together 85 times, so the co-occurrence frequency is 85;

[0092] Among them, the Pearson correlation coefficient is used to calculate the degree of correlation between features. The formula is: Among them, x and y are binary variables with the occurrence of the feature (1 if it occurs, 0 if it does not occur). The matrix element value range is [-1, 1]. The larger the absolute value, the closer the correlation. For example, the correlation coefficient between "clicking at midnight" and "no purchase behavior" is 0.78, indicating a strong positive correlation. xy Represents the Pearson correlation coefficient between feature x and feature y, which is the final calculated value of the degree of correlation. i and y i , where i represents the i-th sample in the data set, and are the average values ​​of feature x and feature y in all samples, This part is the numerator of the formula, which calculates the covariance between feature x and feature y, reflecting whether the changing trends of the two features are consistent.

[0093] This is the denominator of the formula, which plays a normalization role and ensures the correlation coefficient r xy The value of is between [-1,1]. is the variance of feature x, which measures the degree of dispersion of feature x values ​​relative to its mean. is the variance of feature y. The denominator squares the product of the two feature variances, eliminating the influence of the feature's own fluctuation on the correlation calculation, making the correlation coefficients between different features comparable.

[0094] Furthermore, the generated correlation matrix is ​​displayed as a heat map and stored as structured data (e.g., JSON format) for quick and easy subsequent querying. Correlation pairs with a correlation score above 0.5 in the matrix are marked as "strongly correlated feature combinations" and prioritized for matching.

[0095] Extract features (such as click time, IP address, device type, etc.) from the abnormal data to be determined, and convert non-numeric features into one-hot encoding, for example, "click in the early morning" is converted to [1,0], and "not in the early morning" is converted to [0,1].

[0096] Furthermore, for each feature of the data to be determined, the correlation matrix is ​​queried to obtain its correlation weight with other features. For example, if the data to be determined contains the feature "click in the early morning", then its correlation weight with "no purchase behavior" is 0.78. The weighted sum can be expressed by the formula:

[0097] Where n is the number of features of the data to be determined, m is the number of features to determine abnormal data, represents the number of target features corresponding to each row in the association matrix (that is, the number of target pattern dimensions that a single feature needs to be associated with, where the frequency of occurrence of feature i represents the frequency ratio (or absolute number) of the occurrence of the i-th feature in the data, reflecting the activity level of the feature, and the association matrix i, j represents the element value of the i-th row and j-th column in the association matrix, representing the strength of the association between the i-th feature and the j-th target pattern (such as "no purchase behavior"). The correlation matrix i, j represents the sum of the correlations between the i-th feature and all m target patterns, giving the feature's overall correlation strength. For example, if the data to be determined contains two features, and the sum of their correlations with the features of the identified anomaly data is 0.78 and 0.65, respectively, then the score = 1 × 0.78 + 1 × 0.65 = 1.43.

[0098] If the matching score is greater than or equal to the preset threshold (e.g., 1.2), the data to be determined is determined to be associated with the determined abnormal data and is marked as abnormal data;

[0099] If the score is less than the threshold, the system automatically generates a list of abnormal suspicions, including missing key features (such as "no purchasing behavior") and correlation weights. For example: Suspicion 1: Missing the "no purchasing behavior" feature (correlation 0.78); Suspicion 2: Missing the "multiple clicks from the same IP" feature (correlation 0.65).

[0100] In summary, quantifying feature associations and dynamically weighted matching offers multiple, yet unobvious, advantages. First, by constructing a feature correlation matrix, we can uncover hidden patterns of feature co-occurrence within the data, enabling us to accurately capture complex and ever-changing anomaly patterns. For example, we can discover anomalous click associations within specific time periods, regions, and device type combinations, relationships that are difficult to detect through conventional rules or manual analysis. Second, the weighted matching mechanism, combined with dynamic threshold adjustment, empowers the system with adaptive capabilities, enabling it to automatically learn from emerging anomalous feature combinations, such as quickly identifying unique click behaviors generated by new traffic-boosting tools. Furthermore, the generation of a list of suspected anomalies not only enables efficient human-machine collaboration but, more importantly, drives the evolution of the entire anomaly detection model through continuous feedback optimization, forming a self-perfecting intelligent closed loop that fundamentally enhances the intelligence and long-term effectiveness of ad anomaly detection.

[0101] As an optional embodiment, the present invention further includes:

[0102] Perform time series pattern mining on the click time series in the identified abnormal data to generate a time series abnormal pattern library;

[0103] Build user behavior profiles based on historical user click behavior data and extract normal behavior pattern features;

[0104] Obtain the abnormal data to be determined, determine whether its click time series matches any pattern in the time series abnormal pattern library, and at the same time, compare the characteristics of the abnormal data to be determined with the normal pattern characteristics of the user behavior profile, and identify the feature items that deviate from the normal pattern; if the abnormal data to be determined meets any of the following conditions: the click time series matches the time series abnormal pattern; the number of feature items that deviate from the normal pattern reaches a preset threshold; there are specific key features that deviate from the normal pattern, then the abnormal data to be determined is determined to be abnormal data; otherwise, the unmatched features and deviation analysis results are used as optimization feedback data to update the time series abnormal pattern library and user behavior profile.

[0105] Specifically, after identifying abnormal data, extract the click timestamp of each data item (accurate to the minute) and group them by dimensions such as ad slot ID and ad creative type. For example, organize the click times of the same product ad on different days into independent time series, remove abnormal values ​​(such as click times outside normal business hours), and then perform time normalization to unify the time format.

[0106] The Dynamic Time Warping (DTW) algorithm and frequent pattern mining algorithms (such as PrefixSpan) are used to extract frequently occurring abnormal patterns from time series. Specifically, the DTW algorithm is used to calculate the similarity between different time series and group similar click time distributions into a single category. For example, if a similar pattern of frequent clicks on multiple ads between 2:00 AM and 4:00 AM is found, the PrefixSpan algorithm can be used to mine frequently occurring time subsequences, such as a pattern of concentrated clicks at 10:00 AM for three consecutive days.

[0107] Furthermore, the mined abnormal patterns are sorted by confidence and support. Patterns with a confidence level of 80% or higher and a support level of 50% or higher are retained and stored in a time series abnormal pattern library. Each pattern contains information such as time series features (such as start time, interval period, and click frequency) and the associated ad slot ID. For example, the pattern "between 2:00 AM and 4:00 AM, ad slot A has a click frequency 10 times higher than normal" may be present.

[0108] Furthermore, by integrating at least three months of historical click behavior data of users, including click time, clicked ad type, browsing time, purchase behavior, device information, regional information, etc., for example, a user clicked on beauty ads between 7 and 9 p.m. on average every day for the past three months, and 70% of these clicks resulted in purchases.

[0109] Furthermore, clustering algorithms (such as K-Means) and association rule mining algorithms (such as Apriori) are used to extract user behavior features from the data:

[0110] User segmentation: users are divided into different groups based on characteristics such as click time and ad preferences, such as "nighttime active users" and "high-frequency purchasing users";

[0111] Normal pattern extraction: For each user group, we explore their frequently occurring normal behavior patterns. For example, the normal pattern for the "night-active user" group is: clicking ads between 8 PM and midnight, an average browsing time of 3 minutes, and a purchase conversion rate of 15%;

[0112] Furthermore, the extracted normal behavior pattern features are stored as a user behavior profile feature library. Each feature contains information such as group labels, time features, behavior features, and associated weights. For example, the weight of the "nighttime click period" feature of "nighttime active users" is set to 0.8, indicating that this feature has a high reference value in determining whether user behavior is normal;

[0113] Furthermore, the DTW algorithm is used to calculate the similarity between the time series of the data to be determined and each pattern in the pattern library. If the similarity is ≥ 0.7, it is considered a match. For example, if a certain data to be determined has a concentrated click between 3 and 5 am, and the similarity with the "abnormal click pattern in the early morning" in the pattern library reaches 0.8, it will trigger an abnormality judgment;

[0114] Compare the characteristics of the data to be determined (such as click time, ad type, and purchase behavior) with the normal patterns in the user behavior profile feature library and calculate the feature deviation;

[0115] For each feature, calculate its difference from the normal mode, such as the deviation of the actual click time from the normal period, the difference between the purchase conversion rate and the group mean

[0116] When the feature deviation exceeds a preset threshold (such as the click time deviation exceeds 2 hours, the purchase conversion rate is lower than the average by 50%), it is marked as a feature item that deviates from the normal pattern.

[0117] Abnormal data is determined to be abnormal data when it meets any of the following conditions: the click time series matches any pattern in the time series abnormal pattern library; the number of feature items that deviate from the normal pattern reaches the preset threshold (such as 3 or more); there are specific key features that deviate from the normal pattern (such as a user with a high purchase conversion rate suddenly clicks 10 times in a row without purchasing behavior).

[0118] If the above conditions are not met, the unmatched time series features, deviated behavioral features and analysis results will be used as optimization feedback data to automatically update the time series anomaly pattern library and user behavior portrait feature library, such as incorporating newly discovered edge anomaly patterns into the pattern library, or adjusting the normal behavior feature threshold of the user group.

[0119] The above steps break through the limitations of traditional anomaly detection that relies solely on data fluctuations and feature associations, and help achieve in-depth mining and accurate identification of abnormal data from the perspective of time series patterns and user behavior patterns. Through time series pattern mining, hidden abnormal behaviors such as periodic timed brushing and regular click intervals can be captured, which are difficult to detect with isolated click data alone; by building a profile based on historical user behavior and analyzing deviation characteristics, abnormal clicks disguised as normal user behavior can be identified, such as low-frequency abnormal clicks that suddenly appear from highly active users. In addition, the dynamic update mechanism driven by optimized feedback data gives the system the ability to self-evolve and automatically adapt to changes in advertising delivery scenarios, which is conducive to predicting new abnormal patterns in advance. While improving the accuracy of anomaly identification, it helps to provide forward-looking reference for optimizing advertising delivery strategies, creating value far beyond traditional detection methods.

[0120] As an optional embodiment, the present invention further includes:

[0121] Integrate cross-domain data from advertising delivery platforms, including but not limited to ad creative types, delivery channel data, and user conversion behavior data;

[0122] Build a cross-domain correlation graph based on graph neural networks to analyze the potential correlation paths between the abnormal data to be determined and the cross-domain data;

[0123] If it is found that the abnormal data to be determined overlaps with the known abnormal pattern in a cross-domain correlation path, or its correlation path causes an abnormal interruption in the advertising conversion chain, it will be determined as abnormal data;

[0124] Otherwise, the cross-domain association analysis results are included in the optimization feedback data to update the association weights and abnormal path rules of the cross-domain association graph.

[0125] As an optional embodiment, the construction of the cross-domain association graph includes:

[0126] Extract entity data of ad click data, material data, and channel data;

[0127] Mining relationships between entities through association rules and assigning dynamic weights to relationship edges;

[0128] Graph embedding technology is used to convert the graph into a feature vector for rapid retrieval of abnormal association paths.

[0129] Specifically, data is connected to advertising platforms (such as Douyin's Bytedance and Tencent Advertising), user behavior analysis tools (such as Mixpanel), and e-commerce transaction systems (such as Taobao's backend) through API interfaces to collect data including but not limited to advertising creative types, distribution channel data, user conversion behavior data, etc.

[0130] Among them, for key fields such as advertising exposure, multiple filling methods are used (such as random forest prediction of missing values).

[0131] Among them, the timestamps of different platforms are converted into UTC standard time, and the regional information is uniformly mapped into provincial administrative codes.

[0132] Among them, sensitive information such as user ID and mobile phone number are hashed and encrypted.

[0133] The entire advertising delivery process is abstracted into a directed attribute graph, with the following node types: creative node (including creative ID, type, and copy); channel node (including channel ID and traffic source); user behavior node (behavior records such as clicks, add-to-cart, and order placement); time node (time window divided by hours / days); edge type: delivery relationship (creative → delivery channel); behavior association (user click → add-to-cart → order placement); time association (behavior node → time node).

[0134] Among them, association rules are a method for mining hidden associations between entities in data. The core is to analyze the co-occurrence patterns of data items, for example, expressed in the form of "X→Y", where X and Y are non-overlapping entity sets (such as goods, behaviors, attributes, etc.), indicating that "when X appears, Y has a certain probability of appearing", and preset them in the corresponding database.

[0135] Furthermore, the model uses a graph convolutional network (GCN) and a graph attention network (GAT) to encode node attributes (such as ad copy keywords and channel traffic characteristics) into vector representations. Through training on historical normal and abnormal data, the model learns the strength of associations between different nodes. For example, if a channel is frequently associated with high-conversion advertising creatives, the edge weight between the two is higher. The graph structure and weights are incrementally updated every morning based on the latest data to adapt to changes in advertising delivery strategies.

[0136] Furthermore, for the undetermined abnormal data output by step S3, the system uses the Monte Carlo Tree Search (MCTS) algorithm to find associated paths in the graph. The nodes such as the advertising ID and user click behavior of the undetermined data will be used as the search starting point. According to the principle of priority of associated weights, the connection paths with other nodes will be explored, such as "click behavior → delivery channel → similar advertising materials → conversion link". If the search path overlaps with the known abnormal patterns in the graph (such as false clicks caused by a combination of specific channels and low-quality materials) by more than 70%, it is determined to be abnormal. When the associated path shows an abnormal breakpoint in the advertising conversion link (such as no add-to-cart behavior after a large number of clicks), and the node association weight at the breakpoint is higher than the threshold (such as 0.8), it is determined to be abnormal. For example, if the number of clicks on an advertisement on the Douyin channel surges, but there is no corresponding traffic import from the subsequent Taobao store, an abnormal alarm will be triggered.

[0137] Furthermore, if the data to be determined is not judged to be abnormal, its cross-domain correlation analysis results will be used as optimization feedback data to supplement new abnormal path rules, such as incorporating the "high click and low conversion of a certain advertising material in a specific time period + specific channel" pattern into the monitoring scope.

[0138] It should be noted that the above examples are provided for the convenience of understanding the technical solution and have no special meaning.

[0139] Existing methods for identifying abnormal clicks in advertising often focus on analyzing the characteristics of click data itself, such as click frequency and device ID. This solution, however, breaks through the limitations of traditional anomaly detection by leveraging cross-domain data correlation analysis to uncover hidden anomalies across the entire advertising delivery chain, delivering multi-dimensional, yet unnoticeable, value. Firstly, by integrating heterogeneous cross-domain data, including ad creatives, delivery channels, and user conversions, and constructing a dynamic correlation graph, it can uncover deep connections that are difficult to detect with isolated data. For example, this can accurately identify hidden anomaly patterns caused by coordinated "brush traffic networks" across different platforms or the combination of specific channels and low-quality creatives. Secondly, the path analysis and weight learning mechanism based on graph neural networks can adaptively capture abnormal breakpoints in the ad conversion chain, such as the phenomenon of seemingly high clicks but no actual conversions caused by fake traffic. Furthermore, by optimizing feedback-driven graph and dynamic rule updates, the system is self-evolving, enabling it to predict new anomaly patterns and also guiding advertising delivery strategy optimization. This improves anomaly detection accuracy while creating dual value from risk prevention to business growth, enhancing identification accuracy.

[0140] In the second aspect, the present invention also proposes an analysis and anomaly identification system for advertisement click traffic, the method includes the following steps:

[0141] Data acquisition module, which obtains daily advertising click data within a preset time period in real time;

[0142] The data comparison module determines the daily ad click data within the current preset time period and compares it with the daily ad click data of the same time period in the historical data;

[0143] A judgment module, based on the comparison results, marks the amount of ad click data whose fluctuation exceeds a preset threshold as abnormal data. Abnormal data includes confirmed abnormal data and pending abnormal data. Among them, data that meets all abnormal feature judgment criteria is confirmed abnormal data, and data that only meets some abnormal feature judgment criteria is pending abnormal data;

[0144] The association analysis module performs association analysis on confirmed abnormal data and undetermined abnormal data, extracts the co-occurring feature combinations between the two, and analyzes the undetermined abnormal data based on the co-occurring feature combinations to determine whether the undetermined data is confirmed abnormal data;

[0145] The result output module integrates all the confirmed abnormal data to obtain the final judgment result of abnormal ad clicks.

[0146] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A method for analyzing and identifying anomalies in advertisement click traffic, characterized in that: include: S1. Obtain daily ad click data within a preset time period in real time; S2. Determine the daily ad click volume data for the current preset time period and compare it with the daily ad click volume data for the same time period in historical data; S3. Based on the comparison results, the amount of ad click data whose fluctuation exceeds a preset threshold is marked as abnormal data. The abnormal data includes confirmed abnormal data and pending abnormal data. Among them, the abnormal data that meets all the abnormal feature judgment criteria is confirmed abnormal data, and the abnormal data that only meets some of the abnormal feature judgment criteria is pending abnormal data. S4, performing correlation analysis on the determined abnormal data and the abnormal data to be determined, extracting co-occurring feature combinations between the two, analyzing the abnormal data to be determined based on the co-occurring feature combinations, and determining whether the abnormal data to be determined is determined abnormal data; S5. Combining all confirmed abnormal data, a final determination result of abnormal ad clicks is obtained.

2. The method for analyzing and identifying anomalies of advertisement click traffic according to claim 1, characterized in that: The abnormal feature judgment standard includes at least one of high-frequency clicks, false device IDs, and abnormal click time periods.

3. The method for analyzing and identifying anomalies of advertisement click traffic according to claim 1, characterized in that: The S4 step is specifically as follows: Extracting at least two features from the determined abnormal data to form a feature combination; Obtaining abnormal data to be determined and extracting features of the abnormal data to be determined; The features of the abnormal data to be determined are matched with the feature combination of the determined abnormal data. If the features of the abnormal data to be determined include all the features in the feature combination of the determined abnormal data, it is determined that the abnormal data to be determined is associated with the determined abnormal data, and the abnormal data to be determined is determined as abnormal data; if the features of the abnormal data to be determined only include some of the features in the feature combination of the determined abnormal data, the matching status of the features of the abnormal data to be determined and other feature combinations of other determined abnormal data is further determined until the association judgment is completed.

4. The method for analyzing and identifying anomalies of advertisement click traffic according to claim 3, characterized in that: The features include at least two of click time, click device identification, click IP address, and click frequency.

5. The method for analyzing and identifying anomalies of advertisement click traffic according to claim 4, characterized in that: Also includes: Acquire determined abnormal data, perform feature co-occurrence analysis on the determined abnormal data, and generate a feature correlation matrix, wherein element values ​​in the matrix represent the degree of correlation between features; Acquire the abnormal data to be determined, extract its features and perform weighted matching with the feature correlation matrix; Calculate a weighted matching score of the features of the to-be-determined abnormal data and the features of the determined abnormal data. If the score reaches a preset matching threshold, determine that the to-be-determined abnormal data is associated with the determined abnormal data, and determine the to-be-determined abnormal data as abnormal data. If the score does not reach the threshold, a list of abnormal suspicions is generated based on the missing features for subsequent manual review or automatic completion verification.

6. The method for analyzing and identifying anomalies in advertisement click traffic according to claim 5, characterized in that: The calculation of the weighted matching score includes: Determine the corresponding weight of each feature of the abnormal data to be determined in the feature correlation matrix; Calculate the matching score of each feature based on the feature matching situation and corresponding weight; The matching scores of all features are accumulated to obtain the weighted total matching score.

7. The method for analyzing and identifying anomalies of advertisement click traffic according to claim 6, characterized in that: Also includes: Perform time series pattern mining on the click time series in the identified abnormal data to generate a time series abnormal pattern library; Build user behavior profiles based on historical user click behavior data and extract normal behavior pattern features; Obtain the abnormal data to be determined, determine whether its click time series matches any pattern in the time series abnormal pattern library, and at the same time, compare the characteristics of the abnormal data to be determined with the normal pattern characteristics of the user behavior profile to identify the feature items that deviate from the normal pattern; the abnormal data to be determined meets any of the following conditions: the click time series matches the time series abnormal pattern; the number of feature items that deviate from the normal pattern reaches a preset threshold; If there are specific key features that deviate from the normal pattern, the data to be determined as abnormal is determined to be abnormal data; Otherwise, the unmatched features and deviation analysis results are used as optimization feedback data to update the time series anomaly pattern library and user behavior profiles.

8. The method for analyzing and identifying anomalies of advertisement click traffic according to claim 7, characterized in that: Also includes: Integrate cross-domain data from advertising delivery platforms, including but not limited to ad creative types, delivery channel data, and user conversion behavior data; Build a cross-domain correlation graph based on graph neural networks to analyze the potential correlation paths between the abnormal data to be determined and the cross-domain data; If it is found that the abnormal data to be determined overlaps with the known abnormal pattern in a cross-domain correlation path, or its correlation path causes an abnormal interruption in the advertising conversion chain, it will be determined as abnormal data; Otherwise, the cross-domain association analysis results are included in the optimization feedback data to update the association weights and abnormal path rules of the cross-domain association graph.

9. The method for analyzing and identifying anomalies of advertisement click traffic according to claim 8, characterized in that: The construction of the cross-domain association graph includes: Extract entity data of ad click data, material data, and channel data; Mining relationships between entities through association rules and assigning dynamic weights to relationship edges; Graph embedding technology is used to convert the graph into a feature vector for rapid retrieval of abnormal association paths.

10. An advertising click traffic analysis and anomaly identification system, applicable to the advertising click traffic analysis and anomaly identification method according to any one of claims 1 to 9, characterized in that: The system comprises: Data acquisition module, which obtains daily advertising click data within a preset time period in real time; The data comparison module determines the daily ad click data within the current preset time period and compares it with the daily ad click data of the same time period in the historical data; A judgment module, based on the comparison result, marks the amount of advertising click data whose fluctuation exceeds a preset threshold as abnormal data, wherein the abnormal data includes confirmed abnormal data and pending abnormal data. Among them, the abnormal data that meets all the abnormal feature judgment criteria is confirmed abnormal data, and the abnormal data that meets only some of the abnormal feature judgment criteria is pending abnormal data; An association analysis module performs association analysis on the confirmed abnormal data and the abnormal data to be determined, extracts a co-occurring feature combination between the two, analyzes the abnormal data to be determined based on the co-occurring feature combination, and determines whether the abnormal data to be determined is confirmed abnormal data; The result output module integrates all the confirmed abnormal data to obtain the final judgment result of abnormal ad clicks.

Citation Information

Patent Citations

  • Advertisement scalping monitoring identification method and system and readable storage medium

    CN117217830A

  • Internet advertisement identifier distribution platform system and method thereof

    CN118485477A

  • Abnormal request detection method and device, computer equipment and storage medium

    CN118740662A

  • Accurate advertisement putting optimization method and system

    CN119624543A

  • Accurate advertisement putting method based on user behavior pattern mining and clustering analysis

    CN120125300A

Cited By

  • Internet advertisement abnormal flow detection method and detection system

    CN121481640A