Mobile advertisement click rate abnormal data detection method and system
By constructing a multi-dimensional temporal feature system and dynamic threshold adaptation, the problems of misjudgment and missed judgment in the detection of abnormal mobile advertising click volume in the existing technology are solved, and more efficient abnormal data identification is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU HAIYE NETWORK TECHNOLOGY CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for detecting abnormal mobile ad clicks cannot effectively identify complex fraudulent behaviors, and fixed thresholds cannot adapt to normal fluctuations in ad placement, leading to misjudgments and missed detections.
Construct a multi-dimensional time-series feature system, dynamically adapt the anomaly threshold to the business scenario, output a comprehensive anomaly score through a preset anomaly detection model, generate a dynamic anomaly threshold, and perform anomaly detection in conjunction with related business scenario data.
It improves the reliability and robustness of abnormal data detection, reduces the risk of misjudgment and missed judgment caused by traffic fluctuations, and improves the accuracy of abnormal click data detection.
Smart Images

Figure CN121935779A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of internet advertising technology, and in particular to a method and system for detecting abnormal mobile advertising click data. Background Technology
[0002] Currently, mobile advertising has become the core form of advertising in the digital marketing field, and is widely used in various scenarios such as news feeds, splash screens, and app stores. Click volume is a key indicator for measuring the effectiveness of advertising and settling advertising fees. However, due to the influence of malicious fraud in mobile advertising, a large number of fake click data are prone to appear, which affects the evaluation of advertising effectiveness. Therefore, it is necessary to detect abnormal click volume data.
[0003] Existing methods for detecting abnormal click data mostly rely on traditional statistical methods, which identify click data that exceeds the normal fluctuation range by setting a fixed threshold. This method only focuses on the macro-time series changes in click volume and cannot identify complex fraudulent behaviors disguised as normal clicks. At the same time, the fixed threshold cannot adapt to the normal fluctuation scenarios of advertising. For example, legitimate click peaks caused by promotional activities and hot events are easily misjudged as abnormal, while hidden fraudulent behaviors below the threshold will be missed. Summary of the Invention
[0004] The purpose of this application is to provide a method and system for detecting abnormal mobile advertising click volume data. By constructing a multi-dimensional time-series feature system and dynamically adapting the abnormal threshold in conjunction with business scenarios, the reliability of abnormal detection is improved.
[0005] Firstly, this application provides a method for detecting abnormal mobile ad click volume data, including: Acquire raw click data stream and related business scenario data of the target mobile ad within a preset time period; Feature extraction is performed on the raw click data stream to obtain a multi-dimensional time-series feature set, which includes click volume time-series features, click behavior features, and contextual features. Set a detection time window, and based on a multi-dimensional time series feature set, output a comprehensive anomaly score for each detection time window through a preset anomaly detection model; Based on data from related business scenarios, dynamic anomaly thresholds are obtained by pre-setting historical calibration data; The overall anomaly score is compared with the dynamic anomaly threshold to determine whether there is abnormal data. If so, an anomaly detection report is generated based on the abnormal data.
[0006] By integrating click volume time-series features, click behavior features, and contextual features, the core characteristics of fraudulent patterns are fully covered, improving the comprehensiveness of anomaly identification. At the same time, by combining business scenarios and generating dynamic anomaly thresholds, the risk of misjudgment / missed judgment caused by traffic fluctuations can be reduced to a certain extent, which helps to improve the reliability and robustness of click volume anomaly data detection.
[0007] Optionally, the original click data stream includes device ID, IP address, ad ID, and channel ID. The step of extracting features from the original click data stream to obtain a multi-dimensional temporal feature set includes: Based on the raw click data stream, the data is divided into multiple fixed time windows according to a preset time granularity. The click volume of each fixed time window is statistically analyzed to obtain basic statistical characteristics. Based on the raw click data stream, multiple sliding time windows are obtained according to a preset time step. The click volume of each sliding time window is statistically analyzed to obtain trend fluctuation characteristics. The basic statistical features and trend fluctuation features are integrated and denoted as the click volume time series features; Using device ID and ad ID as the behavioral units, click data for each behavioral unit is obtained in each fixed time window, and click behavior characteristics are obtained through data statistics; Using IP address and channel ID as the association unit, click data of the association unit in each fixed time window is obtained, and context association features are obtained through data statistics; Optionally, the preset anomaly detection model includes four types of single detection models representing different quantification dimensions. The setting of detection time windows, based on a multi-dimensional temporal feature set, outputs a comprehensive anomaly score for each detection time window through the preset anomaly detection model, including: Based on a multi-dimensional temporal feature set, the features are aggregated according to the detection time window to generate a structured feature matrix. Based on the structured feature matrix, detection is performed using four types of single detection models, and single model scores are obtained. Based on the detection time window, the corresponding business scenario is matched from the associated business scenario data; Based on business scenarios, dynamic weights are generated for four types of single detection models through a pre-set scenario model mapping library. For the scores of individual models, a weighted calculation is performed by dynamically allocating weights to obtain a comprehensive anomaly score.
[0008] Optionally, based on business scenarios, the step of generating dynamically assigned weights for four types of single detection models through a preset scenario model mapping library includes: The initial weights of four single detection models are obtained by using a pre-set model weight database. Based on business scenarios, the impact coefficients of business scenarios on four types of single detection models are obtained through a pre-set scenario model mapping library. Based on the influence coefficient, the initial allocation weights are adjusted to generate dynamic allocation weights.
[0009] Optionally, the step of obtaining dynamic anomaly thresholds based on associated business scenario data and preset historical calibration data includes: Based on the associated business scenario data, determine the business scenario corresponding to the detection time window, and denot it as the target business scenario; By using historical calibration data, the calibration data corresponding to the detection time window is determined and recorded as the sample calibration data; Based on the sample calibration data, the basic anomaly threshold is determined by a preset quantitative statistical method. Based on the target business scenario, the corresponding adjustment coefficient is matched through a preset scenario coefficient adjustment library; Based on the basic anomaly threshold and the adjustment coefficient, the dynamic anomaly threshold is calculated and obtained.
[0010] Optionally, generating an anomaly detection report based on the abnormal data includes: Determine the detection time window where the abnormal data is located, and based on the detection time window, filter and obtain the corresponding multi-dimensional time series feature set, which is denoted as the abnormal feature subset; Based on the comprehensive anomaly score and dynamic anomaly threshold, the anomaly deviation magnitude is calculated, and the baseline confidence level is calculated based on the anomaly deviation magnitude. Based on a subset of abnormal features, the abnormal contribution features and their corresponding significance coefficients are obtained through a pre-defined method. Based on the abnormal contribution characteristics and the corresponding significance coefficients, the basic confidence level is corrected to obtain the final confidence level; Encode the subset of abnormal features to generate an abnormal feature vector; Based on the abnormal feature vector, similarity matching is performed through a pre-set abnormal feature pattern library to determine the abnormal type; The detection time window, final confidence level, and anomaly type are integrated to generate an anomaly detection report.
[0011] Optionally, the step of obtaining the abnormal contribution features and corresponding significance coefficients based on an abnormal feature subset using a preset method includes: Based on a subset of abnormal features, abnormal contribution features and their corresponding contribution degrees are obtained through a preset method. Normalize the contribution of abnormal contribution characteristics to generate contribution weights; The degree of abnormal deviation of abnormal contribution characteristics is calculated by using preset benchmark reference data; A significance coefficient is generated based on the contribution weight and the degree of abnormal deviation.
[0012] Optionally, the anomaly detection report includes a confidence level. After generating the anomaly detection report based on the anomaly data, it includes: If the confidence level does not reach the preset confidence threshold, the anomaly detection report will be pushed to manual review and feedback results will be received. The feedback results include misjudgment and confirmed anomalies. If the feedback result is a misjudgment, the first prompt message will be output to indicate that the abnormal data was judged incorrectly.
[0013] Secondly, this application provides a mobile advertising click volume anomaly data detection system, comprising: The data acquisition module 101 is used to acquire the raw click data stream and related business scenario data of the target mobile advertisement within a preset time period; The temporal feature extraction module 102 is used to extract features from the original click data stream and obtain a multi-dimensional temporal feature set, which includes click volume time series features, click behavior features, and contextual association features. The anomaly detection module 103 is used to set the detection time window, and based on a multi-dimensional time series feature set, outputs a comprehensive anomaly score for each detection time window through a preset anomaly detection model. The dynamic threshold generation module 104 is used to obtain dynamic anomaly thresholds based on related business scenario data and preset historical calibration data. The detection result generation module 105 is used to compare the comprehensive anomaly score with the dynamic anomaly threshold to determine whether there is abnormal data. If so, an anomaly detection report is generated based on the abnormal data.
[0014] Thirdly, this application provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in the method for detecting abnormal mobile advertising click volume data.
[0015] In summary, this application firstly improves the reliability and robustness of detecting abnormal mobile ad click volume data by constructing a multi-dimensional temporal feature system and dynamically adjusting the anomaly judgment threshold in conjunction with business scenarios. Secondly, by dynamically adjusting the weight of a single detection model based on business scenarios to adapt to the traffic characteristics of different scenarios, it can effectively solve the problem of misjudgment / missed judgment caused by the poor adaptability of fixed weights. Furthermore, by establishing a linkage mechanism between manual review and detection method optimization, the model and parameters can be continuously optimized based on feedback results, which helps to continuously improve the detection accuracy of abnormal click volume data. Attached Figure Description
[0016] Figure 1 This is a flowchart of a method for detecting abnormal mobile advertising click volume provided in an embodiment of this application; Figure 2This is a flowchart provided in an embodiment of the present application for extracting features from the original click data stream to obtain a multi-dimensional temporal feature set; Figure 3 This is a flowchart provided in an embodiment of the present application, which outputs a comprehensive anomaly score for each detection time window using a preset anomaly detection model; Figure 4 This is a flowchart provided in this application embodiment for obtaining dynamic anomaly thresholds based on associated business scenario data and through preset historical calibration data; Figure 5 This is a schematic diagram of a mobile advertising click volume abnormal data detection system provided in an embodiment of this application. Detailed Implementation
[0017] The following is in conjunction with the appendix Figure 1 -Appendix Figure 5 This application will be described in further detail below.
[0018] This application provides a method for detecting abnormal click-through rates in mobile advertising; see [link to relevant documentation]. Figure 1 This includes the following steps: S100: Obtain the raw click data stream and related business scenario data of the target mobile advertisement within a preset time period.
[0019] S200. Extract features from the original click data stream to obtain a multi-dimensional temporal feature set.
[0020] S300: Set the detection time window, and based on the multi-dimensional time series feature set, output the comprehensive anomaly score for each detection time window through the preset anomaly detection model.
[0021] S400: Based on data from related business scenarios, dynamic anomaly thresholds are obtained by pre-setting historical calibration data.
[0022] S500: Compare the comprehensive anomaly score with the dynamic anomaly threshold to determine if there is any abnormal data. If so, generate an anomaly detection report based on the abnormal data.
[0023] In this embodiment of the application, the original click data stream and related business scenario data of the target mobile advertisement within a preset time period are first obtained.
[0024] Among them, the target mobile ad is the mobile ad currently to be detected; since ad click data is time-series data, the goal of anomaly detection is to determine whether the click behavior in a certain period of time deviates from the normal pattern, so a time period is set, that is, the click data of the target mobile ad within the preset time period is obtained and recorded as the raw click data stream.
[0025] The raw click data stream contains basic data such as device ID, IP address, ad ID, channel ID, and click timestamp.
[0026] Related business scenario data refers to a structured data set that is directly related to the entire process of target mobile advertising and is used to characterize the advertising environment, business attributes, channel characteristics, and promotion goals. Specifically, it includes scenario type (such as daily advertising, major promotional events, and new channel promotion), advertising region, target user group, and advertising format (such as splash screen ads and feed ads).
[0027] After obtaining the raw click data stream, abnormal data detection will be performed on the raw click data stream. First, feature extraction will be performed on the raw click data stream to obtain a multi-dimensional time series feature set.
[0028] The multi-dimensional time-series feature set includes click volume time series features, click behavior features, and contextual features.
[0029] Click volume time series features are aggregated features based on time granularity, reflecting the overall trend and fluctuation of ad click volume over time, and are the core of identifying traffic mutations; click behavior features are time-series behavioral features based on individual dimensions such as device / IP, reflecting the click behavior patterns of a single entity (device, IP), and are the core of identifying machine-generated clicks and mass fraud; contextual association features are time-series association features based on the intersection of multiple dimensions, reflecting the relationship between different entities, and are the core of identifying cluster fraud (such as multiple devices under the same IP, or a single IP under the same channel).
[0030] Specifically, see Figure 2 The original click data stream is subjected to feature extraction to obtain a multi-dimensional temporal feature set, including the following steps: S210. Based on the original click data stream, divide it into multiple fixed time windows according to a preset time granularity, perform data statistics on the click volume of each fixed time window, and obtain basic statistical characteristics.
[0031] S220. Based on the original click data stream, multiple sliding time windows are obtained according to a preset time step. The click volume of each sliding time window is statistically analyzed to obtain trend fluctuation characteristics.
[0032] S230. Integrate the basic statistical features with the trend fluctuation features and denote them as the click volume time series features.
[0033] S240. Using device ID and ad ID as the action unit, obtain the click data of the action unit in each fixed time window, and obtain the click behavior characteristics through data statistics.
[0034] S250. Using IP address and channel ID as the association unit, obtain the click data of the association unit in each fixed time window, and obtain the context association features through data statistics.
[0035] First, the extraction of click volume time series features is based on the original click data stream, which is divided into multiple fixed time windows according to a preset time granularity. The preset time granularity, such as 1 minute or 5 minutes, can be adjusted according to the density of ad placement. The fixed time windows do not overlap and are used to provide basic data anchors to ensure that the feature statistics of each time interval are unique and stable.
[0036] By statistically analyzing the click volume for each fixed time window, basic statistical characteristics can be obtained, including the mean, median, maximum, minimum, variance, and standard deviation of the window's click volume.
[0037] Next, based on the original click data stream, multiple sliding time windows are obtained according to a preset time step. The preset time step can be set to, for example, 30 seconds or 1 minute. The time span of the sliding time window is consistent with that of the fixed time window. Adjacent sliding windows have time overlap, which is used to provide trend supplementation and solve the defect of the fixed window that cannot capture gradual anomalies due to boundary issues.
[0038] By statistically analyzing the click volume of each sliding time window, trend fluctuation characteristics can be obtained, including the click volume growth rate between windows, sliding standard deviation, number of consecutive increases / decreases, and frequency of peak occurrences.
[0039] By integrating basic statistical features with trend fluctuation features, a click volume time series feature is formed, which is used to reflect the temporal distribution and fluctuation pattern of click volume.
[0040] For the extraction of click behavior features, the device ID and ad ID are used as the behavior unit (i.e., the set of click behaviors of the same device on the same ad). Click data of the behavior unit in each fixed time window is obtained, and click behavior features are obtained through data statistics. Click behavior characteristics include click frequency, average click interval, shortest click interval, click concentration during different time periods, and click persistence across fixed windows (e.g., clicks occur in three consecutive fixed time windows).
[0041] For the extraction of contextual features, IP address and channel ID are used as the association unit (i.e., the set of click behaviors of the same IP on the same channel). Click data of the association unit in each fixed time window is obtained, and contextual features are obtained through data statistics.
[0042] Contextual features include: single-channel click share (the percentage of total clicks for this IP in the channel), IP geographic concentration (the degree of overlap in the geographical locations of devices corresponding to the same IP), channel click deviation (the difference between the current window's channel clicks and the historical average for the same period), and number of IP-device associations (the number of different device IDs bound to the same IP).
[0043] By integrating click volume time series features, click behavior features, and contextual features, a multi-dimensional time series feature set can be formed.
[0044] Once the multidimensional time series feature set is determined, abnormal data detection can be performed on the multidimensional time series feature set. That is, a detection time window is set, and based on the multidimensional time series feature set, an abnormal detection model is used to output the comprehensive abnormal score for each detection time window.
[0045] Specifically, see Figure 3 The detection time window is set, and based on a multi-dimensional temporal feature set, a pre-set anomaly detection model is used to output a comprehensive anomaly score for each detection time window, including the following steps: S310. Based on a multi-dimensional temporal feature set, aggregate the features according to the detection time window to generate a structured feature matrix.
[0046] S320, based on the structured feature matrix, performs detection using four types of single detection models and obtains single model scores.
[0047] S330: Based on the detection time window, match the corresponding business scenario from the associated business scenario data.
[0048] S340. Based on business scenarios, dynamically assign weights to four types of single detection models through a preset scenario model mapping library.
[0049] S350. For the single model score, a weighted calculation is performed by dynamically allocating weights to obtain a comprehensive anomaly score.
[0050] The preset anomaly detection model is not a single algorithm, but a set of models generated by selecting data in advance and training it according to business needs to identify abnormal patterns in multi-dimensional time-series features. The set of models takes the "feature vector of the detection time window" as input, quantifies the degree of anomaly in each detection time window, and takes the "single model anomaly score" as the direct output. Finally, a comprehensive anomaly score is obtained by fusion.
[0051] Specifically, the preset anomaly detection model includes four types of single detection models, namely statistical models, such as using 3 Models, based on the assumption of normal distribution, identify abnormal data that exceeds the normal fluctuation range; unsupervised clustering models, such as using the Isolation Forest model, determine anomalies by identifying "isolated samples" in a multidimensional feature space; unsupervised reconstruction models, such as using the autoencoder model, calculate reconstruction error by learning the feature patterns of normal data; supervised classification models, such as using the XGBoost model, are trained based on historical cheating / normal samples and output a probability score that a sample belongs to an anomaly.
[0052] Based on the structured feature matrix, detection is performed using four types of single detection models, which can output an anomaly score between 0 and 1, that is, a single model score. The higher the score, the greater the probability of an anomaly.
[0053] First, based on the multi-dimensional temporal feature set, it is aggregated according to the detection time window to generate a structured feature matrix. That is, the click volume time series features, click behavior features, and context association features corresponding to each detection time window are arranged in chronological order, with the detection time window as the row and the multi-dimensional temporal features as the column to form a structured feature matrix.
[0054] Different business scenarios may directly affect the anomaly detection accuracy of a single model. This is because different business scenarios correspond to different traffic distribution characteristics and cheating methods, and different single detection models have different underlying principles and recognition logic, resulting in significant differences in feature adaptability to different scenarios.
[0055] Therefore, in this embodiment of the application, the corresponding business scenario will also be matched from the associated business scenario data based on the detection time window. In this way, the business scenario can provide a scenario-based adaptation basis for the anomaly detection model and eliminate the normal fluctuation differences under different advertising scenarios.
[0056] Once the business scenario is determined, dynamic weights can be generated for the four types of single detection models based on the business scenario and through a preset scenario model mapping library.
[0057] Specifically, based on business scenarios, and through a pre-defined scenario model mapping library, dynamic weights are generated for four types of single detection models, including the following steps: S341. Obtain the initial weight allocation for the four types of single detection models through the preset model weight database.
[0058] S342. Based on business scenarios, obtain the impact coefficients of business scenarios on four types of single detection models through a preset scenario model mapping library.
[0059] S343. Based on the influence coefficient, adjust the initial allocation weights to generate dynamic allocation weights.
[0060] First, the initial weights of the four single detection models will be obtained through a pre-set model weight database.
[0061] The preset model weight database stores the initial weights of four types of single detection models. The initial weights are set based on the actual performance indicators of the four single detection models in normal daily scenarios, and can be continuously updated based on actual detection data.
[0062] First, by using a pre-set model weight database, the initial weights for the four types of single detection models can be obtained. Let the initial weights be denoted as... , =1,2,3,4 .
[0063] Then, based on the business scenario, the influence coefficient of the business scenario on the four types of single detection models is obtained through the preset scenario model mapping library. The influence coefficient is [k1, k2, k3, k4], and the value range is set, for example, [0.5, 1.5], to avoid the influence coefficient being too small, causing the weight of a certain model to approach 0, or too large, causing the weight to be over-concentrated.
[0064] The pre-defined scenario model mapping library stores the impact coefficients of different business scenarios on four types of single detection models. The impact coefficients are used to quantify the objective impact of business scenarios on model performance. For example, in daily delivery scenarios, statistical models have high adaptability, with an impact coefficient of [1.2, 1.0, 1.0, 0.8]; in promotional event scenarios, unsupervised clustering models have high adaptability, with an impact coefficient of [1.0, 1.3, 1.0, 0.7].
[0065] Next, based on the influence coefficient, the initial allocation weights are adjusted to generate dynamic allocation weights. Let the dynamic allocation weights be denoted as... ,but It can be represented as: In addition, to prevent misjudgments caused by a single model, weight constraint rules will be added. For example, the dynamic weight of a single model will not exceed 0.4 at the upper limit and will not be lower than 0.05 at the lower limit. Any excess weight will be distributed to other models proportionally.
[0066] Once the dynamic weight allocation is determined, the scores of individual models can be weighted and calculated using the dynamic weight allocation to obtain the comprehensive anomaly score.
[0067] Let the single model score be... Then the comprehensive abnormality score S can be expressed as: The comprehensive anomaly score S ranges from 0 to 1. The higher the score, the greater the likelihood of anomalies in the click data during that time window.
[0068] After obtaining the comprehensive anomaly score, it needs to be compared with the set anomaly threshold to determine whether it is abnormal data. Similarly, considering that the distribution or manifestation of abnormal data may differ in different business scenarios, a fixed anomaly threshold cannot adapt to differences in business scenarios and fluctuations in advertising traffic.
[0069] For example, in normal daily scenarios, traffic is stable (e.g., daily average click volume fluctuates by ±10%), and cheating often manifests as "sudden traffic surges" (machine-generated traffic). Using a fixed threshold can effectively identify these abnormal surges. However, during major promotional events, normal traffic may surge. If the same fixed threshold is used, it is easy to misjudge "normal traffic peaks" as abnormalities.
[0070] Therefore, in this embodiment of the application, in order to make the anomaly judgment criteria more in line with the actual business environment and reduce the risk of misjudgment and missed judgment, dynamic anomaly thresholds will also be obtained based on related business scenario data and by pre-setting historical calibration data.
[0071] Specifically, see Figure 4 Based on data from related business scenarios, dynamic anomaly thresholds are obtained by pre-setting historical calibration data, including the following steps: S410. Based on the associated business scenario data, determine the business scenario corresponding to the detection time window, and denot it as the target business scenario.
[0072] S420. Determine the calibration data corresponding to the detection time window through historical calibration data, and record it as sample calibration data.
[0073] S430. Based on the sample calibration data, the basic abnormal threshold is determined by a preset quantitative statistical method.
[0074] S440. Based on the target business scenario, match the corresponding adjustment coefficient through a preset scenario coefficient adjustment library.
[0075] S450. Based on the basic anomaly threshold and the adjustment coefficient, calculate and obtain the dynamic anomaly threshold.
[0076] First, since abnormal data detection uses the detection time window as the detection unit, it is necessary to determine the business scenario corresponding to the current detection time window based on the associated business scenario data, and denoted as the target business scenario.
[0077] Then, through historical calibration data, the calibration data corresponding to the detection time window is determined and recorded as sample calibration data. That is, based on the target business scenario and the detection time window, calibration data that is consistent with the target business scenario and has the same time granularity is selected from the historical calibration data and recorded as sample calibration data to ensure the consistency of sample distribution with the current detection data.
[0078] Among them, the preset historical calibration data refers to the set of click volume data that has been marked as "normal" or "abnormal" within a historical time period, which includes the multi-dimensional time series feature set and labeling results of each historical time window.
[0079] Then, based on the sample calibration data, the basic anomaly threshold is determined through a preset quantitative statistical method. This predictive quantitative statistical method, for example, uses the 95th quantile method, which calculates the 95th quantile of the overall anomaly score of normal samples in the sample calibration data, and uses this as the basic anomaly threshold. Or use 3 The method calculates the mean μ and standard deviation of the comprehensive abnormality scores of normal samples in the sample calibration data. Basic anomaly threshold .
[0080] Basic anomaly threshold The value range is 0-1, which is used to reflect the upper limit of abnormal scores in historical normal data.
[0081] Next, based on the target business scenario, the corresponding adjustment coefficient is matched using a preset scenario coefficient adjustment library, and denoted as... .
[0082] The preset scenario coefficient adjustment library stores adjustment coefficients corresponding to different business scenarios, and these adjustment coefficients are determined based on the scenario traffic fluctuation characteristics. The value range can be set, for example, to [0.8, 1.5]. This is suitable for everyday scenarios where traffic fluctuations are small, and the adjustment coefficient is appropriate. =1.0; During major promotional events, traffic fluctuates greatly, so the adjustment coefficient is adjusted accordingly. =1.2; In new channel promotion scenarios, data distribution is unstable, adjustment coefficient =1.3, the range of the adjustment coefficient is; in low-traffic scenarios, the traffic is low, and normal fluctuations are easily amplified, so the threshold needs to be lowered to avoid missed detections, and the adjustment coefficient is... =0.8.
[0083] Finally, based on the basic anomaly threshold and the adjustment coefficient, the dynamic anomaly threshold is calculated and obtained, that is, the dynamic anomaly threshold is T, which can be expressed as: The dynamic anomaly threshold T is limited to a range of 0-1.
[0084] Once the dynamic anomaly threshold is determined, the comprehensive anomaly score can be compared with the dynamic anomaly threshold to determine whether there is abnormal data. If abnormal data is found, an anomaly detection report is generated based on the abnormal data.
[0085] Specifically, generating an anomaly detection report based on the abnormal data includes the following steps: S510. Determine the detection time window where the abnormal data is located, and based on the detection time window, filter and obtain the corresponding multi-dimensional time series feature set, which is denoted as the abnormal feature subset.
[0086] S520. Based on the comprehensive anomaly score and dynamic anomaly threshold, calculate the anomaly deviation magnitude, and calculate the basic confidence level based on the anomaly deviation magnitude.
[0087] S530. Based on the subset of abnormal features, obtain the abnormal contribution features and corresponding significance coefficients through a preset method.
[0088] S540. Based on the abnormal contribution characteristics and the corresponding significance coefficients, the basic confidence level is corrected to obtain the final confidence level.
[0089] S550. Encode the subset of abnormal features to generate an abnormal feature vector.
[0090] S560. Based on the abnormal contribution characteristics, similarity matching is performed through a preset abnormal feature pattern library to determine the abnormal type.
[0091] S570: Integrate the detection time window, confidence level, and anomaly type to generate an anomaly detection report.
[0092] In addition to abnormal data, the anomaly detection report also includes the anomaly type and confidence level corresponding to the abnormal data. Identifying the anomaly type can help determine the cheating pattern of the anomaly and provide direction for subsequent accurate handling; while the confidence level is used to quantify the reliability of the anomaly judgment and optimize the efficiency and accuracy of handling.
[0093] First, the detection time window where the abnormal data is located is determined. All feature data corresponding to this detection time window are selected from the multi-dimensional time-series feature set and recorded as the abnormal feature subset, which serves as the core data foundation for anomaly analysis.
[0094] Then, based on the comprehensive anomaly score and the dynamic anomaly threshold, the anomaly deviation magnitude is calculated, and the baseline confidence level is calculated based on the anomaly deviation magnitude.
[0095] The magnitude of the abnormal deviation is used to reflect the proportion of the overall abnormal score that exceeds the abnormal threshold. Let the magnitude of the abnormal deviation be D. Let the base confidence level be 1. ,but It can be represented as: Basic confidence level The value ranges from 0% to 100%, and is used to initially quantify the reliability of anomaly detection.
[0096] Next, based on the subset of abnormal features, the abnormal contribution features and their corresponding significance coefficients are obtained through a preset method.
[0097] Specifically, based on a subset of abnormal features, the abnormal contribution features and their corresponding significance coefficients are obtained through a pre-defined method, including the following steps: S531. Based on the subset of abnormal features, obtain the abnormal contribution features and corresponding contribution degrees through a preset method.
[0098] S532. Normalize the contribution of abnormal contribution characteristics to generate contribution weights.
[0099] S533. Calculate the degree of abnormal deviation of abnormal contribution characteristics using preset benchmark reference data.
[0100] S534. Generate significance coefficients based on contribution weights and the degree of abnormal deviation.
[0101] The pre-defined method, such as using the SHAP method or the random forest algorithm, evaluates the feature importance of the subset of anomalous features, selecting the top n features by feature importance as anomalous contributing features. The importance value corresponding to each anomalous contributing feature is the contribution score, denoted as . , … For example, n=5 or n=3, which can be flexibly set according to the actual situation.
[0102] After identifying the abnormal contribution characteristics and their corresponding contribution levels, the contribution levels are normalized to generate contribution weights, denoted as the first... The contribution weights of each abnormal contribution feature are: ,but It can be represented as: .
[0103] Then, by using preset benchmark reference data, the degree of abnormal deviation of abnormal contribution features is calculated. Here, preset benchmark reference data refers to the set of feature mean values of normal click data under the target business scenario. That is, by statistically analyzing the feature benchmark values of historical normal data under the same business scenario, a quantitative comparison standard is provided for the degree of deviation of abnormal features, ensuring the objectivity of deviation calculation and scenario adaptability.
[0104] Record No. The degree of abnormal deviation of each abnormal contribution feature is ,but It can be represented as: in, For the first The actual value of each anomalous contribution feature This is the mean of the corresponding feature in the benchmark reference data.
[0105] After determining the degree of abnormal deviation, a significance coefficient can be generated based on the contribution weight and the degree of abnormal deviation, denoted as the first... The significance coefficients of the anomalous contribution features are: Then it can be expressed as: .
[0106] The significance coefficient is used to reflect the degree of influence of abnormal contribution characteristics on the anomaly judgment. The larger the significance coefficient, the more significant the influence.
[0107] Once the significance coefficient is determined, the baseline confidence level can be adjusted based on the abnormal contribution characteristics and the corresponding significance coefficients to obtain the final confidence level.
[0108] Let the final confidence level be... The final confidence level C ranges from 0% to 100%. The confidence level after correction by the significance coefficient of the abnormal contribution feature can more accurately reflect the reliability of the anomaly judgment. For example, an abnormal contribution feature with a high significance coefficient will improve the final confidence level.
[0109] To determine the anomaly type, a subset of anomaly features is encoded to generate anomaly feature vectors. Then, using the anomaly feature vectors as targets, similarity matching is performed through a pre-defined anomaly feature pattern library to determine the anomaly type.
[0110] The preset abnormal feature pattern library stores standard feature vector templates corresponding to four typical abnormal types (machine batch traffic boosting, IP pool / device pool cheating, simulated real person fake clicks, and reuse of historical cheating patterns).
[0111] Using the abnormal feature vector as the target, similarity matching is performed through a preset abnormal feature pattern library. That is, the cosine similarity algorithm is used to calculate the similarity between the abnormal feature vector and each standard template. The standard template with the highest similarity and the similarity reaching the preset similarity threshold is selected as the matching template. The abnormal type corresponding to the matching template is taken as the final abnormal type.
[0112] If all similarities do not exceed the preset similarity threshold, it is determined to be a novel anomaly.
[0113] Finally, the detection time window, confidence level, and anomaly type are integrated to generate an anomaly detection report. In addition, the anomaly detection report can also be supplemented with auxiliary information such as comprehensive anomaly score, dynamic anomaly threshold, and anomaly feature vector to facilitate subsequent analysis.
[0114] In this embodiment of the application, after an anomaly detection report is generated, it will be sent to a human for review, especially for anomaly judgments with low confidence.
[0115] Therefore, in this embodiment of the application, after generating an anomaly detection report based on the abnormal data, the following steps are also included: S610. If the confidence level does not reach the preset confidence threshold, the anomaly detection report will be pushed to manual review and feedback results will be received.
[0116] S620. If the feedback result is a misjudgment, the first prompt message will be output to indicate that the abnormal data was judged incorrectly.
[0117] If the final confidence level does not reach the preset confidence threshold (e.g., set to 60%), the anomaly detection report will be pushed to the manual review terminal; if it exceeds the preset confidence threshold, the report will be randomly sampled and pushed to the manual review terminal at a preset ratio (e.g., set to 10%) to ensure the accuracy of high-confidence anomalies.
[0118] The manual composite terminal will make a final judgment by business experts based on traffic characteristics, channel background, and fraud patterns, and mark the abnormal judgment results to generate feedback results, including confirmed abnormalities and misjudgments.
[0119] If the feedback result is a misjudgment, the first prompt message will be output to indicate that the abnormal data was judged incorrectly, so as to prompt the technicians to review the detection process, analyze the reasons for the misjudgment (such as incomplete feature extraction, unreasonable model parameters, threshold setting deviation, etc.), and optimize the detection method based on the reasons for the misjudgment (such as supplementing feature dimensions, adjusting model weights, correcting threshold calculation logic, etc.).
[0120] This application also provides a mobile advertising click volume abnormal data detection system, see [link]. Figure 5 The system includes: a data acquisition module 101, a time-series feature extraction module 102, an anomaly model detection module 103, a dynamic threshold generation module 104, and a detection result generation module 105.
[0121] The data acquisition module 101 is used to acquire the original click data stream and related business scenario data of the target mobile advertisement within a preset time period.
[0122] The temporal feature extraction module 102 is used to extract features from the original click data stream and obtain a multi-dimensional temporal feature set.
[0123] The anomaly detection module 103 is used to set the detection time window, and based on a multi-dimensional time series feature set, outputs a comprehensive anomaly score for each detection time window through a preset anomaly detection model.
[0124] The dynamic threshold generation module 104 is used to obtain dynamic anomaly thresholds based on related business scenario data and preset historical calibration data.
[0125] The detection result generation module 105 is used to compare the comprehensive anomaly score with the dynamic anomaly threshold to determine whether there is abnormal data. If so, an anomaly detection report is generated based on the abnormal data.
[0126] In this embodiment of the application, the data acquisition module 101 is specifically used to acquire the original click data stream and related business scenario data of the target mobile advertisement within a preset time period.
[0127] The temporal feature extraction module 102 is specifically used to extract features from the raw click data stream acquired by the data acquisition module 101 to obtain a multi-dimensional temporal feature set.
[0128] The anomaly detection module 103 is specifically used to set the detection time window. Based on the multi-dimensional temporal feature set generated by the temporal feature extraction module 102, it outputs the comprehensive anomaly score for each detection time window through a preset anomaly detection model.
[0129] The dynamic threshold generation module 104 is specifically used to obtain dynamic abnormal thresholds based on the related business scenario data obtained by the data acquisition module 101 and by using preset historical calibration data.
[0130] The detection result generation module 105 is specifically used to compare the comprehensive anomaly score generated by the anomaly model detection module 103 with the dynamic anomaly threshold obtained by the dynamic threshold generation module 104 to determine whether there is abnormal data. If so, an anomaly detection report is generated based on the abnormal data.
[0131] This application also provides a computer-readable storage medium storing a computer program that can be loaded by a processor and executed by any of the above-described methods for detecting abnormal mobile advertising click volume data.
[0132] The embodiments described in this application are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the principles of this application should be included within the scope of protection of this application.
Claims
1. A method for detecting abnormal click data in mobile advertising, characterized in that, include: Acquire raw click data stream and related business scenario data of the target mobile ad within a preset time period; Feature extraction is performed on the raw click data stream to obtain a multi-dimensional time-series feature set, which includes click volume time-series features, click behavior features, and contextual features. Set a detection time window, and based on a multi-dimensional time series feature set, output a comprehensive anomaly score for each detection time window through a preset anomaly detection model; Based on data from related business scenarios, dynamic anomaly thresholds are obtained by pre-setting historical calibration data; The overall anomaly score is compared with the dynamic anomaly threshold to determine whether there is abnormal data. If so, an anomaly detection report is generated based on the abnormal data.
2. The method for detecting abnormal mobile advertising click volume data according to claim 1, characterized in that, The original click data stream contains device ID, IP address, ad ID, and channel ID. The feature extraction process for the original click data stream, obtaining a multi-dimensional time-series feature set, includes: Based on the raw click data stream, the data is divided into multiple fixed time windows according to a preset time granularity. The click volume of each fixed time window is statistically analyzed to obtain basic statistical characteristics. Based on the raw click data stream, multiple sliding time windows are obtained according to a preset time step. The click volume of each sliding time window is statistically analyzed to obtain trend fluctuation characteristics. The basic statistical features and trend fluctuation features are integrated and denoted as the click volume time series features; Using device ID and ad ID as the behavioral units, click data for each behavioral unit is obtained in each fixed time window, and click behavior characteristics are obtained through data statistics; Using IP address and channel ID as the association unit, click data of the association unit in each fixed time window is obtained, and context association features are obtained through data statistics.
3. The method for detecting abnormal mobile advertising click volume data according to claim 1, characterized in that, The preset anomaly detection model includes four types of single detection models representing different quantification dimensions. The preset anomaly detection model, based on a multi-dimensional temporal feature set, outputs a comprehensive anomaly score for each detection time window, including: Based on a multi-dimensional temporal feature set, the features are aggregated according to the detection time window to generate a structured feature matrix. Based on the structured feature matrix, detection is performed using four types of single detection models, and single model scores are obtained. Based on the detection time window, the corresponding business scenario is matched from the associated business scenario data; Based on business scenarios, dynamic weights are generated for four types of single detection models through a pre-set scenario model mapping library. For the scores of individual models, a weighted calculation is performed by dynamically allocating weights to obtain a comprehensive anomaly score.
4. The method for detecting abnormal mobile advertising click volume data according to claim 3, characterized in that, Based on business scenarios, and through a pre-set scenario model mapping library, dynamic weights are generated for four types of single detection models, including: The initial weights of four single detection models are obtained by using a pre-set model weight database. Based on business scenarios, the impact coefficients of business scenarios on four types of single detection models are obtained through a pre-set scenario model mapping library. Based on the influence coefficient, the initial allocation weights are adjusted to generate dynamic allocation weights.
5. The method for detecting abnormal mobile advertising click volume data according to claim 1, characterized in that, The process of obtaining dynamic anomaly thresholds based on associated business scenario data and preset historical calibration data includes: Based on the associated business scenario data, determine the business scenario corresponding to the detection time window, and denot it as the target business scenario; By using historical calibration data, the calibration data corresponding to the detection time window is determined and recorded as the sample calibration data; Based on the sample calibration data, the basic anomaly threshold is determined by a preset quantitative statistical method. Based on the target business scenario, the corresponding adjustment coefficient is matched through a preset scenario coefficient adjustment library; Based on the basic anomaly threshold and the adjustment coefficient, the dynamic anomaly threshold is calculated and obtained.
6. The method for detecting abnormal mobile advertising click volume data according to claim 1, characterized in that, The step of generating an anomaly detection report based on the abnormal data includes: Determine the detection time window where the abnormal data is located, and based on the detection time window, filter and obtain the corresponding multi-dimensional time series feature set, which is denoted as the abnormal feature subset; Based on the comprehensive anomaly score and dynamic anomaly threshold, the anomaly deviation magnitude is calculated, and the baseline confidence level is calculated based on the anomaly deviation magnitude. Based on a subset of abnormal features, the abnormal contribution features and their corresponding significance coefficients are obtained through a pre-defined method. Based on the abnormal contribution characteristics and the corresponding significance coefficients, the basic confidence level is corrected to obtain the final confidence level; Encode the subset of abnormal features to generate an abnormal feature vector; Based on the abnormal feature vector, similarity matching is performed through a pre-set abnormal feature pattern library to determine the abnormal type; The detection time window, final confidence level, and anomaly type are integrated to generate an anomaly detection report.
7. The method for detecting abnormal mobile advertising click volume data according to claim 6, characterized in that, The step of obtaining abnormal contribution features and corresponding significance coefficients based on an abnormal feature subset using a preset method includes: Based on a subset of abnormal features, abnormal contribution features and their corresponding contribution degrees are obtained through a preset method. Normalize the contribution of abnormal contribution characteristics to generate contribution weights; The degree of abnormal deviation of abnormal contribution characteristics is calculated by using preset benchmark reference data; A significance coefficient is generated based on the contribution weight and the degree of abnormal deviation.
8. The method for detecting abnormal mobile advertising click volume data according to claim 1, characterized in that, The anomaly detection report includes a confidence level. After generating the anomaly detection report based on the anomaly data, it includes: If the confidence level does not reach the preset confidence threshold, the anomaly detection report will be pushed to manual review and feedback results will be received. The feedback results include misjudgment and confirmed anomalies. If the feedback result is a misjudgment, the first prompt message will be output to indicate that the abnormal data was judged incorrectly.
9. A mobile advertising click-through rate anomaly detection system, characterized in that, include: The data acquisition module (101) is used to acquire the raw click data stream and related business scenario data of the target mobile advertisement within a preset time period; The temporal feature extraction module (102) is used to extract features from the original click data stream and obtain a multi-dimensional temporal feature set, which includes click volume time series features, click behavior features and contextual association features. Anomaly detection module (103) is used to set detection time windows and output a comprehensive anomaly score for each detection time window based on a multi-dimensional time series feature set and a preset anomaly detection model. The dynamic threshold generation module (104) is used to obtain dynamic abnormal thresholds based on related business scenario data and preset historical calibration data; The detection result generation module (105) is used to compare the comprehensive abnormal score with the dynamic abnormal threshold to determine whether there is abnormal data. If so, an abnormal detection report is generated based on the abnormal data.
10. A computer-readable storage medium storing a computer program capable of being loaded by a processor and executing a method for detecting abnormal mobile advertising click volume as described in any one of claims 1 to 8.