Method and device for building classifier and detecting abnormal advertising traffic
Through clustering, acquiring the characteristic distribution prototype of advertising traffic data is built, which solves the problem of time-consuming and labor-consuming manual labeling and reduces costs.
Patent Information
- Application Number
- CN202210012977.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-01-06
AI Technical Summary
In the prior art, it takes a lot of time and labor to manually label all collected samples, resulting in high cost of detecting abnormal advertising traffic.
By obtaining multiple advertising traffic data in the first preset time period, clustering is performed to obtain data feature distributions of normal types and exception types, a clustering center prototype representing these feature distributions is obtained, and an advertising traffic classifier is constructed based on these prototypes.
There is no need to label the collected advertising traffic data, which saves a lot of time and labor and reduces the cost of building advertising traffic classifiers.
Smart Images

Figure CN114358848B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method and apparatus for constructing a classifier and detecting anomalies in advertising traffic. Background Art
[0002] In recent years, with the popularization of mobile Internet, the number of Internet users and the length of Internet time have reached new highs. The problem of advertising fraud that accompanies the prosperity of the online advertising market has appeared from time to time. Some advertisers, advertising operators and advertising publishers will deliberately create false appearances of goods and services, or conceal the truth in advertising activities and take a series of illegal activities. Therefore, it is necessary to detect advertising traffic anomalies. At the same time, with the rapid development of big data technology and artificial intelligence technology, advertising traffic anomaly detection based on artificial intelligence technology has become a hot topic of research in recent years. Advertising traffic includes advertising browsing and advertising clicks. In the existing technology, samples are usually collected in a fully supervised training mode, and all collected samples are manually labeled with category labels. The machine learning model is trained according to the samples with category labels to obtain an advertising traffic anomaly detection model, so as to realize advertising traffic anomaly detection using the advertising traffic anomaly detection model.
[0003] In the process of implementing the embodiments of the present disclosure, it is found that there are at least the following problems in the related art:
[0004] In the prior art, manually labeling all collected samples consumes a lot of time and labor, resulting in high costs for advertising traffic anomaly detection. Summary of the invention
[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0006] The embodiments of the present disclosure provide a method and apparatus, an electronic device, and a storage medium for constructing an advertisement traffic classifier and detecting advertisement traffic anomalies, so as to reduce the cost of constructing the advertisement traffic classifier.
[0007] In some embodiments, the method for constructing an advertising traffic classifier includes: obtaining multiple first advertising traffic data within a first preset time period; clustering the first advertising traffic data to obtain multiple types of advertising traffic data feature distributions; the types include normal types and abnormal types; respectively obtaining prototypes of cluster centers for characterizing the feature distributions of each of the advertising traffic data; and constructing an advertising traffic classifier based on each of the prototypes.
[0008] In some embodiments, the method for detecting anomalies in advertising traffic includes: performing anomaly detection on advertising traffic using an advertising traffic classifier obtained by the above-mentioned method for constructing an advertising traffic classifier.
[0009] In some embodiments, the device for constructing an advertising traffic classifier includes: a first acquisition module, configured to acquire multiple first advertising traffic data within a first preset time period; a clustering module, configured to obtain multiple types of advertising traffic data feature distributions by clustering each of the first advertising traffic data; the types include normal types and abnormal types; a second acquisition module, configured to respectively acquire prototypes of cluster centers for characterizing the feature distribution of each of the advertising traffic data; and a construction module, configured to construct an advertising traffic classifier based on each of the prototypes.
[0010] In some embodiments, the device for detecting advertising traffic anomalies performs advertising traffic anomaly detection using a model constructed using the above-mentioned method for constructing an advertising traffic classifier, including: a third acquisition module, configured to acquire multiple second advertising traffic data within a second preset time period; a feature extraction module, configured to perform feature extraction on each of the second advertising traffic data, and obtain second traffic data features corresponding to each of the second advertising traffic data; a determination module, configured to use the advertising traffic classifier to determine the type of each of the second advertising traffic data according to each of the second traffic data features.
[0011] In some embodiments, the electronic device includes a first processor and a first memory storing program instructions, and the first processor is configured to execute the method for constructing an advertising traffic classifier as described above when running the program instructions.
[0012] In some embodiments, the electronic device includes a second processor and a second memory storing program instructions, and the second processor is configured to execute the above-mentioned method for detecting advertising traffic anomalies when running the program instructions.
[0013] In some embodiments, the storage medium stores program instructions, and when the program instructions are run, the above-mentioned method for constructing an advertisement traffic classifier is executed.
[0014] In some embodiments, the storage medium stores program instructions, and when the program instructions are run, the above-mentioned method for detecting advertising traffic anomalies is executed.
[0015] The method and device, electronic device, and storage medium for constructing an advertising traffic classifier and detecting advertising traffic anomalies provided by the embodiments of the present disclosure can achieve the following technical effects: by obtaining multiple first advertising traffic data within a first preset time period; clustering each first advertising traffic data to obtain multiple types of advertising traffic data feature distributions; the types include normal types and abnormal types; respectively obtaining prototypes of cluster centers for characterizing the feature distribution of each advertising traffic data; and constructing an advertising traffic classifier based on each prototype. By adopting a clustering method to obtain prototypes of cluster centers for characterizing the feature distribution of each advertising traffic data, and constructing an advertising traffic classifier based on each prototype, there is no need to label the collected first advertising traffic data, which saves a lot of time and labor and reduces the cost of constructing an advertising traffic classifier.
[0016] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:
[0018] Figure 1 is a schematic diagram of a method for constructing an advertisement traffic classifier provided by an embodiment of the present disclosure;
[0019] Figure 2 is a schematic diagram of a method for detecting abnormal advertising traffic provided by an embodiment of the present disclosure;
[0020] Figure 3 is a schematic diagram of a device for constructing an advertisement traffic classifier provided by an embodiment of the present disclosure;
[0021] Figure 4 is a schematic diagram of a device for detecting abnormal advertising traffic provided by an embodiment of the present disclosure;
[0022] Figure 5 is a schematic diagram of an electronic device provided by an embodiment of the present disclosure;
[0023] Figure 6 is a schematic diagram of another electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0025] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.
[0026] Unless otherwise stated, the term "plurality" means two or more.
[0027] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.
[0028] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.
[0029] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0030] Combination Figure 1 As shown, the embodiment of the present disclosure provides a method for constructing an advertisement traffic classifier, including:
[0031] Step S101, obtaining a plurality of first advertisement flow data within a first preset time period.
[0032] Step S102, clustering is performed according to each first advertisement traffic data to obtain multiple types of advertisement traffic data feature distributions; the types include normal types and abnormal types.
[0033] Step S103, respectively obtaining the prototypes of the cluster centers for characterizing the characteristic distribution of each advertisement traffic data.
[0034] Step S104: construct an advertisement traffic classifier according to each prototype.
[0035] The method for constructing an advertising traffic classifier provided by the embodiment of the present disclosure is adopted, by obtaining multiple first advertising traffic data within a first preset time period; clustering is performed based on each first advertising traffic data to obtain multiple types of advertising traffic data feature distribution; the types include normal type and abnormal type; respectively obtain the prototypes of the cluster center used to characterize the feature distribution of each advertising traffic data; and construct an advertising traffic classifier based on each prototype. By adopting a clustering method to obtain the prototype of the cluster center used to characterize the feature distribution of each advertising traffic data, and constructing an advertising traffic classifier based on each prototype, there is no need to label the collected first advertising traffic data, which saves a lot of time and labor, and reduces the cost of constructing an advertising traffic classifier.
[0036] Optionally, the first advertisement traffic data within the first preset time period is historical advertisement traffic data.
[0037] In some embodiments, the historical advertising traffic data is evenly divided into 10 parts, and a 10-fold cross validation is performed; wherein, during the training process, 9 parts are selected as training sets and 1 part is selected as a validation set, and this is performed 10 times without repetition.
[0038] Optionally, before clustering the first advertisement traffic data to obtain multiple types of advertisement traffic data feature distributions, the method further includes: performing data preprocessing on the first advertisement traffic data.
[0039] Optionally, data preprocessing includes one or more of data alignment, missing value processing, data minification, data type conversion, and data standard conversion. In this way, by preprocessing the collected first advertising traffic data, the first advertising traffic data with missing and noisy conditions is filtered through data capabilities to ensure the training effect of the advertising traffic classifier, so that the classification accuracy of the advertising traffic classifier obtained through training is higher and the classification effect is better.
[0040] Optionally, clustering is performed on each first advertising traffic data to obtain multiple types of advertising traffic data feature distributions, including: extracting features from each first advertising traffic data to obtain first traffic data features corresponding to each first advertising traffic data; clustering each first traffic data feature to obtain multiple types of advertising traffic data feature distributions.
[0041] Optionally, feature extraction is performed on each first advertising traffic data to obtain first traffic data features corresponding to each first advertising traffic data, including: feature extraction is performed on each first advertising traffic data according to a preset feature extraction model to obtain first traffic data features corresponding to each first advertising traffic data.
[0042] Optionally, the preset feature extraction model is a feature representation model based on a deep MAE (Masked AutoEncoder).
[0043] In some embodiments, the feature representation model based on deep MAE includes an encoder and a decoder. The encoder is a forward feature mapping process. Through local connections of neurons, the input first advertising traffic data is mapped to a high-level feature space, thereby representing it as a more abstract feature, for example: the first traffic data feature; the decoder is a reverse reconstruction input process. Also through local connections of neurons, the first advertising traffic data is reconstructed by the first traffic data feature to obtain a reconstructed feature; the decoder is used to iteratively optimize the encoder. In some embodiments, the numerical similarity between the reconstructed feature and the first traffic data feature is measured by MSE (Mean Square Error). The smaller the mean square error, the greater the numerical similarity; the distribution similarity between the reconstructed feature and the first traffic data feature is measured by MMD (Maximum Mean Discrepancy). The smaller the maximum mean difference, the more similar the distribution; the loss function is calculated by performing a weighted sum operation on it. During the training process, the reconstructed feature is made as close to the first advertising traffic data as possible by minimizing the loss function to learn more abstract and useful information. Finally, the hyperparameters and parameters of the feature representation model are iteratively optimized by the particle swarm optimization algorithm and the Adam optimization algorithm to obtain the feature representation model. Among them, the hyperparameters include the number of layers of the feature representation model, the number of neurons in each layer, and the learning rate of training; the parameters include the weights and biases of the feature representation model.
[0044] Optionally, prototypes of cluster centers for characterizing characteristic distributions of each advertising traffic data are obtained respectively, including: obtaining the posterior probabilities generated by each first advertising traffic data for each advertising traffic data characteristic distribution; obtaining the mean vector of each advertising traffic data characteristic distribution according to each posterior probability; and obtaining the prototype corresponding to each advertising traffic data characteristic distribution according to each mean vector.
[0045] Optionally, the posterior probabilities generated by the characteristic distributions of each first advertising traffic data are obtained respectively, including: obtaining a preset number of Gaussian distributions; initializing each Gaussian distribution and its corresponding weights, and determining the mean vector and covariance matrix corresponding to each Gaussian distribution; and calculating using the weights, mean vectors and covariance matrices according to a preset algorithm to obtain the posterior probability.
[0046] Optionally, the posterior probability is obtained by calculating using the weight, mean vector and covariance matrix according to a preset algorithm, including: Get the posterior probability; where γ jiis the posterior probability generated by the jth first advertising traffic data for the i-th advertising traffic data feature distribution, α i is the weight corresponding to the characteristic distribution of the i-th advertising traffic data, x j is the jth first advertisement traffic data, μ i is the mean vector corresponding to the characteristic distribution of the i-th advertising traffic data, μ l is the mean vector corresponding to the lth advertising traffic data feature distribution, k is the number of advertising traffic data feature distributions, ∑ i is the covariance matrix corresponding to the characteristic distribution of the i-th advertising traffic data, ∑ l is the covariance matrix corresponding to the characteristic distribution of the lth advertising traffic data, p(x j |μ i ,∑ i ) is the probability density of the characteristic distribution of the jth first advertising traffic data corresponding to the ith advertising traffic data, p(x j |μ l ,∑ l ) is the probability density of the characteristic distribution of the j-th first advertising traffic data corresponding to the l-th advertising traffic data.
[0047] Optionally, by calculating Obtain the probability density of the characteristic distribution of the j-th first advertisement traffic data corresponding to the i-th advertisement traffic data; where p(x j |μ i ,∑ i ) is the probability density of the characteristic distribution of the jth first advertisement traffic data corresponding to the ith advertisement traffic data, x j is the jth first advertisement traffic data, μ i is the mean vector corresponding to the characteristic distribution of the i-th advertising traffic data, x is the advertising traffic data, n is the number of variables in the advertising traffic data, and T is the matrix transposition symbol.
[0048] Optionally, a log-likelihood function is determined based on the weights, the mean vector and the covariance matrix. The log-likelihood function is Among them, α i is the weight corresponding to the characteristic distribution of the i-th advertising traffic data, μ i is the mean vector corresponding to the characteristic distribution of the i-th advertising traffic data, ∑ i is the covariance matrix corresponding to the characteristic distribution of the i-th advertising traffic data, and L is the added value of the log-likelihood function.
[0049] Optionally, by calculating Get the mean vector corresponding to the characteristic distribution of the i-th advertising traffic data; where μ i is the mean vector corresponding to the characteristic distribution of the i-th advertising traffic data, γji is the posterior probability generated by the characteristic distribution of the jth first advertising traffic data for the i-th advertising traffic data, x j is the jth first advertisement traffic data, and m is the number of first advertisement traffic data.
[0050] Optionally, by calculating Get the covariance matrix corresponding to the characteristic distribution of the i-th advertising traffic data; where, ∑ i is the covariance matrix corresponding to the characteristic distribution of the i-th advertising traffic data, μ i is the mean vector corresponding to the characteristic distribution of the i-th advertising traffic data, γ ji is the posterior probability generated by the characteristic distribution of the jth first advertising traffic data for the i-th advertising traffic data, x j is the j-th first advertising traffic data, m is the number of first advertising traffic data, and T is the matrix transposition symbol.
[0051] Optionally, by calculating Get the weight corresponding to the characteristic distribution of the i-th advertising traffic data; where α i is the weight corresponding to the characteristic distribution of the i-th advertising traffic data, γ ji is the posterior probability generated by the j-th first advertising traffic data for the characteristic distribution of the ith advertising traffic data, and m is the number of first advertising traffic data.
[0052] In some embodiments, cluster learning is performed on the first traffic data features corresponding to each first advertisement traffic data through GMM (Gaussian Mixture Mode) to obtain a mixed representation of the probability distribution of the multi-dimensional Gaussian mixture model.
[0053] Optionally, obtaining a mean vector of each advertising traffic data characteristic distribution according to each posterior probability includes: performing a mean vector update operation in a loop, the mean vector update operation including: obtaining an alternative mean vector according to each posterior probability; and obtaining an increase in a log-likelihood function corresponding to the alternative mean vector in each mean vector update operation; when the increase in the log-likelihood function converges within a preset numerical range, determining the alternative mean vector corresponding to the increase in the log-likelihood function as the mean vector; or, when the number of cycles of the mean vector update operation is equal to a preset number, determining the alternative mean vector corresponding to the preset number as the mean vector.
[0054] Optionally, the mean vector update operation further includes: obtaining alternative mean vectors, alternative weights and alternative covariance matrices according to the posterior probabilities; and obtaining an increase in the log-likelihood function corresponding to the alternative mean vector during each mean vector update operation.
[0055] Optionally, obtaining a prototype corresponding to each characteristic distribution of the advertisement traffic data according to each mean vector includes: determining each mean vector as a prototype corresponding to each characteristic distribution of the advertisement traffic data.
[0056] Optionally, an advertising traffic classifier is constructed according to each prototype, including: obtaining a first probability that each first advertising traffic data belongs to each type according to the characteristic distribution of each advertising traffic data; obtaining a distance vector between each first probability and each prototype using a preset classifier; obtaining a second probability that each first advertising traffic data belongs to each type according to each distance vector using a preset classifier; determining each first probability as a soft label and each second probability as a predicted label; obtaining a cross entropy between each soft label and each predicted label; and training a preset classifier according to each cross entropy to obtain an advertising traffic classifier.
[0057] Optionally, the preset classifier includes a feature mapping layer, and the preset classifier is used to obtain the distance vector between each first probability and each prototype, including: mapping each first flow data feature and each prototype into the feature space through the feature mapping layer to obtain the distance vector between each first flow data feature and each prototype.
[0058] Optionally, the preset classifier includes an output layer, and the preset classifier is used to obtain the second probability that each first advertising traffic data belongs to each type according to each distance vector, including: performing logical regression on each distance vector through the output layer to obtain the second probability that each first advertising traffic data belongs to each type, and determining the type with the largest probability value as the type of each first advertising traffic data.
[0059] Optionally, the preset classifier further includes an input layer, and the input of the input layer is a first traffic data feature corresponding to each first advertisement traffic data.
[0060] Optionally, a preset classifier is trained according to each cross entropy to obtain an advertising traffic classifier, including: adjusting weight parameters in a feature mapping layer in the preset classifier according to each cross entropy until the number of adjustments reaches a preset number of adjustments.
[0061] In some embodiments, the first traffic data features corresponding to each first advertising traffic data are obtained, and the first probabilities that the first advertising traffic data belongs to each type are obtained based on the distribution of each advertising traffic data feature, and each first probability is determined as a soft label of the first traffic data feature. The advertising traffic classifier takes the first traffic data features corresponding to each first advertising traffic data as input; by mapping each first traffic data feature and each prototype to the feature space, the distance vectors between each first traffic data feature and each prototype are obtained in the feature space; logistic regression is performed on each distance vector to obtain the second probability that each first advertising traffic data belongs to each type, and the type with the largest probability value is determined as the type of each first advertising traffic data.
[0062] In some embodiments, the first probability that the first advertising traffic data belongs to each type includes: the probability that the first advertising traffic data belongs to the normal type is: 90%, the probability that the first advertising traffic data belongs to the abnormal type is: 10%; then the soft label corresponding to the first advertising traffic data is (probability of normal type: 90%, probability of abnormal type: 10%).
[0063] In some embodiments, a logistic regression is performed on each distance vector to obtain a second probability that the first traffic data feature belongs to each type, including: the probability that the first traffic data feature belongs to the normal type is: 95%, and the probability that the first traffic data feature belongs to the abnormal type is: 5%; then the normal type corresponding to the 95% with the largest probability value is determined as the type corresponding to the first advertising traffic data.
[0064] Optionally, after determining the first traffic data features corresponding to each first advertisement traffic data as training samples, the method further includes: oversampling the training samples so that the number of training samples with a high first probability of normal type is equal to the number of training samples with a high first probability of abnormal type. In this way, by oversampling the training samples, the type imbalance problem of the first traffic features can be effectively avoided, and the detection false negative rate of the advertisement traffic classifier due to data imbalance can be reduced.
[0065] In some embodiments, each training sample corresponds to a first probability of a normal type and a first probability of an abnormal type, for example, the probability that the training sample is a normal type is 99%, and the probability that the training sample is an abnormal type is 1%.
[0066] In some embodiments, the number of training samples is 100, of which 90 are normal type training samples with a high first probability, and 10 are abnormal type training samples with a high first probability. The training samples are oversampled through ADASYN (Adaptive Synthetic Sampling), and the number of abnormal type first probability samples is adjusted to 90 to balance the training samples.
[0067] Optionally, the advertisement traffic classifier obtained by the method for constructing an advertisement traffic classifier is used to perform advertisement traffic anomaly detection. In this way, advertisement traffic anomaly detection can be performed, which facilitates identification of abnormal advertisement traffic and realizes advertisement traffic anti-fraud.
[0068] Combination Figure 2 As shown, the embodiment of the present disclosure provides a method for detecting abnormal advertising traffic, including:
[0069] Step S201, obtaining second advertisement flow data within a second preset time period.
[0070] Step S202: extract features from the second advertisement traffic data to obtain second traffic data features corresponding to the second advertisement traffic data.
[0071] Step S203: Determine the type of the second advertisement traffic data according to the second traffic data characteristics by using an advertisement traffic classifier.
[0072] By adopting the method for detecting anomalies in advertising traffic provided by the embodiment of the present disclosure, data mining is realized by acquiring the second traffic data features corresponding to the second advertising traffic data, and the type of the second advertising traffic data is determined according to the second traffic data features by using the advertising traffic classifier. In this way, anomaly detection of advertising traffic data is realized through data capabilities, which facilitates the identification of abnormal advertising traffic and realizes advertising traffic anti-fraud.
[0073] Optionally, an advertising traffic classifier is used to determine the type of each second advertising traffic data according to the characteristics of each second traffic data, including: obtaining the second probability that the second advertising traffic data belongs to each type according to the characteristics of the second traffic data; and determining the type corresponding to the second probability with the largest value as the type corresponding to the second advertising traffic data.
[0074] Optionally, a second probability that the second advertising traffic data belongs to each type is obtained based on the second traffic data feature, including: inputting the second traffic data feature into an advertising traffic classifier, obtaining a first probability that the second traffic data feature belongs to a normal type and a first probability that the abnormal type is obtained, obtaining the distance between each first probability and each prototype, inputting each distance into an output layer of the advertising traffic classifier, obtaining a second probability that the second traffic data feature belongs to a normal type and a second probability that the abnormal type is obtained, and determining the type corresponding to the second probability with the largest value as the type corresponding to the second traffic data feature.
[0075] In some embodiments, the second probability that the second traffic data feature belongs to the normal type is 90%, and the second probability that the second traffic data feature belongs to the abnormal type is 10%, then the normal type is determined as the type corresponding to the second traffic data feature.
[0076] Combination Figure 3As shown, the embodiment of the present disclosure provides a device for constructing an advertising traffic classifier, including: a first acquisition module 301, a clustering module 302, a second acquisition module 303 and a construction module 304; the first acquisition module 301 is configured to acquire multiple first advertising traffic data within a first preset time period, and send each first advertising traffic data to the clustering module; the clustering module 302 is configured to receive the first advertising traffic data sent by the first acquisition module, cluster each first advertising traffic data to obtain multiple types of advertising traffic data feature distribution; the types include normal types and abnormal types; and send the advertising traffic data feature distribution to the second acquisition module, the second acquisition module 303 is configured to receive the advertising traffic data feature distribution sent by the clustering module, and respectively acquire the prototypes of the clustering centers used to characterize the feature distribution of each advertising traffic data, and send each prototype to the construction module; the construction module 304 is configured to receive the prototype sent by the second acquisition module, and construct an advertising traffic classifier according to each prototype.
[0077] The device for constructing an advertising traffic classifier provided by the embodiment of the present disclosure is adopted. A plurality of first advertising traffic data within a first preset time period is acquired through a first acquisition module. A clustering module clusters each first advertising traffic data to obtain characteristic distributions of multiple types of advertising traffic data. A second acquisition module respectively acquires prototypes of cluster centers used to characterize characteristic distributions of each advertising traffic data. A determination module constructs an advertising traffic classifier according to each prototype. By adopting a clustering method to obtain prototypes of cluster centers used to characterize characteristic distributions of each advertising traffic data, and constructing an advertising traffic classifier according to each prototype, there is no need to label the collected first advertising traffic data, thus saving a lot of time and labor, and reducing the cost of constructing an advertising traffic classifier.
[0078] Optionally, the clustering module is configured to obtain multiple types of advertising traffic data feature distributions by clustering each first advertising traffic data in the following manner: extracting features from each first advertising traffic data to obtain first traffic data features corresponding to each first advertising traffic data; clustering each first traffic data feature to obtain multiple types of advertising traffic data feature distributions.
[0079] Optionally, the second acquisition module is configured to respectively acquire the prototypes of the cluster centers used to characterize the characteristic distribution of each advertising traffic data in the following manner: respectively acquire the posterior probabilities generated by each first advertising traffic data for each characteristic distribution of the advertising traffic data; acquire the mean vector of each characteristic distribution of the advertising traffic data according to each posterior probability; and acquire the prototype corresponding to each characteristic distribution of the advertising traffic data according to each mean vector.
[0080] Optionally, obtaining a mean vector of each advertising traffic data characteristic distribution according to each posterior probability includes: performing a mean vector update operation in a loop, the mean vector update operation including: obtaining an alternative mean vector according to each posterior probability; and obtaining an increase in a log-likelihood function corresponding to the alternative mean vector in each mean vector update operation; when the increase in the log-likelihood function converges within a preset numerical range, determining the alternative mean vector corresponding to the increase in the log-likelihood function as the mean vector; or, when the number of cycles of the mean vector update operation is equal to a preset number, determining the alternative mean vector corresponding to the preset number as the mean vector.
[0081] Optionally, the construction module is configured to construct an advertising traffic classifier based on each prototype in the following manner: obtain a first probability that each first advertising traffic data belongs to each type based on the characteristic distribution of each advertising traffic data; use a preset classifier to obtain a distance vector between each first probability and each prototype; use the classifier to obtain a second probability that each first advertising traffic data belongs to each type based on each distance vector; determine each first probability as a soft label, and determine each second probability as a predicted label; obtain the cross entropy between each soft label and each predicted label; train the classifier according to each cross entropy to obtain an advertising traffic classifier.
[0082] Combination Figure 4 As shown, the embodiment of the present disclosure provides a device for detecting advertising traffic anomalies, including: a third acquisition module 401, a feature extraction module 402 and a determination module 403; the third acquisition module 401 is configured to acquire second advertising traffic data within a second preset time period, and send the second advertising traffic data to the feature extraction module; the feature extraction module 402 is configured to receive the second advertising traffic data sent by the third acquisition module, perform feature extraction on the second advertising traffic data, obtain second traffic data features corresponding to each second advertising traffic data, and send the second traffic data features to the determination module; the determination module 403 is configured to receive the second traffic data features sent by the feature extraction module, and use the advertising traffic classifier to determine the type of each second advertising traffic data according to each second traffic data feature.
[0083] By using the device for constructing an advertising traffic classifier provided by the embodiment of the present disclosure, the second advertising traffic data within the second preset time period is obtained through the third acquisition module, the feature extraction module performs feature extraction on the second advertising traffic data, and obtains the second traffic data features corresponding to each second advertising traffic data, and the determination module uses the advertising traffic classifier to determine the type of each second advertising traffic data according to each second traffic data feature. In this way, it is possible to detect anomalies in advertising traffic data, facilitate identification of abnormal advertising traffic, and implement advertising traffic anti-fraud.
[0084] Optionally, the determination module is configured to determine the type of the second advertising traffic data based on the second traffic data characteristics using an advertising traffic classifier in the following manner: obtaining the second probability that the second advertising traffic data belongs to each type based on the second traffic data characteristics; and determining the type corresponding to the second probability with the largest value as the type corresponding to the second advertising traffic data.
[0085] Combination Figure 5 As shown, an embodiment of the present disclosure provides an electronic device, including a first processor (processor) 500 and a first memory (memory) 501. Optionally, the electronic device may also include a first communication interface (CommunicationInterface) 502 and a first bus 503. Among them, the first processor 500, the first communication interface 502, and the first memory 501 can communicate with each other through the first bus 503. The first communication interface 502 can be used for information transmission. The first processor 500 can call the logic instructions in the first memory 501 to execute the method for constructing a polarity fault mode recognition model of a posture control system in the above embodiment.
[0086] The electronic device provided by the embodiment of the present disclosure is adopted to obtain multiple first advertising traffic data within a first preset time period; cluster each first advertising traffic data to obtain multiple types of advertising traffic data feature distribution; the types include normal type and abnormal type; respectively obtain the prototypes of the cluster center used to characterize the feature distribution of each advertising traffic data; and construct an advertising traffic classifier based on each prototype. By adopting a clustering method to obtain the prototype of the cluster center used to characterize the feature distribution of each advertising traffic data, and constructing an advertising traffic classifier based on each prototype, there is no need to label the collected first advertising traffic data, which saves a lot of time and labor and reduces the cost of constructing an advertising traffic classifier.
[0087] Optionally, the electronic device includes a server or a computer, etc.
[0088] In addition, the logic instructions in the first memory 501 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0089] The first memory 501 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The first processor 500 executes the function application and data processing by running the program instructions / modules stored in the first memory 501, that is, the method for building an advertising traffic classifier in the above embodiment is implemented.
[0090] The first memory 501 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the first memory 501 may include a high-speed random access memory and may also include a non-volatile memory.
[0091] An embodiment of the present disclosure provides a storage medium storing program instructions, which, when run, execute the above method for building an advertisement traffic classifier.
[0092] An embodiment of the present disclosure provides a computer program product, which includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the above-mentioned method for constructing an advertising traffic classifier.
[0093] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0094] Combination Figure 6 As shown, an embodiment of the present disclosure provides an electronic device, including a second processor (processor) 600 and a second memory (memory) 601. Optionally, the electronic device may also include a second communication interface (CommunicationInterface) 602 and a second bus 603. Among them, the second processor 600, the second communication interface 602, and the second memory 601 can communicate with each other through the second bus 603. The second communication interface 602 can be used for information transmission. The second processor 600 can call the logic instructions in the second memory 601 to execute the method for identifying the polarity failure mode of the attitude control system of the above embodiment.
[0095] By using the electronic device provided by the embodiment of the present disclosure, the second traffic data feature corresponding to the second advertisement traffic data is obtained, and the type of the second advertisement traffic data is determined according to the second traffic data feature by using the advertisement traffic classifier. In this way, abnormal detection of advertisement traffic data can be achieved, which facilitates identification of abnormal advertisement traffic and realizes advertisement traffic anti-fraud.
[0096] Optionally, the electronic device includes a server, a computer, etc.
[0097] In addition, the logic instructions in the second memory 601 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0098] The second memory 601 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The second processor 600 executes the function application and data processing by running the program instructions / modules stored in the second memory 601, that is, the method for detecting abnormal advertising traffic in the above embodiment is implemented.
[0099] The second memory 601 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the second memory 601 may include a high-speed random access memory and may also include a non-volatile memory.
[0100] The embodiment of the present disclosure provides a storage medium storing program instructions, which, when running, execute the above method for detecting abnormal advertising traffic.
[0101] An embodiment of the present disclosure provides a computer program product, which includes a computer program stored on a computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the above-mentioned method for detecting advertising traffic anomalies.
[0102] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0103] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for enabling a computer device (which may be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes, or a transient storage medium.
[0104] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.
[0105] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods for each specific application to implement the described functions, but such implementations should not be considered to exceed the scope of the embodiments of the present disclosure. The technicians may clearly understand that, for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above may refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.
[0106] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.
[0107] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A method for constructing an advertising traffic classifier, It is characterized in that include: Acquire a plurality of first advertisement flow data within a first preset time period; Clustering is performed according to each of the first advertisement traffic data to obtain multiple types of advertisement traffic data feature distributions; the types include normal types and abnormal types; Respectively obtaining the prototypes of the cluster centers for characterizing the characteristic distribution of each of the advertisement traffic data; Constructing an advertising traffic classifier based on each of the prototypes; Respectively obtaining the prototypes of the cluster centers for characterizing the characteristic distribution of each of the advertisement traffic data, including: Respectively obtain the posterior probability that each of the first advertisement traffic data is generated by the characteristic distribution of each of the advertisement traffic data; Obtaining a mean vector of each of the advertising traffic data feature distributions according to each of the posterior probabilities; The prototype corresponding to each of the advertising traffic data characteristic distributions is obtained according to each of the mean vectors.
2. The method according to claim 1, It is characterized in that Clustering is performed according to each of the first advertisement traffic data to obtain multiple types of advertisement traffic data feature distributions, including: Extracting features from each of the first advertisement traffic data to obtain first traffic data features corresponding to each of the first advertisement traffic data; Clustering is performed on each of the first traffic data features to obtain multiple types of advertisement traffic data feature distributions.
3. The method according to claim 1, It is characterized in that Obtaining the mean vector of each of the advertising traffic data feature distributions according to each of the posterior probabilities includes: The mean vector updating operation is performed cyclically, wherein the mean vector updating operation includes: obtaining a candidate mean vector according to each of the posterior probabilities; and obtaining an increase value of a log-likelihood function corresponding to the candidate mean vector during each mean vector updating operation; When the increase in the log-likelihood function converges within a preset numerical range, the alternative mean vector corresponding to the increase in the log-likelihood function is determined as the mean vector; or, when the number of cycles of the mean vector update operation is equal to a preset number, the alternative mean vector corresponding to the preset number is determined as the mean vector.
4. The method according to claim 2, It is characterized in that An advertising traffic classifier is constructed based on each of the above prototypes, including: Obtaining a first probability that each of the first advertisement traffic data belongs to each of the types according to the characteristic distribution of each of the advertisement traffic data; Using a preset classifier to obtain a distance vector between each of the first probabilities and each of the prototypes; Using the classifier to obtain, according to each of the distance vectors, a second probability that each of the first advertisement traffic data belongs to each of the types; Determine each of the first probabilities as a soft label, and determine each of the second probabilities as a predicted label; Obtaining the cross entropy between each of the soft labels and each of the predicted labels; The classifier is trained according to each cross entropy to obtain the advertisement traffic classifier.
5. A method for detecting anomalies in advertising traffic, It is characterized in that An advertisement traffic classifier obtained by using the method for constructing an advertisement traffic classifier described in any one of claims 1 to 4 is used to perform advertisement traffic anomaly detection.
6. The method according to claim 5, It is characterized in that Using the advertisement traffic classifier to perform advertisement traffic anomaly detection includes: Acquire second advertisement flow data within a second preset time period; Performing feature extraction on the second advertisement traffic data to obtain second traffic data features corresponding to the second advertisement traffic data; The advertisement traffic classifier is used to determine the type of the second advertisement traffic data according to the characteristics of the second traffic data, and the type is used to characterize whether the advertisement traffic data is abnormal.
7. The method according to claim 6, It is characterized in that Determining the type of the second advertisement traffic data according to the second traffic data feature by using the advertisement traffic classifier includes: Acquire a second probability that the second advertisement traffic data belongs to each of the types according to the second traffic data feature; The type corresponding to the second probability with the largest value is determined as the type corresponding to the second advertising traffic data.
8. A device for constructing an advertising traffic classifier, It is characterized in that include: A first acquisition module is configured to acquire a plurality of first advertisement flow data within a first preset time period; A clustering module, configured to cluster the first advertisement traffic data to obtain multiple types of advertisement traffic data feature distributions; the types include normal types and abnormal types; A second acquisition module is configured to respectively acquire the prototypes of the cluster centers for characterizing the characteristic distribution of each of the advertisement traffic data; A construction module, configured to construct an advertisement traffic classifier according to each of the prototypes; The second acquisition module is configured to respectively acquire the prototypes of the cluster centers for characterizing the characteristic distribution of each advertisement traffic data in the following manners: Respectively obtain the posterior probability that each first advertisement traffic data is generated by the characteristic distribution of each advertisement traffic data; Obtaining a mean vector of the characteristic distribution of each advertisement traffic data according to each posterior probability; The prototype corresponding to each advertising traffic data characteristic distribution is obtained according to each mean vector.
9. A device for detecting abnormal advertising traffic, It is characterized in that The advertisement traffic classifier constructed by the method for constructing an advertisement traffic classifier according to any one of claims 1 to 4 performs advertisement traffic anomaly detection, comprising: A third acquisition module is configured to acquire a plurality of second advertisement flow data within a second preset time period; a feature extraction module configured to extract features from each of the second advertisement traffic data to obtain second traffic data features corresponding to each of the second advertisement traffic data; The determination module is configured to use the advertisement traffic classifier to determine the type of each of the second advertisement traffic data according to each of the second traffic data features.
10. An electronic device comprising a first processor and a first memory storing program instructions, It is characterized in that The first processor is configured to execute the method for building an advertisement traffic classifier according to any one of claims 1 to 4 when running the program instructions.
11. An electronic device comprising a second processor and a second memory storing program instructions, It is characterized in that The second processor is configured to execute the method for detecting advertising traffic anomalies as described in any one of claims 5 to 7 when running the program instructions.
12. A storage medium storing program instructions, It is characterized in that When the program instructions are executed, the method for constructing an advertisement traffic classifier as described in any one of claims 1 to 4 is executed.
13. A storage medium storing program instructions, It is characterized in that When the program instructions are executed, the method for detecting advertising traffic anomalies as described in any one of claims 5 to 7 is executed.
Citation Information
Patent Citations
Method, apparatus, computer device and storage medium for detecting abnormal flow
CN109040141A
Abnormal traffic detection method and device, electronic equipment and storage medium
CN113572752A