A method, device, electronic device and storage medium for black production identification
By clustering and matching business traffic data and determining black industry labels, the black industry identification and interception of real-time business traffic is achieved, and the problem of difficult to identify and intercept black industry traffic in the existing technology is solved, and the identification efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202310118548.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-02-02
AI Technical Summary
It is difficult to effectively identify and intercept the traffic of black industry. Black industry often uses IP pools, meat machines, group control and other methods to make batch requests, resulting in problems such as crawling website content and advertising click fraud.
By clustering the business traffic data in the target business scenario, a target cluster cluster is generated, and matching it with the preset reference cluster cluster, the black product label of the target cluster cluster is determined based on the matching results, so that the black product identification and intercepting of real-time business traffic is performed.
It has achieved accurate mining of black industry traffic from massive real-time traffic data, effectively discovered black industry users in network activities, and improved the efficiency and accuracy of black industry identification.
Smart Images

Figure CN116112269B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to network security technologies, and specifically, to a method, apparatus, electronic device, storage medium, and computer program product for identifying black production. Background Art
[0002] Currently, black production often uses methods such as IP pools, meat machines, and mass control to batch drive clients to send requests to sites, so as to achieve purposes such as crawling website content, advertising click fraud, and wool pulling. Among them, black production, that is, the black industry, usually refers to industries that use virus trojans, etc. to obtain benefits. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, electronic device, storage medium, and computer program product for identifying black production.
[0004] According to one aspect of the present disclosure, a method for identifying black production is provided, including:[[]]
[0005] Clustering business traffic data in a current statistical period in a target business scenario to obtain at least one target clustering cluster;
[0006] Matching the target clustering cluster with a reference clustering cluster; the reference clustering cluster includes a set black production label;
[0007] Based on the matching result, determining the black production label of the target clustering cluster according to the black production label of the matched reference clustering cluster;
[0008] According to the target clustering cluster and its black production label, identifying black production for real-time business traffic subsequent to the current statistical period;
[0009] Updating the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster, and entering the next statistical period.
[0010] According to one aspect of the present disclosure, a device for identifying black production is provided, including:[[]]
[0011] A clustering module, configured to cluster business traffic data in a current statistical period in a target business scenario to obtain at least one target clustering cluster;
[0012] A matching module, configured to match the target clustering cluster with a reference clustering cluster; the reference clustering cluster includes a set black production label;
[0013] A label determination module, configured to determine the black production label of the target clustering cluster according to the black production label of the matched reference clustering cluster based on the matching result;
[0014] An identification module, configured to identify black production in the subsequent real-time service traffic of the current statistical period according to the target clustering cluster and its black production label;
[0015] An update module, configured to update the parameters of the matching reference clustering cluster according to the parameters of the target clustering cluster, and enter the next statistical period.
[0016] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the black production identification method of any embodiment of the present disclosure.
[0020] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to execute the black production identification method of any embodiment of the present disclosure.
[0021] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, and the computer program implements the black production identification method of any embodiment of the present disclosure when executed by a processor.
[0022] According to the technology of the present disclosure, black production traffic can be accurately mined from the massive real-time traffic data of the day, and then black production users in network activities can be effectively discovered.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0025] Figure 1 is a flowchart of a black production identification method provided by an embodiment of the present disclosure;
[0026] Figure 2 is a flowchart of another black production identification method provided by an embodiment of the present disclosure;
[0027] Figure 3 is a flowchart of another black production identification method provided by an embodiment of the present disclosure;
[0028] Figure 4It is a schematic flowchart of another black production identification method provided by an embodiment of the present disclosure;
[0029] Figure 5 It is a schematic structural diagram of a black production identification device provided by an embodiment of the present disclosure;
[0030] Figure 6 It is a block diagram of an electronic device for implementing the black production identification method according to an embodiment of the present disclosure. Specific embodiments
[0031] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0032] Figure 1 It is a schematic flowchart of a black production identification method according to an embodiment of the present disclosure. This embodiment is applicable to the situation of identifying black production traffic from the service traffic generated by network activities. This method can be executed by a black production identification device, which is implemented in a software and / or hardware manner and integrated on an electronic device, such as integrated on a server device.
[0033] Specifically, referring to Figure 1 , the black production identification method includes the following steps:
[0034] S101. Cluster the service traffic data in the current statistical period under the target service scenario to obtain at least one target cluster.
[0035] In this embodiment, the target service scenario is exemplarily a machine translation scenario or an experience sharing scenario, and can also be service scenarios such as online voting, group buying, and comments, which are not specifically limited herein. The target service scenario provides a large number of data interfaces for users (including normal users and black production users) to complete corresponding services by requesting these data interfaces, and the traffic data generated by users' requests for data interfaces is the service traffic data under the target service scenario. The service traffic data may include users' interface requests. The statistical period can be selected on a daily basis or on an hourly basis, which is not specifically limited herein. For the service traffic data under the target service scenario statistically in the current statistical period, data clustering can be performed using a perspective clustering method or a single-perspective clustering method to obtain at least one target cluster.
[0036] S102. Match the target cluster with a reference cluster; the reference cluster includes pre-set black production labels.
[0037] In this embodiment, the reference clustering clusters and the black production labels of the reference clustering clusters are determined by performing clustering analysis on the full volume of historical business traffic data in advance or determined manually; the types of the black production labels include at least one of a normal cluster, an abnormal cluster, and an outlier cluster. Among them, if the black production label is a normal cluster, it means that the traffic data under the reference clustering cluster corresponding to the black production label is generated by normal users requesting relevant data interfaces; if the black production label is an abnormal cluster, it means that the traffic data under the reference clustering cluster corresponding to the black production label is generated by black production users requesting relevant data interfaces; if the black production label is an outlier cluster, it means that the traffic data under the reference clustering cluster corresponding to the black production label is generated under extreme circumstances.
[0038] In this embodiment, the target clustering cluster is matched with the reference clustering cluster, and the purpose is to determine the reference clustering cluster most similar to the target clustering cluster, so as to determine the black production label corresponding to the target clustering cluster. In an alternative embodiment, when performing the matching, the clustering center of the target clustering cluster can be matched with the clustering center of the reference clustering cluster. For example, the Euclidean distance between the clustering center of the target clustering cluster and the clustering center of each reference clustering cluster can be calculated, and the similarity size can be determined using the magnitude of the Euclidean distance.
[0039] S103. Based on the matching result, determine the black production label of the target clustering cluster according to the black production label of the matched reference clustering cluster.
[0040] Optionally, based on the matching result, assign the black production label of the matched reference clustering cluster to the black production label of the target clustering cluster.
[0041] S104. According to the target clustering cluster and its black production label, perform black production identification on the real-time business traffic in the subsequent current statistical period.
[0042] Optionally, first determine the target clustering cluster to which the real-time business traffic data belongs; then, according to the black production label of the target clustering cluster, determine whether the real-time business traffic data is black production traffic data.
[0043] Furthermore, after determining the black production traffic data from the real-time business traffic data, corresponding black production intelligence can be generated, and it can be instructed that the security defense service intercepts the corresponding black production traffic data according to the black production intelligence; among them, the black production intelligence can include data interfaces requested by black production users, device information and network information of black production users, etc. In addition, it is also possible to analyze the data interfaces in the target business scenario where black production gathers and is in a high-risk state according to the black production intelligence, so as to focus on maintaining the data interfaces in a high-risk state in the future.
[0044] S105. Update the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster, and enter the next statistical period.
[0045] In this embodiment, according to the parameters of the target clustering cluster, the parameters of the matching reference clustering cluster can be updated by means of mean calculation or assignment. For example, directly replace the parameters of the matching reference clustering cluster with the parameters of the target clustering cluster; or, perform a mean process on the parameters of the target clustering cluster and the parameters of the reference clustering cluster matching the target clustering cluster, and use the processed parameter value as the updated parameter of the reference clustering cluster. Among them, the parameters of the reference clustering cluster to be updated include the clustering center and the search radius.
[0046] It should be noted that by updating the parameters of the reference clustering cluster, the dynamic adjustment of the reference clustering cluster with the black production label is realized. In this way, the black production label corresponding to the target clustering cluster can be accurately determined, and it is avoided that because the parameters of the reference clustering cluster remain fixed, the target clustering cluster that should be given the abnormal cluster label is given the normal cluster label, resulting in the subsequent inability to accurately identify the black production traffic data from the real-time service traffic data. That is to say, if the parameters of the reference clustering cluster are not dynamically adjusted, some black production traffic in the real-time service traffic data cannot be identified.
[0047] In this embodiment, dynamically adjusting the parameters of the reference clustering cluster can ensure the accuracy of the black production label of the determined target clustering cluster; and then ensure the accuracy of identifying the black production traffic from the real-time service traffic based on the target clustering cluster and its black production label; and avoid the problem of black production traffic data caused by the inability to accurately identify some new forms of black production.
[0048] Figure 2 It is a schematic flowchart of another black production identification method according to an embodiment of the present disclosure. Please refer to Figure 2 The black production identification method includes steps S201 - S208.
[0049] In this embodiment, in order to ensure the accuracy of clustering the service traffic data within the current statistical period, a multi-perspective clustering method can be adopted to cluster the service traffic data within the current statistical period in the target service scenario, so as to obtain at least one target clustering cluster. Among them, the multi-perspective clustering method can be optionally to cluster the service traffic data of the current statistical period in the target service scenario by using at least two different clustering methods, and then fuse the clustering results obtained based on different clustering methods to obtain the multi-perspective clustering result. Among them, the multi-perspective clustering result includes at least one target clustering cluster. It should be noted here that the reason for adopting the multi-perspective clustering method for clustering is that each clustering method has its own advantages and disadvantages. For example, for the density-based clustering method, its advantage is that it can find outliers, and its disadvantages are that it cannot control the number of clustering centers and may merge two clustering clusters with obvious boundaries. If a single clustering method is adopted, the deficiencies of the clustering method itself cannot be overcome. On the contrary, if the multi-perspective clustering method is adopted, the advantages of different clustering methods can be taken into account, and the deficiencies of a single clustering method can be eliminated, so as to ensure the accuracy of the clustering result. Optionally, the process of clustering the service traffic data within the current statistical period by using the multi-perspective clustering method can refer to S201-S204.
[0050] S201. According to the preset statistical dimension, count the interface request distribution in the service traffic data within the current statistical period in the target service scenario.
[0051] Among them, the preset statistical dimension can be optionally include Internet Protocol (IP), User Identification (UID), and environmental fingerprint; among them, the environmental fingerprint can be optionally include at least one of device environmental fingerprint and network protocol fingerprint. The service traffic data of the target service scenario includes interface requests, and the interface requests can include information such as user identification and Internet protocol. Therefore, clustering the service traffic data of the target service scenario in the current statistical period is actually clustering the interface request distribution in different dimensions of the service traffic. For the convenience of subsequent clustering processing, data statistics can be performed according to different statistical dimensions before clustering. Exemplarily, referring to Table 1 below, it gives some statistical results.
[0052] Table 1
[0053] Statistical dimension / home / app / json / xxx IP1 0.1 0.2 0.7 IP2 0.3 0.3 0.4 …… …… …… ……
[0054] Among them, / home, / app / json, and / xxx respectively exemplarily represent the path addresses of data interfaces A, B, and C; taking the statistical dimension IP1 as an example, 0.1 indicates that under the IP1 dimension, the proportion of the number of requests to data interface A in the total number of interface requests; similarly, 0.2 indicates the proportion of the number of requests to interface B in the total number of interface requests. And the data 0.1, 0.2, 0.7 represents the interface request distribution under the statistical dimension IP1.
[0055] S202. Cluster the statistically obtained interface request distribution according to the first clustering method and the clustering parameters corresponding to the first clustering method to obtain at least one first clustering cluster.
[0056] In this embodiment, the first clustering method can optionally be a partitioning-based clustering method (kmeans algorithm), and the clustering parameter corresponding to the partitioning-based clustering method is the number of cluster centers. Among them, the number of cluster centers is determined by performing clustering analysis on the historical business data of the target business scenario, for example, by comparing the silhouette coefficients under different numbers of cluster centers; and the clustering analysis of the historical business data can be performed manually, automatically by a machine, or a combination of both, and no specific limitation is made here. When clustering the statistically obtained interface request distribution according to the first clustering method and the clustering parameters corresponding to the first clustering method, multiple clustering clusters are obtained, and each clustering cluster includes multiple sample points, and each sample point represents the interface request distribution under a statistical dimension. Exemplarily, see Table 2:
[0057] Table 2
[0058] Statistical dimension / home / app / json / xxx kmeans cluster id IP1 0.1 0.2 0.7 0 IP2 0.3 0.3 0.4 1 …… …… …… …… ……
[0059] Among them, the kmeans cluster id corresponding to the interface request distribution under each statistical dimension represents the clustering cluster to which this interface request distribution belongs. For example, the interface request distribution under IP1 belongs to the first clustering cluster with a kmeans cluster id of 0 after clustering.
[0060] S203. Cluster the statistically obtained interface request distribution according to the second clustering method and the clustering parameters corresponding to the second clustering method to obtain at least one second clustering cluster.
[0061] In this embodiment, the second clustering method can optionally be a density-based clustering method (dbscan algorithm), and the clustering parameters corresponding to the density-based clustering method include the clustering search radius and the minimum number of sample points in a clustering cluster; the clustering search radius is determined by performing clustering analysis on the historical business data of the target business scenario. After clustering using the second clustering method, at least one second clustering cluster can be obtained, and the sample points belonging to the same second clustering cluster have the same dbscan cluster id.
[0062] S204. Fuse the first clustering cluster and the second clustering cluster to obtain at least one target clustering cluster.
[0063] Optionally, take any second clustering cluster as the current clustering cluster. If there is an intersection between the current clustering cluster and at least one first clustering cluster, re-cluster the sample points in the current clustering cluster according to the at least one first clustering cluster with an intersection to obtain at least one target clustering cluster. Exemplarily, the current clustering cluster is the second clustering cluster with a dbscan cluster id of 3. If the current clustering cluster has intersections with three first clustering clusters with kmeans cluster ids of 4, 5, and 6, that is, the current clustering cluster has the same sample points as the three first clustering clusters with kmeans cluster ids of 4, 5, and 6. At this time, fuse with the first clustering cluster as the standard. For example, for any target sample point in the current clustering cluster, if the target sample point belongs to the first clustering cluster with a kmeans cluster id of 4 at the same time, replace the dbscan cluster id corresponding to the target sample point in the current clustering cluster with the kmeans cluster id with a value of 4 to achieve the purpose of fusion, and then obtain multiple target clustering clusters. It should be noted here that if the number of sample points in a certain first clustering cluster that intersects with the current clustering cluster is less than the preset threshold, fuse this first clustering cluster into the other first clustering cluster closest to it.
[0064] S205. Match the target clustering cluster with a reference clustering cluster; the reference clustering cluster includes pre-set black production labels.
[0065] In this embodiment, the reference clustering cluster and the black production label of the reference clustering cluster are determined by pre-cluster analysis of all historical service traffic data or manually determined; the types of the black production labels include at least one of a normal cluster, an abnormal cluster, and an outlier cluster. Among them, if the black production label is a normal cluster, it means that the traffic data under the reference clustering cluster corresponding to this black production label is generated by normal user requests for relevant data interfaces; if the black production label is an abnormal cluster, it means that the traffic data under the reference clustering cluster corresponding to this black production label is generated by black production user requests for relevant data interfaces; if the black production label is an outlier cluster, it means that the traffic data under the reference clustering cluster corresponding to this black production label is generated by extreme situations.
[0066] In this embodiment, matching the target clustering cluster with the reference clustering cluster aims to determine the reference clustering cluster most similar to the target clustering cluster in order to determine the black production label corresponding to the target clustering cluster. In an optional implementation manner, during matching, the clustering center of the target clustering cluster can be matched with the clustering center of the reference clustering cluster. For example, calculate the Euclidean distance between the clustering center of the target clustering cluster and the clustering center of each reference clustering cluster, and use the size of the Euclidean distance to determine the similarity size.
[0067] S206. Based on the matching result, determine the black production label of the target clustering cluster according to the black production label of the matched reference clustering cluster.
[0068] Optionally, based on the matching result, assign the black production label of the matched reference clustering cluster to the black production label of the target clustering cluster.
[0069] S207. According to the target clustering cluster and its black production label, perform black production identification on the real-time service traffic in the subsequent current statistical period.
[0070] Optionally, first determine the target clustering cluster to which the real-time service traffic data belongs; then, according to the black production label of the target clustering cluster, determine whether the real-time service traffic data is black production traffic data.
[0071] Furthermore, after determining the black production traffic data from the real-time service traffic data, corresponding black production intelligence can be generated, and it can be instructed that the security defense service intercepts the corresponding black production traffic data according to the black production intelligence; among them, the black production intelligence may include data interfaces requested by black production users, device information and network information of black production users, etc. In addition, it is also possible to analyze the data interfaces in the target business scenario where black production gathers and is in a high-risk state according to the black production intelligence, so as to focus on maintaining the data interfaces in a high-risk state in the future.
[0072] S208. Update the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster, and enter the next statistical period.
[0073] In this embodiment, by using the multi-perspective clustering method to cluster the service traffic data in the current statistical period under the target business scenario, the deficiencies of a single clustering method can be overcome, ensuring the accuracy of data clustering. It provides a guarantee for accurately identifying black production data subsequently.
[0074] Figure 3 It is a schematic flowchart of another black production identification method according to an embodiment of the present disclosure. Refer to Figure 3 , and the black production identification method is as follows:
[0075] S301. Cluster the service traffic data in the current statistical period under the target business scenario to obtain at least one target clustering cluster.
[0076] S302. Match the target clustering cluster with a reference clustering cluster; the reference clustering cluster includes pre-set black production labels.
[0077] S303. Based on the matching result, determine the black production label of the target clustering cluster according to the black production label of the matched reference clustering cluster.
[0078] Optionally, based on the matching result, assign the black production label of the matched reference clustering cluster to the black production label of the target clustering cluster.
[0079] After determining the black production label of the target clustering cluster, to identify the black production traffic data in the real-time service traffic data, it is only necessary to determine the target clustering cluster to which the real-time service traffic data belongs. The process of determining the target clustering cluster to which the real-time service traffic data belongs can refer to S304 - S305.
[0080] S304. Determine the similarity between the interface request distribution in the real-time service traffic data and each target clustering cluster.
[0081] Optionally, first determine the cluster vector of each target clustering cluster according to the parameters (such as the clustering center and search radius) and black production label of each target clustering cluster. Further, perform statistics on the interface request distribution in the real-time service traffic data according to a preset statistical dimension. Optionally, perform statistics in a bucketing manner. For one statistical dimension, determine a bucket for each data interface, and this bucket is used to record all interface requests for this data interface. Then, take the ratio of the number of interface requests recorded in each bucket to the total number of interface requests as the interface request distribution. Exemplarily, the statistical result can refer to Table 3:
[0082] Table 3
[0083] Statistical dimension / home / app / json / xxx Total number of requests IP1 0.1 0.2 0.7 50 IP2 0.3 0.3 0.4 100 …… …… …… …… ……
[0084] Among them, the total number of requests 50 is used to represent the total number of all interface requests under the statistical dimension of IP1; 0.1 represents the proportion of the number of requests for the data interface corresponding to / home in the total number of requests.
[0085] Further, when the number of interface requests in any statistical dimension reaches a preset quantity threshold, determine the target vector according to the interface request distribution corresponding to this statistical dimension. Exemplarily, the preset quantity threshold is 100. From Table 3, it can be seen that the total number of interface requests under the statistical dimension of IP2 reaches 100, and form a target vector with 0.3, 0.3, and 0.4. According to the target vector and the cluster vector of each target clustering cluster, determine the similarity between the interface request distribution in this statistical dimension and each target clustering cluster. Exemplarily, calculate the Euclidean distance between the target vector and the cluster vector of each target clustering cluster, and determine the similarity size according to the Euclidean distance size.
[0086] S305. Determine the target clustering cluster to which the real-time service traffic data belongs according to the similarity.
[0087] Optionally, use the target clustering cluster corresponding to the cluster vector closest to the target vector as the clustering cluster to which the real-time service traffic data belongs.
[0088] S306. Determine whether the real-time service traffic data is black production traffic data according to the black production label of the target clustering cluster.
[0089] S307. Update the parameters of the matching reference clustering cluster according to the parameters of the target clustering cluster, and enter the next statistical period.
[0090] In this embodiment, by calculating the similarity between the target vector of the interface request distribution in the statistical dimension and each cluster vector, the clustering cluster to which the interface request distribution belongs can be determined. Furthermore, according to the black production label of the clustering cluster, the black production traffic in the real-time traffic data can be quickly determined, thus improving the efficiency of black production identification.
[0091] Figure 4 It is a schematic flowchart of another black production identification method according to an embodiment of the present disclosure. Refer to Figure 4 , the black production identification method is as follows:
[0092] S401. For the target service scenario, determine the valid data interfaces in the target service scenario through the interface path fuzzy matching algorithm.
[0093] In this embodiment, there are a large number of data interfaces in the target service scenario, but the frequency or quantity of user requests for each data interface is different. The proportion of user requests for some data interfaces is very low, and black production users usually only request data interfaces with higher value. Therefore, in order to reduce the amount of data for subsequent clustering, the data interfaces in the target service scenario can be screened first. For example, the data interfaces with the proportion of user requests exceeding a preset threshold (such as one in ten thousand) can be selected; furthermore, through the interface path fuzzy matching algorithm, the valid data interfaces in the target service scenario can be selected from them. Exemplarily, the number of selected valid data interfaces is generally less than thirty. On this basis, the service traffic data in this embodiment is the traffic data generated by requesting the valid data interfaces. Then, black production identification is performed according to the steps of S402 - S406.
[0094] S402. Cluster the service traffic data in the current statistical period under the target service scenario to obtain at least one target clustering cluster.
[0095] S403. Match the target clustering cluster with the reference clustering cluster; the reference clustering cluster includes the set black production labels.
[0096] S404. Based on the matching result, determine the black production label of the target clustering cluster according to the black production label of the matching reference clustering cluster.
[0097] S405. According to the target clustering cluster and its black production label, perform black production identification on the real-time service traffic subsequent to the current statistical period.
[0098] S406. Update the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster, and enter the next statistical period.
[0099] In this embodiment, valid data interfaces are pre-screened, which can not only reduce the computational complexity of clustering, but also identify black production only for the traffic of some valid data interfaces, ensuring the efficiency of black production identification.
[0100] Figure 5 It is a schematic structural diagram of a black production identification device according to an embodiment of the present disclosure. This embodiment is applicable to the situation of identifying black production traffic from the service traffic generated by network activities. Refer to Figure 5 , the device includes:
[0101] A clustering module 501, configured to cluster the service traffic data in the current statistical period under a target service scenario to obtain at least one target clustering cluster;
[0102] A matching module 502, configured to match the target clustering cluster with a reference clustering cluster; the reference clustering cluster includes a set black production label;
[0103] A label determination module 503, configured to determine the black production label of the target clustering cluster based on the matching result and according to the black production label of the matched reference clustering cluster;
[0104] An identification module 504, configured to perform black production identification on the real-time service traffic subsequent to the current statistical period according to the target clustering cluster and its black production label;
[0105] An update module 505, configured to update the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster, and enter the next statistical period.
[0106] In some embodiments, optionally, the clustering module includes:
[0107] A statistics unit, configured to perform statistics on the interface request distribution in the service traffic data in the current statistical period under a target service scenario according to a preset statistical dimension;
[0108] A first clustering unit, configured to cluster the statistically obtained interface request distribution according to a first clustering method and the clustering parameters corresponding to the first clustering method to obtain at least one first clustering cluster;
[0109] A second clustering unit, configured to cluster the statistically obtained interface request distribution according to a second clustering method and the clustering parameters corresponding to the second clustering method to obtain at least one second clustering cluster;
[0110] A fusion unit, configured to fuse the first clustering cluster and the second clustering cluster to obtain at least one target clustering cluster.
[0111] In some embodiments, optionally, the first clustering method is a partitioning-based clustering method, and the clustering parameter corresponding to the partitioning-based clustering method is the number of cluster centers; the second clustering method is a density-based clustering method, and the clustering parameters corresponding to the density-based clustering method include a clustering search radius and the minimum number of samples in a clustering cluster.
[0112] In some embodiments, optionally, the fusion unit is further configured to:
[0113] Take any second clustering cluster as the current clustering cluster. If there is an intersection between the current clustering cluster and at least one first clustering cluster, re-cluster the sample points in the current clustering cluster according to the at least one first clustering cluster with which there is an intersection, to obtain at least one target clustering cluster.
[0114] In some embodiments, optionally, the label determination module is further configured to:
[0115] Based on the matching result, assign the black production label of the matching reference clustering cluster to the black production label of the target clustering cluster;
[0116] Wherein, the black production label of the reference clustering cluster is determined in advance by performing clustering analysis on all historical service traffic data or determined manually; the types of the black production labels include at least one of a normal cluster, an abnormal cluster, and an outlier cluster.
[0117] In some embodiments, optionally, the recognition module includes:
[0118] A determination unit, configured to determine the target clustering cluster to which the real-time service traffic data belongs;
[0119] An identification unit, configured to determine whether the real-time service traffic data is black production traffic data according to the black production label of the target clustering cluster.
[0120] In some embodiments, optionally, the determination unit includes:
[0121] A similarity determination subunit, configured to determine the similarity between the interface request distribution in the real-time service traffic data and each target clustering cluster;
[0122] A clustering cluster determination subunit, configured to determine the target clustering cluster to which the real-time service traffic data belongs according to the similarity.
[0123] In some embodiments, optionally, the similarity determination subunit is further configured to:
[0124] According to the parameters and black production labels of each target clustering cluster, determine the cluster vector of each target clustering cluster;
[0125] Statistically analyze the interface request distribution in the real-time service traffic data according to a preset statistical dimension;
[0126] When the number of interface requests in any statistical dimension reaches a preset quantity threshold, determine a target vector according to the interface request distribution corresponding to this statistical dimension;
[0127] Determine the similarity between the interface request distribution in this statistical dimension and each target clustering cluster according to the target vector and the cluster vectors of each target clustering cluster.
[0128] In some embodiments, optionally, the matching module is further configured to:
[0129] Perform similarity matching between the clustering center of the target clustering cluster and the clustering center of the reference clustering cluster.
[0130] In some embodiments, optionally, the updating module is further configured to:
[0131] Update the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster by means of mean calculation or assignment; wherein, the parameters of the reference clustering cluster to be updated include the clustering center and the search radius.
[0132] In some embodiments, optionally, it further includes:
[0133] An interface screening module, configured to determine valid data interfaces in the target business scenario through an interface path fuzzy matching algorithm for the target business scenario;
[0134] Correspondingly, the service traffic data is traffic data generated by requesting the valid data interfaces.
[0135] In some embodiments, optionally, the target business scenario is a translation scenario or an experience sharing scenario; the statistical period is at the day level or the hour level.
[0136] The black production identification device provided by the embodiments of the present disclosure can execute the black production identification method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method. The content not described in detail in this embodiment can be referred to the description in any method embodiment of the present disclosure.
[0137] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0138] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0139] Figure 6FIG. shows a schematic block diagram of an exemplary electronic device 600 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0140] As Figure 6 shown, the device 600 includes a computing unit 601 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0141] A plurality of components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0142] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the black production identification method. For example, in some embodiments, the black production identification method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the black production identification method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the black production identification method in any other suitable manner (e.g., by means of firmware).
[0143] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0144] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0145] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0146] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0147] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0148] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0149] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0150] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for identifying black production, including: Clustering the service traffic data in the current statistical period under the target service scenario to obtain at least one target clustering cluster; Matching the target clustering cluster with a reference clustering cluster; the reference clustering cluster includes a set black production label; Based on the matching result, determining the black production label of the target clustering cluster according to the black production label of the matched reference clustering cluster; According to the target clustering cluster and its black production label, identifying black production for the real-time service traffic subsequent to the current statistical period; Updating the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster, and entering the next statistical period; Among them, clustering the service traffic data in the current statistical period under the target service scenario to obtain at least one target clustering cluster includes: Statistically analyzing the interface request distribution in the service traffic data in the current statistical period under the target service scenario according to a preset statistical dimension; Clustering the statistically analyzed interface request distribution according to a first clustering method and the clustering parameters corresponding to the first clustering method to obtain at least one first clustering cluster; Clustering the statistically analyzed interface request distribution according to a second clustering method and the clustering parameters corresponding to the second clustering method to obtain at least one second clustering cluster; Fusing the first clustering cluster and the second clustering cluster to obtain at least one target clustering cluster.
2. The method according to claim 1, wherein, The first clustering method is a partitioning-based clustering method, and the clustering parameter corresponding to the partitioning-based clustering method is the number of cluster centers; the second clustering method is a density-based clustering method, and the clustering parameters corresponding to the density-based clustering method include a clustering search radius and the minimum number of samples in a clustering cluster.
3. The method according to claim 2, wherein, Fusing the first clustering cluster and the second clustering cluster to obtain at least one target clustering cluster includes: Taking any second clustering cluster as the current clustering cluster. If there is an intersection between the current clustering cluster and at least one first clustering cluster, re-clustering the sample points in the current clustering cluster according to the at least one first clustering cluster with an intersection to obtain at least one target clustering cluster.
4. The method according to claim 1, wherein, Based on the matching result, determining the black production label of the target clustering cluster according to the black production label of the matched reference clustering cluster includes: Based on the matching result, assigning the black production label of the matched reference clustering cluster to the black production label of the target clustering cluster; Among them, the black production label of the reference clustering cluster is determined by clustering analysis of the full amount of historical service traffic data in advance or determined manually; the types of black production labels include at least one of a normal cluster, an abnormal cluster, and an outlier cluster.
5. The method according to claim 1, wherein, According to the target clustering cluster and its black production label, identifying black production for the real-time service traffic subsequent to the current statistical period includes: Determining the target clustering cluster to which the real-time service traffic data belongs; According to the black production label of the target clustering cluster, determining whether the real-time service traffic data is black production traffic data.
6. The method according to claim 5, wherein, Determining the target clustering cluster to which the real-time service traffic data belongs includes: Determining the similarity between the interface request distribution in the real-time service traffic data and each target clustering cluster; Determining the target clustering cluster to which the real-time service traffic data belongs according to the similarity.
7. According to the method described in claim 6, wherein, Determining the similarity between the interface request distribution in the real-time service traffic data and each target clustering cluster includes: Determining the cluster vector of each target clustering cluster according to the parameters and black production labels of each target clustering cluster; Statistically analyzing the interface request distribution in the real-time service traffic data according to a preset statistical dimension; When the number of interface requests in any statistical dimension reaches a preset quantity threshold, determining a target vector according to the interface request distribution corresponding to this statistical dimension; Determining the similarity between the interface request distribution in this statistical dimension and each target clustering cluster according to the target vector and the cluster vectors of each target clustering cluster.
8. According to the method described in claim 1, wherein, Matching the target clustering cluster with the reference clustering cluster includes: Performing similarity matching on the clustering centers of the target clustering cluster and the reference clustering cluster.
9. According to the method described in claim 1, wherein, Updating the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster includes: Updating the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster by means of mean calculation or assignment; wherein, the parameters of the reference clustering cluster to be updated include the clustering center and the search radius.
10. According to the method described in claim 1, further including: For the target service scenario, determining the valid data interfaces in the target service scenario through an interface path fuzzy matching algorithm; Correspondingly, the service traffic data is the traffic data generated by requesting the valid data interfaces.
11. According to the method described in claim 1, wherein, The target service scenario is a translation scenario or an experience sharing scenario; the statistical period is at the day level or the hour level.
12. A black production identification device, including: A clustering module, configured to cluster the service traffic data in the current statistical period under the target service scenario to obtain at least one target clustering cluster; A matching module, configured to match the target clustering cluster with a reference clustering cluster; the reference clustering cluster includes pre-set black production labels; A label determination module, configured to determine the black production label of the target clustering cluster based on the matching result according to the black production label of the matched reference clustering cluster; An identification module, configured to perform black production identification on the real-time service traffic subsequent to the current statistical period according to the target clustering cluster and its black production label; An update module, configured to update the parameters of the matched reference clustering cluster according to the parameters of the target clustering cluster and enter the next statistical period; wherein, the clustering module includes: A statistical unit, configured to statistically analyze the interface request distribution in the service traffic data in the current statistical period under the target service scenario according to a preset statistical dimension; A first clustering unit, configured to cluster the statistically analyzed interface request distribution according to a first clustering method and the clustering parameters corresponding to the first clustering method to obtain at least one first clustering cluster; A second clustering unit, configured to cluster the statistically obtained interface request distribution according to a second clustering method and clustering parameters corresponding to the second clustering method, so as to obtain at least one second clustering cluster; A fusion unit, configured to fuse the first clustering cluster and the second clustering cluster to obtain at least one target clustering cluster.
13. The apparatus according to claim 12, wherein, the first clustering method is a partitioning-based clustering method, and the clustering parameter corresponding to the partitioning-based clustering method is the number of cluster centers; the second clustering method is a density-based clustering method, and the clustering parameters corresponding to the density-based clustering method include a clustering search radius and a minimum number of samples of a clustering cluster.
14. The apparatus according to claim 13, wherein, the fusion unit is further configured to: Take any second clustering cluster as the current clustering cluster. If there is an intersection between the current clustering cluster and at least one first clustering cluster, re-cluster the sample points in the current clustering cluster according to the at least one first clustering cluster with an intersection to obtain at least one target clustering cluster.
15. The apparatus according to claim 12, wherein, the label determination module is further configured to: Based on the matching result, assign the black production label of the matching reference clustering cluster to the black production label of the target clustering cluster; wherein, the black production label of the reference clustering cluster is determined in advance by clustering analysis of all historical service traffic data or determined manually; the types of the black production label include at least one of a normal cluster, an abnormal cluster, and an outlier cluster.
16. An electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the black production identification method according to any one of claims 1-11.
17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the black production identification method according to any one of claims 1-11.
18. A computer program product, comprising a computer program which, when executed by a processor, implements the black production identification method according to any one of claims 1-11.
Citation Information
Patent Citations
Text recognition method and device, storage medium and electronic equipment
CN112256880A
Abnormality detection method and device, electronic equipment and computer readable medium
CN114186626A