Abnormal website identification method and system
By dynamically adjusting the clustering center and structure, and optimizing the abnormal website recognition method, the problem of distortion of clustering results in the existing technology is solved, and the accuracy of abnormal website recognition is improved.
Patent Information
- Application Number
- CN202510941818.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-09
AI Technical Summary
In the prior art, the abnormal website recognition method relies on static clustering strategies and cannot dynamically adjust the clustering center and structure, resulting in distortion of clustering results when website traffic or attributes changes, reducing the accuracy of abnormal recognition.
By determining the first clustering result of the initial website and seed website clustering set, the second cluster center website is obtained, and when the distance of the center website exceeds the threshold, the third cluster center website and the target clustering set are calculated, the clustering structure is optimized, and the website data distribution changes are adapted.
The recognition accuracy of abnormal websites is improved, and the stability and accuracy of clustering results are ensured by dynamically adjusting the clustering center and structure.
Smart Images

Figure CN120455169A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Internet technology, and in particular to a method and system for identifying abnormal websites. Background Art
[0002] Abnormal websites, such as malware-spreading websites and websites with illegal content, pose a serious threat to network security and user rights. Prior art methods for identifying abnormal websites can cluster known abnormal websites based on their characteristic information. This can then be used to identify abnormal websites based on the similarity between the cluster set and the characteristic information of the website to be identified. However, existing website clustering methods typically rely on static clustering strategies and are unable to dynamically adjust the cluster center and cluster structure based on the website's characteristic information. Consequently, when website traffic or attributes change, the fixed cluster center struggles to adapt to the new data distribution, leading to distorted clustering results and reduced accuracy in identifying abnormalities.
[0003] Therefore, how to improve the accuracy of identifying abnormal websites has become an urgent problem to be solved. Summary of the Invention
[0004] In response to the above technical problems, the present invention adopts a method for identifying abnormal websites, which includes the following steps: S1. Determine a first clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified within a preset time period and the website information corresponding to the first cluster center website corresponding to each seed website cluster set, wherein the first clustering result includes the corresponding assigned seed website cluster set and the website cluster set that is not assigned to the seed website cluster set, and the seed website cluster set includes several seed websites for characterizing abnormal websites.
[0005] S2. According to the first clustering result corresponding to each initial website, all initial websites corresponding to each seed website cluster set are obtained.
[0006] S3: For any seed website cluster set, obtain the second cluster center website corresponding to the current seed website cluster set based on all initial websites and all seed websites corresponding to the current seed website cluster set.
[0007] S4, obtaining the distance to the center website corresponding to the current seed website cluster set according to the website information of the first cluster center website and the website information of the second cluster center website corresponding to the current seed website cluster set.
[0008] S5. If the distance to the center website corresponding to the current seed website cluster set is greater than the preset first distance threshold, then based on the website information corresponding to all seed websites, the first cluster center website, and the second cluster center website corresponding to the current seed website cluster set, the third cluster center website and the target cluster set corresponding to the current seed website cluster set are obtained.
[0009] S6. Determine the second clustering result corresponding to each initial website according to the website information corresponding to each initial website to be identified and the website information corresponding to each third cluster center website, wherein the second clustering result includes the corresponding assigned target cluster set and the ones not assigned to the target cluster set.
[0010] S7: For any initial website, if the second clustering result corresponding to the current initial website is the corresponding assigned target cluster set, the current initial website is determined to be an abnormal website.
[0011] The present invention also provides an abnormal website identification system, which includes: The first website clustering module is used to determine the first clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified within a preset time period and the website information corresponding to the first cluster center website corresponding to each seed website cluster set, wherein the first clustering result includes the corresponding assigned seed website cluster set and the website cluster set that is not assigned to the seed website cluster set, and the seed website cluster set includes several seed websites for characterizing abnormal websites.
[0012] The cluster set analysis module is used to obtain all initial websites corresponding to each seed website cluster set according to the first clustering result corresponding to each initial website.
[0013] The second cluster center acquisition module is used to acquire the second cluster center website corresponding to any seed website cluster set according to all initial websites and all seed websites corresponding to the current seed website cluster set.
[0014] The center website distance acquisition module is used to acquire the center website distance corresponding to the current seed website cluster set based on the website information of the first cluster center website and the website information of the second cluster center website corresponding to the current seed website cluster set.
[0015] The target cluster set acquisition module is used to obtain the third cluster center website and the target cluster set corresponding to the current seed website cluster set based on the website information corresponding to all seed websites, the first cluster center website and the second cluster center website corresponding to the current seed website cluster set if the distance to the center website corresponding to the current seed website cluster set is greater than the preset first distance threshold.
[0016] The second website clustering module is used to determine the second clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified and the website information corresponding to each third cluster center website, wherein the second clustering result includes the corresponding assigned target cluster set and the ones not assigned to the target cluster set.
[0017] The abnormal website identification module is used to determine, for any initial website, if the second clustering result corresponding to the current initial website is the corresponding assigned target cluster set, that the current initial website is an abnormal website.
[0018] The present invention has at least the following beneficial effects: comparing the initial website with the first cluster center website of the seed website cluster set according to the website information, determining the first clustering result, realizing the preliminary classification of the initial website, and determining the second cluster center website in combination with all the initial websites and all the seed websites corresponding to each seed website cluster set, so as to measure the change of the cluster center after the addition of the initial website according to the center website distance between the first cluster center website and the second cluster center website, thereby characterizing the clustering effect of the seed website cluster set, and when the center website distance is greater than the preset first distance threshold, determining the third cluster center website and the target cluster set by calculating the reference center feature and the fourth relative distance, completing the optimization of the clustering structure, so that the target cluster set can better adapt to the distribution characteristics of the website data, thereby re-clustering the initial website according to the optimized target cluster set, improving the accuracy of the second clustering result, and thereby improving the accuracy of identifying abnormal websites. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0020] Figure 1 This is a flow chart of a method for identifying abnormal websites provided in Example 1 of the present invention; Figure 2 This is a schematic diagram of an abnormal website identification system provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It is understood that, where appropriate, the above-mentioned terms used to distinguish similar objects can be interchanged so that the present invention can also implement other embodiments other than the above-mentioned illustrated embodiments or described embodiments. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] Example 1 This embodiment provides a method for identifying abnormal websites. Figure 1 As shown, the abnormal website identification method includes the following steps: S1. Determine a first clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified within a preset time period and the website information corresponding to the first cluster center website corresponding to each seed website cluster set, wherein the first clustering result includes the corresponding assigned seed website cluster set and the website cluster set that is not assigned to the seed website cluster set, and the seed website cluster set includes several seed websites for characterizing abnormal websites.
[0024] The initial website to be identified is a target for which it has yet to be determined whether it is an abnormal website. Website information includes characteristic data from multiple dimensions, such as the website's network activity, website type, and regional characteristics. The preset time period is a specific time range set before abnormal website identification. By capturing relevant information about each website within the preset time period, the relevant characteristics of each website over a period of time can be understood, providing data support for subsequent cluster analysis and abnormal website identification. The specific duration of the preset time period can be determined by the implementer based on actual needs and analysis objectives. For example, the preset time period can be selected as one day, one week, or one month.
[0025] Seed sites are known anomalous sites and serve as the basis for identifying anomalies in the initial site. A seed site cluster is an initial cluster grouping composed of several seed sites that share high similarity across multiple dimensions, such as network activity, site type, and regional characteristics. The first cluster center site is the representative site for each seed site cluster. This is determined by calculating the feature mean or centroid of all seed sites within the seed site cluster and is used to characterize the typical characteristics of the seed site cluster.
[0026] The first clustering result is a preliminary classification of the initial website based on the similarity between the website information corresponding to the initial website and the website information corresponding to each first cluster center website. A website assigned to a seed website cluster set indicates that the feature similarity between the initial website and the seed website cluster set meets the requirement and is therefore classified into that seed website cluster set. A website not assigned to a seed website cluster set indicates that the feature difference between the initial website and all existing seed website cluster sets is too great to be classified into an existing seed website cluster set.
[0027] In a specific embodiment, the preset time period includes a plurality of preset time slices, the website information includes data packet flow, website address, and time difference, and S1 includes the following steps: S11, according to the data packet flow corresponding to each seed website in each preset time slice, obtain the traffic curve corresponding to each seed website, wherein the traffic curve takes the preset time slice as the independent variable and takes the data packet flow corresponding to each preset time slice as the dependent variable.
[0028] S12, clustering all seed websites according to the traffic curve, website address and time difference corresponding to each seed website, obtaining several seed website cluster sets, and determining the seed website corresponding to the cluster center of each seed website cluster set as the first cluster center website corresponding to each seed website cluster set, wherein the time difference is the time difference between the time zone corresponding to the seed website and the target time zone.
[0029] The preset time period is divided into a plurality of continuous and non-overlapping preset time slices to form a discrete time series, which is convenient for measuring the time distribution of the data packet flow of the seed website.
[0030] Packet traffic is the number and rate of packets transmitted by a website within a preset time slot, reflecting the intensity and pattern of network activity for that website. A website address, consisting of structural information such as the domain name and path, characterizes the website's type and domain. Time difference is the difference between the time zone where the website's server is located and the target time zone, used to analyze the regional characteristics and temporal patterns of website services. The target time zone can be set by the implementer based on actual circumstances, aligning all websites to the same time zone.
[0031] Statistics are collected on the data packet traffic of each seed website in each time slice, and a traffic curve is constructed with the preset time slice as the horizontal axis and the data packet traffic as the vertical axis. The traffic data corresponding to all seed websites are converted into a traffic curve with a unified dimension to facilitate subsequent clustering analysis. The traffic curve is used to intuitively reflect the distribution pattern of the traffic of the corresponding seed website over time, which serves as the basis for identifying abnormal traffic patterns.
[0032] We perform multi-dimensional feature fusion on the traffic curve, website address, and time difference features corresponding to each seed website. Clustering algorithms such as K-means and DBSCAN are then used to cluster the fused features of the seed websites to form a seed website cluster set. The center of each seed website cluster set is calculated by calculating the centroid of all features of all members in the cluster. The seed website closest to the centroid is selected as the first cluster center website, which is used to represent the typical characteristics of the corresponding seed website cluster set.
[0033] As mentioned above, by combining the characteristics of data packet flow, address and time difference, the characteristics of seed websites in multiple dimensions such as network activity, website type, and regional characteristics are characterized, and the characteristics after multi-dimensional fusion are clustered, which improves the clustering accuracy of the seed website cluster set.
[0034] In a specific embodiment, each seed website transmits a number of data packets to a target receiving end within a preset time period, and S11 further includes the following steps: S111, according to the time point when each seed website transmits each corresponding data packet to the target receiving end within the k-1th initial time slice, obtain the time difference when each seed website transmits the corresponding two adjacent data packets to the target receiving end within the k-1th initial time slice, where k is an integer greater than 1, and the initial value is 2, and the time length corresponding to the first initial time slice is a preset length.
[0035] S112 , obtaining the standard deviation of the time differences corresponding to all seed websites in the k-1th initial time slice according to all the time differences corresponding to each seed website in the k-1th initial time slice.
[0036] S113, based on the data packet flow corresponding to each seed website in the 1st to k-1th initial time slices and the time length corresponding to the 1st to k-1th initial time slices, obtain the flow density per unit time corresponding to each seed website in the 1st to k-1th initial time slices.
[0037] S114: Determine the average value of the traffic density per unit time of all seed websites in the 1st to k-1th initial time slices as the predicted traffic density of all seed websites in the kth initial time slice.
[0038] S115. The time length of the k-1th initial time slice is obtained according to the time length of the k-1th initial time slice, the standard deviation of the time difference corresponding to all seed websites in the k-1th initial time slice, the predicted traffic density corresponding to the k-th initial time slice, the preset standard traffic density and the preset adjustment coefficient.
[0039] S116, update k=k+1, repeat S111 until the sum of the time lengths of all initial time slices is greater than or equal to the preset time period, determine the first to the second-to-last initial time slices as the first to the second-to-last preset time slices, and use the end time point of the second-to-last initial time slice as the start time point of the last preset time slice, and use the end time point of the preset time period as the end time point of the last preset time slice, and update to obtain the last preset time slice.
[0040] The time intervals between the arrival of adjacent data packets at each seed website within the k-1th initial time slice are calculated, for example, the time interval between the first and second packets, and the time interval between the second and third packets. These time intervals reflect the burstiness and regularity of the packet traffic corresponding to the seed website. The standard deviation of the time differences for all seed websites within the k-1th initial time slice is calculated to reflect the stability of the packet traffic. Correspondingly, the larger the standard deviation, the worse the stability, and shorter time slices are required to more precisely capture traffic changes. In other words, the larger the standard deviation, the shorter the time span of the corresponding time point.
[0041] The average of the historical traffic density of all seed websites before the kth initial time slice is used as the predicted traffic density for the kth initial time slice. Accordingly, when traffic is dense, a more refined time division is required as the basis for analysis. In other words, the greater the predicted traffic density, the shorter the corresponding time slice length.
[0042] The preset standard traffic density represents a baseline traffic density level and is used to measure the relative situation of the current actual traffic density, so that adjustments can be made during the calculation of the time slice length based on the difference between the traffic density and the standard value. Specifically, the specific value of the preset standard traffic density can be obtained by the implementer based on actual circumstances. For example, it can be directly set based on the experience of network security experts and the needs of specific business scenarios, or it can be obtained based on the average or median traffic density data of multiple seed websites over a period of time.
[0043] The preset adjustment coefficient controls the adjustment range of the time slice. Accordingly, the larger the preset adjustment coefficient, the greater the influence of the standard deviation and traffic density on the time slice length adjustment, and the greater the time slice length change. The smaller the preset adjustment coefficient, the weaker the influence of the standard deviation and traffic density on the time slice length adjustment, and the smaller the time slice length change.
[0044] The time length of the k-1th initial time slice is the basic value for calculating the length of the kth initial time slice, and is enlarged or reduced according to the changes in the preset adjustment coefficient, standard deviation and traffic density.
[0045] For example, according to the time length Δt of the k-1th initial time slice k-1 , the standard deviation of the time difference of all seed websites in the k-1th initial time slice σ, the predicted traffic density ρ corresponding to the kth initial time slice, the preset standard traffic density ρ m And the preset adjustment coefficient α, the time length Δt of the kth initial time slice is obtained k =Δt k-1 ×(1+α×((2 / (1+e -σ ))-1)×(1-ρ / ρ m )).
[0046] As mentioned above, the time length of the kth initial time slice is calculated through a dynamic adjustment function after comprehensively considering the time length of the k-1th initial time slice, the standard deviation of the time difference, the predicted traffic density, the preset standard traffic density and the preset adjustment coefficient. The parameters of multiple dimensions interact with each other and jointly determine the dynamic change of the time slice length, thereby improving the accuracy of the segmentation of the preset time slice and thus improving the versatility of the abnormal website identification method.
[0047] In a specific embodiment, S12 further includes the following steps: S121 , for each seed website, extract features of the traffic curve corresponding to the current seed website to obtain traffic features corresponding to the current seed website.
[0048] S122: Extract features of the website address corresponding to the current seed website to obtain address features corresponding to the current seed website.
[0049] S123: Map the time difference corresponding to the current seed website to a preset numerical range to obtain the time difference feature corresponding to the current seed website.
[0050] S124: Combine the traffic characteristics, address characteristics, and time difference characteristics corresponding to the current seed website to obtain the website characteristics corresponding to the current seed website.
[0051] S125 , clustering all seed websites according to website features corresponding to each seed website to obtain several seed website cluster sets, and determining the seed website corresponding to the cluster center of each seed website cluster set as the first cluster center website corresponding to each seed website cluster set.
[0052] The traffic curve corresponding to the current seed website is extracted to obtain traffic characteristics corresponding to the current seed website, which characterizes the network activity corresponding to the current seed website. For example, the mean, variance, skewness, peak value, and other values of the traffic curve can be calculated to describe the traffic distribution characteristics. The peak time, peak duration, traffic rise rate, and fall rate can be summarized to characterize the traffic fluctuation pattern. Then, the various values are combined to obtain the characteristic vector corresponding to the traffic curve as the traffic feature.
[0053] Extract features from the website address corresponding to the current seed website to obtain the address features corresponding to the current seed website and characterize the network type corresponding to the current seed website. For example, the address structure features can be described based on the domain name length, path depth, and number of parameters corresponding to the website address. The encoding features can be described by processing the domain name using a hash algorithm or one-hot encoding technology. The address structure features and encoding features can then be combined to obtain the address features corresponding to the website address.
[0054] Map the time difference corresponding to the current seed website to a preset numerical range to obtain the time difference feature corresponding to the current seed website, which represents the regional characteristics of the current seed website. For example, the time difference can be normalized and the normalized result can be used as the time difference feature to avoid affecting the clustering effect due to large absolute differences in the time difference.
[0055] Traffic, address, and time difference features are sequentially concatenated into a high-dimensional vector to comprehensively represent the network activity, network type, and regional characteristics of the current seed website. Based on the similarity between website features, the similarity between seed websites in terms of network activity, website type, and regional characteristics is measured, serving as the basis for clustering seed websites.
[0056] Furthermore, clustering algorithms such as K-means and DBSCAN are used to cluster the fused features of the seed sites to form a seed site cluster set. The centroid of the features of all members in each seed site cluster set is calculated, and the seed site closest to the centroid is selected as the first cluster center site, which is used to represent the typical features of the corresponding seed site cluster set.
[0057] As described above, the traffic curve, website address, and time difference corresponding to the seed website are quantified into multi-dimensional features and spliced into a website feature in the form of a high-dimensional vector, which comprehensively represents the network activity, network type, and regional characteristics corresponding to the current seed website, thereby improving the accuracy of clustering the seed website.
[0058] In a specific embodiment, S1 further includes the following steps: S101 , obtaining a traffic curve corresponding to each initial website according to the data packet traffic corresponding to each initial website in each preset time slice.
[0059] S102: Acquire website features corresponding to each initial website based on the traffic curve, website address, and time difference corresponding to each initial website.
[0060] S103 : Acquire a first relative distance between each initial website and each first cluster center website according to the website feature corresponding to each initial website and the website feature corresponding to each first cluster center website.
[0061] S104 , for any seed website cluster set, obtain a second relative distance between each seed website in the current seed website cluster set and the first cluster center website based on the website features of each seed website in the current seed website cluster set and the website features of the first cluster center website.
[0062] S105: Determine the maximum value of the second relative distances between all seed websites in the current seed website cluster set and the first cluster center website as the second distance threshold corresponding to the current seed website cluster set.
[0063] S106: For each initial website, if the first relative distance between the current initial website and the first cluster center website in the current seed website cluster set is less than the second distance threshold corresponding to the current seed website cluster set, the current seed website cluster set is determined as the candidate cluster set corresponding to the current initial website.
[0064] S107, if the number of candidate cluster sets corresponding to the current initial website is greater than 0, then based on the first relative distance between the current initial website and the first cluster center website in each corresponding candidate cluster set, the candidate cluster set corresponding to the smallest first relative distance is determined as the seed website cluster set corresponding to the current initial website, and the first clustering result corresponding to the current initial website is determined as the corresponding assigned seed website cluster set.
[0065] S108: If the number of candidate cluster sets corresponding to the current initial website is 0, it is determined that the first clustering result corresponding to the current initial website is not allocated to the seed website cluster set.
[0066] The method for obtaining the website features corresponding to each initial website is consistent with the method for obtaining the website features corresponding to the seed website.
[0067] Using distance calculation methods such as Euclidean distance, cosine similarity, or Mahalanobis distance, a first relative distance is calculated between the website features of the initial website and the website features of each first cluster center website to characterize the similarity between the initial website and each first cluster center website, and a second relative distance is calculated between each seed website and the first cluster center website to characterize the similarity between each website and the cluster center in the corresponding seed website cluster set. The maximum second relative distance in the current seed website cluster set can be used as a boundary measure for the seed website cluster set, and is used to determine whether the similarity between the initial website and the corresponding seed website cluster set is within an acceptable range by comparing the first relative distance between the initial website and the first cluster center website and the second distance threshold when assigning the initial website to the seed website cluster set.
[0068] Correspondingly, if the first relative distance between the current initial website and the first cluster center website in the current seed website cluster set is less than the second distance threshold corresponding to the current seed website cluster set, indicating that the current initial website is within the acceptable range of the current seed website cluster set, the current seed website cluster set is determined as the candidate cluster set corresponding to the current initial website.
[0069] Furthermore, for each initial website, the candidate cluster set with the closest distance is screened out from all corresponding candidate cluster sets as the seed website cluster set correspondingly assigned to each initial website.
[0070] In the above, based on the similarity between each website in each seed website cluster set and the cluster center, the second distance threshold corresponding to each seed website cluster set is determined, which overcomes the limitation of the fixed threshold when the data is unevenly distributed, and uses the second distance threshold as the boundary measurement basis of the seed website cluster set, thereby improving the screening accuracy of the candidate cluster set, and further improving the accuracy of the first clustering result corresponding to each initial website.
[0071] S2. According to the first clustering result corresponding to each initial website, all initial websites corresponding to each seed website cluster set are obtained.
[0072] Initial websites that have a corresponding seed website cluster set are classified according to the seed website cluster set to which they belong. Initial websites that are not assigned to a seed website cluster set can be individually classified as unrelated initial websites to prevent them from interfering with the analysis results of the clustering effect of subsequent seed website cluster sets.
[0073] Correspondingly, each seed website cluster set can obtain all related initial websites, thereby forming a complete cluster set structure.
[0074] S3: For any seed website cluster set, obtain the second cluster center website corresponding to the current seed website cluster set based on all initial websites and all seed websites corresponding to the current seed website cluster set.
[0075] In a specific embodiment, S3 includes the following steps: S31 , calculating a centroid vector corresponding to the current seed website cluster set based on the website features of all initial websites and the website features of all seed websites corresponding to the current seed website cluster set.
[0076] S32, based on the centroid vector corresponding to the current seed website cluster set and the website features of each initial website and each seed website corresponding to the current seed website cluster set, obtain a third relative distance between the centroid vector corresponding to the current seed website cluster set and the corresponding each initial website and each seed website.
[0077] S33: Determine the initial website or seed website corresponding to the smallest third relative distance as the second cluster center website corresponding to the current seed website cluster set.
[0078] Among them, the centroid vector is the average value of the website characteristics of all initial websites and all seed websites, representing the geometric center of the current seed website cluster set. It can eliminate individual differences in the set and focus on the overall characteristics of the current seed website cluster set.
[0079] Using distance calculation methods such as Euclidean distance, cosine similarity or Mahalanobis distance, the third relative distance between the centroid vector in the current seed website cluster set and each initial website or each seed website is calculated to characterize the similarity between each initial website or each seed website and the overall characteristics of the current seed website cluster set, thereby screening out the initial website or seed website with the smallest third relative distance, that is, the initial website or seed website with the highest similarity to the overall characteristics of the current seed website cluster set, as the second cluster center website corresponding to the current seed website cluster set, that is, screening out the website that can better represent the overall characteristics of the current seed website cluster set.
[0080] As mentioned above, the second cluster center is closer to the overall characteristics of the seed website cluster set after the initial website is added than the first cluster center, avoiding the clustering bias caused by relying solely on the seed website, and providing a data basis for analyzing the clustering effect of the seed website cluster set.
[0081] S4, obtaining the distance to the center website corresponding to the current seed website cluster set according to the website information of the first cluster center website and the website information of the second cluster center website corresponding to the current seed website cluster set.
[0082] The first cluster center is the center obtained by clustering the initial seed sites, representing the typical characteristics or average state of the seed site set. The second cluster center is the center re-determined after considering all the initial sites and all seed sites corresponding to the current seed site cluster set, reflecting the characteristics of the cluster set more comprehensively.
[0083] Use distance calculation methods such as Euclidean distance, cosine similarity or Mahalanobis distance to calculate the distance between the website information of the first cluster center website and the website information of the second cluster center website corresponding to the current seed website cluster set. As the center website distance corresponding to the current seed website cluster set, it can measure the change of the cluster center after the addition of the new initial website, so as to understand the degree of change of the characteristics of the entire cluster set, such as the movement or change amplitude of the center feature, and then characterize the quality of the clustering effect of the seed website cluster set.
[0084] Correspondingly, a larger center-site distance indicates a significant shift in the cluster center after the addition of the initial site, meaning that the characteristics of the cluster set have changed significantly. This may be because the newly added initial site has significant characteristics different from the original seed site, or the cluster division is not reasonable. This indicates that the clustering effect of the seed site cluster set is poor and needs further adjustment. Conversely, a smaller center-site distance indicates that the cluster set is relatively stable, the newly added initial site has little impact on the overall cluster characteristics, the clustering results have a certain degree of reliability and stability, and the clustering effect of the seed site cluster set is better.
[0085] As described above, based on the website information of the first cluster center website and the website information of the second cluster center website, the center website distance corresponding to the current seed website cluster set is obtained, which is used to measure the changes in the cluster center after the addition of the new initial website, and then characterize the clustering effect of the seed website cluster set, providing a data basis for whether to optimize the cluster set in the future.
[0086] S5. If the distance to the center website corresponding to the current seed website cluster set is greater than the preset first distance threshold, then based on the website information corresponding to all seed websites, the first cluster center website, and the second cluster center website corresponding to the current seed website cluster set, the third cluster center website and the target cluster set corresponding to the current seed website cluster set are obtained.
[0087] In a specific embodiment, S5 includes the following steps: S51 , determining an average feature of the website features corresponding to the first cluster center website and the website features corresponding to the second cluster center website corresponding to the current seed website cluster set as a reference center feature corresponding to the current seed website cluster set.
[0088] S52 : Obtain a fourth relative distance between each seed website in the current seed website cluster set and the reference center feature based on the website features and reference center features corresponding to all seed websites corresponding to the current seed website cluster set.
[0089] S53: Determine the seed website corresponding to the smallest fourth relative distance as the third cluster center website corresponding to the current seed website cluster set.
[0090] S54: forming a target cluster set with the third cluster center website as the cluster center based on all seed websites in the current seed website cluster set.
[0091] The fourth relative distance between each seed website and the reference center feature is calculated using a distance calculation method such as Euclidean distance, cosine similarity or Mahalanobis distance.
[0092] The reference center feature balances the initial clustering results based on the seed website and the comprehensive clustering results combined with the initial website. According to the fourth relative distance, the seed website closest to the reference center feature is determined as the third cluster center website, so that the target clustering set with the third cluster center website as the clustering center can not only retain the core information of the seed website, but also absorb the trend characteristics of the initial website. After absorbing the initial website, it has little impact on the overall clustering characteristics, reducing the instability caused by the initial clustering deviation or the initial website fluctuation, and ensuring the reliability and stability of the clustering results.
[0093] As mentioned above, when the difference between the first cluster center and the second cluster center is too large, by calculating the reference center features and selecting the closest seed website as the third cluster center, a more robust target cluster set is generated, which improves the clustering accuracy of the seed website and subsequent initial websites, and thus improves the recognition accuracy of subsequent abnormal websites.
[0094] S6. Determine the second clustering result corresponding to each initial website according to the website information corresponding to each initial website to be identified and the website information corresponding to each third cluster center website, wherein the second clustering result includes the corresponding assigned target cluster set and the ones not assigned to the target cluster set.
[0095] Among them, similar to determining the first clustering result corresponding to each initial website, based on the new target clustering set and the corresponding third clustering center website, the similarity between the website information corresponding to the initial website and the website information corresponding to the third clustering center website is measured, and then the second clustering result corresponding to each initial website is obtained as the clustering result of the optimized initial website.
[0096] As described above, the initial websites are clustered based on the optimized target cluster set and the third cluster center website, thereby improving the accuracy of the second clustering result.
[0097] S7: For any initial website, if the second clustering result corresponding to the current initial website is the corresponding assigned target cluster set, the current initial website is determined to be an abnormal website.
[0098] In a specific embodiment, the abnormal website identification method further includes the following steps: For any initial website, if the second clustering result corresponding to the current initial website is not assigned to the target cluster set, the current initial website is determined to be a normal website.
[0099] In the above, the initial website is compared with the first cluster center website of the seed website cluster set according to the website information to determine the first clustering result, thereby realizing the preliminary classification of the initial website. The second cluster center website is determined in combination with all the initial websites and all the seed websites corresponding to each seed website cluster set, so as to measure the change of the cluster center after the addition of the initial website according to the center website distance between the first cluster center website and the second cluster center website, and characterize the clustering effect of the seed website cluster set. When the center website distance is greater than the preset first distance threshold, the third cluster center website and the target cluster set are determined by calculating the reference center feature and the fourth relative distance, thereby completing the optimization of the clustering structure, so that the target cluster set can better adapt to the distribution characteristics of the website data, thereby re-clustering the initial website according to the optimized target cluster set, improving the accuracy of the second clustering result, and thereby improving the accuracy of identifying abnormal websites.
[0100] Example 2 This embodiment 2 provides a system for identifying abnormal websites. Figure 2 As shown, the abnormal website identification system includes: The first website clustering module 21 is used to determine the first clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified within a preset time period and the website information corresponding to the first cluster center website corresponding to each seed website cluster set, wherein the first clustering result includes the corresponding assigned seed website cluster set and the seed website cluster set that is not assigned to the seed website cluster set, and the seed website cluster set includes several seed websites for characterizing abnormal websites.
[0101] The cluster set analysis module 22 is configured to obtain all initial websites corresponding to each seed website cluster set according to the first clustering result corresponding to each initial website.
[0102] The second cluster center acquisition module 23 is configured to acquire, for any seed website cluster set, a second cluster center website corresponding to the current seed website cluster set based on all initial websites and all seed websites corresponding to the current seed website cluster set.
[0103] The center website distance acquisition module 24 is configured to acquire the center website distance corresponding to the current seed website cluster set based on the website information of the first cluster center website and the website information of the second cluster center website corresponding to the current seed website cluster set.
[0104] The target cluster set acquisition module 25 is used to obtain the third cluster center website and the target cluster set corresponding to the current seed website cluster set based on the website information corresponding to all seed websites, the first cluster center website and the second cluster center website corresponding to the current seed website cluster set if the distance between the center website corresponding to the current seed website cluster set is greater than the preset first distance threshold.
[0105] The second website clustering module 26 is used to determine the second clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified and the website information corresponding to each third cluster center website, wherein the second clustering result includes the corresponding assigned target cluster set and the ones not assigned to the target cluster set.
[0106] The abnormal website identification module 27 is used to determine that the current initial website is an abnormal website if the second clustering result corresponding to the current initial website is the corresponding assigned target cluster set. In a specific embodiment, the preset time period includes a plurality of preset time slices, the website information includes data packet flow, website address, and time difference, and the first website clustering module 21 includes: The first traffic curve acquisition submodule is used to obtain the traffic curve corresponding to each seed website based on the data packet flow corresponding to each seed website in each preset time slice, wherein the traffic curve uses the preset time slice as an independent variable and the data packet flow corresponding to each preset time slice as a dependent variable.
[0107] The first clustering submodule is used to cluster all seed websites according to the traffic curve, website address and time difference corresponding to each seed website, obtain several seed website cluster sets, and determine the seed website corresponding to the cluster center of each seed website cluster set as the first cluster center website corresponding to each seed website cluster set, wherein the time difference is the time difference between the time zone corresponding to the seed website and the target time zone.
[0108] In a specific embodiment, each seed website transmits a number of data packets to a target receiving end within a preset time period, and the first traffic curve acquisition submodule further includes: The time difference acquisition unit is used to obtain the time difference when each seed website transmits the corresponding two adjacent data packets to the target receiving end within the k-1th initial time slice according to the time point when each seed website transmits the corresponding each data packet to the target receiving end within the k-1th initial time slice, wherein k is an integer greater than 1, and the initial value is 2, and the time length corresponding to the first initial time slice is a preset length.
[0109] The standard deviation obtaining unit is used to obtain the standard deviation of the time differences corresponding to all seed websites in the k-1th initial time slice according to all the time differences corresponding to each seed website in the k-1th initial time slice.
[0110] The traffic density acquisition unit is used to obtain the traffic density per unit time corresponding to each seed website in the 1st to k-1th initial time slices based on the data packet traffic corresponding to each seed website in the 1st to k-1th initial time slices and the time length corresponding to the 1st to k-1th initial time slices.
[0111] The predicted traffic density acquisition unit is used to determine the average value of the traffic density of all seed websites in the unit time corresponding to the 1st to k-1th initial time slices as the predicted traffic density of all seed websites corresponding to the kth initial time slice.
[0112] The first time length acquisition unit is used to obtain the time length of the kth initial time slice based on the time length of the k-1th initial time slice, the standard deviation of the time difference corresponding to all seed websites in the k-1th initial time slice, the predicted traffic density corresponding to the kth initial time slice, the preset standard traffic density and the preset adjustment coefficient.
[0113] The second time length acquisition unit is used to update k=k+1, repeatedly execute the time difference acquisition unit until the sum of the time lengths of all initial time slices is greater than or equal to the preset time period, determine the first to the second-to-last initial time slices as the first to the second-to-last preset time slices, and use the end time point of the second-to-last initial time slice as the start time point of the last preset time slice, and use the end time point of the preset time period as the end time point of the last preset time slice to update and obtain the last preset time slice.
[0114] In a specific embodiment, the first clustering submodule further includes: The traffic feature acquisition unit is used to extract features of the traffic curve corresponding to each seed website and obtain the traffic features corresponding to the current seed website.
[0115] The address feature acquisition unit is used to extract features of the website address corresponding to the current seed website and acquire the address features corresponding to the current seed website.
[0116] The time difference feature acquisition unit is used to map the time difference corresponding to the current seed website into a preset numerical range to acquire the time difference feature corresponding to the current seed website.
[0117] The website feature acquisition unit is used to combine the traffic feature, address feature and time difference feature corresponding to the current seed website to obtain the website feature corresponding to the current seed website.
[0118] The first clustering unit is used to cluster all seed websites according to the website features corresponding to each seed website, obtain several seed website cluster sets, and determine the seed website corresponding to the cluster center of each seed website cluster set as the first cluster center website corresponding to each seed website cluster set.
[0119] In a specific embodiment, the first website clustering module 21 further includes: The second traffic curve acquisition submodule is used to acquire the traffic curve corresponding to each initial website according to the data packet flow corresponding to each initial website in each preset time slice.
[0120] The website feature acquisition submodule is used to acquire the website features corresponding to each initial website based on the traffic curve, website address and time difference corresponding to each initial website.
[0121] The first relative distance acquisition submodule is used to acquire the first relative distance between each initial website and each first cluster center website according to the website features corresponding to each initial website and the website features corresponding to each first cluster center website.
[0122] The second relative distance acquisition submodule is used to obtain, for any seed website cluster set, the second relative distance between each seed website in the current seed website cluster set and the first cluster center website based on the website characteristics of each seed website in the current seed website cluster set and the website characteristics of the first cluster center website.
[0123] The second distance threshold acquisition submodule is configured to determine the maximum value of the second relative distances between all seed websites in the current seed website cluster set and the first cluster center website as the second distance threshold corresponding to the current seed website cluster set.
[0124] The candidate cluster set acquisition submodule is used to, for each initial website, determine the current seed website cluster set as the candidate cluster set corresponding to the current initial website if the first relative distance between the current initial website and the first cluster center website in the current seed website cluster set is less than the second distance threshold corresponding to the current seed website cluster set.
[0125] The first result acquisition submodule is used to determine, if the number of candidate cluster sets corresponding to the current initial website is greater than 0, the candidate cluster set corresponding to the smallest first relative distance is determined as the seed website cluster set corresponding to the current initial website based on the first relative distance between the current initial website and the first cluster center website in each corresponding candidate cluster set, and determine the first cluster result corresponding to the current initial website as the corresponding assigned seed website cluster set.
[0126] The second result obtaining submodule is configured to determine that the first clustering result corresponding to the current initial website is not allocated to the seed website clustering set if the number of candidate clustering sets corresponding to the current initial website is 0.
[0127] In a specific embodiment, the second cluster center acquisition module 23 includes: The centroid vector acquisition submodule is used to calculate the centroid vector corresponding to the current seed website cluster set based on the website features of all initial websites and the website features of all seed websites corresponding to the current seed website cluster set.
[0128] The third relative distance acquisition submodule is used to obtain the third relative distance between the centroid vector corresponding to the current seed website cluster set and each corresponding initial website and each seed website based on the centroid vector corresponding to the current seed website cluster set and the website features of each initial website and each seed website corresponding to the current seed website cluster set.
[0129] The second cluster center acquisition submodule is used to determine the initial website or seed website corresponding to the smallest third relative distance as the second cluster center website corresponding to the current seed website cluster set.
[0130] In a specific embodiment, the target cluster set acquisition module 25 includes: The reference center feature acquisition submodule is used to determine the average feature of the website features corresponding to the first cluster center website and the website features corresponding to the second cluster center website corresponding to the current seed website cluster set as the reference center feature corresponding to the current seed website cluster set.
[0131] The fourth relative distance acquisition submodule is used to acquire the fourth relative distance between each seed website in the current seed website cluster set and the reference center feature according to the website features and reference center features corresponding to all seed websites corresponding to the current seed website cluster set.
[0132] The third cluster center acquisition submodule is configured to determine the seed website corresponding to the smallest fourth relative distance as the third cluster center website corresponding to the current seed website cluster set.
[0133] The target cluster set acquisition submodule is used to form a target cluster set with the third cluster center website as the cluster center based on all seed websites in the current seed website cluster set.
[0134] In a specific embodiment, the abnormal website identification system further includes: The normal website identification module is configured to determine, for any initial website, if the second clustering result corresponding to the current initial website is not assigned to the target cluster set, that the current initial website is a normal website.
[0135] It should be noted that the information interaction, execution process and other contents between the above modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0136] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any form. Although the present invention has been disclosed as above in terms of preferred embodiments, they are not intended to limit the present invention. Any technician familiar with this profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A method for identifying abnormal websites, characterized in that: The abnormal website identification method comprises the following steps: S1, determining a first clustering result corresponding to each initial website to be identified based on website information corresponding to each initial website to be identified within a preset time period and website information corresponding to a first cluster center website corresponding to each seed website cluster set, wherein the first clustering result includes a corresponding assigned seed website cluster set and a website cluster set not assigned to the seed website cluster set, and the seed website cluster set includes a plurality of seed websites for characterizing abnormal websites; S2, according to the first clustering result corresponding to each initial website, obtaining all initial websites corresponding to each seed website cluster set; S3, for any seed website cluster set, obtaining the second cluster center website corresponding to the current seed website cluster set based on all initial websites and all seed websites corresponding to the current seed website cluster set; S4, obtaining the distance to the center website corresponding to the current seed website cluster set based on the website information of the first cluster center website and the website information of the second cluster center website corresponding to the current seed website cluster set; S5, if the distance to the center website corresponding to the current seed website cluster set is greater than a preset first distance threshold, then obtaining the third cluster center website and the target cluster set corresponding to the current seed website cluster set based on the website information corresponding to all seed websites, the first cluster center website, and the second cluster center website corresponding to the current seed website cluster set; S6, determining a second clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified and the website information corresponding to each third cluster center website, wherein the second clustering result includes a corresponding assigned target cluster set and websites not assigned to the target cluster set; S7: For any initial website, if the second clustering result corresponding to the current initial website is the corresponding assigned target cluster set, the current initial website is determined to be an abnormal website.
2. The abnormal website identification method according to claim 1, characterized in that: The preset time period includes several preset time slices, and the website information includes data packet flow, website address, and time difference. S1 includes the following steps: S11, obtaining a traffic curve corresponding to each seed website based on the data packet flow corresponding to each preset time slice of each seed website, wherein the traffic curve uses the preset time slice as an independent variable and the data packet flow corresponding to each preset time slice as a dependent variable; S12, clustering all seed websites according to the traffic curve, website address and time difference corresponding to each seed website, obtaining several seed website cluster sets, and determining the seed website corresponding to the cluster center of each seed website cluster set as the first cluster center website corresponding to each seed website cluster set, wherein the time difference is the time difference between the time zone corresponding to the seed website and the target time zone.
3. The abnormal website identification method according to claim 2, characterized in that: Each seed website transmits a number of data packets to a target receiving end within the preset time period. S11 further includes the following steps: S111, based on the time point when each seed website transmits each corresponding data packet to the target receiving end in the k-1th initial time slice, obtain the time difference when each seed website transmits two corresponding adjacent data packets to the target receiving end in the k-1th initial time slice, where k is an integer greater than 1 and has an initial value of 2, and the time length corresponding to the first initial time slice is a preset length; S112, obtaining the standard deviation of the time differences corresponding to all seed websites in the k-1th initial time slice based on all the time differences corresponding to each seed website in the k-1th initial time slice; S113, based on the data packet flow corresponding to each seed website in the first to k-1th initial time slices and the time length corresponding to the first to k-1th initial time slices, obtain the flow density per unit time corresponding to each seed website in the first to k-1th initial time slices; S114, determining the average value of the traffic density per unit time of all seed websites in the first to k-1th initial time slices as the predicted traffic density of all seed websites in the kth initial time slice; S115, obtaining the length of the kth initial time slice based on the length of the k-1th initial time slice, the standard deviation of the time differences corresponding to all seed websites in the k-1th initial time slice, the predicted traffic density corresponding to the kth initial time slice, the preset standard traffic density, and the preset adjustment coefficient; S116, update k=k+1, repeat S111 until the sum of the time lengths of all initial time slices is greater than or equal to the preset time period, determine the first to the second-to-last initial time slices as the first to the second-to-last preset time slices, and use the end time point of the second-to-last initial time slice as the starting time point of the last preset time slice, and use the end time point of the preset time period as the end time point of the last preset time slice, and update to obtain the last preset time slice.
4. The abnormal website identification method according to claim 3, characterized in that: S12 further includes the following steps: S121, for each seed website, extracting features of the traffic curve corresponding to the current seed website to obtain traffic features corresponding to the current seed website; S122, extracting features of the website address corresponding to the current seed website to obtain address features corresponding to the current seed website; S123, mapping the time difference corresponding to the current seed website to a preset numerical range to obtain the time difference feature corresponding to the current seed website; S124, combining the traffic characteristics, address characteristics, and time difference characteristics corresponding to the current seed website to obtain the website characteristics corresponding to the current seed website; S125 , clustering all seed websites according to website features corresponding to each seed website to obtain several seed website cluster sets, and determining the seed website corresponding to the cluster center of each seed website cluster set as the first cluster center website corresponding to each seed website cluster set.
5. The abnormal website identification method according to claim 4, characterized in that: S1 also includes the following steps: S101, obtaining a traffic curve corresponding to each initial website based on the data packet flow corresponding to each initial website in each preset time slice; S102, obtaining website features corresponding to each initial website based on the traffic curve, website address, and time difference corresponding to each initial website; S103, obtaining a first relative distance between each initial website and each first cluster center website based on the website features corresponding to each initial website and the website features corresponding to each first cluster center website; S104: for any seed website cluster set, obtaining a second relative distance between each seed website in the current seed website cluster set and the first cluster center website based on the website characteristics of each seed website in the current seed website cluster set and the website characteristics of the first cluster center website; S105, determining the maximum value of the second relative distances between all seed websites in the current seed website cluster set and the first cluster center website as the second distance threshold corresponding to the current seed website cluster set; S106, for each initial website, if a first relative distance between the current initial website and a first cluster center website in the current seed website cluster set is less than a second distance threshold corresponding to the current seed website cluster set, then determining the current seed website cluster set as a candidate cluster set corresponding to the current initial website; S107: If the number of candidate cluster sets corresponding to the current initial website is greater than 0, then, based on the first relative distances between the current initial website and the first cluster center website in each corresponding candidate cluster set, the candidate cluster set corresponding to the smallest first relative distance is determined as the seed website cluster set corresponding to the current initial website, and the first clustering result corresponding to the current initial website is determined as the corresponding seed website cluster set; S108: If the number of candidate cluster sets corresponding to the current initial website is 0, it is determined that the first clustering result corresponding to the current initial website is not allocated to the seed website cluster set.
6. The abnormal website identification method according to claim 5, characterized in that: S3 includes the following steps: S31, calculating the centroid vector corresponding to the current seed website cluster set based on the website features of all initial websites and the website features of all seed websites corresponding to the current seed website cluster set; S32, obtaining a third relative distance between the centroid vector corresponding to the current seed website cluster set and each corresponding initial website and each seed website based on the centroid vector corresponding to the current seed website cluster set and the website features of each initial website and each seed website corresponding to the current seed website cluster set; S33: Determine the initial website or seed website corresponding to the smallest third relative distance as the second cluster center website corresponding to the current seed website cluster set.
7. The abnormal website identification method according to claim 5, characterized in that: S5 includes the following steps: S51, determining the average feature of the website feature corresponding to the first cluster center website and the website feature corresponding to the second cluster center website corresponding to the current seed website cluster set as the reference center feature corresponding to the current seed website cluster set; S52, obtaining a fourth relative distance between each seed website in the current seed website cluster set and the reference center feature based on the website features and reference center features corresponding to all seed websites corresponding to the current seed website cluster set; S53, determining the seed website corresponding to the smallest fourth relative distance as the third cluster center website corresponding to the current seed website cluster set; S54: forming a target cluster set with the third cluster center website as the cluster center based on all seed websites in the current seed website cluster set.
8. The abnormal website identification method according to claim 1, characterized in that: The abnormal website identification method further comprises the following steps: For any initial website, if the second clustering result corresponding to the current initial website is not assigned to the target cluster set, the current initial website is determined to be a normal website.
9. An abnormal website identification system, characterized in that: The abnormal website identification system includes: a first website clustering module, configured to determine a first clustering result corresponding to each initial website to be identified within a preset time period based on website information corresponding to each initial website to be identified and website information corresponding to a first cluster center website corresponding to each seed website cluster set, wherein the first clustering result includes a corresponding assigned seed website cluster set and websites not assigned to a seed website cluster set, and the seed website cluster set includes a plurality of seed websites for characterizing abnormal websites; A clustering set analysis module is used to obtain all initial websites corresponding to each seed website cluster set based on the first clustering result corresponding to each initial website; The second cluster center acquisition module is used to acquire the second cluster center website corresponding to any seed website cluster set based on all initial websites and all seed websites corresponding to the current seed website cluster set; A center website distance acquisition module is used to acquire the center website distance corresponding to the current seed website cluster set based on the website information of the first cluster center website and the website information of the second cluster center website corresponding to the current seed website cluster set; A target cluster set acquisition module is configured to acquire the third cluster center website and the target cluster set corresponding to the current seed website cluster set based on the website information corresponding to all seed websites, the first cluster center website, and the second cluster center website corresponding to the current seed website cluster set, if the distance to the center website corresponding to the current seed website cluster set is greater than a preset first distance threshold; a second website clustering module, configured to determine a second clustering result corresponding to each initial website based on the website information corresponding to each initial website to be identified and the website information corresponding to each third cluster center website, wherein the second clustering result includes a corresponding assigned target cluster set and websites not assigned to the target cluster set; The abnormal website identification module is used to determine, for any initial website, if the second clustering result corresponding to the current initial website is the corresponding assigned target cluster set, that the current initial website is an abnormal website.
10. The abnormal website identification system according to claim 9, characterized in that: The preset time period includes a plurality of preset time slices, the website information includes data packet flow, website address, and time difference, and the first website clustering module includes: The first traffic curve acquisition submodule is used to acquire a traffic curve corresponding to each seed website based on the data packet flow corresponding to each seed website in each preset time slice, wherein the traffic curve uses the preset time slice as an independent variable and the data packet flow corresponding to each preset time slice as a dependent variable; The first clustering submodule is used to cluster all seed websites according to the traffic curve, website address and time difference corresponding to each seed website, obtain several seed website cluster sets, and determine the seed website corresponding to the cluster center of each seed website cluster set as the first cluster center website corresponding to each seed website cluster set, wherein the time difference is the time difference between the time zone corresponding to the seed website and the target time zone.
Citation Information
Patent Citations
Method and device for analysis based on website clustering result
CN106909569A
Phishing website detection method and system in Android environment
CN107438083A
Malicious website detection method and device, electronic equipment and computer storage medium
CN113901376A
Two-dimensional code malicious website detection method and device based on clustering algorithm
CN115967562A
Domain name generation model training method, phishing website discovery method and related device
CN117892801A