Fraud user identification method and device based on traffic data

By analyzing website traffic data, building an asset tree, calculating path scores and hotspot path deviations, and identifying fraudulent users, the problems of small coverage and low accuracy in existing technologies are solved, and efficient fraudulent user identification is achieved.

CN116055119BActive Publication Date: 2025-09-05CCS TRANSFAR TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211635597.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-19
Publication Date
2025-09-05
Estimated Expiration
2042-12-19

AI Technical Summary

Technical Problem

When identifying fraudulent users, existing technologies have a small coverage and low accuracy, making it difficult to effectively identify fraudulent users mixed in with normal users.

Method used

By analyzing website traffic data, building an asset tree, calculating path scores, screening important paths, performing clustering and hotspot path analysis, calculating hotspot path deviation, and identifying fraudulent users.

Benefits of technology

It improves the accuracy and efficiency of fraudulent user identification, avoids indiscriminate judgment of group users, and greatly improves the fraudulent user identification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116055119B_ABST
    Figure CN116055119B_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure provide a method and apparatus for identifying fraudulent users based on traffic data. The method comprises: determining the paths accessed from a website and their path scores based on the website's traffic data; screening the traffic data from the website's traffic data, including traffic data for paths whose access paths have a path score greater than or equal to a preset threshold; clustering the filtered traffic data based on the access paths; determining the hotspot paths for each class based on the access paths of each traffic data in each class; calculating the hotspot path deviation of each traffic data in each class based on the hotspot paths of each class; and identifying fraudulent users based on the hotspot path deviation of each traffic data in each class. In this way, the effectiveness of identifying fraudulent users can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of network security technology, and in particular to a method and device for identifying fraudulent users based on traffic data. Background Art

[0002] With the continuous development of the Internet, in order to enable normal users to better enjoy services, it is necessary to identify users and determine fraudulent users mixed in with normal users.

[0003] Currently, most internet companies rely on rule-based engines and offline investigations to identify fraudulent users. However, this approach suffers from limited coverage and low accuracy. Therefore, improving the effectiveness of fraudulent user identification has become a pressing technical challenge. Summary of the Invention

[0004] The present disclosure provides a method and device for identifying fraudulent users based on traffic data.

[0005] In a first aspect, an embodiment of the present disclosure provides a method for identifying fraudulent users based on traffic data, the method comprising:

[0006] Determine the path the website is accessed and its path score based on the website's traffic data;

[0007] Filtering, from the traffic data of the website, access paths including traffic data of paths having a path score higher than or equal to a preset threshold;

[0008] Clustering the filtered traffic data according to the access paths of the filtered traffic data;

[0009] Determine the hotspot path for each class based on the access path of each traffic data in each class;

[0010] According to the hotspot path of each class, the hotspot path deviation of each traffic data in each class is calculated;

[0011] Identify fraudulent users based on the hotspot path deviation of each traffic data in each class.

[0012] In some implementations of the first aspect, determining, based on website traffic data, a path through which a website is accessed and its path score includes:

[0013] Perform asset tree sorting on the website's traffic data to obtain the website's asset tree, where the asset tree represents the path of website access in a tree structure;

[0014] Calculate the basic score, algorithm score, adjustment score, and supplementary information score for each path in the asset tree based on the traffic data corresponding to the path in the asset tree;

[0015] The path score of each path in the asset tree is calculated based on the basic score, algorithm score, adjustment score, and supplementary information score of each path in the asset tree.

[0016] In some implementations of the first aspect, clustering the filtered traffic data according to access paths of the filtered traffic data includes:

[0017] According to the session ID of the filtered traffic data, the traffic data of the same session ID is aggregated into the overall traffic data;

[0018] Normalizing and vectorizing the path jumps of each integrated flow data to obtain the path access features of each integrated flow data;

[0019] Clustering is performed on each integrated traffic data according to the path access characteristics of each integrated traffic data.

[0020] In some implementations of the first aspect, determining a hotspot path for each class based on access paths for each traffic data in each class includes:

[0021] Count the number of traffic data involved in each access path in its corresponding class;

[0022] Divide the number of traffic data involved in each access path in its corresponding class by the total number of traffic data in the corresponding class to obtain the hotspot coefficient of each access path in its corresponding class;

[0023] An access path corresponding to each class whose hotspot coefficient is greater than or equal to a preset threshold is determined as a hotspot path.

[0024] In some implementations of the first aspect, calculating the hotspot path deviation of each flow data in each class according to the hotspot path of each class includes:

[0025] Statistics are collected for each traffic data in each class to obtain statistical data for each traffic data in each class. The statistical data include: hotspot path time consumption, hotspot path access ratio, the ratio of independent hotspot paths to the total number of hotspot paths, and the total number of hotspot path visits.

[0026] Based on the statistical data of each traffic data in each class, calculate the deviation of the hotspot path time consumption, the deviation of the hotspot path access ratio, the deviation of the ratio of the number of independent hotspot paths to the total number of hotspot paths, and the deviation of the total number of hotspot path visits for each traffic data in each class;

[0027] According to the deviation of each flow data in each class, the hotspot path deviation of each flow data in each class is calculated.

[0028] In some implementations of the first aspect, calculating the hotspot path deviation of each flow data in each class according to each deviation of each flow data in each class includes:

[0029] For any flow data in each class, the deviations are weighted and summed to obtain the hotspot path deviation of the flow data.

[0030] In some implementations of the first aspect, identifying fraudulent users based on hotspot path deviations of traffic data in each class includes:

[0031] The user corresponding to the traffic data whose hotspot path deviation is greater than or equal to the preset threshold is determined as a fraudulent user.

[0032] In a second aspect, an embodiment of the present disclosure provides a fraudulent user identification device based on traffic data, the device comprising:

[0033] A determination module, used to determine the path of website access and its path score based on the website's traffic data;

[0034] A screening module, configured to screen, from the traffic data of the website, access paths including traffic data of paths having a path score higher than or equal to a preset threshold;

[0035] A clustering module, for clustering the filtered traffic data according to access paths of the filtered traffic data;

[0036] The determination module is further used to determine the hotspot path of each class based on the access path of each flow data in each class;

[0037] The calculation module is used to calculate the hotspot path deviation of each flow data in each class based on the hotspot path of each class;

[0038] The identification module is used to identify fraudulent users based on the hotspot path deviation of each traffic data in each class.

[0039] In a third aspect, an embodiment of the present disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0040] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to execute the method described above.

[0041] In the embodiments of the present disclosure, the importance of different paths for accessing the website can be determined based on the website's traffic data, and then the traffic data for irrelevant paths can be filtered out. On this basis, the filtered traffic data can be clustered, so as to more specifically determine the group users, and based on the hot paths of the class, the fraudulent users in the group users can be efficiently identified, avoiding indiscriminate determination of the group users, thereby greatly improving the identification effect of fraudulent users.

[0042] It should be understood that the contents described in the Summary of the Invention section are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. The accompanying drawings are provided for a better understanding of the present disclosure and do not constitute a limitation of the present disclosure. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, among which:

[0044] Figure 1 A flowchart of a method for identifying fraudulent users based on traffic data provided by an embodiment of the present disclosure is shown;

[0045] Figure 2 A schematic diagram of a pseudo-static resource provided by an embodiment of the present disclosure is shown;

[0046] Figure 3 A structural diagram of a fraudulent user identification device based on traffic data provided by an embodiment of the present disclosure is shown;

[0047] Figure 4 A structural diagram of an exemplary electronic device capable of implementing an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0048] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present disclosure.

[0049] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.

[0050] In response to the problems encountered in the background technology, the embodiments of the present disclosure provide a method and apparatus for identifying fraudulent users based on traffic data. Specifically, based on the website's traffic data, the paths accessed by the website and their path scores are determined; from the website's traffic data, traffic data for access paths including paths with path scores greater than or equal to a preset threshold are screened; based on the access paths of the screened traffic data, the screened traffic data are clustered; based on the access paths of each traffic data in each class, the hotspot paths of each class are determined; based on the hotspot paths of each class, the hotspot path deviation of each traffic data in each class is calculated; and based on the hotspot path deviation of each traffic data in each class, fraudulent users are identified.

[0051] In this way, the importance of different paths for accessing the website can be determined based on the website's traffic data, and then the traffic data for irrelevant paths can be filtered out. On this basis, the filtered traffic data can be clustered, so that group users can be determined more specifically. According to the hot paths of the class, fraudulent users in the group users can be efficiently identified, avoiding indiscriminate judgment of group users, which greatly improves the identification effect of fraudulent users.

[0052] The following describes in detail the method and apparatus for identifying fraudulent users based on traffic data provided by the embodiments of the present disclosure through specific embodiments in conjunction with the accompanying drawings.

[0053] Figure 1 FIG. 1 shows a flow chart of a fraudulent user identification method based on traffic data provided by an embodiment of the present disclosure, such as Figure 1 As shown, the fraud user identification method 100 may include the following steps:

[0054] S110: Determine the path along which the website is accessed and its path score based on the website traffic data.

[0055] In some embodiments, the full traffic data of the website may be obtained, and then the traffic data of the website may be sorted into an asset tree to obtain the asset tree of the website, wherein the asset tree represents the path of the website being accessed in a tree structure.

[0056] Based on the traffic data corresponding to the paths in the asset tree, the basic score, algorithm score, adjustment score, and supplementary information score of each path in the asset tree are calculated.

[0057] The path score of each path in the asset tree is calculated based on the basic score, algorithm score, adjustment score, and supplementary information score of each path in the asset tree. For example, for any path in the asset tree, the path score is calculated by taking the weighted sum of its basic score, algorithm score, adjustment score, and supplementary information score.

[0058] In this way, the path score can be calculated from multiple dimensions, which improves the credibility of the path score.

[0059] S120 , filtering access paths from the traffic data of the website to include traffic data of paths whose path scores are higher than or equal to a preset threshold.

[0060] Simply put, it is to filter out traffic data for important paths (i.e. paths with higher importance) based on the scores (i.e. importance) of different paths visited by the website, and at the same time filter out traffic data for irrelevant paths (i.e. paths with lower importance), thereby effectively improving the accuracy of subsequent fraudulent user identification.

[0061] S130: Clustering the filtered traffic data according to the access paths of the filtered traffic data.

[0062] In some embodiments, the traffic data with the same session ID may be aggregated into integrated traffic data according to the session ID of the filtered traffic data. That is, the traffic data with the same session ID may be aggregated into an integrated data corresponding to the session ID.

[0063] Normalize and vectorize the path jumps of each integrated traffic data set to obtain the path access features of each integrated traffic data set. For example, use the multi-hot algorithm to normalize and vectorize the path jumps (including URLs and referers). That is, the path jumps within a request are used as a dimension, and the path jumps of all requests in the traffic data are calculated as the corresponding dimension features, namely the path access features.

[0064] Clustering is performed on each integrated flow data according to the path access features of each integrated flow data. For example, the DBSCAN algorithm is used to cluster the path access features of each integrated flow data, thereby achieving clustering of each integrated flow data.

[0065] S140 , determining a hotspot path for each class according to the access path of each flow data in each class.

[0066] In some embodiments, the number of traffic data involved in each access path in its corresponding class can be counted, and the number of traffic data involved in each access path in its corresponding class can be divided by the total number of traffic data in the corresponding class to obtain the hotspot coefficient of each access path in its corresponding class, and the access path whose hotspot coefficient in the access path corresponding to each class is greater than or equal to a preset threshold is determined as a hotspot path.

[0067] That is to say, if a path appears in most of the traffic data within a class, it can be considered as a hotspot path of the class.

[0068] S150 , calculating the hotspot path deviation of each flow data in each class according to the hotspot path of each class.

[0069] In some embodiments, statistics can be performed on the traffic data in each class to obtain statistical data of the traffic data in each class, wherein the statistical data may include: hotspot path time consumption, hotspot path access ratio, the ratio of the number of independent hotspot paths to the total number of hotspot paths, and the total number of hotspot path visits.

[0070] Based on the statistical data of each traffic data in each class, the deviation of the hotspot path time consumption, the deviation of the hotspot path access ratio, the deviation of the ratio of the number of independent hotspot paths to the total number of hotspot paths, and the deviation of the total number of hotspot path visits of each traffic data in each class are calculated.

[0071] Based on the deviations of each flow data in each class, the hotspot path deviation of each flow data in each class is calculated. For example, for any flow data in each class, the deviations are weighted and summed to obtain the hotspot path deviation of the flow data.

[0072] In this way, the hotspot path deviation of traffic data can be calculated from multiple dimensions, thereby improving the credibility of the hotspot path deviation.

[0073] S160: Identify fraudulent users based on the hotspot path deviation of each traffic data in each class.

[0074] Specifically, the user corresponding to the traffic data whose hotspot path deviation is greater than or equal to a preset threshold can be determined as a fraudulent user.

[0075] In the embodiments of the present disclosure, the importance of different paths for accessing the website can be determined based on the website's traffic data, and then the traffic data for irrelevant paths can be filtered out. On this basis, the filtered traffic data can be clustered, so as to more specifically determine the group users, and based on the hot paths of the class, the fraudulent users in the group users can be efficiently identified, avoiding indiscriminate determination of the group users, thereby greatly improving the identification effect of fraudulent users.

[0076] The following is a detailed introduction to the fraudulent user identification method provided by the embodiment of the present disclosure in conjunction with a specific embodiment, as follows:

[0077] The asset tree combing for traffic data mainly involves the processing of static resources, pseudo-static resources, dynamic resources and ordinary resources, including the merging of static resources, standardization of pseudo-static resource parameters and the setting of dynamic resources.

[0078] Pseudo-static resources such as Figure 2 As shown, pseudo-static resources refer to paths that contain parameters, such as Figure 2 The path shown in the Baidu search for keyword 1 is a parameter. Different search contents only have different parameters after the keyword. Therefore, when processing traffic data, it is necessary to standardize the parameters in the path to make the analysis of user behavior through the path more accurate.

[0079] The process of standardizing pseudo-static resources is as follows: 1. Change the path to the form shown in the figure, and express the content of each part in the form of key and value. For example, https is expressed as scheme:https. Because a certain part may have multiple combinations, according to the different combinations in the data, if the key has the same name, it is arranged in order by the size of the number. Otherwise, it is constructed according to the actual parameter name; 2. After constructing the key and value combination, merge them according to the same key. If the number of different value values ​​corresponding to the same key in the entire data set exceeds a certain number, the corresponding value will be judged as a parameter value, and the relevant content needs to be standardized; 3. Generate corresponding standardization rules for the path containing the parameter, and standardize other subsequent paths.

[0080] Static resources mainly refer to the paths of specific images and file types. These paths are finally named with specific file suffixes. When processing static resources, the relevant paths are standardized directly based on whether the path suffix contains ico, gif, bmp, jpg, css, etc., and static resources of different categories are formally classified as one type of asset.

[0081] Dynamic resources mainly refer to paths with specific functions, including paths containing sensitive information such as login, logout, and payment. These paths can reflect more specific behaviors from the path content. These paths are treated as dynamic resources during the processing process and are assets with relatively high weights.

[0082] Aside from the three paths mentioned above, all other paths are considered common resources. Once the path resources are organized, a tree structure of all paths is formed, revealing all the behavioral patterns of the entire traffic data. This provides clearer data for subsequent user behavior analysis and fraud ring account detection algorithms.

[0083] In the asset tree, each path from the root node to the leaf node needs to be scored, that is, the sorted paths are weighted and scored. The scoring is mainly done by setting corresponding scoring rules in the relevant field information in the traffic data that can reflect whether the request is normal or not, so as to obtain the weight of the path in the entire data. During the scoring process, the basic score and the algorithm score will be calculated first and the total score will be added up, and then the weight adjustment stage will be carried out (the weight adjustment will only involve the basic score, the algorithm score and the total score). First, the weight will be adjusted according to the asset classification (for example, static resources will be significantly downgraded), and then the weight will be downgraded according to the proportion of ajax in the path (non-ajax will not be downgraded). Then, for two situations with high visits and few response types, and high visits and low importance of jump logic, the weight will be downgraded according to the corresponding proportions, and finally the total score will be obtained.

[0084] Basic Score: Reflects the impact of basic status on path assets in a single request. Basic scoring is based on field values ​​in the data, and can also be configured to perform secondary regularization extraction on the fields. After scoring different values ​​according to the configuration, the scores are then weighted again using the TF-IDF algorithm, with rarer values ​​receiving higher weights. Basic scoring includes a behavior score, which determines the method used in traffic data requests, including GET and POST. Different methods have different preset scores, such as 10 for GET and 15 for POST. A status score, which determines the score of different status codes, is calculated based on the number the status code begins with. Status codes beginning with 2 indicate a normal connection and receive a lower score, while status codes beginning with other numbers indicate a connection problem and receive a higher score. A cookie score, which determines whether the cookie field in the data contains content. A session score, which uses regularization to detect relevant session information. An error score, which determines whether specific status codes, such as 404, are present.

[0085] Because each score is static and subject to certain deviations in real-world data, dynamic adjustments are necessary based on specific data. This adjustment utilizes the TF-IDF algorithm, which weights the static scores of individual rule scores within each scoring item based on the number of occurrences of content within each scoring item and whether the content is present in all paths (i.e., its rarity). If a scoring item appears in the majority of requests, its weight is lowered, minimizing its impact on the final score.

[0086] Each request in the traffic data is assigned one of the five weighted scores described above. However, specific assets may appear multiple times in the traffic data, so the scores for each asset must be aggregated accordingly. For the behavior score, the score for a specific path asset is the maximum score across all corresponding requests. For the status score, the average score across all corresponding requests is used. For the cookie score, the maximum score is used for aggregation. For both the session score and the error status score, the average score is used.

[0087] Algorithm Score: Reflects the impact of path pointing in traffic data on path assets. Source Score and Visit Score: Log the number of sources / visits and multiply by a fixed coefficient to ensure that the data volume itself does not significantly affect the score. Jump Logic Score: Calculated using the Page Rank algorithm, this value primarily calculates the interconnectedness of paths and reflects the importance of the path in the data. If many paths point to a single path, then this path is likely more important than other paths.

[0088] Adjustment Score: Reflects the impact of path responses on path assets. Response Score: Response content is hashed using MD5, the number of unique hash values ​​is calculated, and then the log is processed and multiplied by a fixed coefficient. Ajax Score: Each log entry is determined to be an Ajax call (specifically, using X-Requested-With == XMLRequest). Finally, the percentage of paths that are Ajax calls is calculated.

[0089] Supplementary information score: reflects the impact of other relevant information on traffic data. API identification score: identifies whether the path may be an API interface based on whether the content type field is in the hmtl / xml format, and finally assigns a corresponding score.

[0090] First, accesses with asset scores below a certain threshold, i.e., traffic data, are filtered to exclude accesses to irrelevant paths, thus reducing the impact of many noise paths on subsequent algorithms.

[0091] Since data generally does not contain user IDs, we extract the session ID from the cookie and aggregate multiple accesses from the same session into a single session to represent the access behavior of the session. To help focus on valid sessions during aggregation, the following settings are used:

[0092] Because some sessions are much longer than others, you can split the long session into multiple sessions. If the time between two consecutive operations is more than 30 minutes, the session will be cut off from this point in time.

[0093] The session can then be further filtered, with the following filter options:

[0094] (1) How many requests for key paths does the session need to cover?

[0095] (2) The key paths covered by the session must include specific keywords, that is, some sensitive paths.

[0096] After performing the above operations, several sessions are generated. Each session ID is the concatenation of the extracted session ID and the truncation time. The session contains all relevant key asset accesses and has a certain length. Subsequent clustering and other operations will be performed on a session basis. The multi-hot algorithm is used to normalize and vectorize path jumps (including URLs and referers). Specifically, the path jumps within a request are used as a dimension, and the path jumps of all requests within the session are calculated as the corresponding dimension features.

[0097] After preprocessing, the algorithm clusters the vectorized sessions using the DBSCAN algorithm, using cosine distance to determine the degree of deviation between sessions. The algorithm connects sessions whose distance is less than a threshold. If a session is connected to multiple sessions simultaneously, a group, or cluster, is formed. It should be noted that, unlike most clustering algorithms, this algorithm also generates sessions that do not form a group and are counted as group -1. The clustering algorithm requires the following information:

[0098] (1) The distance threshold between sessions;

[0099] (2) The minimum number of connections between sessions required for clustering.

[0100] After generating a session group, if a path appears in more than a certain threshold percentage of sessions within the group, it is considered a hot path for that session. In the algorithm, it is entirely possible that the same path appears in the hot path list of multiple groups.

[0101] After clustering, the generated hotspot paths and related statistical data will be used. The statistical data include the hotspot path time of each session, the hotspot path access ratio of each session, the ratio of the number of independent hotspot paths of each session to the total number of hotspot paths, and the total number of hotspot path visits.

[0102] During the actual scoring, each session will be scored for the group's statistical data.

[0103] Items include:

[0104] 0(1) Deviation of hotspot path time;

[0105] (2) Deviation of the hotspot path access ratio;

[0106] (3) Deviation of the ratio of the number of independent hotspot paths to the total number of hotspot paths;

[0107] (4) Deviation of the total number of visits to the hotspot path.

[0108] Deviation is defined as (session statistic value - group statistic average) / 5 group statistic standard deviations. Time deviation is multiplied by -1 to indicate that shorter time periods are more abnormal. The final total score is the weighted sum of each deviation, plus the following two additional scores:

[0109] (1) In the deviation of the proportion of the number of independent hotspot paths accessed by the session to the total number of hotspot paths and

[0110] When the "time deviation on the hotspot path" is greater than 0, additional points will be added according to the deviation of the two; 0(2) When the total number of visits on the hotspot path exceeds a certain higher threshold, additional weight will be reduced.

[0111] This is common when users repeatedly refresh the same business or engage in other unskilled behaviors.

[0112] The hotspot path deviation degree reflecting user behavior is scored, and a corresponding threshold is set. The user corresponding to the session with a hotspot path deviation degree greater than or equal to the threshold is identified as a fraudulent user.

[0113] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0114] The above is an introduction to the method embodiment. The following is a further explanation of the solution disclosed in the present disclosure through an apparatus embodiment.

[0115] Figure 3 FIG. 1 shows a structural diagram of a fraudulent user identification device based on traffic data provided by an embodiment of the present disclosure, such as Figure 3 As shown, the fraud user identification device 300 may include:

[0116] The determination module 310 is used to determine the path of website access and its path score based on the website traffic data.

[0117] The screening module 320 is configured to screen, from the traffic data of the website, access paths including traffic data of paths having a path score higher than or equal to a preset threshold.

[0118] The clustering module 330 is configured to cluster the filtered traffic data according to the access paths of the filtered traffic data.

[0119] The determination module 310 is further configured to determine a hotspot path for each class based on the access path of each flow data in each class.

[0120] The calculation module 340 is used to calculate the hotspot path deviation of each flow data in each class based on the hotspot path of each class.

[0121] The identification module 350 is used to identify fraudulent users based on the hotspot path deviation of each traffic data in each class.

[0122] In some embodiments, the determination module 310 is specifically configured to:

[0123] Perform asset tree sorting on the website's traffic data to obtain the website's asset tree, where the asset tree represents the path of website access in a tree structure;

[0124] Calculate the basic score, algorithm score, adjustment score, and supplementary information score for each path in the asset tree based on the traffic data corresponding to the path in the asset tree;

[0125] The path score of each path in the asset tree is calculated based on the basic score, algorithm score, adjustment score, and supplementary information score of each path in the asset tree.

[0126] In some embodiments, the clustering module 330 is specifically configured to:

[0127] According to the session ID of the filtered traffic data, the traffic data of the same session ID is aggregated into the overall traffic data;

[0128] Normalizing and vectorizing the path jumps of each integrated flow data to obtain the path access features of each integrated flow data;

[0129] Clustering is performed on each integrated traffic data according to the path access characteristics of each integrated traffic data.

[0130] In some embodiments, the determination module 310 is specifically configured to:

[0131] Count the number of traffic data involved in each access path in its corresponding class;

[0132] Divide the number of traffic data involved in each access path in its corresponding class by the total number of traffic data in the corresponding class to obtain the hotspot coefficient of each access path in its corresponding class;

[0133] An access path corresponding to each class whose hotspot coefficient is greater than or equal to a preset threshold is determined as a hotspot path.

[0134] In some embodiments, the calculation module 340 is specifically configured to:

[0135] Statistics are collected for each traffic data in each class to obtain statistical data for each traffic data in each class. The statistical data include: hotspot path time consumption, hotspot path access ratio, the ratio of independent hotspot paths to the total number of hotspot paths, and the total number of hotspot path visits.

[0136] Based on the statistical data of each traffic data in each class, calculate the deviation of the hotspot path time consumption, the deviation of the hotspot path access ratio, the deviation of the ratio of the number of independent hotspot paths to the total number of hotspot paths, and the deviation of the total number of hotspot path visits for each traffic data in each class;

[0137] According to the deviation of each flow data in each class, the hotspot path deviation of each flow data in each class is calculated.

[0138] In some embodiments, the calculation module 340 is specifically configured to:

[0139] For any flow data in each class, the deviations are weighted and summed to obtain the hotspot path deviation of the flow data.

[0140] In some embodiments, the identification module 350 is specifically configured to:

[0141] The user corresponding to the traffic data whose hotspot path deviation is greater than or equal to the preset threshold is determined as a fraudulent user.

[0142] It is understandable that Figure 3 Each module / unit in the fraud user identification device 300 has the function of realizing Figure 1 The functions of the various steps in the fraudulent user identification method 100 shown and their ability to achieve corresponding technical effects are not described here for the sake of brevity.

[0143] Figure 4A block diagram of an exemplary electronic device capable of implementing embodiments of the present disclosure is shown. Electronic device 400 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 400 may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0144] like Figure 4 As shown, the electronic device 400 may include a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the electronic device 400 may also be stored in the RAM 403. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0145] Multiple components in the electronic device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0146] The computing unit 401 may be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as method 100. For example, in some embodiments, the method 100 may be implemented as a computer program product, including a computer program tangibly embodied in a computer-readable medium, such as a storage unit 408. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the method 100 described above may be performed. Alternatively, in other embodiments, the computing unit 401 may be configured to perform the method 100 in any other appropriate manner (e.g., by means of firmware).

[0147] The various embodiments described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0148] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0149] In the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a computer-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0150] It should be noted that the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute method 100 and achieve the corresponding technical effects achieved by executing the method in the embodiment of the present disclosure. For the sake of brevity, they will not be repeated here.

[0151] In addition, the present disclosure also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the method 100 is implemented.

[0152] To provide interaction with a user, the embodiments described above may be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0153] The embodiments described above can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0154] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0155] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0156] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A fraudulent user identification method based on traffic data, characterized in that: The method comprises: Determine the path the website is accessed and its path score based on the website's traffic data; Filtering, from the traffic data of the website, access paths including traffic data of paths having a path score higher than or equal to a preset threshold; Clustering the filtered traffic data according to the access paths of the filtered traffic data; Determine the hotspot path for each class based on the access path of each traffic data in each class; According to the hotspot path of each class, the hotspot path deviation of each traffic data in each class is calculated; Identify fraudulent users based on the hotspot path deviation of each traffic data in each class; Clustering the filtered traffic data according to the access path of the filtered traffic data includes: According to the session ID of the filtered traffic data, the traffic data of the same session ID is aggregated into the overall traffic data; Normalizing and vectorizing the path jumps of each integrated flow data to obtain the path access features of each integrated flow data; Clustering each integrated flow data according to its path access characteristics; The hotspot path deviation of each flow data in each class is calculated based on the hotspot path of each class, including: Statistics are collected for each traffic data in each class to obtain statistical data for each traffic data in each class, wherein the statistical data include: hotspot path time consumption, hotspot path access ratio, ratio of the number of independent hotspot paths to the total number of hotspot paths, and the total number of hotspot path visits; Based on the statistical data of each traffic data in each class, calculate the deviation of the hotspot path time consumption, the deviation of the hotspot path access ratio, the deviation of the ratio of the number of independent hotspot paths to the total number of hotspot paths, and the deviation of the total number of hotspot path visits for each traffic data in each class; According to the deviation of each flow data in each class, the hotspot path deviation of each flow data in each class is calculated.

2. The method according to claim 1, characterized in that Determining the access path of the website and its path score based on the website traffic data includes: Performing asset tree sorting on the website's traffic data to obtain an asset tree of the website, wherein the asset tree represents a path through which the website is accessed in a tree structure; Calculate the basic score, algorithm score, adjustment score, and supplementary information score of each path in the asset tree based on the traffic data corresponding to the path in the asset tree; The path score of each path in the asset tree is calculated according to the basic score, algorithm score, adjustment score, and supplementary information score of each path in the asset tree.

3. The method according to claim 1, characterized in that Determining the hotspot path of each class based on the access path of each flow data in each class includes: Count the number of traffic data involved in each access path in its corresponding class; Divide the number of traffic data involved in each access path in its corresponding class by the total number of traffic data in the corresponding class to obtain the hotspot coefficient of each access path in its corresponding class; An access path corresponding to each class whose hotspot coefficient is greater than or equal to a preset threshold is determined as a hotspot path.

4. The method according to claim 1, wherein The calculation of the hotspot path deviation of each flow data in each class according to each deviation of each flow data in each class includes: For any flow data in each class, the deviations are weighted and summed to obtain the hotspot path deviation of the flow data.

5. The method according to any one of claims 1 to 4, characterized in that The method of identifying fraudulent users based on the hotspot path deviation of each traffic data in each class includes: The user corresponding to the traffic data whose hotspot path deviation is greater than or equal to the preset threshold is determined as a fraudulent user.

6. A fraudulent user identification device based on traffic data, characterized in that: The device comprises: A determination module, used to determine the path of website access and its path score based on the website's traffic data; A screening module, configured to screen, from the traffic data of the website, access paths including traffic data of paths having a path score higher than or equal to a preset threshold; A clustering module, for clustering the filtered traffic data according to access paths of the filtered traffic data; The determination module is further configured to determine a hotspot path for each class based on the access path of each flow data in each class; The calculation module is used to calculate the hotspot path deviation of each flow data in each class based on the hotspot path of each class; The identification module is used to identify fraudulent users based on the hotspot path deviation of each traffic data in each class; The clustering module is specifically used for: According to the session ID of the filtered traffic data, the traffic data of the same session ID is aggregated into the overall traffic data; Normalizing and vectorizing the path jumps of each integrated flow data to obtain the path access features of each integrated flow data; Clustering each integrated flow data according to its path access characteristics; The calculation module is specifically used for: Statistics are collected for each traffic data in each class to obtain statistical data for each traffic data in each class, wherein the statistical data include: hotspot path time consumption, hotspot path access ratio, ratio of the number of independent hotspot paths to the total number of hotspot paths, and the total number of hotspot path visits; Based on the statistical data of each traffic data in each class, calculate the deviation of the hotspot path time consumption, the deviation of the hotspot path access ratio, the deviation of the ratio of the number of independent hotspot paths to the total number of hotspot paths, and the deviation of the total number of hotspot path visits for each traffic data in each class; According to the deviation of each flow data in each class, the hotspot path deviation of each flow data in each class is calculated.

7. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Abnormal behavior identification method and device

    CN111865941A