Abnormal information identification method and device, electronic equipment, computer readable medium
By acquiring user log data, determining traffic characteristics and sequence statistical dimensions, and using a gradient boosting decision tree model for anomaly detection, combined with naming rule analysis, the problem of incomplete data features in the identification of black and gray market users was solved, thus improving identification efficiency and accuracy.
Patent Information
- Application Number
- CN202310280301.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-21
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2043-03-21
AI Technical Summary
Existing technologies for identifying black and gray market users suffer from incomplete data features, insufficient recall, low accuracy, high cost of manual rule maintenance, and easy circumvention.
By acquiring user log data, we determine traffic characteristics and sequence statistical dimensions, use a gradient boosting decision tree model for anomaly detection, and combine naming rules to analyze and identify abnormal users.
It improves the efficiency and accuracy of abnormal user identification, enhances recall accuracy, reduces the maintenance cost of manual rules, and effectively identifies black and gray market users.
Smart Images

Figure CN116346453B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer application, in particular to the technical field of information security, and especially to an abnormal information identification method and device, an electronic device, a computer readable medium, and a computer program product. BACKGROUND
[0002] With the explosive development of Internet business, especially mobile Internet of Things, black and gray production users have evolved from the traditional routine of "attacking and penetrating the system for profit" to the mode of "taking advantage of the lack of business risk control for large-scale profit", and gradually formed a large-scale, well-defined black industry chain. Black and gray production users destroy the platform ecology and take away platform discounts, which is seriously harmful.
[0003] At present, the mining scheme of black and gray production users generally mines based on business data and attack environment. Due to the characteristics of large scale and clear division of labor of black and gray production, only mining abnormal user information from business data or attack environment may have the problems of incomplete data characteristics, insufficient recall rate of mining, and low accuracy of mining. SUMMARY
[0004] An abnormal information identification method and device, an electronic device, a computer readable storage medium, and a computer program product are provided.
[0005] According to a first aspect, an abnormal information identification method is provided, which includes: obtaining user log data in at least one scenario; determining traffic features in each scenario in a set time period based on the user log data; obtaining risk log data based on the traffic features and the user log data, the risk log data being user log data with abnormal values of the traffic features; obtaining sequence statistical dimension features of a user using a resource based on the risk log data; and obtaining an abnormal detection result of the user based on the sequence statistical dimension features and the risk log data.
[0006] According to a second aspect, an abnormal information identification device is provided, which includes: an obtaining unit configured to obtain user log data in at least one scenario; a determining unit configured to determine traffic features in each scenario in a set time period based on the user log data; a screening unit configured to obtain risk log data based on the traffic features and the user log data, the risk log data being user log data with abnormal values of the traffic features; an obtaining unit configured to obtain sequence statistical dimension features of a user using a resource based on the risk log data; and a detection unit configured to obtain an abnormal detection result of the user based on the sequence statistical dimension features and the risk log data.
[0007] According to a third aspect, an electronic device is provided, comprising at least one processor; and a memory connected with the at least one processor in communication, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any implementation of the first aspect.
[0008] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause a computer to perform the method according to any implementation of the first aspect.
[0009] According to a fifth aspect, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any implementation of the first aspect.
[0010] Embodiments of the present disclosure provide an abnormal information identification method and device, which first acquires user log data in at least one scenario; secondly, determines traffic features in each scenario in a set time period based on the user log data; thirdly, obtains risk log data based on the traffic features and the user log data, the risk log data being user log data with abnormal values of the traffic features; fourthly, obtains sequence statistical dimension features when a user uses a resource based on the risk log data; and finally, obtains an abnormal detection result of the user based on the sequence statistical dimension features and the risk log data. Thus, the user log data in at least one scenario is collected, ensuring the comprehensiveness of the analysis of the user behavior sequence in various scenarios; the risk log data obtained based on the traffic features of the user log data can effectively circumscribe the risk traffic; based on the circumscribed risk traffic, the sequence statistical dimension features that are difficult to capture by artificial expert rules are calculated, the cheating network abnormal user is mined, and the abnormal user identification efficiency, mining accuracy and recall accuracy are improved.
[0011] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0013] Figure 1 is a flowchart of one embodiment of the abnormal information identification method according to the present disclosure;
[0014] Figure 2 is a flowchart of another embodiment of the abnormal information identification method according to the present disclosure;
[0015] Figure 3is a framework diagram of one embodiment of an abnormal information identification method according to the present disclosure;
[0016] Figure 4 is a structural schematic diagram of one embodiment of an abnormal information identification device according to the present disclosure;
[0017] Figure 5 is a block diagram of an electronic device for implementing the abnormal information identification method of the embodiments of the present disclosure. DETAILED DESCRIPTION
[0018] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are included to provide a thorough understanding of embodiments of the present disclosure, and are taken as illustrative only. Therefore, it will be understood by those of ordinary skill in the art that various changes and modifications to the embodiments described herein can be made without departing from the scope and spirit of the present disclosure. Also, in the following description, descriptions of well-known functions and constructions are omitted for clarity and conciseness.
[0019] In the present embodiments, "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implying the number of the technical features indicated. Therefore, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features.
[0020] The black production cheating has a large-scale and clear division of labor black industry chain. It is difficult to discover user cheating from the source by data mining from the business data generated from traditional technology, that is, it is difficult to discover cheating in the number raising and number selling stages. This leads to incomplete analysis of data characteristics of black production, and insufficient recall rate of mining. At the same time, the rules, word tables and the like mined need to be defined by business experts, which has high maintenance cost and is easy to be bypassed by variant attacks.
[0021] Attack environment association mining has obvious scene isolation, simple features, and is easy to be bypassed by black production users. In addition, fields such as IP are shared by normal users and black production, which will lead to insufficient precision of mining and easy to cause injury to normal users.
[0022] In order to effectively analyze black production users, the behavior log data of all users is counted. From the statistical results, it can be seen that the black production users have many differences with normal users in user behavior sequence.
[0023] For example, black production users generally use machines to cheat in batches, and their behavior efficiency is generally higher than that of normal users, which is specifically reflected in that a certain resource of black production, such as a device, IP, or mobile phone number, will be used multiple times in a short period of time.
[0024] For example, since the machine batch cheating is different from the manual operation in nature, the sequence statistical dimension feature has characteristics, which are embodied in that the distribution of the entropy value of the time difference of the black production operation is obviously different from that of the normal user.
[0025] For example, in order to hide the IP and bypass the IP interception strategy, the black and gray production user often uses proxy IP, second dialing IP and other means to attack, and there will be space-time dimension features in the user behavior sequence, which are embodied in that the black and gray production user may operate in different geographical locations within a short time, which is obviously different from the normal user.
[0026] In view of the above analysis, the present disclosure provides an abnormal information recognition method which can effectively recognize the black and gray production user, Figure 1 The flow 100 of one embodiment of the abnormal information recognition method according to the present disclosure is shown, and the above abnormal information recognition method comprises the following steps:
[0027] Step 101, acquiring user log data in at least one scene.
[0028] In this embodiment, the user log data is also called user behavior trajectory and traffic log, and the user log data is the trajectory data recorded by the machine reflecting the user's operation on the resources on the network. Specifically, the user log data can be at least one kind of behavior trajectory data of the user operating the resources in real time.
[0029] In this embodiment, the collection, storage, use, processing, transmission, provision and disclosure of the user log data comply with the relevant laws and regulations and do not violate public order and good customs.
[0030] In this embodiment, the at least one scene is the scene of the user operating the resources, and the at least one scene can include: the scene of acquiring the user behavior data, including the account behavior scene data such as user registration, login, mobile phone binding / unbinding, setting user name, etc.; the at least one scene can also include: the user's business scene and the information state scene related to the user, resources, etc., wherein the business scene includes account active business scene data such as user private message, comment, attention, like, transaction, etc.; the information state scene includes: user identity verification state, email binding state, security protection state, bank card binding state, etc. account information state scene data.
[0031] In this embodiment, the user log data is the data of the user operating the resources at different times and in different places. Specifically, the user log data includes: account ID, account name, account binding mobile phone and email, account operation name, account operation time, account operation environment, account operation device, account operation proxy, account operation browser, account operation system, account operation terminal, etc.
[0032] Step 102, determine the traffic characteristics in each scene in the set time period based on the user log data.
[0033] In this embodiment, the traffic characteristics are data reflecting the traffic state of the user log data, and the change of the user log data can be effectively determined through the value of the traffic characteristics.
[0034] In this embodiment, the traffic characteristics are used to reflect the specific characteristics of the traffic, and it can be determined through the traffic characteristics which traffic is abnormal traffic (the average value of the traffic changes several times or several tens of times in the set time period) and which is normal traffic. The user log data with added traffic characteristics corresponding to the abnormal traffic is the risk log data, and the risk log data can not only determine the user behavior data, but also determine the real-time traffic volume when the user operates the resource.
[0035] Step 103, obtain the risk log data based on the traffic characteristics and the user log data.
[0036] In this embodiment, the risk log data is the user log data with abnormal value of the traffic characteristics. The user log data is screened, the user log data with abnormal value of the traffic characteristics is screened out, and the value of the traffic characteristics is added to the user log data to obtain the risk log data.
[0037] In this embodiment, the risk log data is also the user log data with added value of the traffic characteristics, and the added value of the traffic characteristics is greater than the normal value of the traffic characteristics.
[0038] Step 104, obtain the sequence statistical dimension characteristics of the user using the resource based on the risk log data.
[0039] In this embodiment, the sequence statistical dimension characteristics are used to reflect the behavior sequence change trend of the user using the resource. Since the sequence statistical dimension characteristics can be a specific value, the behavior sequence change trend of the user can be quantified through the specific value, and the differences of the behavior sequence change trends of the users can be compared.
[0040] In this embodiment, the risk log data has the time, place, and resource volume of the user operating the resource. The time sequence for operating different resources is determined through the time in the risk log data, and the geographical position coordinates for operating various resources are determined based on the time sequence. For each resource, the sequence of geographical position coordinates under the time sequence condition of the resource is determined, the change trend of each geographical position coordinate in the sequence of geographical position coordinates is determined and counted, and the sequence statistical dimension characteristics are obtained.
[0041] Step 105, obtain the abnormal detection result of the user based on the sequence statistical dimension characteristics and the risk log data.
[0042] In this embodiment, the abnormality detection result of the user includes: a probability that the user in the risk log data is an abnormal user. The abnormal user can be a black and gray production user.
[0043] In this embodiment, the traffic feature in the risk log data is used to reflect the longitudinal change amount of the user log data, and the sequence statistical dimension feature is used to reflect the transverse change amount of the user log data. When the longitudinal change amount and the transverse change amount of any user in the risk log data both exceed the respective change amount threshold, the user is determined to be an abnormal user.
[0044] Optionally, the step 105 includes: determining the value of the traffic feature of each user in the risk log data; determining the value of the sequence statistical dimension feature of each user in the risk log data; for each user, taking the square root of the difference between the square of the value of the traffic feature of the user and the square of the value of the sequence statistical dimension feature, to obtain a user dimension value, the user dimension value is used to reflect the characteristics of the user in the entire scene, and in response to the user dimension value being greater than a preset dimension threshold, the user is determined to be an abnormal user.
[0045] The abnormal information identification method provided by the embodiments of the present disclosure first acquires user log data in at least one scene; second, determines traffic features in each scene in a set time period based on the user log data; third, obtains risk log data based on the traffic features and the user log data, the risk log data being user log data with abnormal values of the traffic features; fourth, obtains sequence statistical dimension features of a user when using a resource based on the risk log data; and finally, obtains an abnormality detection result of the user based on the sequence statistical dimension features and the risk log data. In this way, the user log data in at least one scene is collected, ensuring the comprehensiveness of the analysis of the behavior sequence of the user in various scenes; the risk log data obtained based on the traffic features of the user log data can effectively circumscribe the risk traffic; based on the circumscribed risk traffic, the sequence statistical dimension features that are difficult to capture by artificial expert rules are calculated, the abnormal user of the cheating network is mined, and the efficiency of abnormal user identification, the mining accuracy, and the recall accuracy are improved.
[0046] In view of the fact that in some scenes, abnormal users (such as black and gray production) generally register accounts in batches, and the abnormal users generate usernames by using some username tools, the username generation of this kind is completely the same as that of normal users, and generally has a specific rule to follow, in order to better identify abnormal users, in some optional implementation manners of this embodiment, the at least one scene includes an account behavior scene, and the abnormality detection result includes a user identifier, and the method further includes: acquiring a username corresponding to the user identifier in the account behavior scene; performing naming rule analysis on the username to obtain a naming analysis result; and based on the naming analysis result, detecting whether the user corresponding to the user identifier is an abnormal user.
[0047] In this embodiment, the anomaly detection result includes: a user identifier and a confidence degree corresponding to the user identifier, where the confidence degree represents a probability that the user is an abnormal user.
[0048] In this embodiment, the account behavior scenario refers to a scenario related to an account in a user operation resource behavior, and the account behavior scenario includes: a user registration, a user login, a user mobile phone unbinding, a user name setting, and the like.
[0049] In this embodiment, the naming rule analysis refers to a naming manner of a user name, for example, some user names include letters, numbers, symbols, and the like. By analyzing the naming rules of a large number of users, it can be determined whether the user name is generated by a naming tool, thereby providing a reliable means for effectively checking abnormal users.
[0050] The anomaly information identification method provided in this embodiment includes: obtaining a user name corresponding to a user identifier in an anomaly detection result in an account behavior scenario; performing naming rule analysis on the user name to obtain a naming analysis result; and determining, based on the naming analysis result, whether a user corresponding to the user identifier is an abnormal user, thereby providing a reliable implementation manner for screening abnormal users in the account behavior scenario and improving the accuracy of abnormal user screening.
[0051] In some optional implementation manners of this embodiment, the naming rule analysis on the user name to obtain the naming analysis result includes: performing character segmentation on the user name to obtain a character segmentation result; and calculating a proportion of rare characters in the character segmentation result to obtain the naming analysis result including a rare character proportion. The determination, based on the naming analysis result, of whether the user corresponding to the user identifier is an abnormal user includes: determining whether the rare character proportion is greater than a preset rare proportion threshold; and in response to a determination that the rare character proportion is greater than the preset rare proportion threshold, determining that the user corresponding to the user identifier is an abnormal user.
[0052] In this embodiment, after obtaining the user identifier in the anomaly detection result, the user name corresponding to the user identifier is obtained; the user name is subjected to character segmentation to obtain a character segmentation result, where the character segmentation result is a character sequence composed of multiple individual characters; a proportion of rare characters in the character segmentation result is calculated to obtain a naming analysis result including a rare character proportion; and it is determined whether the rare character proportion is greater than a preset rare proportion threshold, and if the rare character proportion is greater than the preset rare proportion threshold, it is determined that the user corresponding to the user identifier is an abnormal user.
[0053] In the embodiment, the proportion of rare characters is equal to the number of rare characters in the character sequence divided by the number of all characters in the character sequence. The rare characters can be obtained by comparing each character in the character sequence with the rare characters in the preset rare character dictionary. The naming analysis result of the proportion of rare characters includes a user identifier, a user name, and the proportion of rare characters.
[0054] In the embodiment, the rare character proportion threshold can be adaptively set based on the account behavior scenario. For example, in the user name registration scenario, the rare character proportion threshold can be a value between 70% and 100%.
[0055] The method for detecting an abnormal user provided by the optional implementation manner is based on the split character result of the user name. Whether the proportion of rare characters is greater than the preset rare character proportion threshold is detected. When the proportion of rare characters is greater than the preset rare character proportion threshold, it is determined that the user with the user name is an abnormal user. Thus, a reliable implementation manner is provided for determining an abnormal user.
[0056] In some optional implementation manners of the embodiment, the naming rule analysis on the user name to obtain the naming analysis result includes: performing word segmentation on the user name to obtain a word segmentation result; and calculating the proportion of word groups in the word segmentation result to obtain a naming analysis result including a word group proportion. Whether the user corresponding to the user identifier is an abnormal user is detected based on the naming analysis result, including: detecting whether the word group proportion is greater than a preset word group proportion threshold; and in response to detecting that the word group proportion is greater than the preset word group proportion threshold, determining that the user corresponding to the user identifier is an abnormal user.
[0057] In the embodiment, after obtaining the user identifier of the abnormal detection result, the user identifier corresponding to the user name is obtained; the user name is subjected to word segmentation to obtain a word segmentation result, wherein the word segmentation result is a word sequence composed of multiple word groups and multiple individual characters; the proportion of word groups in the word segmentation result is calculated to obtain a naming analysis result including a word group proportion; and whether the word group proportion is greater than a preset word group proportion threshold is detected. If the word group proportion is greater than the preset word group proportion threshold, it is determined that the user corresponding to the user identifier is an abnormal user.
[0058] In the embodiment, the word group proportion is equal to the number of word groups in the word sequence divided by the number of all units (characters, word groups) in the word sequence. The word groups can be obtained by comparing each unit in the word sequence with the word groups in the preset word group dictionary. The naming analysis result of the word group proportion includes a user identifier, a user name, and the word group proportion.
[0059] In the embodiment, the word group proportion threshold can be adaptively set based on the account behavior scenario. For example, in the user name registration scenario, the word group proportion threshold can be a value between 50% and 100%.
[0060] The method for detecting an abnormal user provided by the optional implementation manner detects whether the proportion of the word group is greater than a preset word group proportion threshold based on the word segmentation result after the username is word segmented, and determines that the user with the username is an abnormal user when the proportion is greater than the preset word group proportion threshold, thereby providing another reliable implementation manner for determining an abnormal user.
[0061] In another embodiment of the present disclosure, the at least one scenario is an account behavior scenario, and the abnormal detection result includes a user identifier. The abnormal information identification method further includes: obtaining behavior log data corresponding to the user identifier in a related scenario; extracting related features of the behavior log data; and detecting whether the user corresponding to the user identifier is an abnormal user based on the related features.
[0062] In the embodiment, the related scenario is a scenario related to the account behavior scenario and having the same user, that is, the user in the related scenario and the account behavior scenario has the same user identifier. Specifically, the related scenario can be a private message scenario, a transaction scenario, or an interaction scenario.
[0063] In the embodiment, the behavior log data refers to user log data of the user in the related scenario, and obtaining the behavior log data corresponding to the user identifier in the related scenario refers to extracting behavior log data having the same user identifier as the account behavior scenario in the related scenario.
[0064] In the embodiment, the related features are features related to the related scenario. For example, when the related scenario is a transaction scenario, the related features are transaction features, which can include the number of transactions, transaction address information (multiple addresses in multiple transactions), and transaction amount distribution. For another example, when the related scenario is an interaction scenario, the related features are interaction features, which can include the proportion of followed accounts, the number of followings, and the number of C-end (Customer, usually a consumer, a personal terminal user uses a client) users followed.
[0065] In the embodiment, detecting whether the user corresponding to the user identifier is an abnormal user based on the related features includes detecting whether the value of the related features is greater than a preset feature threshold. If the value is greater than the preset feature threshold, the user corresponding to the user identifier is determined to be an abnormal user. For example, in the transaction scenario, the number of transactions of the same user identifier is greater than the number threshold (for example, 100 times), which indicates the possibility of brushing, and the user corresponding to the user identifier is determined to be an abnormal user.
[0066] Optionally, the detecting whether the user corresponding to the user identifier is an abnormal user based on the related features further includes: detecting whether the related features satisfy a preset feature distribution, and determining that the user corresponding to the user identifier is not an abnormal user if the related features satisfy the preset feature distribution. For example, in a transaction scenario, if the transaction amount distribution of the same user identifier belongs to a normal distribution, it is determined that the user corresponding to the user identifier is not an abnormal user.
[0067] The abnormal information identification method provided in this embodiment can obtain behavior log data in a related scenario based on the user identifier obtained in the account behavior scenario, extract related features, and detect abnormal users through the related features. The abnormal detection result in the account behavior scenario can be applied to the related scenario, and the comprehensiveness of abnormal user detection is improved.
[0068] In this embodiment, the sequence statistical dimension feature is a feature that cannot be captured by artificial expert rules. The sequence statistical dimension feature can be used to analyze the user behavior trajectory distribution in the risk log data in different dimensions. Based on the difference between the behavior trajectory distribution of the black and gray production and the normal user, the sequence statistical dimension feature can be determined from the time and space dimensions to deeply analyze the black and gray production user.
[0069] In some optional implementations of this embodiment, the sequence statistical dimension feature includes a geographical position sequence difference statistical value. The sequence statistical dimension feature of the user using the resource is obtained based on the risk log data, including: sorting a plurality of geographical position coordinates of the user using the resource in a time sequence based on the risk log data to obtain a geographical position coordinate sequence; starting from the first geographical position coordinate in the geographical position coordinate sequence, sequentially differencing adjacent two geographical position coordinates to obtain a geographical position difference sequence corresponding to the geographical position coordinate sequence; and counting all differences in the geographical position difference sequence to obtain a geographical position sequence difference statistical value.
[0070] In this optional implementation, the resource used or operated by the user is different, and the subject corresponding to the geographical position coordinate is different. When the resource used or operated by the user is an IP, the geographical position coordinate includes an IP home coordinate. When the resource used or operated by the user is a mobile phone, the geographical position coordinate includes a mobile phone number home coordinate and a user residence coordinate.
[0071] In this optional implementation, a statistical formula can be used to obtain the geographical position sequence difference statistical value through all differences in the geographical position difference sequence. The sequence difference statistical value is a value reflecting the behavior state of the user. The calculated sequence difference statistical value is compared with the sequence difference statistical value of the normal user. If the difference is small, it can be determined that the user corresponding to the sequence difference statistical value is a normal user, otherwise, it is an abnormal user.
[0072] In this embodiment, the geographical position sequence difference statistical value includes: mean value, maximum value, minimum value, median, mode, entropy, kurtosis, skewness. The above statistical geographical position difference value sequence all difference values, the geographical position sequence difference statistical value includes: the average of all difference values in the geographical position difference sequence; all difference values in the geographical position difference sequence are sorted in order of value size to obtain the maximum value and the minimum value; all difference values in the geographical position difference sequence are sorted in order of value size, and the middle value of all difference values is selected to obtain the median of the geographical position difference sequence; the difference value of the number of occurrences in the geographical position difference sequence is selected to obtain the mode; all difference values in the geographical position difference sequence are calculated based on the entropy calculation formula, as shown in formula (1), to obtain the entropy; all difference values in the geographical position difference sequence are calculated based on the kurtosis calculation formula, as shown in formula (2), to obtain the peak value; all difference values in the geographical position difference sequence are calculated based on the skewness calculation formula, as shown in formula (3), to obtain the skewness.
[0073]
[0074]
[0075]
[0076] In formula (1), n is the number of difference values in the geographical position difference sequence, and xi is each difference value in the geographical position difference sequence.
[0077] In formula (2), σ is the standard deviation of the geographical position difference sequence, μ is the mean value of the geographical position difference sequence, x' is each difference value in the geographical position difference sequence, and the peak value represents the characteristic number of the peak height of the probability density distribution curve at the mean value.
[0078] In formula (3), σ is the standard deviation of the geographical position difference sequence, μ is the mean value of the geographical position difference sequence, T is each difference value in the geographical position difference sequence, and the skewness is a measure of the skewness direction and degree of the statistical data distribution, and is a numerical feature of the degree of asymmetry of the statistical data distribution.
[0079] The sequence statistical dimension feature obtaining method provided by the optional implementation manner sorts the geographical position coordinates when the user uses the resource to obtain the geographical position coordinate sequence; the difference between two adjacent geographical position coordinates in the geographical position coordinate sequence is obtained to obtain the geographical position difference sequence corresponding to the geographical position coordinate sequence, and the geographical position sequence difference statistical value is obtained by statistical analysis of the geographical position difference sequence, which can deeply analyze the geographical position characteristics of the user and provide reliable identification basis for identifying abnormal users.
[0080] In some optional implementations of the embodiment, the sequence statistical dimension feature includes a time information sequence difference statistical value, and the sequence statistical dimension feature of the user using the resource is obtained based on the risk log data, including: sorting, based on the risk log data, a plurality of use time stamps of the user operating the resource in time sequence to obtain a use time stamp sequence; starting from a first use time stamp in the use time stamp sequence, sequentially differencing adjacent two use time stamps to obtain a time difference value sequence corresponding to the use time stamp sequence; and statistically counting all difference values in the time difference value sequence to obtain the time information sequence difference statistical value.
[0081] In the optional implementation, the resource can be an IP, a mobile phone number, an account, or a device, and the use time stamp is a time point of the user operating the resource, for example, the use time stamp is a time stamp of the user logging in or registering an account.
[0082] In the optional implementation, after the use time stamps are sorted in time sequence, the first use time stamp in the use time stamp sequence is the earliest time stamp of use, and for the first use time stamp, the difference between the adjacent two use time stamps is the second use time stamp minus the first use time stamp, to obtain a difference value between the two, which is the first difference value in the time difference value sequence.
[0083] In the optional implementation, statistically counting all difference values in the time difference value sequence is to analyze the distribution probability of the difference values in the time difference value sequence by using a statistical method to obtain an index result, find out the rules and characteristics of the user operating the resource, and determine the specific characteristics of each user operating the resource.
[0084] The method for obtaining the sequence statistical dimension feature provided in the optional implementation sorts the use time stamps of the user using the resource to obtain a use time stamp sequence, differs the adjacent two use time stamps in the use time stamp sequence to obtain a time difference value sequence corresponding to the use time stamp sequence, and statistically counts the time information sequence difference statistical value from the time difference value sequence, which can deeply analyze the operation time characteristics of the user and provide a reliable identification basis for identifying abnormal users.
[0085] Optionally, the sequence statistical dimension feature comprises: a time information sequence difference statistical value and a geographic position sequence difference statistical value, and the sequence statistical dimension feature of the user using the resource is obtained based on the risk log data, comprising: first, obtaining the geographic position sequence difference statistical value in the manner of the above embodiment; determining whether the geographic position sequence difference statistical value is less than a pre-obtained geographic position sequence difference statistical threshold of a normal user; in response to the geographic position sequence difference statistical value being less than the geographic position sequence difference statistical threshold of the normal user, sorting, based on the risk log data, a plurality of use time stamps of the user operating the resource in chronological order to obtain a use time stamp sequence; and starting from a first use time stamp in the use time stamp sequence, sequentially taking differences between adjacent two use time stamps to obtain a time difference value sequence corresponding to the use time stamp sequence; and obtaining a time information sequence difference statistical value by counting all the difference values in the time difference value sequence.
[0086] In this optional implementation, in view of the principle that the geographic position information has large distinguishability, the geographic position sequence difference statistical value is first obtained, and then the time information sequence difference statistical value is calculated based on the fact that the geographic position sequence difference statistical value cannot distinguish the abnormal user, and subsequently, the geographic position sequence difference statistical value and the time information sequence difference statistical value are simultaneously used to improve the efficiency of abnormal user identification.
[0087] In some optional implementations of this embodiment, the time information sequence difference statistical value comprises: an entropy value, a kurtosis value, a skewness value, a maximum value, a minimum value, and all the difference values in the statistical time difference value sequence, and obtaining the time information sequence difference statistical value comprises at least one of the following:
[0088] Based on the entropy calculation formula, all the difference values in the time information sequence are calculated to obtain the entropy value of the time information sequence, and specifically, the entropy calculation formula is shown in the above formula (1).
[0089] Based on the kurtosis calculation formula, all the difference values in the time information sequence are calculated to obtain the kurtosis value of the time information sequence, and specifically, the kurtosis calculation formula is shown in the above formula (2).
[0090] Based on the skewness calculation formula, all the difference values in the time information sequence are calculated to obtain the skewness value of the time information sequence, and specifically, the skewness calculation formula is shown in the above formula (3).
[0091] All the difference values in the time information sequence are sorted in order of value size to obtain the maximum value of the time information sequence, wherein the size order can be from large to small or from small to large, and when the difference values in the time information sequence are the same, the time indicated by the use time stamp corresponding to the difference value is sorted in size.
[0092] Sort all the difference values in the time information sequence in order of value size to obtain the minimum value of the time information sequence.
[0093] The optional implementation provides a difference value statistical value of the time information sequence, and the statistical value is based on entropy, kurtosis, skewness, maximum value and minimum value in statistics, so that the calculation amount is small and the implementation is simple.
[0094] Figure 2 A flow 200 of one embodiment of the abnormal information identification method according to the present disclosure is shown, and the abnormal information identification method includes the following steps:
[0095] In step 201, user log data in at least one scenario is obtained.
[0096] In step 202, based on the user log data, traffic features in each scenario in a set time period are determined.
[0097] In step 203, based on the traffic features and the user log data, risk log data is obtained.
[0098] In step 204, based on the risk log data, sequence statistical dimension features when the user uses the resource are obtained.
[0099] It should be understood that the operations and features in steps 201-204 above correspond to the operations and features in steps 101-104, respectively, and therefore the description of the operations and features in steps 101-104 above also applies to steps 201-204, and will not be repeated here.
[0100] In step 205, based on the risk log data, category features, numerical features, time features and spatial features are calculated respectively.
[0101] In an embodiment of the present disclosure, the category features, numerical features, time features and spatial features are specific features of the user when operating or using the resource.
[0102] The category features are the user type to which the user belongs in the risk log data, and the category features can be user information thickness, for example, whether the user has a username, whether the user is bound to a mobile phone, whether the user is bound to an email, whether the user sets a password, whether the user is bound to a bank card, whether the user completes identity real-name authentication, and whether the user completes face real-name verification.
[0103] The numerical feature is a behavior frequency feature of a user using a resource or operating a resource, and is used to reflect a user dimension cross statistical value. The numerical feature includes: a behavior frequency of a user in a period of time, a behavior frequency of an IP used by the user in a period of time, a behavior frequency of a device used by the user in a period of time, how many de-duplicated users use the IP used by the user in a period of time, how many device uses the IP used by the user in a period of time, and a mobile phone use frequency of the IP used by the user in a period of time.
[0104] The time feature is a time feature of a user using a resource or operating a resource, and the space feature is a geographical feature of a user using a resource or operating a resource. The time feature includes: whether the user operation is in the early morning, whether the user operation is in the office hours, whether the user operation is in the weekend, whether the user operation is in the holiday, a maximum value and a minimum value of a time difference between multiple operations of the user.
[0105] The size category feature, the numerical feature, the time feature, and the space feature based on the risk log data can exist simultaneously, and inputting the four features into the gradient boosting decision tree model can enable the gradient boosting decision tree model to effectively distinguish the user behavior information in the risk log data.
[0106] Optionally, when the information amount in the risk log data is limited, only one or more of the category feature, the numerical feature, the time feature, and the space feature can be used to input the gradient boosting decision tree model.
[0107] In step 206, an aggregated statistical dimension feature is obtained based on the risk log data.
[0108] In this embodiment, the aggregated statistical dimension feature is a feature obtained after data aggregation. Specifically, the aggregated statistical dimension feature can include: a total use frequency of an IP used by a user in a period of time, and a use user number of a device used by the user in a period of time. The aggregated statistical dimension feature can aggregate multiple attributes of a user together to obtain a comprehensive attribute value of the user, or aggregate multiple attributes of a device together to obtain a comprehensive attribute value of the device.
[0109] In step 207, the category feature, the numerical feature, the time feature, the space feature, the sequence statistical dimension feature, and the aggregated statistical dimension feature are simultaneously input into a gradient boosting decision tree model that has been pre-trained, to obtain an abnormal detection result of a user output by the gradient boosting decision tree model.
[0110] In this embodiment, the gradient boosting decision tree model is a pre-trained model, and the training process of the gradient boosting decision tree model is as follows:
[0111] (1) Construct positive and negative samples, and divide a training set and a test set.
[0112] The positive and negative sample construction includes: setting the black and gray production data or black and gray production suspicious data as positive samples, and the label is 1; setting the normal user or high value user as negative sample, and the label is 0.
[0113] The training set and test set division includes: dividing the data of the training model into a training set, and dividing the data of the test model into a test set. The division method can use the K-fold cross-validation method, and the training set is divided into k (k>1) sub-samples, and a single sub-sample is reserved as the data of the verification model, and the other k-1 samples are used for training. In actual training, k can be set to 5.
[0114] (2) Set the training parameters
[0115] The gradient boosting decision tree model is constructed, and the following training parameters need to be set:
[0116] Model category, learning rate, boosting method, maximum depth of tree model, number of leaf nodes on a tree, feature sampling ratio (randomly selected feature ratio per iteration), data sampling ratio (randomly selected data row ratio per iteration).
[0117] (3) Model training:
[0118] Use GBDT tools, such as LightGBM (Light Gradient Boosting Machine, a framework for implementing GBDT algorithm) to construct GBDT (Gradient Boosting Decision Tree, a histogram-based decision tree algorithm). The basic idea of the histogram algorithm is: first, discretize the continuous floating point feature value into k integers, and construct a histogram with a width of k. When traversing the data, use the discretized value as the index to accumulate the statistics in the histogram. After traversing the data once, the histogram accumulates the required statistics, and then according to the discrete value of the histogram, traverse to find the optimal split point.
[0119] On the basis of the histogram algorithm, the GBDT tool further optimizes it. First, it discards the layer-by-layer growth decision tree growth strategy used by most GBDT tools, and uses a leaf-by-leaf growth algorithm with depth limitation. Compared with layer-by-layer growth, leaf-by-leaf growth has the following advantages: under the same number of splits, it can reduce more error and get better accuracy; the disadvantage of leaf-by-leaf growth is that it may grow a relatively deep decision tree, causing overfitting. Therefore, the GBDT tool adds a maximum depth limit to the leaf-by-leaf generation algorithm to prevent overfitting while ensuring efficiency.
[0120] (4) Model evaluation
[0121] After the model training is completed, the recall rate, the precision rate and the F1-score of the model result are calculated to evaluate the model training effect. The precision rate is equal to the number of correctly classified positive samples divided by the sum of the correctly classified positive samples and the incorrectly classified positive samples. The recall rate is equal to the number of correctly classified positive samples divided by the sum of the correctly classified positive samples and the incorrectly classified negative samples. The F1-score is equal to twice the precision rate multiplied by the recall rate divided by the sum of the precision rate and the recall rate.
[0122] After the gradient boosting decision tree model training is completed, the strategy can output a risk score between 0 and 1. The higher the risk score, the higher the user abnormality degree. According to business requirements, the abnormality degree threshold or the proportion of abnormal users can be defined to realize dynamic management and corresponding risk strategies, and further disposal can be performed on the black and gray production users that meet the risk strategies.
[0123] In this embodiment, in terms of algorithm model, the GBDT tool model is used, which has the advantages of simple operation and stronger interpretability. Moreover, the model has the advantages of good training effect and not easy to overfit. Faster training speed and lower memory consumption are required for the business of black and gray production mining, which requires high timeliness and interpretability. In terms of evaluation method, the risk score between 0 and 1 is output, and the risk strategy and threshold can be dynamically adjusted. Different thresholds can be adopted in different disposal means, which is very flexible.
[0124] The abnormal information recognition method provided in this embodiment obtains risk log data, calculates sequence statistical dimension features, aggregation statistical dimension features, category features, numerical features, time features and space features based on the risk log data, inputs the category features, numerical features, time features, space features, sequence statistical dimension features and aggregation statistical dimension features into the gradient boosting decision tree model that has been pre-trained, so that the gradient boosting decision tree model uses statistical features and machine model strategies that are rich in data scenarios, comprehensive in features, complex in feature construction mode and difficult to capture by artificial expert rules in terms of data and features, outputs flexible risk strategies and thresholds, and increases the recall rate of mining compared with traditional single rule recognition.
[0125] In some optional implementation manners of this embodiment, based on the risk log data, the category features, numerical features, time features and space features are respectively calculated, including the following four items:
[0126] Based on the risk log data, the behavior categories of the behaviors of each user in each scenario are calculated to obtain the category features of each user. The behaviors of the user include username setting, secret authentication and resource attack history, and the user categories include whether the user exists, whether the user completes the secret authentication and whether the resource used by the user has an attack history.
[0127] Based on the risk log data, the number of behaviors of each user in each scene and in a set time period is calculated to obtain numerical features of each user; wherein the number of behaviors of the user is a value obtained after multiple operations in the risk log are counted.
[0128] Based on the risk log data, the time points of behaviors of each user in each scene and the time point difference between at least two operations are calculated to obtain time features of each user.
[0129] Based on the risk log data, the geographic positions of each user when using different resources in each scene are calculated to obtain spatial features of each user.
[0130] In the optional implementation manner, based on the risk log data, the category features, the numerical features, the time features and the spatial features are respectively calculated to reflect the risk information in the risk log data from multiple dimensions, thereby providing a reliable means for extracting abnormal information of abnormal users in all directions.
[0131] In some embodiments of the present disclosure, the above abnormal information identification method comprises: obtaining user log data in at least one scene; determining traffic features in each scene in a set time period based on the user log data; obtaining risk log data based on the traffic features and the user log data; obtaining sequence statistical dimension features of the user when using resources based on the risk log data; obtaining aggregated statistical dimension features based on the risk log data; and inputting the sequence statistical dimension features and the aggregated statistical dimension features into a gradient boosting decision tree model which is pre-trained to obtain an abnormal detection result of the user output by the gradient boosting decision tree model.
[0132] The abnormal information identification method provided in the embodiments determines the abnormal detection result of the user through the risk log data, the sequence statistical dimension features and the aggregated statistical dimension features, thereby providing a reliable implementation manner for identifying abnormal users.
[0133] In some embodiments of the present disclosure, the above abnormal information identification method comprises: obtaining user log data in at least one scene; determining traffic features in each scene in a set time period based on the user log data; obtaining risk log data based on the traffic features and the user log data; calculating category features, numerical features, time features and spatial features based on the risk log data; obtaining aggregated statistical dimension features based on the risk log data; and inputting the calculated category features, numerical features, time features, spatial features and aggregated statistical dimension features into a gradient boosting decision tree model which is pre-trained to obtain an abnormal detection result of the user output by the gradient boosting decision tree model.
[0134] The abnormal information identification method provided by the embodiment determines the abnormal detection result of the user by the risk log data, the sequence statistical dimension feature, the category feature, the numerical value feature, the time feature and the space feature, and provides another reliable implementation manner for identifying the abnormal user.
[0135] In order to analyze the black and gray production, the amount of data of the obtained user log data is large, in order to effectively analyze the gray and black production user, the risk log data in the user log data can be circled by using a feature with a small feature dimension, and the feature with a small feature dimension can be a traffic feature. Alternatively, the feature with a small feature dimension can also include a scene feature. The scene feature can exclude scenes in which black and gray production is not prone to occur.
[0136] In some optional implementations of the present disclosure, the traffic feature includes page access volume and independent visitor access number. The process feature in each scene in the set time period is determined based on the user log data, including determining the earliest time and the latest time of the user accessing different pages (with time records in the user log data), subtracting the earliest time from the latest time to obtain a set time period, counting the number of accesses to each page in the set time period based on the user log data, and taking the maximum access number as the page access volume; and counting the maximum number of users accessing the same page in the set time period as the independent visitor access number.
[0137] The risk log data is obtained based on the traffic feature and the user log data, including at least one or more of the following: in response to detecting that the page access volume corresponding to a page in the user log data is greater than a first preset threshold, adding the page access volume in the user log data corresponding to the page; and in response to detecting that the independent visitor access number corresponding to an IP in the user log data is greater than a second preset threshold, adding the page access volume in the user log data corresponding to the IP.
[0138] In this optional implementation, the black and gray production users generally use machines to cheat in batches, and the behavior efficiency is generally higher than that of normal users. Therefore, the first preset threshold and the second preset threshold can be experience values obtained based on analysis of black and gray production access traffic.
[0139] In this optional implementation, the user log data to which the page access volume is added and the user log data to which the page access volume is added are used as the risk log data.
[0140] The method for obtaining risk log data provided by the optional implementation manner takes the page access volume and the independent visitor access number as the traffic features, filters the user log data, can highlight the user log data with a large access volume, and therefore adds the traffic features to the user log data with the large access volume to obtain the risk log data. The method can focus on the user log data with risks in the user log data with a large data volume, improves the data processing efficiency, and improves the efficiency of the abnormal user analysis.
[0141] In some optional implementation manners of the embodiment, the traffic features further include: a factor aggregation degree of each resource; the factor aggregation degree is the usage amount of each resource of a user under limited available resources, and the factor aggregation degree includes: the number of registered users under an IP and the number of registered accounts under a device. A normal user can generally register seven or eight users per day under an IP, and if 20,000 users are registered under an IP per day, the user log data of the user with the IP is determined as the log data with risks.
[0142] The above determining the flow features in each scene in the set time period based on the user log data includes: determining a factor aggregation degree of each resource based on the user log data.
[0143] The above obtaining the risk log data based on the traffic features and the user log data includes: calculating a value of the factor aggregation degree of each resource in the user log data, and detecting whether the value of the factor aggregation degree of the resource is greater than an aggregation degree threshold; and in response to detecting that the value of the factor aggregation degree of the resource is greater than the aggregation degree threshold, adding the value of the factor aggregation degree in the user log data corresponding to the resource.
[0144] In the optional implementation manner, the aggregation degree threshold is an experience value obtained based on the analysis of the resource usage of the black and gray production.
[0145] In the optional implementation manner, when the value of the factor aggregation degree of the resource is greater than the aggregation degree threshold, the value of the factor aggregation degree is added in the user log corresponding to the resource, and the user log with the added value of the factor aggregation degree is taken as the risk log data.
[0146] The method for obtaining risk log data provided by the optional implementation manner takes the factor aggregation degree of each resource as the traffic features, filters the user log data, can highlight the user log data with a large resource usage amount, and therefore adds the traffic features to the user log data with the large resource usage amount to obtain the risk log data. The method can focus on the user log data with risks in the user log data with a large data volume, improves the data processing efficiency, and improves the efficiency of the abnormal user analysis.
[0147] In some optional implementations of the embodiment, the traffic feature further includes a resource operation success rate; and determining the traffic feature in each scenario in the set time period based on the user log data includes: determining, based on the user log data, a total amount of resource operations and a successful amount of resource operations of the user in each scenario in the set time period. For example, in the login scenario, the total number of times that the user logs into a certain page in the set time period is taken as the total amount of resource operations, and the number of times that the user successfully logs into the page is taken as the successful amount of resource operations.
[0148] Based on the traffic feature and the user log data, the risk log data is obtained by: calculating, for each resource in the user log data, a value of the resource operation success rate, and detecting whether the value of the resource operation success rate is less than a success threshold corresponding to the resource; and in response to detecting that the value of the resource operation success rate is less than the success threshold corresponding to the resource, adding the value of the success rate in the user log data of the resource.
[0149] In this optional implementation, the resource operation success rate is equal to the successful amount of resource operations divided by the total amount of resource operations. Since the probability of a black and gray production hitting a risk control rule is relatively high, the success rate of the black and gray production user when operating the resource is relatively low. Therefore, the success threshold is an empirical value obtained based on the analysis of the resource operation of the black and gray production user.
[0150] In this optional implementation, when the value of the resource operation success rate is less than the success threshold corresponding to the resource, the value of the success rate is added in the user log data corresponding to the resource, and the user log data with the added value of the success rate is taken as the risk log data.
[0151] The method for obtaining risk log data provided in this optional implementation takes the resource operation success rate as a traffic feature to filter the user log data, which can highlight the user log data of successful resource operations. Therefore, after adding the traffic feature to the user log data of successful resource operations, the user log data with the risk log data can be focused on, the data processing efficiency is improved, and the efficiency of analyzing abnormal users is improved.
[0152] In one embodiment of the present disclosure, as shown in Figure 3 FIG. 1 is a framework diagram of an embodiment of an abnormal information identification method of the present disclosure, in which Figure 3 In the embodiment, the user log data obtained is user log data in at least one scenario, and the scenario can be: a user registration scenario, a user login scenario, an account information state data scenario, and an account active business scenario. Therefore, the obtained user log data includes: user registration data, user login data, account information state data, and account active business scenario data. Figure 3In the method, the risk flow is obtained by calculating the user log data with abnormal flow characteristics, also referred to as risk log data. Figure 3 In the method, the feature construction is to calculate any one or more of the sequence statistical dimension feature, the category feature, the numerical feature, the time feature, the space feature, and the aggregated statistical dimension feature based on the risk log data, to obtain the feature information fully representing the black and gray production user. Figure 3 In the method, the anomaly detection is to analyze and determine the abnormal user in the risk log data by using machine learning, artificial rules, and other methods. The machine learning model can predict the user behavior by using the model, and output an abnormal risk score. The higher the risk score, the higher the degree of user abnormality. The abnormal information identification method can effectively identify the black and gray production user, and prevent the black and gray production user from causing damage and destruction to the business.
[0153] In one example of the present disclosure, the abnormal information identification method can be applied to the identification of black and gray production users. The identified black and gray production users can guide the anti-cheating business of different systems or platforms. Specifically, an anti-cheating server (a server executing the method of the present disclosure) obtains user log data of a plurality of users operating different resources in account behavior, user transaction, user interaction, and other scenarios. The account behavior can be user registration, login, setting a user name, and the like. The user transaction scenario is a scenario in which the user trades information such as goods. The user interaction scenario is a scenario in which the user comments on and likes the resource. The anti-cheating server calculates the page access volume, the number of independent visitor accesses, and the factor aggregation degree in each scenario in a certain time period based on the obtained user log data. The user log data that does not satisfy any one or more of the page access volume, the number of independent visitor accesses, and the factor aggregation degree is taken as risk log data. The risk log data is subjected to geographical position sequence difference value statistics and time information sequence difference value statistics to obtain geographical position sequence difference value statistics and time information sequence difference value statistics. The detection result of the user in the risk log data is obtained by analyzing the geographical position sequence difference value statistics and the time information sequence difference value statistics. The detection result includes the user identifier of the user and the abnormal risk score of each user. The abnormal risk score is between 0 and 1. The higher the abnormal risk score, the higher the degree of the user belonging to the black and gray production.
[0154] In this example, when the abnormal risk score is greater than a preset score threshold (for example, 0.8), the user with the abnormal risk score greater than the preset score threshold is marked and added to the cheating account library, thereby effectively preventing the black and gray production user from causing damage and destruction to the business.
[0155] Further referring to Figure 4 , as an implementation of the method shown in the above figures, the present disclosure provides one embodiment of an abnormal information identification device. The device embodiment and the method embodiment are the same asFigure 1 The method embodiments shown correspond to the device, which can be applied to various electronic devices.
[0156] As Figure 4 shown, the abnormal information recognition device 400 provided by the embodiment includes an acquisition unit 401, a determination unit 402, a screening unit 403, an obtaining unit 404, and a detection unit 405. The acquisition unit 401 can be configured to acquire user log data in at least one scenario. The determination unit 402 can be configured to determine traffic features in each scenario in a set time period based on the user log data. The screening unit 403 can be configured to obtain risk log data based on the traffic features and the user log data, the risk log data being user log data with abnormal values of the traffic features. The obtaining unit 404 can be configured to obtain sequence statistical dimension features of a user when using a resource based on the risk log data. The detection unit 405 can be configured to obtain an abnormal detection result of the user based on the sequence statistical dimension features and the risk log data.
[0157] In the embodiment, the specific processing of the acquisition unit 401, the determination unit 402, the screening unit 403, the obtaining unit 404, and the detection unit 405 in the abnormal information recognition device 400 and the technical effects brought by the specific processing can be respectively referred to the related description of steps 101, 102, 103, 104, and 105 in the corresponding embodiment. Figure 1 The related description of steps 101, 102, 103, 104, and 105 in the corresponding embodiment, which will not be repeated here.
[0158] In some optional implementation manners of the embodiment, the at least one scenario includes an account behavior scenario, and the abnormal detection result includes a user identifier. The device 400 further includes an analysis unit (not shown in the figure). The analysis unit can be configured to acquire a username corresponding to the user identifier in the account behavior scenario; perform naming rule analysis on the username to obtain a naming analysis result; and detect whether the user corresponding to the user identifier is an abnormal user based on the naming analysis result.
[0159] In some optional implementation manners of the embodiment, the analysis unit is further configured to divide the username into characters to obtain a character division result; calculate a proportion of rare characters in the character division result to obtain a naming analysis result including a rare character proportion; detect whether the rare character proportion is greater than a preset rare proportion threshold; and in response to detecting that the rare character proportion is greater than the preset rare proportion threshold, determine that the user corresponding to the user identifier is an abnormal user.
[0160] In some optional implementations of the present embodiment, the analysis unit is further configured to: perform word segmentation on the username to obtain a word segmentation result; calculate a proportion of a word group in the word segmentation result to obtain a naming analysis result including the proportion of the word group; detect whether the proportion of the word group is greater than a preset word group proportion threshold; and in response to detecting that the proportion of the word group is greater than the preset word group proportion threshold, determine that the user corresponding to the user identifier is an abnormal user.
[0161] In some optional implementations of the present embodiment, the at least one scenario is an account behavior scenario, and the abnormal detection result includes a user identifier. The apparatus 400 further includes a screening unit (not shown in the figure). The screening unit can be configured to: obtain behavior log data corresponding to the user identifier in a related scenario, the related scenario being a scenario related to the account behavior scenario and having the same user; extract a related feature of the behavior log data, the related feature being a feature related to the related scenario; and detect whether the user corresponding to the user identifier is an abnormal user based on the related feature.
[0162] In some optional implementations of the present embodiment, the sequence statistical dimension feature includes a geographical position sequence difference statistical value. The obtaining unit 404 is further configured to: sort, based on the risk log data, a plurality of geographical position coordinates in a time sequence when the user uses a resource, to obtain a geographical position coordinate sequence; sequentially subtract, starting from a first geographical position coordinate in the geographical position coordinate sequence, two adjacent geographical position coordinates to obtain a geographical position difference sequence corresponding to the geographical position coordinate sequence; and statistically count all difference values in the geographical position difference sequence to obtain the geographical position sequence difference statistical value.
[0163] In some optional implementations of the present embodiment, the sequence statistical dimension feature includes a time information sequence difference statistical value. The obtaining unit 404 is further configured to: sort, based on the risk log data, a plurality of use time stamps in a time sequence when the user operates a resource, to obtain a use time stamp sequence; sequentially subtract, starting from a first use time stamp in the use time stamp sequence, two adjacent use time stamps to obtain a time difference sequence corresponding to the use time stamp sequence; and statistically count all difference values in the time difference sequence to obtain the time information sequence difference statistical value.
[0164] In some optional implementations of the present disclosure, the obtaining unit 404 is further configured to perform at least one of the following five items:
[0165] Based on the entropy calculation formula, all the difference values in the time information sequence are calculated to obtain the entropy value of the time information sequence. Based on the kurtosis calculation formula, all the difference values in the time information sequence are calculated to obtain the kurtosis value of the time information sequence. Based on the skewness calculation formula, all the difference values in the time information sequence are calculated to obtain the skewness value of the time information sequence. All the difference values in the time information sequence are sorted in order of value size to obtain the maximum value of the time information sequence. All the difference values in the time information sequence are sorted in order of value size to obtain the minimum value of the time information sequence.
[0166] In some optional implementations of the present disclosure, the detection unit 405 includes an aggregation module (not shown in the figure), a first detection module (not shown in the figure). The aggregation module can be configured to obtain aggregated statistical dimension features based on the risk log data. The first detection module can be configured to input the sequence statistical dimension features and the aggregated statistical dimension features into a pre-trained gradient boosting decision tree model to obtain an abnormal detection result of the user output by the gradient boosting decision tree model.
[0167] In some optional implementations of the present disclosure, the detection unit 405 includes a calculation module (not shown in the figure), a second detection module (not shown in the figure). The calculation module can be configured to calculate category features, numerical features, time features, and spatial features based on the risk log data. The second detection module can be configured to input the category features, numerical features, time features, spatial features, and sequence statistical dimension features into a pre-trained gradient boosting decision tree model to obtain an abnormal detection result of the user output by the gradient boosting decision tree model.
[0168] In some optional implementations of the present disclosure, the calculation module is further configured to calculate, based on the risk log data, behavior categories to which behaviors of each user in each scenario belong, to obtain category features of each user; calculate, based on the risk log data, behavior frequencies of each user in each scenario and within a set time period, to obtain numerical features of each user; calculate, based on the risk log data, behavior time points of each user in each scenario and time point difference values between at least two operations, to obtain time features of each user; and calculate, based on the risk log data, geographic positions of each user when using different resources in each scenario, to obtain spatial features of each user.
[0169] In some optional implementations of the present disclosure, the traffic features include: page access volume, independent visitor access number. The screening unit 403 is further configured to perform any one or both of the following: in response to detecting that the page access volume corresponding to a page in the user log data is greater than a first preset threshold, adding the page access volume in the user log data corresponding to the page. In response to detecting that the independent visitor access number corresponding to an IP in the user log data is greater than a second preset threshold, adding the page access volume in the user log data corresponding to the IP.
[0170] In some optional implementations of the present disclosure, the traffic features include: factor aggregation degree of each resource; the screening unit 403 is further configured to: for each resource in the user log data, calculate the value of the factor aggregation degree of the resource, and detect whether the value of the factor aggregation degree of the resource is greater than an aggregation degree threshold; in response to detecting that the value of the factor aggregation degree of the resource is greater than the aggregation degree threshold, adding the value of the factor aggregation degree in the user log data corresponding to the resource.
[0171] In some optional implementations of the present disclosure, the traffic features include: resource operation success rate; the screening unit 403 is further configured to: for each resource in the user log data, calculate the value of the resource operation success rate, and detect whether the value of the resource operation success rate is less than a success threshold corresponding to the resource; in response to detecting that the value of the resource operation success rate is less than the success threshold corresponding to the resource, adding the value of the success rate in the user log data of the resource.
[0172] The abnormal information identification device provided by the embodiments of the present disclosure first acquires the user log data in at least one scenario by the acquisition unit 401; secondly, determines the traffic features in each scenario in a set time period based on the user log data by the determination unit 402; thirdly, obtains the risk log data based on the traffic features and the user log data by the screening unit 403, the risk log data being the user log data with abnormal values of the traffic features; fourthly, obtains the sequence statistical dimension features when the user uses the resource based on the risk log data by the obtaining unit 404; and finally, obtains the abnormal detection result of the user based on the sequence statistical dimension features and the risk log data by the detection unit 405. Thus, the user log data in at least one scenario is collected, ensuring the comprehensiveness of the sequence analysis of the user behavior in various scenarios; the risk log data obtained based on the traffic features of the user log data can effectively circumscribe the risk traffic; based on the circumscribed risk traffic, the sequence statistical dimension features that are difficult to capture by artificial expert rules are calculated, the cheating network abnormal user is mined, and the abnormal user identification efficiency, mining accuracy, and recall accuracy are improved.
[0173] In the technical solutions of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0174] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0175] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0176] As shown in Figure 5 The electronic device 500 includes a determination unit 501 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 can also be stored. The determination unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0177] A plurality of components in the electronic device 500 are connected to the I / O interface 505, including an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, a speaker, etc.; the storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the electronic device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0178] The determining unit 501 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the determining unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various determining units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The determining unit 501 performs various methods and processes described above, such as the abnormal information identification method. For example, in some embodiments, the abnormal information identification method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the determining unit 501, one or more steps of the abnormal information identification method described above can be performed. Alternatively, in other embodiments, the determining unit 501 can be configured to perform the abnormal information identification method by any other appropriate means, such as by means of firmware.
[0179] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0180] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, implements the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0181] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0182] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0183] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0184] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0185] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0186] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. An abnormal information identification method, the method comprising: obtaining user log data in at least one scenario; determining traffic features in each scenario in a set time period based on the user log data; obtaining risk log data based on the traffic features and the user log data, the risk log data being user log data with abnormal values of traffic features; obtaining sequence statistical dimension features of a user using a resource based on the risk log data; obtaining an abnormal detection result of the user based on the sequence statistical dimension features and the risk log data; the sequence statistical dimension features include a geographical position sequence difference statistical value, and the obtaining of the sequence statistical dimension features of the user using the resource based on the risk log data comprises: ordering, based on the risk log data, a plurality of geographical position coordinates of the user using the resource in time sequence to obtain a geographical position coordinate sequence; starting from a first geographical position coordinate in the geographical position coordinate sequence, sequentially differencing adjacent two geographical position coordinates to obtain a geographical position difference value sequence corresponding to the geographical position coordinate sequence; and statistically counting all difference values in the geographical position difference value sequence to obtain a geographical position sequence difference statistical value.
2. The method of claim 1, wherein, the at least one scenario includes an account behavior scenario, and the abnormal detection result includes a user identifier, and the method further comprises: obtaining a username corresponding to the user identifier in the account behavior scenario; performing naming rule analysis on the username to obtain a naming analysis result; detecting whether a user corresponding to the user identifier is an abnormal user based on the naming analysis result.
3. The method of claim 2, wherein, the performing of the naming rule analysis on the username to obtain the naming analysis result comprises: segmenting the username to obtain a segmentation result; calculating a proportion of rare characters in the segmentation result to obtain a naming analysis result including a rare character proportion ratio; the detecting of whether the user corresponding to the user identifier is the abnormal user based on the naming analysis result comprises: detecting whether the rare character proportion ratio is greater than a preset rare character proportion threshold value; in response to detecting that the rare character proportion ratio is greater than the preset rare character proportion threshold value, determining that the user corresponding to the user identifier is the abnormal user.
4. The method of claim 2, wherein, the performing of the naming rule analysis on the username to obtain the naming analysis result comprises: segmenting the username to obtain a segmentation result; calculating a proportion of word groups in the segmentation result to obtain a naming analysis result including a word group proportion ratio; the detecting of whether the user corresponding to the user identifier is the abnormal user based on the naming analysis result comprises: detecting whether the word group proportion ratio is greater than a preset word group proportion threshold value; in response to detecting that the word group proportion ratio is greater than the preset word group proportion threshold value, determining that the user corresponding to the user identifier is the abnormal user.
5. The method of claim 1, wherein, the at least one scenario is an account behavior scenario, the abnormal detection result includes a user identifier, and the method further comprises: obtaining behavior log data corresponding to the user identifier in a related scenario, the related scenario being a scenario related to the account behavior scenario and having the same user; extracting a related feature of the behavior log data, the related feature being a feature related to the related scene; detecting whether the user corresponding to the user identifier is an abnormal user based on the related feature.
6. The method of claim 1, wherein, The sequence statistical dimension feature includes: a time information sequence difference value statistical value, and the sequence statistical dimension feature of the user using the resource is obtained based on the risk log data, and includes: sequentially sorting a plurality of use time stamps of the user operating the resource based on the risk log data, to obtain a use time stamp sequence; starting from a first use time stamp in the use time stamp sequence, sequentially taking a difference between adjacent two use time stamps to obtain a time difference value sequence corresponding to the use time stamp sequence; statistically counting all difference values in the time difference value sequence to obtain a time information sequence difference value statistical value.
7. The method of claim 6, wherein, The statistical counting of all difference values in the time difference value sequence to obtain the time information sequence difference value statistical value includes at least one of the following: calculating all difference values in the time information sequence based on an entropy calculation formula to obtain an entropy value of the time information sequence; calculating all difference values in the time information sequence based on a kurtosis calculation formula to obtain a kurtosis value of the time information sequence; calculating all difference values in the time information sequence based on a skewness calculation formula to obtain a skewness value of the time information sequence; sequentially sorting all difference values in the time information sequence according to value size to obtain a maximum value of the time information sequence; sequentially sorting all difference values in the time information sequence according to value size to obtain a minimum value of the time information sequence.
8. The method of claim 1, wherein, The sequence statistical dimension feature and the risk log data are used to obtain an abnormal detection result of the user, and include: obtaining an aggregated statistical dimension feature based on the risk log data; inputting the sequence statistical dimension feature and the aggregated statistical dimension feature into a pre-trained gradient boosting decision tree model to obtain the abnormal detection result of the user output by the gradient boosting decision tree model.
9. The method of claim 1, wherein, The sequence statistical dimension feature and the risk log data are used to obtain an abnormal detection result of the user, and include: respectively calculating a category feature, a numerical feature, a time feature and a space feature based on the risk log data; simultaneously inputting the category feature, the numerical feature, the time feature, the space feature and the sequence statistical dimension feature into a pre-trained gradient boosting decision tree model to obtain the abnormal detection result of the user output by the gradient boosting decision tree model.
10. The method of claim 9, wherein, The risk log data is used to respectively calculate a category feature, a numerical feature, a time feature and a space feature, and includes: based on the risk log data, calculating a behavior category to which a behavior of each user under each scene belongs to obtain a category feature of each user; based on the risk log data, calculating a behavior frequency of each user in a set time period under each scene to obtain a numerical feature of each user; based on the risk log data, calculating a behavior time point of each user under each scene and a time point difference between at least two operations to obtain a time feature of each user; Based on the risk log data, geographical positions of each user using different resources in each scenario are calculated to obtain spatial features of each user.
11. The method according to one of claims 1-10, wherein, The traffic features include: page access volume, independent visitor access number; The risk log data is obtained based on the traffic features and the user log data, including at least one or more of the following: In response to detecting that the page access volume corresponding to a page in the user log data is greater than a first preset threshold, the page access volume is added to the user log data corresponding to the page; In response to detecting that the independent visitor access number corresponding to an IP in the user log data is greater than a second preset threshold, the page access volume is added to the user log data corresponding to the IP.
12. The method according to one of claims 1-10, wherein, The traffic features further include: factor aggregation degree of each resource; The risk log data is obtained based on the traffic features and the user log data, including: For each resource in the user log data, the value of the factor aggregation degree of the resource is calculated, and it is detected whether the value of the factor aggregation degree of the resource is greater than an aggregation degree threshold; In response to detecting that the value of the factor aggregation degree of the resource is greater than the aggregation degree threshold, the value of the factor aggregation degree is added to the user log data corresponding to the resource.
13. The method according to one of claims 1-10, wherein, The traffic features further include: resource operation success rate; The risk log data is obtained based on the traffic features and the user log data, including: For each resource in the user log data, the value of the resource operation success rate is calculated, and it is detected whether the value of the resource operation success rate is less than a success threshold corresponding to the resource; In response to detecting that the value of the resource operation success rate is less than the success threshold corresponding to the resource, the value of the success rate is added to the user log data of the resource.
14. An abnormal information identification device, the device comprising: an acquisition unit configured to acquire user log data in at least one scenario; a determination unit configured to determine traffic features in each scenario in a set time period based on the user log data; a screening unit configured to obtain risk log data based on the traffic features and the user log data, the risk log data being user log data with abnormal values of traffic features; an obtaining unit configured to obtain sequence statistical dimension features of a user using a resource based on the risk log data; a detection unit configured to obtain an abnormal detection result of the user based on the sequence statistical dimension features and the risk log data; The sequence statistical dimension features include: geographical position sequence difference statistical value, and the obtaining unit is further configured to: based on the risk log data, a plurality of geographical position coordinates of a user using a resource are sorted in time sequence to obtain a geographical position coordinate sequence; starting from the first geographical position coordinate in the geographical position coordinate sequence, two adjacent geographical position coordinates are sequentially subtracted to obtain a geographical position difference sequence corresponding to the geographical position coordinate sequence; All differences in the geographical position difference sequence are counted to obtain a geographical position sequence difference statistical value.
15. The apparatus of claim 14, wherein, The sequence statistical dimension feature comprises a time information sequence difference value statistical value, and the obtaining unit is further configured to: sort, based on the risk log data, a plurality of use time stamps of the user operating the resource in time sequence to obtain a use time stamp sequence; sequentially subtract adjacent two use time stamps from a first use time stamp in the use time stamp sequence to obtain a time difference value sequence corresponding to the use time stamp sequence; and count all difference values in the time difference value sequence to obtain a time information sequence difference value statistical value.
16. The apparatus of claim 15, wherein, The obtaining unit is further configured to perform at least one of the following five operations: calculating all difference values in the time information sequence based on an entropy calculation formula to obtain an entropy value of the time information sequence; calculating all difference values in the time information sequence based on a kurtosis calculation formula to obtain a kurtosis value of the time information sequence; calculating all difference values in the time information sequence based on a skewness calculation formula to obtain a skewness value of the time information sequence; sorting all difference values in the time information sequence in value size order to obtain a maximum value of the time information sequence; and sorting all difference values in the time information sequence in value size order to obtain a minimum value of the time information sequence.
17. The apparatus of claim 14, wherein, The detection unit further comprises: an aggregation module configured to obtain an aggregated statistical dimension feature based on the risk log data; a first detection module configured to input the sequence statistical dimension feature and the aggregated statistical dimension feature into a pre-trained gradient boosting decision tree model to obtain an abnormality detection result of the user output by the gradient boosting decision tree model.
18. The apparatus of claim 17, wherein, The detection unit further comprises: a calculation module configured to calculate, based on the risk log data, a category feature, a numerical feature, a time feature, and a space feature respectively; a second detection module configured to input the category feature, the numerical feature, the time feature, the space feature, and the sequence statistical dimension feature into a pre-trained gradient boosting decision tree model at the same time to obtain an abnormality detection result of the user output by the gradient boosting decision tree model.
19. The apparatus of claim 18, wherein the calculation module is further configured to: calculate, based on the risk log data, a behavior category to which a behavior of each user belongs in each scenario to obtain a category feature of each user; calculate, based on the risk log data, a behavior frequency of each user in each scenario and within a set time period to obtain a numerical feature of each user; calculate, based on the risk log data, a behavior time point of each user in each scenario and a time point difference value between at least two operations to obtain a time feature of each user; and calculate, based on the risk log data, a geographic location of each user when using different resources in each scenario to obtain a space feature of each user.
20. The apparatus of any of claims 14-19, wherein, The traffic features include: page access volume, independent visitor access number; the screening unit is further configured to perform any one or both of the following: in response to detecting that the page access volume corresponding to a page in the user log data is greater than a first preset threshold, adding the page access volume in the user log data corresponding to the page; in response to detecting that the independent visitor access number corresponding to an IP in the user log data is greater than a second preset threshold, adding the page access volume in the user log data corresponding to the IP.
21. The apparatus of one of claims 14-19, wherein, The traffic features include: factor aggregation degree of each resource; the screening unit is further configured to: for each resource in the user log data, calculate the value of the factor aggregation degree of the resource, and detect whether the value of the factor aggregation degree of the resource is greater than an aggregation degree threshold; in response to detecting that the value of the factor aggregation degree of the resource is greater than the aggregation degree threshold, adding the value of the factor aggregation degree in the user log data corresponding to the resource.
22. The apparatus of one of claims 14-19, the flow characteristics further comprising: Resource operation success rate; The screening unit is further configured to: for each resource in the user log data, calculate the value of the resource operation success rate, and detect whether the value of the resource operation success rate is less than the success threshold corresponding to the resource; in response to detecting that the value of the resource operation success rate is less than the success threshold corresponding to the resource, adding the value of the success rate in the user log data of the resource.
23. An electronic device, comprising: comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.
24. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-13.
25. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-13.
Citation Information
Patent Citations
CDN attack detection method and device, storage medium and electronic equipment
CN112367324A
Method, device and equipment for identifying access log and computer readable medium
CN112579418A
Abnormality detection method and device, equipment and medium
CN114218283A