A Community-Based Behavior Detection Method and Device
By constructing a similarity matrix and iterative information entropy to calculate the number of clubs and assessing user behavior, the problem of not having both coverage and manslaughter rate of order brushing detection in the existing technology is solved, and higher detection accuracy and flexible business adaptability are achieved.
Patent Information
- Application Number
- CN201810419220.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2018-05-04
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2038-05-04
AI Technical Summary
The aggregation strategy and classification prediction methods used in the prior art for order brushing detection cannot take into account both the coverage rate and the manslaughter rate, and the detection accuracy depends on the quality of the training data set, resulting in inaccurate detection.
By constructing a similarity matrix of user behavior characteristic data, calculating strong connectivity graphs and iterating information entropy to determine the number of clubs, evaluating user behaviors within clubs, and using modular functions for scoring to realize club division and real-time monitoring.
It improves the accuracy of order-brushing detection, reduces the manslaughter rate, and can be applied in a graded manner according to business needs to meet the needs of different scenarios.
Smart Images

Figure CN110443265B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a community-based behavior detection method and apparatus. Background Art
[0002] Currently, online shopping has become an important consumption habit in people's lives. Users can only judge whether a store is trustworthy through information such as the store's reputation, transaction volume, and buyer evaluations. These data will affect users' trust in merchants and directly determine whether a buyer will make a purchase at the store. However, a large number of fraudulent behaviors have emerged in these data that are supposed to truly reflect the business conditions of merchants, affecting users' judgment and achieving the merchants' purposes.
[0003] With the continuous improvement of risk control detection methods, the behavior of brushing orders has changed from traditional machine-based batch brushing to hiring people for false transactions. To some extent, traditional risk control strategies based on rules have certain limitations. In the prior art, clustering strategies and classification prediction methods are used to detect the above events. Among them, the clustering strategy evaluates the validity of an order by setting thresholds in one or a few dimensions such as IP, device number, and device fingerprint clustering degree. Classification prediction predicts the validity of a user in dimensions composed of the user's basic characteristics and behavior characteristics through training labeled training sample data.
[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0005] The clustering strategy for brushing order detection often cannot have both high coverage rate and low false positive rate, and cannot balance the two well. In addition, the premise of classification prediction is that there is a labeled data set. To some extent, the training data set determines the prediction effect, so the detection accuracy cannot be guaranteed. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a community-based behavior detection method and apparatus, which can solve the problem of inaccurate detection of abnormal events in the prior art.
[0007] To achieve the above object, according to one aspect of an embodiment of the present invention, there is provided a community-based behavior detection method, including obtaining user behavior feature data to construct a similarity matrix; calculating a strongly connected graph in the similarity matrix, merging pairwise users with a similarity equal to a preset threshold into the same community, and obtaining the final number of communities by iteratively calculating the information entropy of the similarity matrix; evaluating the user behavior within the community.
[0008] Optionally, obtaining the final number of communities by iteratively calculating the information entropy of the similarity matrix includes:
[0009] The information entropy formula of the similarity matrix:
[0010]
[0011] Among them, C win represents the sum of similarities within the w-th community; C wout represents the sum of similarities between the w-th community and other communities; ∑S ij represents the sum of similarities in the overall community matrix; m is the number of communities in the similarity matrix;
[0012] And
[0013] Among them, D i , D j represent the number of users connected to the i-th and j-th users respectively, and N i ∩N j represents the number of common behaviors of the i-th and j-th users.
[0014] Optionally, evaluating the user behaviors within the community includes:
[0015] Assume there are m communities and M evaluation dimensions. Calculate the total score of the evaluation dimensions as the result of evaluating the community. The formula is as follows:
[0016]
[0017] Among them, w = 1, 2, ……, m, Dallk represents the number of records excluding the white list for the k-th dimension, Dk represents the number of records obtained after removing duplicates after excluding the white list for the k-th dimension; h(t) is the step function, and the formula is as follows:
[0018]
[0019] Optionally, before constructing the similarity matrix, it includes:
[0020] Clean the user behavior data to fill in missing data and exclude abnormal data;
[0021] Preprocess the cleaned user behavior data to obtain the filtered user behavior feature data.
[0022] Optionally, it also includes:
[0023] According to the evaluation of the user behaviors within the community, conduct real-time monitoring of the user behaviors within the community.
[0024] In addition, according to an aspect of an embodiment of the present invention, a community-based behavior detection device is provided, including a construction module configured to obtain user behavior feature data to construct a similarity matrix; an evaluation module configured to calculate a strongly connected graph in the similarity matrix, merge pairs of users with a similarity equal to a preset threshold into the same community, and obtain the final number of communities by iteratively calculating the information entropy of the similarity matrix; and evaluate the user behaviors within the communities.
[0025] Optionally, the evaluation module obtaining the final number of communities by iteratively calculating the information entropy of the similarity matrix includes:
[0026] The information entropy formula of the similarity matrix:
[0027]
[0028] where C win represents the sum of similarities within the w-th community; C wout represents the sum of similarities between the w-th community and other communities; ∑S ij represents the sum of similarities in the overall community matrix; and m is the number of communities in the similarity matrix.
[0029] And
[0030] where D i , D j represent the number of users connected to the i-th and j-th users respectively, and N i ∩N j represents the number of common behaviors of the i-th and j-th users.
[0031] Optionally, the evaluation module evaluating the user behaviors within the communities includes:
[0032] Assume there are m communities and M evaluation dimensions. Calculate the total score of the evaluation dimensions as the result of evaluating the communities. The formula is as follows:
[0033]
[0034] where w = 1, 2, ……, m, Dall k represents the number of records excluding the whitelist for the k-th dimension, and D k represents the number of records obtained after removing duplicates after excluding the whitelist for the k-th dimension; h(t) is a step function, and the formula is as follows:
[0035]
[0036] Optionally, before the construction module constructs the similarity matrix, it includes:
[0037] Clean the user behavior data to fill in the missing data and exclude the abnormal data;
[0038] Preprocess the cleaned user behavior data to obtain the filtered user behavior feature data.
[0039] Optionally, the evaluation module is further configured to:
[0040] According to the evaluation of the user behavior within the community, perform real-time monitoring on the user behavior within the community.
[0041] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including:
[0042] One or more processors;
[0043] A storage device for storing one or more programs,
[0044] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the above-mentioned community-based behavior detection embodiments.
[0045] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any of the above-mentioned community-based behavior detection embodiments is implemented.
[0046] One of the above embodiments of the invention has the following advantages or beneficial effects: The present invention constructs a similarity matrix between users using the behavior characteristics of users, discovers user groups in the network, and is beneficial to user profiling. At the same time, a new modular function is proposed. By iteratively calculating the modular function value, the community division reaches the optimal, and conventional clustering algorithms can be avoided. By scoring the degree of cheating in the community, it can be applied at different levels according to different business needs to meet the needs of the business side.
[0047] The further effects of the above-mentioned non-conventional optional methods will be described in combination with specific embodiments below. Description of the Drawings
[0048] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:
[0049] Figure 1 is a schematic diagram of the main process of the community-based behavior detection method according to the embodiments of the present invention;
[0050] Figure 2 is a schematic diagram of the main process of the community-based behavior detection method according to the reference embodiments of the present invention;
[0051] Figure 3 Schematic diagram of the main modules of the community-based behavior detection device according to an embodiment of the present invention;
[0052] Figure 4 Exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0053] Figure 5 Schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. Detailed implementation manners
[0054] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0055] Figure 1 Community-based behavior detection method according to an embodiment of the present invention, as Figure 1 shown, the community-based behavior detection method includes:
[0056] Step S101, obtaining user behavior feature data to construct a similarity matrix.
[0057] Preferably, after obtaining the user behavior feature data, the user behavior data can be cleaned to fill in missing data and exclude abnormal data. Of course, the cleaned user behavior data can also be preprocessed to obtain filtered user behavior feature data, so that the user behavior feature data used to construct the similarity matrix is more accurate.
[0058] Step S102, performing community division according to the similarity matrix to evaluate the user behavior within the community.
[0059] Among them, the community (the community is Community, an attribute in complex network analysis) reflects the local characteristics of individual behaviors in the network and their mutual association relationships. Studying the communities in the network plays a crucial role in understanding the structure and function of the entire network, and can analyze and predict the interaction relationships of all elements in the entire network. The community structure in the network means that the vertices in the network can be divided into communities, the connections between vertices within the community are relatively dense, and the connections between vertices between communities are relatively sparse. The Community Detection community discovery algorithm is to distinguish groups or communities with relatively tight connections in relationship networks such as social networks.
[0060] In an embodiment, the specific implementation process of community division according to the similarity matrix includes:
[0061] Calculate the strongly connected graph in the similarity matrix to merge two users with a similarity equal to a preset threshold into the same community. Then, iteratively calculate the information entropy of the similarity matrix to obtain the final number of communities. Preferably, calculate the strongly connected graph in the similarity matrix to merge two users with a similarity of 1 into the same community.
[0062] Further, iteratively calculating the information entropy of the similarity matrix to obtain the final number of communities includes:
[0063] The information entropy formula of the similarity matrix:
[0064]
[0065] where C win represents the sum of similarities within the w-th community; C wout represents the sum of similarities between the w-th community and other communities; ∑S ij represents the sum of similarities in the overall community matrix; m is the number of communities in the similarity matrix;
[0066] while
[0067] where D i , D j represent the number of users connected to the i-th and j-th users respectively, and N i ∩N j represents the number of common behaviors of the i-th and j-th users.
[0068] In another embodiment, the specific implementation process of evaluating the behaviors of users within the community includes:
[0069] Suppose there are m communities and M evaluation dimensions. Calculate the total score of the evaluation dimensions as the result of evaluating the community. The formula is as follows:
[0070]
[0071] where w = 1, 2, ……, m, Dall k represents the number of records excluding the whitelist for the k-th dimension, and D k represents the number of records obtained after de-duplication of the k-th dimension after excluding the whitelist; h(t) is a step function, and the formula is as follows:
[0072]
[0073] It is also worth noting that after evaluating the user behavior within the community, the user behavior within the community can be monitored in real time, and situations such as scalpers hoarding goods and spike-order brushing can be controlled.
[0074] According to the various embodiments above, it can be seen that the described community-based behavior detection method can utilize the website behavior characteristics of users to construct a similarity matrix between users, discover user groups in the network, and facilitate user profiling. At the same time, a new modular function is proposed. By iteratively calculating the modular function value, the community division can reach the optimal state and can avoid conventional clustering algorithms. By scoring the cheating degree of the community, it can be applied at different levels according to different business needs to meet the requirements of the business side. By calculating the feature uniqueness score of users within the community to evaluate the severity of cheating in the community, of course, it can also be used in other scenarios besides order brushing. In addition, traditional rule methods belong to hard division, and to some extent, misclassification is inevitable. At the same time, this method belongs to soft division, and to some extent, it can reduce the misclassification rate.
[0075] Figure 2 It is a schematic diagram of the main process of the community-based behavior detection method according to the reference embodiments of the present invention. The community-based behavior detection method may include:
[0076] Step S201, collect user behavior feature data.
[0077] Among them, the collection of user behavior feature data is mainly through data logging to collect relevant information on user behavior, which may include: user registration feature data, browsing feature data, order placement feature data, etc. Preferably, the entire behavior of the user since opening the website will be logged, and relevant data when reporting to the user access interface at certain key positions (registration, coupon collection, order placement, etc.) will be recorded.
[0078] Step S202, clean the user behavior feature data.
[0079] In the embodiment, the cleaning of the user behavior feature data is mainly the filling of missing data and the exclusion of significantly abnormal data. Among them, the significantly abnormal data refers to data that violates the norm, such as the situation where the user's stay duration on the website is negative.
[0080] Preferably, when filling in the missing data, the missing value is filled as 0 or the median according to the situation. For example, if the PV (PV is a term in website analysis used to measure the number of web pages visited by website users) of the user on the web page is empty, it can be set to 0.
[0081] In addition, it is also worth noting that due to the different ways of user access behavior and certain reasons such as the network, user data may be missing or underreported. For such situations, the missing ratio can also be used for evaluation. When the missing ratio is greater than the preset threshold, the dimension can be directly discarded. Among them, the missing ratio refers to the proportion of the missing data in this dimension to the total data.
[0082] Step S203: Preprocess the cleaned user behavior feature data to obtain valuable user behavior feature data.
[0083] In the embodiment, mainly the collected data is subjected to feature screening to screen out valuable feature data. This part needs to consider the discrimination degree of each dimension and the correlation between dimensions (feature screening).
[0084] Preferably, when preprocessing, the method of calculating the dimension information entropy is used to eliminate the dimensions that are too scattered or concentrated (such as accessing the home page, because the home page of the entrance website is for the vast majority of users).
[0085] As a specific embodiment: Process the original data. Assume that the existing processed data is: {a1, a2, …, a 13}, with a total of 13 users. Among them, the devices of (a1, a2, a3, a4) four users when logging in on the same day are all A, the devices of (a5, a6, a7, a8) four users when logging in on the same day are all B, the devices of (a9, a 10 , a 11 , a 12 ) when logging in on the same day are all C, and the device of a 13 when logging in on the same day is D. There is also order placement IP information: The order placement IPs of (a1, a2, a3, a4) four users on the same day are all IP1, the order placement IPs of (a5, a6, a7) three users on the same day are all IP2, the order placement IPs of (a8, a9) two users on the same day are all IP3, and the order placement IPs of (a 10 , a 11 , a 12 ) three users on the same day are all IP4, and the order placement IP of a 13 on the same day is IP5. In addition, a1, a2, a3, a4, a5, a6, a7, a8 users have no browsing records, and a9, a 10 , a 11 , a 12 , a 13 all have browsing records.
[0086] Step S204: Construct a similarity matrix according to the user behavior feature data.
[0087] As an example, there are n users, and N behavioral characteristics of the users are taken (the behavioral characteristics may include: registration, browsing, placing an order, searching, etc.), and the behavioral similarity between users is calculated:
[0088]
[0089] where D i , D j represents the number of users connected to the i-th and j-th users respectively, and N i ∩N j represents the number of common behaviors of the i-th and j-th users. Thus, the similarity between users can be constructed, and at the same time, this matrix is a sparse matrix.
[0090] It should be noted that the so-called connected users refer to users who are consistent in a certain behavior, that is, connected in that behavior. And the so-called common behavior refers to the behavior of users being similar at a certain time node (for example, accessing a certain product, clicking on a certain page).
[0091] As a specific example: According to the example given in step S203, the following similarity matrix can be obtained in step S204:
[0092]
[0093] Step S205, calculate the strongly connected graph in the matrix to merge all pairs of users with a similarity of a preset threshold into the same community.
[0094] In the embodiment, in order to simplify the calculation and reduce the algorithm complexity, the strongly connected graph in the matrix is calculated before merging. Among them, the strongly connected graph itself has very strong properties, and the probability that users in the strongly connected graph belong to a community is extremely high. Preferably, in the obtained strongly connected graph, all users corresponding to a similarity value of 1 are merged into one community.
[0095] For example: User A and users B and C have bought the same product on a certain website, and have visited several pages (such as excluding the home page, the detail page of the purchased product, etc.), and have used the same IP address, shipping address, and all behavioral paths on the website, then it is considered that these users belong to the same community.
[0096] Therefore, step S205 can form some independent communities in the entire matrix, thereby improving the calculation efficiency of this implementation process.
[0097] As a specific example, according to the similarity matrix obtained in step S204, through the merging of the strongly connected graph, the following groups are obtained: (a1, a2, a3, a4), (a5, a6, a7), (a 10 , a11 , a 12 ), a8, a9, a 13 . There are a total of 6 small communities. Next, calculate the H value by successively merging the communities (independent nodes do not need to be merged, such as node a 13 ), and take the community division result when the maximum value is obtained.
[0098] Step S206, calculate the information entropy of the similarity matrix.
[0099] Among them, the information entropy represents the amount of information contained in the information. Preferably, the optimal number of communities is selected by iteratively calculating the information entropy of the similarity matrix.
[0100] The information entropy formula of the similarity matrix:
[0101] Where
[0102] Where, C win represents the sum of similarities within the w-th community. C wout represents the sum of similarities between the w-th community and other communities. ∑S ij represents the sum of similarities in the overall community matrix. m is the number of communities in the similarity matrix.
[0103] It should be noted that C wout can be obtained through the following process:
[0104] Calculate the total similarity value of the w-th community, and calculate the similarity value between users in other communities and users in the w-th community. Then sum the above two.
[0105] In addition, in the process of calculating the information entropy of the similarity matrix, the prerequisite for each merge is C win > C wout . Take the community division result corresponding to the maximum entropy value to obtain the optimal number of communities.
[0106] As a specific example, the community situations in various possible cases can be calculated according to the formula
[0107] respectively:
[0108] (1) Entropy value of the initial community: (a1, a2, a3, a4), (a5, a6, a7), (a 10 , a 11 , a 12 ), a8, a9, a 13 when the H value is: 0.82.
[0109] (2) Since the situations of a8 and a9 are the same, calculate the situation of a8, and the same is true for a9. Assign a8 to (a5, a6, a7), and the community is divided into: (a1, a2, a3, a4), (a5, a6, a7, a8), (a 10 ,a 11 ,a 12 )、a9、a 13 There are five communities in total, and the H value at this time is: 0.94.
[0110] (3) Assign a8 and a9 to (a5, a6, a7) and (a 10 ,a 11 ,a 12 ), the community is now divided into: (a1, a2, a3, a4), (a5, a6, a7, a8), (a9, a 10 ,a 11 ,a 12 ), a 13 There are four communities in total, and the H value at this time is: 1.07.
[0111] (4)a 13 It is an independent node and has no similarity nodes with the other three communities. The operation ends here.
[0112] Step S207: obtaining qualified communities from the community division results according to a preset threshold.
[0113] In the embodiment, the community division result obtained in step S206 includes multiple communities of different sizes, a threshold is set to extract certain communities that meet the conditions, and then a specific scenario is selected for application.
[0114] Among them, the conditions are as follows: to some extent, a community of a single node or a few nodes may appear, which is likely to be the result of a family member. A threshold is set here to avoid this situation.
[0115] As a specific example, through the calculation in step S206, it can be found that the H value is the largest when it is divided into 4 communities, so the division result at this time is taken.
[0116] Step S208, evaluating the obtained communities that meet the conditions.
[0117] In the embodiment, the obtained qualified communities are evaluated by the uniqueness principle. Further, the principle of the uniqueness principle is that for normal users, their behaviors are diverse, while abnormal users always have some common features (for example, multiple users have the same behavior in the same dimension).
[0118] Among them, the dimensions can be dimensions such as the IP dimension of community users, registration time, etc. When the IP dimension is selected for determination, the IP whitelist needs to be excluded. The IP whitelist refers to that the users in the whitelist will be processed preferentially. In addition, the IP dimension can refer to the behavior of access IP values such as the order placement IP, login IP, registration IP, etc.
[0119] It should be noted that the user dimension selected when judging the cheating degree of the community should be different from the user dimension (i.e., behavior characteristics) selected when calculating the similarity.
[0120] In a preferred embodiment, there are m communities and M evaluation dimensions. The total score of the dimensions is calculated as the result of evaluating the communities. The formula is as follows:
[0121]
[0122] Where w = 1, 2, ……, m, Dallk represents the number of records of the kth dimension after excluding the whitelist, and Dk represents the number of records obtained after removing duplicates after excluding the whitelist for the kth dimension; h(t) is a step function, and the formula is as follows:
[0123]
[0124] In this way, the cumulative uniqueness of the communities in these dimensions can be calculated. The higher the score, the lower the possibility of cheating of the community. Thus, each community has a different risk score.
[0125] As a specific example, since the users in the same community in the community division result are not necessarily problematic users, it is necessary to judge the community at this time.
[0126]
[0127] This formula is used to calculate whether each community is a cheating community. Calculate this value for the four communities respectively, and the calculation results are as follows:
[0128] The score of the community (a1, a2, a3, a4) is 4 / 13 * 4 + 5 / 13 * 4 = 36 / 13;
[0129] The score of the community (a5, a6, a7, a8) is 4 / 13 * 4 + 5 / 13 * 3 = 31 / 13;
[0130] (a9, a 10 , a 11 , a 12 ) The score of this community is 0. In this case, this reason needs to be verified, which may be due to reasons such as the IP being a public exit, etc.;
[0131] a 13It is an independent node with a score of 0. (In real data, this quantity should exist in large numbers, accounting for more than 50% of the whole). So far, communities with a score greater than 0 can be simply regarded as cheating groups (the specific threshold value should be determined according to the business situation).
[0132] Step S209, according to the evaluation results of the communities, monitor the users within the communities in real time.
[0133] In the embodiment, the evaluation results of the data are pushed to the real-time system to monitor the linkage of users within the same community in real time, and situations such as scalpers hoarding goods and flash sale brushing can be controlled.
[0134] As a specific embodiment, according to the calculation results in the specific embodiment of step S208, it can be seen that the possibilities (higher scores) of (a1, a2, a3, a4) and (a5, a6, a7, a8) belonging to the cheating group are higher than those of the other communities, and (a1, a2, a3, a4) is higher than (a5, a6, a7, a8). On the one hand, when the business side uses it, it can be used for business according to the risk level to ensure the normal operation of the business. On the other hand, since the accounts within the community often have linkage, when a certain account within the community has a certain behavior, the other associated accounts can be observed and restricted in order to make timely responses.
[0135] Figure 3 is a community-based behavior detection device according to an embodiment of the present invention, as Figure 3 shown, the community-based behavior detection device 300 includes a construction module 301 and an evaluation module 302. Among them, the construction module 301 obtains user behavior feature data to construct a similarity matrix. Then, the evaluation module 302 performs community division according to the similarity matrix to evaluate the user behavior within the community.
[0136] Furthermore, after obtaining the user behavior feature data, the construction module 301 can clean the user behavior data to fill in missing data and exclude abnormal data. Of course, the cleaned user behavior data can also be preprocessed to obtain the screened user behavior feature data, so that the user behavior feature data used for constructing the similarity matrix is more accurate.
[0137] As a preferred embodiment, the specific implementation process of the evaluation module 302 performing community division according to the similarity matrix includes:
[0138] Calculate the strongly connected graph in the similarity matrix to merge two users with a similarity of a preset threshold into the same community. Then, obtain the final number of communities by iteratively calculating the information entropy of the similarity matrix.
[0139] Further, obtaining the final number of communities by iteratively calculating the information entropy of the similarity matrix includes:
[0140] The information entropy formula of the similarity matrix:
[0141]
[0142] where C win represents the sum of similarities within the w-th community; C wout represents the sum of similarities between the w-th community and other communities; ∑S ij represents the sum of similarities in the overall community matrix; m is the number of communities in the similarity matrix;
[0143] And
[0144] where D i , D j represent the number of users connected to the i-th and j-th users respectively, and N i ∩N j represents the number of common behaviors of the i-th and j-th users.
[0145] In another embodiment, the specific implementation process of evaluating the behaviors of users within the community includes:
[0146] Suppose there are m communities and M evaluation dimensions, and calculate the total score of the evaluation dimensions as the result of evaluating the communities. The formula is as follows:
[0147]
[0148] where w = 1, 2, ……, m, Dall k represents the number of records excluding the white list in the k-th dimension, and D k represents the number of records obtained after removing duplicates after excluding the white list in the k-th dimension; h(t) is a step function, and the formula is as follows:
[0149]
[0150] It is also worth noting that after the evaluation module 302 evaluates the behaviors of users within the community, it can monitor the behaviors of users within the community in real time and control the occurrence of situations such as scalpers hoarding goods and spamming during flash sales.
[0151] It should be noted that the specific implementation content of the community-based behavior detection device described in the present invention has been described in detail in the above-mentioned community-based behavior detection method, so the repeated content will not be described here.
[0152] Figure 4Fig. 400 shows an exemplary system architecture to which the community-based behavior detection method or the community-based behavior detection device according to an embodiment of the present invention can be applied. Or Figure 4 Fig. 400 shows an exemplary system architecture to which the community-based behavior detection method or the community-based behavior detection device according to an embodiment of the present invention can be applied.
[0153] As Figure 4 shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405. The network 404 is used to provide a medium for communication links between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0154] Users can use the terminal devices 401, 402, 403 to interact with the server 405 through the network 404 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 401, 402, 403, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0155] The terminal devices 401, 402, 403 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0156] The server 405 may be a server providing various services, such as a background management server (only as an example) that supports shopping websites browsed by users using the terminal devices 401, 402, 403. The background management server may analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - only as examples) to the terminal devices.
[0157] It should be noted that the community-based behavior detection method provided by the embodiments of the present invention is generally executed by the server 405. Correspondingly, the community-based behavior detection device is generally disposed in the server 405.
[0158] It should be understood that Figure 4 the numbers of the terminal devices, the network, and the server in
[0159] are merely illustrative. According to actual needs, there may be any number of terminal devices, networks, and servers. Figure 5 Fig. 500 shows a schematic structural diagram of a computer system of a terminal device suitable for implementing the embodiments of the present invention. Figure 5The terminal device shown is only an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0160] As Figure 5 shown, computer system 500 includes a central processing unit (CPU) 501 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or programs loaded from a storage section 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0161] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as required. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as required so that a computer program read from it can be installed into the storage section 508 as required.
[0162] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509 and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present invention are executed.
[0163] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0165] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a construction module and an evaluation module. Among them, the names of these modules do not constitute a limitation to the modules themselves in some cases.
[0166] As another aspect, the present invention further provides a computer-readable medium, which may be included in the device described in the above embodiments; or may exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by a device, the device includes: obtaining user behavior characteristic data to construct a similarity matrix; calculating a strongly connected graph in the similarity matrix, merging pairs of users with a similarity of 1 into the same community, and obtaining the final number of communities by iteratively calculating the information entropy of the similarity matrix; evaluating the user behaviors within the community. According to the technical solution of the embodiments of the present invention, the problem of inaccurate detection of abnormal events in the prior art can be solved.
[0167] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A community-based behavior detection method, characterized in that, Including: Obtain user behavior characteristic data, calculate the behavior similarity between users based on the user behavior characteristic data to construct a similarity matrix; The user behavior characteristic data includes user registration characteristic data, browsing characteristic data, and order placement characteristic data; Calculate the strongly connected graph in the similarity matrix, merge two users with a similarity equal to a preset threshold into the same community, iteratively calculate the information entropy of the similarity matrix through the sum of behavior similarities within the community, the sum of behavior similarities between communities, the sum of behavior similarities in the overall similarity matrix, and the number of communities in the similarity matrix, take the community division result corresponding to when the information entropy is maximized, and obtain the final number of communities by obtaining the communities that meet the conditions in the community division result according to a preset threshold; The information entropy formula of the similarity matrix: Among them, C win represents the sum of similarities within the w-th community; C wout represents the sum of similarities between the w-th community and other communities; ∑S ij represents the sum of similarities in the overall community matrix; m is the number of communities in the similarity matrix; while Among them, D i , D j represents the number of users connected to the \(i\)-th and \(j\)-th users respectively, and \(N i \cap N j represents the number of common behaviors of the \(i\)-th and \(j\)-th users; Evaluate the user behavior within the community; the user behavior includes fraud behavior and brushing behavior.
2. The method according to claim 1, wherein Evaluating the user behavior within the community includes: Suppose there are m communities and M evaluation dimensions, calculate the total score of the evaluation dimensions as the result of evaluating the community, and the formula is as follows: where w = 1, 2, ……, m, Dall k represents the number of records with the white list removed for the k-th dimension, D k represents the number of records obtained after duplicate removal after removing the white list for the k-th dimension; h(t) is a step function, and the formula is as follows:
3. The method according to claim 1, characterized in that Before constructing the similarity matrix, including: Clean the user behavior data to fill in missing data and exclude abnormal data; Preprocess the cleaned user behavior data to obtain the filtered user behavior characteristic data.
4. The method according to claim 1, wherein Also including: According to the evaluation of the user behavior within the community, conduct real-time monitoring of the user behavior within the community.
5. A community-based behavior detection device, characterized in that, Including: A construction module for obtaining user behavior characteristic data and calculating the behavior similarity between users based on the user behavior characteristic data to construct a similarity matrix; The user behavior characteristic data includes user registration characteristic data, browsing characteristic data, and order placement characteristic data; An evaluation module for calculating the strongly connected graph in the similarity matrix, merging two users with a similarity equal to a preset threshold into the same community, iteratively calculating the information entropy of the similarity matrix through the sum of behavior similarities within the community, the sum of behavior similarities between communities, the sum of behavior similarities in the overall similarity matrix, and the number of communities in the similarity matrix, taking the community division result corresponding to when the information entropy is maximized, and obtaining the final number of communities by obtaining the communities that meet the conditions in the community division result according to a preset threshold; The information entropy formula of the similarity matrix: Among them, C win represents the sum of similarities within the w-th community; C wout represents the sum of similarities between the w-th community and other communities; ∑S ij represents the sum of similarities in the overall community matrix; m is the number of communities in the similarity matrix; while Among them, D i , D j represents the number of users connected to the \(i\)-th and \(j\)-th users respectively, and \(N i \cap N j represents the number of common behaviors of the \(i\)-th and \(j\)-th users; evaluate the user behaviors within the community; The user behavior includes fraud behavior and brushing behavior.
6. The device according to claim 5, characterized in that, The evaluation module evaluates the user behavior within the community, including: Suppose there are m communities and M evaluation dimensions, calculate the total score of the evaluation dimensions as the result of evaluating the community, and the formula is as follows: where w = 1, 2, ……, m, Dall k represents the number of records with the white list removed for the k-th dimension, D k represents the number of records obtained after deduplication of the k-th dimension after removing the white list; h(t) is a step function, and the formula is as follows:
7. The device according to claim 5, characterized in that, Before the construction module constructs the similarity matrix, including: Clean the user behavior data to fill in missing data and exclude abnormal data; Preprocess the cleaned user behavior data to obtain the filtered user behavior characteristic data.
8. The device according to claim 5, characterized in that The evaluation module is also used for: According to the evaluation of the user behavior within the community, conduct real-time monitoring of the user behavior within the community.
9. An electronic device, characterized in that, Including: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-4.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, the method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Complex network community mining method based on local minimum edges
CN104102745A