Method, apparatus, medium and computer device for restricting the behavior of user groups

By building a classifier and calculating the risk feature vector of user groups in the live broadcast platform, and determining high-risk user groups, the problem of insufficient identification accuracy in the existing technology is solved, more accurate behavioral restrictions are achieved, and the platform rights and interests are protected.

CN114065083BActive Publication Date: 2025-07-11WUHAN DOUYU NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010757412.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-31
Publication Date
2025-07-11
Estimated Expiration
2040-07-31

AI Technical Summary

Technical Problem

The prior art restriction identification accuracy of user group behavior in live broadcast platforms is insufficient, resulting in misjudgment and loss of rights and interests.

Method used

By extracting the user's risk characteristics, building a classifier, determining individual risk scores, splicing risk feature vectors, forming user groups, and calculating group risk scores and group differences, and determining whether the consensus score exceeds the threshold to identify high-risk user groups.

Benefits of technology

The accuracy of identification of high-risk user groups has been improved, the rights and interests of live broadcast platforms have been ensured, and the coverage has been increased by 10%-15%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114065083B_ABST
    Figure CN114065083B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, medium and computer device for restricting the behavior of user groups. For any business scenario in a live broadcast platform, risk characteristics of all users are extracted, and individual risk scores of each user in the business scenario are determined based on the risk characteristics; at least one user group is determined based on the risk characteristics of each user in all business scenarios to ensure the accuracy of risk judgment; in order to improve the determination accuracy, the group risk score and group divergence degree of the user group in the business scenario are determined, a consensus score of the user group in the business scenario is determined based on the group risk score and group divergence degree, and it is judged whether the user group is a high-risk group based on the consensus score; considering the batch behavior and internal divergence degree of the user group comprehensively, it is possible to accurately identify user groups with batch behavior and high synchronization in a certain business scenario, accurately restrict the behavior of the user group, and thus ensure the rights and interests of the live broadcast platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of risk control, and particularly to a method, device, medium and computer device for restricting the behavior of user groups. Background Art

[0002] On live streaming platforms, there are usually some user groups with similar characteristics. A user group may be formed by several small accounts registered by a natural person, or may be formed by illegal users through machine operations.

[0003] If a user group uses abnormal means to profit purposefully in certain business scenarios on the live streaming platform, it is necessary to restrict the behavior of the user group. However, to avoid misjudgment, generally only when it is determined that a certain user group is a high-risk user group in a certain business scenario, will the behavior of the user group be restricted.

[0004] In the prior art, when restricting the behavior of user groups, it is generally achieved by behavior synchronization and formulating a blacklist of users in business scenarios. However, if the behavior synchronization of some users does not indicate the relevance of these users, it may just be a coincidence. Therefore, this way of determining whether a user group is a high-risk user group in this scenario is likely to have misjudgments. When formulating a blacklist of group users for business scenarios, since there are many business scenarios on the live streaming platform, it is difficult to completely cover all business scenarios, so there may also be misjudgments. Summary of the Invention

[0005] In view of the problems existing in the prior art, embodiments of the present invention provide a method, device, medium and computer device for restricting the behavior of user groups, which are used to solve the technical problem that in the prior art, when restricting the behavior of user groups in business scenarios on a live streaming platform, due to the inability to ensure the recognition accuracy of whether a user group in a business scenario is a high-risk user group, the behavior of the user group cannot be accurately restricted, resulting in damage to the interests of the live streaming platform.

[0006] In the first aspect of the present invention, there is provided a method for restricting the behavior of user groups, which is applied to a live streaming platform. The method includes:

[0007] For any business scenario in the live streaming platform, extract the risk characteristics of all users;

[0008] Construct a classifier based on the risk characteristics, and based on the classifier and the risk characteristics of the users, determine the individual risk scores of each user in the business scenario; each user corresponds to one of the individual risk scores;

[0009] For each of the users, obtain the risk characteristics of the user in all business scenarios of the live streaming platform, splice the risk characteristics to obtain a risk characteristic vector;

[0010] Determine at least one user group based on the risk characteristic vectors of the users;

[0011] For each of the user groups, determine the importance of each user in the user group; based on the importance of each user in the user group and the individual risk score of each user in the business scenario, determine the group risk score of the user group in the business scenario;

[0012] Determine the group divergence degree of the user group in the business scenario; the group divergence degree is the deviation between the group risk score of the user group and the weighted risk scores of the users, and the weighted risk score is determined according to the individual risk score of the user and the importance of the user in the user group;

[0013] Determine the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence degree of the user group in the business scenario;

[0014] Judge whether the consensus score exceeds a preset consensus score threshold. If it exceeds, determine that the user group is a high-risk user group in the business scenario;

[0015] In the business scenario, restrict the behavior of the high-risk user group based on a preset behavior restriction strategy.

[0016] Optionally, the constructing a classifier based on the risk characteristics includes:

[0017] Determine sample data according to the risk characteristics; the sample data includes: positive sample data and negative sample data; the positive sample data includes the risk characteristics corresponding to users with historical risk behaviors, and the negative sample data includes the risk characteristics corresponding to users without historical risk behaviors;

[0018] Train the sample data based on a preset machine learning algorithm to obtain the classifier.

[0019] Optionally, the determining at least one user group based on the risk characteristic vectors of the users includes:

[0020] Step a, randomly select one of the risk characteristic vectors as the starting center vector C among all the risk characteristic vectors;

[0021] Step b, for any current risk characteristic vector F among the remaining risk characteristic vectorsu , determine whether the current feature vector satisfies the formula ||C - F u ||2 <= r; if it satisfies, store the user corresponding to the current risk feature vector into the user set; the r is a preset radius;

[0022] Step c, for the user set, according to the formula determine the offset vector S of the user set; the u is the user in the user set U;

[0023] Step d, determine the gradient and norm of the offset vector, and determine the first central vector C' according to the formula C' = Cgrad(S)||S||2;

[0024] Step e, iterate Steps b to d until ||S||2 < ε; the ε is an empirical coefficient, 0.01 < ε < 0.1; where, a corresponding user set is generated at the end of each iteration, and the user set corresponds to the user group.

[0025] Optionally, determining the importance of each user in the user group in the user group includes:

[0026] According to the formula determine the importance w(u1, g) of each user in the user group in the user group; where, the c(u1, g) is the activity of user u1 in the user group g, and the activity is the number of actions of user u1 in the scenario; the u1 is any target user in the user group g; the v is any user in the user group g.

[0027] Optionally, determining the group risk score of the user group in the business scenario based on the importance of each user in the user group in the user group and the individual risk score of each user in the business scenario includes:

[0028] According to the formula determine the group risk score gr(g, s) of the user group; where, the s is the business scenario, the g is any user group in the business scenario s, the u1 is any user in the user group g in the business scenario s; the r(u1, s) is the individual risk score of any user in the user group g in the business scenario s; w(u1, g) is the importance of any user in the user group g in the user group g.

[0029] Optionally, determining the group divergence degree of the user group in the business scenario includes:

[0030] According to the formula Determine the group divergence diff(g, s) of the user group in the business scenario; where w(u1, g)r(u1, s) is the weighted risk score of any user u1 in the user group g; gr(g, s) is the group risk score of the user group g; s is the business scenario, g is the user group, and u1 is any user in the user group g.

[0031] Optionally, determining the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence of the user group in the business scenario includes:

[0032] Determine the consensus score F(g, s) of the user group in the business scenario according to the formula F(g, s)=w*gr(g, s)+(1 - w)*(1 - diff(g, s)); where w is the weight coefficient, 0 < w < 1; gr(g, s) is the group risk score of the user group g; diff(g, s) is the group divergence of the user group g in the business scenario s.

[0033] In a second aspect of the present invention, there is provided a device for restricting the behavior of a user group, the device including:

[0034] An extraction unit for extracting the risk characteristics of all users for any business scenario in the live broadcast platform;

[0035] A determination unit for constructing a classifier based on the risk characteristics, and determining the individual risk score of each user in the business scenario based on the classifier and the risk characteristics of the user; each user corresponds to one individual risk score;

[0036] An acquisition unit for, for each user, acquiring the risk characteristics of the user in all business scenarios of the live broadcast platform, splicing the risk characteristics to obtain a risk feature vector;

[0037] Determine at least one user group based on the risk feature vectors of each user;

[0038] For each user group, determine the importance of each user in the user group in the user group; determine the group risk score of the user group in the business scenario based on the importance of each user in the user group in the user group and the individual risk score of each user in the business scenario;

[0039] Determine the group divergence degree of the user group in the business scenario; the group divergence degree is the deviation between the group risk score of the user group and the weighted risk scores of each user, and the weighted risk score is determined according to the individual risk score of the user and the importance of the user in the user group;

[0040] Determine the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence degree of the user group in the business scenario;

[0041] A judgment unit, configured to judge whether the consensus score exceeds a preset consensus score threshold. If it exceeds, it is determined that the user group is a high-risk user group in the business scenario;

[0042] A restriction unit, configured to restrict the behavior of the high-risk user group in the business scenario based on a preset behavior restriction policy.

[0043] In a third aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any one of the first aspects is implemented.

[0044] In a third aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in any one of the first aspects is implemented.

[0045] The present invention provides a method, device, medium, and computer device for restricting the behavior of user groups. For any business scenario in a live broadcast platform, risk characteristics of all users are extracted, and individual risk scores of each user in the business scenario are determined based on the risk characteristics; at least one user group is determined based on the risk characteristics of each user in all business scenarios; in this way, the user group is determined based on the risk characteristics of each user, so the accuracy can be ensured when making risk judgments; in order to further improve the accuracy of judgment, for each user group, the group risk score and group divergence degree of the user group in the business scenario are determined, and the consensus score of the user group in the business scenario is determined based on the group risk score and group divergence degree. If the consensus score exceeds the threshold, it is determined that the user group is a high-risk group; in this way, it is equivalent to comprehensively considering the batch behavior and internal divergence degree of the user group, so that user groups with batch behavior and high synchronization in a certain business scenario can be accurately identified, the identification accuracy is improved, the behavior of the user group can be accurately restricted, and thus the rights and interests of the live broadcast platform can be ensured. Description of the Drawings

[0046] Figure 1Schematic flowchart of a method for restricting the behavior of user groups provided in Embodiment 1 of the present invention;

[0047] Figure 2 Schematic structural diagram of a device for restricting the behavior of user groups provided in Embodiment 2 of the present invention;

[0048] Figure 3 Schematic structural diagram of a computer device for restricting the behavior of user groups provided in Embodiment 3 of the present invention;

[0049] Figure 4 Schematic structural diagram of a computer medium for restricting the behavior of user groups provided in Embodiment 3 of the present invention. Detailed implementation manners

[0050] In order to solve the technical problem in the prior art that when restricting the behavior of user groups in a business scenario on a live broadcast platform, due to the inability to ensure the recognition accuracy of whether a user group in a business scenario is a high-risk user group, the behavior of user groups cannot be accurately restricted, and thus the rights and interests of the live broadcast platform are damaged. The present invention provides a method, device, medium and computer device for restricting the behavior of user groups.

[0051] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Embodiment 1

[0053] This embodiment provides a method for restricting the behavior of user groups. As Figure 1 shown, the method includes:

[0054] S110, for any business scenario in the live broadcast platform, extract the risk characteristics of all users;

[0055] The business scenarios of the live streaming platform include multiple ones, such as: login scenario, bullet screen scenario, treasure box scenario, lottery scenario, and so on. In order to improve the accuracy of risk identification, for any business scenario in the live streaming platform, the risk characteristics of all users are extracted. Among them, the risk characteristics include: the total number of behaviors of users in the business scenario, the number of IPs used by users in the business scenario, the number of devices used by users in the business scenario, and the number of clients used by users in the business scenario, and so on. Because when an abnormal user group conducts abnormal operations, devices and IPs are repeatedly used among users many times, and the behavior of each user is often high-frequency. Therefore, the above several risk characteristics are essential information parameters for improving the accuracy of identifying high-risk user groups. They are the traces left by users after use, objectively existing, not selected by subjective human factors, but obtained for solving technical problems regarding the total number of behaviors of the above users in the business scenario, the number of IPs used by users in the business scenario, the number of devices used by users in the business scenario, and the number of clients used by users in the business scenario (that is, selected in line with natural laws).

[0056] S111, construct a classifier based on the risk characteristics, and determine the individual risk score of each user in the business scenario based on the classifier and the risk characteristics of the user; each user corresponds to one individual risk score;

[0057] After extracting the risk characteristics of all users in each business scenario, construct a classifier based on the risk characteristics; specifically including:

[0058] Determine sample data according to the risk characteristics; the sample data includes: positive sample data and negative sample data; the positive sample data includes the risk characteristics corresponding to users with historical risk behaviors, and the negative sample data includes the risk characteristics corresponding to users without historical risk behaviors;

[0059] Train the sample data based on a preset machine learning algorithm to obtain a classifier.

[0060] After the classifier is constructed, determine the individual risk score of each user in the business scenario based on the classifier and all the risk characteristics of the user in this scenario; each user corresponds to one individual risk score.

[0061] For example, if all the risk characteristics of user u in business scenario s are f s , so the individual risk score r(u, s) of user u in business scenario s can be determined according to the formula r(u, s) = C s (f s ); where C s is the classifier. User u is any user in scenario s.

[0062] S112. For each of the users, obtain the risk characteristics of the user in all business scenarios of the live streaming platform, splice the risk characteristics to obtain a risk characteristic vector.

[0063] For each user, obtain the risk characteristics of the user in all business scenarios of the live streaming platform, splice the risk characteristics to obtain a risk characteristic vector; all the risk characteristic vectors form a user vector space.

[0064] For example, if the live streaming platform includes 2 scenarios, user u has 10 risk characteristics in the first scenario s and 10 risk characteristics in the second scenario t, then splice the risk characteristics in the two scenarios to form a 20-dimensional risk characteristic vector F. u 。

[0065] S113. Determine at least one user group based on the risk characteristic vectors of the users.

[0066] After the risk characteristic vectors of each user are determined, it is necessary to determine at least one user group based on the risk characteristic vectors of the users. Specifically, it includes:

[0067] Step a. Randomly select a risk characteristic vector from all the risk characteristic vectors as the starting center vector C.

[0068] Step b. For any current risk characteristic vector F among the remaining risk characteristic vectors u , determine whether the current characteristic vector satisfies the formula ||C - F u ||² ≤ r; if it is satisfied, store the user corresponding to the current risk characteristic vector in the user set; r is a preset radius.

[0069] Step c. For the user set, determine the offset vector S of the user set according to the formula ; at this time, user u is a user in the user set.

[0070] Step d. Determine the gradient and modulus length of the offset vector, and determine the first center vector C' according to the formula C' = C + grad(S)||S||²; the gradient of the offset vector is grad(S), and the modulus length of the offset vector is ||S||².

[0071] Step e, iterate through Step b to Step d until ||S||2 < ε; ε is an empirical coefficient, 0.01 < ε < 0.1; where, a corresponding user set is generated at the end of each iteration, and the user set is the user group. The preferred value of ε can be 0.5; in order to ensure the accuracy of user clustering, therefore, the value of ε in this embodiment needs to be greater than 0.01; in order to reduce the iteration error and thus ensure the determination accuracy of the user group, the value of ε needs to be less than 0.1. If the value of ε is less than 0.01, the user groups will be divided very finely, and there will be many small groups with low confidence, resulting in inaccurate group identification; if the value of ε is greater than 0.1, the scale of the user group is too large, and some user groups are not split, which will also cause inaccurate group identification.

[0072] This way of continuously iterating to generate each user group can improve the accuracy of user clustering based on risk characteristics, ensure that the behaviors of users in each user group are consistent, and thus ensure the determination accuracy of the user group.

[0073] S114, for each of the user groups, determine the importance of each user in the user group in the user group; based on the importance of each user in the user group in the user group and the individual risk score of each user in the business scenario, determine the group risk score of the user group in the business scenario.

[0074] After determining at least one user group, for each user group, determine the importance of each user in the user group in the user group; based on the importance of each user in the user group in the user group and the individual risk score of each user in the business scenario, determine the group risk score of the user group in the business scenario.

[0075] As an alternative embodiment, determining the importance of each user in the user group in the user group includes:

[0076] According to the formula Determine the importance w(u1,g) of each user in the user group g in the user group; where, c(u1,g) is the activity of user u1 in the user group g, is the sum of the activities of all users in the user group g; g is any user group; the activity is the number of behaviors of user u1 in the business scenario; user u1 is any target user in the user group g, user v is any user in the user group g, and user v includes user u1.

[0077] Here, the activity of user u1 in the user group g can be determined according to the formula Determine; where, s is any business scenario; X is all business scenarios in the live broadcast platform, and c(u1,g) is the number of behaviors of user u1 in the business scenario s.

[0078] The calculation principle of w(u1, g) is as follows: First, determine the activity of any user u1 in the user group g and the total activity of all users in the group. Then, take the ratio of the activity of user u1 in the group to the total activity as the importance of each user in the user group. It can be seen that the higher the activity of each user in the user group, the higher the importance of the user in the user group.

[0079] After determining the importance of each user in the user group, according to the formula Determine the group risk score gr(g, s) of the user group g; where s is the business scenario, g is any user group in the business scenario s, u1 is any user in the user group g within the business scenario s; r(u1, s) is the individual risk score of any user in the user group g in the business scenario s; w(u1, g) is the importance of any user in the user group g in the user group g, and |g| is the number of users in the user group g.

[0080] Here, the calculation principle of the group risk score gr(g, s) of the user group is: perform a weighted average on the individual risk scores of the users within the user group, so the sum of the weighted risk scores of all users within the user group can be obtained. w(u1, g) can be used as a weight, so the group risk score can also be understood as the weighted average score of user risks.

[0081] S115. Determine the group divergence of the user group in the business scenario; the group divergence is the deviation between the group risk score of the user group and the weighted risk scores of each user, and the weighted risk score is determined according to the individual risk score of the user and the importance of the user in the user group.

[0082] After determining the group risk scores of each user group, according to the formula Determine the group divergence diff(g, s) of the user group g in the business scenario s; where w(u1, g)r(u1, s) is the weighted risk score of any user u1 in the user group g; gr(g, s) is the group risk score of the user group g; s is the business scenario, g is any user group in the business scenario s, and u1 is any user in the user group g.

[0083] Here, the calculation principle of the group divergence diff(g, s) is as follows: for any user group, the group divergence is the deviation between the group risk score gr(g, s) of the user group and the weighted risk scores w(u1, g)r(u1, s) of each user. The greater the difference between the two, the greater the divergence between the user risk and the group risk. Use (w(u1, g)r(u1, s) - gr(g, s)) 2 to measure the divergence of a single user, and the sum of the divergences of individual users is used to obtain the final group divergence diff(g, s).

[0084] S116. Determine the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence of the user group in the business scenario;

[0085] After the group divergence of the user group is determined, the consensus score of the user group in the business scenario is determined according to the group risk score of the user group and the group divergence of the user group in the business scenario.

[0086] Specifically, the consensus score F(g, s) of the user group in the business scenario is determined according to the formula F(g, s) = w * gr(g, s) + (1 - w) * (1 - diff(g, s)); where w is the weight coefficient, 0 < w < 1; gr(g, s) is the group risk score of the user group g; diff(g, s) is the group divergence of the user group g in the business scenario s.

[0087] The calculation principle of F(g, s) is as follows: determine the weight coefficient of the group divergence and the weight coefficient of the group risk score, and then obtain the consensus score of the user group based on the group divergence, the group risk score of the user group, and the corresponding weight coefficients. Among them, the weight coefficient is determined according to the tolerance degree of the business scenario for the group risk. The higher the tolerance degree, the lower the corresponding group risk score; the lower the tolerance degree, the higher the corresponding group risk score.

[0088] It can be seen that the higher the group risk score and the lower the group divergence, the higher the degree of risk behavior and consistency of the group in the business scenario. At this time, the probability that the user group is identified as a high-risk group in this business scenario is also greater.

[0089] When determining the consensus score of a user group in this embodiment, both the group risk score and the group divergence are taken into account, which can ensure the accuracy of determining high-risk user groups. Because if only the group risk score is considered, if the group risk score is high, it may be caused by some high-scoring users in the user group, and the score differences among users are relatively large. Then, there is no batch behavior in this business scenario, and it is not suitable to punish all users in the user group. Therefore, this user group cannot be directly determined as a high-risk group. If only the group divergence is considered, if the group divergence is small, it is also possible that the group risk score is relatively small. This situation indicates that although the users in the user group have a certain degree of synchronization in behavior, there is no high risk. At this time, it is also not suitable to punish the users in the user group.

[0090] S117. Determine whether the consensus score exceeds a preset consensus score threshold. If it exceeds, determine that the user group is a high-risk user group in the business scenario.

[0091] After the consensus score of the user group in the business scenario is determined, determine whether the consensus score exceeds the preset consensus score threshold. If it exceeds, determine that the user group is a high-risk user group in the business scenario. At this time, it is necessary to punish the users in the user group.

[0092] For example, in a login scenario, a user group includes 4 users. The individual risk scores of these 4 users are 0.75, 0.6, 0.7, and 0.8 respectively. The importance levels of these 4 users in the user group are 0.3, 0.2, 0.3, and 0.2 respectively.

[0093] Then, the group risk score of the user group is:

[0094] gr(g, s) = 0.75 * 0.3 + 0.6 * 0.2 + 0.7 * 0.3 + 0.8 * 0.2 = 0.715;

[0095] The group divergence is:

[0096] diff(g, s) = 0.3 * (0.75 - 0.715) 2 + 0.2 * (0.6 - 0.715) 2 + 0.3 * (0.7 - 0.715) 2 + 0.2 * (0.8 - 0.715) 2 = 0.0045

[0097] Take w = 0.5;

[0098] The consensus score is:

[0099] F(g, s) = 0.5 * 0.715 + 0.5 * (1 - 0.0045) = 0.86

[0100] If the threshold of the consensus score is 0.7, it indicates that this user group is a high-risk user group in the login business scenario, and it is necessary to restrict the behaviors of the users in this user group, such as restricting logins, etc.

[0101] Using the method provided in this embodiment, high-risk user groups and high-risk users can be accurately identified. Compared with the prior art, this embodiment takes into account both the group risk score and the group divergence. In this way, the accuracy of the determined high-risk user groups can be ensured, and furthermore, the coverage of high-risk users can be improved. Compared with the identification method of the prior art, the coverage of high-risk users can be increased by 10% - 15%.

[0102] S118. In the business scenario, restrict the behaviors of the high-risk user group based on a preset behavior restriction policy.

[0103] If it is determined that a user group is a high-risk user group in a certain business scenario, then restrict the behaviors of the high-risk user group based on a preset behavior restriction policy.

[0104] For example, in the lottery scenario, if it is determined that a certain user group is a high-risk user group in the lottery scenario, then temporarily block the IPs of all users in this user group and prohibit lottery participation.

[0105] The method for restricting the behaviors of user groups provided in this embodiment extracts the risk characteristics of all users for any business scenario in the live broadcast platform, and determines the individual risk scores of each user in the business scenario based on the risk characteristics; determines at least one user group based on the risk characteristics of each user in all business scenarios; in this way, the user group is determined based on the risk characteristics of each user, so the accuracy can be ensured when making risk judgments; to further improve the accuracy of the judgment, for each user group, determine the group risk score and the group divergence of the user group in the business scenario, determine the consensus score of the user group in the business scenario based on the group risk score and the group divergence, and if the consensus score exceeds the threshold, determine that this user group is a high-risk group; in this way, it is equivalent to comprehensively considering the batch behaviors and internal divergences of this user group, so user groups with batch behaviors and high synchronization in a certain business scenario can be accurately identified, the identification accuracy is improved, the behaviors of user groups can be accurately restricted, and thus the rights and interests of the live broadcast platform can be ensured.

[0106] Based on the same inventive concept, the present invention also provides a device for restricting the behaviors of user groups. For details, see Embodiment 2.

[0107] Embodiment 2

[0108] This embodiment provides a device for restricting the behavior of user groups, such as Figure 2 described, the device includes: an extraction unit 21, a determination unit 22, a judgment unit 23, and a restriction unit 24; wherein,

[0109] The extraction unit 21 is used to extract the risk characteristics of all users for any business scenario in the live broadcast platform; the risk characteristics include: the total number of behaviors of the user in the business scenario, the number of IPs used by the user in the business scenario, the number of devices used by the user in the business scenario, and the number of clients used by the user in the business scenario;

[0110] The determination unit 22 is used to construct a classifier based on the risk characteristics, and determine the individual risk score of each user in the business scenario based on the classifier and the risk characteristics of the user; each user corresponds to one individual risk score;

[0111] For each user, obtain the risk characteristics of the user in all business scenarios of the live broadcast platform, splice the risk characteristics to obtain a risk characteristic vector;

[0112] Determine at least one user group based on the risk characteristic vectors of each user;

[0113] For each user group, determine the importance of each user in the user group in the user group; determine the group risk score of the user group in the business scenario based on the importance of each user in the user group and the individual risk score of each user in the business scenario;

[0114] Determine the group divergence degree of the user group in the business scenario; the group divergence degree is the deviation between the group risk score of the user group and the weighted risk scores of each user, and the weighted risk score is determined according to the individual risk score of the user and the importance of the user in the user group;

[0115] Determine the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence degree of the user group in the business scenario;

[0116] The judgment unit 23 is used to judge whether the consensus score exceeds a preset consensus score threshold. If it exceeds, it is determined that the user group is a high-risk user group in the business scenario;

[0117] The restriction unit 24 is used to restrict the behavior of the high-risk user group in the business scenario based on a preset behavior restriction strategy.

[0118] Specifically, the business scenarios of the live streaming platform include multiple ones, such as: login scenario, bullet screen scenario, treasure chest scenario, lottery scenario, and so on. In order to improve the accuracy of risk identification, for any business scenario in the live streaming platform, the extraction unit 21 is used to extract the risk characteristics of all users. Among them, the risk characteristics include: the total number of behaviors of users in the business scenario, the number of IPs used by users in the business scenario, the number of devices used by users in the business scenario, and the number of clients used by users in the business scenario, and so on. Because when an abnormal user group conducts abnormal operations, devices and IPs will be repeatedly used among users, and the behavior of each user is often high-frequency. Therefore, the above several risk characteristics are indispensable information parameters for improving the accuracy of identifying high-risk user groups. They are the traces left by users after use and objectively exist. They are not selected subjectively by people, but are obtained for the above-mentioned total number of behaviors of users in the business scenario, the number of IPs used by users in the business scenario, the number of devices used by users in the business scenario, and the number of clients used by users in the business scenario in order to solve technical problems (that is, selected in line with natural laws).

[0119] After the extraction unit 21 extracts the risk characteristics of all users, the determination unit 22 is used to construct a classifier based on the risk characteristics; specifically including:

[0120] Determine sample data according to the risk characteristics; the sample data includes: positive sample data and negative sample data; the positive sample data includes the risk characteristics corresponding to users with historical risk behaviors, and the negative sample data includes the risk characteristics corresponding to users without historical risk behaviors;

[0121] Train the sample data based on a preset machine learning algorithm to obtain a classifier.

[0122] After the classifier is constructed, based on the classifier and all the risk characteristics of users in this scenario, determine the individual risk scores of each user in the business scenario; each user corresponds to an individual risk score.

[0123] For example, if all the risk characteristics of user u in business scenario s are f s , so the individual risk score r(u, s) of user u in business scenario s can be determined according to the formula r(u, s) = C s (f s ); where C s is the classifier. User u is any user in scenario s.

[0124] The determination unit 23 is used to, for each user, obtain the risk characteristics of the user in all business scenarios of the live streaming platform, splice the risk characteristics to obtain a risk feature vector; all the risk feature vectors constitute a user vector space.

[0125] For example, if the live streaming platform includes 2 scenarios, user u has 10 risk features in the first scenario s and 10 risk features in the second scenario t, then the risk features in the two scenarios are concatenated to form a 20-dimensional risk feature vector F u 。

[0126] After the risk feature vectors of each user are determined, the determination unit 23 needs to determine at least one user group based on the risk feature vectors of each user, specifically including:

[0127] Step a, randomly select a risk feature vector from all the risk feature vectors as the starting center vector C;

[0128] Step b, for any current risk feature vector F in the remaining risk feature vectors u , determine whether the current feature vector satisfies the formula ||C - F u ||2 <= r; if it satisfies, store the user corresponding to the current risk feature vector in the user set; r is a preset radius;

[0129] Step c, for the user set, according to the formula determine the offset vector S of the user set; at this time, user u is the user in the user set;

[0130] Step d, determine the gradient and norm of the offset vector, and determine the first center vector C' according to the formula C' = C grad(S) ||S||2; the gradient of the offset vector is grad(S), and the norm of the offset vector is ||S||2.

[0131] Step e, iterate steps b to d until ||S||2 < ε; ε is an empirical coefficient, 0.01 < ε < 0.1; where, a corresponding user set is generated at the end of each iteration, and the user set is the user group. The preferred value of ε can be 0.5; in order to ensure the accuracy of user clustering, so in this embodiment, the value of ε needs to be greater than 0.01; in order to reduce the iteration error and thus ensure the determination accuracy of the user group, the value of ε needs to be less than 0.1.

[0132] This way of generating each user group through continuous iteration can improve the accuracy of user clustering based on risk features, ensure that the behaviors of users in each user group are consistent, and thus ensure the determination accuracy of the user group.

[0133] After determining at least one user group, for each user group, the determination unit 23 needs to determine the importance of each user in the user group within the user group; based on the importance of each user in the user group within the user group and the individual risk score of each user in the business scenario, determine the group risk score of the user group in the business scenario.

[0134] As an optional embodiment, the determination unit 23 determines the importance of each user in the user group within the user group, including:

[0135] According to the formula Determine the importance w(u1, g) of each user in the user group g within the user group; where c(u1, g) is the activity of user u1 in user group g, Is the sum of the activities of all users in user group g; g is any user group; the activity is the number of behaviors of user u1 in the business scenario; user u1 is any target user in user group g, user v is any user in user group g, and user v includes user u1.

[0136] Here, the activity of user u1 in user group g can be calculated according to the formula Where s is any business scenario; X is all scenarios in the live broadcast platform, and c(u1, g) is the number of behaviors of user u1 in business scenario s.

[0137] The calculation principle of w(u1, g) is: first determine the activity of any user u1 in user group g in the group and the total activity of all users in the group, and then take the ratio of the activity of user u1 in the group and the total activity as the importance of each user in the user group. It can be seen that the higher the activity of each user in the user group, the higher the importance of the user in the user group. After determining the importance of each user in the user group within the user group, the determination unit 23 determines according to the formula Determine the group risk score gr(g, s) of user group g; where s is the business scenario, g is any user group in business scenario s, and u1 is any user in user group g within business scenario s; r(u1, s) is the individual risk score of any user in user group g in business scenario s; w(u1, g) is the importance of any user in user group g within user group g, and |g| is the number of users in user group g.

[0138] Here, the calculation principle of the group risk score gr(g, s) of the user group is: perform a weighted average on the individual risk scores of the users in the user group, so the sum of the weighted risk scores of all users in the user group can be obtained w(u1,g) can be used as a weight, so the group risk score can also be understood as the weighted average score of user risks.

[0139] After the group risk scores of each user group are determined, the determination unit 23 determines the group divergence diff(g, s) of the user group g in the service scenario s according to the formula where w(u1,g)r(u1,s) is the weighted risk score of any user u1 in the user group g; gr(g, s) is the group risk score of the user group g; s is the service scenario, g is any user group in the service scenario s, and u1 is any user in the user group g.

[0140] Here, the calculation principle of the group divergence diff(g, s) is as follows: for any user group, the group divergence is the deviation between the group risk score gr(g, s) of the user group and the weighted risk scores w(u1,g)r(u1,s) of each user. The greater the difference between the two, the greater the divergence between the user risk and the group risk; use (w(u1,g)r(u1,s) - gr(g, s)) 2 to measure the divergence of a single user, and the sum of the divergences of individual users is added up to obtain the final group divergence diff(g, s).

[0141] After the group divergence of the user group is determined, the determination unit 23 determines the consensus score of the user group in the service scenario according to the group risk score of the user group and the group divergence of the user group in the service scenario.

[0142] Specifically, the consensus score F(g, s) of the user group in the service scenario is determined according to the formula F(g, s) = w * gr(g, s) + (1 - w) * (1 - diff(g, s)); where w is the weight coefficient, 0 < w < 1, the weight coefficient; gr(g, s) is the group risk score of the user group g; diff(g, s) is the group divergence of the user group g in the service scenario s.

[0143] The calculation principle of F(g, s) is as follows: determine the weight coefficient of the group divergence and the weight coefficient of the group risk score, and then obtain the consensus score of the user group based on the group divergence, the group risk score of the user group, and the corresponding weight coefficients. Among them, the weight coefficient is determined according to the tolerance degree of the service scenario for the group risk. The higher the tolerance degree, the lower the corresponding group risk score; the lower the tolerance degree, the higher the corresponding group risk score.

[0144] It can be seen that the higher the group risk score and the lower the group divergence, the higher the degree of risk behavior and consistency of the group in the service scenario. At this time, the probability that the user group is identified as a high-risk group in this service scenario is also greater.

[0145] When determining the consensus score of the user group in this embodiment, the group risk score and the group divergence degree are taken into consideration simultaneously, which can ensure the accuracy of identifying high-risk user groups. Because if only the group risk score is considered, if the group risk score is high, it may be caused by some high-scoring users in the user group, and the score differences among users are relatively large. Then there is no batch behavior in this business scenario, and it is not suitable to punish all users in the user group. Therefore, this user group cannot be directly determined as a high-risk group. If only the group divergence degree is considered, if the group divergence degree is small, it is also possible that the group risk score is relatively small. This situation indicates that although the users in the user group have a certain degree of synchronization in behavior, there is no high risk. At this time, it is also not suitable to punish the users in the user group.

[0146] After the consensus score of the user group in the business scenario is determined, the judgment unit 23 is used to judge whether the consensus score exceeds the preset consensus score threshold. If it exceeds, it is determined that the user group is a high-risk user group in the business scenario, and at this time, the users in the user group need to be punished.

[0147] For example, in a certain user group in the login scenario, there are 4 users. The individual risk scores of these 4 users are: 0.75, 0.6, 0.7, 0.8; the importance degrees of these 4 users in the user group are: 0.3, 0.2, 0.3, 0.2.

[0148] Then, the group risk score of the user group is:

[0149] gr(g,s) = 0.75 * 0.3 + 0.6 * 0.2 + 0.7 * 0.3 + 0.8 * 0.2 = 0.715;

[0150] The group divergence degree is:

[0151] diff(g,s) = 0.3 * (0.75 - 0.715) 2 + 0.2 * (0.6 - 0.715) 2 + 0.3 * (0.7 - 0.715) 2 + 0.2 * (0.8 - 0.715) 2 = 0.0045

[0152] Take w = 0.5;

[0153] The consensus score is:

[0154] F(g,s) = 0.5 * 0.715 + 0.5 * (1 - 0.0045) = 0.86

[0155] If the threshold of the consensus score is 0.7, it indicates that this user group is a high-risk user group in the login business scenario, and it is necessary to restrict the behaviors of the users in this user group, such as restricting logins, etc.

[0156] By using the method provided in this embodiment, high-risk user groups and high-risk users can be accurately identified. Compared with the prior art, this embodiment takes into account both the group risk score and the group divergence degree, which can ensure the accuracy of the determined high-risk user groups, and further improve the coverage of high-risk users. Compared with the identification method of the prior art, the coverage of high-risk users can be increased by 10%-15%.

[0157] If it is determined that a user group is a high-risk user group in a certain business scenario, the restriction unit 24 is used to restrict the behaviors of the high-risk user group based on a preset behavior restriction strategy.

[0158] For example, in a lottery scenario, if it is determined that a certain user group is a high-risk user group in the lottery scenario, then the IPs of all users in this user group are temporarily blocked and lottery participation is prohibited.

[0159] The method for restricting the behaviors of user groups provided in this embodiment extracts the risk characteristics of all users for any business scenario in a live streaming platform, and determines the individual risk scores of each user in the business scenario based on the risk characteristics; at least one user group is determined based on the risk characteristics of each user in all business scenarios; in this way, the user group is determined based on the risk characteristics of each user, so the accuracy can be ensured when making risk judgments; to further improve the accuracy of the judgment, for each user group, the group risk score and the group divergence degree of the user group in the business scenario are determined, and the consensus score of the user group in the business scenario is determined based on the group risk score and the group divergence degree. If the consensus score exceeds the threshold, it is determined that this user group is a high-risk group; in this way, it is equivalent to comprehensively considering the batch behaviors and internal divergence degrees of this user group, so user groups with batch behaviors and high synchronization in a certain business scenario can be accurately identified, the identification accuracy is improved, the behaviors of user groups can be accurately restricted, and thus the rights and interests of the live streaming platform are ensured.

[0160] Embodiment III

[0161] This embodiment provides a computer device, as Figure 3 shown, including a memory 310, a processor 320, and a computer program 311 stored on the memory 310 and executable on the processor 320. When the processor 320 executes the computer program 311, the following steps are implemented:

[0162] For any business scenario in the live streaming platform, extract the risk characteristics of all users; the risk characteristics include: the total number of behaviors of the user in the business scenario, the number of IPs used by the user in the business scenario, the number of devices used by the user in the business scenario, and the number of clients used by the user in the business scenario;

[0163] Build a classifier based on the risk characteristics, and based on the classifier and the risk characteristics of the user, determine the individual risk score of each user in the business scenario; each user corresponds to one individual risk score;

[0164] For each user, obtain the risk characteristics of the user in all business scenarios of the live streaming platform, splice the risk characteristics to obtain a risk feature vector;

[0165] Determine at least one user group based on the risk feature vectors of the users;

[0166] For each user group, determine the importance of each user in the user group in the user group; based on the importance of each user in the user group in the user group and the individual risk score of each user in the business scenario, determine the group risk score of the user group in the business scenario;

[0167] Determine the group divergence of the user group in the business scenario; the group divergence is the deviation between the group risk score of the user group and the weighted risk scores of the users, and the weighted risk score is determined according to the individual risk score of the user and the importance of the user in the user group;

[0168] Determine the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence of the user group in the business scenario;

[0169] Judge whether the consensus score exceeds a preset consensus score threshold. If it exceeds, determine that the user group is a high-risk user group in the business scenario;

[0170] In the business scenario, restrict the behaviors of the high-risk user group based on a preset behavior restriction strategy.

[0171] In the specific implementation process, when the processor 320 executes the computer program 311, any implementation manner in the first embodiment can be implemented.

[0172] Since the computer device introduced in this embodiment is the device adopted for implementing the method for restricting the behavior of user groups in Embodiment 1 of the present application, based on the method introduced in Embodiment 1 of the present application, those skilled in the art can understand the specific implementation manner and various variations of the computer device in this embodiment. Therefore, the implementation of how this server realizes the method in the embodiments of the present application will not be described in detail here. As long as the device adopted by those skilled in the art to implement the method in the embodiments of the present application falls within the scope of protection of the present application.

[0173] Based on the same inventive concept, the present application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.

[0174] Embodiment 4

[0175] This embodiment provides a computer-readable storage medium 400, as Figure 4 shown, on which a computer program 411 is stored. When the computer program 411 is executed by a processor, the following steps are implemented:

[0176] For any business scenario in the live streaming platform, extract the risk characteristics of all users; the risk characteristics include: the total number of behaviors of the user in the business scenario, the number of IPs used by the user in the business scenario, the number of devices used by the user in the business scenario, and the number of clients used by the user in the business scenario;

[0177] Based on the risk characteristics, construct a classifier, and based on the classifier and the risk characteristics of the user, determine the individual risk score of each user in the business scenario; each user corresponds to one individual risk score;

[0178] For each user, obtain the risk characteristics of the user in all business scenarios of the live streaming platform, splice the risk characteristics to obtain a risk characteristic vector;

[0179] Based on the risk characteristic vectors of each user, determine at least one user group;

[0180] For each user group, determine the importance of each user in the user group in the user group; based on the importance of each user in the user group in the user group and the individual risk score of each user in the business scenario, determine the group risk score of the user group in the business scenario;

[0181] Determine the group divergence degree of the user group in the business scenario; the group divergence degree is the deviation between the group risk score of the user group and the weighted risk scores of each user, and the weighted risk score is determined according to the individual risk score of the user and the importance of the user in the user group;

[0182] Determine the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence degree of the user group in the business scenario;

[0183] Judge whether the consensus score exceeds a preset consensus score threshold. If it exceeds, determine that the user group is a high-risk user group in the business scenario;

[0184] In the business scenario, restrict the behavior of the high-risk user group based on a preset behavior restriction strategy.

[0185] In the specific implementation process, when the computer program 411 is executed by a processor, any implementation manner in Embodiment 1 can be implemented.

[0186] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0187] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0188] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions in the process Figure 1one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0189] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 the functions specified in one block or multiple blocks.

[0190] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be interpreted to include the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0191] As described above, it is only the preferred embodiment of the present invention, and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for restricting the behavior of a user group, characterized in that, Applied in a live streaming platform, the method includes: For any business scenario in the live streaming platform, extract the risk characteristics of all users; Based on the risk characteristics, construct a classifier, and based on the classifier and the risk characteristics of the users, determine the individual risk scores of each user in the business scenario; each user corresponds to one individual risk score; For each user, obtain the risk characteristics of the user in all business scenarios of the live streaming platform, splice the risk characteristics to obtain a risk feature vector; Determine at least one user group based on the risk feature vectors of each user; For each user group, determine the importance of each user in the user group in the user group; based on the importance of each user in the user group and the individual risk score of each user in the business scenario, determine the group risk score of the user group in the business scenario; Determine the group divergence degree of the user group in the business scenario; the group divergence degree is the deviation between the group risk score of the user group and the weighted risk scores of each user, and the weighted risk score is determined according to the individual risk score of the user and the importance of the user in the user group; Determine the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence degree of the user group in the business scenario; Judge whether the consensus score exceeds a preset consensus score threshold. If it exceeds, determine that the user group is a high-risk user group in the business scenario; In the business scenario, restrict the behavior of the high-risk user group based on a preset behavior restriction strategy; The determining at least one user group based on the risk feature vectors of each user includes: Step a, randomly select one of the risk feature vectors as the starting center vector C from all the risk feature vectors; Step b, for any current risk feature vector F in the remaining risk feature vectors u , determine whether the current risk feature vector satisfies the formula ; if it is satisfied, store the user corresponding to the current risk feature vector in the user set; where r is a preset radius. Step c, for the user set, according to the formula to determine the offset vector S of the user set; where u is a user in the user set U; Step d, determine the gradient and the norm of the offset vector, and determine the first central vector according to the formula ;​ Step e, iterate steps b to d until ; the is an empirical coefficient, 0.01 < < 0.1; wherein, a corresponding user set is generated at the end of each iteration, and the user set corresponds to the user group; The determining the importance of each user in the user group in the user group includes: According to the formula determine the importance of each user in the user group within the user group ; where the is the activity level of user u1 in the user group g, and the activity level is the number of actions of user u1 in the scenario; u1 is any target user in the user group g; v is any user in the user group g The determining the group divergence degree of the user group in the business scenario includes: According to the formula determine the group divergence of the user group in the business scenario ; where, the is the weighted risk score of any user u1 in the user group g; the is the group risk score of the user group g; the s is the business scenario, the g is the user group, and the u1 is any user in the user group g.

2. The method according to claim 1, wherein The constructing a classifier based on the risk characteristics includes: Determine sample data according to the risk characteristics; the sample data includes: positive sample data and negative sample data; the positive sample data includes the risk characteristics corresponding to users with historical risk behaviors, and the negative sample data includes the risk characteristics corresponding to users without historical risk behaviors; Train the sample data based on a preset machine learning algorithm to obtain the classifier.

3. The method according to claim 1, characterized in that, The determining the group risk score of the user group in the business scenario based on the importance of each user in the user group and the individual risk score of each user in the business scenario includes: According to the formula to determine the group risk score of the user group ; where s is the business scenario, g is any user group in the business scenario s, u1 is any user in the user group g within the business scenario s; the is the individual risk score of any user in the user group g in the business scenario s; is the importance of any user in the user group g within the user group g.

4. The method according to claim 1, characterized in that, The determining the consensus score of the user group in the business scenario according to the group risk score of the user group and the group divergence degree of the user group in the business scenario includes: According to the formula determine the consensus score of the user group in the business scenario ; where, the w is the weight coefficient, 0 < w < 1; the is the group risk score of the user group g; the is the group divergence of the user group g in the business scenario s.

5. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1 to 4.

6. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the method described in any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Risk control method and device, medium and apparatus

    CN110599004A

  • Method and device for determining high-risk user

    WO2019196549A1