A Reputation Ranking Method Based on User Rating Patterns and Rating Bias
By calculating the Gini coefficient and range of users and combining the changes in the Gini coefficient and range, a reputation ranking method is established, which solves the problem of identifying fraudulent users in the rating system and achieves efficient and accurate fraudulent user detection, especially with excellent performance in sparse networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-10
- Publication Date
- 2026-04-03
AI Technical Summary
In existing rating systems, fraudulent users, through tactics such as hiring fake accounts and random ratings, severely undermine the fairness and credibility of the rating system, making it difficult for existing methods to effectively distinguish between legitimate and fraudulent users.
By calculating the Gini coefficient and range of users, and combining the changes in the Gini coefficient and range, a reputation ranking method is established to identify the ladder rating behavior of normal users and the non-ladder rating behavior of fraudulent users. The changes in the Gini coefficient and range are used to measure the ladder distribution of user ratings, thus distinguishing between fraudulent users and normal users.
It accurately and quickly identifies fraudulent users, demonstrating significant accuracy and robustness, especially excelling in random bot screening. It shortens the algorithm's convergence time and exhibits high detection efficiency in sparse networks.
Smart Images

Figure CN116029753B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information technology. It utilizes big data technology to perform precise analysis of user rating patterns and relates to a reputation ranking method based on user rating patterns and rating biases. Background Technology
[0002] E-commerce and online shopping have gradually become an indispensable part of people's daily lives. Rating systems play a crucial role in consumer decision-making, as customers often refer to product ratings before purchasing. However, some merchants, for their own profit, often hire large numbers of online trolls to manipulate product ratings. This seriously undermines the fairness and credibility of the rating system. Furthermore, some fraudulent users may randomly rate products without careful consideration, which can also confuse the rating system. These two types of fraudulent users are widespread in rating systems, and consumers may be misled by false reviews and ratings.
[0003] To minimize the negative impact of fake reviews and ratings on rating systems, numerous fraud detection methods have been proposed to identify fraudulent users. The earliest was a correlation-based ranking algorithm (CR), which assigns reputation to each user based on the correlation between user ratings and product ratings. Specifically, a user's reputation is determined by the correlation coefficient between the user's rating vector and the corresponding product rating vector. Later, Gao et al. proposed a group-based ranking algorithm (GR) and an iterative group-based ranking algorithm (IGR). Both methods are based on the assumption that users belonging to larger groups have more reliable reputations, assigning higher reputation values to users belonging to larger groups. Furthermore, Lee et al. found that product ratings almost follow a normal distribution. Therefore, they proposed a bias-based ranking algorithm (DR), where Z-scores are used to assign corresponding reputation scores to users. Wu et al. proposed an iterative balanced ranking algorithm (IBR) to eliminate rating bias in the rating system. Despite the numerous rating ranking algorithms proposed, the underlying mechanisms of fraudulent user formation and the fundamental differences between fraudulent and legitimate users remain unclear. Summary of the Invention
[0004] The purpose of this invention is to provide a reputation ranking method based on user rating patterns and rating biases, which can identify the tiered rating behavior of normal users and the non-tiered rating behavior of fraudulent users by utilizing changes in the Gini coefficient and range. This method can not only identify fraudulent users more accurately and quickly than existing mainstream methods, but it can also be applied to large sparse bipartite rating networks in a short time, even if the network is relatively sparse.
[0005] To achieve the above objectives, the solution of the present invention is:
[0006] A reputation ranking method based on user rating patterns and rating biases, characterized by the following steps:
[0007] Step 1: Establish a triplet data structure for users, products, and ratings. Where i represents the user ID and j represents the product ID. This represents user i's rating of product j;
[0008] Step 2: Calculate the corresponding Gini coefficient for each user based on their historical rating records; the method for calculating the Gini coefficient for each user is as follows:
[0009] Step A21: Count the number of times a user gives a rating, the average score of user ratings, and the sum of the differences in ratings given by users, and calculate the user's Gini coefficient;
[0010] Step A22, using the processing function Amplify the Gini signal of each user;
[0011] Step 3: Calculate the user's range based on the number of times the user has rated the user. ;
[0012] Step 4, set the user's range Combined with the amplified Gini coefficient, it forms an index;
[0013] Step 5: Apply this metric to the bias-based ranking algorithm (DR) and calculate the reputation score for each user; the specific steps for applying this metric to the bias-based ranking algorithm (DR) are as follows:
[0014] Step A51: Assuming each user has a discriminant function, and that most users' ratings of the target product are close to its true quality, calculate the rating distribution for each user. ;
[0015] Step A52: Calculate the average quality of the products based on their ratings. ;
[0016] Step A53, assuming the number of users is large enough, use the average quality of the goods. Substitute for the absolute quality of a product Then update the rating distribution for each user. ;
[0017] Step A54: Update the rating distribution for each user based on the probability distribution of the average product quality and the user rating distribution that the average product quality follows. ;
[0018] Step A55: Calculate the Z-score for each user according to the Central Limit Theorem. ;
[0019] Step A56: Calculate each user's reputation score based on their Z-score and metrics. ;
[0020] Step 6: Sort the users' reputation scores and select the N users with the lowest reputation scores as fraudulent users;
[0021] Step 7: Select Recall and AUC as metrics to compare the accuracy and robustness of the PDDR method with the DR, IGR and IBR methods.
[0022] In step A21, the Gini coefficient for each user is calculated according to the following formula. :
[0023]
[0024] in, This represents the number of ratings given by user i. This represents the average score given by user i. This represents user i's rating of product j.
[0025] In step A22, the processing function is used according to the following formula. Magnified Gini coefficient :
[0026]
[0027] in, This represents the unamplified Gini coefficient.
[0028] In step 3, the user's range is calculated according to the following formula. :
[0029]
[0030] in, This represents the maximum number of ratings given by user i. This represents the minimum number of ratings for user i.
[0031] In step 4, the user's Gini coefficient and range are calculated according to the following formula. Component Indicators :
[0032]
[0033] in, The Gini coefficient representing user i. This represents the range of user i.
[0034] In step A51, the rating distribution of user i is represented by the following formula:
[0035]
[0036] in, Representative products absolute quality Represents a normal distribution. and These represent the mean and standard deviation, respectively. The discriminant function representing user i;
[0037] In step A52, the average mass of the target product is calculated according to the following formula. :
[0038]
[0039] in, The representative gave a score of The number of users, Representing user i's view on the product The rating given;
[0040] In step A53, the user rating distribution is updated according to the following formula. :
[0041]
[0042] in, Representing user i's view on the product The rating given Represents standard deviation;
[0043] In step A54, the user rating distribution is updated according to the following formula. :
[0044]
[0045] Where 's' represents the user's rating of the product. This represents the number of products rated 's' by user i. Represents standard deviation, The probability distribution representing the average quality of the goods;
[0046] In step A55, the Z-score for each user is calculated according to the following formula. :
[0047]
[0048] in, Represents the user's rating score. The standard deviation of user i, The probability distribution representing the average quality of the goods. The number of products for which user i gave a rating of s;
[0049] In step A56, the Z-fraction is used according to the following formula. and indicators Calculate user reputation :
[0050]
[0051] in, This represents the number of products rated 's' by user i. The probability distribution representing the average quality of the goods. The metric representing user i consists of the user's Gini coefficient and range. composition.
[0052] In step 6, the user reputation values obtained by the PDDR method, DR method, IGR method and IBR method are sorted using the bubble sort method, and the top N users of each method are selected as shills.
[0053] In step 7, the Recall value is used to measure the accuracy of the PDDR, DR, IGR, and IBR methods. The Recall value measures the degree to which malicious users are detected in the top-L ranking list, and its form is as follows:
[0054]
[0055] in N represents the number of paid users; d'(L) represents the number of paid users detected in list L; a higher Recall value indicates a higher number of paid users detected in list L, thus indicating higher accuracy of the ranking method; additionally, the AUC value is used to measure robustness to the PDDR, DR, IGR, and IBR methods. The AUC value is the probability that a randomly selected paid user ranks higher than a randomly selected normal user, and its form is as follows:
[0056]
[0057] Where N represents the number of independent comparisons, This represents the number of times that the reputation of the online trolls is relatively low. The AUC value represents the number of times that the reputation of paid users and normal users are the same in the comparison; the higher the AUC value, the stronger the stability of the algorithm, and the higher the robustness of the algorithm.
[0058] This invention proposes a new hypothesis based on user rating models: fraudulent users and legitimate users can be distinguished by changes in a combination of the Gini coefficient and the range. This combination can reflect the rating behavior of ordinary users with personal preferences or the ratings of fraudulent users with inappropriate purposes. Specifically, we combine the Gini coefficient and the range to model the indicators based on users' historical rating records.
[0059] The principle of this invention is to use the changes in the Gini coefficient and range to measure the step distribution of user ratings. If a user rating conforms to the step distribution, it is considered a normal user; otherwise, it is considered a fraudulent user. The reputation ranking method (PDDR) based on user rating patterns and rating bias described in this invention was tested on three real datasets (MoiveLens, Netflix, and MoiveLens_100) and compared with other existing methods, such as bias-based ranking algorithms (DR), iterative group-based ranking algorithms (IGR), and iterative balance ranking algorithms (IBR). Experimental results show that the method of this invention (PDDR) can not only identify fraudulent users more accurately and quickly than existing mainstream methods, but also can be applied to large sparse bipartite rating networks in a short time, even with relatively sparse networks.
[0060] The beneficial effects of this invention are:
[0061] (1) This invention can accurately and quickly identify fraudulent users. It further distinguishes between fraudulent users and normal users by the changes in the Gini coefficient and range. It has significant accuracy and robustness, especially for screening random online trolls.
[0062] (2) This invention has higher detection efficiency in the early stage of the algorithm, which greatly shortens the convergence time of the algorithm. (3) This invention has strong versatility and strong application prospects. In the future, it will be combined with graph neural networks for further in-depth research.
[0063] This invention not only has low time and space complexity, strong versatility, and easy expansion, but can also be applied to e-commerce platform recommendation systems, online "water army" identification, and water army detection. Attached Figure Description
[0064] Figure 1 This is a flowchart of the present invention.
[0065] Figure 2 This is a flowchart of the PDDR algorithm.
[0066] Figure 3 The present invention presents the recall rate R(L) curve as a function of L under different datasets.
[0067] Figure 4This invention measures the recall rate of different numbers of online trolls on different datasets.
[0068] Figure 5 The present invention presents the AUC (Average Recall) curve as a function of p values under different datasets.
[0069] Figure 3 In each subplot, the horizontal axis L represents the number of users selected from the top-L list, with L ranging from 0 to 250. The vertical axis R(L) represents the proportion of detected fraudulent users, with R(L) ranging from 0 to 1. Subplots (a), (b), and (c) represent the recall rate of detecting malicious bots. Subplots (d), (e), and (f) represent the recall rate of detecting random bots. The number of bots is 50. The results of each experiment will be repeated 100 times, and the experiments for each method are independent of each other.
[0070] Figure 4 In each subplot, the horizontal axis p represents the proportion of fake users to the total number of users, with a value ranging from 0 to 0.25. The vertical axis R(L) represents the proportion of detected fraudulent users, with a value ranging from 0 to 1. Subplots (a), (b), and (c) represent the recall values for detecting malicious fake users. Subplots (d), (e), and (f) represent the recall values for detecting random fake users. In this experiment, the number of fake users varied as the experiment progressed. The results of each experiment were repeated 100 times, and the experiments for each method were independent of each other.
[0071] Figure 5 In each subplot, the horizontal coordinate p is calculated using the formula: p = d / m, with a value ranging from 0 to 0.25. The vertical coordinate AUC represents the probability that a randomly selected bot user ranks higher than a randomly selected normal user, with an AUC value ranging from 0 to 1. Subplots (a), (b), and (c) represent the AUC values used to detect malicious bots. Subplots (d), (e), and (f) represent the AUC values used to detect random bots. The results of each experiment will be repeated 100 times, with each experiment being independent of the others. Detailed Implementation
[0072] The technical solution and beneficial effects of the present invention will be described in detail below with reference to the accompanying drawings.
[0073] like Figure 1 As shown, a reputation ranking method based on user rating patterns and rating bias is described below:
[0074] (1) Establish a triplet data structure for users, products, and ratings. Where i represents the user ID and j represents the product ID. The rating of user i for product j is represented by a triplet. This method of constructing a dataset can effectively solve problems such as insufficient computer memory and data sparsity.
[0075] (2) Based on each user’s historical rating records, calculate the sum of the number of ratings, average rating score and rating difference given by each user, and calculate the corresponding Gini coefficient.
[0076] (3) Use processing functions Amplify the Gini coefficient for each user to enhance its ability to identify the stepped distribution of user ratings.
[0077] (4) Count the maximum and minimum number of ratings given by users, and calculate the range of users to distinguish the rating distribution and rating bias of each user;
[0078] (5) Combine the user's range with the amplified Gini coefficient into a single index;
[0079] (6) Apply the indicators to the user reputation formula in the bias-based ranking algorithm (DR) and recalculate the user reputation value of the DR method after the indicators are applied;
[0080] (7) Calculate the user reputation scores using the PDDR, DR, IGR, and IBR methods respectively, and perform bubble sort to sort the user reputation scores from smallest to largest. After sorting, select the N users with the lowest reputation scores as fraudulent users.
[0081] (10) Recall and AUC were selected as evaluation metrics. Recall measures the degree to which malicious users are detected in the top-L ranking list, and AUC represents the probability that randomly selected bot users rank higher than randomly selected normal users. The Recall and AUC values of the PDDR, DR, IGR, and IBR methods were calculated, and the identification efficiency of the PDDR, DR, IGR, and IBR methods for bots was compared.
[0082] Specifically, this embodiment mainly includes the following steps, which will be described in conjunction with the accompanying drawings for ease of understanding:
[0083] Step 1: Construct triples for data storage. Since user ratings for products are too sparse, storing them in a matrix manner would lead to wasted space. Therefore, a triplet data structure was constructed based on user, product, and rating information. It uses a smaller space for data storage. The rating given by user i to product j is .exist Figure 2 , Figure 3 and Figure 4 In order to facilitate a detailed explanation of the method's flow, information is stored in a matrix format.
[0084] Step 2: Calculate the sum of the number of ratings, average rating score, and rating difference for each user, and then calculate the Gini coefficient for each user using the following formula. The calculation method is shown in formula (1):
[0085] Formula (1)
[0086] In formula (1), This represents the number of ratings given by user i. This represents the average score given by user i. This represents user i's rating of product j.
[0087] Step 3: Since the Gini signal generated in Step 2 is relatively weak, a processing function is used. Amplified Gini signal, processing function Inspired by the Sigmoid function. The magnified Gini coefficients are labeled as follows. The calculation method is shown in formula (2):
[0088] Formula (2)
[0089] In formula (2), This represents the unamplified Gini coefficient.
[0090] Step 4: Count the maximum and minimum number of ratings given by users, and calculate the user's range. ;
[0091] Formula (3)
[0092] In formula (3), This represents the maximum number of ratings given by user i. This represents the minimum number of ratings for user i.
[0093] Step 5: Magnify the Gini coefficients Extreme difference with users Combined into an indicator ,index It is used to measure whether a user rating pattern conforms to a hierarchical rating distribution. If the indicator... The smaller the value, the more likely the user's rating pattern follows a stepped distribution, and the more likely they are a normal user. Conversely, the higher the value, the more likely they are a paid user. (Indicator) The calculation method is shown in formula (4):
[0094] Formula (4)
[0095] In formula (4), The amplified Gini coefficient representing user i. This represents the range of user i.
[0096] Step 6: Set the indicators User reputation applied to bias-based ranking (DR) algorithms In the process, the user reputation value of the DR method is recalculated after the application of the indicators. ;
[0097] First, bias-based ranking (DR) algorithms are based on the assumption that each user has a unique discriminant function, and that most users' ratings of the target product are close to its true quality. Therefore, the distribution of user ratings... As shown in formula (5):
[0098] Formula (5)
[0099] In formula (5), Representative products absolute quality Represents a normal distribution. and These represent the mean and standard deviation, respectively. The discriminant function representing user i.
[0100] Secondly, based on the ratings received for each product, the average quality of the target product is calculated. As shown in formula (6):
[0101] Formula (6)
[0102] In formula (6), The representative gave a score of The number of users, Representing user i's view on the product The rating given.
[0103] Secondly, when the number of users is large enough, the average quality of the goods... Approximate absolute quality of the product Then formula (5) can be rewritten as formula (7):
[0104] Formula (7)
[0105] In formula (7), Representing user i's view on the product The rating given.
[0106] Next, based on the probability distribution of the average quality of the goods. With the average quality of the goods If all follow the user rating distribution, then formula (7) can be rewritten as formula (8):
[0107] Formula (8)
[0108] In formula (8), s represents the rating given by the user to the product. This represents the number of products rated 's' by user i.
[0109] Then, according to the central limit theorem in statistics, calculate the Z-score for each user, denoted as . As shown in formula (9):
[0110] Formula (9)
[0111] In formula (9), Represents the user's rating score. The standard deviation of user i, The probability distribution representing the average quality of the goods. This represents the number of items that user i rated s.
[0112] Finally, based on each user's Z-score and metrics Seeking to gain the trust of each user As shown in formula (10):
[0113] Formula (10)
[0114] In formula (10), This represents the number of products rated 's' by user i. The probability distribution representing the average quality of the goods. The metric representing user i consists of the user's Gini coefficient and range. Composition. Application of indicators in the DR method, such as Figure 2 As shown.
[0115] Step 7: Use the Recall value to measure the accuracy of the PDDR, DR, IGR, and IBR methods. The Recall value measures the degree to which malicious users are detected in the top-L ranking list, and its form is shown in formula (21):
[0116] Formula (11)
[0117] In formula (11), The value of N represents the number of paid users. d'(L) represents the number of paid users detected in list L. A higher Recall value indicates a higher number of paid users detected in list L, thus indicating higher accuracy of the ranking method. Additionally, we use the AUC value to measure the robustness of the PDDR, DR, IGR, and IBR methods. The AUC value is the probability that a randomly selected paid user ranks higher than a randomly selected normal user, and its form is shown in formula (12):
[0118] Formula (12)
[0119] Where N represents the number of independent comparisons, This represents the number of times that the reputation of the online trolls is relatively low. The AUC value represents the number of times that the reputation of paid users and normal users are the same in the comparison. The higher the AUC value, the stronger the stability of the algorithm, and thus the higher its robustness.
[0120] The efficiency of the PDDR, DR, IGR, and IBR methods in identifying fraudulent users was verified using three real-world datasets (MovieLens, Netflix, and MovieLens_100). The MovieLens and MovieLens_100 datasets are from the GroupLens project at the University of Minnesota. The Netflix dataset (http: / / pan.baidu.com / s / 1dDtmbW9) was provided by a recommendation system competition launched by a DVD rental company. The Netflix and MovieLens datasets use a 5-point scale, with 1 representing the worst and 5 representing the best. The MovieLens_100 dataset uses a 10-point scale, with 1 representing the worst and 10 representing the best. Detailed information about the datasets is shown in Table 1. Table 1: Statistical information for the MovieLens, Netflix, and MovieLens_100 datasets. Where m represents the number of users, n represents the number of items, ... For the sparsity of bipartite graphs, where The number of user ratings.
[0121] Table 1
[0122] Data set m n l S MovieLens 943 1682 60 0.0630 Netfix 3000 2779 71 0.0237 MovieLens_100 7120 130642 1048575 0.00113
[0123] The recall rate curves of the method of this invention as L increases in the three datasets in Table 1 are shown below. Figure 3As shown in the figures, experimental results demonstrate that, as seen in subgraphs (a), (b), and (c), the PDDR and IBR methods significantly outperform other methods in detecting malicious users. Subgraphs (d), (e), and (f) show that the PDDR method is significantly more accurate than the DR, IBR, and IGR methods in detecting random users. Particularly in subgraph (e), even with 3 million scores, our PDDR method still achieves higher accuracy than DR, IGR, and IBR, exceeding DR by approximately 50.27%.
[0124] The recall rate curves of the method of this invention as the fraud ratio p increases in the three datasets in Table 1 are shown below. Figure 4 As shown in the figure. Experimental results demonstrate that when the number of fraudulent users is small, the proposed PDDR method has a significant advantage in detecting fraudulent users, especially in detecting random users. Furthermore, as the number of fraudulent users increases, both the PDDR and IGR methods proposed in this invention consistently maintain high accuracy. In particular, when p = 0.03, the recall values of PDDR and IGR exceed 0.8.
[0125] The AUC curves of the method of this invention in the three datasets shown in Table 1 are as follows: Figure 5 As shown in the figure. Experimental results demonstrate that the PDDR method proposed in this invention exhibits significant robustness in detecting fraudulent users, particularly in detecting random users. Furthermore, the recall value of the PDDR method proposed in this invention remains consistently around 0.98 or higher, while the mean recall values of the DR, IBR, and IGR methods are around 0.96 or higher.
[0126] The above analysis reveals the following advantages of this invention: (1) The test results from the three datasets show that this method can effectively identify the rating ladder distribution of normal users and fraudulent users, and has good accuracy and robustness, especially in the detection of random users. (2) This invention has higher detection efficiency in the early stages of the algorithm, greatly shortening the convergence time. (3) This invention has strong versatility and promising application prospects, and will be further studied in the future in combination with graph neural networks.
[0127] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention.
Claims
1. A reputation ranking method based on user rating patterns and rating bias, characterized in that... Includes the following steps: Step 1: Establish a triplet data structure for users, products, and ratings. Where i represents the user ID and j represents the product ID. This represents user i's rating of product j; Step 2: Calculate the corresponding Gini coefficient for each user based on their historical rating records; the method for calculating the Gini coefficient for each user is as follows: Step A21: Count the number of times a user gives a rating, the average score of user ratings, and the sum of the differences in ratings given by users, and calculate the user's Gini coefficient; Step A22, using the processing function Amplify the Gini signal of each user; Step 3: Calculate the user's range based on the number of times the user has rated the user. ; Step 4, set the user's range Combined with the amplified Gini coefficient, it forms an index; Step 5: Apply this indicator to the bias-based ranking algorithm and calculate the reputation score for each user; the specific steps for applying this indicator to the bias-based ranking algorithm are as follows: Step A51: Assuming each user has a discriminant function, and that most users' ratings of the target product are close to its true quality, calculate the rating distribution for each user. ; Step A52: Calculate the average quality of the products based on their ratings. ; Step A53, assuming the number of users is large enough, use the average quality of the goods. Substitute for the absolute quality of a product Then update the rating distribution for each user. ; Step A54: Update the rating distribution for each user based on the probability distribution of the average product quality and the user rating distribution that the average product quality follows. ; Step A55: Calculate the Z-score for each user according to the Central Limit Theorem. ; Step A56: Calculate each user's reputation score based on their Z-score and metrics. ; Step 6: Sort the users' reputation scores and select the N users with the lowest reputation scores as fraudulent users; Step 7: Select Recall and AUC as metrics to compare the accuracy and robustness of the PDDR method with the DR, IGR and IBR methods.
2. The reputation ranking method based on user rating patterns and rating bias as described in claim 1, characterized in that: In step A21, the Gini coefficient for each user is calculated according to the following formula. : ; in, This represents the number of ratings given by user i. This represents the average score given by user i. This represents user i's rating of product j.
3. The reputation ranking method based on user rating patterns and rating bias as described in claim 1, characterized in that: In step A22, the processing function is used according to the following formula. Magnified Gini coefficient : ; in, This represents the unamplified Gini coefficient.
4. The reputation ranking method based on user rating patterns and rating bias as described in claim 1, characterized in that: In step 3, the user's range is calculated according to the following formula. : ; in, This represents the maximum number of ratings given by user i. This represents the minimum number of ratings for user i.
5. The reputation ranking method based on user rating patterns and rating bias as described in claim 1, characterized in that: In step 4, the user's Gini coefficient and range are calculated according to the following formula. Component Indicators : ; in, The Gini coefficient representing user i. This represents the range of user i.
6. The reputation ranking method based on user rating patterns and rating bias as described in claim 1, characterized in that: In step A51, the rating distribution of user i is represented by the following formula: ; in, Representative products absolute quality Represents a normal distribution. and These represent the mean and standard deviation, respectively. The discriminant function representing user i; In step A52, the average mass of the target product is calculated according to the following formula. : ; in, The representative gave a score of The number of users, Representing user i's view on the product The rating given; In step A53, the user rating distribution is updated according to the following formula. : ; in, Representing user i's view on the product The rating given Represents standard deviation; In step A54, the user rating distribution is updated according to the following formula. : ; Where 's' represents the user's rating of the product. This represents the number of products rated 's' by user i. Represents standard deviation, The probability distribution representing the average quality of the goods; In step A55, the Z-score for each user is calculated according to the following formula. : ; in, Represents the user's rating score. The standard deviation of user i, The probability distribution representing the average quality of the goods. The number of products for which user i gave a rating of s; In step A56, the Z-fraction is used according to the following formula. and indicators Calculate user reputation : ; in, This represents the number of products rated 's' by user i. The probability distribution representing the average quality of the goods. The metric representing user i consists of the user's Gini coefficient and range. composition.
7. The reputation ranking method based on user rating patterns and rating bias as described in claim 1, characterized in that: In step 6, the user reputation values obtained by the PDDR method, DR method, IGR method and IBR method are sorted using the bubble sort method, and the top N users of each method are selected as shills.
8. The reputation ranking method based on user rating patterns and rating bias as described in claim 1, characterized in that: In step 7, the Recall value is used to measure the accuracy of the PDDR, DR, IGR, and IBR methods. The Recall value measures the degree to which malicious users are detected in the top-L ranking list, and its form is as follows: ; in This represents the number of online trolls, which is the value of N. d'(L) represents the number of fake users detected in list L; a higher Recall value indicates a higher number of fake users detected in list L, thus indicating higher accuracy of the ranking method; additionally, the AUC value is used to measure robustness to the PDDR, DR, IGR, and IBR methods. The AUC value is the probability that a randomly selected fake user ranks higher than a randomly selected normal user, and its form is as follows: ; Where N represents the number of independent comparisons, This represents the number of times that the reputation of the online trolls is relatively low. This represents the number of times that paid users and normal users have the same reputation in a comparison. A higher AUC value indicates stronger algorithm stability, which in turn indicates higher algorithm robustness.
Citation Information
Patent Citations
E-commerce navy identification method based on range
CN111275526A
E-commerce water army identification method based on interval segmentation
CN113674045A