A robust score aggregation method based on secondary review

The robust scoring aggregation method based on secondary review solves the problems of unfair scoring, sparsity of scoring information, and heterogeneity of groups in scoring aggregation, and achieves fairness and robustness in cross-group comparisons. It is applicable to various scoring scenarios such as academic conferences and project reviews.

CN122134209APending Publication Date: 2026-06-02BEIJING NORMAL UNIV AT ZHUHAI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING NORMAL UNIV AT ZHUHAI
Filing Date
2026-05-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies suffer from problems such as unfair scoring having a significant impact, sparse scoring information, and heterogeneous grouping in score aggregation, leading to unstable and unfair aggregation results.

Method used

A robust scoring aggregation method based on secondary review is adopted. This method involves parallel group scoring modeling, initial scoring aggregation and standardization within groups, construction of secondary review tasks and scaling adjustments, and finally output of scoring results that can be compared across groups.

Benefits of technology

It effectively eliminates group heterogeneity, improves cross-group comparability, reduces the impact of unfair scoring, enhances the robustness and fairness of scoring results, and is highly efficient and applicable to various scoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134209A_ABST
    Figure CN122134209A_ABST
Patent Text Reader

Abstract

This invention discloses a robust scoring aggregation method based on secondary review, belonging to the field of scoring decision technology, including the following steps: parallel group scoring modeling; initial aggregation and standardization within groups; secondary review task construction; secondary aggregation and scale readjustment; and final score output. This invention employs the aforementioned robust scoring aggregation method based on secondary review to address the inter-group heterogeneity problem caused by inconsistencies in scoring standards and variances among different groups in parallel group scoring activities. It proposes an aggregation method that can automatically readjust the scoring scales of each group, enabling fair comparison between objects from different groups. It improves the robustness of the aggregation results when unfair scoring exists, such as malicious or arbitrary scoring, reducing the impact of abnormal scoring on the results. While ensuring the overall efficiency of the scoring task, it establishes connections between groups by designing a small number of secondary review tasks, achieving objective correction of the initial intra-group aggregation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of scoring decision technology, and in particular to a robust scoring aggregation method based on secondary review. Background Technology

[0002] In various online and offline review and recommendation scenarios, rating is one of the most common means of expressing preferences. Activities such as movie ratings, product reviews, project evaluations, and competition scoring all require aggregating scores from multiple raters into a comprehensive score for ranking, filtering, or decision-making. Traditionally, this approach typically involves simply calculating the mean or median of all ratings, assuming all raters are equally trustworthy. While simple to implement, this method suffers from the following problems: (1) The impact of unfair ratings is significant: some raters may give extreme scores due to carelessness, bias or malice, and simple averaging is easily “biased” by these abnormal ratings.

[0003] (2) Sparse rating information: When the number of objects is large and the rater can only rate a small portion of them, the number of ratings a single object has is limited, and the aggregation result is unstable.

[0004] (3) Parallel Group Heterogeneity Problem: In large-scale scoring tasks, objects are often divided into multiple groups, and different groups of scorers review them in parallel to improve efficiency. For example, academic conferences, review competitions, and project bidding all use group review. There are often differences in scoring standards and score dispersion between different groups, i.e., "group heterogeneity". The result is that if a group as a whole "scores strictly", the average score of excellent objects may be lower than that of ordinary objects in another group that "scores leniently", thus leading to unfair cross-group comparisons.

[0005] Existing robust rating aggregation methods assign different weights to raters based on their overall behavior to mitigate the impact of malicious ratings. However, most of these methods assume all objects are on the same scale, failing to effectively handle mean and variance differences arising from parallel grouping. In this situation, simply performing robust aggregation within each group cannot solve the cross-group comparability problem; a new method that can automatically calibrate the scale between groups is needed. Summary of the Invention

[0006] The purpose of this invention is to provide a robust scoring aggregation method based on secondary review, thereby solving the problems mentioned in the background art.

[0007] To achieve the above objectives, this invention provides a robust scoring aggregation method based on secondary review, comprising the following steps: S1. Parallel group scoring modeling: By defining the set of raters and objects and corresponding parameters, the set is divided into non-overlapping groups that meet the total quantity constraint. Within each group, the raters' scores for the objects are limited to a preset range. S2. Initial Rating Aggregation and Standardization within Groups: Robustly aggregate the ratings of objects within each group to obtain the initial aggregated ratings; calculate the statistics for each group and complete the Z-score standardization; summarize the variances of each group to obtain the global mean squared error estimate. S3. Construction of Secondary Review Task: To achieve cross-group comparability of scores from different groups, a secondary review process of appropriate scale is set up. The secondary review task is constructed by selecting the secondary review subjects and scorers. S4. Secondary Aggregation and Scale Rebalancing: For the subjects subject to secondary review, calculate the secondary aggregated score and mean, and combine the global variance to achieve scale rebalancing of the score, so as to obtain the final aggregated score that can be compared across groups. S5. Final Score Output: After mapping the final aggregated score to the score range required for actual application, the final score result is output.

[0008] Preferably, the specific steps of S1 are as follows: S11. Define the basic set and parameters: Clarify the set of raters and the set of objects to be rated; S12. Divide all objects and raters into several non-overlapping groups, configure the corresponding rater set and object set for each group, and satisfy the quantity constraint conditions. S13. Establish the rating relationship between the rater and the subject within each group, set the range of rating values, and allow for some missing ratings.

[0009] Preferably, the specific steps of S2 are as follows: S21. Initial aggregation: For each object in each group, the scores of the raters in the group are aggregated using the average or robust aggregation method to obtain the initial aggregated score, which suppresses the influence of unfair scoring. S22. Within-group statistics calculation: Calculate the mean and standard deviation of the initial aggregated scores of all objects in each group, which are used to quantitatively characterize the overall level and dispersion of the group's scores. S23, Z-score standardization: Based on the initial aggregated score, mean and standard deviation, calculate the standardized Z-score for each group and each object to eliminate inter-group scale differences; S24. Global Variance Statistics: Summarize the variances of all groups, calculate the global mean squared error estimate, and efficiently characterize the overall dispersion level of the overall scoring system.

[0010] Preferably, the specific steps of S3 are as follows: S31. Secondary Object Selection: Based on each group of object sets, select a secondary review set and construct a global secondary review object set; S32. Secondary rating selection: Select a portion of the ratings from all ratings or from each group of ratings to form a secondary rating set, and control the selection scale by proportion. S33. Secondary scoring task assignment: Assign cross-group objects to secondary scorers to achieve cross-scoring; and reconstruct the expected quality based on the relationship between the scorer and the object's group, using either the original scoring error model or the inter-group standard deviation.

[0011] Preferably, the specific steps of S31 are as follows: S311. From the object set of each group, select objects for secondary review randomly or according to a preset ratio to form a secondary review object set. S312. For the selected secondary review objects, take the union of the sets to form a global secondary review object set, and determine the size of the global secondary review object set.

[0012] Preferably, the specific steps of S33 are as follows: S331. Assign objects from different groups to raters in the set to form cross-group ratings; S332. If the rater and the subject are from the same original group, their original rating error model is used directly; if they are from different groups, the expected quality is regenerated based on the standard deviation parameters of the target group and the source group to characterize their rating habits.

[0013] Preferably, the specific steps of S4 are as follows: S41. Secondary aggregate score calculation: For each secondary review object, collect the corresponding secondary scores of the secondary reviewers, and use average or robust aggregation to obtain its secondary aggregate score. S42. Secondary group mean estimation: Calculate the secondary aggregated score mean of each group of secondary review subjects as the overall score level after cross-group score correction. S43. Scale readjustment: Based on global variance estimation and the mean of the secondary aggregated scores, the standardized scores are restored to a uniform scale to obtain the final aggregated scores.

[0014] Preferably, the specific steps of S43 are as follows: S431. Using a unified global variance estimate and group-level corrected mean, the standardized scores are restored to a unified scale to obtain the final aggregated scores of the objects. S432. Linearly map the final aggregated score to the target score interval to meet application requirements.

[0015] Therefore, the robust scoring aggregation method based on secondary review described above, as used in this invention, has the following beneficial effects: (1) Eliminate group heterogeneity and ensure cross-group comparability: This invention establishes cross-links between different groups through secondary review, and readjusts the scoring scale of each group by using the global variance and group-level corrected mean with better statistical significance, so that objects in different groups can be compared under a unified scale, avoiding the phenomenon of "strict groups being disadvantaged and lenient groups being advantageous".

[0016] (2) Robustness to unfair scoring: This invention can use existing robust aggregation methods in both the initial and secondary aggregation stages. By assigning weights to raters or individual scores, the impact of malicious or arbitrary scoring is reduced. Simulation experiments show that, in the presence of a large proportion of unfair scoring, the correlation between the ranking obtained by this method and the true quality of the object is significantly higher than that of the direct aggregation method.

[0017] (3) Efficiency and cost controllable: The secondary review of this invention only samples some objects and raters. The sampling ratio can be flexibly set according to resource and accuracy requirements. It can improve overall fairness without repeating the review of all objects.

[0018] (4) Solid theoretical foundation: This invention proves through unbiased estimators and efficiency analysis that the group-level mean and variance estimation used in this method is statistically more effective, providing theoretical support for engineering implementation.

[0019] (5) Wide range of applicable scenarios: This invention is applicable to all scenarios where there is a need for group review and scoring aggregation, such as academic conference paper review, project establishment review, competition scoring, and online platform product / content evaluation.

[0020] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0021] Figure 1 This is a schematic flowchart of an embodiment of a robust scoring aggregation method based on secondary review according to the present invention. Detailed Implementation

[0022] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0023] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0024] Example Please see Figure 1 This invention provides a robust scoring aggregation method based on secondary review, comprising the following steps: S1, Parallel group scoring modeling.

[0025] S11. Let the set of raters be... U The object collection is V The number of raters is N =| U |, the number of objects is M =| V |

[0026] S12, Divide all objects into K The first non-overlapping group: k The group's object collection is V k The set of raters is U k . No. k Group contains p k =| U k | Individual ratings and q k =| V k | objects that satisfy: .

[0027] S13, No. k In the group, the first i Ratings For the first j objects The given score is recorded as The ratings are limited to a given range (e.g., [0,5] or [0,100]). A sparse rating matrix is ​​permissible, meaning that some ratings may be sparsity. Missing.

[0028] S2, Initial scoring aggregation and standardization within the group.

[0029] S21, Initial Assembly.

[0030] For each group k Each object in The initial aggregated score is obtained by aggregating the scores from all raters within the group. The aggregation function can be a simple average, or it can employ existing robust aggregation methods (such as the correlation-based CR method, the Z-score-based ZR method, the Beta distribution-based BR method, etc.) to mitigate the impact of unfair scoring.

[0031] S22. Calculation of within-group statistics.

[0032] For the first k Group, calculate the mean of the initial aggregated scores of all objects in the group. with mean squared deviation These characters respectively characterize the overall scoring level and dispersion of the group.

[0033] S23, Z-value standardization.

[0034] For the first k Each object in the group Calculate its standardized score (Z-score): ; The Z-score can eliminate differences in the mean and variance of ratings within each group, mapping the quality of an object to a uniform relative scale.

[0035] S24, Global Variance Estimation.

[0036] variance of all groups Summarize the results and calculate the global mean squared error estimator. This estimator reflects the overall variance level of the entire system. It is statistically more meaningful than that of a single group. More effective.

[0037] S3, Secondary Review Task Construction.

[0038] To establish comparability between different groups, this invention designs a moderately sized secondary review stage, with the following specific steps: S31, Secondary object selection.

[0039] S311, from each group k collection of objects V k In proportion (e.g., 10%-70%) Selected randomly or by importance These objects constitute the set of objects for secondary review. .

[0040] S312. All selected objects form a global secondary review object set. Its scale is .

[0041] S32, Secondary rating selection.

[0042] Select a number of raters from the entire set of raters or from each group of raters to form a secondary set of raters. It can be achieved through proportion θControl the selection size (e.g., select a number of experienced raters for each group).

[0043] S33, Secondary scoring task allocation.

[0044] S331, is a set The raters are assigned to subjects from different groups, forming a cross-group "cross-rating".

[0045] S332. If the rater and the subject are from the same original group, their original rating error model can be used directly; if they are from different groups, the expected quality can be regenerated based on the standard deviation parameters of the target group and the source group to characterize their scoring habits.

[0046] It is worth mentioning that in the secondary review task construction stage of this invention, a certain number of objects and raters are currently randomly or proportionally selected from each group for secondary review. In practical applications, the sampling strategy for secondary review can be adjusted according to the specific requirements of the scoring task. For example, raters with high credibility or scoring experience can be selected, or sampling can be based on the correlation between raters and objects, thereby improving the efficiency and accuracy of secondary review and reducing potential biases in the review process.

[0047] S4, Secondary scoring aggregation and group-level scale readjustment.

[0048] S41, Secondary aggregation score calculation.

[0049] For each object undergoing secondary review Collect all selected raters Second rating The secondary aggregation score of the object is obtained through average or robust aggregation. .

[0050] S42, Estimation of the mean of the second group.

[0051] For the first k For all individuals within the group who underwent a second review, calculate the average of their aggregated second review scores: ; This mean reflects the percentage after introducing cross-group scoring. k The "corrected" overall level of the group objects is used to replace the initial aggregation phase. .

[0052] S43. Rescaling.

[0053] S431. Using a unified global variance estimate and group-level adjusted mean Standardized scoring By restoring the scale to a uniform level, the final aggregate score of the object is obtained: .

[0054] S432. To meet application requirements, it is possible to further... Linearly mapped to the target score interval (e.g., [0,5]).

[0055] After treating object quality as a population random variable, the traditional and and the improved methods and These are all unbiased estimators of the population expectation and variance. Based on this, we prove using the central limit theorem and analysis of variance that... By constructing a sampling method using grouping and secondary sampling, and meeting certain sampling ratio conditions (such as...), It has a smaller variance when it is ); at the same time, By using optimal weighted combination of variance estimates for each group, variance minimization was achieved, thus resulting in a more straightforward approach. It has higher estimation efficiency. Therefore, compared to direct score aggregation methods, it is more efficient. and The evaluation results are more stable and accurate under appropriate object sampling ratios, and have better performance than... and s k Higher estimation efficiency, therefore, the final It effectively approximates the true quality of an object in a statistical sense.

[0056] It is important to note that, in the initial and secondary aggregation stages within a group, in addition to the averaging method, various existing or future robust scoring aggregation algorithms can be used, such as correlation-based iterative weighting methods, probabilistic aggregation methods based on Beta or Dirichlet distributions, to further improve robustness against unfair scoring.

[0057] To further optimize the scoring aggregation process, a multi-round secondary review mechanism can be considered. Under this mechanism, the weights of the raters and the scoring scales are gradually adjusted based on the results of each round of review, ultimately achieving more accurate scoring aggregation. Furthermore, an adaptive mechanism can be employed to automatically adjust sampling ratios, rater selection strategies, and other tactics based on feedback from each round of review, thereby improving the adaptability of the entire review process and ensuring efficiency and robustness across different scoring scenarios.

[0058] The scoring aggregation method of this invention is currently mainly applicable to scoring tasks with scoring ranges within a fixed interval (such as scores of 0-5 or 0-100). However, in practical applications, the types of scoring tasks and scoring scales may be more diverse. Therefore, the aggregation method of this invention can be further extended to adapt to different scoring scales (such as scoring tasks with negative scores or a wider range) and task types (such as a combination of qualitative and quantitative scoring). Through adaptive extension to scoring scales and task types, the robust scoring aggregation method of this invention can be effective in more practical application scenarios.

[0059] Example 1 The implementation of the method of the present invention will be described below with reference to the embodiments.

[0060] S1, Basic Settings.

[0061] S11. In this embodiment, there are 15 raters and 15 subjects to be evaluated, and the score is [0, 50].

[0062] S12. Divide the objects evenly into 3 groups, with 5 objects in each group; assign 5 raters to each group, and the raters in each group do not overlap.

[0063] S2, Parallel group scoring and initial aggregation.

[0064] S21. Each group of raters only evaluates 5 objects in their own group, resulting in 3 5×5 rating matrices.

[0065] S22. For each group, calculate the mean score as the initial aggregation result. At this point, it can be seen that the overall score of Group 1 is relatively high and varies greatly, while the overall score of Group 3 is relatively low and fluctuates greatly, indicating that there are significant differences in the standard deviation and variance between groups.

[0066] S3, Intra-group standardization.

[0067] S31. Calculate the mean and variance of the initial aggregated scores for each group, and obtain... , , as well as s 1 , s 2 , s 3 .

[0068] S32. Further calculate the standardized score for each object. And obtain the global variance estimate. .

[0069] S4. Construct a secondary review.

[0070] S41. Randomly select one object from each group to participate in the second review, for a total of 3 objects.

[0071] S42. Select two raters from each group to form a secondary rater set. These raters cross-evaluate the three selected subjects.

[0072] S43. Average the secondary scores of each selected object to obtain the secondary aggregation result, and calculate the secondary group mean of each group accordingly. , , .

[0073] S5. Scale readjustment and result comparison.

[0074] use , and each object Calculate the final aggregate score for all 15 objects. It also provides a global ranking.

[0075] The comparison revealed that the object with the highest score within a group in the initial aggregation was no longer ranked first globally in the final ranking, but was replaced by an object from another group; similarly, the object with the lowest score in the initial aggregation also saw a significant improvement in its final ranking. This demonstrates that the proposed method effectively eliminates the interference of group heterogeneity on the ranking results, making the final ranking closer to the true quality of the objects.

[0076] In larger-scale simulation experiments, comparing the Kendall correlation coefficients of the direct aggregation method and the method of the present invention, it was found that as the standard difference and variance difference between groups increase, the ranking accuracy advantage of the method of the present invention becomes more and more obvious; even when there is unfair scoring, the robust aggregation method can still maintain high robustness.

[0077] Therefore, this invention adopts the aforementioned robust scoring aggregation method based on secondary review. Addressing the heterogeneity problem between groups caused by inconsistencies in scoring standards and variances among different groups in parallel group scoring activities, it proposes an aggregation method that can automatically readjust the scoring scales of each group, enabling fair comparison between objects from different groups. In the presence of unfair scoring (malicious scoring, arbitrary scoring), it improves the robustness of the aggregation results and reduces the impact of abnormal scoring on the results. While ensuring the overall efficiency of the scoring task, it establishes connections between groups by designing a small number of secondary review tasks, achieving objective correction of the initial intra-group aggregation results.

[0078] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A robust scoring aggregation method based on secondary review, characterized in that, Includes the following steps: S1. Parallel group scoring modeling: By defining the set of raters and objects and corresponding parameters, the set is divided into non-overlapping groups that meet the total quantity constraint. Within each group, the raters' scores for the objects are limited to a preset range. S2, Initial Rating Aggregation and Standardization within Groups: Robust aggregation of the ratings of objects within each group is performed to obtain the initial aggregated ratings; Calculate the statistics for each group and standardize the Z-score; summarize the variances of each group to obtain the global mean squared error estimate; S3. Construction of Secondary Review Task: To achieve cross-group comparability of scores from different groups, a secondary review process of appropriate scale is set up. The secondary review task is constructed by selecting the secondary review subjects and scorers. S4. Secondary Aggregation and Scale Rebalancing: For the subjects subject to secondary review, calculate the secondary aggregated score and mean, and combine the global variance to achieve scale rebalancing of the score, so as to obtain the final aggregated score that can be compared across groups. S5. Final Score Output: After mapping the final aggregated score to the score range required for actual application, the final score result is output.

2. The robust scoring aggregation method based on secondary review according to claim 1, characterized in that, The specific steps of S1 are as follows: S11. Define the basic set and parameters: Clarify the set of raters and the set of objects to be rated; S12. Divide all objects and raters into several non-overlapping groups, configure the corresponding rater set and object set for each group, and satisfy the quantity constraint conditions. S13. Establish the rating relationship between the rater and the subject within each group, set the range of rating values, and allow for some missing ratings.

3. The robust scoring aggregation method based on secondary review according to claim 1, characterized in that, The specific steps of S2 are as follows: S21. Initial aggregation: For each object in each group, the scores of the raters in the group are aggregated using the average or robust aggregation method to obtain the initial aggregated score, which suppresses the influence of unfair scoring. S22. Within-group statistics calculation: Calculate the mean and standard deviation of the initial aggregated scores of all objects in each group, which are used to quantitatively characterize the overall level and dispersion of the group's scores. S23, Z-score standardization: Based on the initial aggregated score, mean and standard deviation, calculate the standardized Z-score for each group and each object to eliminate inter-group scale differences; S24. Global Variance Statistics: Summarize the variances of all groups, calculate the global mean squared error estimate, and efficiently characterize the overall dispersion level of the overall scoring system.

4. The robust scoring aggregation method based on secondary review according to claim 1, characterized in that, The specific steps of S3 are as follows: S31. Secondary Object Selection: Based on each group of object sets, select a secondary review set and construct a global secondary review object set; S32. Secondary rating selection: Select a portion of the ratings from all ratings or from each group of ratings to form a secondary rating set, and control the selection scale by proportion. S33. Secondary scoring task assignment: Assign cross-group objects to secondary scorers to achieve cross-scoring; and reconstruct the expected quality based on the relationship between the scorer and the object's group, using either the original scoring error model or the inter-group standard deviation.

5. A robust scoring aggregation method based on secondary review according to claim 4, characterized in that, The specific steps of S31 are as follows: S311. From the object set of each group, select objects for secondary review randomly or according to a preset ratio to form a secondary review object set. S312. For the selected secondary review objects, take the union of the sets to form a global secondary review object set, and determine the size of the global secondary review object set.

6. The robust scoring aggregation method based on secondary review according to claim 4, characterized in that, The specific steps of S33 are as follows: S331. Assign objects from different groups to raters in the set to form cross-group ratings; S332. If the rater and the subject are from the same original group, their original rating error model is used directly; if they are from different groups, the expected quality is regenerated based on the standard deviation parameters of the target group and the source group to characterize their rating habits.

7. A robust scoring aggregation method based on secondary review according to claim 1, characterized in that, The specific steps of S4 are as follows: S41. Secondary aggregate score calculation: For each secondary review object, collect the corresponding secondary scores of the secondary reviewers, and use average or robust aggregation to obtain its secondary aggregate score. S42. Secondary group mean estimation: Calculate the secondary aggregated score mean of each group of secondary review subjects as the overall score level after cross-group score correction. S43. Scale readjustment: Based on global variance estimation and the mean of the secondary aggregated scores, the standardized scores are restored to a uniform scale to obtain the final aggregated scores.

8. A robust scoring aggregation method based on secondary review according to claim 7, characterized in that, The specific steps of S43 are as follows: S431. Using a unified global variance estimate and group-level corrected mean, the standardized scores are restored to a unified scale to obtain the final aggregated scores of the objects. S432. Linearly map the final aggregated score to the target score interval to meet application requirements.