Competition review method and device based on standard score adjustment and medium

By quantifying expert review preferences and dynamically adjusting scoring parameters, the problem of differences in expert review preferences and the impact of extreme values ​​in large-scale innovation competitions is solved, thus achieving fairness and accuracy in review results and adapting to the needs of different competition scenarios.

CN120849778APending Publication Date: 2025-10-28NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511004479.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In large-scale innovation competitions, the sheer number of entries and the significant differences in the preferences of expert judges make fair grading and ranking difficult. Existing standard score adjustment methods fail to effectively quantify expert scoring tendencies and remove extreme values, thus affecting the accuracy of the ranking.

Method used

By identifying the intersection of the reviewed works of any two experts, the difference in review scores is calculated to quantify the experts' review bias. This bias is then introduced as an adjustment factor into the standard score calculation model. Combined with a genetic algorithm to optimize parameters, the mean and standard deviation of the scores are dynamically adjusted to weaken the impact of abnormal scores.

Benefits of technology

Effectively quantify expert subjective biases, dynamically adjust scoring parameters, enhance the ability to handle abnormal data, improve the fairness and accuracy of review results, and achieve adaptive matching of the model to different competition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849778A_ABST
    Figure CN120849778A_ABST
Patent Text Reader

Abstract

The invention relates to a competition review method and device based on standard score adjustment and a medium. The method comprises the steps that the intersection of review works of any two experts is determined; calculating a review score difference of each work in the review work intersection under the score of any two experts; according to all other experts having the review work intersection with the target expert, adding and averaging the review score difference of the target expert, and quantifying the review tendency of the target expert; constructing a standard score calculation model, introducing the quantized review tendency as an adjustment factor into standard score mean value calculation, and generating an adjusted standard score calculation formula; and setting a value range of influence parameters in a standard score calculation formula, and optimizing the parameters based on a genetic algorithm to maximize a correlation coefficient of adjusted work sorting and expert negotiation sorting. According to the method, subjective tendency of experts is systematically quantified, scoring parameters are dynamically adjusted, and abnormal data processing capability is enhanced, so that limitation of a traditional standard scoring model is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of competition judging, specifically to a competition judging method, equipment, and medium based on standard score adjustment. Background Art

[0002] In large-scale innovation competitions, the sheer number of entries and the significant differences in expert reviewer preferences make fair grading and ranking crucial issues that urgently need to be addressed. Traditional grading methods typically rely on subjective expert scoring, but due to limitations in workload, time, and resources, it is difficult to ensure that each expert provides a thorough and consistent review of all entries. Although existing technologies employ standardized score adjustment methods (such as Z-score standardization) to mitigate the impact of differences in expert scoring styles, they still have the following limitations: Existing standard scoring models adjust scores by fixing the mean and standard deviation, without systematically quantifying the inherent "positive" or "conservative" scoring tendencies of experts, resulting in the ranking of the true quality of works still being affected by subjective bias. Traditional methods do not incorporate strategies to remove extreme values ​​(such as the highest and lowest scores), and outlier scores may significantly affect the final score of the work. The existing standard scoring formula has fixed parameters, which cannot dynamically adapt to the complex changes in the distribution of work quality and the behavior of expert judges in different competition scenarios.

[0003] To address the aforementioned issues, some improvement methods attempt to optimize scoring by "removing the highest and lowest scores" or introducing expert consultation mechanisms. However, these methods suffer from high computational complexity, low efficiency, and difficulty in scaling up. Furthermore, existing research lacks systematic modeling of the synergistic effects of multiple factors; for example, it does not comprehensively consider the impact of key factors such as work quality, expert bias, and the dispersion of score distribution on the final ranking. Summary of the Invention

[0004] This invention provides a competition review method, equipment, and medium based on standard score adjustment, which aims to solve the problems of fairness in expert scoring and accuracy in ranking entries in large-scale innovation competitions.

[0005] To achieve the above objectives, the first aspect of the present invention provides a competition evaluation method based on standard score adjustment, comprising the following steps: Determine the intersection of the reviewed works of any two experts, where the intersection of the reviewed works is the set of works jointly reviewed by the two experts; Calculate the difference in review scores between any two experts for each work in the intersection of the reviewed works; The review scores of the target expert are summed and averaged based on the reviews of all other experts whose works overlap with those of the target expert, thus quantifying the review bias of the target expert. A standard score calculation model is constructed, and the quantified review tendency is introduced as an adjustment factor into the calculation of the standard score mean, thereby generating an adjusted standard score calculation formula. The range of values ​​for the influencing parameters in the standard score calculation formula is set, and the parameters are optimized based on a genetic algorithm to maximize the correlation coefficient between the adjusted work ranking and the expert consultation ranking.

[0006] Furthermore, the intersection of the reviewed works of any two experts is determined through set operations, wherein the intersection of the reviewed works includes all works that have been jointly rated by the two experts.

[0007] Furthermore, the method for calculating the review score difference includes: for each work in the intersection of the reviewed works, calculating the average score of each work from any two experts, and using the difference between the average scores as the review score difference.

[0008] Furthermore, the formula for quantifying the review preferences of target experts is as follows:

[0009] in, For target experts The evaluation bias quantification value represents the overall score deviation of this expert relative to other experts; For target experts and experts The difference in mean scores among works reviewed jointly; To meet with the target experts The total number of other experts whose reviewed works overlap; To limit the summation range to excluding the target expert Other experts besides [them].

[0010] Furthermore, methods for generating the adjusted standard score calculation formula include: For each expert's review of all works, remove their highest and lowest scores to obtain the remaining score set; Calculate the mean and standard deviation of the remaining rating set as the adjusted rating mean and benchmark standard deviation; The quantified target expert review tendency is used as an adjustment factor and linearly combined with the adjusted mean score to generate a dynamic mean parameter. Based on the dynamic mean parameter and the preset benchmark standard deviation, a standard score calculation formula is constructed, which is the standardized offset of the adjusted score from the dynamic mean parameter.

[0011] Furthermore, the generation of the dynamic mean parameter also includes: dynamically adjusting the weight of the adjustment factor based on the number of works reviewed by experts, wherein the weight is negatively correlated with the number of works reviewed.

[0012] Furthermore, the method for optimizing the parameters based on genetic algorithms includes: Define an initial population, which contains several individuals, each representing a set of parameter combinations, including adjustment factors, mean score weights, standard deviation weights, and reviewer bias weights. For each individual's parameter combination, the scores of all works are adjusted based on the standard score calculation formula to generate an adjusted ranking of the works; The Pearson correlation coefficient between the adjusted work ranking and the expert-deliberated ranking was calculated as the fitness value of the individual. Based on fitness values, individuals in the population are selected, crossovered, and mutated to generate a new generation of population; Repeat the above steps until the preset number of iterations or the fitness value meets the convergence condition, and output the optimal parameter combination.

[0013] Furthermore, when quantifying review bias, if the target expert has no overlap in reviewing works with other experts, their review bias is estimated based on historical review data, which includes the expert's scoring distribution characteristics in similar competitions.

[0014] To achieve the above objectives, a second aspect of the present invention provides an electronic device including a memory and a processor, the memory being used to store a program that supports the processor in executing the competition evaluation method based on standard score adjustment, the processor being configured to execute the program stored in the memory.

[0015] To achieve the above objectives, a third aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the competition evaluation method based on standard score adjustment.

[0016] The beneficial effects of this invention are: Compared with existing technologies, this invention provides a competition review method, device, and medium based on standard score adjustment. By systematically quantifying expert subjective biases, dynamically adjusting scoring parameters, and enhancing abnormal data processing capabilities, it effectively solves the limitations of traditional standard score models. Specifically, firstly, by calculating the average difference in scores between any two experts on jointly reviewed works, an expert review bias quantification model is constructed. This transforms subjective scoring styles into quantifiable "positive" or "conservative" bias values, which are then embedded as adjustment factors in the standard score calculation formula, directly correcting the interference of inherent expert preferences on the mean. Secondly, a strategy of removing the highest and lowest scores is introduced into the standard score model. Combined with the dynamic adjustment of the mean and standard deviation parameters based on expert bias, the impact of abnormal scores on the ranking of works is weakened. Furthermore, a genetic algorithm is used to perform multi-objective optimization of parameters such as weight coefficients and bias correction factors in the model to maximize the correlation between the adjusted work ranking and the expert negotiation results, achieving adaptive matching of the model to different competition scenarios. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below.

[0018] Figure 1 This is a flowchart of a competition review method based on standard score adjustment disclosed in an embodiment of the present invention.

[0019] Figure 2 This is a diagram showing the relationship between the score difference, the review portfolio, the score vector, and the mean, which is disclosed in an embodiment of the present invention.

[0020] Figure 3 This is a description diagram of an expert scoring style disclosed in an embodiment of the present invention.

[0021] Figure 4 This is a diagram showing the original score distribution of each expert in the first round of review, as disclosed in an embodiment of the present invention.

[0022] Figure 5 This is a diagram showing the distribution of the original scores of each expert in the zero-second round of review, as disclosed in an embodiment of the present invention.

[0023] Figure 6 This is a diagram showing the distribution of standard scores for each expert in the first round of review, as disclosed in one embodiment of the present invention.

[0024] Figure 7 This is a diagram showing the distribution of standard scores for each expert in the second round of review, as disclosed in one embodiment of the present invention.

[0025] Figure 8This is a diagram showing the distribution of standard scores for each expert in the first round of review, according to a second embodiment of the present invention.

[0026] Figure 9 This is a diagram showing the distribution of standard scores for each expert in the second round of review, as disclosed in Scheme 2 of this invention.

[0027] Figure 10 This is a diagram showing the distribution of standard scores for each expert in the first round of review, according to Scheme 3 disclosed in this invention.

[0028] Figure 11 This is a diagram showing the distribution of standard scores for each expert in the second round of review, according to one of the embodiments of the present invention.

[0029] Figure 12 This is a diagram showing the distribution of standard scores for each expert in the first round of review, according to one of the four schemes disclosed in this invention.

[0030] Figure 13 This is a diagram showing the distribution of standard scores for each expert in the second round of review, according to one of the embodiments of the present invention.

[0031] Figure 14 This is a diagram showing the distribution of standard scores for each expert in the first round of review, according to one of the embodiments of the present invention.

[0032] Figure 15 This is a diagram showing the distribution of standard scores for each expert in the second round of review, according to one of the embodiments of the present invention.

[0033] Figure 16 This is a diagram showing the distribution of standard scores for each expert in the first round of review, according to one of the six schemes disclosed in this invention.

[0034] Figure 17 This is a diagram showing the distribution of standard scores for each expert in the second round of review, according to one of the six schemes disclosed in this invention.

[0035] Figure 18 This is a diagram showing the distribution of standard scores for each expert in the first round of review, according to one of the seven schemes disclosed in this invention.

[0036] Figure 19 This is a diagram showing the distribution of standard scores for each expert in the second round of review, according to one of the embodiments of the present invention (Scheme 7).

[0037] Figure 20 This is a diagram showing the distribution of standard scores for each expert in the first round of review, as disclosed in one embodiment of the present invention.

[0038] Figure 21 This is a diagram showing the distribution of standard scores for each expert in the second round of review, as disclosed in one embodiment of the present invention.

[0039] Figure 22 This is a graph showing the average distribution of scores for each work in the first round of judging, as disclosed in an embodiment of the present invention.

[0040] Figure 23 This is a graph showing the standard deviation distribution of scores for each work in the first round of judging, as disclosed in an embodiment of the present invention.

[0041] Figure 24 This is a graph showing the average distribution of scores for each work in the second round of judging, as disclosed in an embodiment of the present invention.

[0042] Figure 25 This is a graph showing the standard deviation distribution of scores for each work in the second round of judging, as disclosed in an embodiment of the present invention.

[0043] Figure 26 This is a comparison chart of the ranking of works based on expert consultation, original scores, existing solutions, and self-selected solutions, as disclosed in an embodiment of the present invention. Detailed Implementation

[0044] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0045] According to embodiments of the present invention, it should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the following manufacturing method, in some cases the steps shown or described may be performed in a different order than that shown here.

[0046] like Figure 1 As shown, this invention provides a competition evaluation method based on standard score adjustment, comprising the following steps: Step S100: Determine the intersection of the reviewed works of any two experts, wherein the intersection of the reviewed works is the set of works jointly reviewed by the two experts; Each expert is assigned a number of works during the review process, forming their own independent review portfolio. For any two experts (such as Expert A and Expert B), the intersection of their reviewed works is defined as the works that appear repeatedly in each expert's review portfolio. For example, if Expert A's review portfolio is {Work 1, Work 2, Work 3} and Expert B's review portfolio is {Work 2, Work 3, Work 4}, then the intersection of their reviewed works is {Work 2, Work 3}. This process is implemented through set operations, specifically: For any expert and Record the intersection of their reviewed works as ,Right now: (1) in, and They represent experts respectively. and The collection of works reviewed.

[0047] Step S200: Calculate the difference in review scores between any two experts for each work in the intersection of the reviewed works; For the intersection of the reviewed works of two identified experts (such as Expert A and Expert B) (e.g., Work 2 and Work 3), the scoring data of the two experts for these shared works needs to be extracted separately. Taking Work 2 as an example, if Expert A's score is 85 points and Expert B's score is 78 points, then the score difference for a single work is 7 points; similarly, the score difference for all intersecting works is calculated. However, step S200 does not directly calculate the score difference for a single work, but rather uses the mean difference to comprehensively reflect the overall scoring tendency of the two experts on the shared works.

[0048] In practice, the first step is to compile the scoring data for all works in the overlapping pool of reviewed submissions. Assuming that for experts... and Its overlap with the reviewed works The works are divided into sets of 100. and Then experts Compared to Regarding the intersection of reviewed works The difference in review scores is: (2) in, and They represent experts respectively. and Intersection of judged works The average score of the works submitted by the judges.

[0049] gather and Expert and The portfolios for judging are divided into two sets, where each set contains the specific works being judged. Corresponding to these two portfolios, the scoring vectors are... and It records the scores given by experts to all works in their respective review portfolios; essentially, it is a collection of scores. For example, Experts The portfolio of his review The Middle The scoring results for each work.

[0050] Furthermore, and Scoring vectors and The mean of is given as a numerical value. It is defined as the difference between the average scores given by two experts in jointly reviewing a work, and is used to quantify the experts' scores. Compared to The overall scoring tendencies differ.

[0051] The above logical relationship is as follows: Figure 2 Portfolio shown and Generate scoring vectors through rating mapping. and The mean was calculated to obtain and Ultimately, the difference between the two forms .

[0052] Step S300: Based on all other experts whose reviewed works overlap with the target expert, sum and average the differences in the review scores of the target expert to quantify the review tendency of the target expert. After determining the difference in the intersection scores of the works of any two experts, the overall review score difference between that expert and all other experts can be determined by summing and averaging. Specifically, let's assume the target expert is... All of them The other group of experts whose reviewed works overlap is: For each of the other experts Step S200 has calculated the difference in the mean scores of the two authors on their joint work. At this point, it is necessary to... For all To summarize and reflect The tendency to give an average rating relative to all experts.

[0053] In the specific operation, first iterate through all pairs of... There are experts who jointly reviewed the works. Extract the mean difference of the corresponding scores one by one. Then, the arithmetic mean of these differences is calculated using the following formula: (3) in, For target experts The evaluation bias quantification value represents the overall score deviation of this expert relative to other experts; For target experts and experts The difference in mean scores among works reviewed jointly; To meet with the target experts The total number of other experts whose reviewed works overlap; To limit the summation range to excluding the target expert Other experts besides [them].

[0054] Formula (3) can quantify the differences in scoring preferences among different experts, enabling the assessment of all experts' scores. Figure 3 It is quantitatively represented on the "activity level axis".

[0055] This step, by summarizing data from multiple experts, transforms scattered scoring differences into quantifiable expert preference indicators, providing a direct basis for dynamic adjustments in the subsequent standardized scoring model. This reduces the interference of individual scoring styles on the final ranking and improves the fairness of the review results.

[0056] Step S400: Construct a standard score calculation model, and introduce the quantified review tendency as an adjustment factor into the calculation of the standard score mean to generate an adjusted standard score calculation formula. This step introduces quantified expert review preferences as an adjustment factor to weaken the influence of subjective expert scoring habits on the final ranking of works. This step is an improvement on the traditional standard score formula (Formula 4), replacing the fixed average standard score (50 in the formula) with a dynamic parameter, allowing it to reflect the combined effect of the actual quality of the work and expert scoring preferences. The traditional standard score formula is as follows: (4) In formula (4), Experts On the work The standard score is represented by 50, where 50 represents the mean of the changed standard score and 10 represents the variance of the changed standard score. Experts On the work The original rating, Experts The average score of all reviewed entries. Experts The sample standard deviation of the scoring. When considering the modification plan, we start with the mean, turning the fixed value of 50 into a variable that changes with the score of the work. Therefore, the standard score formula (4) is modified into formula (5-11), as shown in Table 1 below: Table 1. Eight review schemes generated by changing the standard score calculation formula.

[0057] Specifically, the traditional standardized score formula uses raw scores... The scores are converted to standardized scores with a mean of 50 and a standard deviation of 10 to eliminate differences in expert rating scales. However, this model does not consider inherent expert rating biases (such as "positive" or "conservative"). Therefore, step S400 proposes to incorporate expert review biases. (Obtained through quantification in step S300) is introduced as an adjustment factor into the mean calculation. For example, in the formula of Scheme Six, the mean part is modified to... ,in, Indicates the work The mean score is calculated by removing the highest and lowest scores from all expert ratings. For experts The review bias value, with the negative sign indicating a reverse correction to the expert bias.

[0058] This improvement achieves a dual objective: firstly... By eliminating extreme scores, the quality level of the work itself is reflected more consistently; secondly, The mean benchmark is dynamically adjusted by introducing a weighting coefficient (0.6) for expert bias. For example, if an expert has a "conservative" bias (i.e., ... The scoring benchmark will be lowered, thereby indirectly improving the standardized results of its original low scores, and vice versa. At the same time, the standard deviation part still uses the traditional method to maintain the consistency of the dispersion of the score distribution.

[0059] Ultimately, the revised standard scoring formula not only retains the core functions of standardized processing but also integrates both the quality of the work and expert bias through a dynamic mean parameter, thus more objectively reflecting the true level of the work. This model design effectively reduces ranking bias caused by differences in experts' subjective scoring styles, improving the fairness and reliability of the review results.

[0060] Step S500: Set the range of values ​​for the influencing parameters in the standard score calculation formula, and optimize the parameters based on the genetic algorithm to maximize the correlation coefficient between the adjusted work ranking and the expert consultation ranking.

[0061] Based on the constructed dynamic standard score model, by introducing adjustable parameters and utilizing historical or actual review data, the optimal parameter combination is automatically found to maximize the correlation between the adjusted work ranking and the expert consensus ranking. The specific implementation process is as follows: The parameters in the standard score model mainly include coefficients affecting mean adjustment, standard deviation weighting, and expert review bias. For example, based on Scheme Six, additional parameters include... (Weighted average of work ratings) (Weighted average of expert scores) (Reviewer preference adjustment factor) and (Standard deviation scaling factor). For example, the improved model in formula (12) is expressed as: (12) The following section uses historical expert scoring data and a genetic algorithm to optimize parameters, further improving the scientific rigor and accuracy of the standardized scoring model. It should be noted that the standard deviation of the expert review score distribution is not adjusted in this model. Since it is assumed that all works are randomly assigned, the quality distribution of works reviewed by a unified set of experts is generally considered to be consistent. Furthermore, setting the standard deviation of all expert review score distributions to be the same can avoid score distribution anomalies caused by individual expert factors.

[0062] Let's assume the Pearson correlation coefficient between the ranking scheme obtained after standard score calculation and the expert consensus ranking scheme. Analysis of Scheme Six reveals that the main factors influencing expert review score adjustments include the mean of the work scores, expert review bias, expert score mean, and expert score standard deviation. The mean of the work scores and expert review bias together determine the overall upper and lower levels of the expert score distribution. The standard deviation measures the distance of expert review scores from the mean, reflecting the dispersion of the review scores, and determines the width of the distribution curve. When the standard deviation is large, the curve is flatter, and the distribution is more dispersed; when the standard deviation is small, the curve is steeper, and the distribution is more concentrated. To make the model closer to the expert consensus result, this method introduces parameters for the above four factors, aiming to maximize the Pearson correlation coefficient by optimizing these parameters. Therefore, the optimized model is established as follows: (13) in, Let be the objective function value, and represent the ranking vector based on standard scores. Consult with experts to sort the vector The Pearson correlation coefficient is used to measure the consistency between the two. A ranking vector of works calculated based on the standard score formula. The Each element reflects the work. The standardized ranking is the result of ranking the works after the review scheme was adjusted. Represents the consensus vector of the works. The Each element reflects the work's collective approval by experts. The ranking Represents the sorting vector The length of the mold, Represents the expert consultation ranking vector The length of the mold, and These represent the work ranking vector calculated based on the standard score formula and the work ranking vector agreed upon by experts, respectively. This indicates the number of works and the upper limit of the summation calculation range. The constraints respectively constrain the parameters and The range of values ​​is .

[0063] Using the optimization model described above, the optimal standard score calculation model can be generated by optimizing the relevant parameters.

[0064] Similarly, to solve the above optimization model using a genetic algorithm, we first define the population structure of the genetic algorithm, with the number of individuals in the population being... Each individual represents a parameter setting scheme, describing the parameter values ​​for four influencing factors. The genes on each individual represent the values ​​of the corresponding factors. Because excessively large parameter values ​​would cause the original scoring to lose its objective assessment of the work's quality by experts, this method sets its value range to be [insert range here]. .

[0065] For each individual, i.e., each review task allocation scheme, calculate its fitness, which is an indicator of the quality of the individual's solution. The fitness function is referenced in formula (14). The optimization objective is defined as: (14) in, This represents the objective function value, i.e., the Pearson correlation coefficient. This is used to measure the consistency between two sets of ranking results. This is the ranking result of the works after the review scheme was adjusted. This is the standard ranking result agreed upon by multiple experts in the second phase. Represents data group The sample mean, this formula is used to measure the ranking of works after adjustments under different review schemes ( Standard ranking agreed upon by experts ( The similarity between ( ), Pearson correlation coefficient The higher the value, the closer the ranking result of the scheme is to the expert consensus.

[0066] To further explain why Scheme 6 was chosen as the improvement in this invention, the following experiment will be conducted: Based on the existing standard score calculation model (Formula (4)), and considering the quality of the work itself and the expert scoring bias, seven new standard score calculation models (Formulas (5-11)) are proposed, specifically for the standard score mean. Table 2 below explains the practical implications of these eight formulas: Table 2. Practical implications of the eight standard score calculation models

[0067] Mean minus review bias in the model Its practical significance lies in eliminating the influence of experts' subjective scoring biases on the ranking of works' true quality, through changes... Multiplier factors are used to change the degree of influence, thereby generating different solutions.

[0068] For the control experiment, this method adds scheme zero to Table 2, which represents the scheme without any adjustments, as detailed in Table 3.

[0069] Table 3 adds a solution of zero without any processing.

[0070] The nine schemes in total, consisting of eight schemes with zero-sum design, were examined in the first and second stages (the first stage was online review, and the second stage was on-site review). The changes in expert scores and project grades after adjusting the standard scores were investigated.

[0071] Figures 4-26 The presentation shows the distribution of each reviewer's raw scores and standardized scores for schemes zero through nine. From... Figures 4-21 The 18 box plots clearly show that standardized score models can reduce the impact of the range. Standardized data does not reduce the range (maximum minus minimum), but it can reduce the impact of the range on data analysis. Standardized data focuses more on the relative positions between data points, rather than being affected by absolute values.

[0072] pass Figures 22-26 Image analysis revealed that schemes two, four, six, and eight were more effective at handling outlier data. This is because these four models incorporated a review bias. This indicates that the factor of expert scoring habits may be reduced, so the expert scores are closer to the actual situation, and outliers are usually easier to identify.

[0073] By comparing the distribution of scores before and after different scheme adjustments and the changes in the scores of the works, the following three conclusions can be drawn.

[0074] (1) All the adjusted scores are closer to the consensus of the experts in the ranking of the works than the original scores. Therefore, it is necessary to adjust the original scores and then determine the final ranking.

[0075] (2) Taking Schemes 3 and 4, 5 and 6, 7 and 8 as examples, all the ranking schemes obtained by the standard score calculation models that consider the expert scoring bias are closer to the ranking results agreed upon by the experts than the schemes that do not consider the expert scoring bias. It can be concluded that considering the expert scoring bias is meaningful. Although the correlation coefficient between Scheme 2 and the ranking results agreed upon by the experts is lower than that of the basic standard score model, the difference is not significant.

[0076] (3) Among all the standard score calculation models, the average score of the works after removing the highest and lowest scores is the mean, and the standard score calculation scheme that considers the expert scoring bias has the highest correlation with the final ranking scheme determined by the experts. In particular, the correlation coefficients of schemes five and six are higher than those of schemes seven and eight. It can be seen that the average score of the works after removing the highest and lowest scores can better reflect the quality of the works themselves and weaken the influence of the review allocation scheme and other factors on the misjudgment of the quality of the works.

[0077] The Pearson correlation coefficient is a statistical indicator used to measure the strength and direction of the linear relationship between two numerical variables. It is calculated based on the concepts of covariance and standard deviation. Specifically, the Pearson correlation coefficient represents the degree of linear correlation between two variables.

[0078] Pearson correlation coefficient is expressed as The calculation formula is as follows: (16) Where X and Y represent two numerical vectors, n is the sample size, and |X| and |Y| are the means of X and Y, respectively.

[0079] The Pearson correlation coefficient r ranges from -1 (perfectly negative correlation) to 1 (perfectly positive correlation). When r is 0, it indicates that there is no linear correlation between the two variables.

[0080] Since the first-prize winning entries, selected by multiple experts, have a high degree of credibility, the ranking of the winning entries by multiple experts will be used as the standard ranking.

[0081] The scores of the works after adjustments from scheme zero to scheme eight were ranked, and the correlation coefficient between them and the standard ranking was calculated. After several stages, the schemes were ranked from largest to smallest according to their correlation coefficients, as shown in Table 4.

[0082] Table 4. Scheme ranking based on correlation coefficient

[0083] Table 4 shows that the ranking of the options is as follows: Option 6 > Option 1 > Option 2 > Option 4 > Option 5 > Option 8 > Option 7 > Option 3 > Option 0. Since Option 6 has the highest accuracy, the aforementioned improvements were made to the standard score model corresponding to Option 6.

[0084] This embodiment proposes the traditional standard score formula (4) to address the impact of differences in expert scoring styles on the comparability of scores. Formula (4) eliminates the systematic bias of different expert scoring scales by converting the original scores into standard scores with a mean of 50 and a standard deviation of 10, thus standardizing the scores. However, in large-scale competitions, the overlap of expert review portfolios is low, and the assumption that the quality distribution of portfolios is the same in the traditional standard score does not hold, thus limiting its applicability. To address this, by adjusting the calculation method of the standard score mean (such as introducing expert bias, average work score, and removing extreme values), eight improved schemes and a control scheme zero were generated, totaling nine schemes. These schemes aim to explore standardization methods that are more in line with actual review scenarios. By comparing the score distribution before and after adjustment and the correlation coefficient with the expert-negotiated ranking, the optimization effect of different models is verified. In particular, Scheme Six uses the "average work score after removing the highest and lowest scores" as the standard score mean and introduces an expert scoring bias adjustment term. This design avoids the interference of extreme scores on the mean and weakens the influence of the expert's subjective scoring style, making the adjusted score closer to the true quality of the work. Table 4 shows that Scheme 6 has the highest Pearson correlation coefficient (0.6459) with the expert consultation ranking, indicating that its ranking result is closest to the multi-expert consensus, verifying the effectiveness of the model in balancing fairness, anti-interference and accuracy.

[0085] According to another aspect of the embodiments of this application, an electronic device is also provided, including a processor and a memory, wherein the processor is configured to implement the steps of the method when executing a computer program stored in the memory.

[0086] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0087] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0088] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0089] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0090] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A competition judging method based on standard score adjustment, characterized in that, The steps include: Determine the intersection of the reviewed works of any two experts, where the intersection of the reviewed works is the set of works jointly reviewed by the two experts; Calculate the difference in review scores between any two experts for each work in the intersection of the reviewed works; The review scores of the target expert are summed and averaged based on the reviews of all other experts whose works overlap with those of the target expert, thus quantifying the review bias of the target expert. A standard score calculation model is constructed, and the quantified review tendency is introduced as an adjustment factor into the calculation of the standard score mean, thereby generating an adjusted standard score calculation formula. The range of values ​​for the influencing parameters in the standard score calculation formula is set, and the parameters are optimized based on a genetic algorithm to maximize the correlation coefficient between the adjusted work ranking and the expert consultation ranking.

2. The competition judging method based on standard score adjustment as described in claim 1, characterized in that, The intersection of the reviewed works of any two experts is determined by set operations, wherein the intersection of the reviewed works includes all works that are jointly rated by the two experts.

3. The competition judging method based on standard score adjustment as described in claim 1, characterized in that, The method for calculating the difference in review scores includes: for each work in the intersection of the reviewed works, calculating the average score of each work from any two experts, and taking the difference between the average scores as the difference in review scores.

4. The competition judging method based on standard score adjustment as described in claim 1, characterized in that, The formula for quantifying the review preferences of target experts is: in, For target experts The evaluation bias quantification value represents the overall score deviation of this expert relative to other experts; For target experts and experts The difference in mean scores among works reviewed jointly; To meet with the target experts The total number of other experts whose reviewed works overlap; To limit the summation range to excluding the target expert Other experts besides [them].

5. The competition judging method based on standard score adjustment as described in claim 1, characterized in that, Methods for generating the adjusted standard score calculation formula include: For each expert's review of all works, remove their highest and lowest scores to obtain the remaining score set; Calculate the mean and standard deviation of the remaining rating set as the adjusted rating mean and benchmark standard deviation; The quantified target expert review tendency is used as an adjustment factor and linearly combined with the adjusted mean score to generate a dynamic mean parameter. Based on the dynamic mean parameter and the preset benchmark standard deviation, a standard score calculation formula is constructed, which is the standardized offset of the adjusted score from the dynamic mean parameter.

6. The competition judging method based on standard score adjustment as described in claim 5, characterized in that, The generation of the dynamic mean parameter also includes: dynamically adjusting the weight of the adjustment factor based on the number of works reviewed by experts, wherein the weight is negatively correlated with the number of works reviewed.

7. The competition judging method based on standard score adjustment as described in claim 1, characterized in that, Methods for optimizing the parameters based on genetic algorithms include: Define an initial population, which contains several individuals, each representing a set of parameter combinations, including adjustment factors, mean score weights, standard deviation weights, and reviewer bias weights. For each individual's parameter combination, the scores of all works are adjusted based on the standard score calculation formula to generate an adjusted ranking of the works; The Pearson correlation coefficient between the adjusted work ranking and the expert-deliberated ranking was calculated as the fitness value of the individual. Based on fitness values, individuals in the population are selected, crossovered, and mutated to generate a new generation of population; Repeat the above steps until the preset number of iterations or the fitness value meets the convergence condition, and output the optimal parameter combination.

8. The competition judging method based on standard score adjustment as described in claim 1, characterized in that, When quantifying review bias, if the target expert has no overlap in reviewing works with other experts, their review bias is estimated based on historical review data, which includes the expert's scoring distribution characteristics in similar competitions.

9. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store programs that support the processor in executing the competition evaluation method based on standard score adjustment as described in any of claims 1-8, and the processor is configured to execute the programs stored in the memory.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is run by the processor, it performs the steps of the competition evaluation method based on standard score adjustment as described in any one of claims 1-8.