Trusted evaluation weight distribution method and system based on large model
By constructing an indicator evaluation model and an expert role scoring process based on a large language model, the reliability assessment method solves the problems of stability, efficiency and adaptability in reliability assessment, and realizes efficient and reliable aggregation and automated processing of assessment results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies suffer from insufficient stability and robustness, low alignment with human review, poor efficiency and scalability, weak auditability, poor adaptability and portability, and insufficient theoretical interpretability in trust assessment.
We employ a large language model-based approach to construct an indicator evaluation model, generate prompt templates and scoring scales for expert roles, organize round-based questioning and review processes, calculate consistency, calibration improvement, and stability indicators, dynamically adjust weights, and aggregate scores by combining objective and subjective weights.
It improves the stability and robustness of evaluation results, enhances relevance and consistency with authoritative reviews, automates processes, reduces human and time costs, supports the expansion of massive objects and multi-dimensional data, and maintains robustness across different data scales and domains.
Smart Images

Figure CN121787970A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of trustworthiness assessment technology, specifically to a trustworthiness assessment weight allocation method and system based on a large model. Background Technology
[0002] Currently, traditional methods have the following drawbacks: Regarding stability and robustness: The failure to distinguish between high-quality and low-quality opinions and the lack of suppression for highly relevant dimensions or experts resulted in low-reliability or high-redundancy information affecting the results, leading to large variance and instability in the results, and a lack of robustness to conflicts and noise.
[0003] Regarding alignment with human reviewers: the weights are unrelated to post-debate calibration improvements, making it difficult to reflect improvements in evidence quality; the final rankings differ significantly from those of authoritative reviewers, and the relevance or consistency indicators are low.
[0004] In terms of efficiency and scalability: manual AHP or Delphi is highly subjective, time-consuming, and has poor cross-task migration capabilities.
[0005] In terms of auditability and reproducibility: the lack of structured records of evidence, calculation, weighting and conclusion makes it impossible to trace the source of the authority and makes it difficult to support compliance review and result verification.
[0006] In terms of adaptability and portability: fixed weights or fixed fusion ratios fail in new scenarios and are difficult to maintain robustness and portability across different data scales and domains.
[0007] In terms of theoretical interpretability: the use of black-box average voting makes it difficult to explain the reasons for each step of weighting. Summary of the Invention
[0008] To achieve the above objectives, the present invention provides the following technical solution: a reliable evaluation weight allocation method based on a large model, comprising: An indicator evaluation model is constructed based on a large language model. The project workpiece is analyzed based on the indicator evaluation model to obtain the original scoring data and related evidence data. Generate and manage prompt templates and scoring scales for multiple expert roles; trigger each expert role to score the evaluation object by dimension according to the prompt templates and scoring scales, and generate corresponding textual evidence; wherein, the textual evidence is the basis for each expert role's scoring; Organize a round-based questioning and review process; in each round, each expert role conducts questioning based on the scores and textual evidence of other expert roles, and the questioned expert role reviews and adjusts the scores and textual evidence; at the same time, control the termination conditions and convergence criteria of the questioning and review process; the termination condition is reaching a preset number of questions and reviews or the difference in scores among expert roles is less than a preset threshold, and the convergence criterion is that the scores of each expert role tend to stabilize. Based on the final scores from each expert role, consistency index, calibration improvement index, and stability index are calculated, and these indices are mapped to reliability factors. Based on the final scores from each expert role, the standard deviation, correlation matrix, and information content are calculated; based on the standard deviation, correlation matrix, and information content, the objective weights are determined; based on the objective weights, the subjective weights of each expert role, and the reliability factor, the final weights are output; wherein, the subjective weights are the weights assigned by each expert role to each evaluation dimension based on their own experience and professional knowledge. The scores given by each expert role to the evaluated object are aggregated based on the final weights, and the aggregated score result is output.
[0009] Preferably, the indicator evaluation model includes a demand analysis agent, a knowledge retrieval agent, and a fact verification agent; the original scoring data is used to characterize the initial scores of different evaluation objects under each evaluation dimension, and the relevant evidence data is used to characterize various types of data in the original scoring data; The requirement parsing agent is used to parse project documents, identify the core business objectives in the project documents using hierarchical extraction technology, recursively decompose the core business objectives into a set of candidate attributes and supporting evidence types, identify and define fuzzy and lacking evidence indicators in the candidate attribute set, and generate a missing information set. The project documents include a requirement specification and an architecture design document. The knowledge retrieval agent is used to search for missing information sets in an external general knowledge base, supplement the missing information sets, obtain supplementary information, and classify the retrieval results into four levels according to the evidence reliability label: authoritative confirmation, domain consensus, expert experience, and weak inference, and assign corresponding reliability weights. The fact-verifying agent is used to check whether there are logical contradictions between the supplementary information and the project documents.
[0010] Preferably, after analyzing the project artifacts based on the indicator evaluation model to obtain the original scoring data and relevant evidence data, the method further includes: Define the scope of standardization; For each rating in the original rating data, a linear transformation is performed according to the relative position of the rating among all ratings, based on a defined standardization range, to generate a standardized matrix; Create an index for each piece of raw scoring data to record the corresponding position of the raw scoring data in the relevant evidence data, thus forming an evidence index; The standardization matrix is used to unify the original scoring data to a specific range, and the evidence index is used to record the location information of the relevant evidence data corresponding to each original scoring data.
[0011] Preferably, the prompt template is used to guide the large model to simulate different expert roles for evaluation, and the scoring scale is the standard for each expert role to score the evaluation object in different dimensions; Generate and manage prompt templates and scoring scales for multiple expert roles, including: Determine the required type of expert role based on the type of assessment object and the assessment purpose; For each type of expert role, a corresponding prompt template is generated, wherein the prompt template includes background information of the evaluation object, description of the evaluation dimensions, and guiding questions; Develop a scoring scale for each expert role type.
[0012] Preferably, each expert role is triggered to score the evaluation object according to the prompt template and scoring scale, and generate corresponding textual evidence, including: Input the prompt template and scoring scale into the large model to simulate the scoring of the evaluation object by various expert roles; The large model scores the evaluated object across various evaluation dimensions based on the guiding questions in the prompt template and the scoring scale. At the same time, the large model generates corresponding textual evidence based on the scoring criteria, explaining the reasons and basis for the scoring.
[0013] Preferably, the organization of a round-robin questioning and review process includes: At the start of each round, one expert role is randomly selected as the questioner, and the other expert roles are questioned. The questioning party raises questions based on the scores given by the questioned party and written evidence; The party being questioned shall review the questions raised and make adjustments if it finds that there are problems with the scoring or the textual evidence. Repeat the above questioning and review process until the termination conditions and convergence criteria are met.
[0014] Preferably, based on the final scoring data from each expert role, consistency indicators, calibration improvement indicators, and stability indicators are calculated, and these indicators are mapped to reliability factors, including: Consistency indicators are determined by calculating the correlation coefficients or differences between the scores given by different expert roles. Compare the initial score with the final score, calculate the percentage or extent of score improvement, and determine the calibration improvement indicators; Observe the fluctuations in scores given by each expert role during multiple rounds of questioning and review, calculate the fluctuation range, and determine the stability index; The consistency index, calibration improvement index, and stability index are converted into reliability factors according to the preset mapping rules.
[0015] Preferably, based on the final scoring data of each expert role, the standard deviation, correlation matrix, and information content are calculated; based on the standard deviation, correlation matrix, and information content, objective weights are determined, including: For each evaluation dimension, the standard deviation of the scores given by each expert role is calculated, and the standard deviation is used to indicate the degree of dispersion of the scores; Calculate the correlation coefficients between the scores of each evaluation dimension to form a correlation matrix, which is used to indicate the correlation between the dimensions; Based on the standard deviation and correlation matrix of the scores for each evaluation dimension, the information content of each evaluation dimension is calculated. The greater the information content, the richer the information contained in the score of that dimension. Based on the amount of information in each evaluation dimension, the objective weights are determined according to the principle that the greater the amount of information, the higher the weight. Based on objective weights, subjective weights for each expert role, and reliability factors, the final weights are output, including: Determine the integration ratio of subjective weights and objective weights; Based on the fusion ratio, the subjective and objective weights of each expert role are weighted and summed. Based on the reliability factor, the weights after weighted summation are adjusted, and the final weights are output. The higher the reliability factor, the closer the adjusted weights are to the weighted summation result. The consistency index is used to indicate the degree of consistency in the scores given by each expert role, the calibration improvement index is used to indicate the degree of improvement in the scores after questioning and review, and the stability index is used to indicate the fluctuation of the scores in multiple rounds of questioning and review.
[0016] Preferably, the scores given by each expert role to the evaluated object are aggregated according to the final weight, and the aggregated score result is output, including: Multiply the scores given by each expert role to the evaluated object under each evaluation dimension by the corresponding final weight; The weighted scores under each evaluation dimension are summed to obtain the comprehensive score of the evaluated object under each expert role evaluation. The comprehensive scores from each expert role are aggregated to output the final aggregated score.
[0017] A credibility assessment weight allocation system based on a large model, applicable to the aforementioned credibility assessment weight allocation method based on a large model, includes: The document parsing unit is used to construct an indicator evaluation model based on a large language model, and to parse the project workpiece based on the indicator evaluation model to obtain the original scoring data and related evidence data. An independent review unit is used to generate and manage prompt templates and scoring scales for multiple expert roles; it triggers each expert role to score the evaluation object by dimension according to the prompt templates and scoring scales, and generates corresponding textual evidence; wherein, the textual evidence is the basis for each expert role's scoring. The group debate unit is used to organize a round-based questioning and review process. In each round, each expert role questions based on the scores and textual evidence of other expert roles. The expert role being questioned reviews the questions and adjusts the scores and textual evidence. At the same time, the termination conditions and convergence criteria of the questioning and review process are controlled. The termination condition is that a preset number of questions and reviews is reached or the difference in scores among expert roles is less than a preset threshold. The convergence criterion is that the scores of each expert role tend to stabilize. The reliability quantification unit is used to calculate the consistency index, calibration improvement index and stability index based on the final scoring data of each expert role, and to map the consistency index, calibration improvement index and stability index to the reliability factor. The weight fusion unit is used to calculate the standard deviation, correlation matrix, and information content based on the final scoring data of each expert role; determine the objective weight based on the standard deviation, correlation matrix, and information content; and output the final weight based on the objective weight, the subjective weight of each expert role, and the reliability factor; wherein, the subjective weight is the weight assigned by each expert role to each evaluation dimension based on their own experience and professional knowledge. The scoring aggregation unit is used to aggregate the scores of each expert role for the evaluation object according to the final weight, and output the aggregated scoring result.
[0018] Compared with the prior art, the beneficial effects of the present invention are: This invention introduces a reliability factor for dynamic weighting, which reduces the weight of low-reliability and highly redundant information, thereby reducing the variance of the results and improving the stability and robustness of decision-making, effectively offsetting noise and conflict. The calibration enhancement is mapped into the weights, which strengthens the weights of experts and dimensions that show significant improvement after debate, resulting in a final ranking that is closer to authoritative review, improving relevance and consistency.
[0019] This invention replaces manual review with a large-scale model and multi-expert intelligent agent, achieving automated parallel processing of the process, significantly reducing manpower and time costs, supporting massive objects and multi-dimensional data expansion, and ensuring traceability by recording standardized parameters, correlation matrices, reliability factors, debate records and final results throughout the entire process, supporting compliance review and result verification. Moreover, the parameters can be adaptively learned on the validation set or configured according to the scenario, so as to maintain robustness in different data scales and domains and enhance portability. Attached Figure Description
[0020] Figure 1 This is a schematic flowchart of the overall method in one embodiment of the present invention; Figure 2 This is a schematic diagram of the overall system architecture in one embodiment of the present invention.
[0021] In the diagram: 1. Document parsing unit; 2. Independent review unit; 3. Group debate unit; 4. Reliability quantification unit; 5. Weight fusion unit; 6. Scoring aggregation unit. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Example 1, please refer to Figure 1 This invention provides a technical solution: a reliable evaluation weight allocation method based on a large model, comprising: S1. Construct an indicator evaluation model based on a large language model, analyze the project artifacts based on the indicator evaluation model, and obtain the original scoring data and relevant evidence data. S2. Generate and manage prompt templates and scoring scales for multiple expert roles; trigger each expert role to score the evaluation object by dimension according to the prompt templates and scoring scales, and generate corresponding textual evidence; wherein, the textual evidence is the basis for each expert role's scoring. It should be noted that, from the software quality model and historical project documents, LLM is used to extract the concerns of different stakeholders and summarize them into a three-dimensional feature vector. : Attribute Priority The most emphasized attributes of a character (such as safety and efficiency).
[0024] Risk appetite Tolerance for uncertainty (such as risk aversion, neutrality, and risk preference).
[0025] Evidence preference Preferred types of evidence (e.g., priority given to standard clauses, measured data, and user feedback).
[0026] Role instantiation: Based on the above feature vectors, the abstract role configuration is encoded into structured prompt words and instantiated into a set of expert intelligent agents with independent personalities and reasoning abilities (such as security expert agents, performance engineers, product managers, and QA experts).
[0027] S3. Organize a round-based questioning and review process; in each round, each expert role conducts questioning based on the scores and textual evidence of other expert roles, and the questioned expert role conducts review and adjusts the scores and textual evidence; at the same time, control the termination conditions and convergence criteria of the questioning and review process; the termination condition is reaching a preset number of questions and reviews or the difference in scores among expert roles is less than a preset threshold, and the convergence criterion is that the scores of each expert role tend to stabilize. It should be noted that, in response to cognitive biases and conflicts arising from subjective empowerment, this step designs a multi-round structured debate mechanism managed by a debate coordinator: Modeling: Introduction The proposed binary pair To decouple the characterization expert evaluation. Among them, It is a fuzzy number (such as a triangular fuzzy number) that describes the weight distribution. It is a scalar measure describing the level of expert confidence. The value is determined by the model's internal confidence level (based on...) It is calculated by weighting probability and external evidence support (based on the level of evidence cited).
[0028] Semantic entropy-driven: After each round of debate, all experts' natural language arguments are mapped into semantic vectors, and then... Clustering is used to form opinion clusters. Semantic entropy is used to quantify the dispersion of the current group's opinions.
[0029] The structured debate process comprises three sub-stages: position statement, cross-examination, and consensus convergence. The coordinator, based on the opposition of semantic clusters, instructs experts with different positions to attack each other's weaknesses in their arguments (such as unreliable evidence or ignoring specific risks). Experts must dynamically adjust their viewpoints (revise A) or reduce their confidence (reduce B) during the debate. The debate stops when the semantic entropy falls below a preset threshold and the fuzzy variance meets the convergence condition, and the subjective consensus weight and subjective reliability are output.
[0030] S4. Based on the final scoring data of each expert role, calculate the consistency index, calibration improvement index and stability index, and map the consistency index, calibration improvement index and stability index to the reliability factor. S5. Calculate the standard deviation, correlation matrix, and information content based on the final scoring data of each expert role; determine the objective weights based on the standard deviation, correlation matrix, and information content; output the final weights based on the objective weights, the subjective weights of each expert role, and the reliability factor; where the subjective weights are the weights assigned to each evaluation dimension by each expert role based on their own experience and professional knowledge. It should be noted that, in order to organically combine data-driven approaches with expert experience, this step introduces an objective analytical agent: Data reliability modeling: A data reliability model is established while calculating objective weights using the CRITIC method. This comprehensively considers missing data rates, high-frequency noise, and time lag to calculate the reliability score of the objective data. .
[0031] Reliability gating mechanism: This mechanism incorporates subjective reliability... With objective reliability Perform isomorphic alignment, using Construct a nonlinear gating function to calculate the subjective attention coefficient. and objective attention coefficient This mechanism can automatically suppress the contribution of a source to the final result when the quality of the source deteriorates (such as when data is severely missing or experts have significant disagreements), thus achieving adaptive rebalancing of weights.
[0032] S6. Aggregate the scores of each expert role for the evaluated object according to the final weight, and output the aggregated score result.
[0033] In an optional embodiment, the indicator evaluation model includes a demand analysis agent, a knowledge retrieval agent, and a fact verification agent; the original scoring data is used to characterize the initial scores of different evaluation objects under each evaluation dimension, and the relevant evidence data is used to characterize various types of information in the original scoring data; The requirements analysis agent is used to parse project documents, use hierarchical extraction technology to identify the core business objectives in the project documents, recursively decompose the core business objectives into a set of candidate attributes and supporting evidence types, identify and define fuzzy and lacking evidence indicators in the candidate attribute set, and generate a missing information set. The project documents include the requirements specification and architecture design document. The knowledge retrieval agent is used to search for missing information sets in external general knowledge bases, supplement the missing information sets, obtain supplementary information, and classify the retrieval results into four levels according to the evidence reliability label: authoritative confirmation, domain consensus, expert experience and weak inference, and assign corresponding reliability weights. The fact-verifying agent is used to check for logical contradictions between supplementary information and project documents.
[0034] It should be noted that the working logic of the requirement parsing agent is as follows: The core task of the requirements analysis agent is to transform unstructured project documents into structured initial evaluation models. Input: Project artifact set (Including Requirements Specification (SRS), Architectural Design Document (SDD), and a collection of related software standard documents) (e.g., GB / T25000.10); Processing procedure: Goal identification: The requirements analysis agent uses a prompt (such as "You are a senior requirements analyst...") to extract the project's core business goals from the document. Hierarchical decomposition: Using chain-like thinking, the Goal is decomposed into a set of candidate top-level attributes. (such as security, reliability, performance efficiency); Recursive extraction: Further extract the set of sub-attributes S under each top-level attribute and the set of potential evidence types E supporting these attributes (such as "server logs" and "static code scan report"). Missing Nodes Identification: The requirement parsing agent performs a self-reflection mechanism, traversing each node in the model and calculating its definition clarity. If an attribute is vaguely defined (e.g., only mentioning "the system must have high availability" but without providing quantitative indicators such as "annual downtime"), or lacks corresponding evidence, AR marks it as a missing item and generates a missing information set. ; Mathematical expression: ; in The preset confidence threshold; The working logic of the knowledge retrieval agent is as follows: The task of a knowledge retrieval agent is to fill cognitive gaps using external knowledge and to assign credibility labels to information; Input: Missing information set ; Processing procedure: External retrieval: The knowledge retrieval agent will retrieve missing items. The mapping is to a retrieval query, which performs a search in an external general knowledge base K (containing industry white papers, academic papers, and standard specification libraries); Evidence reliability grading: The knowledge retrieval agent classifies retrieved information into discrete ordinal sets based on the authority of the source. and assign weights : (Authoritative Verification): Benchmark data from national standards (GB / ISO) or projects with clearly defined sources; weighting ; (Domain consensus): Derived from widely accepted industry best practices or technical white papers; Weighting ; (Expert Experience): The large-scale model originates from logical reasoning based on pre-trained knowledge, which falls under general common sense; weights ; (Weak inference): Vague speculation or analogy based solely on context, lacking a clear source; weighting. ; Output: Complete set Each message All of them are accompanied by reliability tags. ; The working logic of the fact-verifying agent is as follows: The fact-verifying agent is responsible for ensuring that the introduced external knowledge does not conflict with the actual situation of the project. Input: complete set and project original documents ; Processing procedure: Logical consistency check: Fact-verifying agent check Does it contradict the initial model? For example, the knowledge retrieval agent introduces the "stateless service recoverability" metric, but the fact-verifying agent finds that the architecture document explicitly describes the system as "stateful," in which case the fact-verifying agent will mark the metric as "not applicable." Implementation check: Check whether there are corresponding entities in the project documentation to support the data collection for this metric; Model update: Validated information is incorporated into the model. ;in, This represents a fact-verifying agent; This closed loop continues to iterate until the size of the missing information set is reached. Or reach the maximum number of iterations The final generated indicator set This will serve as the benchmark for subsequent weight allocation.
[0035] In an optional embodiment, after analyzing the project artifacts based on the indicator evaluation model to obtain the original scoring data and relevant evidence data, the method further includes: Define the scope of standardization; For each rating in the original rating data, a linear transformation is performed according to the relative position of the rating among all ratings, based on a defined standardization range, to generate a standardized matrix; Create an index for each piece of raw scoring data to record the corresponding position of the raw scoring data in the relevant evidence data, thus forming an evidence index; The standardization matrix is used to unify the original scoring data to a specific range, and the evidence index is used to record the location information of the relevant evidence data corresponding to each original scoring data.
[0036] In an optional embodiment, the prompt template is used to guide the large model to simulate different expert roles in evaluation, and the scoring scale is the standard by which each expert role scores the evaluation object in different dimensions; Generate and manage prompt templates and scoring scales for multiple expert roles, including: Determine the required type of expert role based on the type of assessment object and the assessment purpose; For each type of expert role, a corresponding prompt template is generated. The prompt template includes background information of the evaluation object, explanation of the evaluation dimensions, and guiding questions. Develop a scoring scale for each expert role type.
[0037] It should be noted that LLM was used to perform paragraph-level analysis on a large number of software engineering documents, quality models, and project documents to identify the differences in focus among different roles during evaluation. These differences were then summarized into three core dimensions to form a feature vector. : Attribute Priority The quality attribute that this character is most concerned about. Risk appetite The character's attitude towards uncertainty (Risk Appetite); Evidence preference This role tends to rely on a particular form of evidence when making decisions. Based on the above feature vectors, this embodiment constructs the following typical expert roles and injects them into the LLM through PromptEngineering: Example of a Prompt: "You are now a [security expert]; your core concern is [system confidentiality and integrity]; you have an [extremely risk-averse] attitude, preferring to sacrifice performance for security; you prioritize citing [these principles] when arguing your points."
[0038] In an optional embodiment, each expert role is triggered to score the evaluation object according to the prompt template and scoring scale by dimension, and generate corresponding textual evidence, including: Input the prompt template and scoring scale into the large model to simulate the scoring of the evaluation object by various expert roles; The large model scores the evaluated object across various evaluation dimensions based on the guiding questions in the prompt template and the scoring scale. At the same time, the large model generates corresponding textual evidence based on the scoring criteria, explaining the reasons and basis for the scoring.
[0039] In an optional embodiment, a round-robin questioning and review process is organized, including: At the start of each round, one expert role is randomly selected as the questioner, and the other expert roles are questioned. The questioning party raises questions based on the scores given by the questioned party and written evidence; The party being questioned shall review the questions raised and make adjustments if it finds that there are problems with the scoring or the textual evidence. Repeat the above questioning and review process until the termination conditions and convergence criteria are met.
[0040] It should be noted that during the debate, each expert... For attributes Output a ; Fuzzy weights : Due to the ambiguity of language descriptions (such as "very important"), the Triangular Fuzzy Number (TFN) is used to represent the weights. , representing the lower bound, most likely value, and upper bound of the weight, respectively; Reliability : This represents the experts' self-evaluation. The degree of confidence is defined as a weighted average of "intra-model confidence" and "external evidence support". Intra-model reliability Through analysis Text generation (Probability distribution) calculation; to avoid stop words To mitigate interference, a masking function is introduced. Only content words are counted. The average log probability, and using The function performs calibration: ,in The baseline threshold is set at 1.5 to 2.0 (based on experience). Sensitivity coefficient; external evidence support. According to the level of evidence cited by the experts in their arguments: calculate; The more authoritative the evidence cited (e.g.) (Level), the higher the S value; balance factor As the debate progressed... The increase in [something] suggests that experts should rely more on objective evidence than intuition, therefore Set as random The decay function; To determine whether the debate should stop, semantic entropy (SE) is introduced as a criterion for consensus convergence [1]; Reason embedding: Using the Sentence-BERT model to embed all experts on attributes Textual Reasons Mapped to a high-dimensional vector ; Semantic clustering: performing operations on vectors. Clustering groups semantically similar viewpoints into a single category, resulting in semantic clusters. ; Entropy calculation: ; in It represents the percentage of each opinion cluster within the expert group; High entropy (HighSE): This means that opinions are extremely divergent, with multiple equally matched opposing viewpoints that require further debate. Low entropy (LowSE): This means that opinions tend to converge and most experts reach a consensus; Position Statement (Phase 1): All experts output initial positions based on their own Personas. and the reasons; the moderator checks whether the source of the evidence is legal; Cross-challenge (Phase 2): The Moderator identifies opposing semantic clusters (e.g., Cluster 1 emphasizes security, Cluster 2 emphasizes performance); it instructs the Cluster 1 agent to attack the weaknesses in Cluster 2's arguments (e.g., "outdated evidence" or "failure to consider high-concurrency scenarios"); the Cluster 2 agent must defend itself, and if it cannot effectively refute the arguments, it must lower its credibility. Or adjust the weight Move closer to the other party; Consensus convergence (Phase 3): Calculated after each round. ; Stop condition: And the variance of fuzzy numbers ; If the stopping condition is met, proceed to the aggregation phase; otherwise, proceed to the next round of debate (until the maximum number of rounds is reached). The final subjective weighting is not a simple average, but rather based on the quality of the experts' performance in the debate. Weighted: ; in Experts The number of times a viewpoint is cited or echoed by other experts (calculated by semantic similarity matching); Finally, calculate the normalized weights. And aggregate to obtain subjective fuzzy weights. and subjective reliability .
[0041] In an optional embodiment, based on the final scoring data of each expert role, a consistency index, a calibration improvement index, and a stability index are calculated, and these indices are mapped to a reliability factor, including: Consistency indicators are determined by calculating the correlation coefficients or differences between the scores given by different expert roles. Compare the initial score with the final score, calculate the percentage or extent of score improvement, and determine the calibration improvement indicators; Observe the fluctuations in scores given by each expert role during multiple rounds of questioning and review, calculate the fluctuation range, and determine the stability index; The consistency index, calibration improvement index, and stability index are converted into reliability factors according to the preset mapping rules.
[0042] In one optional embodiment, the standard deviation, correlation matrix, and information content are calculated based on the final scoring data of each expert role; based on the standard deviation, correlation matrix, and information content, objective weights are determined, including: For each evaluation dimension, the standard deviation of the scores given by each expert role is calculated. The standard deviation is used to indicate the degree of dispersion of the scores. Calculate the correlation coefficients between the scores of each evaluation dimension to form a correlation matrix, which is used to indicate the correlation between the dimensions; Based on the standard deviation and correlation matrix of the scores for each evaluation dimension, the information content of each evaluation dimension is calculated. The greater the information content, the richer the information contained in the score of that dimension. Based on the amount of information in each evaluation dimension, the objective weights are determined according to the principle that the greater the amount of information, the higher the weight. Based on objective weights, subjective weights for each expert role, and reliability factors, the final weights are output, including: Determine the integration ratio of subjective weights and objective weights; Based on the fusion ratio, the subjective and objective weights of each expert role are weighted and summed. Based on the reliability factor, the weights after weighted summation are adjusted, and the final weights are output. The higher the reliability factor, the closer the adjusted weights are to the weighted summation result. Among them, the consistency index is used to indicate the degree of consistency of scores given by different expert roles, the calibration improvement index is used to indicate the degree of improvement of scores after questioning and review, and the stability index is used to indicate the fluctuation of scores in multiple rounds of questioning and review.
[0043] It should be noted that, in order to address the potential blind spots in subjective judgment, an objective data analysis and adaptive fusion mechanism is introduced; Data preprocessing and weight calculation: Input the measured data matrix (e.g., McCabe complexity, LOC lines of code, etc. in the NASAMDP dataset); the objective weights are calculated using the CRITIC method (Criteria Importance Through Intercriteria Correlation). The CRITIC method takes into account both the comparative strength (standard deviation) of the indicators and the conflict between the indicators (correlation coefficient). Calculate reliability for objective data To measure the "quality" of the data; ; The percentage of null / missing values in the dataset for this metric; : The degree of high-frequency noise in the data (normalized variance); : The difference between the data generation time and the current evaluation time (penalty for outdated data); Thus, the objective result is also formalized into ,in It can be regarded as a degenerate fuzzy number (i.e. ); Now possessing subjectivity and objective ;Use the B value to construct dynamic gating for fusion; Conflict detection: First, use the center of gravity method to... Defuzzing yields the scalar expectation. ; Calculate the subjective and objective deviations ; Gating coefficient calculation: The confidence score is converted into attention coefficients using the Softmax function. ; ; in It is a sensing factor. when hour, It degenerates into the arithmetic mean; when It degenerates into the arithmetic mean; In this case, the side with the higher reliability is completely trusted (Winner-Take-All). Recommended settings It provides smooth yet distinctive transitions; Adaptive fusion and conflict correction: ; Special case handling: If and Both are very high (high reliability) and Significant discrepancies (strong conflicts) usually indicate that experts possess tacit knowledge not reflected in the data, or that the data reveals deep-seated patterns unnoticed by the experts. In such cases, the system will trigger an alert, recommending manual intervention, or defaulting to a maximum-pooling strategy to maintain a conservative assessment. .
[0044] In an optional embodiment, the scores given by each expert role to the evaluated object are aggregated according to the final weight, and the aggregated score result is output, including: Multiply the scores given by each expert role to the evaluated object under each evaluation dimension by the corresponding final weight; The weighted scores under each evaluation dimension are summed to obtain the comprehensive score of the evaluated object under each expert role evaluation. The comprehensive scores from each expert role are aggregated to output the final aggregated score.
[0045] Example 2, please refer to Figure 2 This invention provides a technical solution: a reliable evaluation weight allocation system based on a large model, applicable to the aforementioned reliable evaluation weight allocation method based on a large model, comprising: Document parsing unit 1 is used to construct an indicator evaluation model based on a large language model, and to parse the project artifacts based on the indicator evaluation model to obtain the original scoring data and related evidence data. Independent review unit 2 is used to generate and manage prompt templates and scoring scales for multiple expert roles; it triggers each expert role to score the evaluation object by dimension according to the prompt templates and scoring scales, and generates corresponding textual evidence; among which, the textual evidence is the explanation of the basis for each expert role's scoring. Group debate unit 3 is used to organize a round-based questioning and review process. In each round, each expert role questions based on the scores and textual evidence of other expert roles. The expert role being questioned reviews the questions and adjusts the scores and textual evidence. At the same time, the termination conditions and convergence criteria of the questioning and review process are controlled. The termination condition is that the preset number of questions and reviews is reached or the difference in scores among expert roles is less than a preset threshold. The convergence criterion is that the scores of each expert role tend to stabilize. The reliability quantification unit 4 is used to calculate the consistency index, calibration improvement index and stability index based on the final scoring data of each expert role, and to map the consistency index, calibration improvement index and stability index to the reliability factor. The weight fusion unit 5 is used to calculate the standard deviation, correlation matrix and information content based on the final scoring data of each expert role; determine the objective weight based on the standard deviation, correlation matrix and information content; and output the final weight based on the objective weight, the subjective weight of each expert role and the reliability factor; wherein, the subjective weight is the weight assigned by each expert role to each evaluation dimension based on their own experience and professional knowledge. The scoring aggregation unit 6 is used to aggregate the scores of each expert role for the evaluation object according to the final weight, and output the aggregated scoring result.
[0046] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. A reliable evaluation weight allocation method based on a large model, characterized in that, include: An indicator evaluation model is constructed based on a large language model. The project workpiece is analyzed based on the indicator evaluation model to obtain the original scoring data and related evidence data. Generate and manage prompt templates and scoring scales for multiple expert roles; trigger each expert role to score the evaluation object by dimension according to the prompt templates and scoring scales, and generate corresponding textual evidence; wherein, the textual evidence is the basis for each expert role's scoring; Organize a round-based questioning and review process; in each round, each expert role conducts questioning based on the scores and textual evidence of other expert roles, and the questioned expert role reviews and adjusts the scores and textual evidence; at the same time, control the termination conditions and convergence criteria of the questioning and review process; the termination condition is reaching a preset number of questions and reviews or the difference in scores among expert roles is less than a preset threshold, and the convergence criterion is that the scores of each expert role tend to stabilize. Based on the final scores from each expert role, consistency index, calibration improvement index, and stability index are calculated, and these indices are mapped to reliability factors. Based on the final scores from each expert role, the standard deviation, correlation matrix, and information content are calculated; based on the standard deviation, correlation matrix, and information content, the objective weights are determined; based on the objective weights, the subjective weights of each expert role, and the reliability factor, the final weights are output; wherein, the subjective weights are the weights assigned by each expert role to each evaluation dimension based on their own experience and professional knowledge. The scores given by each expert role to the evaluated object are aggregated based on the final weights, and the aggregated score result is output.
2. The method for weight allocation in a reliable evaluation based on a large model according to claim 1, characterized in that, The indicator evaluation model includes a demand analysis agent, a knowledge retrieval agent, and a fact verification agent; the original scoring data is used to represent the initial scores of different evaluation objects under each evaluation dimension, and the relevant evidence data is used to represent various types of materials in the original scoring data. The requirement parsing agent is used to parse project documents, identify the core business objectives in the project documents using hierarchical extraction technology, recursively decompose the core business objectives into a set of candidate attributes and supporting evidence types, identify and define fuzzy and lacking evidence indicators in the candidate attribute set, and generate a missing information set. The project documents include a requirement specification and an architecture design document. The knowledge retrieval agent is used to search for missing information sets in an external general knowledge base, supplement the missing information sets, obtain supplementary information, and classify the retrieval results into four levels according to the evidence reliability label: authoritative confirmation, domain consensus, expert experience, and weak inference, and assign corresponding reliability weights. The fact-verifying agent is used to check whether there are logical contradictions between the supplementary information and the project documents.
3. The method for weight allocation in a reliable evaluation based on a large model according to claim 2, characterized in that, After analyzing the project artifacts based on the aforementioned indicator evaluation model to obtain the original scoring data and relevant evidence data, the method further includes: Define the scope of standardization; For each rating in the original rating data, a linear transformation is performed according to the relative position of the rating among all ratings, based on a defined standardization range, to generate a standardized matrix; Create an index for each piece of raw scoring data to record the corresponding position of the raw scoring data in the relevant evidence data, thus forming an evidence index; The standardization matrix is used to unify the original scoring data to a specific range, and the evidence index is used to record the location information of the relevant evidence data corresponding to each original scoring data.
4. The method for weight allocation in a reliable evaluation based on a large model according to claim 3, characterized in that, The prompt template is used to guide the large model to simulate different expert roles to conduct evaluations, and the scoring scale is the standard for each expert role to score the evaluation object in different dimensions. Generate and manage prompt templates and scoring scales for multiple expert roles, including: Determine the required type of expert role based on the type of assessment object and the assessment purpose; For each type of expert role, a corresponding prompt template is generated, wherein the prompt template includes background information of the evaluation object, description of the evaluation dimensions, and guiding questions; Develop a scoring scale for each expert role type.
5. The method for weight allocation in a reliable evaluation based on a large model according to claim 4, characterized in that, This triggers each expert role to score the evaluation object according to the prompt template and scoring scale, and generates corresponding textual evidence, including: Input the prompt template and scoring scale into the large model to simulate the scoring of the evaluation object by various expert roles; The large model scores the evaluated object across various evaluation dimensions based on the guiding questions in the prompt template and the scoring scale. At the same time, the large model generates corresponding textual evidence based on the scoring criteria, explaining the reasons and basis for the scoring.
6. The method for weight allocation in a reliable evaluation based on a large model according to claim 5, characterized in that, Organize a round-robin questioning and review process, including: At the start of each round, one expert role is randomly selected as the questioner, and the other expert roles are questioned. The questioning party raises questions based on the scores given by the questioned party and written evidence; The party being questioned shall review the questions raised and make adjustments if it finds that there are problems with the scoring or the textual evidence. Repeat the above questioning and review process until the termination conditions and convergence criteria are met.
7. The method for weight allocation in a reliable evaluation based on a large model according to claim 6, characterized in that, Based on the final scores from each expert role, consistency indicators, calibration improvement indicators, and stability indicators are calculated. These indicators are then mapped to reliability factors, including: Consistency indicators are determined by calculating the correlation coefficients or differences between the scores given by different expert roles. Compare the initial score with the final score, calculate the percentage or extent of score improvement, and determine the calibration improvement indicators; Observe the fluctuations in scores given by each expert role during multiple rounds of questioning and review, calculate the fluctuation range, and determine the stability index; The consistency index, calibration improvement index, and stability index are converted into reliability factors according to the preset mapping rules.
8. The method for weight allocation in a reliable evaluation based on a large model according to claim 7, characterized in that, Based on the final scores from each expert role, calculate the standard deviation, correlation matrix, and information content; Based on standard deviation, correlation matrix, and information content, objective weights are determined, including: For each evaluation dimension, the standard deviation of the scores given by each expert role is calculated, and the standard deviation is used to indicate the degree of dispersion of the scores; Calculate the correlation coefficients between the scores of each evaluation dimension to form a correlation matrix, which is used to indicate the correlation between the dimensions; Based on the standard deviation and correlation matrix of the scores for each evaluation dimension, the information content of each evaluation dimension is calculated. The greater the information content, the richer the information contained in the score of that dimension. Based on the amount of information in each evaluation dimension, the objective weights are determined according to the principle that the greater the amount of information, the higher the weight. Based on objective weights, subjective weights for each expert role, and reliability factors, the final weights are output, including: Determine the integration ratio of subjective weights and objective weights; Based on the fusion ratio, the subjective and objective weights of each expert role are weighted and summed. Based on the reliability factor, the weights after weighted summation are adjusted, and the final weights are output. The higher the reliability factor, the closer the adjusted weights are to the weighted summation result. The consistency index is used to indicate the degree of consistency in the scores given by each expert role, the calibration improvement index is used to indicate the degree of improvement in the scores after questioning and review, and the stability index is used to indicate the fluctuation of the scores in multiple rounds of questioning and review.
9. The method for weight allocation in a reliable evaluation based on a large model according to claim 8, characterized in that, The scores given by each expert role to the evaluated object are aggregated according to the final weights, and the aggregated score results are output, including: Multiply the scores given by each expert role to the evaluated object under each evaluation dimension by the corresponding final weight; The weighted scores under each evaluation dimension are summed to obtain the comprehensive score of the evaluated object under each expert role evaluation. The comprehensive scores from each expert role are aggregated to output the final aggregated score.
10. A credibility assessment weight allocation system based on a large model, applicable to the credibility assessment weight allocation method based on a large model as described in any one of claims 1-9, characterized in that, include: The document parsing unit is used to construct an indicator evaluation model based on a large language model, and to parse the project workpiece based on the indicator evaluation model to obtain the original scoring data and related evidence data. An independent review unit is used to generate and manage prompt templates and scoring scales for multiple expert roles; it triggers each expert role to score the evaluation object by dimension according to the prompt templates and scoring scales, and generates corresponding textual evidence; wherein, the textual evidence is the basis for each expert role's scoring. The group debate unit is used to organize a round-based questioning and review process. In each round, each expert role questions based on the scores and textual evidence of other expert roles. The expert role being questioned reviews the questions and adjusts the scores and textual evidence. At the same time, the termination conditions and convergence criteria of the questioning and review process are controlled. The termination condition is that a preset number of questions and reviews is reached or the difference in scores among expert roles is less than a preset threshold. The convergence criterion is that the scores of each expert role tend to stabilize. The reliability quantification unit is used to calculate the consistency index, calibration improvement index and stability index based on the final scoring data of each expert role, and to map the consistency index, calibration improvement index and stability index to the reliability factor. The weight fusion unit is used to calculate the standard deviation, correlation matrix, and information content based on the final scoring data of each expert role; determine the objective weight based on the standard deviation, correlation matrix, and information content; and output the final weight based on the objective weight, the subjective weight of each expert role, and the reliability factor; wherein, the subjective weight is the weight assigned by each expert role to each evaluation dimension based on their own experience and professional knowledge. The scoring aggregation unit is used to aggregate the scores of each expert role for the evaluation object according to the final weight, and output the aggregated scoring result.