A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution
The network public opinion feedback evaluation method based on multi-agent collaborative self-evolution solves the problems of inconsistent evaluation results, low efficiency, and poor rule adaptability in existing technologies. It achieves high-precision, stable, and multi-dimensional evaluation results, thereby improving the timeliness and adaptability of network public opinion governance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TRS INFORMATION TECH CO LTD
- Filing Date
- 2026-04-27
- Publication Date
- 2026-07-24
AI Technical Summary
Existing methods for evaluating online public opinion feedback suffer from several drawbacks: strong subjectivity in evaluation results, insufficient stability, low processing efficiency, difficulty in accumulating structured knowledge and evolutionary capabilities, limited ability to express rules, difficulty in depicting complex semantics, poor adaptability, lack of dynamic learning and self-evolution capabilities, and difficulty in achieving multi-dimensional collaborative decision-making.
A multi-agent collaborative self-evolution method is adopted. An evaluation agent cluster is generated through a differentiated large language model to conduct parallel evaluation and cross-review. The arbitration agent adjusts the weights, and the memory agent optimizes the evaluation details to achieve multi-dimensional evaluation and rule updates.
It improved the accuracy and stability of online public opinion feedback assessment, reduced score fluctuations, enhanced the reproducibility and objectivity of assessment results, achieved continuous optimization of rules, and improved assessment efficiency and accuracy.
Smart Images

Figure CN122453205A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network public opinion feedback evaluation technology, and in particular to a network public opinion feedback evaluation method based on multi-agent collaborative self-evolution. Background Technology
[0002] In the existing technology, the existing methods for evaluating online public opinion feedback mainly include: manual evaluation and system evaluation based on rule engines.
[0003] Manual evaluation typically involves dedicated evaluators reviewing and scoring public opinion events based on established evaluation criteria. This includes assessing indicators such as response time, completeness of feedback content, appropriateness of handling measures, and standardization of information dissemination. Some organizations may also organize expert panels for comprehensive review to arrive at the final evaluation result.
[0004] The advantage of manual assessment lies in its ability to make comprehensive judgments based on actual circumstances, possessing a certain degree of flexibility and experience-based judgment. However, it has the following obvious drawbacks:
[0005] 1. The evaluation results are highly subjective and lack stability.
[0006] Differences in professional background, experience level, and risk appetite among evaluators lead to significant fluctuations in evaluation results for the same public opinion feedback. Particularly between newly hired staff and senior experts, discrepancies exist in the understanding of evaluation criteria and standards, making it difficult to guarantee consistency and reproducibility of scores.
[0007] 2. Low processing efficiency, making it difficult to support high-concurrency scenarios.
[0008] During periods of high public opinion intensity or major emergencies, the number of public opinion events surges, making it difficult for manual assessments to complete large-scale review tasks in a timely manner, which seriously affects the timeliness of public opinion governance and the efficiency of the response loop.
[0009] 3. Difficulty in accumulating structured knowledge and evolutionary capabilities
[0010] The experience accumulated during manual evaluation often remains at the individual level, lacking a systematic modeling and knowledge accumulation mechanism, and thus failing to form a sustainable and optimized evaluation model.
[0011] System evaluation methods based on rule engines typically break down the detailed rules for public opinion feedback evaluation into structured rules (such as scoring formulas, weight configurations, and time threshold judgments), and achieve automatic scoring through rule matching and logical calculations. For example, time interval rules can be set for "response timeliness"; field coverage rules can be configured for "information completeness"; and quantity threshold rules can be set for "number of posts," etc. The system uses a predefined rule base to match and calculate public opinion feedback data to generate evaluation results.
[0012] While rule-based system evaluation methods have improved evaluation efficiency to some extent, they still have the following shortcomings:
[0013] 1. Rules have limited expressive power and are difficult to describe complex semantics.
[0014] Public opinion feedback often contains a large amount of unstructured text, emotional expressions, and contextual information. Traditional rule engines, which are mainly based on explicit logical judgments, struggle to accurately identify the true attitude, risk level, and social impact of the feedback.
[0015] 2. High rule maintenance costs and poor adaptability.
[0016] The public opinion environment and ways of expression are constantly changing, and fixed rules are difficult to adapt to new communication models and styles in a timely manner. The addition, modification, and weight adjustment of rules rely on manual configuration, which is costly to maintain and results in delayed updates.
[0017] 3. Lack of dynamic learning and self-evolution capabilities
[0018] Existing rule engine systems typically lack automatic learning capabilities and cannot optimize models based on historical evaluation results, expert corrections, or actual handling effects. As a result, the evaluation system remains in a "static rule-driven" state for a long time.
[0019] 4. Difficulty in achieving multi-dimensional collaborative decision-making
[0020] Public opinion feedback assessment often involves multiple dimensions such as response efficiency, emotional control effectiveness, information transparency, and the impact of public opinion trends. A single rule system is difficult to dynamically weigh and collaboratively reason about multi-dimensional and complex indicators.
[0021] Therefore, the ability to respond to and manage public opinion events needs to be further improved. Summary of the Invention
[0022] To address the issues of low accuracy, inconsistent results, lack of dynamic iteration capabilities, and difficulty in continuous optimization of existing online public opinion feedback evaluation methods, a new online public opinion feedback evaluation method based on multi-agent collaborative self-evolution is proposed, which solves the aforementioned problems.
[0023] The technical solution of this invention is as follows:
[0024] A method for evaluating online public opinion feedback based on multi-agent collaborative self-evolution, comprising the following steps:
[0025] S1: Parallel evaluation by multiple evaluation agents: various evaluation agents are generated based on the differentiated large language model, model parameters and prompt word strategy; an evaluation agent cluster is constructed through the generated evaluation agents; the evaluation agent cluster evaluates the network public opinion feedback data according to the public opinion feedback evaluation rules and generates the first evaluation result of each evaluation agent.
[0026] S2: Cross-review by multiple evaluation agents: The evaluation agents review the first evaluation results to obtain the second evaluation results of each evaluation agent; the first evaluation results and the second evaluation results of each evaluation agent are compared to obtain the review results of each evaluation agent.
[0027] S3: Fusion of evaluation results from multiple evaluation agents: The arbitration agent adjusts the initial evaluation weights of each evaluation agent based on the review results of each evaluation agent to obtain the corrected evaluation weights; the first evaluation results are weighted and summed according to the corrected evaluation weights to generate the system evaluation result;
[0028] S4: Update the evaluation agent cluster: The memory agent extracts optimized core dimensions from long-term memory public opinion feedback data; updates the public opinion feedback evaluation rules and initial evaluation weights according to the optimized core dimensions; the optimized core dimensions include: high-frequency misjudgment roles, high-frequency misjudgment rules, new scoring dimensions, and deviation scenarios.
[0029] Preferably, before the parallel evaluation by multiple evaluation agents, the network public opinion feedback data is cleaned and formatted to obtain preprocessed network public opinion feedback data, which is then input into the evaluation agent cluster for evaluation.
[0030] Preferably, the online public opinion feedback data includes the original text, audio, video and structured metadata published on the platform; the structured metadata includes, but is not limited to: response time, handling department, handling information, dissemination platform, dissemination speed, dissemination breadth and interaction volume.
[0031] Preferably, the method for parallel evaluation by multiple evaluation agents is as follows:
[0032] S11: Construct an evaluation agent cluster: Construct a large language model configuration matrix based on the combination of large language model type, model parameters, and prompt word strategy; generate various evaluation agents based on the large language model configuration matrix; construct an evaluation agent cluster through the generated evaluation agents; the prompt word strategy includes: the role of the large language model, scoring rules and standards, relevant materials, and output dimensions.
[0033] S12: Each evaluation agent evaluates the preprocessed online public opinion feedback data according to the public opinion feedback evaluation rules, and obtains the first evaluation result of each evaluation agent; the first evaluation result includes: comprehensive score, comprehensive score basis, score of each unit's handling situation and score basis of each unit's handling situation; the public opinion feedback evaluation rules include: comprehensive score rules, score rules of the competent unit and score rules of the collaborating unit.
[0034] Preferably, the method by which the evaluation agent evaluates the preprocessed online public opinion feedback data according to the public opinion feedback evaluation rules is as follows:
[0035] A1: Obtain real-time metadata corresponding to the original text published on the platform in the online public opinion feedback data through the real-time data acquisition interface; calculate the public opinion index based on the real-time metadata;
[0036] A2: Calculate the score of the competent authority's handling of the situation based on real-time metadata, public opinion index, and public opinion feedback assessment rules, and provide the basis for the score of the competent authority;
[0037] A3: Determine the handling score of the collaborating unit and the basis for the scoring of the collaborating unit based on the online public opinion feedback data, the public opinion compliance rule library, response time, public opinion index, and handling information;
[0038] A4: Based on the public opinion feedback assessment rules, the scores of the handling by the collaborating units and the handling by the supervising unit are weighted and summed to obtain a comprehensive score; the comprehensive score basis is obtained by integrating the scoring basis of the supervising unit and the scoring basis of the collaborating units through a large model.
[0039] A5: The first assessment result is constructed by a comprehensive score, including the comprehensive score basis, the score of the competent authority's handling, the score basis of the competent authority, the score of the handling of the collaborating units, and the score basis of the collaborating units.
[0040] Preferably, the method for determining the handling score of the collaborating unit and the basis for the collaborating unit score is as follows:
[0041] B1: The evaluation agent determines the compliance rules for the corresponding posting area based on the original text and public opinion compliance rule library published by the platform, and completes the compliance judgment of the information handling based on the compliance rules to obtain the compliance result;
[0042] B2: The large model determines the handling score of the collaborating unit and the basis for the scoring of the collaborating unit based on the public opinion index, the response timeliness, the content of the handling and the compliance results.
[0043] Preferably, the public opinion index includes: public opinion heat index, public opinion sentiment index, and public opinion risk index;
[0044] The public opinion heat index is:
[0045] (1)
[0046] Where Volume represents the total amount of original text information published on the platform in the public opinion assessment data; Speed represents the speed of dissemination after being published on the platform; Spred represents the breadth of dissemination after being published on the platform; wi, wz, and ws are weighting coefficients;
[0047] The public opinion sentiment index is:
[0048] (2)
[0049] The public opinion risk index is:
[0050] (3)
[0051] Where wa and wb are weighting coefficients.
[0052] Preferably, the method for cross-review of multiple evaluation agents is as follows:
[0053] S21: In the evaluation agent cluster, the prompt words of the agent to be reviewed are extracted in a loop and used as input prompt words for other evaluation agents in the evaluation agent cluster to obtain the second evaluation result of other evaluation agents for the prompt words;
[0054] S22: Compare the first evaluation result and the second evaluation result obtained in the review evaluation agent by inputting the prompt word to obtain the review result; the review result includes: inconsistencies in scoring logic, rule reference, semantic understanding deviation and evaluation result inconsistencies between the first evaluation result and the second evaluation result.
[0055] Preferably, the method for fusing the evaluation results of the multiple evaluation agents is as follows:
[0056] S31: Evaluate the correctness of the review result through the evaluation model in the arbitration agent; if the evaluation model determines that the review result is correct, adjust the initial evaluation weight of the evaluation agent corresponding to the review result to obtain the corrected evaluation weight of the evaluation agent; the initial evaluation weight of the evaluation agent includes: the weight of the model used in the evaluation agent, the thinking weight, and the role weight.
[0057] S32: Based on the modified evaluation weights of each evaluation agent, the scores in the first evaluation results of each evaluation agent are weighted and summed to obtain the scores of each item in the system evaluation results;
[0058] S33: The scoring criteria in the first evaluation results of each evaluation agent are integrated by artificial intelligence or large model to obtain the scoring criteria for each item in the system evaluation results.
[0059] Preferably, the weighted summation method is as follows:
[0060] (4)
[0061] Where n represents the total number of evaluation agents in the evaluation agent cluster. This represents the weights of the model used by the i-th evaluating agent; This represents the thinking weight of the i-th evaluating agent; This represents the role weight of the i-th evaluating agent; This represents the score in the first evaluation result of the i-th evaluating agent.
[0062] Preferably, the method for updating the public opinion feedback evaluation rules is as follows:
[0063] S41: Data Preprocessing and Feature Extraction: Analyze the long-term memory public opinion feedback data to identify the unit roles whose scores were modified most frequently in the manually corrected results, thus obtaining high-frequency misjudged roles; analyze the long-term memory public opinion feedback data using a large model to obtain new scoring dimensions and bias scenarios;
[0064] S42: Analysis and Update Suggestions: The large model generates update suggestions based on the high-frequency misjudgment roles, high-frequency misjudgment rules, new scoring dimensions, and deviation scenarios; the large model updates the public opinion feedback evaluation details through the update suggestions, generating a draft update suggestion and new evaluation weights for each evaluation agent; the update suggestions include: adjusting the description of easily misjudged rules, adding new evaluation dimensions, and modifying the weights of the calculated scoring dimensions.
[0065] S43: Update the public opinion feedback evaluation rules: The draft update suggestion is manually reviewed; the approved draft update suggestion is used as the new public opinion feedback evaluation rules; the public opinion feedback evaluation rules in the evaluation agent cluster are replaced with the manually reviewed new public opinion feedback evaluation rules; the initial evaluation weights of each evaluation agent in the arbitration agent are replaced with the new evaluation weights.
[0066] Preferably, the manual correction results are obtained by experts manually correcting the system evaluation results based on online public opinion feedback data; the manual correction results include: comprehensive scoring rules, scoring rules of the supervising unit and scoring rules of collaborating units, modification items, and basis for modification.
[0067] Beneficial effects:
[0068] This invention belongs to the field of online public opinion feedback evaluation technology, and proposes an online public opinion feedback evaluation method based on multi-agent collaborative self-evolution. By conducting online public opinion feedback evaluation through a cluster of evaluation agents, and through cross-review among the evaluation agents within the cluster, the accuracy of online public opinion feedback evaluation is improved, and the subjective bias of a single model or single evaluator is reduced. The arbitration fusion mechanism effectively integrates the evaluation results of all evaluation agents, effectively reducing score fluctuations and improving score stability and reproducibility, avoiding the problem of unreproducible evaluation results caused by subjective evaluations of a single model or single evaluator. By using a memory agent to analyze historical data and optimize rules, continuous rule optimization is achieved, breaking through the static limitations of traditional rule engines.
[0069] The evaluation agents in the cluster employ different large models, or large models with the same structure but different parameter settings. Each agent also has a different prompting word strategy, enabling multi-faceted evaluation of online public opinion feedback data. Cross-review among the agents avoids subjective biases inherent in single models or individual evaluators, improving the objectivity and accuracy of online public opinion feedback evaluation. By setting different prompting word strategies—that is, by assigning different roles, scoring rules, and standards to the large language models—the models can score the handling of each unit from different perspectives (such as timeliness and risk sensitivity). This avoids the limitations of single models, which often have a narrow focus and cannot achieve multi-dimensional thinking, reducing the risk of single-model bias and improving the objectivity and accuracy of online public opinion feedback evaluation.
[0070] In the arbitration agent, the weights of each evaluation agent are adjusted based on the review results, and then a weighted sum is used to obtain the final system evaluation score. This effectively reduces score fluctuations and improves score stability and reproducibility. When performing the weighted summation, the weights of each evaluation agent include three dimensions: the weight of the model used, the weight of the thought process, and the weight of the role. Compared to assigning a weight to each evaluation agent in only one dimension, this approach more accurately determines the importance of each evaluation agent, achieves a comprehensive balance among multiple evaluation agents, and improves the accuracy of the system evaluation results obtained from the arbitration agent.
[0071] After the arbitration agent's results are integrated, they still need to be manually revised. This dual evaluation by both the system and humans not only improves evaluation efficiency but also enhances accuracy. Furthermore, the manual evaluation results provide a basis for revising subsequent public opinion feedback evaluation rules. This is because relying solely on the evaluation agent makes it difficult to identify problems or incompleteness in the evaluation rules. Human evaluation, however, can consider dimensions not addressed in the rules and uncover errors. Therefore, manual evaluation provides a basis for revising the rules of the memory agent, further improving the accuracy of subsequent system evaluation results.
[0072] The memory agent statistically analyzes long-term memory public opinion feedback data to obtain frequently misjudged roles, frequently misjudged rules, new scoring dimensions, and deviation scenarios. This enables a multi-dimensional assessment of the problems existing in the currently used public opinion feedback evaluation rules. Then, the large model modifies and adds rules in the public opinion feedback evaluation rules based on frequently misjudged roles, frequently misjudged rules, new scoring dimensions, and deviation scenarios. This not only achieves continuous rule optimization, breaking through the static limitations of traditional rule engines, but also ensures the adaptability, fairness, and accuracy of the updated public opinion feedback evaluation rules, while reducing the pressure of manual maintenance. Attached Figure Description
[0073] Figure 1 This is a flowchart of a network public opinion feedback evaluation method based on multi-agent collaborative self-evolution.
[0074] Figure 2 This is an architecture diagram of a network public opinion feedback evaluation method based on multi-agent collaborative self-evolution.
[0075] Figure 3 Update the process for public opinion feedback evaluation details for memory-based intelligent agents. Detailed Implementation
[0076] Example 1
[0077] like Figure 1 As shown, a network public opinion feedback evaluation method based on multi-agent collaborative self-evolution includes the following steps:
[0078] S1: Parallel evaluation by multiple evaluation agents: various evaluation agents are generated based on the differentiated large language model, model parameters and prompt word strategy; an evaluation agent cluster is constructed through the generated evaluation agents; the evaluation agent cluster evaluates the network public opinion feedback data according to the public opinion feedback evaluation rules and generates the first evaluation result of each evaluation agent.
[0079] S2: Cross-review among multiple evaluation agents: Evaluation agents review the first evaluation result to obtain the second evaluation result of each evaluation agent;
[0080] S3: Multi-evaluation agent evaluation result fusion: The arbitration agent adjusts the initial evaluation weights of each evaluation agent based on the first and second evaluation results of each evaluation agent to obtain the modified evaluation weights of each evaluation agent, and performs a weighted summation of the first evaluation results of all evaluation agents based on the modified evaluation weights to generate the system evaluation result.
[0081] S4: Improve the evaluation agent cluster: The memory agent extracts rules to optimize the core dimensions from long-term memory public opinion feedback data; updates the public opinion feedback evaluation details and initial evaluation weights according to the optimized core dimensions; the optimized core dimensions include: high-frequency misjudgment roles, high-frequency misjudgment rules, new scoring dimensions, and deviation scenarios.
[0082] Example 2
[0083] like Figure 2 As shown, a closed-loop evaluation architecture of "multi-agent collaboration—arbitration fusion—human correction—memory evolution" includes the following core technical means:
[0084] (1) Parallel evaluation mechanism of multiple evaluation agents
[0085] Construct 3-5 AI evaluation agents based on different model architectures or parameter strategies, and independently analyze and score the same public opinion feedback content under the constraints of unified public opinion feedback evaluation rules.
[0086] Different intelligent agents possess the following:
[0087] Model differentiation: different large language models or different reasoning strategies are used; the different reasoning strategies refer to allowing the large model to perform short-term thinking, medium-term thinking, long-term thinking or very long-term thinking.
[0088] Weight configuration differences: Configuring different random sampling parameters for large models, such as temperature settings, will cause the model to output diverse evaluation results for the same task, making comprehensive evaluation more complete.
[0089] Different focuses: By using different prompting word strategies, the large model is given different thinking focuses, such as time-sensitive, semantic depth, risk-sensitive, etc.
[0090] The parallel evaluation mechanism of multiple evaluation agents achieves "model diversity game" by employing different large language models or inference strategies, setting different model weights, and using different prompt word strategies. Different large language models, based on their respective training data and model weights, can complement each other in scoring from different knowledge frameworks and reasoning approaches. By setting different prompt words (such as different scoring criteria), the large models can be guided to score from different emphases, forming a multi-dimensional collaborative evaluation. The "model diversity" formed in the evaluation agent cluster reduces the risk of single-model bias and improves the comprehensiveness, accuracy, and reliability of the evaluation of public opinion feedback results.
[0091] Specifically, the input strategy for evaluating the agent is as follows:
[0092] I. Given the roles and task requirements of the large model
[0093] As a public opinion expert, you are required to calculate and evaluate the scores of the responsible units and collaborating units involved in the handling of this public opinion incident based on the materials provided by the user. Then, you need to calculate the comprehensive score according to the rules and output only the score results and corresponding analysis without any additional expansion.
[0094] II. Specific scoring rules for a given large model
[0095] 1. Overall score = (Score of the supervising unit + Sum of scores of all collaborating units) ÷ Number of participating units (including the supervising unit), round the result to the nearest integer.
[0096] 2. Supervisory unit score = (Timeliness of early warning discovery score + Total score for task assignment) ÷ 2.
[0097] 3. Collaborative unit score = Sum of scores from multiple evaluations ÷ Number of evaluations.
[0098] III. Specific scoring criteria for the given large model
[0099] (a) Overall score = (score of the supervising unit + total score of all collaborating units) ÷ number of participating units (including the supervising unit), and the result is rounded to the nearest integer.
[0100] 1. The score of the supervising unit = (score for timeliness of early warning discovery + score for accuracy of task assignment) ÷ 2; the score for accuracy of task assignment = score for timeliness of task assignment ± 1 point (add 1 point if there is no supplementary task, subtract 1 point if there is a supplementary task).
[0101] 1. Early warning detection timeliness score: assess the time it takes from information being exposed online to early warning detection.
[0102] (1) 10 points are awarded if the time taken is less than 5 minutes;
[0103] (2) 9 points are awarded if the time taken is 5-10 minutes;
[0104] (3) 8 points are awarded for taking 10-30 minutes;
[0105] (4) 7 points are awarded for taking 30-60 minutes;
[0106] (5) 6 points are awarded for taking 60-120 minutes (6 points is the baseline score);
[0107] (6) If the time exceeds 120 minutes: 1 point will be deducted from the base score for every 1-60 minutes of additional time (e.g., 5 points for 130 minutes).
[0108] (7) When the calculated public opinion risk index value exceeds the preset value and the time exceeds 30 minutes, 1 point will be deducted from the time score for every 10 minutes of time spent. For example, if the time spent is 60 minutes and the public opinion risk index is 0.8, which exceeds the preset value of 0.75, then the current early warning detection time score is 4 points.
[0109] 2. Total score for task assignment = Timeliness score for task assignment ± 1 point
[0110] 3. Task assignment timeliness
[0111] (1) 10 points are awarded if the time taken is less than 5 minutes;
[0112] (2) 9 points are awarded if the time taken is 5-10 minutes;
[0113] (3) 8 points are awarded for taking 10-30 minutes;
[0114] (4) 7 points are awarded for taking 30-60 minutes;
[0115] (5) 6 points are awarded for taking 60-120 minutes (6 points is the baseline score);
[0116] (6) If the time exceeds 120 minutes: 1 point will be deducted from the base score for every 1-60 minutes of additional time (e.g., 5 points for 130 minutes).
[0117] (7) If there is no supplementary task, it means that the task was accurately assigned, and 1 point will be added; if there is a supplementary task, it may mean that the initial task was not accurate enough, and 1 point will be deducted.
[0118] (ii) Collaborative unit score = Sum of scores from multiple evaluations ÷ Number of evaluations.
[0119] 1. Each evaluation agent analyzes and scores the handling of the collaborative units from different perspectives, obtaining the score and AI analysis of each collaborative unit's feedback for each instance; the scores and AI analysis of each feedback obtained from the analysis are saved in (iv) related materials. The perspectives include: feedback efficiency and the accuracy of handling information.
[0120] When evaluating feedback efficiency, the public opinion risk index of the current platform's posts should be considered. If the public opinion risk index value exceeds the preset value, the score should be reduced.
[0121] 2. Evaluate the agent's existing evaluations of each collaborating unit in the analysis of relevant materials. If a unit has multiple evaluations, the evaluation agent needs to combine the scores corresponding to each feedback with the AI analysis to form a comprehensive scoring basis.
[0122] IV. Relevant Materials
[0123] <Content>
[0124] (I) Task Title: Netizens' Feedback on Problems with the Renovation Project of Middle School A
[0125] (ii) Supervising unit
[0126] 1. Warning detection time: 25 minutes and 19 seconds
[0127] 2. Time taken to handle public opinion: 15 minutes
[0128] 3. Are there any supplementary tasks? No.
[0129] (III) Collaborating Units: Collaborating Unit A
[0130] First feedback
[0131] Score: 6 points
[0132] AI analysis: 6 points. The feedback was too brief and lacked detailed explanations of the specific circumstances of the incident and the measures taken. Although the feedback took far less than 10 minutes, demonstrating a rapid response, as the "first" feedback, it only mentioned "transferred to relevant departments for verification" without providing specific time, parties involved, or follow-up plans, resulting in incomplete information.
[0133] Second feedback
[0134] Score: 7 points
[0135] AI analysis: 7 points. The feedback was quick, meeting the criteria for bonus points and demonstrating efficient rapid response. Although brief, the feedback clearly identified the responsible department and the appropriate measures, fulfilling the basic requirement of multiple feedbacks, and no further points were deducted.
[0136] Third feedback
[0137] Score: 5 points
[0138] AI Analysis: 5 points. Multiple feedback reports from the responding unit were received, but the feedback lacked basic information, leading to the deduction of points. The feedback primarily suggested transferring the project to collaborating unit D for handling, but failed to provide details about the current situation or appropriate measures, and lacked a comprehensive analysis of public opinion and specific response plans.
[0139] (iv) Collaborating Unit: Collaborating Unit B
[0140] First feedback
[0141] Score: 7 points
[0142] AI Analysis: 7 points. Collaborating Unit B's feedback was excellent in terms of detail and handling measures, but points were deducted due to the lengthy feedback time. Although the feedback exceeded the time limit, it provided a detailed explanation of the project background and the school's responsibilities, and actively communicated with the construction unit, demonstrating its attention to public opinion and proactive response.
[0143] The feedback detailed the project's specific timeline, the parties involved, and the current status of the incident. It clearly defined the school's responsibilities as the end-user and promptly contacted the construction contractor, requesting their intervention to verify and resolve the issue. This proactive communication and clear division of responsibilities facilitated the rapid identification and resolution of the problem, demonstrating a high level of public opinion management capability. Although the feedback was time-consuming, its depth and relevance compensated for the time constraint, ultimately achieving an effective response to public opinion.
[0144] (v) Collaborating Unit: Collaborating Unit C
[0145] First feedback
[0146] Score: 7 points
[0147] AI Analysis: 7 points. The feedback was relatively detailed, providing basic information about the event, including the specific project name, time, and parties involved. Although the feedback was completed within the time limit, it failed to detail the specific handling of the public opinion issue, hence the slight bonus points on top of the baseline score.
[0148] V. Limiting the Output Dimensions and Format of Large Models
[0149] (a) Output Dimension
[0150] 1. First, give the score for each section, then describe the scoring criteria.
[0151] 2. When outputting the comprehensive scoring criteria and scoring basis, the evaluation should be carried out from a business perspective, and should not simply reflect the addition or subtraction of scores or calculation formulas.
[0152] 3. Do not mention rules that do not affect the score, do not elaborate on the rules, do not give suggestions, and do not make summaries.
[0153] 4. The entire text should be kept within 300 words, with clear logic and concise wording.
[0154] (ii) Output according to the format below
[0155] Overall score: X points
[0156] Overall scoring criteria: XXXX
[0157] Supervisory Unit: X points, XXXX
[0158] XXXX[Collaborating Unit 1]: X points, XXXX
[0159] XXXX[Collaborative Unit 2]: X points, XXXX
[0160] (III) Output Examples of Large Models
[0161] Overall rating: 8 points
[0162] Overall evaluation criteria: While all units demonstrated a positive attitude during this public opinion crisis management process, there is still room for improvement in information transparency and the details of the handling. The rapid response and accurate task allocation by the supervising unit laid the foundation for subsequent handling, while the differences in information sharing and problem-solving among collaborating units highlighted the importance of multi-departmental collaboration.
[0163] Supervisory Unit: 8 points. The supervisory unit demonstrated high timeliness in early warning and task assignment during this public opinion crisis. The early warning was detected within 25 minutes and 19 seconds, and the task assignment was completed within 15 minutes. This efficient early warning and task allocation ensured that relevant departments could intervene quickly and effectively control the spread of public opinion.
[0164] Collaborating Unit A: 6 points. Collaborating Unit A demonstrated a rapid response, but information transparency needs improvement. The first feedback only mentioned "transferred to relevant departments for verification," lacking specific timelines and involved parties, making the information too brief. The second feedback clearly identified the competent department and handling measures, but the content remained brief. The third feedback suggested transferring the project to Collaborating Unit D for handling, but did not elaborate on the current situation or handling measures, lacking a comprehensive analysis of public opinion.
[0165] Collaborating Unit B: 7 points. Collaborating Unit B demonstrated a high level of responsibility and initiative in handling this public opinion incident; however, the lengthy feedback process affected the efficiency of information dissemination. The feedback detailed the project background and the school's responsibilities, and actively communicated with the construction contractor to promote problem-solving. This proactive communication and clear division of responsibilities helps to quickly locate the problem and drive its resolution; however, the lengthy feedback process may affect the timeliness of information dissemination.
[0166] (2) Cross-Review Mechanism
[0167] Each evaluation agent reviews the evaluation results of other agents, pointing out the following issues: inconsistencies in scoring logic, incorrect rule citations, semantic comprehension biases, and the risk of scoring too high or too low.
[0168] Specifically, the cross-review mechanism involves inputting the prompts from evaluation agent A into evaluation agents B and C to review the results. Because different evaluation agents use different large-scale model settings and have different inference logic, if the outputs of multiple evaluation agents are consistent (i.e., high similarity), it indicates that the accuracy and reliability of the first evaluation result from evaluation agent A are high. If discrepancies occur, it indicates that the evaluation result from evaluation agent A is incorrect. This avoids errors in the evaluation results of a single evaluation agent from affecting the accuracy of the final evaluation result.
[0169] The cross-review mechanism forms a peer review system for intelligent agents, which structurally reduces the probability of misjudgment by individual models.
[0170] (3) Arbitration Agent Fusion Mechanism
[0171] An arbitration agent is constructed; the arbitration agent adjusts the weights of each evaluation agent by integrating the original scores of each evaluation agent, the review opinions between each agent, and the results of score difference analysis; finally, a multi-dimensional weight fusion algorithm is used to generate the final evaluation version.
[0172] Specifically, the steps for the fusion of arbitration agents are as follows:
[0173] The first step is to evaluate the correctness of the second evaluation result through the evaluation model in the arbitration agent; if the evaluation model determines that the evaluation of the second evaluation result is correct, the initial evaluation weight of the evaluation agent corresponding to the first evaluation result is reduced to obtain the corrected evaluation weight of the evaluation agent; the initial evaluation weight includes: the weight of the model used in the evaluation agent, the thinking weight, and the role weight.
[0174] First evaluation weight = Original evaluation weight × (1 - Penalty coefficient)
[0175] The penalty coefficient is set based on the degree of difference between the scores of the first and second assessment results in the second assessment results;
[0176] The second step is to perform a weighted sum of the scores in the first evaluation results of each evaluation agent according to the modified evaluation weights of each evaluation agent, so as to obtain the scores of each item in the system evaluation results.
[0177] Specifically, the method for obtaining the score through multi-dimensional weight fusion is as follows: (4)
[0178] Where n represents the total number of evaluation agents in the evaluation agent cluster. This represents the weights of the model used by the i-th evaluating agent; This represents the thinking weight of the i-th evaluating agent; This represents the role weight of the i-th evaluating agent; This represents the score in the first evaluation result of the i-th evaluating agent.
[0179] Specifically, the weights of the evaluation agent are set according to the size of the model; the thinking weights of the evaluation agent are set according to the reasoning strategy selected by the large model; and the role weights of the evaluation agent are determined according to whether the role selection of the large model is for experts from the supervising unit or experts from collaborating units.
[0180] The third step involves synthesizing the analysis of the first evaluation results of each evaluation agent by humans or large models to obtain the score analysis of each item in the system evaluation results.
[0181] Specifically, the comprehensive evaluation results are as follows:
[0182] Overall rating: 8 points
[0183] Overall evaluation criteria: While all units demonstrated a positive attitude during this public opinion crisis management process, there is still room for improvement in information transparency and the details of the handling. The rapid response and accurate task allocation by the supervising unit laid the foundation for subsequent handling, while the differences in information sharing and problem-solving among collaborating units highlighted the importance of multi-departmental collaboration.
[0184] Supervisory Unit: 8 points. The supervisory unit demonstrated high timeliness in early warning and task assignment during this public opinion crisis. The early warning was detected within 25 minutes and 19 seconds, and the task assignment was completed within 15 minutes. This efficient early warning and task allocation ensured that relevant departments could intervene quickly and effectively control the spread of public opinion.
[0185] Collaborating Unit A: 6 points. Collaborating Unit A demonstrated a rapid response, but information transparency needs improvement. The first feedback only mentioned "transferred to relevant departments for verification," lacking specific timelines and involved parties, making the information too brief. The second feedback clearly identified the competent department and handling measures, but the content remained brief. The third feedback suggested transferring the project to Collaborating Unit D for handling, but did not elaborate on the current situation or handling measures, lacking a comprehensive analysis of public opinion.
[0186] Collaborating Unit B: 7 points. Collaborating Unit B demonstrated a high level of responsibility and initiative in handling this public opinion incident; however, the lengthy feedback process affected the efficiency of information dissemination. The feedback detailed the project background and the school's responsibilities, and actively communicated with the construction contractor to promote problem-solving. This proactive communication and clear division of responsibilities helps to quickly locate the problem and drive its resolution; however, the lengthy feedback process may affect the timeliness of information dissemination.
[0187] (4) Artificially Interventional Correction Mechanism
[0188] Human intervention is permitted to modify and optimize the final evaluation version. Manual modifications will be recorded in a structured manner, including: the modification item, the reason for the modification, the original evaluation agent cluster's score, and the final human score.
[0189] The manually modified assessment results are saved to the long-term memory module for subsequent evolutionary analysis and updates to the public opinion feedback assessment rules.
[0190] (5) Memory-based intelligent agents and rule self-evolution mechanism
[0191] like Figure 3 As shown, a Memory Agent is set up to analyze the data in the long-term memory module on a weekly basis. The analysis includes: statistical analysis of the differences between the model evaluation version and the manually corrected version, identification of high-frequency misjudgment rules, attribution analysis of contextual bias, and identification of new public opinion features, i.e., new evaluation dimensions.
[0192] Specifically, the method of memory-based intelligent agents and rule self-evolution is as follows:
[0193] The first step is data preprocessing and feature extraction: The unit roles with the most score modifications in the manually corrected results are identified to obtain high-frequency misjudged roles; the rules with the most modified references in the manually corrected results are identified to obtain high-frequency misjudged rules; and the large-scale model analyzes the unknown labels in the long-term memory public opinion feedback data to obtain new scoring dimensions and bias scenarios.
[0194] Specifically, the long-term memory public opinion feedback data stored in the long-term memory module includes: complete contextual information of a single assessment, including: real-time data obtained through external tools, dissemination data, policies and regulations, multi-round feedback information, as well as manual assessment results and system assessment results.
[0195] Specifically, the large model extracts evaluation dimensions from the manually corrected results and handling dimensions from the handling information of each unit, such as sensitivity and unit coordination, from the long-term memory public opinion feedback data. The extracted handling dimensions are compared with the public opinion feedback evaluation rules to obtain evaluation dimensions that are not included in the public opinion feedback evaluation rules, thus obtaining new evaluation dimensions.
[0196] For example, prompts for new evaluation dimensions of a large model could be:
[0197] Input: Manually corrected results from long-term memory public opinion feedback data, handling information from various units, and detailed rules for evaluating public opinion feedback used;
[0198] Required analytical operations for large models:
[0199] 1) Extract the evaluation dimensions used in the manual evaluation from the results of manual correction, such as the timeliness of the handling, the suitability of the handling measures, and the sensitivity to public opinion;
[0200] 2) Extract handling dimensions from the handling information of each unit, such as the efficiency of coordination and linkage between units, the accuracy of cross-departmental resource allocation, and the consistency of emergency response decisions;
[0201] 3) Integrate the assessment dimensions and the treatment dimensions to obtain the first assessment dimension set;
[0202] 4) Compare the first set of evaluation dimensions with the dimensions in the public opinion feedback evaluation rules to obtain the second set of evaluation dimensions that are not in the public opinion feedback evaluation dimensions but are in the first set of evaluation dimensions;
[0203] Output: Output the set of the second evaluation dimensions.
[0204] Specifically, the deviation scenario refers to the actual assessment deviating from the expected target or standard rules, which leads to multiple manual modifications and makes accurate assessment impossible. Therefore, a large model is used to extract the deviation scenario in the assessment process from the reasons for manual modifications, thereby supplementing the missing assessment dimensions in the assessment details. The deviation scenario includes, for example, information asymmetry in multi-unit collaboration or insufficient risk sensitivity.
[0205] The second step is analysis and update recommendations: The large model generates update recommendations based on the high-frequency misjudgment roles, high-frequency misjudgment rules, new scoring dimensions, and deviation scenarios; the large model updates the public opinion feedback evaluation details through the update recommendations, generating a draft update recommendation and new evaluation weights for each evaluation agent; the update recommendations include: adjusting the description of easily misjudged rules, adding new evaluation dimensions, modifying the weights of the calculated scoring dimensions, and the evaluation weights for each evaluation agent.
[0206] By setting large model prompts, the system can generate update suggestions and update the detailed rules for evaluating public opinion feedback.
[0207] The goal of setting up a large model is to optimize core dimensions (including high-frequency misjudged roles, high-frequency misjudged rules, new scoring dimensions, and deviation scenarios) based on the input rules. First, the model generates update suggestions by optimizing the core dimensions according to the rules. Then, it updates the public opinion feedback evaluation details based on the update suggestions to obtain a draft of the update suggestions.
[0208] Inputs to the large model: high-frequency misjudged roles, high-frequency misjudgment rules, new scoring dimensions and bias scenarios obtained in the first step, and detailed rules for public opinion feedback evaluation;
[0209] Output of the large model: Output update suggestions, and the specific modifications made in the update suggestions and the basis for the modifications.
[0210] The third step is to update the public opinion feedback evaluation rules: the draft update proposal is manually reviewed; the approved draft update proposal is used as the new public opinion feedback evaluation rules; the public opinion feedback evaluation rules in the evaluation agent cluster are replaced with the manually reviewed new public opinion feedback evaluation rules; and the weights of each evaluation agent in the arbitration agent are replaced with the new evaluation weights.
[0211] The large model in the memory agent automatically generates updated suggestions for public opinion feedback assessment rules based on the results of analyzing data in the long-term memory module, and submits them for manual review and confirmation. After manual confirmation, the version of the public opinion feedback assessment rules is upgraded, forming a closed-loop evolutionary system of "rules-model-data".
[0212] (6) External tool invocation mechanism (MCP invocation capability)
[0213] Each agent (such as the evaluation agent and the arbitration agent) supports calling external tool interfaces (ModelCapability Platform, MCP), including:
[0214] Real-time data acquisition interface: When the evaluation agent is conducting public opinion feedback evaluation, it is used to acquire online public opinion feedback information and other related real-time posting data;
[0215] Public opinion dissemination trend analysis tool: When the evaluation agent is conducting public opinion feedback evaluation, it is used to obtain the real-time online dissemination of the original text corresponding to the online public opinion feedback information;
[0216] Data statistics and calculation script execution: When the evaluation agent is conducting public opinion feedback evaluation, it calculates the public opinion index based on real-time network dissemination and dissemination analysis data; the public opinion index calculation involves:
[0217] 1. Public opinion heat index (Hotness) reflects the degree of attention an event receives.
[0218] Calculation formula:
[0219] (1)
[0220] Where Volume represents the total amount of information, Speed represents the speed of propagation, Spreda represents the breadth of propagation, and wi, wz, and ws are weights.
[0221] 2. Sentiment Index reflects public opinion trends.
[0222] Calculation formula:
[0223] (2)
[0224] Note: Results range from -1 to 1. The closer the result is to -1, the more critical the public opinion situation.
[0225] 3. Public Opinion Risk Index (Risk): The overall score is based on the degree of harm.
[0226] Calculation formula: (3)
[0227] The higher the Risk score, the greater the popularity and the stronger the negative sentiment, and the higher the risk index.
[0228] Policy and Regulation Retrieval System: When the evaluation agent is conducting public opinion feedback evaluation, it queries the compliance rules corresponding to the original posting field based on the posting field of the original posting text; the large model judges whether the operation in the feedback information is compliant based on the compliance rules.
[0229] Semantic enhancement analysis tool: When the evaluation agent is conducting public opinion feedback evaluation, it performs semantic enhancement on the original post text to find post information related to the original post text; this post information is used as the basis for the evaluation agent's evaluation. This post information serves as enhancement for the retrieval content parameters of the "real-time data acquisition interface" and the policy and regulation retrieval system mentioned above.
[0230] The assessment agent, through an external tool invocation mechanism, obtains external knowledge and real-time dissemination and handling status corresponding to the original text, enabling more accurate assessment of online public opinion feedback.
[0231] Example 3
[0232] A closed-loop evaluation framework based on a multi-agent system. Its core logic lies in eliminating the randomness bias of a single model through game theory and mutual review among models, and utilizing the concept of long short-term memory to achieve automatic iteration of evaluation rules through a memory agent.
[0233] (I) Method Implementation: A Public Opinion Assessment Method Based on Multi-Agent Collaborative Self-Evolution
[0234] Based on the system flowchart, the specific steps of this embodiment are as follows:
[0235] 1. Data Input and Preprocessing
[0236] The system receives public opinion feedback data to be evaluated. This data includes not only raw text but also structured metadata, such as response time, handling department, dissemination platform, and interaction volume. The input module cleans and formats the data to ensure it conforms to the input specifications of the subsequent intelligent agent. The raw text includes text content, images, and videos.
[0237] Specifically, an example of the formatted content is as follows:
[0238]
[0239] 2. Parallel evaluation by multiple evaluation agents
[0240] The system schedules 3 to 5 evaluation agents with differentiated configurations.
[0241] Implementation: Each agent loads a different Large Language Model (LLM) platform (such as GPT-4, Claude 3, Llama 3, etc.), or uses the same platform but is configured with different prompt strategies.
[0242] Differentiated focus: For example, assessment agent A focuses on semantic sensitivity, assessment agent B focuses on compliance review, and assessment agent C focuses on propagation risk prediction. This differentiated focus is achieved by setting different prompt word strategies.
[0243] Specifically, by introducing "model diversity game", the scoring "illusion" or systematic bias caused by training data bias of a single model is effectively avoided.
[0244] 3. Cross-review mechanism
[0245] After generating preliminary results, each evaluation agent needs to exchange results for "peer review".
[0246] Implementation method: Agent A reads Agent B's rating reasons, checks whether it has misread a certain clause in the public opinion feedback evaluation rules, and generates "agree" or "disagree" opinions.
[0247] Specifically, by simulating the consultation mechanism of human experts, logical contradictions or obvious semantic misunderstandings are corrected within the system first.
[0248] 4. Arbitration Integration and Initial Output
[0249] The arbitration agent collects all original scores and cross-review opinions.
[0250] Implementation method: A weighted fusion algorithm is used. If a certain agent's score is challenged by a majority of other agents with sufficient reasons, the arbitrator agent will lower the weight of that agent.
[0251] Specifically, the large model used by the arbitration agent is an evaluation model trained through fine-tuning. Therefore, the arbitration agent can effectively judge whether the agent's review results are accurate and whether the reasons are sufficient.
[0252] Enhanced capabilities: In this process, intelligent agents can call external tools (such as real-time public opinion heat query API and legal database) through the capability platform to provide factual basis for arbitration.
[0253] 5. Manual correction and data structuring
[0254] The system outputs the arbitration-integrated evaluation version. Then, it enters the manual intervention process, where experts can make corrections through an interactive interface. In the interface, experts evaluate the system's evaluation data based on public opinion feedback, and then indicate their agreement or disagreement; if they disagree, they provide specific reasons and make modifications.
[0255] Key point: The system not only records the final corrected score, but also extracts the "reason for modification" through natural language processing technology (e.g., the rule was interpreted too strictly, or a specific context was ignored).
[0256] 6. Memory Evolution and Rule Upgrade
[0257] The memory agent scans the historical database periodically (e.g., weekly).
[0258] Implementation method: Compare the differences between the "system arbitration version" and the "human final version". If it is found that in a certain type of public opinion (e.g., involving specific industry terms), the human frequently lowers the score, the memory agent will identify the ambiguities in the description of the public opinion feedback evaluation rules.
[0259] Self-evolution process: New rule description suggestions are generated and submitted to the administrator for review. Upon approval, the public opinion feedback evaluation rule library is updated, and all agents will follow these rules in the next cycle. Version running.
[0260] It should be noted that the specific embodiments described above enable those skilled in the art to more fully understand the present invention, but do not limit the present invention in any way. Therefore, although the present invention has been described in detail with reference to the accompanying drawings and embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention. In short, all technical solutions and improvements that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention patent.
Claims
1. A network public opinion feedback evaluation method based on multi-agent collaborative self-evolution, characterized in that, S1: Parallel evaluation by multiple evaluation agents: various evaluation agents are generated based on the differentiated large language model, model parameters and prompt word strategy; an evaluation agent cluster is constructed through the generated evaluation agents; the evaluation agent cluster evaluates the network public opinion feedback data according to the public opinion feedback evaluation rules and generates the first evaluation result of each evaluation agent. S2: Cross-review by multiple evaluation agents: The evaluation agents review the first evaluation results to obtain the second evaluation results of each evaluation agent; the first evaluation results and the second evaluation results of each evaluation agent are compared to obtain the review results of each evaluation agent. S3: Fusion of evaluation results from multiple evaluation agents: The arbitration agent adjusts the initial evaluation weights of each evaluation agent based on the review results of each evaluation agent to obtain the corrected evaluation weights; The system evaluation results are generated by weighting and summing the results of each first evaluation based on the revised evaluation weights. S4: Update the evaluation agent cluster: The memory agent extracts and optimizes the core dimensions from long-term memory public opinion feedback data; The public opinion feedback evaluation rules and initial evaluation weights are updated according to the core optimization dimensions. The core optimization dimensions include: high-frequency misjudgment roles, high-frequency misjudgment rules, new scoring dimensions, and deviation scenarios.
2. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 1, characterized in that, Before the parallel evaluation by multiple evaluation agents, the network public opinion feedback data is cleaned and formatted to obtain preprocessed network public opinion feedback data, which is then input into the evaluation agent cluster for evaluation.
3. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 2, characterized in that, The online public opinion feedback data includes the original text, audio, video and structured metadata published on the platform; the structured metadata includes, but is not limited to: response time, handling department, handling information, dissemination platform, dissemination speed, dissemination breadth and interaction volume.
4. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 3, characterized in that, The method for parallel evaluation by multiple evaluation agents is as follows: S11: Construct an evaluation agent cluster: Construct a large language model configuration matrix based on the combination of large language model type, model parameters and prompt word strategy, and generate various evaluation agents based on the large language model configuration matrix; An evaluation agent cluster is constructed using the various generated evaluation agents; the prompt word strategy includes: the role of the large language model, scoring rules and standards, relevant materials, and output dimensions; S12: Each evaluation agent evaluates the preprocessed online public opinion feedback data according to the public opinion feedback evaluation rules, and obtains the first evaluation result of each evaluation agent; the first evaluation result includes: comprehensive score, comprehensive score basis, score of each unit's handling situation and score basis of each unit's handling situation; the public opinion feedback evaluation rules include: comprehensive score rules, score rules of the competent unit and score rules of the collaborating unit.
5. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 4, characterized in that, The method by which the evaluation agent evaluates the preprocessed online public opinion feedback data according to the public opinion feedback evaluation rules is as follows: A1: Obtain real-time metadata corresponding to the original text published on the platform in the online public opinion feedback data through the real-time data acquisition interface; calculate the public opinion index based on the real-time metadata; A2: Calculate the score of the competent authority's handling of the situation based on real-time metadata, public opinion index, and public opinion feedback assessment rules, and provide the basis for the score of the competent authority; A3: Determine the handling score of the collaborating unit and the basis for the scoring of the collaborating unit based on the online public opinion feedback data, the public opinion compliance rule library, response time, public opinion index, and handling information; A4: Based on the public opinion feedback assessment rules, the scores of the handling by the collaborating units and the handling by the supervising unit are weighted and summed to obtain a comprehensive score; the comprehensive score basis is obtained by integrating the scoring basis of the supervising unit and the scoring basis of the collaborating units through a large model. A5: The first assessment result is constructed by a comprehensive score, including the comprehensive score basis, the score of the competent authority's handling, the score basis of the competent authority, the score of the handling of the collaborating units, and the score basis of the collaborating units.
6. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 5, characterized in that, The method for determining the handling score of the collaborating unit and the basis for the collaborating unit score is as follows: B1: The evaluation agent determines the compliance rules for the corresponding posting area based on the original text and public opinion compliance rule library published by the platform, and completes the compliance judgment of the information handling based on the compliance rules to obtain the compliance result; B2: The large model determines the handling score of the collaborating unit and the basis for the scoring of the collaborating unit based on the public opinion index, the response timeliness, the content of the handling and the compliance results.
7. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 5, characterized in that, The public opinion indices include: public opinion heat index, public opinion sentiment index, and public opinion risk index; The public opinion heat index is: (1) Where Volume represents the total amount of original text information published on the platform in the public opinion assessment data; Speed represents the speed of dissemination after being published on the platform; Spred represents the breadth of dissemination after being published on the platform; wi, wz, and ws are weighting coefficients; The public opinion sentiment index is: (2) The public opinion risk index is: (3) Where wa and wb are weighting coefficients.
8. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 1, characterized in that, The method for cross-review of multiple evaluation agents is as follows: S21: In the evaluation agent cluster, the prompt words of the agent to be reviewed are extracted in a loop and used as input prompt words for other evaluation agents in the evaluation agent cluster to obtain the second evaluation result of other evaluation agents for the prompt words; S22: Compare the first evaluation result and the second evaluation result obtained in the review and evaluation agent by comparing the prompt word input to obtain the review result; The review results include: inconsistencies in scoring logic, rule referencing, semantic understanding, and evaluation results between the first and second evaluation results.
9. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 1, characterized in that, The method for fusing the evaluation results of the multiple evaluation agents is as follows: S31: Evaluate the correctness of the review result through the evaluation model in the arbitration agent; if the evaluation model determines that the review result is correct, adjust the initial evaluation weight of the evaluation agent corresponding to the review result to obtain the corrected evaluation weight of the evaluation agent. The initial evaluation weights of the evaluation agent include: the weight of the model used by the evaluation agent, the weight of thinking, and the weight of the role; S32: Based on the modified evaluation weights of each evaluation agent, the scores in the first evaluation results of each evaluation agent are weighted and summed to obtain the scores of each item in the system evaluation results; S33: The scoring criteria in the first evaluation results of each evaluation agent are integrated by artificial intelligence or large model to obtain the scoring criteria for each item in the system evaluation results.
10. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 9, characterized in that, The weighted summation method is as follows: (4) Where n represents the total number of evaluation agents in the evaluation agent cluster. This represents the weights of the model used by the i-th evaluating agent; This represents the thinking weight of the i-th evaluating agent; This represents the role weight of the i-th evaluating agent; This represents the score in the first evaluation result of the i-th evaluating agent.
11. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 1, characterized in that, The method for updating the public opinion feedback assessment details is as follows: S41: Data Preprocessing and Feature Extraction: Analyze the long-term memory public opinion feedback data to identify the unit roles whose scores were modified most frequently in the manually corrected results, thus obtaining high-frequency misjudged roles; analyze the long-term memory public opinion feedback data using a large model to obtain new scoring dimensions and bias scenarios; S42: Analysis and Update Suggestions: The large model generates update suggestions based on the high-frequency misjudged roles, high-frequency misjudgment rules, new scoring dimensions, and bias scenarios. The large model updates the public opinion feedback evaluation rules based on the update suggestions, generating a draft update suggestion and new evaluation weights for each evaluation agent; the update suggestions include: adjusting the description of rules prone to misjudgment, adding new evaluation dimensions, and modifying the weights of the scoring dimensions; S43: Update the public opinion feedback evaluation rules: The draft update suggestion is manually reviewed; the approved draft update suggestion is used as the new public opinion feedback evaluation rules; the public opinion feedback evaluation rules in the evaluation agent cluster are replaced with the manually reviewed new public opinion feedback evaluation rules; the initial evaluation weights of each evaluation agent in the arbitration agent are replaced with the new evaluation weights.
12. A network public opinion feedback evaluation method based on multi-agent cooperative self-evolution as described in claim 11, characterized in that, The manual correction results were obtained by experts manually correcting the system evaluation results based on online public opinion feedback data. The manual correction results include: comprehensive scoring rules, scoring rules of the supervising unit and scoring rules of collaborating units, modification items, and the basis for modification.