A government affair reply quality evaluation method
The government response quality assessment method, which employs multi-dimensional feature extraction, clustering, and iterative threshold adjustment, addresses the issues of single feature dimensions and inconsistent assessments in existing technologies, thereby achieving in-depth characterization and efficient evaluation of government response quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 湖南工商大学
- Filing Date
- 2026-06-10
- Publication Date
- 2026-07-17
AI Technical Summary
Existing methods for assessing the quality of government responses suffer from limitations such as a single feature dimension, global standardization leading to assessment distortion, and a lack of adaptability and consistency verification, resulting in poor assessment efficiency and accuracy.
A large language model is used for multi-dimensional feature extraction. After clustering, Z-scores are standardized. A cross-cluster bidirectional cross-comparison and a three-level priority hybrid comparison are constructed. Combined with an iterative threshold adjustment mechanism, a machine learning model is trained for quality assessment.
It enables in-depth characterization of the quality of government responses, eliminates distribution bias between clusters, ensures the consistency and adaptability of evaluation standards, and improves the accuracy and reliability of evaluation.
Smart Images

Figure CN122414569A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method for evaluating the quality of government response. Background Technology
[0002] Currently, against the backdrop of the deepening digital transformation of government services, government responses, as a core link between the government and the public, directly impact the level of government services and the government's credibility. With the development of large language model technology, automated text evaluation has become possible. However, existing methods for evaluating the quality of government responses suffer from the following significant technical shortcomings: (1) Single feature dimension and lack of depth: Existing methods mostly stay at the surface semantic analysis and cannot accurately quantify the performance of government responses in terms of deep government logic such as "specificity, relevance, information richness, appropriateness and readability"; (2) Global standardization leads to distorted evaluation: Existing solutions often use global standardization when processing large-scale government data. Due to the inherent differences in the difficulty of responding to different types of questions and the benchmark, global standardization will seriously mask the true quality distribution within clusters, leading to an imbalance in the classification of levels; (3) Lack of global consistency verification and adaptive correction mechanism: Existing evaluation models often "evaluate to the end" without verifying the rationality of the evaluation results. Especially between different topic clusters, it is easy to see the phenomenon of "quality inversion" where "the actual quality of high-scoring samples in low-difficulty clusters is lower than that of low-scoring samples in high-difficulty clusters", and the existing system does not have the ability to adaptively adjust the threshold to correct this defect.
[0003] It is evident that there is an urgent need for a method to assess the quality of government responses that is efficient, accurate, and adaptable. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a method for evaluating the quality of government responses, which at least partially solves the problems of poor evaluation efficiency, accuracy and adaptability in the prior art.
[0005] In a first aspect, embodiments of the present invention provide a method for evaluating the quality of government response, including: Step 1: Use a large language model to extract features from the acquired government response data to obtain multi-dimensional features; Step 2: Cluster the government response data according to question type and topic to obtain multiple clusters; Step 3: Within each cluster, quality assessment is performed using a large language model to obtain raw scores, and Z-score standardization is performed independently to convert the raw scores into in-cluster Z-scores. Based on a standard deviation principle, the in-cluster Z-scores are classified into quality levels. Step 4: Construct cross-cluster bidirectional cross-comparison pairs, wherein the bidirectional cross-comparison pairs include cross-comparisons of the best and worst samples, and a three-level priority hybrid comparison method is used to determine the rationality of the bidirectional cross-comparison pairs; Step 5: For clusters with unreasonable conditions, use an iterative threshold adjustment mechanism to calculate the threshold adjustment amount to update the quality level division threshold, and re-divide the quality levels until all bidirectional cross-comparison pairs are reasonable. Step 6: Use multi-dimensional features as input and manually labeled quality levels as tags to train a machine learning model to output the quality assessment results of government response.
[0006] According to a specific implementation of an embodiment of the present invention, the multi-dimensional feature is a five-dimensional feature, including specificity feature, relevance feature, information richness feature, appropriateness feature, and readability feature.
[0007] According to a specific implementation of the present invention, the expression for the intra-cluster Z-score is: ; in, Indicates the first The first cluster The sample at the th Intra-cluster Z-scores on each metric Indicates the first The first cluster The sample at the th The raw scores on each indicator Indicates the first The first cluster The mean of each indicator within the cluster Indicates the first The first cluster The standard deviation of each indicator within the cluster.
[0008] According to a specific implementation of an embodiment of the present invention, the quality level classification of the intra-cluster Z-score based on a standard deviation principle is specifically as follows: When the Z-score within a cluster is less than negative one standard deviation, it is considered low quality; When the in-cluster Z-score is between negative one standard deviation and positive one standard deviation, it is considered to be of medium quality; When the Z-score within a cluster is greater than one standard deviation, it is considered to be of high quality.
[0009] According to a specific implementation of an embodiment of the present invention, the construction of cross-cluster bidirectional cross-comparison pairs includes: Within each cluster, the highest and lowest score samples for each evaluation metric are extracted, and cross-cluster best and worst bidirectional cross-comparison pairs are constructed based on the highest and lowest score samples between different clusters.
[0010] According to a specific implementation of an embodiment of the present invention, the three-level priority hybrid comparison method specifically includes: The first priority is to compare the quality levels between samples; The second priority is to compare the intra-cluster Z-scores among samples when the quality levels are the same; The third priority is to compare the original scores between samples when the Z scores are the same within the cluster; If the highest-scoring sample is not inferior to the lowest-scoring sample in all three priority comparisons mentioned above, the comparison is deemed reasonable; otherwise, it is deemed unreasonable.
[0011] According to a specific implementation of an embodiment of the present invention, the step of calculating the threshold adjustment amount using an iterative threshold adjustment mechanism to update the quality level classification threshold includes: Identify clusters with unreasonable comparison pairs, calculate the difference between the original score of the lowest-scoring sample in the cluster and the original score of the highest-scoring sample in the better cluster, and combine the preset learning rate and the number of better clusters to obtain the average weighted adjustment amount as the threshold adjustment amount. The threshold adjustment amount is subtracted from the original low-quality upper limit threshold, and the threshold adjustment amount is added to the original high-quality lower limit threshold to complete the update of the quality level classification threshold.
[0012] The government response quality assessment scheme in this embodiment of the invention includes: Step 1, using a large language model to extract features from the acquired government response data to obtain multi-dimensional features; Step 2, clustering the government response data according to question type and topic to obtain multiple clusters; Step 3, performing quality assessment within each cluster using a large language model to obtain raw scores, and independently performing Z-score standardization to convert the raw scores into in-cluster Z-scores, and classifying the in-cluster Z-scores into quality levels based on a standard deviation principle; Step 4, constructing cross-cluster bidirectional cross-comparison pairs, wherein the bidirectional cross-comparison pairs include cross-comparison of the best and worst samples, and using a three-level priority mixed comparison method to judge the rationality of the bidirectional cross-comparison pairs; Step 5, for clusters with unreasonable conditions, using an iterative threshold adjustment mechanism to calculate the threshold adjustment amount to update the quality level classification threshold, and reclassifying the quality levels until all bidirectional cross-comparison pairs are reasonable; Step 6, using multi-dimensional features as input and manually labeled quality levels as labels to train a machine learning model to output the government response quality assessment results.
[0013] The beneficial effects of the embodiments of the present invention are as follows: 1. It achieves a deep characterization of the professional logic of government affairs. By constructing a five-dimensional feature system that includes specificity, relevance, information richness, appropriateness, and readability, it achieves a comprehensive characterization of the quality of government affairs responses from surface semantics to deep government affairs logic, effectively solving the problem that existing technologies have only one feature dimension and cannot represent the professionalism of government affairs.
[0014] 2. Eliminate distribution bias between clusters. A technique is proposed that combines independent Z-score standardization within clusters with a standard deviation principle. This enables the quality level classification to accurately match the actual data distribution characteristics of different government affairs themes, overcomes the evaluation distortion of global standardization under non-homogeneous data distribution, and significantly improves the local objectivity of the evaluation results.
[0015] 3. Ensures global logical consistency of assessment scale. Innovatively introduces cross-cluster bidirectional cross-comparison and three-level priority hybrid comparison mechanism. Through the step-by-step verification logic of "grade - Z score - raw score", it effectively eliminates the "quality inversion" phenomenon in cross-cluster assessment and ensures the uniformity of assessment standards among different subject clusters.
[0016] 4. It possesses adaptive closed-loop correction and industrial-grade reliability. Combined with an iterative threshold adjustment mechanism based on learning rate control, it achieves automatic optimization and closed-loop correction of the partitioning threshold, greatly enhancing the system's adaptability to data from new domains. Machine learning verification and sensitivity analysis demonstrate that the technical system of this application achieves a high-precision approximation of manually labeled data, possesses strong generalization ability and robustness, effectively avoids the risk of overfitting, and fundamentally improves the accuracy and reliability of government response quality assessment. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart illustrating a method for assessing the quality of government responses provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the specific implementation process of a government response quality assessment method provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of a five-dimensional feature evaluation system provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of a three-level priority hybrid comparison method provided in an embodiment of the present invention; Figure 5A flowchart of an iterative threshold adjustment mechanism provided in an embodiment of the present invention. Detailed Implementation
[0019] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0020] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0021] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this invention, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0022] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. The illustrations only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.
[0023] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0024] This invention provides a method for evaluating the quality of government responses, which can be applied to the automated text evaluation process in digital government service scenarios.
[0025] See Figure 1This is a flowchart illustrating a method for assessing the quality of government responses provided in an embodiment of the present invention. Figure 1 As shown, the method mainly includes the following steps: Step 1: Use a large language model to extract features from the acquired government response data to obtain multi-dimensional features; Step 2: Cluster the government response data according to question type and topic to obtain multiple clusters; Step 3: Within each cluster, quality assessment is performed using a large language model to obtain raw scores, and Z-score standardization is performed independently to convert the raw scores into in-cluster Z-scores. Based on a standard deviation principle, the in-cluster Z-scores are classified into quality levels. Step 4: Construct cross-cluster bidirectional cross-comparison pairs, wherein the bidirectional cross-comparison pairs include cross-comparisons of the best and worst samples, and a three-level priority hybrid comparison method is used to determine the rationality of the bidirectional cross-comparison pairs; Step 5: For clusters with unreasonable conditions, use an iterative threshold adjustment mechanism to calculate the threshold adjustment amount to update the quality level division threshold, and re-divide the quality levels until all bidirectional cross-comparison pairs are reasonable. Step 6: Use multi-dimensional features as input and manually labeled quality levels as tags to train a machine learning model to output the quality assessment results of government response.
[0026] In specific implementation, such as Figure 2 As shown, this embodiment of the invention provides a method for evaluating the quality of government response, the process of which includes: Step 101: Extract five-dimensional features using a large language model. Specifically, this involves using a large language model to extract multi-dimensional features from the acquired government response data, resulting in features in five dimensions: specificity, relevance, information richness, appropriateness, and readability.
[0027] Step 102: Clustering by question type and topic. Specifically, this involves clustering the aforementioned government response data by question type and topic to obtain multiple clusters, facilitating subsequent benchmark comparisons within the same cluster.
[0028] Step 103: Calculate intra-cluster Z-score standardization. Specifically, this is used to construct a reference sample set within each cluster, perform intra-cluster quality assessment using a large language model, obtain the scores of samples within each cluster on the five-dimensional indicators, and then independently perform Z-score standardization within each cluster, calculating the mean and standard deviation within each cluster, thus converting the original scores into intra-cluster Z-scores.
[0029] Step 104: Classify quality levels (high / medium / low). Specifically, based on the intra-cluster Z-scores calculated above, a standard deviation principle is used to classify the government response data into quality levels, including low quality, medium quality, and high quality.
[0030] Step 105: Construct and validate two-way cross-comparison pairs. Specifically, this involves selecting the highest and lowest score samples for each metric within each cluster to construct two-way cross-comparison pairs for "best vs. worst" and "worst vs. best".
[0031] Step 106: Determine if the comparison pairs are reasonable. Specifically, a three-level priority hybrid comparison method is used (first priority is comparing quality levels, second priority is comparing Z-scores within the cluster, and third priority is comparing raw scores) to determine if each comparison pair is reasonable. If the determination result is negative (i.e., there are unreasonable cases), proceed to step 107; if the determination result is positive (i.e., all comparison pairs are reasonable), proceed to step 108.
[0032] Step 107, Iterative Threshold Adjustment Mechanism. When an unreasonable comparison pair is identified, an iterative threshold adjustment mechanism is used to calculate the threshold adjustment amount. The original threshold is updated by adjusting the low-quality threshold downward or the high-quality threshold upward. Then, the process returns to step 104 to reclassify the quality levels based on the new threshold until all comparison pairs are determined to be reasonable.
[0033] Step 108: Train a machine learning model to verify its effectiveness. Specifically, this involves using the five-dimensional features as input and manually labeled quality levels as tags to train a machine learning classification model to verify the effectiveness and generalization ability of the five-dimensional features.
[0034] Step 109: Output the quality assessment results. Specifically, after the above-mentioned feature extraction, grouping, standardization, rationality verification, and iterative adjustment, the final government response quality assessment results are output.
[0035] The following section provides a detailed explanation of the underlying algorithm execution details and parameter configurations for steps 101 to 109, using specific algorithm formulas and data processing examples: Step 1: Data acquisition, preprocessing, and clustering.
[0036] Step 1.1: Data Collection and Preprocessing. Using the Scrapy web scraping framework, 913 pieces of judicial service Q&A data were collected from the consultation and complaint section of a government affairs platform. Each piece of data contains the question text and its corresponding reply text, forming a government affairs Q&A dataset. : ; in, The total number of samples, For the first The problem text of the record, For the first The response text of the record.
[0037] Step 1.2: Question-answer text clustering.
[0038] Constructing text vector representations of question-answer pairs: ; in, For vector encoders, For the first The question and answer are correct. 3D text vector, for A dimensional real vector space.
[0039] Initial text vectors are constructed using a TF-IDF weighted model, and terms are... In the reply text The formula for calculating the weights in the formula is: ; in, Indicates terms In the reply text The frequency in To the total number of reply texts, For included terms The number of response texts.
[0040] The core objective of K-Means clustering is to find the optimal cluster partition by minimizing the sum of squared distances from all sample points to their respective cluster centers. The formula is as follows: ; in, It is the pre-defined number of clusters; It is the first A cluster; It belongs to a cluster A sample vector; It is a cluster The centroid, which is the mean vector of all sample points in the cluster, is calculated using the following formula: ; It is the Euclidean norm. This indicates the value of the independent variable that minimizes the objective function.
[0041] Determining the optimal number of clusters using silhouette coefficients .sample contour coefficient Defined as: ; in, It is a sample The average distance to other samples in the same cluster (intra-cluster dissimilarity). It is a sample The average distance to all samples in the nearest other cluster (inter-cluster dissimilarity).
[0042] The average silhouette coefficient of the entire dataset is ; Calculate different using the elbow rule The sum of squared errors (SSE) corresponding to the values: ; Experimental results show that when This approach achieves the optimal balance between fine-grained clustering and topic interpretability. The core keywords and sample size for each cluster are shown in Table 1. Table 1
[0043] like Figure 3 As shown in the figure, this application embodiment provides a five-dimensional feature evaluation system 200 for government response quality, which includes: Specificity feature 201 is used to assess whether the response content is closely related to the specific scenario of the problem and to conduct targeted analysis. Relevance feature 202 is used to determine whether the response content is directly related to issues in the judicial field and to provide useful information; Information richness feature 203 is used to evaluate whether the response provides detailed background information in addition to answering the question; Appropriateness feature 204 is used to assess whether the tone of the response is appropriate, the attitude is friendly, and the expression is standard. Readability feature 205 is used to determine whether a response is logically clear, well-written, and easy to understand.
[0044] Step 2: Construction of evaluation indicators and standards for the quality of government response.
[0045] Step 2.1: Construction of Quality Assessment Indicators for Government Responses. An assessment indicator set is constructed from five dimensions: specificity, relevance, information richness, appropriateness, and readability. The specific meaning is as follows: Specificity. Determine whether the response is closely related to the specific context of the problem and conducts targeted analysis, rather than applying a generic template.
[0046] Relevance. This determines whether the response is directly related to "issues in the judicial field"; issues outside the judicial field are excluded.
[0047] Information richness. This assesses whether the response provides supplementary background information beyond the question itself, excluding responses that merely address the surface of the question without offering additional information.
[0048] Appropriateness. This refers to whether the response is clear and direct, avoids evasiveness or perfunctory responses, uses formal language consistent with the rigor of government policy responses, and avoids being stiff or mechanical, making the questioner feel valued and receiving effective guidance.
[0049] Readability. This assesses whether the response is logically clear, well-written, and free of redundant information, allowing the average questioner to clearly understand the core viewpoints and supporting evidence.
[0050] Step 2.2: Constructing the response quality assessment criteria.
[0051] index Specificity Set a question The specific set of elements included is .reply The coverage of these elements is defined as: ; in, It is an indicator function; if it responds... The text mentions or answers the key elements. If the condition is met, return 1; otherwise, return 0.
[0052] index Correlation First, define the government affairs problem identification function. : ; Regarding government affairs The relevance score is defined as the coverage rate of government keywords in the responses. ; in, For a predefined set of keywords in the government affairs field, This is an indicator function. If... If the sample is invalid, it can be discarded or processed in another way.
[0053] index Information richness Break down the reply text into basic information and additional information The richness of information is a combination of the richness of legal basis and the richness of interpretation. ; in, , It is a subset of laws directly related to the issue. The text is the legal basis cited in the reply. Indicates text length; , It is additional information in the reply; This refers to the proportion of additional information in the entire reply text. The actual length of the legal basis text provided in the response, as a percentage of the "ideally required length of the legal basis text." and These are the weighting coefficients, and .
[0054] index Appropriateness Appropriateness score consists of three parts: no excuses index, clear conclusion, and guidance. ; in, These are the weighting coefficients, and . , A collection of excuses Number of times it appears in the replies; This indicator is a binary variable (0 or 1) used to measure whether the response provides a clear and unambiguous conclusion; , The higher the proportion of guiding content in the entire response, the more it indicates that the response not only explains the problem but also provides actionable guidelines.
[0055] index :readability The readability score is a linear combination of multiple positive and negative indicators: ; Among them, each component represents structural clarity, logical connector density, terminology standardization, semantic ambiguity, and redundancy, respectively. .
[0056] Each indicator is evaluated using a three-level system: ; in Indicates the first The first reply was in The evaluation results under each indicator. The five-dimensional characteristic evaluation indicators and quantitative system for government response quality are shown in Table 2.
[0057] Table 2
[0058] Step 3: LLM-based response quality assessment guided by positive and negative reference sets.
[0059] Step 3.1: Construct a high-quality evaluation reference set.
[0060] To avoid inconsistencies in the evaluation criteria of Large Language Models (LLMs) across different topic clusters, this invention uses manually labeled ideal responses (positive references) and negative ideal responses (negative references) in each cluster as evaluation references for LLMs.
[0061] For clusters ( , ),definition: Positive Reference Set :Depend on An ideal response and its thought process chain ; in, For the first The ideal text content for a reply; This is the corresponding thought chain explanation vector. It is an indicator Obtain a high-quality (3 points) reasoning text; the ideal response scores 0.5 on all five metrics. (high quality).
[0062] Negative reference set :Depend on A negative ideal response and its thought chain composition ; in, For the first The text content of a negative ideal response; This is the corresponding thought chain explanation vector. It is an indicator Obtained low-quality (1 point) reasoning text; negative ideal response scored zero on all five metrics. (Low quality).
[0063] Reference set construction scale.
[0064] To balance evaluation efficiency and representativeness, the positive and negative reference sets for each cluster are set to be of equal size and satisfy the following conditions: ; That is, the number of positive and negative reference samples in each cluster does not exceed 10, and does not exceed 10% of the total number of samples in that cluster.
[0065] The reference set is generated through manual annotation: randomly selected from each cluster. For each sample, a five-dimensional score is assigned by a human based on the evaluation criteria in step 2, and a detailed explanation of the thought process is written to ensure that the positive reference sample scores 3 points in all dimensions and the negative reference sample scores 1 point in all dimensions.
[0066] Step 3.2: Intragroup assessment.
[0067] A set of positive and negative references is designed in the prompt words. The large language model evaluates the response quality within each cluster and conducts the evaluation based on the evaluation indicators and standards designed in step 2.
[0068] 1. Define the evaluation function for the large language model. .
[0069] For clusters Any question and answer pair in ( The evaluation process is as follows: ; The output The first LLM is given based on the reference set and index definition. 1 reply in the indicator The evaluation level is (1=low quality, 2=medium quality, 3=high quality).
[0070] 2. Evaluation prompt design.
[0071] LLM input prompts Functions constructed from prompt words generate: ; Its structured content includes: 1. Task Description: It explicitly requires a multi-dimensional quality assessment of government responses.
[0072] 2. Definition of evaluation indicators: Explaining each one to The meaning and three-level scoring criteria.
[0073] 3. Correct reference example: exhibit One or two samples were selected, including their responses, five-dimensional scores (out of 3), and explanations of their thought processes.
[0074] 4. Negative Reference Example: Demonstration One or two samples were selected, including their responses, five-dimensional scores (all 1 point), and explanations of their thought processes.
[0075] 5. Issues to be evaluated: Provide the issues currently awaiting evaluation. and reply .
[0076] 6. Output format requirements: The LLM should output the five-dimensional scores (each dimension being an integer of 1 / 2 / 3) and a brief explanation of the reasoning in JSON format.
[0077] Cluster All of them After each response is evaluated, a five-dimensional rank vector is obtained for each response.
[0078] To characterize the central tendency and dispersion of scores within a cluster, the central tendency and dispersion of all responses within that cluster are calculated based on the index. mean of grade scores with standard deviation : ; in, It is the first Clusters ( , ), It is a cluster The number of samples included. It is a cluster The first One reply in the metrics The original score on It is a cluster All samples in the index The arithmetic mean of the above.
[0079] ; in, It is the first The difference between the original score of each sample and the mean of the cluster (i.e., the deviation).
[0080] Thus, we obtain the standardized version Fraction : ; Standardized score It approximately follows a standard normal distribution .
[0081] The quality level can then be determined based on the degree to which the score deviates from the mean.
[0082] Define low quality threshold With high quality threshold as follows: ; ; Based on the above threshold, the cluster The response within the indicator The quality is divided into three levels.
[0083] Therefore, the first is defined 1 reply in the indicator The assessment level for: ; in These correspond to low, medium, and high quality, respectively.
[0084] Step 3.3: Intergroup comparison.
[0085] The highest and lowest scores of each group are used to construct a set, which is then used to form a comparison matrix. The rows are the sets of the highest quality responses in each category that are classified as high quality, and the columns are the sets of the lowest quality responses in each category that are classified as low quality. This matrix is used to perform a difference analysis.
[0086] Furthermore, a method was designed to identify unreasonable evaluation categories. By comparing across clusters, the rationality of the hierarchical division within each cluster was examined, and clusters that may have biases were identified.
[0087] 1. Construction of inter-group comparison pairs.
[0088] For each metric From each cluster Selecting extreme value samples: High quality representative : cluster The medium grade is "high quality" ( )and The highest-scoring response has its original score recorded as follows: .
[0089] Low quality representative : cluster The medium grade is "low quality" ( )and The lowest-scoring response has its original score recorded as follows: .
[0090] Define cluster and In indicators One of the comparisons above is .
[0091] For all Clusters can be constructed A comparison.
[0092] 2. Mixed comparison method.
[0093] like Figure 4 As shown, this application embodiment provides a three-level priority hybrid comparison method. The judgment process 301 of the three-level priority hybrid comparison method includes: The first priority comparison step 302 is used to initially compare whether the quality level L between extreme value samples is different. If they are different, the reasonableness judgment result 305 is output. The second priority comparison step 303 is used to further compare whether the intra-cluster Z scores of each sample are different when the quality levels are the same. If they are different, the reasonableness judgment result 306 is output. The third priority comparison step 304 is used to compare the original scores S of the samples and output the final reasonableness judgment result 307 when the Z scores are also the same.
[0094] Introducing a three-level priority comparison function Determine the reasonableness of the comparison pair. For the comparison pair... The logic for its rationality judgment is as follows: ; If it appears If the best response is not better than the worst response, then the comparison is deemed unreasonable.
[0095] The logic for determining the rationality of cross-cluster comparison in the three-level priority system is shown in Table 3.
[0096] Table 3
[0097] 3. Threshold iterative optimization mechanism. For example... Figure 5 As shown in the figure, this application embodiment provides an execution flow of an iterative threshold adjustment mechanism, which includes: Identification step 401 is used to identify clusters with unreasonable comparisons; Calculation step 402 is used to calculate the threshold adjustment amount Δi based on the original score difference; Update step 403 is used to update the low-quality threshold tiL and the high-quality threshold tiH according to the adjustment amount; Step 404 is used to reclassify the quality levels of samples within a cluster based on a new threshold. Verification step 405 is used to reconstruct the comparison pair and perform a reasonableness verification again; The determination step 406 is used to determine whether all comparison pairs are reasonable. If the determination result is "no", the closed-loop logic is triggered to return to the identification step 401 to continue the iteration.
[0098] Unreasonably defined cluster set This includes clusters containing unreasonable comparison pairs. For Adjust the threshold for its quality grade classification. The original threshold was defined as: ; ; in, The upper limit of low quality, This represents the lower limit of low quality.
[0099] The adjusted new threshold is: ; ; Among the adjustment amounts Calculation based on the difference between this cluster and a better cluster: ; here To control the learning rate, adjust the step size; Those whose optimal responses are better than clusters The set of clusters with the worst responses; and All scores are raw scores.
[0100] Re-divide clusters using the adjusted threshold. The quality level within the range was determined, and new comparison pairs were constructed for verification. This process was iterated until all comparison pairs were deemed reasonable. Experimental verification showed that after this adjustment, the five indicators... One comparison pair (actually, after deduplication, there are 10 pairs) The reasonable rate of (number) reached .
[0101] Step 4: Machine Learning-Based Quality Rating Prediction Model Step 4.1: Supervised Dataset Construction. The five-dimensional index score vector obtained after iterative optimization in Step 3 is used as the input feature. For the... A sample, defined as: ; The overall quality level marked by humans is ,in , , These correspond to "low quality," "medium quality," and "high quality," respectively. All The supervised dataset consists of 10 samples: ; Step 4.2: Dataset Splitting and Feature Engineering. The dataset is split into a training set and a... and test set The ratio is Let the number of samples in the training set be... The number of samples in the test set is ,satisfy The five-dimensional score vector is a numerical discrete feature and can be directly used as input to the classification model.
[0102] Step 4.3: Classification Model Construction. Multi-class logistic regression is chosen as the base prediction model. For a given input feature vector... The model assumes that the first kind( The conditional probability of ) is: ; in, For the first Class weight vector, This is the bias term. Model parameters. By minimizing the band Estimating the loss function using regularized cross-entropy: ; in, For indicator functions, The regularization coefficient is obtained by... Cross-validation (taking) Determine the optimal value.
[0103] Step 4.4: Model Evaluation and Selection. To select the optimal model, five algorithms were compared: logistic regression, K-nearest neighbors, random forest, support vector machine, and Naive Bayes. On the test set... The performance of each model was evaluated, with metrics including accuracy and macro-average F1. ; in, The category predicted by the model.
[0104] ; Experiments show that, in a preferred embodiment of the present invention, the performance comparison table of machine learning classification models is shown in Table 4. The random forest model achieves better prediction results, reaching [a certain level] on the test set. The accuracy. Using the best model on the full dataset. Predicting from 100 samples, the final overall prediction accuracy is 100%. The prediction accuracy rates for each level are as follows: Level 1 (low quality) Grade 2 (Medium quality) Level 3 (High Quality) .
[0105] Table 4
[0106] The government response quality assessment method provided in this embodiment constructs a five-dimensional feature system encompassing specificity, relevance, information richness, appropriateness, and readability. This system comprehensively characterizes the quality of government responses from surface semantics to deep government logic, effectively solving the problem of existing technologies having a single feature dimension and being unable to represent the professionalism of government affairs. It proposes a technical approach combining independent Z-score standardization within clusters with a single standard deviation principle, enabling quality level classification to accurately match the actual data distribution characteristics of different government themes. This overcomes the assessment distortion caused by global standardization under non-homogeneous data distributions, significantly improving the local objectivity of the assessment results. The method introduces a cross-cluster bidirectional cross-comparison and a three-level priority hybrid comparison mechanism. Through a step-by-step verification logic of "level - Z-score - original score," it effectively eliminates the "quality inversion" phenomenon in cross-cluster assessments, ensuring the uniformity of assessment standards between different theme clusters. Combined with an iterative threshold adjustment mechanism based on learning rate control, it achieves automatic optimization and closed-loop correction of the classification threshold, greatly enhancing the system's adaptability to new domain data. Through machine learning verification and sensitivity analysis, it has been demonstrated that the technical system of this application achieves a high-precision approximation of manually labeled tags, possesses strong generalization ability and robustness, effectively avoids the risk of overfitting, and fundamentally improves the accuracy and reliability of government response quality assessment.
[0107] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof.
[0108] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for evaluating the quality of government response, characterized in that, include: Step 1: Use a large language model to extract features from the acquired government response data to obtain multi-dimensional features; Step 2: Cluster the government response data according to question type and topic to obtain multiple clusters; Step 3: Within each cluster, quality assessment is performed using a large language model to obtain raw scores, and Z-score standardization is performed independently to convert the raw scores into in-cluster Z-scores. Based on a standard deviation principle, the in-cluster Z-scores are classified into quality levels. Step 4: Construct cross-cluster bidirectional cross-comparison pairs, wherein the bidirectional cross-comparison pairs include cross-comparisons of the best and worst samples, and a three-level priority hybrid comparison method is used to determine the rationality of the bidirectional cross-comparison pairs; Step 5: For clusters with unreasonable conditions, use an iterative threshold adjustment mechanism to calculate the threshold adjustment amount to update the quality level division threshold, and re-divide the quality levels until all bidirectional cross-comparison pairs are reasonable. Step 6: Use multi-dimensional features as input and manually labeled quality levels as tags to train a machine learning model to output the quality assessment results of government response.
2. The method according to claim 1, characterized in that, The multi-dimensional features are five-dimensional features, including specificity features, relevance features, information richness features, appropriateness features, and readability features.
3. The method according to claim 1, characterized in that, The expression for the intra-cluster Z-score is: in, Indicates the first The first cluster The sample at the th Intra-cluster Z-scores on each metric Indicates the first The first cluster The sample at the th The raw scores on each indicator Indicates the first The first cluster The mean of each indicator within the cluster Indicates the first The first cluster The standard deviation of each indicator within the cluster.
4. The method according to claim 3, characterized in that, The quality level classification of the intra-cluster Z-scores based on a standard deviation principle is as follows: When the Z-score within a cluster is less than negative one standard deviation, it is considered low quality; When the in-cluster Z-score is between negative one standard deviation and positive one standard deviation, it is considered to be of medium quality; When the Z-score within a cluster is greater than one standard deviation, it is considered to be of high quality.
5. The method according to claim 1, characterized in that, The construction of cross-cluster bidirectional cross-comparison pairs includes: Within each cluster, the highest and lowest score samples for each evaluation metric are extracted, and cross-cluster best and worst bidirectional cross-comparison pairs are constructed based on the highest and lowest score samples between different clusters.
6. The method according to claim 5, characterized in that, The three-level priority hybrid comparison method specifically includes: The first priority is to compare the quality levels between samples; The second priority is to compare the intra-cluster Z-scores among samples when the quality levels are the same; The third priority is to compare the original scores between samples when the Z scores are the same within the cluster; If the highest-scoring sample is not inferior to the lowest-scoring sample in all three priority comparisons mentioned above, the comparison is deemed reasonable; otherwise, it is deemed unreasonable.
7. The method according to claim 6, characterized in that, The step of calculating the threshold adjustment amount using an iterative threshold adjustment mechanism to update the quality level classification threshold includes: Identify clusters with unreasonable comparison pairs, calculate the difference between the original score of the lowest-scoring sample in the cluster and the original score of the highest-scoring sample in the better cluster, and combine the preset learning rate and the number of better clusters to obtain the average weighted adjustment amount as the threshold adjustment amount. The threshold adjustment amount is subtracted from the original low-quality upper limit threshold, and the threshold adjustment amount is added to the original high-quality lower limit threshold to complete the update of the quality level classification threshold.