A multi-granularity hierarchical uncertainty evaluation method and system for long text generation

CN122817418APending Publication Date: 2026-09-25QUAN CHENG LABORATORY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611299013.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-26
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0008]针对现有技术的不足,本发明提供一种面向长文本生成的多粒度层次化不确定性评估方法及系统,解决现有长文本生成不确定性评估方法多针对单一粒度建模,难以同时刻画原子事实、句子结构和整体回答分布的问题;现有方法难以有效利用同一输入问题下多次采样生成结果之间的跨回答一致性信息问题;现有方法在长链路推理和开放域长文本场景中,对局部错误、事实漂移和推理不稳定现象的识别能力有限问题

Benefits of technology

1、通过对同一输入问题进行多次采样并联合分析,可更充分地利用模型输出波动信息,提高不确定性评估的稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817418A_ABST
    Figure CN122817418A_ABST
Patent Text Reader

Abstract

The application relates to a multi-granularity hierarchical uncertainty evaluation method and system for long text generation, and belongs to the technical field of artificial intelligence and natural language processing. The steps comprise: acquiring an input question, setting a sampling number, a temperature parameter and a decoding control parameter; performing multi-time sampling generation on the input question to obtain N candidate answers and corresponding token probability sequences; performing atomic fact extraction on the candidate answers to calculate atomic fact granularity uncertainty FU(x); performing step-level or sentence-level segmentation on the candidate answers to calculate sentence granularity uncertainty SU(x); aggregating the confidence of a single candidate answer according to a semantic equivalent answer cluster to obtain overall generation granularity uncertainty GU(x); and averaging FU(x), SU(x) and GU(x) to obtain a final question-level uncertainty RU(x).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multi-granular hierarchical uncertainty assessment method and system for long text generation, belonging to the field of artificial intelligence and natural language processing technology. Background Technology

[0002] With the widespread application of large language models in tasks such as question answering, retrieval enhancement generation, chained reasoning, summary generation, and knowledge-based writing, the quality of model output has become a key factor affecting the reliability of applications. Especially in long text generation scenarios, the same answer often contains multiple factual fragments, multiple reasoning steps, and a final conclusion; once some of the content is biased, fabricated, or logically mismatched, it may make the overall answer unreliable.

[0003] Most existing uncertainty assessment methods are designed for short texts or single-sentence answers. They typically estimate model confidence using only token probability, single-answer score, or coarse-grained semantic consistency. These methods are effective in single-sentence classification or short question-and-answer scenarios, but they still have significant shortcomings in long text generation scenarios.

[0004] On the one hand, long text answers usually have a hierarchical structure. For example, chained reasoning answers contain multiple intermediate reasoning steps as well as the final answer; open-domain long text answers contain not only several atomic facts, but also sentences and paragraphs formed by combining multiple facts. Existing methods, if they only model at the token level or the whole sentence level, often cannot accurately reflect the consistency and volatility between information of different granularities.

[0005] On the other hand, for the same input question, large language models can usually generate multiple candidate answers under random sampling conditions. Different candidate answers may exhibit the following characteristics: the final answer is consistent but the reasoning process is different; the reasoning process is locally consistent but the final answer is different; some facts are stable while some facts are unstable. Existing methods lack a unified mechanism to incorporate the consistency of these multiple sampling results at different granularities into the same question-level uncertainty assessment framework.

[0006] Furthermore, some existing methods, when estimating uncertainty in long texts, only focus on whether atomic facts are supported by other answers, or only focus on semantic matching between sentences, or only focus on the semantic clustering distribution of the final answer. They lack a technical solution that integrates the atomic fact layer, sentence layer, and overall answer layer. As a result, in complex reasoning and long text generation scenarios, the output uncertainty index has insufficient stability, limited interpretability, and is difficult to accurately reflect the model's overall grasp of the current problem.

[0007] Therefore, a new technical solution is urgently needed to perform multi-granular hierarchical modeling of the multiple sampling results of the same input problem, calculate uncertainty at the atomic fact layer, sentence layer and overall generation layer respectively, and then unify and fuse the multi-granular results to obtain the overall uncertainty score for the current problem, thereby improving the risk identification capability, stability and interpretability in long text generation scenarios. Summary of the Invention

[0008] To address the shortcomings of existing technologies, this invention provides a multi-granular hierarchical uncertainty assessment method and system for long text generation. It solves the problems of existing long text generation uncertainty assessment methods, which often focus on single-granularity modeling and struggle to simultaneously characterize atomic facts, sentence structure, and overall response distribution; existing methods also struggle to effectively utilize cross-response consistency information between multiple samples generated from the same input question; and existing methods have limited ability to identify local errors, fact drift, and inference instability in long-link reasoning and open-domain long text scenarios.

[0009] The technical solution of the present invention is as follows: A multi-granular hierarchical uncertainty assessment method for long text generation includes: (1) Obtain the input question to be evaluated, and call the target large language model to sample the input question multiple times to generate N candidate answers; (2) Perform atomic fact extraction on N candidate answers and calculate the atomic fact uncertainty score based on the implication relationship of atomic facts to other candidate answers; (3) Perform sentence-level or reasoning step-level segmentation on N candidate answers, and calculate sentence uncertainty score based on the implication relationship of the maximum matching across answer sentences; (4) Extract the final cleaned answer from the N candidate answers, perform semantic clustering based on the bidirectional implication relation, and obtain multiple semantically equivalent answer clusters; (5) Extract keywords and keyword contribution from the reasoning steps of each candidate answer, and backfill the token probability corresponding to the keyword from the token probability sequence in the candidate answer generation process; (6) Take the minimum value of the multiple token probabilities corresponding to each keyword as the probability representation of the keyword, and then combine it with the keyword contribution to obtain the answer confidence of each candidate answer; (7) Aggregate the confidence scores of each candidate answer according to the semantically equivalent answer cluster to obtain the uncertainty score at the overall generation granularity; (8) Take the average of the atomic fact uncertainty score, sentence uncertainty score and overall generation uncertainty score to obtain the final uncertainty score of the current input question.

[0010] According to a preferred embodiment of the present invention, in step (1), the method of generating the data through multiple sampling is as follows: Set the sampling number, temperature parameter, and decoding control parameter of the target large language model. Call the target large language model to sample the input question multiple times to generate N candidate answers. For each candidate answer, save at least the following information: candidate answer text, candidate answer token sequence, and the conditional probability of each token at the time of generation. For chain reasoning tasks, further save the cleaned final answer text for subsequent semantic clustering.

[0011] According to a preferred embodiment of the present invention, in step (1), the N candidate answers are denoted as:

[0012] Where, r i As candidate answers, Each candidate answer r i The system includes a corresponding token sequence and token conditional probabilities. The final output is not a score for a single answer, but rather a result reflecting the overall uncertainty of the current input question.

[0013] According to a preferred embodiment of the present invention, in step (2), the atomic fact extraction model is invoked for each candidate answer to split the candidate answer into several independent atomic facts. An atomic fact is a factual expression that contains a single information point and can be independently determined to be true or consistent. Let the set of atomic facts extracted from the i-th candidate answer be:

[0014] Among them, M i Let m represent the number of atomic facts corresponding to the i-th candidate answer, and m be the atomic fact index. ; For any atomic fact in the i-th candidate answer The implication model is used to calculate its overall text for the j-th candidate answer. The implied probability is denoted as:

[0015] The atomic fact consistency score of the i-th candidate answer relative to the j-th candidate answer can be expressed as:

[0016] The local atomic fact uncertainty of the i-th candidate answer is defined as:

[0017] Where N represents the total number of candidate answers generated by sampling the same input question; Ultimately, the uncertainty at the atomic fact granularity of the current input problem is defined as:

[0018] At the implementation level, the implication model can adopt potsawee / deberta-v3-large-mnli, for any text pair First, input both into the implication model to obtain binary classification logits, then use softmax to obtain the probability distributions for the "implication / contradiction" classes, and take the probability corresponding to the implication class as... In this application implementation, the implied probability is denoted as:

[0019] According to a preferred embodiment of the present invention, in step (3), each candidate answer is segmented according to the reasoning step structure. For chain-reasoning answers, step-level segments Step1, Step2, and Final Answer are taken as sentence-level units. The set of sentence units after segmentation of the i-th candidate answer is denoted as:

[0020] Among them, L i This represents the total number of sentence-level units obtained after segmenting the i-th candidate answer; For any sentence-level unit in the i-th candidate answer , Calculate the implication probability of each sentence-level unit in the j-th candidate answer, and denote the set of sentence-level units after segmentation of the j-th candidate answer as . , Let represent the sentence-level unit index in the j-th candidate answer; then, the maximum value is taken as the best matching score.

[0021] The sentence consistency score of the i-th candidate answer relative to the j-th candidate answer can be expressed as:

[0022] The definition of local sentence uncertainty in the i-th candidate answer corresponds to the definition of local atomic fact uncertainty, the difference being that sentence-level units are used as the comparison objects here, i.e.:

[0023] Finally, the sentence-level uncertainty of the current input question is defined as:

[0024] According to a preferred embodiment of the present invention, in step (4), the final answer after cleaning is first extracted for each candidate answer, and then a bidirectional implication judgment is performed on any two final answers; if and only if both directions satisfy the implication condition, the two final answers are determined to be semantically equivalent and are classified into the same semantically equivalent answer cluster; Let the final set of answers after cleaning be:

[0025] in, Indicates the first The final answer for each candidate answer; The bidirectional implication clustering method described above can aggregate candidate answers with the same final answer or semantic equivalence together, thereby forming multiple semantically equivalent answer clusters.

[0026] According to a preferred embodiment of the present invention, in step (6), the i-th candidate answer is first parsed to identify the structured segments corresponding to Step 1 to Step n and the Final Answer, and the m-th step text of the i-th candidate answer is denoted as... For the text of this step, extract keywords and their corresponding contributions. If a step is marked as NO ANSWER, then that step will not participate in subsequent calculations. Keywords represent words, phrases, or sub-expressions that have a major impact on the reasoning conclusion in this step; keyword contribution indicates the degree of importance of the keyword's semantic contribution to the current step, assuming it is derived from the step text. The CCP extracted If there are 10 keywords, then the set of keywords corresponding to this step is denoted as:

[0027] in, Let the keyword be the u-th keyword extracted in the m-th step of the i-th candidate answer. Corresponding in the generated results If there are 10 tokens, then the set of probabilities corresponding to those tokens is denoted as:

[0028] in, Keywords The probability of generating the corresponding t-th token is calculated by taking the minimum probability of multiple tokens corresponding to the same keyword, and this minimum probability is used as the keyword-level probability representation of the keyword in this step.

[0029] Let the text of the m-th step of the i-th candidate answer be... For the keywords extracted in this step First, locate the start and end positions of the token in the corresponding token sequence of the step text. Then, backfill the original generated token sequence to obtain the token ID sequence corresponding to the keyword and the generation probability of each token. If the keyword fails to match in the corresponding step text, skip the keyword. For keywords that match successfully... Keywords The set of all token generation probabilities corresponding to the occurrence of this step, and This represents the keyword-level probability representation of the keyword in the current step, obtained by taking the minimum value of the token probability set.

[0030] Keyword confidence is achieved by first aggregating repeated keywords and then calculating the confidence of a single answer. Specifically, if the same keyword appears multiple times in different steps, the keyword-level probabilities obtained from each occurrence are first aggregated. Let H be the total number of occurrences of keyword k. k The probability of each occurrence of the corresponding keyword level is as follows: Then the aggregation probability after its recurrence is expressed as:

[0031] Correspondingly, if the same keyword corresponds to multiple contribution scores in different steps... The aggregation contribution is obtained by taking their average value.

[0032] After aggregating duplicate keywords, the confidence score of the i-th candidate answer is calculated based on the deduplicated keyword set.

[0033] in, Let represent the set of keywords after deduplication in the i-th candidate answer. If the denominator is 0, it degenerates into taking the average of the probabilities of all valid keywords as the confidence backoff value.

[0034] The above implementation corresponds to the main aggregation path in the code: first, keywords and their contribution are extracted from each step; then, keyword probabilities are obtained based on precise token alignment; subsequently, aggregation is performed on the keyword text, and secondary aggregation of probabilities and averaging of contribution are performed on repeated keywords; finally, the result corresponding to a single candidate answer is obtained. Therefore, in this application Instead of simply averaging all the steps, the answer-level confidence score is calculated by extracting data within each step, removing duplicates between steps, and then aggregating duplicates.

[0035] According to a preferred embodiment of the present invention, in step (7), after obtaining the confidence scores of all candidate answers, the confidence scores of candidate answers belonging to the same semantic cluster are summed according to the semantically equivalent answer clusters to obtain the first... Cluster scores of semantic clusters:

[0036] in, Indicates the first Let K be the total number of semantic clusters. Overall generation granularity uncertainty fraction Defined as:

[0037]

[0038] Therefore, the overall generation granularity uncertainty score takes into account both the reliability of key tokens within candidate answers and the distribution of candidate answers among semantically equivalent answer clusters.

[0039] According to a preferred embodiment of the present invention, in step (8), the atomic fact granularity uncertainty of the current input problem is... Sentence granularity uncertainty and overall generation granularity uncertainty We directly take the average to obtain the final problem-level uncertainty score:

[0040] The final question-level uncertainty score is used to represent the overall uncertainty of the target large language model for the current input question. The larger the value, the higher the instability of the model output in multiple samplings, multi-granularity structures, and overall answer distribution.

[0041] This application also provides a multi-granularity hierarchical uncertainty assessment system for long text generation, the system comprising: The sampling and generation module is used to obtain the input question to be evaluated and call the target large language model to sample and generate N candidate answers multiple times. The atomic fact processing module is used to extract atomic facts from N candidate answers and calculate the atomic fact uncertainty score based on the implication relationship of atomic facts to other candidate answers. The sentence processing module is used to perform sentence-level or reasoning step-level segmentation on N candidate answers and calculate sentence uncertainty scores based on the implication relations of the maximum matching across answer sentences. The semantic clustering module is used to extract the final cleaned answer from N candidate answers. It performs semantic clustering based on bidirectional implication relations to obtain multiple semantically equivalent answer clusters. The keyword extraction and probability backfilling module is used to extract keywords and keyword contribution from the reasoning steps of each candidate answer, and backfill the token probability corresponding to the keyword from the token probability sequence in the candidate answer generation process; The generation-level aggregation module is used to take the minimum value of the probabilities of multiple tokens corresponding to each keyword as the probability representation of the keyword, and then combine the keyword contribution to obtain the answer confidence of each candidate answer. The confidence of each candidate answer is aggregated according to the semantically equivalent answer cluster to obtain the uncertainty score of the overall generation granularity. The fusion output module is used to average the uncertainty scores of atomic facts, sentences, and the overall generation uncertainty score to obtain the final uncertainty score of the current input question.

[0042] In another embodiment, this application also provides an apparatus including a processor and a memory, wherein the memory stores a computer program that, when executed on the processor, implements the above-described method.

[0043] In another embodiment, this application also provides a computer-readable storage medium having a computer program stored thereon that, when executed on a processor, implements the above-described method.

[0044] The beneficial effects of this invention are as follows: 1. By sampling and jointly analyzing the same input problem multiple times, the fluctuation information of the model output can be utilized more fully, thereby improving the stability of uncertainty assessment.

[0045] 2. By modeling at the atomic fact granularity, sentence granularity, and overall generation granularity respectively, the sources of instability within long text responses can be more comprehensively characterized.

[0046] 3. By combining semantically equivalent answer clusters with keyword-level confidence, the reliability of the final answer distribution and key reasoning components can be considered simultaneously.

[0047] 4. By integrating multi-granularity uncertainties into a unified question-level score, more interpretable evaluation results can be provided for hallucination detection, risk alerts, human-machine collaborative auditing, and high-reliability generation applications. Attached Figure Description

[0048] Figure 1The HIDE flowchart of this invention generates multiple candidate answers for the same input question and extracts uncertainty information from the atomic fact granularity, sentence granularity and overall generation granularity respectively. Finally, the results are fused to obtain the question-level uncertainty result. HIDE is the framework name of the uncertainty measurement method: Hierarchical-DEcomposition enhanced UQ (Uncertainty Quantification).

[0049] Figure 2 This is a graph showing the trend of uncertainty assessment results under different sampling numbers.

[0050] Figure 3 The figure shows the results of the ablation experiment with multiple particle size fractions.

[0051] Figure 4 The figure shows the experimental results of the generator-level signal combination.

[0052] Figure 5 This is a schematic diagram of the method framework of the present invention.

[0053] Figure 6 This is a flowchart of the method of the present invention.

[0054] Figure 7 This is a schematic diagram illustrating the calculation of uncertainty in atomic fact granularity and sentence granularity according to the present invention.

[0055] Figure 8 This is a schematic diagram illustrating the calculation of generation granularity uncertainty in this invention.

[0056] Figure 9 This is a schematic diagram illustrating the final uncertainty fusion of the present invention. Detailed Implementation

[0057] The present invention will be further described below with reference to the embodiments and accompanying drawings, but is not limited thereto.

[0058] Example 1: This embodiment provides a multi-granularity hierarchical uncertainty assessment system for long text generation. The system may include an input management unit, a sampling generation module, an atomic fact processing module, a sentence processing module, a semantic clustering module, a keyword extraction and probability backfilling module, a generation-level aggregation module, and a fusion output module.

[0059] The input management unit is used to receive the problem to be evaluated, the calling parameters, the sampling configuration, and the output requirements; The sampling and generation module is used to obtain the input question to be evaluated and call the target large language model to sample and generate N candidate answers multiple times. The atomic fact processing module is used to extract atomic facts from N candidate answers and calculate the atomic fact uncertainty score based on the implication relationship of atomic facts to other candidate answers. The sentence processing module is used to perform sentence-level or reasoning step-level segmentation on N candidate answers and calculate sentence uncertainty scores based on the implication relations of the maximum matching across answer sentences. The semantic clustering module is used to extract the final cleaned answer from N candidate answers. It performs semantic clustering based on bidirectional implication relations to obtain multiple semantically equivalent answer clusters. The keyword extraction and probability backfilling module is used to extract keywords and keyword contribution from the reasoning steps of each candidate answer, and backfill the token probability corresponding to the keyword from the token probability sequence in the candidate answer generation process; The generation-level aggregation module is used to take the minimum value of the probabilities of multiple tokens corresponding to each keyword as the probability representation of the keyword, and then combine the keyword contribution to obtain the answer confidence of each candidate answer. The confidence of each candidate answer is aggregated according to the semantically equivalent answer cluster to obtain the uncertainty score of the overall generation granularity. The fusion output module is used to average the uncertainty scores of atomic facts, sentences, and the overall generation uncertainty score to obtain the final uncertainty score of the current input question, such as... Figure 1 As shown.

[0060] Overall process like Figures 5-6 As shown, the method in this embodiment includes the following steps: S101 acquires the input problem and sets the number of samples, temperature parameters, and decoding control parameters; S102 calls the target large language model to sample the input question multiple times to generate N candidate answers and their corresponding token probability sequences; S103 performs atomic fact extraction on candidate answers and calculates the atomic fact granularity uncertainty FU(x); S104 performs step-level or sentence-level segmentation on candidate answers and calculates the sentence-level uncertainty SU(x); S105 performs bidirectional implication clustering on the final answers after candidate answer cleaning to form semantically equivalent answer clusters; S106 extracts keywords and keyword contribution from the reasoning steps of candidate answers, and calculates the confidence of a single candidate answer by combining the token probability; S107 aggregates the confidence scores of individual candidate answers according to semantically equivalent answer clusters to obtain the overall generation granularity uncertainty GU(x); S108 takes the average of FU(x), SU(x), and GU(x) to obtain the final problem-level uncertainty RU(x), as follows: Figure 9 As shown.

[0061] In this embodiment, the N candidate answers obtained for the input question x are denoted as: R(x) = {r_1, r_2, ..., r_N} Each candidate answer r_i is accompanied by a corresponding token sequence and a token conditional probability sequence. The final output target is not the score of a single answer, but the overall uncertainty result for the current input question.

[0062] Implementation methods for multiple sampling In this embodiment, multiple candidate answers are generated for the same input question using random sampling. This sampling generation can be achieved by setting temperature parameters, top-k parameters, top-p parameters, or other decoding control parameters. Preferably, the number of samplings... The value is 5, but this application is not limited to this value. In one set of implementations, the following parameter combinations can be used:

[0063]

[0064] Of course, those skilled in the art can adjust the above parameters according to the difficulty of the task, the size of the model, and the response length. Figure 2 To illustrate the changing trends of uncertainty assessment results under different sampling numbers, CoT-UQ (Chain-of-Thoughtenhanced Uncertainty Quantification) is an enhanced uncertainty quantification method based on thought chains. Experiments show that as the number of samplings increases from 2 to 5, the assessment effect usually improves significantly; however, as the number of samplings continues to increase, the gains gradually slow down.

[0065] Figure 3 The results of the ablation experiment for multi-granularity fractions show that the fused result is usually better than the single-granularity fraction, indicating that the granularity of atomic facts, sentence granularity and overall generation granularity are complementary.

[0066] In a set of experimental implementations, in addition to the multiple sampling generated In addition to the candidate answers, one additional greedy answer can be generated for the same input question, denoted as . Greedy answers are primarily used in the experimental evaluation phase to construct correct / incorrect labels, while the candidate answer set obtained from multiple samplings... Only used for calculation , , and the final uncertainty score .

[0067] For each candidate answer, the system saves at least the following information: candidate answer text, candidate answer token sequence, and the conditional probability of each token at the time of generation. For chain reasoning tasks, the cleaned final answer text can also be saved for subsequent semantic clustering.

[0068] Atomic fact granularity In one implementation, an atomic fact extraction model is invoked for each candidate answer to break it down into several independent atomic facts. Each atomic fact is a factual expression that contains a single piece of information and can be independently determined to be true or consistent. (Note: The original text contains a typo and can be left as is.) The set of atomic facts obtained from the candidate answers is as follows:

[0069]

[0070] Among them, M i Let m represent the number of atomic facts corresponding to the i-th candidate answer, and m be the atomic fact index. ; For any atomic fact in the i-th candidate answer The implication model is used to calculate its overall text for the j-th candidate answer. The implied probability is denoted as:

[0071] The atomic fact consistency score of the i-th candidate answer relative to the j-th candidate answer can be expressed as:

[0072] Furthermore, the first The local atomic fact uncertainty of a candidate answer is defined as:

[0073] Ultimately, the uncertainty at the atomic fact granularity of the current input problem is defined as:

[0074] At the implementation level, the implied model can be 'potsawee / deberta-v3-large-mnli'. For any text pair First, input both into the implication model to obtain binary classification logits, then use softmax to obtain the probability distributions for the "implication / contradiction" classes, and take the probability corresponding to the implication class as... In the code implementation of this application, the implied probability can be denoted as:

[0075]

[0076] In other words, in the implementation of AU / FU and SU, it is preferable not to explicitly use neutral categories, but to directly use the implied class probability to measure whether the current fact fragment or step fragment can be supported by another candidate answer; the higher the implied probability, the more stable the fragment is across multiple samples.

[0077] Atomic fact extraction models can employ external instruction-following models, such as instruction-based or dialogue-based models, to break down the entire response into several facts using a unified prompt. Preferably, the extraction results are output using structured labels, such as Facts. <begin> ... <end>And agreed to distinguish between different facts by <split>Separate; then use regular expressions to separate... <begin>and <end>Extract text between, then press <split>Segmentation is performed. To improve feasibility, a retry mechanism can be set up for each candidate answer; when an extraction fails, enhanced prompts with fixed output format constraints are sent until at least one atomic fact is successfully extracted or the maximum number of retries is reached.

[0078] In one set of implementations, atomic fact extraction can be performed in parallel, i.e., for Each candidate answer initiates a concurrent request to obtain If the first If no valid atomic facts are extracted from a candidate response, then... It can be set to 0. Furthermore, the implied probability matrix between each atomic fact and the full text of the remaining candidate answers can be saved as a basis for subsequent debugging, visualization, or result interpretation.

[0079] Sentence granularity uncertainty implementation method; In one implementation, each candidate answer is segmented according to the reasoning step structure. For chain-reasoning answers, step-level segments such as Step 1, Step 2, and Final Answer are treated as sentence-level units. The set of sentence units after segmentation of the i-th candidate answer is denoted as:

[0080] Among them, L i This represents the total number of sentence-level units obtained after segmenting the i-th candidate answer; For any sentence-level unit in the i-th candidate answer , Calculate the implication probability of each sentence-level unit in the j-th candidate answer, and denote the set of sentence-level units after segmentation of the j-th candidate answer as . , Let represent the sentence-level unit index in the j-th candidate answer; then, the maximum value is taken as the best matching score.

[0081] The sentence consistency score of the i-th candidate answer relative to the j-th candidate answer can be expressed as:

[0082] The definition of local sentence uncertainty in the i-th candidate answer corresponds to the definition of local atomic fact uncertainty, the difference being that sentence-level units are used as the comparison objects here, i.e.:

[0083] Finally, the sentence-level uncertainty of the current input question is defined as:

[0084] In one set of implementations, sentence segmentation preferably uses a segmentation method based on inference step markers, such as using 'Stepi:' and 'Final Answer:' as natural boundaries; in another set of implementations, a general sentence segmenter can also be used. Preferably, to ensure subsequent token alignment and the traceability of the original text, the internal character order of candidate answers is not changed during sentence segmentation; segmentation boundaries are only set before 'Step i:' or 'Final Answer:', so that each sentence-level unit remains a substring of the original answer text.

[0085] Figure 7 This is a schematic diagram illustrating the calculation of uncertainty in atomic fact granularity and sentence granularity according to the present invention.

[0086] In a set of implementations, if the first The set of sentences obtained by segmenting the candidate answers is empty, or the first... If the set of sentences obtained from segmenting a candidate answer is empty, then the corresponding sentence-level similarity can be set to 0. Furthermore, for each sentence-level unit... In addition to recording its best matching sentence and maximum implied probability, it can also retain its relationship with candidate answers. The implied probabilities of all sentence units are listed to form a sentence-level matching matrix, which is used to explain which reasoning steps diverge across multiple samples.

[0087] In one implementation, the cleaned final answer is first extracted from each candidate answer. Then, a bidirectional implication judgment is performed on any two final answers; if and only if the implication condition is satisfied in both directions, the two final answers are determined to be semantically equivalent and are classified into the same semantically equivalent answer cluster.

[0088] Let the final set of answers after cleaning be:

[0089] in, Indicates the first The final answer for each candidate response. Bidirectional implication clustering can be implemented using the following pseudocode:

[0090] Input: cleaned answers Y = {y_1, y_2, ..., y_N}Initialize parent[i] =ifor i in 1..N: for j in i+1..N: label_ij = NLI(y_i, y_j) label_ji = NLI(y_j,y_i) if label_ij == entailment and label_ji == entailment: Union(i, j)for iin 1..N: cluster_id[i]= Find(i)Output: semantic_set_ids_entailment ={cluster_id[1], ..., cluster_id[N]} The aforementioned bidirectional entailment clustering method can aggregate candidate answers with the same final answer or semantic equivalence together to form multiple semantic clusters. Preferably, the bidirectional entailment judgment model adopts 'microsoft / deberta-large-mnli', which can output three types of results: "entailment / neutrality / contradiction". In this application implementation, the disjoint-set union operation is performed only when the predicted labels in both directions are 'entailment', but this application is not limited to this model.

[0091] Implementation of generating overall granularity uncertainty; In one implementation, the i-th candidate answer is first parsed to identify the structured segments corresponding to Step 1 to Step n and the Final Answer. The m-th step text of the i-th candidate answer is denoted as... For the text of this step, extract keywords and their corresponding contributions. If a step is marked as NO ANSWER, then that step will not participate in subsequent calculations. Keywords represent words, phrases, or sub-expressions that have a major impact on the reasoning conclusion in this step; keyword contribution indicates the degree of importance of the keyword's semantic contribution to the current step, assuming it is derived from the step text. The CCP extracted If there are 10 keywords, then the set of keywords corresponding to this step is denoted as:

[0092] in, Let the keyword be the u-th keyword extracted in the m-th step of the i-th candidate answer. Corresponding in the generated results If there are 10 tokens, then the set of probabilities corresponding to those tokens is denoted as:

[0093] in, Keywords The probability of generating the corresponding t-th token is calculated by taking the minimum probability of multiple tokens corresponding to the same keyword, and this minimum probability is used as the keyword-level probability representation of the keyword in this step.

[0094] Let the text of the m-th step of the i-th candidate answer be... For the keywords extracted in this step First, locate the start and end positions of the token in the corresponding token sequence of the step text. Then, backfill the original generated token sequence to obtain the token ID sequence corresponding to the keyword and the generation probability of each token. If the keyword fails to match in the corresponding step text, skip the keyword. For keywords that match successfully... Keywords The set of all token generation probabilities corresponding to the occurrence of this step, and This represents the keyword-level probability representation of the keyword in the current step, obtained by taking the minimum value of the token probability set.

[0095] In a preferred embodiment of this application, keyword confidence is achieved through a path of "aggregating repeated keywords first, then calculating the confidence of a single answer." Specifically, if the same keyword appears multiple times in different steps, the keyword-level probabilities obtained from each occurrence are first aggregated a second time. Let the keyword... Appeared in The probability of each occurrence of the corresponding keyword level is as follows: Then the aggregation probability after its repeated occurrence can be expressed as:

[0096]

[0097] Correspondingly, if the same keyword corresponds to multiple contribution scores in different steps... The aggregation contribution can then be obtained by taking their average value.

[0098] After completing the aggregation of duplicate keywords, the number of duplicate keywords is calculated based on the deduplicated keyword set. Confidence of each candidate answer:

[0099] in, Indicates the first The set of keywords after deduplication from the candidate answers. If the denominator is 0, it can be degenerated into taking the average of the probabilities of all valid keywords as the confidence backoff value.

[0100] The above implementation corresponds to the main aggregation path in the code: first, keywords and their contribution are extracted from each step; then, keyword probabilities are obtained based on precise token alignment; subsequently, aggregation is performed on the keyword text, and secondary aggregation of probabilities and averaging of contribution are performed on repeated keywords; finally, the result corresponding to a single candidate answer is obtained. Therefore, in this application Instead of simply averaging all the steps, the answer-level confidence score is calculated after "intra-step extraction, inter-step deduplication, and aggregation of duplicates".

[0101] After obtaining the confidence scores of all candidate answers, the confidence scores of candidate answers belonging to the same semantic cluster are summed according to semantically equivalent answer clusters to obtain the cluster score of the c-th semantic cluster:

[0102] in, Indicates the first Let K be the total number of semantic clusters. In one implementation, the overall generation granularity uncertainty can be defined as:

[0103]

[0104] Therefore, 'GU' simultaneously considers the reliability of key tokens within candidate answers and the distribution of candidate answers among semantically equivalent answer clusters.

[0105] In a preferred embodiment of this example, the keyword extraction model can be configured to retries multiple times; it is considered valid only if the returned result meets the following conditions: first, the output does not contain an additional 'Final Answer:' section; second, it can successfully parse the keywords and contribution corresponding to each step; and third, the number of steps parsed is consistent with the number of steps identified in the candidate answer. If a valid result cannot be obtained after multiple retries, the generation-level score of the candidate answer can be marked as invalid and removed or processed according to preset rules during the semantic cluster aggregation stage. Figure 4 The results of the generation-level signal combination experiment are presented. Semantic Entropy is used for semantic entropy, and SentenceSAR is used for discrimination and loss prevention. The experiment shows that the overall generation-granularity signal based on keyword token probability and semantic clustering distribution can effectively complement the fact-granularity and sentence-granularity.

[0106] The final implementation method incorporates uncertainties; In one implementation, there is uncertainty regarding the atomic fact granularity of the current input problem. Sentence granularity uncertainty and overall generation granularity uncertainty Taking the average directly yields the final problem-level uncertainty:

[0107] The final question-level uncertainty can be used to represent the overall uncertainty of the target large language model with respect to the current input question. The larger the value, the higher the instability of the model output in multiple samplings, multi-granularity structures, and overall answer distribution.

[0108] Figure 8 This is a schematic diagram illustrating the calculation of generation granularity uncertainty in this invention.

[0109] Prompt words and implementation details; In one implementation, multiple sampling generates prompt templates tailored to different task types. For chained reasoning tasks, the model can be required to output the reasoning process step by step and explicitly provide the final answer. To improve the readability of the patent document, only the typical prompt structure used in the experiments is shown below:

[0110] Please reason the following question step by step.Label eachreasoning step as "Step i:".Each step should build on previous steps andcontribute to the final answer.After all reasoning steps, output:"FinalAnswer: <answer>"Example:Question: Emily picked 36 apples. She gave 12 apples to her friendand then picked 9 more apples. How many apples does Emily havenow?Response: Let's think step by step.Step 1: Emily started with 36apples.Step 2: She gave away 12 apples, so 36 - 12 = 24.Step 3: She then picked 9 more apples, so 24 + 9 = 33.Final Answer: 33Question: John drives for 3 hours at 60 mph and then turns around. He spends the first 2 hours instandstill traffic, the next 0.5 hourat 30 mph, and the remaining time at 80mph. How far is he from home at the end of those 4 hours?Response: Let's thinkstep by step. For multi-hop question-and-answer tasks, the same step-by-step reasoning prompt can be used, only the question content needs to be replaced. For example, for the HotpotQA question "Pilot is the first episode of the drama developed by whom?", the above template can be reused directly, simply replacing the question with the corresponding multi-hop question.

[0111] For open-domain biography generation tasks, a concise query-based prompt can be used. An example is shown below:

[0112] Tell me a bio of Carolina Portesi Peroni concisely. In another set of implementations, to reduce the risk of fabricating uncertain knowledge, it can be further extended as follows: Please write a concise biography of {topic}.Keep the answer brief and factual.If uncertain about a fact, avoid fabrication and prefer conservative wording. For atomic fact extraction, a unified instruction template can be used to decompose the entire candidate answer into atomic facts and output them using fixed markers, as shown in the following example: You are given a sentence. Your task is to break the sentence downinto a list of atomic facts.An atomic fact is a sentence containing a singular piece of information.Each atomic fact in the outputted list should check a different piece of information.Output format:Facts: <begin>fact_1 <split>fact_2 <split>fact_3 <end> In one implementation, the atomic fact extraction prompt can be further supplemented with multiple few-shot examples to cover different types of responses such as open-domain question answering, mathematical reasoning, and multi-step logical reasoning. This allows the model to learn to break down complete reasoning answers into "independently decidable" atomic facts, rather than mechanically breaking them down line by line.

[0113] For keyword extraction, the model can be required to output keywords and their contribution scores step by step, for example: Extract relevant keywords from each reasoning step.Score eachkeyword's importance on a 1-10 scale.Format:keyword ( / score / )Separatemultiple keywords with ';'Use 'NO ANSWER' for non-contributing steps. The keyword extraction prompt can also incorporate the following constraints: First, the extracted keywords must be completely consistent with the original step text, including capitalization, numerical form, and mathematical symbols; second, if a step merely repeats the question conditions or makes no substantial contribution to the final answer, then 'Step i: NO ANSWER' should be returned; third, the contribution reflects both the semantic importance of the keyword in the current step and its degree of determination on the final answer. For different tasks such as mathematical reasoning, multi-hop question answering, and medical question answering, different one-shot examples can be configured to improve the consistency between keyword extraction and scoring.

[0114] Experimental and verification implementation methods; To verify the effectiveness of the method in this embodiment, experiments were conducted on chained reasoning tasks and open-domain long text generation tasks.

[0115] In a set of experiments, the following datasets can be selected: 1. GSM8K, used for mathematical reasoning and chain reasoning scenarios, 300 samples can be randomly selected in the experiment; 2. HotpotQA, used for multi-hop question answering scenarios, can randomly select 300 samples in the experiment; 3. BIOS, used for open-domain biography generation scenarios, can use 181 samples.

[0116] In a set of experiments, the following basic generative models can be selected:

[0117] In one implementation, semantic clustering uses 'microsoft / deberta-large-mnli', the atomic fact granularity and sentence granularity of the implication model use 'potsawee / deberta-v3-large-mnli', and the keyword extraction model can use 'Llama-3.1-8B-Instruct' or an equivalent instruction model. 'microsoft / deberta-large-mnli' can output three categories of labels: "implication / neutral / contradictory," suitable for bidirectional implication clustering; 'potsawee / deberta-v3-large-mnli' in this implementation is used to output two categories of probabilities: "implication / contradictory," and the implication probability is taken as the fragment support of AU / FU and SU.

[0118] In a set of experiments, the default sample size can be taken as... In the sample number sensitivity analysis, the sample number can be further set to... To observe the changing trends of problem-level uncertainty indicators under different sampling budgets.

[0119] For GSM8K, the numerical answer corresponding to the 'Final Answer' can be extracted from the greedy answers and compared with the standard answer to obtain the correct or incorrect label. For HotpotQA, the final answer text of the greedy answers can be normalized and matched with the reference answer to obtain the correct or incorrect label. For BIOS, FACTSCORE can be calculated from the greedy biography results, and the Pearson correlation coefficient between the uncertainty score and FACTSCORE can be calculated.

[0120] For GSM8K and HotpotQA, AUROC can be used as an uncertainty assessment metric. Preferably, the problem-level uncertainty score is used. As a ranking score, the degree of error in a greedy answer is defined as a binary label. In other words, it is obtained from multiple samplings. Used to estimate the model around the greedy response The uncertainty of the answer is addressed by taking the AUROC label from the final correctness of the greedy answer. Specifically, binary labels can be defined:

[0121]

[0122] In GSM8K, the final answer can be determined by extracting the value corresponding to the greedy answer's 'Final Answer' and comparing it with the standard numerical answer. In HotpotQA, the final answer to a greedy answer can be normalized for case, whitespace, and other necessary string standardization before being matched with the reference answer to determine its final value. Under this definition, if a certain problem The higher the value, the more likely it is to be ranked before the incorrect greedy samples, thus achieving a higher AUROC.

[0123] For BIOS, the Pearson correlation coefficient between the uncertainty score and FACTSCORE can be used as an evaluation metric. Preferably, following the approach in the original FactScore paper, each atomic fact in the generated biography can be verified to ensure it is supported by reliable knowledge sources. Specifically, a set of atomic facts can be extracted from the generated biography text first. Then, with the help of external evidence sources, retrieved documents, knowledge bases, or fact-checking modules, it is determined whether each atomic fact is supported. Let the first atomic fact be... The verification results of the atomic facts are as follows:

[0124]

[0125] Then the FACTSCORE corresponding to the biography text can be defined as:

[0126] In other words, a higher FACTSCORE indicates a higher proportion of supported facts in the generated biography, and thus stronger factual accuracy. This definition aligns with the basic idea of ​​the FactScore method, which "evaluates the factual accuracy of long texts based on atomic fact support rates." In this embodiment, since higher question-level uncertainty generally implies weaker factual accuracy, therefore... and The Pearson correlation coefficient is usually negative; the larger the absolute value of the negative correlation, the more stable the method of this embodiment is in characterizing factual risks.

[0127] Table 1 compares the main experiment results on different models and datasets. GSM8K and HotpotQA report the AUROC metric, while BIOS reports the PCC metric.

[0128] Table 1: Comparison of main experiment results on different models and datasets

[0129] The main experiment results show that the method in this embodiment exhibits strong stability across multiple models and datasets. In particular, when atomic fact granularity, sentence granularity, and overall generation granularity are fused together, the resulting question-level uncertainty score typically reflects the reliability of the model's current answer better than any single granularity; on most models, the fused granularity... Indicators are better than individual ones , or This indicates that the three granularity signals are significantly complementary.

[0130] Application and implementation methods; The method in this embodiment can be used for chained reasoning tasks, such as mathematical reasoning, multi-hop question answering, and logical deduction tasks, as well as for open-domain long text generation tasks, such as character profile generation, knowledge-based long answer generation, and complex explanatory text generation tasks.

[0131] In the application process, the final problem-level uncertainty score can be further used for: 1. High-risk issues trigger manual review; 2. Add risk warnings or confidence indicators to the generated results; 3. Sort and filter multiple candidate answers; 4. As a trigger condition for hallucination detection, automatic rejection, or secondary search; 5. As an auxiliary control signal in training, evaluation, or online inference systems.

[0132] The above embodiments highlight the core technical ideas of this application. Any variations, such as model replacement, sampling parameter adjustment, keyword extraction model replacement, clustering implementation method replacement, and fusion method, as long as they do not deviate from the multi-granularity hierarchical uncertainty assessment concept of this application, should fall within the protection scope of this invention.< / end> < / split> < / split> < / begin> < / answer> < / split> < / end> < / begin> < / split> < / end> < / begin>

Claims

1. A multi-granularity hierarchical uncertainty assessment method for long text generation, characterized in that, include: (1) Obtain the input question to be evaluated, and call the target large language model to sample the input question multiple times to generate N candidate answers; (2) Perform atomic fact extraction on N candidate answers and calculate the atomic fact uncertainty score based on the implication relationship of atomic facts to other candidate answers; (3) Perform sentence-level or reasoning step-level segmentation on N candidate answers, and calculate sentence uncertainty score based on the implication relationship of the maximum matching across answer sentences; (4) Extract the final cleaned answer from the N candidate answers, perform semantic clustering based on the bidirectional implication relation, and obtain multiple semantically equivalent answer clusters; (5) Extract keywords and keyword contribution from the reasoning steps of each candidate answer, and backfill the token probability corresponding to the keyword from the token probability sequence in the candidate answer generation process; (6) Take the minimum value of the multiple token probabilities corresponding to each keyword as the probability representation of the keyword, and then combine it with the keyword contribution to obtain the answer confidence of each candidate answer; (7) Aggregate the confidence scores of each candidate answer according to the semantically equivalent answer cluster to obtain the uncertainty score at the overall generation granularity; (8) Take the average of the atomic fact uncertainty score, sentence uncertainty score and overall generation uncertainty score to obtain the final uncertainty score of the current input question.

2. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 1, characterized in that, In step (1), the method of generating multiple samples is as follows: Set the sampling number, temperature parameter, and decoding control parameter of the target large language model. Call the target large language model to sample the input question multiple times to generate N candidate answers. For each candidate answer, save at least the following information: candidate answer text, candidate answer token sequence, and the conditional probability of each token at the time of generation. For chain reasoning tasks, further save the cleaned final answer text for subsequent semantic clustering.

3. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 2, characterized in that, In step (1), the N candidate answers are denoted as: ; Where, r i As candidate answers, Each candidate answer r i The system includes a corresponding token sequence and token conditional probabilities. The final output is not a score for a single answer, but rather a result reflecting the overall uncertainty of the current input question.

4. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 3, characterized in that, In step (2), the atomic fact extraction model is called for each candidate answer to split the candidate answer into several independent atomic facts. An atomic fact is a factual expression that contains a single information point and has a single truth value or consistency. Let the set of atomic facts extracted from the i-th candidate answer be: ; Among them, M i Let m represent the number of atomic facts corresponding to the i-th candidate answer, and m be the atomic fact index. ; For any atomic fact in the i-th candidate answer The implication model is used to calculate its overall text for the j-th candidate answer. The implied probability is denoted as: ; The atomic fact consistency score of the i-th candidate answer relative to the j-th candidate answer is expressed as: ; The local atomic fact uncertainty of the i-th candidate answer is defined as: ; Where N represents the total number of candidate answers generated by sampling the same input question; Ultimately, the uncertainty at the atomic fact granularity of the current input problem is defined as: 。 5. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 4, characterized in that, In step (3), each candidate answer is segmented according to the reasoning step structure. For chain-reasoning answers, the step-level segments Step1, Step2, and Final Answer are taken as sentence-level units. The set of sentence units after segmentation of the i-th candidate answer is denoted as: ; Among them, L i This represents the total number of sentence-level units obtained after segmenting the i-th candidate answer; For any sentence-level unit in the i-th candidate answer , Calculate the implication probability of each sentence-level unit in the j-th candidate answer, and denote the set of sentence-level units after segmentation of the j-th candidate answer as . , Let represent the sentence-level unit index in the j-th candidate answer; then, the maximum value is taken as the best matching score. ; The sentence consistency score of the i-th candidate answer relative to the j-th candidate answer is expressed as: ; The definition of local sentence uncertainty in the i-th candidate answer corresponds to the definition of local atomic fact uncertainty, the difference being that sentence-level units are used as the comparison objects here, i.e.: ; Finally, the sentence-level uncertainty of the current input question is defined as: 。 6. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 5, characterized in that, In step (4), the final cleaned answer is first extracted for each candidate answer. Then, a bidirectional implication judgment is performed on any two final answers. If and only if both directions satisfy the implication condition, the two final answers are determined to be semantically equivalent and are classified into the same semantically equivalent answer cluster. Let the final set of answers after cleaning be: ; in, Indicates the first The final answer for each candidate answer; Bidirectional implication clustering aggregates candidate answers with the same final answer or semantic equivalence together, thus forming multiple semantically equivalent answer clusters.

7. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 6, characterized in that, In step (6), the i-th candidate answer is first parsed to identify the structured segments corresponding to Step 1 to Step n and the Final Answer. The m-th step text of the i-th candidate answer is denoted as... For the text of this step, extract keywords and their corresponding contributions. If a step is marked as NO ANSWER, then that step will not participate in subsequent calculations. Keywords represent words, phrases, or sub-expressions that have a major impact on the reasoning conclusion in this step; keyword contribution indicates the degree of importance of the keyword's semantic contribution to the current step, assuming it is derived from the step text. The CCP extracted If there are 10 keywords, then the set of keywords corresponding to this step is denoted as: ; in, Let the keyword be the u-th keyword extracted in the m-th step of the i-th candidate answer. Corresponding in the generated results If there are 10 tokens, then the set of probabilities of the corresponding tokens is denoted as: ; in, Keywords The probability of generating the corresponding t-th token is calculated by taking the minimum probability of multiple tokens corresponding to the same keyword, and this minimum probability is used as the keyword-level probability representation of the keyword in this step. ; Let the text of the m-th step of the i-th candidate answer be... For the keywords extracted in this step First, locate the start and end positions of the token in the corresponding token sequence of the step text. Then, backfill the original generated token sequence to obtain the token ID sequence corresponding to the keyword and the generation probability of each token. If the keyword fails to match in the corresponding step text, skip the keyword. For keywords that match successfully... Keywords The set of all token generation probabilities corresponding to the occurrence of this step, and This represents the keyword-level probability representation of the keyword in the current step, obtained by taking the minimum value of the token probability set; Keyword confidence is achieved by first aggregating repeated keywords and then calculating the confidence of a single answer. Specifically, if the same keyword appears multiple times in different steps, the keyword-level probabilities obtained from each occurrence are first aggregated. Let H be the total number of occurrences of keyword k. k The probability of each occurrence of the corresponding keyword level is as follows: Then the aggregation probability after its recurrence is expressed as: ; Correspondingly, if the same keyword corresponds to multiple contribution scores in different steps... The aggregation contribution is obtained by taking their average value. ; After aggregating duplicate keywords, the confidence score of the i-th candidate answer is calculated based on the deduplicated keyword set. ; in, Let represent the set of keywords after deduplication in the i-th candidate answer. If the denominator is 0, it degenerates into taking the average of the probabilities of all valid keywords as the confidence backoff value.

8. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 7, characterized in that, In step (7), after obtaining the confidence scores of all candidate answers, the confidence scores of candidate answers belonging to the same semantic cluster are summed according to the semantically equivalent answer clusters to obtain the first... Cluster scores of semantic clusters: ; in, Indicates the first Let K be the total number of semantic clusters. Overall generation granularity uncertainty score Defined as: ; 。 9. The multi-granularity hierarchical uncertainty assessment method for long text generation as described in claim 8, characterized in that, In step (8), the uncertainty of the atomic fact granularity of the current input problem is addressed. Sentence granularity uncertainty and overall generation granularity uncertainty We directly take the average to obtain the final problem-level uncertainty score: ; The final question-level uncertainty score is used to represent the overall uncertainty of the target large language model for the current input question. The larger the value, the higher the instability of the model output in multiple samplings, multi-granularity structures, and overall answer distribution.

10. A multi-granularity hierarchical uncertainty assessment system for long text generation, characterized in that, include: The sampling and generation module is used to obtain the input question to be evaluated and call the target large language model to sample and generate N candidate answers multiple times. The atomic fact processing module is used to extract atomic facts from N candidate answers and calculate the atomic fact uncertainty score based on the implication relationship of atomic facts to other candidate answers. The sentence processing module is used to perform sentence-level or reasoning step-level segmentation on N candidate answers and calculate sentence uncertainty scores based on the implication relations of the maximum matching across answer sentences. The semantic clustering module is used to extract the final cleaned answer from N candidate answers. It performs semantic clustering based on bidirectional implication relations to obtain multiple semantically equivalent answer clusters. The keyword extraction and probability backfilling module is used to extract keywords and keyword contribution from the reasoning steps of each candidate answer, and backfill the token probability corresponding to the keyword from the token probability sequence in the candidate answer generation process; The generation-level aggregation module is used to take the minimum value of the probabilities of multiple tokens corresponding to each keyword as the probability representation of the keyword, and then combine the keyword contribution to obtain the answer confidence of each candidate answer. The confidence of each candidate answer is aggregated according to the semantically equivalent answer cluster to obtain the uncertainty score of the overall generation granularity. The fusion output module is used to average the uncertainty scores of atomic facts, sentences, and the overall generation uncertainty score to obtain the final uncertainty score of the current input question.