Multi-dimensional evaluation system and method for reasoning consistency of large language model

By constructing a multidimensional evaluation system and method, the problem of ignoring the reasoning consistency of large language models in existing technologies is solved, and a comprehensive and fine-grained evaluation of the model reasoning process and results is achieved, which improves the comprehensiveness and accuracy of the evaluation.

CN120632394APending Publication Date: 2025-09-12SHANXI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510762856.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing evaluation methods focus on the accuracy of large language models and ignore consistency, especially the consistency of the reasoning process, resulting in an inability to comprehensively evaluate the model's cognitive intelligence level.

Method used

A multidimensional evaluation system and method for the reasoning consistency of large language models was constructed, including a consistency identification module, an evaluation dataset generation module, and a consistency discrimination module. The reasoning consistency of the model was evaluated through a multi-dimensional and multi-granularity evaluation system. A three-stage generation strategy was adopted to construct the evaluation dataset, and the task features were automatically identified in combination with the task form to evaluate the consistency of the reasoning process and results.

Benefits of technology

It realizes multi-angle and fine-grained evaluation of large language models, and can evaluate the output consistency of the model from multiple dimensions such as factual consistency and logical consistency, thereby improving the comprehensiveness and accuracy of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632394A_ABST
    Figure CN120632394A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-dimensional evaluation system and method for reasoning consistency of a large language model, and belongs to the technical field of natural language processing. Aiming at the problem that the evaluation method is not comprehensive enough due to the fact that large language model consistency evaluation is limited to reasoning result and fact consistency, a system comprising a consistency identification module, an evaluation data set generation module, a consistency judgment module and a consistency evaluation module is designed; the method comprises the following steps: analyzing and determining a consistency type required to be checked in combination with a task form, generating an instruction based on the consistency type required to be checked in combination with an existing data sample, automatically constructing an evaluation data set, and automatically evaluating the consistency type required to be evaluated based on the constructed data set according to the consistency type required to be evaluated and logic requirements thereof. And evaluating the consistency of a reasoning process and a reasoning result generated by the large language model on the constructed consistency evaluation data set, and giving a model consistency capability comprehensive evaluation report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a multidimensional evaluation system and method for reasoning consistency of a large language model. Background Art

[0002] Large language models, due to their powerful semantic understanding and complex reasoning capabilities, are widely used in high-value scenarios such as machine translation and intelligent question-answering. However, large language models still face the challenge of inconsistency. Specifically, large language models often produce "hallucinatory" outputs that are inconsistent with common facts or the given context, a phenomenon known as factual inconsistency. When faced with similar questions, large language models often output contradictory results, a phenomenon known as logical inconsistency. This inconsistency not only affects the reliability of large language models but also limits their application in some key areas.

[0003] Researchers have evaluated the various performance characteristics of large language models, but existing evaluation methods focus on factual consistency, that is, evaluating the accuracy of model outputs. These methods use a single evaluation metric and ignore the evaluation of consistency across multiple model outputs. Furthermore, existing evaluation methods focus on inference results and lack an evaluation of the consistency of the model's inference process, making it impossible to obtain fine-grained consistency evaluation information. This makes evaluation methods for the inference consistency of large language models incomplete, and language intelligence evaluation indicators cannot effectively evaluate the model's cognitive intelligence level from a consistency perspective. Therefore, how to build a comprehensive evaluation system to achieve a multi-angle and systematic evaluation of the inference consistency of large language models has become a pressing issue. Summary of the Invention

[0004] In response to the problem that the evaluation of large language models focuses on accuracy but ignores consistency, and the evaluation method focuses on the inference results but ignores the inference process, which leads to the problem that the evaluation method is not comprehensive, the present invention provides a multi-dimensional evaluation system and method for the inference consistency of large language models.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions:

[0006] A multi-dimensional evaluation system for large language model reasoning consistency, comprising a consistency identification module, an evaluation data set generation module, a consistency determination module, and a consistency evaluation module;

[0007] The consistency identification module combines task form analysis to determine the consistency type that the large language model must meet. While existing evaluation methods can detect consistency along a single dimension, they struggle to adapt to the diverse needs of different task scenarios. This module primarily automatically identifies task features based on task form, matching them to detection dimensions such as factual consistency and logical consistency.

[0008] The evaluation data set generation module generates instructions based on the consistency type to be verified and in combination with existing data samples, and automatically constructs the evaluation data set.

[0009] The consistency judgment module: based on the constructed evaluation data set, according to the consistency type and definition required to be evaluated, evaluates the consistency of the reasoning process and reasoning results generated by the large language model based on the constructed evaluation data set.

[0010] The consistency evaluation module provides a comprehensive assessment report on the consistency capabilities of large language models. Breaking away from the single scoring model of existing methods, this module implements a multi-dimensional and multi-granular evaluation of the output consistency of large language models, forming a three-level evaluation system: task-level, attribute-level, and system-level.

[0011] The evaluation dataset generation module adopts a three-stage generation strategy: 1) primary instruction generation: using a large language model to automatically generate instructions; 2) feedback optimization: feeding back the quality of data generated based on the instructions to the original model, and optimizing the original instructions accordingly; 3) dataset generation: constructing the dataset required for consistency evaluation based on the original dataset and instructions.

[0012] A multi-dimensional evaluation method for reasoning consistency of a large language model, comprising the following steps:

[0013] Step 1: The task text dataset Input the large language model, combine various consistency categories and definitions, and prompt the large language model to determine the categories that meet the consistency requirements, forming a set of consistency categories to be tested;

[0014] Step 2: The original dataset Task text in Input the large language model, perturb the original input according to the logical consistency property to be detected, and obtain the input text pairs that meet the logical consistency property;

[0015] Step 3: Input the obtained text pair into the large language model, requiring the large language model to provide the corresponding reasoning process while outputting the answer. Combined with the logical consistency attribute, it is judged in turn whether the reasoning process and results output by the large language model meet the various consistency types required for evaluation;

[0016] Step 4: Combine the performance of the large language model on various tasks to produce a large language model consistency capability assessment report, including forming a three-level evaluation system at the task level, attribute level, and system level from the perspectives of reasoning process and reasoning results.

[0017] Furthermore, the consistency types in step 1 include two major categories:

[0018] Factual consistency and logical consistency ;

[0019] Factual consistency means that the output of the large language model satisfies the facts and does not contradict world knowledge; logical consistency means that there can be no logical contradictions between the multiple outputs of the large language model;

[0020] According to different logical relationships, it is further divided into negative consistency , symmetric consistency , transitive consistency , additive consistency ;

[0021] Since different tasks satisfy different types of logical consistency, it is necessary to combine the task form to determine the type of consistency it satisfies and evaluate the consistency of large language model reasoning accordingly.

[0022] The prompts must include the following information: (1) detailed definitions of factual consistency and logical consistency; (2) relevant examples of factual consistency and logical consistency; (3) task format; and (4) relevant instructions for determining the type of consistency.

[0023] The input text pair in step 2 needs to meet the following requirements:

[0024] For negative consistency, it is necessary to Construct its corresponding negative form text , forming an input text pair ;

[0025] For symmetric consistency, it is necessary to base it on the original task text. Perturb the order to form the original task text Symmetrical text , and form a text pair ,in, and Represents text Two orthogonal semantic units that meet symmetric consistency, such as in the textual entailment task, and Corresponding to the premise and assumption parts respectively;

[0026] For transitive consistency, it is necessary to satisfy Original task text and Recombining to obtain , and form text pairs ;in, 、 and 、 Text and Orthogonal semantic units that satisfy transitive consistency; Represents text-based and New task text constructed to test transfer consistency;

[0027] For additive consistency, it is necessary to base it on the original text and Constructing new text , and form a text pair ;in, Represents text-based and A new task text constructed to test additive consistency;

[0028] A three-stage generation strategy is used when constructing the text pairs. The three-stage generation strategy is as follows: 1) Primary instruction generation: using the large language model to automatically generate instructions based on the consistency definition. 2) Feedback optimization: using text pairs that meet the requirements as context examples, the large language model combines the primary instructions and imitates the context examples to generate text pairs that meet the requirements. The quality of the generated text pairs is evaluated. If the qualified rate of the text pairs is lower than , this is fed back to the model to adjust the primary instructions and obtain the adjusted optimal instructions. The examples used in this step can be manually generated, or a large language model can be used to generate candidate text pairs with zero-shot prompts and combined with manually screened qualified text pairs as contextual examples. 3) Dataset Construction: Generate the required evaluation dataset based on the optimized instructions.

[0029] The standard for consistency capability evaluation in step 4 is: Represents the large language model under test, Represents the large language model under test For the prediction of inference results, Represents the large language model under test explanations generated simultaneously with predictions;

[0030] For factual consistency, the reasoning process and results output by the large language model must be consistent with world knowledge;

[0031] In logical consistency, for negation consistency, the large language model is used for text pairs. The output must meet , that is, the task text When negated, the prediction of the large language model is opposite to the original one, while the interpretation remains semantically unchanged; Represents the language model to be tested; and Represent the models under test For task text The reasoning results and corresponding explanations given; and Represent the models under test For task text The reasoning results and corresponding explanations given; Indicates logical negation, that is, the truth or false value of a proposition is negated;

[0032] For symmetric consistency, large language models are used for text pairs. The output should satisfy , that is, the prediction results of the large language model remain unchanged after the input order is swapped, and the explanation should remain semantically unchanged; among them, and Represent the models under test For task text The reasoning results and corresponding explanations given;

[0033] For transitive consistency, large language models are used for text pairs. The output must meet , that is, the large language model is The prediction results should satisfy the task-related logical transfer relationship, and the explanation should be and Zhongyu and relevant parts; among which, and Represent the models under test For task text The given reasoning results and the corresponding reasoning process; and Represents the model under test For task text The given reasoning results and the corresponding reasoning process; and Represents the model under test For task text The given reasoning results and the corresponding reasoning process;

[0034] Taking textual implication as an example, when and When both represent implication relations, It should also be an implication relationship; when and When they are respectively implication and contradiction relations, It should be a contradictory relationship, etc.

[0035] For additive consistency, the large language model is used for text pairs. The output should satisfy , that is, the large language model under test is The prediction results should be consistent with the and Remain consistent when forecasting alone, while explaining Should be compared with large language models and The union of the separately produced interpretations remains semantically unchanged.

[0036] Methods for determining semantic equivalence of explanations: When evaluating the consistency of explanations, it is necessary to determine whether two explanations are semantically equivalent. Here, we introduce the methods for determining this. First, the longer explanation text is divided into finer-grained atomic units. These atomic units are the smallest indivisible semantic units in the text, that is, the smallest factual units. Dependency parsing can be used to extract the subject-verb-object structure, or a trained large language model for sequence annotation can be used to annotate the boundaries of the atomic units (B / I / O labels), or context learning can be combined with a large language model. On this basis, each atomic unit is judged in turn to determine whether there is a contradiction, thereby improving the accuracy of contradiction detection. This process can be achieved using a trained large language model for textual implication.

[0037] The three-level evaluation system in step 4 is: Composition set , The consistency type used for evaluation is indicated by the indicator vector Indicates that, , To represent a dataset Whether it is used to verify the consistency of facts, To represent a dataset Whether to use it to check for negation consistency, To represent a dataset Whether to use it to check symmetry consistency, To represent a dataset Whether it is used to check the delivery consistency, To represent a dataset whether to test additive consistency;

[0038] Assuming that the original dataset is used Evaluation consistency type When the large language model For text The resulting reasoning process and reasoning results are expressed as and , the consistency index of the inference results and inference process generated by the model is recorded as and , then for the dataset , fact consistency index of the reasoning results of the tested model and factual consistency index of the reasoning process The calculation is as follows:

[0039]

[0040]

[0041] in, is the gold standard answer for the original dataset, Gold standard explanations of the original dataset or world knowledge gained through retrieval, Representation dataset The amount of data in It is a semantic equivalence discrimination language model used to indicate and Are they semantically equivalent?

[0042] The reasoning results of the tested model negate the consistency index and negation consistency index of the reasoning process The calculation is as follows:

[0043]

[0044]

[0045] Symmetric consistency index of the inference results of the tested model Symmetric consistency index of the reasoning process The calculation is as follows:

[0046]

[0047]

[0048] The consistency index of the inference results of the tested model and transitive consistency index of the reasoning process The calculation is as follows:

[0049]

[0050]

[0051] Additive consistency index of the inference results of the tested model and additive consistency index of the reasoning process The calculation is as follows:

[0052]

[0053]

[0054] At the task level, for the task dataset , the consistency performance score of the large language model is expressed as:

[0055]

[0056]

[0057] in, and Respectively in the task dataset The ability of the model under test to produce consistent reasoning results and reasoning processes;

[0058] At the attribute level, for This consistency,consistency performance score of the large language model is expressed as:

[0059]

[0060]

[0061] in, and Respectively expressed in This type of consistency refers to the ability of the tested model to produce consistent reasoning results and reasoning processes. Indicates the number of data sets;

[0062] At the large language model level, the overall consistency score of the evaluated large language model is expressed as:

[0063]

[0064]

[0065] in, and They represent the overall ability of the tested model to produce consistent reasoning results and the reasoning process, respectively.

[0066] Compared with the prior art, the present invention has the following advantages:

[0067] The method proposed in this paper can measure the consistency of large language model outputs from multiple dimensions, including factual consistency and various logical consistency. Furthermore, by evaluating the consistency of the interpretations of large language model outputs, the consistency of large language model reasoning can be evaluated at a more granular level, beyond the inference results. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1This is a flowchart of an evaluation system for consistency of reasoning of large language models;

[0069] Figure 2 Generate module diagrams for the evaluation dataset;

[0070] Figure 3 Schematic diagram of the evaluation method framework for consistency of reasoning of large language models. DETAILED DESCRIPTION

[0071] To gain a deeper understanding of the present invention, we will provide a comprehensive and detailed description thereof. However, the present invention has various implementations and is not limited to the specific examples listed herein. These examples are presented to enhance a comprehensive understanding of the present disclosure.

[0072] A multi-dimensional evaluation system for large language model reasoning consistency, comprising a consistency identification module, an evaluation data set generation module, a consistency determination module, and a consistency evaluation module;

[0073] The consistency identification module combines task form analysis to determine the consistency type that the large language model must meet. While existing evaluation methods can detect consistency along a single dimension, they struggle to adapt to the diverse needs of different task scenarios. This module primarily automatically identifies task features based on task form, matching them to detection dimensions such as factual consistency and logical consistency.

[0074] The evaluation data set generation module generates targeted evaluation instructions based on the consistency type to be tested and in combination with existing data samples, and automatically constructs the evaluation data set.

[0075] The consistency judgment module: based on the constructed evaluation data set, according to the consistency type and definition required to be evaluated, evaluates the consistency of the reasoning process and reasoning results generated by the large language model based on the constructed evaluation data set.

[0076] The consistency evaluation module provides a comprehensive assessment report on the consistency capabilities of large language models. Breaking away from the single scoring model of existing methods, this module implements a multi-dimensional and multi-granular evaluation of the output consistency of large language models, forming a three-level evaluation system: task-level, attribute-level, and system-level.

[0077] The evaluation dataset generation module adopts a three-stage generation strategy: 1) primary instruction generation: using a large language model to automatically generate instructions; 2) feedback optimization: feeding back the quality of data generated based on the instructions to the original model, and optimizing the original instructions accordingly; 3) dataset generation: constructing the dataset required for consistency evaluation based on the original dataset and instructions.

[0078] A multi-dimensional evaluation method for reasoning consistency of a large language model, comprising the following steps:

[0079] Step 1: The task text dataset Input the large language model, combine various consistency categories and definitions, and prompt the large language model to determine the categories that meet the consistency requirements, forming a set of consistency categories to be tested;

[0080] Step 2: The original dataset Text in Input the large language model, perturb the original input according to the logical consistency property to be detected, and obtain the input text pairs that meet the logical consistency property;

[0081] Step 3: Input the obtained text pair into the large language model, requiring the large language model to provide the corresponding reasoning process while outputting the answer. Combined with the logical consistency attribute, it is judged in turn whether the reasoning process and results output by the large language model meet the various consistency types required for evaluation;

[0082] Step 4: Combine the performance of the large language model on various tasks to produce a large language model consistency capability assessment report, including forming a three-level evaluation system at the task level, attribute level, and system level from the perspectives of reasoning process and reasoning results.

[0083] Furthermore, the consistency types in step 1 include two major categories:

[0084] Factual consistency and logical consistency ;

[0085] Factual consistency means that the output of the large language model satisfies the facts and does not contradict world knowledge; logical consistency means that there can be no logical contradictions between the multiple outputs of the large language model;

[0086] According to different logical relationships, it is further divided into negative consistency , symmetric consistency , transitive consistency , additive consistency ;

[0087] Since different tasks satisfy different types of logical consistency, it is necessary to combine the task form to determine the type of consistency it satisfies and evaluate the consistency of large language model reasoning accordingly.

[0088] The prompts must include the following information: (1) detailed definitions of factual consistency and logical consistency; (2) relevant examples of factual consistency and logical consistency; (3) task format; and (4) relevant instructions for determining the type of consistency.

[0089] The input text pair in step 2 needs to meet the following requirements:

[0090] For negative consistency, it is necessary to Construct its corresponding negative form text , forming an input text pair ;

[0091] For symmetric consistency, it is necessary to base it on the original input text Perturb the order to form a new symmetrical text , and form a text pair ,in, and Represents text Two orthogonal semantic units that meet symmetric consistency, such as in the textual entailment task, and Corresponding to the premise and assumption parts respectively;

[0092] For transitive consistency, it is necessary to satisfy The original input text and Recombining to obtain , and form text pairs ;in, 、 and 、 Text and Orthogonal semantic units that satisfy transitive consistency; Represents text-based and New task text constructed to test transfer consistency;

[0093] For additive consistency, it is necessary to base it on the original text and Constructing new text , and form a text pair ;in, Represents text-based and A new task text constructed to test additive consistency;

[0094] A three-stage generation strategy is used when constructing the text pairs. The three-stage generation strategy is as follows: 1) Primary instruction generation: using the large language model to automatically generate instructions based on the consistency definition. 2) Feedback optimization: using text pairs that meet the requirements as context examples, the large language model combines the primary instructions and imitates the context examples to generate text pairs that meet the requirements. The quality of the generated text pairs is evaluated. If the qualified rate of the text pairs is lower than , this is fed back to the model to adjust the primary instructions and obtain the adjusted optimal instructions. The examples used in this step can be manually generated, or a large language model can be used to generate candidate text pairs with zero-shot prompts and combined with manually screened qualified text pairs as contextual examples. 3) Dataset Construction: Generate the required evaluation dataset based on the optimized instructions.

[0095] The standard for consistency capability evaluation in step 4 is: Represents the large language model under test, Represents the large language model under test For the prediction of inference results, Represents the large language model under test explanations generated simultaneously with predictions;

[0096] For factual consistency, the reasoning process and results output by the large language model must be consistent with world knowledge;

[0097] In logical consistency, for negation consistency, the large language model is used for text pairs. The output must meet , that is, the task text When negated, the prediction of the large language model is opposite to the original one, while the interpretation remains semantically unchanged; Represents the large language model to be tested; and Represent the models under test For task text The reasoning results and corresponding explanations given; and Represent the models under test For task text The reasoning results and corresponding explanations given; Indicates logical negation, that is, the truth or false value of a proposition is negated;

[0098] For symmetric consistency, large language models are used for text pairs. The output should satisfy , that is, the prediction results of the large language model remain unchanged after the input order is swapped, and the explanation should remain semantically unchanged; among them, and Represent the models under test For task text The reasoning results and corresponding explanations given;

[0099] For transitive consistency, large language models are used for text pairs. The output must meet , that is, the large language model is The prediction results should satisfy the task-related logical transfer relationship, and the explanation should be and Zhongyu and relevant parts; among which, and Represent the models under test For task text The given reasoning results and the corresponding reasoning process; and Represents the model under test For task text The given reasoning results and the corresponding reasoning process; and Represents the model under test For task text The given reasoning results and the corresponding reasoning process;

[0100] Taking textual implication as an example, when and When both represent implication relations, It should also be an implication relationship; when and When they are respectively implication and contradiction relations, It should be a contradictory relationship, etc.

[0101] For additive consistency, the large language model is used for text pairs. The output should satisfy , that is, the large language model is The prediction results should be consistent with the and Remain consistent when forecasting alone, while explaining Should be compared with large language models and The union of the separately produced interpretations remains semantically unchanged.

[0102] Methods for determining semantic equivalence of explanations: When evaluating the consistency of explanations, it is necessary to determine whether two explanations are semantically equivalent. Here, we introduce the methods for determining this. First, the longer explanation text is divided into finer-grained atomic units. These atomic units are the smallest indivisible semantic units in the text, that is, the smallest factual units. Dependency parsing can be used to extract the subject-verb-object structure, or a trained large language model for sequence annotation can be used to annotate the boundaries of the atomic units (B / I / O labels), or context learning can be combined with a large language model. On this basis, each atomic unit is judged in turn to determine whether there is a contradiction, thereby improving the accuracy of contradiction detection. This process can be achieved using a trained large language model for textual implication.

[0103] The three-level evaluation system in step 4 is: Composition set , The consistency type used for evaluation is indicated by the indicator vector Indicates that, , To represent a dataset Whether it is used to verify the consistency of facts, To represent a dataset Whether to use it to check for negation consistency, To represent a dataset Whether to use it to check symmetry consistency, To represent a dataset Whether it is used to check the delivery consistency, To represent a dataset whether to test additive consistency;

[0104] Assuming that the original dataset is used Evaluation consistency type When the large language model For text The resulting reasoning process and reasoning results are expressed as and , the consistency index of the inference results and inference process generated by the model is recorded as and , then for the dataset , fact consistency index of the reasoning results of the tested model and factual consistency index of the reasoning process The calculation is as follows:

[0105]

[0106]

[0107] in, is the gold standard answer for the original dataset, Gold standard explanations of the original dataset or world knowledge gained through retrieval, Representation dataset The amount of data in It is a semantic equivalence discrimination language model used to indicate and Are they semantically equivalent?

[0108] The reasoning results of the tested model negate the consistency index and negation consistency index of the reasoning process The calculation is as follows:

[0109]

[0110]

[0111] Symmetric consistency index of the inference results of the tested model Symmetric consistency index of the reasoning process The calculation is as follows:

[0112]

[0113]

[0114] The consistency index of the inference results of the tested model and transitive consistency index of the reasoning process The calculation is as follows:

[0115]

[0116]

[0117] Additive consistency index of the inference results of the tested model and additive consistency index of the reasoning process The calculation is as follows:

[0118]

[0119]

[0120] At the task level, for the task dataset , the consistency performance score of the large language model is expressed as:

[0121]

[0122]

[0123] in, and Respectively in the task dataset The ability of the model under test to produce consistent reasoning results and reasoning processes;

[0124] At the attribute level, for This consistency,consistency performance score of the large language model is expressed as:

[0125]

[0126] in, and Respectively expressed in This type of consistency refers to the ability of the tested model to produce consistent reasoning results and reasoning processes. Indicates the number of data sets;

[0127] At the large language model level, the overall consistency score of the evaluated large language model is expressed as:

[0128]

[0129] in, and They represent the overall ability of the tested model to produce consistent reasoning results and the reasoning process, respectively.

[0130] Any matters not described in detail in this specification are prior art known to those skilled in the art. Although the above description of the present invention is based on specific embodiments to facilitate understanding of the present invention by those skilled in the art, it should be understood that the present invention is not limited to the scope of the specific embodiments. As long as various modifications are within the spirit and scope of the present invention as defined and determined by the appended claims, such modifications will be obvious to those skilled in the art, and all inventions and creations utilizing the concepts of the present invention are protected.

Claims

1. A multi-dimensional evaluation system for the consistency of reasoning of large language models, characterized by: The system includes a consistency identification module, an evaluation data set generation module, a consistency determination module and a consistency evaluation module; The consistency identification module determines the consistency type that the large language model needs to meet by combining task form analysis; The evaluation data set generation module generates targeted evaluation instructions based on the consistency type to be tested and in combination with existing data samples, and automatically constructs the evaluation data set; The consistency determination module: based on the constructed evaluation data set, according to the consistency type and definition required to be evaluated, evaluates the consistency of the reasoning process and reasoning results generated by the large language model based on the constructed evaluation data set; The consistency evaluation module provides a comprehensive evaluation report on the consistency capability of the large language model.

2. A multi-dimensional evaluation system for large language model reasoning consistency according to claim 1, characterized in that: The evaluation dataset generation module adopts a three-stage generation strategy: 1) primary instruction generation: using a large language model to automatically generate instructions; 2) feedback optimization: feeding back the quality of data generated based on the instructions to the original model, and optimizing the original instructions accordingly; 3) dataset generation: constructing the dataset required for consistency evaluation based on the original dataset and instructions.

3. A multi-dimensional evaluation method for the consistency of reasoning in large language models, characterized by: The method comprises the following steps: Step 1: The task text dataset Input the large language model, combine various consistency categories and definitions, and prompt the large language model to determine the categories that meet the consistency requirements, forming a set of consistency categories to be tested; Step 2: The original dataset Text in Input the large language model, perturb the original input according to the logical consistency property to be detected, and obtain the input text pairs that meet the logical consistency property; Step 3: Input the obtained text pair into the large language model, requiring the large language model to provide the corresponding reasoning process while outputting the answer. Combined with the logical consistency attribute, it is judged in turn whether the reasoning process and results output by the large language model meet the various consistency types required for evaluation; Step 4: Combine the performance of the large language model on various tasks to produce a large language model consistency capability assessment report, including forming a three-level evaluation system at the task level, attribute level, and system level from the perspectives of reasoning process and reasoning results.

4. The multi-dimensional evaluation method for large language model reasoning consistency according to claim 3, characterized in that: The consistency types in step 1 include two categories: Factual consistency and logical consistency ; Factual consistency means that the output of the large language model satisfies the facts and does not contradict world knowledge; logical consistency means that there can be no logical contradictions between the multiple outputs of the large language model; According to different logical relationships, it is further divided into negative consistency , symmetric consistency , transitive consistency , additive consistency ; The prompts must include the following information: (1) detailed definitions of factual consistency and logical consistency; (2) relevant examples of factual consistency and logical consistency; (3) task format; and (4) relevant instructions for determining the type of consistency.

5. The multi-dimensional evaluation method for large language model reasoning consistency according to claim 3, characterized in that: The input text pair in step 2 needs to meet the following requirements: For negative consistency, it is necessary to Construct its corresponding negative form text , forming an input text pair ; For symmetric consistency, it is necessary to base it on the original input text Perturb the order to form the original text Symmetrical text , and form a text pair ,in, and Represents text Two orthogonal semantic units that satisfy symmetric consistency; For transitive consistency, it is necessary to satisfy The original input text and Recombining to obtain , and form text pairs ;in, 、 and 、 Text and Orthogonal semantic units that satisfy transitive consistency; Represents text-based and New task text constructed to test transfer consistency; For additive consistency, it is necessary to base it on the original text and Constructing new text , and form a text pair ;in, Represents text-based and A new task text constructed to test additive consistency; A three-stage generation strategy is adopted when constructing the text pairs; the three-stage generation strategy is: 1) primary instruction generation; 2) feedback optimization; 3) data set construction.

6. The multi-dimensional evaluation method for large language model reasoning consistency according to claim 3, characterized in that: The standard for consistency capability evaluation in step 4 is: Represents the large language model under test, Representing a large language model For the prediction of inference results, Represents the large language model under test explanations generated simultaneously with predictions; For factual consistency, the reasoning process and results output by the large language model must be consistent with world knowledge; In logical consistency, for negation consistency, the large language model is used for text pairs. The output must meet , that is, the task text When negated, the prediction of the large language model is opposite to the original one, while the interpretation remains semantically unchanged; Represents the language model to be tested; and Represent the models under test For task text The reasoning results and corresponding explanations given; and Represent the models under test For task text The reasoning results and corresponding explanations given; Indicates logical negation, that is, the truth or false value of a proposition is negated; For symmetric consistency, large language models are used for text pairs. The output should satisfy , that is, the prediction results of the large language model remain unchanged after the input order is swapped, and the explanation should remain semantically unchanged; among them, and Represent the models under test For task text The reasoning results and corresponding explanations given; For transitive consistency, large language models are used for text pairs. The output must meet , that is, the large language model is The prediction results should satisfy the task-related logical transfer relationship, and the explanation should be and Zhongyu and relevant parts; among which, and Represent the models under test For task text The given reasoning results and the corresponding reasoning process; and Represents the model under test For task text The given reasoning results and the corresponding reasoning process; and Represents the model under test For task text The given reasoning results and the corresponding reasoning process; For additive consistency, the large language model is used for text pairs. The output should satisfy , that is, the large language model is The prediction results should be consistent with the and Remain consistent when forecasting alone, while explaining Should be compared with large language models and The union of the separately produced interpretations remains semantically unchanged.

7. The multi-dimensional evaluation method for large language model reasoning consistency according to claim 3, characterized in that: The three-level evaluation system in step 4 is: Composition set , The consistency type used for evaluation is indicated by the indicator vector Indicates that, , To represent a dataset Whether it is used to verify the consistency of facts, To represent a dataset Whether to use it to check for negation consistency, To represent a dataset Whether to use it to check symmetry consistency, To represent a dataset Whether it is used to check the delivery consistency, To represent a dataset whether to test additive consistency; Assuming that the original dataset is used Evaluation consistency type When the large language model For text The resulting reasoning process and reasoning results are expressed as and , the consistency index of the inference results and inference process generated by the model is recorded as and , then for the dataset , fact consistency index of the reasoning results of the tested model and factual consistency index of the reasoning process The calculation is as follows: , ,in, is the gold standard answer for the original dataset, Gold standard explanations of the original dataset or world knowledge gained through retrieval, Representation dataset The amount of data in It is a semantic equivalence discrimination language model used to indicate and Are they semantically equivalent? The reasoning results of the tested model negate the consistency index and negation consistency index of the reasoning process The calculation is as follows: , , the symmetric consistency index of the inference results of the tested model Symmetric consistency index of the reasoning process The calculation is as follows: , , the consistency index of the inference results of the tested model and transitive consistency index of the reasoning process The calculation is as follows: , , additive consistency index of the inference results of the tested model and additive consistency index of the reasoning process The calculation is as follows: , , at the task level, for the task dataset , the consistency performance score of the large language model is expressed as: , ,in, and Respectively in the task dataset The ability of the model under test to produce consistent reasoning results and reasoning processes; At the attribute level, for This consistency,consistency performance score of the large language model is expressed as: , ,in, and Respectively expressed in This type of consistency refers to the ability of the tested model to produce consistent reasoning results and reasoning processes. Indicates the number of data sets; At the large language model level, the overall consistency score of the evaluated large language model is expressed as: , ,in, and They represent the overall ability of the tested model to produce consistent reasoning results and the reasoning process, respectively.

Citation Information

Cited By

  • Method for measuring model value reasoning reliability

    CN121144067A