Method and device for evaluating planning capability of large language model, electronic equipment, storage medium and computer program product

By calculating the causal impact of each layer extraction rate, detection accuracy, information flow score and historical steps of the large language model, the problem of difficulty in determining the prospective planning ability of the large language model is solved, and a multi-directional proof of the model planning ability is achieved.

CN120011770AInactive Publication Date: 2025-05-16INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510140942.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

It is difficult for the prior art to determine whether large language models have forward-looking planning capabilities.

Method used

By obtaining multiple samples, inputting large language models, calculating the extraction rate and detection accuracy of each layer, evaluating the causal impact of information flow scores and historical steps, and comprehensively evaluating the planning ability of large language models.

Benefits of technology

The proof of the forward-looking planning ability of large language models in multiple directions is realized, and the explanation is provided that the model has short-term forward-looking future decision-making ability in global observable planning tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011770A_ABST
    Figure CN120011770A_ABST
Patent Text Reader

Abstract

The invention relates to a big language model planning capability evaluation method and device, electronic equipment, a storage medium and a computer program product, and the method comprises the steps: inputting a plurality of samples into a big language model, obtaining a representation vector of each sample in each layer, and calculating the extraction rate and detection accuracy of the layer; calculating an information flow score of each type of component contained in each sample, and evaluating the possibility that the type of component is used as an information source; and obtaining a shielding prediction result after the operation result of the target execution operation contained in each sample is shielded and an unshielded prediction result before shielding, and evaluating the influence of the target execution operation on an output result. Therefore, by calculating the extraction rate, the detection accuracy, the information flow score and the causality influence of the historical steps of the model, theoretical support is provided for the model to have the interpretability of the short-term prospective future decision-making ability in the global observable planning task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and more specifically, to a method, device, electronic device, storage medium, and computer program product for evaluating the planning ability of a large language model. Background Art

[0002] Planning is the process of formulating a series of actions to transform a given initial state into a desired target state. As a core module of intelligent agents, planning has been widely used in many fields, such as embodied intelligence, network navigation, tool use, etc.

[0003] In related technologies, researchers are committed to stimulating and evaluating the planning ability of large language models. For example, they proposed hint engineering and instruction fine-tuning to improve the planning ability of large language models. In addition, some researchers have built benchmarks to evaluate the planning ability of large language models, etc. However, although researchers have made some progress in the above areas, the underlying mechanism behind the planning ability of large language models is still largely an unexplored frontier. Therefore, exploring the potential planning mechanism of large language models to determine whether they have forward-looking planning capabilities has become a problem that needs to be solved urgently. Summary of the invention

[0004] The present disclosure provides a method, device, electronic device, storage medium and computer program product for evaluating the planning capability of a large language model, so as to at least solve the problem in the above-mentioned related technologies that it is difficult to determine whether a large language model has forward-looking planning capabilities.

[0005] According to a first aspect of an embodiment of the present disclosure, a method for evaluating the planning ability of a large language model is provided, comprising: obtaining a plurality of samples, wherein each sample contains a plurality of types of components, and the sample is a text type sample; inputting the plurality of samples into the large language model to obtain a representation vector of each sample in each of a plurality of layers contained in the large language model; based on the representation vector of each layer, calculating the extraction rate and detection accuracy of the layer as a first evaluation parameter, wherein the extraction rate is used to characterize the possibility that the decoding head of the layer accurately predicts the operation result of a single execution operation among a plurality of execution operations contained in the each sample, and the detection accuracy is used to characterize the possibility that the probe of the layer accurately predicts a plurality of operation results corresponding one to one to a plurality of execution operations contained in the each sample; for For each sample, calculate the information flow score of each type of component among the multiple types of components contained in the sample; based on the information flow score, evaluate the possibility of this type of component as the information source of the output result of the large language model as a second evaluation parameter; for each sample, obtain the masked prediction result output by the large language model after masking the operation result of the target execution operation among the multiple execution operations contained in the sample and the unmasked prediction result output by the large language model before masking; based on the masked prediction result and the unmasked prediction result, evaluate the influence of the target execution operation on the output result of the large language model as a third evaluation parameter; based on the first evaluation parameter, the second evaluation parameter and the third evaluation parameter, evaluate the planning ability of the large language model.

[0006] Optionally, the calculation of the extraction rate of the layer based on the representation vector of each layer includes: inputting the representation vector of the last word of each execution operation among the multiple execution operations contained in each sample at the layer into the decoding head of the layer to obtain the decoding result corresponding to the layer; inputting the representation vector of the last word of each execution operation among the multiple execution operations contained in each sample at the layer into the decoding head of the last layer to obtain the decoding result corresponding to the last layer; for the i-th execution operation contained in each sample, based on the decoding result corresponding to the layer and the decoding result corresponding to the last layer, determining the consistency of the i-th execution operation contained in each sample at the layer, wherein i=1, 2,…, N, N represents the number of the multiple execution operations contained in each sample, and N is a positive integer; based on the consistency of the i-th execution operation contained in each sample at the layer and the number of the multiple samples, determining the extraction rate of the i-th execution operation at the layer.

[0007] Optionally, the evaluation method further includes: for each layer, calculating the average extraction rate of each execution operation at the layer as a fourth evaluation parameter for evaluating the planning capability of the large language model.

[0008] Optionally, the calculation of the detection accuracy of the layer based on the representation vector of each layer includes: inputting the representation vector of the last word of the first execution operation among the multiple execution operations contained in each sample into the probe at the layer to obtain the detection result of the sample at the layer, wherein the detection result includes the prediction result of each execution operation among the multiple execution operations contained in the sample; inputting the representation vector of the last word of the first execution operation among the multiple execution operations contained in each sample into the probe at the last layer of the multiple layers to obtain the detection result of the sample at the last layer; calculating the detection accuracy of the layer based on the detection result of each sample in the multiple samples at the layer and the detection result of the sample at the last layer, wherein the detection accuracy is the ratio of the number of samples in the multiple samples whose detection results in the layer are consistent with the detection results in the last layer to the multiple samples.

[0009] Optionally, for each sample, calculating the information flow score of each type of component among the multiple types of components contained in the sample, includes: for each sample, calculating the average of multiple word segmentation information flow scores of multiple word segmentations contained in each type of component among the multiple types of components contained in the sample, as the information flow score of this type of component.

[0010] Optionally, the evaluating, based on the information flow score, the possibility of the component of this type being the information source of the output result of the large language model includes: determining the highest component with the highest information flow score corresponding to the multiple types of components contained in each sample of the multiple samples; determining the quantitative ratio of each type of the highest component among the multiple highest components corresponding to the multiple samples, wherein the larger the quantitative ratio, the higher the possibility of the corresponding type of component being the information source of the output result of the large language model.

[0011] Optionally, the evaluating the influence of the target execution operation on the output result of the large language model based on the masked prediction result and the unmasked prediction result includes: calculating a first proportion of the number of samples whose masked prediction results are consistent with the true output result label of the large language model in the multiple samples; calculating a second proportion of the number of samples whose unmasked prediction results are consistent with the true output result label of the large language model in the multiple samples; calculating the gap between the first proportion and the second proportion; and based on the gap, evaluating the influence of the target execution operation on the output result of the large language model, wherein the larger the gap, the greater the influence of the target execution operation on the output result of the large language model.

[0012] According to a second aspect of an embodiment of the present disclosure, there is provided an evaluation device for the planning capability of a large language model, comprising: a sample acquisition module, configured to acquire a plurality of samples, wherein each sample comprises a plurality of types of components, and the sample is a text type sample; a first evaluation parameter calculation module, configured to input the plurality of samples into the large language model, and obtain a representation vector of each sample in each of a plurality of layers contained in the large language model; based on the representation vector of each layer, calculating an extraction rate and a detection accuracy rate of the layer as a first evaluation parameter, wherein the extraction rate is used to characterize the possibility that a decoding head of the layer accurately predicts an operation result of a single execution operation among a plurality of execution operations contained in the each sample, and the detection accuracy rate is used to characterize the possibility that a probe of the layer accurately predicts a plurality of operation results corresponding one to one to a plurality of execution operations contained in the each sample; a second evaluation parameter calculation module, The method comprises the following steps: a) calculating the information flow score of each type of component among the multiple types of components contained in the sample; and based on the information flow score, evaluating the possibility of this type of component as the information source of the output result of the large language model as a second evaluation parameter; b) obtaining, for each sample, a shielded prediction result output by the large language model after shielding the operation result of the target execution operation among the multiple execution operations contained in the sample, and b) obtaining the unshielded prediction result output by the large language model before shielding; and based on the shielded prediction result and the unshielded prediction result, evaluating the influence of the target execution operation on the output result of the large language model as a third evaluation parameter; and c) evaluating the planning ability of the large language model based on the first evaluation parameter, the second evaluation parameter and the third evaluation parameter.

[0013] Optionally, the first evaluation parameter calculation module is configured to: input the representation vector of the last word of each execution operation among the multiple execution operations contained in each sample at the layer into the decoding head of the layer to obtain the decoding result corresponding to the layer; input the representation vector of the last word of each execution operation among the multiple execution operations contained in each sample at the last layer among the multiple layers into the decoding head of the last layer to obtain the decoding result corresponding to the last layer; for the i-th execution operation contained in each sample, based on the decoding result corresponding to the layer and the decoding result corresponding to the last layer, determine the consistency of the i-th execution operation contained in each sample at the layer, where i=1, 2,…,N, N represents the number of multiple execution operations contained in each sample, and N is a positive integer; based on the consistency of the i-th execution operation contained in each sample at the layer and the number of the multiple samples, determine the extraction rate of the i-th execution operation at the layer.

[0014] Optionally, the device for evaluating the planning capability of the large language model also includes: a fourth evaluation parameter calculation module, configured to calculate, for each layer, the mean extraction rate of each execution operation at the layer, as a fourth evaluation parameter for evaluating the planning capability of the large language model.

[0015] Optionally, the first evaluation parameter calculation module is configured to: input the representation vector of the last participle of the first execution operation among the multiple execution operations contained in each sample at the layer into the probe to obtain the detection result of the sample at the layer, wherein the detection result includes the prediction result of each execution operation among the multiple execution operations contained in the sample; input the representation vector of the last participle of the first execution operation among the multiple execution operations contained in each sample at the last layer of the multiple layers into the probe to obtain the detection result of the sample at the last layer; calculate the detection accuracy of the layer based on the detection result of each sample in the multiple samples at the layer and the detection result of the sample at the last layer, wherein the detection accuracy is the ratio of the number of samples in the multiple samples whose detection results in the layer are consistent with the detection results in the last layer.

[0016] Optionally, the second evaluation parameter calculation module is configured to: for each sample, calculate the average of multiple word segmentation information flow scores of the multiple word segmentations contained in each type of component among the multiple types of components contained in the sample, as the information flow score of this type of component.

[0017] Optionally, the second evaluation parameter calculation module is configured to: determine the highest component with the highest information flow score among the multiple types of components contained in each sample of the multiple samples; determine the quantitative ratio of each type of the highest component among the multiple highest components corresponding to the multiple samples, wherein the larger the quantitative ratio, the higher the possibility that the corresponding type of component is the information source of the output result of the large language model.

[0018] Optionally, the third evaluation parameter calculation module is configured to: calculate a first proportion of the number of samples whose masked prediction results are consistent with the true output result label of the large language model in the multiple samples; calculate a second proportion of the number of samples whose unmasked prediction results are consistent with the true output result label of the large language model in the multiple samples; calculate the gap between the first proportion and the second proportion; and based on the gap, evaluate the influence of the target execution operation on the output result of the large language model, wherein the larger the gap, the greater the influence of the target execution operation on the output result of the large language model.

[0019] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement a method for evaluating the planning capability of a large language model according to the present disclosure.

[0020] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform a method for evaluating the planning capability of a large language model according to the present disclosure.

[0021] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements a method for evaluating the planning capability of a large language model according to the present disclosure.

[0022] The technical solution provided by the embodiments of the present disclosure brings at least the following beneficial effects: In the present disclosure, by calculating the model's extraction rate, detection accuracy, information flow score, and causal influence of historical steps, the model's forward-looking planning capability is proven in multiple directions, providing theoretical support for the interpretability of the model's short-term forward-looking future decision-making capability in globally observable planning tasks.

[0023] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0025] Figure 1 is a comparative schematic diagram showing the planning process of the block world in the related art and the present disclosure; Figure 2 is a flowchart illustrating a method for evaluating the planning ability of a large language model according to an exemplary embodiment of the present disclosure; Figure 3 is a schematic diagram showing a sample according to an exemplary embodiment of the present disclosure; Figure 4 is a schematic diagram showing means and variances of extraction rates at different layers of a large language model according to an exemplary embodiment of the present disclosure; Figure 5 is a schematic diagram showing a comparison of detection results using a linear probe and a nonlinear probe according to an exemplary embodiment of the present disclosure; Figure 6 is a schematic diagram showing a comparison of the accuracy of detection using a linear probe and a nonlinear probe according to an exemplary embodiment of the present disclosure; Figure 7 is a schematic diagram showing the result of exploring future steps using a linear probe according to an exemplary embodiment of the present disclosure; Figure 8 is a distribution diagram showing information flows of various types of information blocks according to an exemplary embodiment of the present disclosure; Fig. 9 is a schematic diagram showing the results of a single-step intervention analysis according to an exemplary embodiment of the present disclosure; Fig.10 is a block diagram illustrating an evaluation apparatus for the planning capability of a large language model according to an exemplary embodiment of the present disclosure; Fig.11 is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation methods described in the following examples do not represent all implementation methods consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.

[0028] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three types of parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three types of parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.

[0029] As mentioned above, the underlying mechanism behind the planning ability of large language models is still a largely unexplored frontier. Therefore, exploring the underlying planning mechanism of large language models to determine whether they have forward-looking planning capabilities has become a pressing issue.

[0030] In order to solve the above-mentioned problems existing in the relevant technologies, the present invention provides a method, device, electronic device, storage medium and computer program product for evaluating the planning ability of a large language model. By calculating the extraction rate, detection accuracy, information flow score and causal influence of historical steps of the model, it realizes the proof of the model's forward-looking planning ability in multiple directions, that is, it provides theoretical support for the explainability of the model's short-term forward-looking future decision-making ability in globally observable planning tasks.

[0031] Specifically, we can study the classic planning task Block World, which is a fully observable setting task. The states of all entities are known from the initial state and the goal state, so no exploration is required. Given the initial state and goal state of Block World, the model can only pick up or put down a block, so the model must generate a sequence of actions to transform the initial state into the goal state. Figure 1 2 is a schematic diagram showing a comparison of the planning process of the block world in the related art and the present disclosure. Figure 1, the green path is the correct path to transform the initial state of the block world into the target state. However, it is still unclear whether the model greedily considers only the action of t+1 at step t, or proactively considers the actions of t+2 and beyond. The inspiration from psychology is that humans think proactively when making plans. Based on this, the present disclosure proposes a hypothesis of model proactive planning.

[0032] The existence hypothesis of forward-looking planning decisions: In the planning task, the large language model is given a rule, an initial state, a target state, and a task description prompt. At the current step, the model needs to predict the next action, and when planning is successful in a fully observable environment, the probe can detect the short-term future decision steps in the internal representation to a certain extent.

[0033] Furthermore, the present disclosure designs a two-stage paradigm to verify the above assumptions, which can be specifically divided into the "information flow search stage" and the "internal representation detection stage". In addition, the "information flow search stage" mainly analyzes the information flow and component functions in the planning process; the "internal representation detection stage" mainly checks whether the model stores future information in the internal representation.

[0034] Specifically, in the "information flow search phase", the present disclosure studies how planning is performed internally by analyzing the multi-layer perceptron and multi-head attention components at the last token. By calculating the extraction rate, it is found that the multi-head attention output of the middle layer at the last token can directly decode the correct decision to a certain extent. Based on this discovery, the present disclosure further calculates the information source of the multi-head attention to track the source of the decision, and finds that planning mainly depends on the target state and recent historical steps.

[0035] In the "internal representation detection stage", the present disclosure studies what information is encoded in the information flow and whether this information has been considered in advance for future decisions. For the existence of future decisions, the present disclosure uses detection methods to detect future decisions and reveals that when planning is successful, some short-term future decisions are encoded in the middle and upper levels. For the causal relationship of historical steps, the present disclosure explores the impact of different historical steps on the final decision by blocking the information flow from historical steps. In this way, the above two-stage paradigm can be used to verify the existence hypothesis of forward-looking planning decisions.

[0036] Figure 2 is a flowchart illustrating a method for evaluating the planning ability of a large language model according to an exemplary embodiment of the present disclosure.

[0037] Reference Figure 2 In step 201, multiple samples may be obtained, wherein each sample may contain multiple types of components, and the sample may be a text type sample. Figure 3is a schematic diagram showing a sample according to an exemplary embodiment of the present disclosure. Figure 3 , a sample can contain four types of components, namely: "rule", "init state", "goal state" and "plan". "Rule" is used to indicate the rules that need to be followed when moving blocks; "init state" is used to indicate the placement of each block in the initial state; "goalstate" is used to indicate the desired placement state of each block; "plan" is used to indicate the moving step strategy for each block in the process of adjusting the placement of blocks from the initial state to the target state.

[0038] In step 202, multiple samples may be input into a large language model to obtain a representation vector of each sample in each of the multiple layers included in the large language model. Then, based on the representation vector of each layer, the extraction rate and detection accuracy of the layer may be calculated as the first evaluation parameter, wherein the extraction rate may be used to characterize the possibility that the decoding head of the layer accurately predicts the operation result of a single execution operation among the multiple execution operations (steps) included in each sample; the detection accuracy may be used to characterize the possibility that the probe of the layer accurately predicts the multiple operation results corresponding to the multiple execution operations included in each sample.

[0039] For the "extraction rate", the extraction rate can be calculated to analyze the functions of the multi-head attention component and the multi-layer feedforward network component in forward planning. Specifically, at the position of the last token of the model, the role of the multi-head attention component and the multi-layer perceptron component in the answer generation process can be studied, and then the different components of different layers of the model can be directly decoded to the final answer. At this point, the "extraction rate" indicator can be used to quantitatively analyze the functions of different components.

[0040] The calculation method of "extraction rate" can be: first calculate the token corresponding to the maximum value of the probability distribution of the output vocabulary of the last layer at the last token position as the actual extraction result; then, the token corresponding to the maximum value of the internal representation after being decoded by the decoder of the large language model can be solved as the predicted extraction result. If the actual extraction result is consistent with the predicted extraction result, it can be considered that an extraction event has occurred on the internal representation. That is, the extraction rate is to compare the token with the highest probability of output of the last layer of the model with the token decoded by the decoding head of a certain intermediate layer representation. If the two are consistent, it is considered that an "extraction event" has occurred in this layer. In this way, it is found that multi-head self-attention has a higher extraction rate, and the extraction has already occurred in the intermediate layer. This shows that the multi-head self-attention of the intermediate layer can directly decode the planning decision to a certain extent.

[0041] According to an exemplary embodiment of the present disclosure, the representation vector of the last word of each execution operation in the multiple execution operations contained in each sample at this layer can be input into the decoding head of this layer to obtain the decoding result corresponding to this layer. Figure 3 , Figure 3 The sample in contains 4 execution operations, namely step1, step2, step3 and step4. In addition, the last participle of the first execution operation step1 is "up", the last participle of the second execution operation step2 is "of", the last participle of the third execution operation step3 is "up", and the last participle of the fourth execution operation step4 is "of".

[0042] Then, the representation vector of the last word segment of each execution operation in the multiple execution operations included in each sample in the last layer among the multiple layers can be input into the decoding head of the last layer to obtain the decoding result corresponding to the last layer.

[0043] Next, for the i-th execution operation contained in each sample, the consistency of the i-th execution operation contained in each sample at the layer can be determined based on the decoding result corresponding to the layer and the decoding result corresponding to the last layer, where i=1, 2, ..., N, N can represent the number of multiple execution operations contained in each sample, and N is a positive integer. For example, return to reference Figure 3 ,against Figure 3 The sample shown in contains the first execution operation (step1), assuming that the decoding result corresponding to step1 in the last layer is "white". If the decoding result corresponding to step1 in this layer is "white", which is the same as the decoding result "white" corresponding to step1 in the last layer, it means that the consistency result of step1 in this layer is "consistent"; or, if the decoding result corresponding to step1 in this layer is "blue", which is different from the decoding result "white" corresponding to step1 in the last layer, it means that the consistency result of step1 in this layer is "inconsistent".

[0044] In addition, for example, assuming that there are 100 samples in total, and assuming i=1, the 100 samples will include 100 step1s in total. Since each step1 corresponds to a consistency result, the 100 step1s can correspond to 100 consistency results. In addition, some of these 100 consistency results may be "consistent", while some may be "inconsistent".

[0045] Then, the extraction rate of the i-th execution operation at the layer can be determined based on the consistency of the i-th execution operation at the layer contained in each sample and the number of multiple samples. Exemplarily, as mentioned above, assuming i=1, the ratio of the number of "consistent" results in the aforementioned 100 consistency results to the 100 consistency results can be calculated as the extraction rate of the first execution operation at the layer.

[0046] Thus, in the present disclosure, it is possible to quantitatively analyze whether the multi-head attention component and the multi-layer feedforward network component can directly decode the planning decision through the decoding head at different layers when generating the decision, and then determine the function of the multi-head attention component and the multi-layer feedforward network component in forward planning based on the analysis results. It turns out that the multi-head self-attention of the middle layer can directly decode the planning decision to a certain extent.

[0047] According to an exemplary embodiment of the present disclosure, for each layer, the average of the extraction rates of each execution operation at the layer may be calculated as a fourth evaluation parameter for evaluating the planning capability of the large language model.

[0048] Exemplary, return reference Figure 3 , Figure 3 The samples in contain a total of 4 execution operations, namely step1, step2, step3 and step4. And, assuming that there are a total of 100 samples, the 100 step1s contained in the 100 samples can correspond to one extraction rate, the 100 step2s contained in the 100 samples can correspond to one extraction rate, the 100 step3s contained in the 100 samples can correspond to one extraction rate, and the 100 step4s contained in the 100 samples can correspond to one extraction rate. In this way, for each layer, the mean of the 4 extraction rates corresponding to the above 4 execution operations in the layer can be calculated as the fourth evaluation parameter for evaluating the planning ability of the large language model. In addition, the variance of the above extraction rate can also be calculated to evaluate the planning ability of the large language model.

[0049] Figure 4 is a schematic diagram showing the mean and variance of the extraction rate at different layers of a large language model according to an exemplary embodiment of the present disclosure. Figure 4 , the answer extraction rate of Multi-Head Self-Attention (MHSA) is higher than that of Multilayer Perceptron (MLP), indicating that the attention mechanism is the main responsible module for answer extraction. And, referring to Figure 4, the outputs of the middle and upper layers of the model gradually form stable answers. In these layers, the extraction rate (extraction rate) of multi-head self-attention is significantly higher than that of multi-layer perceptron, indicating that multi-head self-attention plays a major role in the decision-making stage. In addition, compared with multi-layer perceptron, the variance of the extraction rate of multi-head self-attention at different steps is smaller, indicating that the multi-head self-attention layer shows higher consistency at different steps.

[0050] Regarding the “detection accuracy”, linear probes and nonlinear probes can be used to analyze the existence of the current world state and future planning decisions in the internal representation of the model. That is, the internal representation vectors of the model to be analyzed at different stages can be obtained, including the initial state, target state, and intermediate decision steps. In addition, for each layer of the internal representation of the model, linear probes and nonlinear probes can be constructed and trained to enable them to predict the characteristics of the current world state and future decision characteristics. In this way, by testing the prediction results of the trained probes, it is found that the state and decision information gradually become richer as the number of internal layers of the model increases, and there are representations of the current world state and future decisions.

[0051] Specifically, it is possible to check whether the internal representation encodes two types of information: the current block state and future decisions. For example, the internal representations of the initial state, the target state, and each step can be detected, and linear detectors and nonlinear detectors can be trained for each block and each layer. The input of the linear detector can be the model hidden layer representation, and the output can be a matrix representing the probability; the structure of the nonlinear detector can be: the normalized exponential function acts on the weight matrix × the corrected linear unit function acts on the weight matrix × the input result. For the current block state, the detector output can be the probability matrix of the color of the upper and lower blocks of each color block; for future decisions, the detector can output the predicted color of each step. In addition, the samples can be divided into training sets and test sets, and the weighted F1 accuracy of the current block state and the accuracy of future decisions can be calculated for evaluation.

[0052] Figure 5 is a schematic diagram showing a comparison of detection results using a linear probe and a nonlinear probe according to an exemplary embodiment of the present disclosure; Figure 6 is a schematic diagram showing a comparison of the accuracy of detection using a linear probe and a nonlinear probe according to an exemplary embodiment of the present disclosure; Figure 7 is a schematic diagram showing the result of exploring future steps using a linear probe according to an exemplary embodiment of the present disclosure.

[0053] Reference Figure 5 and Figure 6 ,In terms of the world state, the detection accuracy gradually improves as the number of layers increases, which indicates that the early layers of the model are enriching the representation of the current state. Figure 5The black line in the figure (step 6) has a lower detection accuracy than the light blue line (step 2), indicating that the model has difficulty maintaining the current placement representation of the block as the planning step progresses. In addition, by comparing the linear detection and nonlinear detection in the figure, it can be found that both have the same trend, indicating that the model internally stores the current state in a linear manner. Moreover, the figure also shows that the "future decision" of the action has a similar trend, which reveals that the forward-looking decision is enriched in the early layers.

[0054] In future decision-making, refer to Figure 7 , the following can be observed: For row 6, column 1, the detector can predict the future sixth step with an accuracy of 0.51 in the first step, which indicates that the model stores information about future decisions in advance, supporting the hypothesis of forward planning. Moreover, for each row, the values ​​increase from left to right. For example, the accuracy of the fifth column of the sixth row is higher than that of the first column, which means that the model is more certain about the output of the sixth step in the fifth step compared to the first step, indicating that the model has difficulty in long-distance planning. In addition, the accuracy of the first column indicates that the prediction accuracy of the next five steps after the initial step shows a downward trend, indicating that the model stores future decision information in advance, supporting the hypothesis that forward planning decisions exist.

[0055] According to an exemplary embodiment of the present disclosure, the representation vector of the last word segment of the first execution operation among the multiple execution operations contained in each sample at this layer can be input into the probe to obtain the detection result of the sample at this layer, wherein the detection result can include the prediction result of each execution operation among the multiple execution operations contained in the sample.

[0056] Exemplary, return reference Figure 3 , Figure 3 The sample in contains a total of 4 execution operations, namely step1, step2, step3 and step4. In addition, the last word of the first execution operation step1 is "up". The representation vector of the last word "up" in this layer can be input into the probe to obtain the detection result of the sample in this layer, wherein the detection result can include the prediction result of each of the 4 execution operations included in the sample, for example, the detection result can be (white-table-gray-black).

[0057] Then, the representation vector of the last word segment of the first execution operation among the multiple execution operations contained in each sample at the last layer among the multiple layers may be input into the probe to obtain the detection result of the sample at the last layer.

[0058] Next, the detection accuracy of the layer can be calculated based on the detection result of each sample in the multiple samples at the layer and the detection result of the sample in the last layer, wherein the detection accuracy can be the ratio of the number of samples in the multiple samples whose detection results in the layer are consistent with the detection results in the last layer.

[0059] For example, return reference Figure 3 ,against Figure 3 For the sample shown in , suppose that the representation vector of the last word "up" of the first execution operation among the four execution operations contained in the sample is input into the probe at the last layer, and the detection result of the sample at the last layer is: (white-table-gray-black). If the detection result of the sample at this layer is "(white-table-gray-black)", which is the same as the detection result of the sample at the last layer "(white-table-gray-black)", it means that the consistency result of the sample at this layer is "consistent"; or, if the detection result of the sample at this layer is "(black-table-white-red)", which is different from the detection result of the sample at the last layer "(white-table-gray-black)", it means that the consistency result of the sample at this layer is "inconsistent".

[0060] In addition, for example, assuming that there are 100 samples in total, these 100 samples will correspond to 100 consistency results in total. Moreover, some of these 100 consistency results may be "consistent", while some of them may be "inconsistent". At this time, the ratio of the number of "consistent" results in the aforementioned 100 consistency results to the 100 consistency results can be calculated as the detection accuracy of this layer.

[0061] In this way, by using linear probes and nonlinear probes to analyze the existence of the current world state and future planning decisions in the internal representation of the model, it can be proved that the state and decision information gradually become richer as the number of internal layers of the model increases. That is, by using linear probes and nonlinear probes to detect the information representation content of the current state and future decisions in the internal representation at a certain moment, the existence of advance planning can be proved.

[0062] In step 203, for each sample, the information flow score of each type of component among the multiple types of components contained in the sample can be calculated. Then, the possibility of the component of this type as the information source of the output result of the large language model can be evaluated based on the information flow score as a second evaluation parameter.

[0063] Specifically, the input sample can be decomposed into several blocks, and then the information source of the multi-head attention extraction planning decision can be traced by analyzing the information flow at the block granularity to determine which block the information source of the multi-head attention extraction answer mainly depends on. Furthermore, the token-level information flow can be calculated first, and then the block-level information flow can be calculated based on the token-level information flow and the calculated block-level information flow can be used to calculate the contribution of different semantic blocks to the decision.

[0064] According to an exemplary embodiment of the present disclosure, for each sample, the average of multiple word segmentation information flow scores corresponding to multiple word segmentations contained in each type of component of the multiple types of components contained in the sample can be calculated as the information flow score of the component of this type. That is, the information flow of the token granularity can be calculated first, and then the average value of the information flow of different tokens in the same block can be taken to represent the information flow of the block granularity.

[0065] Specifically, first, the input data can be divided into different information blocks, including: special tokens, initial states, target states, different historical steps, action prompts, and so on. Then, for each layer, the gradient score of the loss function to the attention matrix can be calculated to obtain the information flow score of each token. For example, the gradient and attention score can be element-wise multiplied and summed over all attention heads and then the absolute value can be calculated to obtain the information flow score at the token granularity. Next, the information flow between information blocks can be calculated based on the token information flow. Exemplarily, the information flow scores between information blocks can be obtained by averaging all token information flows within the information block.

[0066] For example, the information flow score at the token granularity can be calculated by the following formula:

[0067] in, It is The attention score of the layer, i.e. The information flow fraction of the layer, Indicates Head, is the loss function, Indicates that from token Flow to token The information flow score.

[0068] The information flow score at the block granularity can be calculated by the following formula:

[0069] in, For Block The interval To another block The interval The information flow score.

[0070] It should be noted that in this disclosure, it is possible to calculate from blocks To block And, due to the limitation of the causal attention mechanism, only the information flow in the lower triangular matrix can be calculated.

[0071] Exemplary, return reference Figure 3 , Figure 3 The sample in shows four types of components, namely: "rule", "init state", "goal state" and "plan", and each type of component contains multiple word segments. For example, the component of type "rule" contains word segments such as "you", "can", "pick-up", etc.; the component of type "initstate" contains word segments such as "empty", "black", "white", etc.; the component of type "goal state" contains word segments such as "gray", "on", "black", etc.; the component of type "plan" contains word segments such as "step1", "pick-up", "answer", etc. In this way, by analyzing the information flow of different input blocks, it can be found that multi-head attention mainly extracts information from the target state and recent historical steps.

[0072] According to an exemplary embodiment of the present disclosure, the highest component with the highest corresponding information flow score among the multiple types of components contained in each sample of the multiple samples can be determined. Then, the quantitative ratio of each type of the highest component among the multiple highest components corresponding to the multiple samples can be determined, wherein the larger the quantitative ratio, the higher the possibility that the corresponding type of component is the information source of the output result of the large language model.

[0073] For example, assuming that there are 100 samples in total, and each sample contains 4 types of components, then each sample can correspond to a highest component with the highest information flow score. In this way, the 100 samples can correspond one by one to the 100 highest components with the highest information flow scores. At this time, the quantitative ratio of each type of highest component in the 100 highest components corresponding to the 100 samples can be determined, wherein the larger the quantitative ratio, the higher the possibility that the corresponding type of component is the information source of the output result of the large language model.

[0074] Figure 8 2 is a schematic diagram showing the distribution of information flows of various types of information blocks according to an exemplary embodiment of the present disclosure. Figure 8 , the darker the color, the more attention it pays. It can be seen that the multi-head attention mechanism mainly extracts information from the target state, indicating that it mainly depends on the target state; and the multi-head attention mechanism mainly depends on recent historical steps rather than earlier historical steps.

[0075] In step 204, for each sample, a masked prediction result output by the large language model after masking the operation result of the target execution operation among the multiple execution operations included in the sample and an unmasked prediction result output by the large language model before masking can be obtained. Then, based on the masked prediction result and the unmasked prediction result, the influence of the target execution operation on the output result of the large language model can be evaluated as a third evaluation parameter.

[0076] Specifically, in the multi-head attention mechanism, the causal influence of historical steps on future decisions can be detected by setting the key of the key-value pair to zero using a causal analysis method. For example, first, the key information in all historical decision steps can be masked to block the influence of past decision information on the current decision; then, only the key information of a certain historical decision step can be kept visible and the information of other historical steps can be masked at the same time to obtain the decision probability when the step is visible. Finally, the influence of a single historical decision step on the current decision step can be calculated by comparing the difference between the decision probabilities when all historical steps are masked and when only one historical step is kept visible. The larger the impact value, the greater the influence of the historical decision step on the current decision step, which can prove the causal influence of the historical step on future decisions.

[0077] That is, in the present disclosure, the impact of each historical decision step on a specific future decision step can be quantitatively evaluated by selectively shielding historical decision information in the key value of the attention mechanism and comparing the changes in the predicted probability of the model for future decisions before and after shielding. Specifically, for a sequence decision task with a length of several steps, the above quantitative evaluation process can be divided into three steps: first, the key values ​​corresponding to all decision information tags in the historical sequence can be set to zero in each layer of the attention mechanism to block the past decision information from affecting the current step. At this time, the decision prediction probability of the last step can be obtained; then, only the decision information tag of a certain historical step can be restored to be visible to keep the tags of other historical steps in a shielded state, and the prediction can be re-performed to obtain a new decision prediction probability for the last step; finally, the difference between the two predicted probabilities can be compared as a quantitative indicator of the impact of the historical step on the last step, and the larger the value of the indicator, the greater the impact of the historical decision step on the last decision step.

[0078] According to an exemplary embodiment of the present disclosure, a first proportion of the number of samples whose shielded prediction results are consistent with the true output result label of the large language model in multiple samples can be calculated. Exemplarily, assuming that there are 100 samples in total, the first proportion of the number of samples whose shielded prediction results output by the large language model are consistent with the true output result label of the large language model after shielding the operation result of the first execution operation among the multiple execution operations contained in each sample can be calculated.

[0079] Then, the second proportion of the number of samples whose unshielded prediction results are consistent with the true output result label of the large language model in the multiple samples can also be calculated. Exemplarily, the second proportion of the number of samples whose unshielded prediction results output by the large language model are consistent with the true output result label of the large language model in the 100 samples can also be calculated when the operation result of the first execution operation among the multiple execution operations contained in each sample is not shielded.

[0080] Next, the gap between the first ratio and the second ratio can be calculated. Then, based on the gap, the influence of the target execution operation on the output result of the large language model can be evaluated. For example, based on the gap, the influence of the first execution operation on the output result of the large language model can be evaluated, wherein the larger the gap, the greater the influence of the target execution operation on the output result of the large language model.

[0081] Fig. 9 is a schematic diagram showing the results of a single-step intervention analysis according to an exemplary embodiment of the present disclosure. Fig. 9 , we can see that the model is not greedy, that is, it is not limited to preparing for the next step, which proves the conclusion of forward planning from the causal relationship. In addition, by observing the values ​​of each column, it is found that the most important steps for prediction are often the later steps, which shows that the forward planning ability of the large language model is still relatively preliminary.

[0082] In step 205, the planning capability of the large language model may be evaluated based on the first evaluation parameter, the second evaluation parameter, and the third evaluation parameter.

[0083] Exemplarily, the larger the first evaluation parameter is, that is, the larger the extraction rate and detection accuracy are, the stronger the forward-looking planning ability of the large language model is; the larger the second evaluation parameter is, that is, the larger the information flow score is, the higher the possibility that the corresponding type of component is the information source of the output result of the large language model is; the larger the third evaluation parameter is, the greater the influence of the corresponding execution operation on the output result of the large language model is.

[0084] In the present disclosure, in order to solve the problem of unknown planning mechanism, a forward planning mechanism interpretability method for large language models is designed. Specifically, in terms of information flow search, the information flow of extraction rate and block granularity can be used to analyze the information flow within the large language model in the planning task; in terms of internal representation detection, linear probes and nonlinear probes can be used to prove the existence of future planning decisions; in addition, the causal influence of historical steps on future planning decisions can be proved by setting the keys in multi-head attention to zero and causally intervening. It turns out that in globally observable planning tasks, there are short-term future decisions within the large language model, which provides a theoretical basis for the interpretability of the planning mechanism of the large language model, that is, it proves that the forward planning decision-making ability of the large language model really exists.

[0085] Fig.10 is a block diagram illustrating an evaluation apparatus 1000 for a planning capability of a large language model according to an exemplary embodiment of the present disclosure.

[0086] Reference Fig.10 The planning ability evaluation device 1000 of the large language model may include a sample acquisition module 1001, a first evaluation parameter calculation module 1002, a second evaluation parameter calculation module 1003, a third evaluation parameter calculation module 1004 and a planning ability evaluation module 1005.

[0087] The sample acquisition module 1001 may acquire multiple samples, wherein each sample may contain multiple types of components, and the sample may be a text type sample.

[0088] The first evaluation parameter calculation module 1002 can input multiple samples into the large language model to obtain the representation vector of each sample in each of the multiple layers contained in the large language model. Then, based on the representation vector of each layer, the extraction rate and detection accuracy of the layer can be calculated as the first evaluation parameter, wherein the extraction rate can be used to characterize the possibility of the decoding head of the layer to accurately predict the operation result of a single execution operation among the multiple execution operations (steps) contained in each sample; the detection accuracy can be used to characterize the possibility of the probe of the layer to accurately predict the multiple operation results corresponding to the multiple execution operations contained in each sample.

[0089] According to an exemplary embodiment of the present disclosure, the first evaluation parameter calculation module 1002 may input the representation vector of the last word of each execution operation in the multiple execution operations contained in each sample in the layer into the decoding head of the layer to obtain the decoding result corresponding to the layer. Then, the first evaluation parameter calculation module 1002 may also input the representation vector of the last word of each execution operation in the multiple execution operations contained in each sample in the last layer in the multiple layers into the decoding head of the last layer to obtain the decoding result corresponding to the last layer.

[0090] Next, for the i-th execution operation contained in each sample, the first evaluation parameter calculation module 1002 can determine the consistency of the i-th execution operation contained in each sample at the layer based on the decoding result corresponding to the layer and the decoding result corresponding to the last layer, where i=1, 2, ..., N, N can represent the number of multiple execution operations contained in each sample, and N is a positive integer. Then, the first evaluation parameter calculation module 1002 can determine the extraction rate of the i-th execution operation at the layer based on the consistency of the i-th execution operation contained in each sample at the layer and the number of multiple samples.

[0091] Thus, in the present disclosure, it is possible to quantitatively analyze whether the multi-head attention component and the multi-layer feedforward network component can directly decode the planning decision through the decoding head at different layers when generating the decision, and then determine the function of the multi-head attention component and the multi-layer feedforward network component in forward planning based on the analysis results. It turns out that the multi-head self-attention of the middle layer can directly decode the planning decision to a certain extent.

[0092] According to an exemplary embodiment of the present disclosure, the above-mentioned large language model planning ability evaluation device 1000 may further include a fourth evaluation parameter calculation module. For each layer, the fourth evaluation parameter calculation module may calculate the mean of the extraction rate of each execution operation at the layer as a fourth evaluation parameter for evaluating the planning ability of the large language model.

[0093] According to an exemplary embodiment of the present disclosure, the first evaluation parameter calculation module 1002 may input the representation vector of the last word segment of the first execution operation among the multiple execution operations contained in each sample at this layer into the probe to obtain the detection result of the sample at this layer, wherein the detection result may include the prediction result of each execution operation among the multiple execution operations contained in the sample.

[0094] Then, the first evaluation parameter calculation module 1002 may input the representation vector of the last word segment of the first execution operation among the multiple execution operations included in each sample at the last layer among the multiple layers into the probe to obtain the detection result of the sample at the last layer.

[0095] Next, the first evaluation parameter calculation module 1002 can calculate the detection accuracy of the layer based on the detection results of each sample in the multiple samples at the layer and the detection results of the sample in the last layer, wherein the detection accuracy can be the ratio of the number of samples in the multiple samples whose detection results in the layer are consistent with the detection results in the last layer.

[0096] In this way, by using linear probes and nonlinear probes to analyze the existence of the current world state and future planning decisions in the internal representation of the model, it can be proved that the state and decision information gradually become richer as the number of internal layers of the model increases. That is, by using linear probes and nonlinear probes to detect the information representation content of the current state and future decisions in the internal representation at a certain moment, the existence of advance planning can be proved.

[0097] The second evaluation parameter calculation module 1003 can calculate the information flow score of each type of component among the multiple types of components contained in the sample for each sample. Then, the second evaluation parameter calculation module 1003 can evaluate the possibility of the component of this type as the information source of the output result of the large language model based on the information flow score as the second evaluation parameter.

[0098] Specifically, the input sample can be decomposed into several blocks, and then the information source of the multi-head attention extraction planning decision can be traced by analyzing the information flow at the block granularity to determine which block the information source of the multi-head attention extraction answer mainly depends on. Furthermore, the token-level information flow can be calculated first, and then the block-level information flow can be calculated based on the token-level information flow and the calculated block-level information flow can be used to calculate the contribution of different semantic blocks to the decision.

[0099] According to an exemplary embodiment of the present disclosure, the second evaluation parameter calculation module 1003 can calculate, for each sample, the average of multiple word segmentation information flow scores of multiple word segmentations contained in each type of component of the multiple types of components contained in the sample as the information flow score of the component of this type. That is, the information flow of the token granularity can be calculated first, and then the average value of the information flow of different tokens in the same block can be taken to represent the information flow of the block granularity.

[0100] Specifically, first, the input data can be divided into different information blocks, including: special tokens, initial states, target states, different historical steps, action prompts, and so on. Then, for each layer, the gradient score of the loss function to the attention matrix can be calculated to obtain the information flow score of each token. For example, the gradient and attention score can be element-wise multiplied and summed over all attention heads and then the absolute value can be calculated to obtain the information flow score at the token granularity. Next, the information flow between information blocks can be calculated based on the token information flow. Exemplarily, the information flow scores between information blocks can be obtained by averaging all token information flows within the information block.

[0101] According to an exemplary embodiment of the present disclosure, the second evaluation parameter calculation module 1003 can determine the highest component with the highest corresponding information flow score among the multiple types of components contained in each sample of the multiple samples. Then, the second evaluation parameter calculation module 1003 can determine the quantitative ratio of each type of the highest components among the multiple highest components corresponding to the multiple samples, wherein the larger the quantitative ratio, the higher the possibility that the corresponding type of component is the information source of the output result of the large language model.

[0102] For example, assuming that there are 100 samples in total, and each sample contains 4 types of components, then each sample can correspond to a highest component with the highest information flow score. In this way, the 100 samples can correspond one by one to the 100 highest components with the highest information flow scores. At this time, the quantitative ratio of each type of highest component in the 100 highest components corresponding to the 100 samples can be determined, wherein the larger the quantitative ratio, the higher the possibility that the corresponding type of component is the information source of the output result of the large language model.

[0103] The third evaluation parameter calculation module 1004 can obtain, for each sample, a shielded prediction result output by the large language model after shielding the operation result of the target execution operation among the multiple execution operations included in the sample, and an unshielded prediction result output by the large language model before shielding. Then, the third evaluation parameter calculation module 1004 can evaluate the influence of the target execution operation on the output result of the large language model based on the shielded prediction result and the unshielded prediction result as the third evaluation parameter.

[0104] Specifically, in the multi-head attention mechanism, the causal influence of historical steps on future decisions can be detected by setting the key of the key-value pair to zero using a causal analysis method. For example, first, the key information in all historical decision steps can be masked to block the influence of past decision information on the current decision; then, only the key information of a certain historical decision step can be kept visible and the information of other historical steps can be masked at the same time to obtain the decision probability when the step is visible. Finally, the influence of a single historical decision step on the current decision step can be calculated by comparing the difference between the decision probabilities when all historical steps are masked and when only one historical step is kept visible. The larger the impact value, the greater the influence of the historical decision step on the current decision step, which can prove the causal influence of the historical step on future decisions.

[0105] According to an exemplary embodiment of the present disclosure, the third evaluation parameter calculation module 1004 can calculate the first proportion of the number of samples whose shielded prediction results are consistent with the true output result label of the large language model in multiple samples. Exemplarily, assuming that there are 100 samples in total, the first proportion of the number of samples whose shielded prediction results output by the large language model are consistent with the true output result label of the large language model after shielding the operation result of the first execution operation among the multiple execution operations contained in each sample can be calculated.

[0106] Then, the third evaluation parameter calculation module 1004 may also calculate the second proportion of the number of samples whose unshielded prediction results are consistent with the true output result label of the large language model in the multiple samples. Exemplarily, the second proportion of the number of samples whose unshielded prediction results output by the large language model are consistent with the true output result label of the large language model in the 100 samples may also be calculated when the operation result of the first execution operation among the multiple execution operations contained in each sample is not shielded.

[0107] Next, the third evaluation parameter calculation module 1004 may calculate the gap between the first ratio and the second ratio. Then, the third evaluation parameter calculation module 1004 may evaluate the influence of the target execution operation on the output result of the large language model based on the gap. For example, the influence of the first execution operation on the output result of the large language model may be evaluated based on the gap, wherein the larger the gap, the greater the influence of the target execution operation on the output result of the large language model.

[0108] The planning ability evaluation module 1005 may evaluate the planning ability of the large language model based on the first evaluation parameter, the second evaluation parameter, and the third evaluation parameter.

[0109] Exemplarily, the larger the first evaluation parameter is, that is, the larger the extraction rate and detection accuracy are, the stronger the forward-looking planning ability of the large language model is; the larger the second evaluation parameter is, that is, the larger the information flow score is, the higher the possibility that the corresponding type of component is the information source of the output result of the large language model is; the larger the third evaluation parameter is, the greater the influence of the corresponding execution operation on the output result of the large language model is.

[0110] Fig.11 is a block diagram illustrating an electronic device 1100 according to an exemplary embodiment of the present disclosure.

[0111] Reference Fig.11The electronic device 1100 includes at least one memory 1101 and at least one processor 1102, wherein the at least one memory 1101 stores instructions, and when the instructions are executed by the at least one processor 1102, a method for evaluating the planning capability of a large language model according to an exemplary embodiment of the present disclosure is executed.

[0112] As an example, the electronic device 1100 may be a PC, a tablet device, a personal digital assistant, a smart phone, or other device capable of executing the above instructions. Here, the electronic device 1100 is not necessarily a single electronic device, but may also be any device or circuit collection capable of executing the above instructions (or instruction sets) individually or in combination. The electronic device 1100 may also be part of an integrated control system or system manager, or may be configured as a portable electronic device interconnected with a local or remote (e.g., via wireless transmission) interface.

[0113] In the electronic device 1100, the processor 1102 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller or a microprocessor. As an example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0114] The processor 1102 may execute instructions or codes stored in the memory 1101, wherein the memory 1101 may also store data. Instructions and data may also be sent and received over a network via a network interface device, wherein the network interface device may employ any known transmission protocol.

[0115] The memory 1101 may be integrated with the processor 1102, for example, by placing RAM or flash memory within an integrated circuit microprocessor or the like. In addition, the memory 1101 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 1101 and the processor 1102 may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor 1102 can read files stored in the memory.

[0116] In addition, the electronic device 1100 may further include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the electronic device 1100 may be connected to each other via a bus and / or a network.

[0117] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided, and when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is able to perform the above-mentioned method for evaluating the planning ability of the large language model. Examples of computer-readable storage media here include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disk storage, hard disk drive (HDD), solid state drive (SSD), card storage (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device, any other device is configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the above-mentioned computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system, so that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers.

[0118] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program, which implements the method for evaluating the planning ability of a large language model according to the present disclosure when the computer program is executed by a processor.

[0119] According to the method, device, electronic device, storage medium and computer program product for evaluating the planning ability of a large language model disclosed in the present invention, by calculating the extraction rate, detection accuracy, information flow score and causal influence of historical steps of the model, it is possible to prove that the model has forward-looking planning ability in multiple directions, that is, it provides theoretical support for the interpretability of the model's short-term forward-looking future decision-making ability in globally observable planning tasks.

[0120] According to an exemplary embodiment of the present disclosure, it is possible to quantitatively analyze whether the multi-head attention component and the multi-layer feedforward network component can directly decode the planning decision through the decoding head at different layers when generating the decision, and then determine the function of the multi-head attention component and the multi-layer feedforward network component in forward planning based on the analysis results. It turns out that the multi-head self-attention of the middle layer can directly decode the planning decision to a certain extent.

[0121] According to an exemplary embodiment of the present disclosure, by using linear probes and nonlinear probes to analyze the existence of the current world state and future planning decisions in the internal representation of the model, it can be proved that the state and decision information are gradually enriched as the number of internal layers of the model increases. That is, by detecting the information representation content of the current state and future decisions in the internal representation at a certain moment through linear probes and nonlinear probes, the existence of advance planning can be proved.

[0122] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0123] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A method for evaluating the planning ability of a large language model, characterized in that: include: Acquire multiple samples, wherein each sample contains multiple types of components, and the samples are text type samples; Input the multiple samples into the large language model to obtain a representation vector of each sample in each of the multiple layers included in the large language model; based on the representation vector of each layer, calculate the extraction rate and detection accuracy of the layer as the first evaluation parameter, wherein the extraction rate is used to characterize the possibility that the decoding head of the layer accurately predicts the operation result of a single execution operation among the multiple execution operations included in each sample, and the detection accuracy is used to characterize the possibility that the probe of the layer accurately predicts the multiple operation results corresponding to the multiple execution operations included in each sample; For each sample, calculating the information flow score of each type of component among the multiple types of components contained in the sample; based on the information flow score, evaluating the possibility of the component of this type as the information source of the output result of the large language model as a second evaluation parameter; For each of the samples, obtaining a masked prediction result output by the large language model after masking an operation result of a target execution operation among multiple execution operations included in the sample, and an unmasked prediction result output by the large language model before masking; based on the masked prediction result and the unmasked prediction result, evaluating the influence of the target execution operation on the output result of the large language model as a third evaluation parameter; The planning capability of the large language model is evaluated based on the first evaluation parameter, the second evaluation parameter, and the third evaluation parameter.

2. The evaluation method according to claim 1, characterized in that: The step of calculating the extraction rate of each layer based on the representation vector of the layer includes: Inputting the representation vector of the last word of each execution operation in the multiple execution operations included in each sample at the layer into the decoding head of the layer to obtain the decoding result corresponding to the layer; Inputting the representation vector of the last word of each execution operation in the multiple execution operations included in each sample into the decoding head of the last layer in the multiple layers to obtain a decoding result corresponding to the last layer; For the i-th execution operation included in each sample, based on the decoding result corresponding to the layer and the decoding result corresponding to the last layer, determine the consistency of the i-th execution operation included in each sample at the layer, where i=1, 2, ..., N, N represents the number of multiple execution operations included in each sample, and N is a positive integer; Based on the consistency of the i-th execution operation included in each sample at the layer and the number of the multiple samples, an extraction rate of the i-th execution operation at the layer is determined.

3. The evaluation method according to claim 2, characterized in that: The evaluation method also includes: For each layer, the average extraction rate of each execution operation at the layer is calculated as a fourth evaluation parameter for evaluating the planning ability of the large language model.

4. The evaluation method according to claim 1, characterized in that: The calculating the detection accuracy of each layer based on the representation vector of each layer includes: Inputting the representation vector of the last word of the first execution operation among the multiple execution operations included in each sample at this layer into the probe to obtain the detection result of the sample at this layer, wherein the detection result includes the prediction result of each execution operation among the multiple execution operations included in the sample; Inputting the representation vector of the last word of the first execution operation among the multiple execution operations included in each sample into the probe of the last layer among the multiple layers to obtain the detection result of the sample at the last layer; Based on the detection result of each sample in the multiple samples at the layer and the detection result of the sample in the last layer, the detection accuracy of the layer is calculated, wherein the detection accuracy is the ratio of the number of samples in the multiple samples whose detection results in the layer are consistent with the detection results in the last layer to the number of samples in the multiple samples.

5. The evaluation method according to claim 1, characterized in that: The step of calculating, for each sample, the information flow score of each type of component among the multiple types of components contained in the sample comprises: For each sample, the average of the multiple word segmentation information flow scores of the multiple word segmentations contained in each type of component in the multiple types of components contained in the sample is calculated as the information flow score of the component of this type.

6. The evaluation method according to claim 1, characterized in that: The evaluating, based on the information flow score, the possibility of the component of this type being the information source of the output result of the large language model comprises: Determine a highest component having a highest information flow score among the multiple types of components included in each sample of the multiple samples; Determine the quantitative ratio of each type of the highest components among the multiple highest components corresponding to the multiple samples, wherein the larger the quantitative ratio is, the higher the possibility that the corresponding type of component serves as the information source of the output result of the large language model.

7. The evaluation method according to claim 1, characterized in that: The evaluating, based on the masked prediction result and the unmasked prediction result, the influence of the target execution operation on the output result of the large language model includes: Calculate a first ratio of the number of samples whose masked prediction results are consistent with the true output result label of the large language model to the plurality of samples; Calculate a second ratio of the number of samples whose unshielded prediction results are consistent with the true output result label of the large language model to the plurality of samples; calculating a difference between the first ratio and the second ratio; Based on the gap, the influence of the target execution operation on the output result of the large language model is evaluated, wherein the larger the gap is, the greater the influence of the target execution operation on the output result of the large language model is.

8. A device for evaluating the planning ability of a large language model, characterized in that: include: A sample acquisition module is configured to acquire a plurality of samples, wherein each sample includes a plurality of types of components, and the samples are text type samples; A first evaluation parameter calculation module is configured to input the multiple samples into the large language model, obtain a representation vector of each sample in each of the multiple layers included in the large language model; based on the representation vector of each layer, calculate the extraction rate and detection accuracy of the layer as the first evaluation parameter, wherein the extraction rate is used to characterize the possibility that the decoding head of the layer accurately predicts the operation result of a single execution operation among the multiple execution operations included in each sample, and the detection accuracy is used to characterize the possibility that the probe of the layer accurately predicts the multiple operation results corresponding to the multiple execution operations included in each sample; A second evaluation parameter calculation module is configured to calculate, for each sample, an information flow score of each type of component among the multiple types of components contained in the sample; and evaluate, based on the information flow score, the possibility of the component of this type as an information source of the output result of the large language model as a second evaluation parameter; A third evaluation parameter calculation module is configured to obtain, for each sample, a masked prediction result output by the large language model after masking an operation result of a target execution operation among multiple execution operations included in the sample, and an unmasked prediction result output by the large language model before masking; based on the masked prediction result and the unmasked prediction result, evaluate the influence of the target execution operation on the output result of the large language model as a third evaluation parameter; A planning ability evaluation module is configured to evaluate the planning ability of the large language model based on the first evaluation parameter, the second evaluation parameter and the third evaluation parameter.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method for evaluating the planning capability of a large language model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method for evaluating the planning ability of a large language model as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for evaluating the planning ability of a large language model according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Method and system for interpretability test and evaluation of large language model

    CN118861594A