Large language model interpretability analysis method and device, equipment and storage medium

By generating mapping matrix and filtering specific word tuples, the problem of insufficient interpretability of large language models in complex tasks is solved, and more accurate and simplified interpretability analysis is achieved, which improves the credibility and transparency of the model.

CN120218047AInactive Publication Date: 2025-06-27PENG CHENG LAB
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510157320.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Due to the ‘black box’ characteristics in complex tasks, the reasoning process, internal knowledge structure and decision-making basis are difficult to analyze, resulting in insufficient interpretability.

Method used

A large language model interpretability analysis method is proposed. By obtaining corpus data, generating mapping matrix, calculating neighborhood average parameters and significance parameters of word tuples, and screening out specific word tuples to improve the accuracy of the analysis and simplify the process.

Benefits of technology

This method can quickly locate specific word tuples closely related to model interpretability, reduce analysis complexity, improve the accuracy of results, and thus improve the credibility and transparency of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218047A_ABST
    Figure CN120218047A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a large language model interpretability analysis method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the steps of obtaining input data and a reasoning result of a large language model, obtaining relevant parameters of lexical tuple groups formed by input lexical units and result lexical units to generate a mapping matrix, calculating a neighborhood average parameter of each lexical tuple group in the mapping matrix, selecting part of lexical tuple groups of which the relevant parameters are greater than the neighborhood average parameter as candidate lexical tuple groups, and selecting the lexical tuple groups as candidate lexical tuple groups. And calculating a significance parameter corresponding to the mapping matrix based on the candidate lexical tuple group, if the significance parameter is greater than or equal to a preset significance index, taking the candidate lexical tuple group as a specific lexical tuple group, and obtaining an interpretability analysis result of the large language model according to all the specific lexical tuple groups. According to the method, the specific lexical tuple closely related to the model interpretability is rapidly positioned, the analysis workload and complexity are reduced, the result accuracy is improved, and the credibility and transparency of the large language model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical Field

[0002] This application relates to the field of artificial intelligence technology, and in particular to methods, devices, equipment, and storage media for analyzing the interpretability of large language models. Background Art

[0003] When dealing with complex natural language tasks such as complex text generation and in-depth semantic understanding, large language models demonstrate powerful generation capabilities and can generate coherent and reasonable text based on the input. Although large language models perform well in text generation and handling complex semantic relationships, due to their "black box" nature, the reasoning process, internal knowledge structure, and decision-making basis of the models are often difficult to be clearly analyzed. Especially in complex task fields such as molecular design and medical reasoning, there are often problems with insufficient model interpretability.

[0004] In related technologies, the implicit relationship between the input and output of large language models is understood by analyzing intermediate variables. This method has a relatively complicated operation process and requires in-depth understanding of the model architecture and data flow to accurately select intermediate variables and conduct complex analysis, which is time-consuming. Moreover, if the intermediate variables are not selected reasonably, the interpretability results will be inaccurate. Summary of the Invention

[0005] The main purpose of the embodiments of this application is to propose methods, devices, equipment, and storage media for analyzing the interpretability of large language models, which simplify the complexity of the interpretability analysis process of large language models and improve the accuracy of the interpretability results.

[0006] To achieve the above purpose, the first aspect of the embodiments of this application proposes a method for analyzing the interpretability of a large language model, including:

[0007] Obtain the corpus data of the large language model, where the corpus data includes input data and inference results;

[0008] Select input tokens from the input data and result tokens from the inference results, obtain the relevant parameters of the token groups formed by the input tokens and the result tokens, and generate a mapping matrix according to the relevant parameters;

[0009] Calculate the neighborhood average parameter of each token group in the mapping matrix, and select some of the token groups whose relevant parameters are greater than the neighborhood average parameter as candidate token groups;

[0010] Based on the candidate token groups, calculate the significance parameter corresponding to the mapping matrix. If the significance parameter is greater than or equal to the preset significance index, use the candidate token groups as specific token groups;

[0011] Obtain the interpretability analysis result of the large language model based on all the specific word tuples.

[0012] In some embodiments, before calculating the neighborhood average parameter of each word tuple in the mapping matrix, the method further includes:

[0013] Obtain the row sum corresponding to each row of the mapping matrix, and rearrange the rows of the mapping matrix in descending order according to the row sum to obtain an intermediate mapping matrix;

[0014] Obtain the column sum corresponding to each column of the intermediate mapping matrix, and rearrange the columns of the intermediate mapping matrix in descending order according to the column sum to obtain the updated mapping matrix.

[0015] In some embodiments, calculating the neighborhood average parameter of each word tuple in the mapping matrix includes:

[0016] Obtain the neighborhood parameter of each word tuple in the mapping matrix;

[0017] Calculate the mean parameter and standard deviation parameter corresponding to the word tuple according to the neighborhood parameter;

[0018] Calculate the product of the standard deviation parameter and the preset threshold and then add the mean parameter to obtain the neighborhood average parameter.

[0019] In some embodiments, calculating the significance parameter corresponding to the mapping matrix based on the candidate word tuple includes:

[0020] Obtain the actual proportion of the candidate word tuple in the mapping matrix;

[0021] Obtain the distribution parameter of the mapping matrix, and calculate the expected proportion based on the distribution parameter and the candidate word tuple;

[0022] Calculate the significance parameter according to the expected proportion and the actual proportion.

[0023] In some embodiments, the distribution parameter includes an overall mean parameter and an overall variance parameter, and calculating the expected proportion based on the distribution parameter and the candidate word tuple includes:

[0024] Obtain the neighborhood average parameter of each word tuple, subtract the overall mean parameter from the neighborhood average parameter to obtain a first intermediate value, and calculate the quotient of the first intermediate value and the overall variance parameter;

[0025] Input the quotient into the cumulative distribution function of the standard normal distribution to obtain an intermediate probability value, and calculate the result of one minus the intermediate probability value as the expected proportion.

[0026] In some embodiments, calculating the significance parameter according to the expected ratio and the actual ratio includes:

[0027] Obtaining the matrix dimension parameter of the mapping matrix;

[0028] Calculating the difference between the actual ratio and the expected ratio to obtain a second intermediate value, calculating the product of the expected ratio and the intermediate probability value to obtain a third intermediate value, calculating the quotient of the third intermediate value and the matrix dimension parameter to obtain a fourth intermediate value, and calculating the quotient of the second intermediate value and the square root of the fourth intermediate value to obtain the significance parameter.

[0029] In some embodiments, selecting the input token from the input data and selecting the result token from the inference result includes:

[0030] Obtaining a plurality of initial input tokens from the input data and the first word frequency corresponding to each initial input token;

[0031] Obtaining a plurality of initial result tokens from the inference result and the second word frequency corresponding to each initial result token;

[0032] Selecting the initial input tokens whose first word frequency is greater than or equal to a preset word frequency as the input tokens, and selecting the initial result tokens whose second word frequency is greater than or equal to the preset word frequency as the result tokens.

[0033] To achieve the above object, a second aspect of the embodiments of the present application proposes an interpretability analysis device for large language models, including:

[0034] A data acquisition module: used to acquire the corpus data of the large language model, and the corpus data includes input data and inference results;

[0035] A matrix generation module: used to select input tokens from the input data and select result tokens from the inference result, obtain the relevant parameters of the token group composed of the input tokens and the result tokens, and generate a mapping matrix according to the relevant parameters;

[0036] A token screening module: used to calculate the neighborhood average parameter of each token group in the mapping matrix, and select some of the token groups whose relevant parameters are greater than the neighborhood average parameter as candidate token groups;

[0037] A significance analysis module: used to calculate the significance parameter corresponding to the mapping matrix based on the candidate token groups, and if the significance parameter is greater than or equal to a preset significance index, use the candidate token groups as specific token groups;

[0038] Result generation module: configured to obtain the interpretability analysis result of the large language model based on all the specific word tuples.

[0039] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0040] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, which is a storage medium that stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0041] The method, device, equipment, and storage medium for analyzing the interpretability of a large language model proposed in the embodiments of the present application obtain the corpus data of the large language model. The corpus data includes input data and inference results. Input word tokens are selected from the input data and result word tokens are selected from the inference results. The relevant parameters of the word token tuples formed by the input word tokens and the result word tokens are obtained. A mapping matrix is generated based on the relevant parameters. The neighborhood average parameter of each word token tuple in the mapping matrix is calculated. Some word token tuples whose relevant parameters are greater than the neighborhood average parameter are selected as candidate word token tuples. Based on the candidate word token tuples, the corresponding significance parameter of the mapping matrix is calculated. If the significance parameter is greater than or equal to the preset significance index, the candidate word token tuples are used as specific word token tuples, and the interpretability analysis result of the large language model is obtained based on all the specific word token tuples. In this embodiment, considering that the corpus data can intuitively reflect the inference process of the large language model, a mapping matrix is generated based on this, and the mapping matrix is used to display the association information between word tokens in the input and output processes in a structured manner. Then, according to the relationship between the neighborhood average parameter and the relevant parameter, the word token tuples that are specific to other word token tuples are screened out and used as candidate word token tuples. This operation filters out the word token tuples that are irrelevant to a specific task, making the analysis process more focused on the "specific" word token tuples with high context relevance. Subsequently, based on the significance analysis result, the importance of the specific word token tuples in the inference process is further clarified to ensure that they have statistical significance, so as to determine whether the candidate word token tuples can be used to evaluate the interpretability of the large language model. Through this analysis process, the specific word token tuples that are closely related to the model interpretability can be quickly located, reducing the analysis workload and complexity. At the same time, since the neighborhood average parameter is derived based on the relative position and correlation of the word token tuples in the entire mapping matrix, the selected specific word token tuples can reflect the actual inference process of the model. Therefore, the interpretability analysis result obtained based on these specific word token tuples can more accurately reflect the internal mechanism and decision logic of the large language model, thereby improving the accuracy of the result and ultimately enhancing the credibility and transparency of the large language model. Description of the Drawings

[0042] Figure 1 is a flowchart of the method for analyzing the interpretability of large language models provided by the embodiments of the present application.

[0043] Figure 2 is a flowchart of selecting input tokens from input data and result tokens from inference results provided by the embodiments of the present application.

[0044] Figure 3 is a flowchart of updating the mapping matrix provided by the embodiments of the present application.

[0045] Figure 4 is a flowchart of calculating the neighborhood average parameter of each token group in the mapping matrix provided by the embodiments of the present application.

[0046] Figure 5 is a flowchart of calculating the significance parameter corresponding to the mapping matrix based on candidate token groups provided by the embodiments of the present application.

[0047] Figure 6 is a flowchart of calculating the expected proportion based on the distribution parameter and candidate token groups provided by the embodiments of the present application.

[0048] Figure 7 is a flowchart of calculating the significance parameter based on the expected proportion and the actual proportion provided by the embodiments of the present application.

[0049] Figure 8 is a schematic diagram of the application process of the method for analyzing the interpretability of large language models provided by the embodiments of the present application.

[0050] Figure 9 is a structural block diagram of the device for analyzing the interpretability of large language models provided by another embodiment of the present application.

[0051] Figure 10 is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. Detailed Embodiments

[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the flowchart.

[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs. The terms used herein are for the purpose of describing embodiments of this application only and are not intended to limit this application.

[0055] First, several terms involved in this application are analyzed as follows:

[0056] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information processes of human consciousness and thinking. Artificial intelligence also refers to the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0057] When dealing with complex natural language tasks, such as complex text generation, deep semantic understanding, etc., large language models demonstrate powerful generation capabilities and can generate coherent and reasonable text based on the input. Although large language models perform excellently in text generation and handling complex semantic relationships, due to their "black box" characteristics, the reasoning process, internal knowledge structure, and decision-making basis of the models are often difficult to be clearly analyzed. Especially in complex task fields such as molecular design and medical reasoning, there are often problems of insufficient model interpretability. This lack of transparency not only increases the difficulty of model debugging and optimization but also limits its wide application in high-risk application fields.

[0058] In related technologies, the implicit relationship between the input and output of large language models is understood by analyzing intermediate variables. This way has a relatively complicated operation process, requires in-depth understanding of the model architecture and data flow to accurately select intermediate variables and conduct complex analysis, and the process is time-consuming. Moreover, if the intermediate variables are not selected reasonably, the interpretability results will be inaccurate.

[0059] Based on this, the embodiments of the present application provide a method, apparatus, device, and storage medium for analyzing the interpretability of large language models. Considering that the corpus data can intuitively reflect the reasoning process of large language models, a mapping matrix is generated based on this, and the correlation information between tokens during the input and output processes is displayed in a structured manner by means of this mapping matrix. Then, according to the relationship between the neighborhood average parameter and the relevant parameters, the token groups that are specific to other token groups are screened out and used as candidate token groups. This operation is used to filter out the token groups that are irrelevant to a specific task, making the analysis process more focused on the "specific" token groups with high context relevance. Subsequently, according to the significance analysis results, the importance of the specific token groups in the reasoning process is further clarified to ensure its statistical significance, so as to determine whether the candidate token groups can be used to evaluate the interpretability of large language models. Through this analysis process, the specific token groups closely related to the model interpretability can be quickly located, reducing the analysis workload and complexity. At the same time, since the neighborhood average parameter is derived based on the relative position and correlation of the token groups in the entire mapping matrix, the selected specific token groups can reflect the actual reasoning process of the model. Therefore, the interpretability analysis results obtained based on these specific token groups can more accurately reflect the internal mechanism and decision-making logic of large language models, thereby improving the accuracy of the results and ultimately enhancing the credibility and transparency of large language models.

[0060] The embodiments of the present application provide a method, apparatus, device, and storage medium for analyzing the interpretability of large language models, which will be specifically described through the following embodiments. First, the method for analyzing the interpretability of large language models in the embodiments of the present application will be described.

[0061] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0062] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0063] The method for analyzing the interpretability of a large language model provided by the embodiments of this application relates to the field of artificial intelligence technology. The method for analyzing the interpretability of a large language model provided by the embodiments of this application can be applied to a terminal, or to a server, or can be a computer program running on a terminal or a server. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client that supports the analysis of the interpretability of a large language model, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module, or plug-in. Among them, the terminal communicates with the server through a network. The method for analyzing the interpretability of a large language model can be executed by the terminal or the server, or jointly executed by the terminal and the server.

[0064] In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, or a smart watch, etc. The server can be an independent server, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (Content Delivery Network, CDN), and big data and artificial intelligence platforms; it can also be a service node in a blockchain system, and the service nodes in this blockchain system form a peer-to-peer (Peer To Peer, P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP) protocol. The terminal and the server can be connected through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not limit this here.

[0065] This application can be used in numerous general - purpose or special - purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi - processor systems, microprocessor - based systems, set - top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer - executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0066] It should be noted that in each specific embodiment of this application, when it comes to relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when an embodiment of this application needs to obtain sensitive personal information of a user, it will obtain the user's separate permission or separate consent through methods such as pop - up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user - related data for the normal operation of the embodiment of this application will be obtained.

[0067] The following describes the method for analyzing the interpretability of the large - language model in the embodiments of this application.

[0068] Figure 1 is an optional flowchart of the method for analyzing the interpretability of the large - language model provided by the embodiments of this application. Figure 1 The method in may include but is not limited to steps 110 to 150. At the same time, it can be understood that this embodiment does not specifically limit the order of steps 110 to 150 in Figure 1 and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.

[0069] Step 110: Obtain the corpus data of the large - language model.

[0070] In one embodiment, the large language model, which is the large language model to be subjected to interpretability analysis, and the corpus data is the historical data of the large language model, including input data and inference results. Among them, the input data is the text information provided by the user to the large language model, which can be simple questions, complex instructions, long articles, etc. The inference result is the text content output by the large language model after reasoning based on the input data by using the knowledge and algorithms it has learned.

[0071] Step 120: Select input tokens from the input data and result tokens from the inference results, obtain the relevant parameters of the token pairs composed of the input tokens and the result tokens, and generate a mapping matrix according to the relevant parameters.

[0072] In one embodiment, tokens need to be extracted from the corpus data for subsequent analysis. Refer to Figure 2 , Figure 2 is the flowchart for selecting input tokens from the input data and result tokens from the inference results provided by the embodiments of the present application, which specifically includes the following steps:

[0073] Step 210: Obtain multiple initial input tokens from the input data and the first token frequency corresponding to each initial input token.

[0074] In one embodiment, the input data is one or more segments of text, and these texts need to be segmented into smaller units, namely tokens. Here, the tokens can be words, sub-words, characters, etc., which are specifically determined according to the actual tokenization method adopted. For example, for the input data: "I like reading books", the obtained initial input tokens can be: "I", "like", "reading", "books".

[0075] After obtaining the initial input tokens, count the number of times each initial input token appears in the entire set of input data to obtain the first token frequency corresponding to each initial input token. For example, in an input data containing multiple texts, "reading" appears 10 times, then the first token frequency of "reading" is 10.

[0076] Step 220: Obtain multiple initial result tokens from the inference results and the second token frequency corresponding to each initial result token.

[0077] In one embodiment, in the above-mentioned manner, multiple initial result tokens are obtained from the inference results, and at the same time, the second token frequency corresponding to each initial result token is obtained.

[0078] Step 230: Select the initial input tokens with the first word frequency greater than or equal to the preset word frequency as input tokens, and select the initial result tokens with the second word frequency greater than or equal to the preset word frequency as result tokens.

[0079] In one embodiment, in order to simplify data processing and improve subsequent processing efficiency, this embodiment needs to select high-frequency tokens for analysis. Therefore, a preset word frequency is set as the screening criterion to select tokens with higher frequencies from a large number of initial input tokens and initial result tokens as high-frequency tokens. The importance of high-frequency tokens in the reasoning process is usually higher than that of low-frequency tokens. Therefore, this screening process can highlight key information and improve the analysis efficiency and accuracy.

[0080] In one embodiment, the first word frequency of each initial input token is compared with the preset word frequency. Suppose the preset word frequency is 5. If the first word frequency of an initial input token "reading" is 10, since 10≥5, "reading" will be selected as an input token; while if the first word frequency of another initial input token "I" is 3, since 3<5, "I" will not be selected. Similarly, the second word frequency of each initial result token is compared with the preset word frequency. Then, n initial input tokens with the first quantity are obtained as input tokens, and m initial result tokens with the second quantity are obtained as result tokens.

[0081] It can be understood that the embodiments of the present application further screen the input tokens and result tokens, and eliminate meaningless tokens or tokens irrelevant to the current task, so as to be more conducive to accurately capturing the key knowledge relied on by the model in the reasoning process of a specific task.

[0082] Next, obtain the relevant parameters of the token groups composed of input tokens and result tokens, and generate a mapping matrix according to the relevant parameters.

[0083] Among them, the relevant parameters characterize the correlation between the input tokens and the result tokens, and their values can be determined according to the frequency of co-occurrence of the input tokens and the result tokens in the corpus data. Given that there is a corresponding relationship between the input data and the reasoning result, and the input data may cover multiple different texts. For example, assume that the input data contains 10 texts, and there are 10 corresponding reasoning result texts. If an input token appears in 8 texts of the input data, and among these 8 texts, 6 texts contain a certain result token, at this time, it can be determined that the relevant parameter between the input token and the result token is 0.6.

[0084] After obtaining the relevant parameters between each input token and different result tokens in the above manner, a mapping matrix is generated based on these relevant parameters. Among them, the relevant parameters serve as the elements at the corresponding positions in the mapping matrix. Suppose the number of input tokens is the first quantity \(n\), and the number of result tokens is the second quantity \(m\), then the dimension of the mapping matrix is \(n\times m\). With the help of the mapping matrix, the correlation between the input tokens and the result tokens can be visually presented. If the value of the element at the corresponding position in the mapping matrix is large, it can be determined that the token pair composed of the input token and the result token at that position has a high correlation.

[0085] In one embodiment, after obtaining the mapping matrix, a further filtering operation is performed based on this matrix. Specifically, the "general" token pairs that frequently appear generally are filtered out, and key attention is paid to the "specific" token pairs that only appear in specific task contexts. Such token pairs generally carry more abundant and valuable context information. For example, common words like "of", "is", "in" appear frequently in almost all types of texts. However, for explaining the reasoning process of the large language model in a specific task, the help they can provide is relatively limited. In the reasoning process of certain specific tasks, the frequency of occurrence of some token pairs may be low, but as long as they appear, they symbolize the core knowledge of the task. Taking the reasoning process of a medical diagnosis task as an example, the combination of token pairs such as "symptom", "disease name", "diagnostic index", although their frequency of occurrence may not be as high as that of general token pairs, represents the core knowledge of this task and is crucial for understanding the decision-making process of the model in medical diagnosis.

[0086] Therefore, the embodiments of the present application also need to use a local specificity filtering method to eliminate the token pairs in the mapping matrix that are not important in a specific context, ensuring that only the token pairs that are crucial in a specific context are finally retained, thereby improving the accuracy of the interpretability analysis results.

[0087] In one embodiment, in order to perform local specificity filtering on the mapping matrix, it is first necessary to rearrange it. Referring to Figure 3 , Figure 3 is the flowchart for updating the mapping matrix provided by the embodiments of the present application, which specifically includes the following steps:

[0088] Step 310: Obtain the row sum corresponding to each row of the mapping matrix, and rearrange the rows of the mapping matrix in descending order of the row sum to obtain an intermediate mapping matrix.

[0089] In one embodiment, for the previously generated mapping matrix with dimensions n×m, that is, it includes n rows. For the elements in each row, calculate the sum of its elements as the row total of that row. Suppose a row of data in the mapping matrix is [0.2, 0.8, 0.1, 0.4], then the row total of this row is: 0.2 + 0.3 + 0.1 + 0.4 = 1.5.

[0090] After calculating the row totals of all rows, sort the row totals in descending order. After sorting, rearrange the rows of the mapping matrix according to the new order of the row totals to obtain an intermediate mapping matrix. That is to say, in the intermediate mapping matrix, the rows with larger row totals are arranged in the front. A larger row total indicates that the input token corresponding to that row has a higher total correlation with all result tokens, that is, this input token has a stronger association in the input-output relationship of the model. Through this rearrangement, the input tokens that play a more crucial role in the model inference process can be concentrated in the front part of the mapping matrix.

[0091] Step 320: Obtain the column total corresponding to each column of the intermediate mapping matrix, and rearrange the columns of the intermediate mapping matrix in descending order of the column total to obtain an updated mapping matrix.

[0092] In one embodiment, for the intermediate mapping matrix, calculate the column total of each column in the same way as calculating the row total. When the column totals of all columns are calculated, sort these column totals in descending order. Subsequently, rearrange the columns of the intermediate mapping matrix according to the new order generated by the column total sorting, thereby obtaining an updated mapping matrix. In this process, the larger the value of the column total, the higher the total correlation between the result token corresponding to that column and all input tokens, that is, this result token is associated with more different input tokens in the input-output relationship of the model. Therefore, in this embodiment, through the column rearrangement operation, the result tokens that are closely related to more input tokens can be arranged in the front part of the mapping matrix.

[0093] It can be seen that after the above rearrangement and update steps of the mapping matrix, the updated mapping matrix can more intuitively display the association between input tokens and result tokens. The input tokens and result tokens corresponding to the rows and columns in the front of the mapping matrix play a more important role in the model inference process, and adjacent elements are more likely to come from rows or columns with similar totals, thereby enhancing the similarity of adjacent tokens. The updated mapping matrix helps to quickly determine which input tokens are more likely to cause which result tokens to appear, thereby facilitating the understanding of the model's decision-making mechanism and enhancing the interpretability of the model.

[0094] In one embodiment, assume element A in the mapping matrix ijFollowing the independent and identically distributed (such as normal distribution), in the updated mapping matrix A', the row sum R i and the column sum C j will be similar to the normal distribution. Due to sorting, the relationship between adjacent elements in the updated mapping matrix is closer, which can be expressed as:

[0095] P(A k ′ l ≈A k ′ ±1,l±1 |R i ≈R i±1 ,C j ≈C j±1 )>P(A ij ≈A i±1,j±1 )

[0096] The meaning of the above formula is: in the updated mapping matrix, if the row sum R i of a certain row is similar to the row sum R i±1 of the adjacent row, and the column sum C j of a certain column is similar to the column sum C j±1 of the adjacent column, then the element A k ′ l at the position (k, l) in the mapping matrix is more relevant to the element A k ′ ±1,l±1 at its adjacent position (k±1, l±1) than the element A ij at the position (i, j) in the non-updated mapping matrix is to the element A i±1,j±1 at its adjacent position (i±1, j±1).

[0097] The above formula shows that in the sorted mapping matrix, the values between adjacent elements are more likely to be close, while in the initial mapping matrix, the similarity between adjacent elements is weak. Therefore, in the embodiment of the present application, through the sorting operation, the updated mapping matrix enhances the similarity between adjacent word tuples, which helps to discover word tuples that have a specific impact on model inference.

[0098] Step 130: Calculate the neighborhood average parameter of each word tuple in the mapping matrix, and select some word tuples whose relevant parameters of the word tuples are greater than the neighborhood average parameter as candidate word tuples.

[0099] In one embodiment, it is necessary to screen specific word tuples from the mapping matrix. Therefore, it is necessary to determine whether the word tuples in the mapping matrix deviate significantly from the distribution of their surrounding word tuples. Referring to Figure 4 ,[[]]END]] Figure 4 is the flowchart for calculating the neighborhood average parameter of each word tuple in the mapping matrix provided by the embodiment of the present application, which specifically includes the following steps:

[0100] Step 410: Obtain the neighborhood parameters of each word tuple in the mapping matrix.

[0101] In one embodiment, since each word tuple corresponds to an element in the mapping matrix. For example, if the input word and the result word form a word tuple, the element value at their intersection position in the mapping matrix (assuming the input word corresponds to the row and the result word corresponds to the column) represents the correlation parameter between these two words. Therefore, taking the position of the word tuple in the mapping matrix as the center, select some tuples as neighborhood word tuples in the up, down, left, and right directions, and use the correlation parameters corresponding to the neighborhood word tuples as neighborhood parameters. For example, the word tuple corresponding to the position (i, j), and its neighborhood word tuples can be the word tuples corresponding to the positions (i±a, j±b), where a and b are set according to actual requirements.

[0102] Step 420: Calculate the mean parameter and the standard deviation parameter corresponding to the word tuple according to the neighborhood parameters.

[0103] In one embodiment, assume that the set of neighborhood parameters is N = {n1, n2, …, nk}, where k is the number of neighborhood word tuples. At this time, the mean parameter μ neighbor refers to the average value of the neighborhood parameters, and the standard deviation parameter σ neighbor is used to measure the degree of dispersion of the neighborhood parameters relative to the mean parameter. When calculating, first, calculate the square of the difference between each neighborhood parameter and the mean parameter, then find the average of these squared values, and finally take the square root of this average to obtain the standard deviation parameter.

[0104] In the embodiments of the present application, the mean parameter provides a benchmark for the neighborhood correlation of the word tuple. If the mean parameter is high, it indicates that the correlation between the word tuple and its surrounding neighborhood word tuples is generally strong, indicating that in the model, the input-output relationship represented by the word tuple has a high consistency within its neighborhood semantic range. For example, in a text classification task, if the mean parameter of a word tuple representing a specific category (such as the "technology" category) is high, it means that the neighborhood semantics related to "technology" (such as "electronic devices", "software", etc.) have a high degree of association with the "technology" category in the model, and the model processes these related semantics more consistently.

[0105] The standard deviation parameter reflects the stability of the model when processing neighborhood semantics. When the standard deviation parameter is low, it indicates that when the model processes the neighborhood semantics related to the word tuple, its correlation is relatively stable, and the decision-making of the model has a certain reliability. A high standard deviation parameter means that there is a large uncertainty in the model when processing neighborhood semantics, and it is necessary to further analyze the processing logic of the model on these semantics to optimize the model performance or improve its interpretability.

[0106] Step 430: Calculate the product of the standard deviation parameter and the preset threshold, and then add the mean parameter to obtain the neighborhood average parameter.

[0107] In one embodiment, the neighborhood average parameter is expressed as:

[0108] μ neighbor + T·σ neighbor

[0109] where T represents the preset threshold, and the preset threshold can be set according to the actual situation.

[0110] At this time, if the relevant parameter of the token tuple is greater than the neighborhood average parameter, then this token tuple is considered to be statistically significantly deviated from the distribution of its surrounding token tuples. Therefore, this part of the token tuples is used as candidate token tuples. That is, taking the token tuple corresponding to the position (i, j) as an example, its relevant parameter is A ij , if it satisfies: A ij > μ neighbor + T ·σ neighbor , then the token tuple corresponding to the position (i, j) is considered a candidate token tuple.

[0111] According to the above process, each token tuple in the updated mapping matrix is judged, and the token tuples that meet the preset threshold conditions are selected as candidate token tuples.

[0112] Step 140: Based on the candidate token tuples, calculate the significance parameter corresponding to the mapping matrix. If the significance parameter is greater than or equal to the preset significance index, the candidate token tuples are used as specific token tuples.

[0113] In one embodiment, after obtaining the candidate token tuples, the embodiment of the present application also needs to perform a significance test on the mapping matrix. According to the significance analysis result, further clarify the importance of the specific token tuples obtained from the mapping matrix in the inference process, ensure that it has statistical significance, and thus determine whether the candidate token tuples can be used to evaluate the interpretability of the large language model. Refer to Figure 5 , Figure 5 is the flowchart for calculating the significance parameter corresponding to the mapping matrix based on the candidate token tuples provided by the embodiment of the present application, which specifically includes the following steps:

[0114] Step 510: Obtain the actual proportion of the candidate token tuples in the mapping matrix.

[0115] In one embodiment, the actual proportion refers to the ratio of the number of candidate token tuples to the total number of token tuples in the mapping matrix, and is expressed as:

[0116]

[0117] where P actualIndicates the actual proportion, I ij Is an indicator function, which is 1 when the token tuple is a candidate token tuple and 0 otherwise. n×m represents the matrix dimension parameter of the mapping matrix, indicating the total number of token tuples in the mapping matrix.

[0118] Step 520: Obtain the distribution parameters of the mapping matrix, and calculate the expected proportion based on the distribution parameters and the candidate token tuple.

[0119] In one embodiment, since the mapping matrix satisfies independent and identical distribution and can be a normal distribution, the mean and variance of the mapping matrix can be calculated as the distribution parameters. At this time, the distribution parameters include the overall mean parameter μ and the overall variance parameter σ.

[0120] Refer to Figure 6 , Figure 6 Is the flowchart for calculating the expected proportion based on the distribution parameters and the candidate token tuple provided by the embodiment of the present application, specifically including the following steps:

[0121] Step 610: Obtain the neighborhood average parameter of each token tuple, subtract the overall mean parameter from the neighborhood average parameter to obtain the first intermediate value, and calculate the quotient of the first intermediate value and the overall variance parameter.

[0122] In one embodiment, the first intermediate value is expressed as:

[0123] μ neighbor +T·σ neighbor -μ

[0124] The quotient is expressed as:

[0125]

[0126] Step 620: Input the quotient into the cumulative distribution function of the standard normal distribution to obtain the intermediate probability value, and calculate the result of one minus the intermediate probability value as the expected proportion.

[0127] In one embodiment, the intermediate probability value is expressed as:

[0128]

[0129] Among them, Φ represents the cumulative distribution function of the standard normal distribution.

[0130] Furthermore, the expected proportion P expected Is expressed as:

[0131]

[0132] Step 530: Calculate the significance parameter according to the expected proportion and the actual proportion.

[0133] In one embodiment, refer to Figure 7, Figure 7 is a flowchart for calculating a significance parameter based on an expected ratio and an actual ratio provided by an embodiment of the present application, which specifically includes the following steps:

[0134] Step 710: Obtain the matrix dimension parameter of the mapping matrix.

[0135] Step 720: Calculate the difference between the actual ratio and the expected ratio to obtain a second intermediate value, calculate the product of the expected ratio and the intermediate probability value to obtain a third intermediate value, calculate the quotient of the third intermediate value and the matrix dimension parameter to obtain a fourth intermediate value, and calculate the quotient of the second intermediate value and the square root of the fourth intermediate value to obtain the significance parameter.

[0136] In one embodiment, the second intermediate value is expressed as:

[0137] P actual -P expected

[0138] The third intermediate value is expressed as:

[0139] P expected (1 - P expected )

[0140] The fourth intermediate value is expressed as:

[0141]

[0142] The significance parameter is expressed as:

[0143]

[0144] After obtaining the significance parameter corresponding to the mapping matrix, it is judged against a preset significance index. If the significance parameter is greater than or equal to the preset significance index, it is considered that the result of the mapping matrix conforms to statistical laws, and it is considered that the candidate token group is significantly higher than the random distribution statistically, and it plays an important role in the model learning process. All candidate token groups are used as specific token groups.

[0145] Step 150: Obtain the interpretability analysis result of the large language model based on all specific token groups.

[0146] In one embodiment, after the above-mentioned local specificity filtering and statistical significance analysis, the selected specific token groups are used as the interpretability analysis results and presented to the user, so that information mining can be performed on the prediction data and the large language model itself based on the specific token groups, so that it can be viewed through a graphical interface or other visualization methods which token groups play a decisive role in the inference results of the model, thereby helping the user understand the decision-making process of the large language model, providing guidance for further optimizing the model, and guiding the user on how to adjust the model parameters or select appropriate training data.

[0147] In the large language model interpretability analysis method provided by the embodiment of the present application, a mapping matrix is first generated, and then local specificity filtering is used to screen the high-frequency token groups in the mapping matrix, removing those "general" token groups that are irrelevant to model inference, which can accurately extract the knowledge mapping pairs of the model and provide a clear interpretability framework, using the relationship between the input tokens and the result tokens to reveal how the model processes and learns the input data, and extracting meaningful knowledge from it. Then, statistical significance analysis is used to further confirm whether the selected candidate token groups are statistically significant by calculating the significance parameters, avoiding the interference caused by ordinary token groups, and ensuring that the extracted features are more scientific and representative. Furthermore, interpretability analysis is performed using specific token groups to achieve the interpretability analysis of the large language model (LLM).

[0148] In one embodiment, taking the description of biochemical small molecules as an example of the application scenario of the large model, the input data of the large language model is molecular modal data, such as SMILES, SELFIES or IUPAC name, and the output data of the large language model is molecular description data. Refer to Figure 8 , Figure 8 is a schematic diagram of the application process of the large language model interpretability analysis method provided by the embodiment of the present application.

[0149] First, data preparation is performed to ensure that the input data meets the requirements of large language model analysis. Among them, the data preparation process includes multiple sub-steps such as data preprocessing, large model construction, and model output collation, aiming to convert the original input data into a format that can be input into the large language model and prepare for subsequent analysis.

[0150] In data preparation, it is necessary to ensure that the input data meets the requirements of large language model analysis. Among them, the data preparation process covers multiple sub-steps such as data preprocessing, large model construction, and model output collation, aiming to convert the original input data into a format that can be input into the large language model and lay a foundation for subsequent analysis. Specifically, in the data preprocessing step, the original data needs to be cleaned and formatted. This includes removing noisy data, unifying the data format, and converting text data into a format suitable for input into the large language model (such as converting to standard text or molecular representation, etc.). Data preprocessing aims to ensure the quality of the input data and reduce the interference of irrelevant information. Subsequently, the large language model is constructed. After completing the data preparation, the input data is input into the pre-trained large language model. This step is a key link in constructing the model inference process, ensuring that the model can carry out effective learning and inference based on the input data. Finally, the model output collation work is implemented. The inference results generated by the model need to be further collated and formatted to facilitate subsequent analysis. This mainly involves converting the output of the model's inference results into token form and removing irrelevant noise tokens to ensure the regularity of the output.

[0151] Subsequently, a mapping matrix is generated. This step is the basis for conducting interpretability analysis. In this process, first, the input data and the inference results are respectively converted into token sequences. Then, tokens are screened. Multiple initial input tokens and the corresponding first token frequencies of each initial input token are obtained from the input data, and multiple initial result tokens and the corresponding second token frequencies of each initial result token are obtained from the inference results. Then, the initial input tokens with the first token frequency greater than or equal to the preset token frequency are selected as input tokens, and the initial result tokens with the second token frequency greater than or equal to the preset token frequency are selected as result tokens. After that, a mapping matrix is constructed based on the relevant parameters corresponding to the input tokens and the result tokens to show the relationship between the input data and the inference results. Each element in the mapping matrix represents the matching degree of the corresponding token group and also reveals the way the model infers based on the input.

[0152] Subsequently, it enters the local specificity filtering stage. In this step, a rearrangement operation is performed on the mapping matrix. Specifically, first obtain the row sum corresponding to each row of the mapping matrix, and then rearrange the rows of the mapping matrix in descending order according to the row sum to obtain an intermediate mapping matrix. After that, obtain the column sum corresponding to each column of the intermediate mapping matrix, and rearrange the columns of the intermediate mapping matrix in descending order according to the column sum to obtain an updated mapping matrix. After completing the update of the mapping matrix, the screening work is carried out based on this updated mapping matrix. The specific method is to calculate the neighborhood average parameter of each token group in the mapping matrix, and then select some token groups with relevant parameters greater than the neighborhood average parameter as candidate token groups. Through this operation, high-frequency token groups without specific context can be eliminated, and token groups with higher relevance to the task can be retained as candidate token groups.

[0153] Then, a statistical significance analysis is conducted. The purpose is to verify whether the selected candidate token groups are statistically significant through mathematical and statistical methods, so as to ensure that the extracted candidate token groups truly represent the knowledge learning preferences of the model. In this process, the significance parameter needs to be determined based on the calculated mean parameter and standard deviation parameter of each one, and whether the candidate token groups can be used as specific token groups is determined according to the preset significance index and significance parameter to improve the model transparency.

[0154] After completing the statistical significance analysis, it enters the result output and visualization stage, where the specific token groups are analyzed, or they are summarized to generate the results of interpretability analysis. They can also be visualized to intuitively display the analysis results and help users more intuitively understand the knowledge learning preferences of the model.

[0155] After obtaining the results of interpretability analysis, they are applied and adjusted. Based on this, the response pattern of the model to a specific task can be analyzed. The large language model can be optimized based on the analysis results by modifying the training parameters or strategies of the model to improve the performance of the model in a specific task. Or, based on the analysis results, certain specific tokens or token groups can be analyzed in depth to further explore the internal reasoning mechanism of the large language model, which helps to discover the potential rules in the reasoning process of the model and enhance the credibility and controllability of the model.

[0156] The technical solution provided by the embodiment of the present application obtains the corpus data of the large language model. The corpus data includes input data and inference results. Select input tokens from the input data and result tokens from the inference results, obtain the relevant parameters of the token groups formed by the input tokens and the result tokens, generate a mapping matrix according to the relevant parameters, calculate the neighborhood average parameters of each token group in the mapping matrix, and select some token groups whose relevant parameters are greater than the neighborhood average parameters as candidate token groups. Based on the candidate token groups, calculate the significance parameter corresponding to the mapping matrix. If the significance parameter is greater than or equal to the preset significance index, regard the candidate token groups as specific token groups, and obtain the interpretability analysis result of the large language model according to all the specific token groups. In this embodiment, considering that the corpus data can intuitively reflect the inference process of the large language model, a mapping matrix is generated based on this, and the correlation information between tokens in the input and output processes is displayed in a structured manner by means of this mapping matrix. Then, according to the relationship between the neighborhood average parameter and the relevant parameter, filter out the token groups that are specific to other token groups and regard them as candidate token groups. Use this operation to filter out the token groups irrelevant to a specific task, so that the analysis process focuses more on the "specific" token groups with high context relevance. Subsequently, according to the significance analysis result, further clarify the importance of the specific token groups in the inference process to ensure that they have statistical significance, and thus determine whether the candidate token groups can be used to evaluate the interpretability of the large language model. Through this analysis process, it is possible to quickly locate the specific token groups closely related to the model interpretability, reduce the analysis workload and complexity. At the same time, since the neighborhood average parameter is derived based on the relative position and correlation of the token groups in the entire mapping matrix, the selected specific token groups can reflect the actual inference process of the model. Therefore, the interpretability analysis result obtained based on these specific token groups can more accurately reflect the internal mechanism and decision logic of the large language model, thereby improving the accuracy of the result and ultimately enhancing the credibility and transparency of the large language model.

[0157] The embodiment of the present application also provides an interpretability analysis device for a large language model, which can implement the above-mentioned interpretability analysis method for the large language model. Refer to Figure 9 and this device includes:

[0158] A data acquisition module 910: used to acquire the corpus data of the large language model, and the corpus data includes input data and inference results.

[0159] A matrix generation module 920: used to select input tokens from the input data and result tokens from the inference results, obtain the relevant parameters of the token groups formed by the input tokens and the result tokens, and generate a mapping matrix according to the relevant parameters.

[0160] Token screening module 930: It is used to calculate the neighborhood average parameter of each token group in the mapping matrix, and select some token groups whose relevant parameters of the token group are greater than the neighborhood average parameter as candidate token groups.

[0161] Significance analysis module 940: It is used to calculate the significance parameter corresponding to the mapping matrix based on the candidate token groups. If the significance parameter is greater than or equal to the preset significance index, the candidate token groups are used as specific token groups.

[0162] Result generation module 950: It is used to obtain the interpretability analysis result of the large language model according to all specific token groups.

[0163] The specific implementation manner of the large language model interpretability analysis device in this embodiment is basically the same as that of the above large language model interpretability analysis method, and will not be elaborated here.

[0164] This application embodiment also provides an electronic device, including:

[0165] At least one memory;

[0166] At least one processor;

[0167] At least one program;

[0168] The program is stored in the memory, and the processor executes the at least one program to implement the large language model interpretability analysis method described above in this application. This electronic device can be any intelligent terminal including mobile phones, tablet computers, personal digital assistants (PDAs), in-vehicle computers, etc.

[0169] Please refer to Figure 10 , Figure 10 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0170] Processor 1001, which can be implemented in ways such as a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by this application embodiment;

[0171] The memory 1002 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1002 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1002 and are called by the processor 1001 to execute the large language model interpretability analysis method of the embodiments of this application;

[0172] The input / output interface 1003 is used to implement information input and output;

[0173] The communication interface 1004 is used to implement communication interaction between this device and other devices. It can communicate through a wired method (such as USB, network cable, etc.) or through a wireless method (such as a mobile network, WIFI, Bluetooth, etc.); and

[0174] The bus 1005 transmits information between the various components of the device (such as the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004);

[0175] Among them, the processor 1001, the memory 1002, the input / output interface 1003, and the communication interface 1004 are communicatively connected to each other inside the device through the bus 1005.

[0176] The embodiments of this application also provide a storage medium. The storage medium is a storage medium that stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned large language model interpretability analysis method.

[0177] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0178] The method, apparatus, device, and storage medium for analyzing the interpretability of a large language model proposed in the embodiments of the present application obtain the corpus data of the large language model, where the corpus data includes input data and inference results. Input tokens are selected from the input data, and result tokens are selected from the inference results. The relevant parameters of the token pairs formed by the input tokens and the result tokens are obtained, and a mapping matrix is generated according to the relevant parameters. The neighborhood average parameter of each token pair in the mapping matrix is calculated, and some of the token pairs whose relevant parameters are greater than the neighborhood average parameter are selected as candidate token pairs. Based on the candidate token pairs, the significance parameter corresponding to the mapping matrix is calculated. If the significance parameter is greater than or equal to the preset significance index, the candidate token pairs are used as specific token pairs, and the interpretability analysis result of the large language model is obtained according to all the specific token pairs. In this embodiment, considering that the corpus data can intuitively reflect the inference process of the large language model, a mapping matrix is generated based on this, and the association information between tokens in the input and output processes is displayed in a structured manner by means of this mapping matrix. Then, according to the relationship between the neighborhood average parameter and the relevant parameter, the token pairs that are specific to other token pairs are screened out and used as candidate token pairs. This operation is used to filter out the token pairs that are irrelevant to a specific task, making the analysis process more focused on the "specific" token pairs with higher context relevance. Subsequently, according to the significance analysis result, the importance degree of the specific token pairs in the inference process is further clarified to ensure that it has statistical significance, so as to determine whether the candidate token pairs can be used to evaluate the interpretability of the large language model. Through this analysis process, the specific token pairs closely related to the model interpretability can be quickly located, reducing the analysis workload and complexity. At the same time, since the neighborhood average parameter is derived based on the relative position and correlation of the token pairs in the entire mapping matrix, the selected specific token pairs can reflect the actual inference process of the model. Therefore, the interpretability analysis result obtained based on these specific token pairs can more accurately reflect the internal mechanism and decision logic of the large language model, thereby improving the accuracy of the result and ultimately enhancing the credibility and transparency of the large language model.

[0179] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0180] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine some steps, or different steps.

[0181] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0182] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0183] In the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0184] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0185] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0186] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0187] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0188] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0189] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.

Claims

1. A large language model interpretability analysis method, characterized in that: include: Acquire corpus data of a large language model, wherein the corpus data includes input data and inference results; Selecting an input word-gram from the input data and selecting a result word-gram from the inference result, obtaining relevant parameters of a word-gram group consisting of the input word-gram and the result word-gram, and generating a mapping matrix according to the relevant parameters; Calculating the neighborhood average parameter of each word tuple in the mapping matrix, and selecting some of the word tuples whose relevant parameters are greater than the neighborhood average parameter as candidate word tuples; Based on the candidate word tuple, calculating the significance parameter corresponding to the mapping matrix, and if the significance parameter is greater than or equal to a preset significance index, taking the candidate word tuple as a specific word tuple; The interpretability analysis result of the large language model is obtained according to all the specific word tuples.

2. The large language model interpretability analysis method according to claim 1, characterized in that: Before calculating the neighborhood average parameter of each word tuple in the mapping matrix, the method further includes: Obtaining a row sum corresponding to each row of the mapping matrix, and rearranging the rows of the mapping matrix in descending order of the row sums to obtain an intermediate mapping matrix; The column sum corresponding to each column of the intermediate mapping matrix is ​​obtained, and the columns of the intermediate mapping matrix are rearranged in descending order of the column sums to obtain an updated mapping matrix.

3. The large language model interpretability analysis method according to claim 1, characterized in that: The calculating of the neighborhood average parameter of each word tuple in the mapping matrix includes: Obtaining a neighborhood parameter of each of the word tuples in the mapping matrix; Calculate the mean parameter and standard deviation parameter corresponding to the word tuple according to the neighborhood parameter; The product of the standard deviation parameter and a preset threshold is calculated and then added to the mean parameter to obtain the neighborhood average parameter.

4. The large language model interpretability analysis method according to claim 1, characterized in that: The calculating, based on the candidate word tuple, the significance parameter corresponding to the mapping matrix comprises: Obtaining the actual proportion of the candidate word tuple in the mapping matrix; Obtaining distribution parameters of the mapping matrix, and calculating expected proportions based on the distribution parameters and the candidate word tuples; The significance parameter is calculated according to the expected ratio and the actual ratio.

5. The large language model interpretability analysis method according to claim 4, characterized in that: The distribution parameters include an overall mean parameter and an overall variance parameter, and the calculation of the expected proportion based on the distribution parameters and the candidate word tuples includes: Obtaining the neighborhood average parameter of each word tuple, subtracting the overall mean parameter from the neighborhood average parameter to obtain a first intermediate value, and calculating a quotient of the first intermediate value and the overall variance parameter; The quotient is input into the cumulative distribution function of the standard normal distribution to obtain an intermediate probability value, and a result of subtracting the intermediate probability value is calculated as the expected proportion.

6. The large language model interpretability analysis method according to claim 5, characterized in that: The calculating the significance parameter according to the expected ratio and the actual ratio comprises: Obtaining matrix dimension parameters of the mapping matrix; Calculate the difference between the actual ratio and the expected ratio to obtain a second intermediate value, calculate the product of the expected ratio and the intermediate probability value to obtain a third intermediate value, calculate the quotient of the third intermediate value and the matrix dimension parameter to obtain a fourth intermediate value, and calculate the quotient of the second intermediate value and the square root of the fourth intermediate value to obtain the significance parameter.

7. The large language model interpretability analysis method according to any one of claims 1 to 6, characterized in that: The selecting an input word from the input data and selecting a result word from the inference result includes: Obtaining a plurality of initial input word units and a first word frequency corresponding to each of the initial input word units from the input data; Obtaining a plurality of initial result word-grams and a second word frequency corresponding to each of the initial result word-grams from the inference result; The initial input word-element whose first word frequency is greater than or equal to the preset word frequency is selected as the input word-element, and the initial result word-element whose second word frequency is greater than or equal to the preset word frequency is selected as the result word-element.

8. A large language model interpretability analysis device, characterized in that: include: Data acquisition module: used to acquire corpus data of a large language model, wherein the corpus data includes input data and inference results; A matrix generation module: used for selecting input word-grams from the input data and selecting result word-grams from the inference results, obtaining relevant parameters of a word-gram group composed of the input word-grams and the result word-grams, and generating a mapping matrix according to the relevant parameters; A word tuple screening module is used to calculate the neighborhood average parameter of each word tuple in the mapping matrix, and select the word tuples whose related parameters are greater than the neighborhood average parameter as candidate word tuples; A significance analysis module: used for calculating the significance parameter corresponding to the mapping matrix based on the candidate word tuple, and if the significance parameter is greater than or equal to a preset significance index, the candidate word tuple is regarded as a specific word tuple; Result generation module: used to obtain the interpretability analysis result of the large language model based on all the specific word tuples.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the large language model interpretability analysis method described in any one of claims 1 to 7 when executing the computer program.

10. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the large language model interpretability analysis method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Interpretable deep learning method, interpretable deep learning device, computer and medium

    CN114492417A

  • Inplausible method based on few-sample relation prediction model

    CN114860953A

  • Interpretable machine learning for large scale data

    CN117616431A

  • Inplanatable analysis method, device, equipment and medium

    CN117709471A

  • KR20210148873A