Interpretability prediction method, system and device and computer readable storage medium
By obtaining input samples and user background data, screening key sample features and generating prompt words, and using a large language model to output natural language explanation text, the problem of abstract and difficult-to-understand explanation content in existing technologies is solved, and the model prediction results are made easy to understand and explain.
Patent Information
- Application Number
- CN202511139932.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-28
AI Technical Summary
Existing model interpretability methods are highly technical in expression and cannot meet the understandability needs of business personnel and users. Traditional methods explain content in an abstract and difficult way, and are not easy to understand.
By obtaining input samples and user background data, screening key sample features, generating prompt words and inputting them into a large language model, and outputting natural language explanation text, the feature contribution is converted into a natural language explanation that fits the user's cognition by combining user background data and the capabilities of the large language model.
The interpretability of model prediction results has been improved, allowing users to more intuitively understand the basis for model decisions. The explanation content fits the user's cognitive framework and meets the understandability needs of business personnel and users.
Smart Images

Figure CN120851228A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an interpretable prediction method, system, device, and computer-readable storage medium. Background Technology
[0002] In today's complex and ever-changing business scenarios, especially in critical sectors such as healthcare and finance, the interpretability of machine learning models is crucial. For example, in credit risk assessment, business personnel need to clearly understand the specific basis for the model's rejection of loan applications to verify the rationality of the decision and meet compliance requirements.
[0003] However, traditional model interpretability methods, such as Local Interpretable Model-agnostic Explanations (LIME) and Shapley Additive Explanations (SHAP), while providing explanations to some extent, still have significant shortcomings. Specifically, these methods typically present explanations in the form of feature contribution values or graphical visualizations, which are rather abstract and difficult for ordinary business users to understand. Summary of the Invention
[0004] The main objective of this application is to provide an interpretable prediction method, system, device, and computer-readable storage medium, aiming to solve the technical problem that existing model interpretability methods are highly technical and difficult to meet the requirements of understandability.
[0005] To achieve the above objectives, this application provides an interpretability prediction method, which includes:
[0006] In response to a prediction result interpretation request, input samples and user background data are obtained based on the prediction result interpretation request, wherein the input samples include multiple sample features and their corresponding feature values;
[0007] The input sample is fed into a pre-trained prediction mini-model to obtain the prediction result;
[0008] Determine the local contribution of each of the sample features to the prediction result, and select key sample features based on the local contribution.
[0009] Based on the key sample features and the user background data, prompt words are generated. The prompt words and the prediction results are input into a large language model, and a natural language explanation text for the prediction results is output.
[0010] In one embodiment, the step of generating prompt words based on the key sample features and the user background data includes:
[0011] Obtain the feature values, global importance values, and feature sensitivity values of the key sample features;
[0012] By concatenating the key sample features, the feature values of the key sample features, the local contribution of the key sample features, the global importance value, the feature sensitivity value, the user background data, and the predefined template guidance information, prompt words are obtained. The predefined template guidance information includes one or more of the following: explanatory tone, depth of detail, and domain terminology specifications.
[0013] In one embodiment, after the step of inputting the input sample into a pre-trained prediction mini-model to obtain the prediction result, the method further includes:
[0014] The input samples are vectorized to obtain the prediction results;
[0015] Check if a target feature vector matching the prediction result exists in the preset cache interpretation database;
[0016] If it exists, then the explanatory text corresponding to the target feature vector in the preset cache explanation database is determined to be the explanatory text for the prediction result;
[0017] If not, then perform the step of determining the local contribution of each of the sample features to the prediction result, and filtering key sample features based on the local contribution.
[0018] In one embodiment, after the step of inputting the prompt word and the prediction result into a large language model and outputting a natural language explanation text for the prediction result, the method further includes:
[0019] The natural language interpretation text is input into a pre-trained compliance detector, and the compliance detection result is output.
[0020] If the compliance detection result indicates that the natural language interpretation text is non-compliant, then the preset security instruction prompt word is input into the large language model so that the large language model can regenerate the natural language interpretation text;
[0021] If the compliance test result indicates that the natural language interpretation text is compliant, then the interpretability prediction process ends.
[0022] In one embodiment, before the step of inputting the natural language interpreted text into a pre-trained compliance detector and outputting a compliance detection result, the method further includes:
[0023] Obtain the compliance detector to be trained and the original training samples;
[0024] The original training samples are perturbed to obtain adversarial training samples, wherein the perturbation includes one or more of the following: sensitive word replacement, insertion of sensitive word templates, and semantic perturbation.
[0025] The compliance detector is trained adversarially based on the adversarial training samples to obtain the trained compliance detector.
[0026] In one embodiment, after the step of inputting the prompt word and the prediction result into a large language model and outputting a natural language explanation text for the prediction result, the method further includes:
[0027] In response to a request for unsatisfactory interpretation of the natural language interpreted text, feedback content is obtained based on the request for unsatisfactory interpretation.
[0028] The large language model is incrementally trained or fine-tuned using reinforcement learning based on the feedback content.
[0029] In one embodiment, the local contribution is the SHAP value, and the step of screening key sample features based on the local contribution includes:
[0030] The SHAP values are sorted in descending order to obtain a sorted queue;
[0031] Select the sample features corresponding to the first preset number of SHAP values in the sorting queue as key sample features.
[0032] Furthermore, to achieve the above objectives, this application also provides an interpretability prediction system, the interpretability prediction system comprising:
[0033] The response module is used to respond to a prediction result interpretation request and obtain input samples and user background data according to the prediction result interpretation request, wherein the input samples include multiple sample features and their corresponding feature values;
[0034] The prediction module is used to input the input samples into a pre-trained prediction mini-model to obtain prediction results;
[0035] A filtering module is used to determine the local contribution of each of the sample features to the prediction result, and to filter key sample features based on the local contribution.
[0036] The explanation module is used to generate prompt words based on the key sample features and the user background data, input the prompt words and the prediction results into the large language model, and output natural language explanation text for the prediction results.
[0037] In addition, to achieve the above objectives, this application also provides an interpretability prediction device, the interpretability prediction device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the interpretability prediction method as described above.
[0038] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of the interpretability prediction method as described above.
[0039] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the interpretability prediction method described above.
[0040] One or more technical solutions proposed in this application have at least the following technical effects:
[0041] This application effectively solves the core problem of abstract and difficult-to-understand explanations in existing technologies by integrating user background data and the capabilities of a large language model, transforming traditional feature contribution into natural language explanations that align with user cognition. Specifically, after obtaining input samples (such as customer credit data) and user background data (such as user role as "branch risk control specialist" and knowledge tag containing "regional economic indicators"), key sample features that contribute highly to the prediction results are first extracted. Then, semantic prompts are generated based on the key sample features and user background data. By introducing user background data, the key sample features are parsed within the user's specific business context, providing a semantic context for parsing. This allows features to move beyond abstract numerical representations and anchor to the user's specific business entity space, dynamically driving the large language model to generate natural language explanation text that aligns with the user's cognitive framework. This enables users to more intuitively understand the model's decision-making basis, thereby improving the understandability of the interpretable expression of the model's prediction results. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1This is a flowchart illustrating the first embodiment of the interpretability prediction method of this application;
[0045] Figure 2 This is a flowchart illustrating the third embodiment of the interpretability prediction method of this application;
[0046] Figure 3 This is a schematic diagram of the interpretation process involved in an embodiment of the interpretability prediction method of this application;
[0047] Figure 4 This is a schematic diagram of the system architecture of the interpretability prediction system of this application;
[0048] Figure 5 This is a schematic diagram of the hardware operating environment involved in the interpretability prediction method apparatus in the embodiments of this application.
[0049] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0050] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0051] In the banking and finance sector, with the widespread application of artificial intelligence and machine learning technologies, model interpretability has become a critical technical issue. Financial institutions extensively use complex machine learning models in business processes such as credit approval, risk assessment, and customer segmentation. While these models can process complex data and make highly accurate predictions, their internal logic is often difficult to understand, figuratively referred to as a "black box." For example, in credit risk assessment, banks use complex machine learning models to comprehensively analyze an applicant's credit history, income, debt level, and other multi-dimensional data to determine whether to grant a loan and to determine the loan amount and interest rate. However, the decision-making process of these models is difficult to explain to loan officers and customers, leading to a lack of trust in the model's output among business personnel, and customers struggling to understand the reasons for loan rejection or granting specific loan conditions.
[0052] Existing model interpretability techniques, such as Local Model Interpretation (LIME) and Shapley Explanation (SHAP), still have several shortcomings. First, these methods typically present explanations in the form of feature contribution values or graphical visualizations, which are abstract and difficult for ordinary business users to understand intuitively. Second, local interpretation methods like LIME are highly dependent on sampling and perturbations, easily leading to unstable interpretation results; even small changes in the input samples can result in significantly different interpretations. Third, these posterior interpretation methods often focus on the local effects of a single predicted instance, lacking the ability to generalize and generalize to the global decision-making logic; for example, LIME can only construct a local linear surrogate model near the target sample, and its interpretation results cannot be directly generalized to other samples. Finally, these methods often have high computational costs, limiting their support for real-time and large-scale scenarios. In summary, existing interpretability methods are technically demanding, volatile, and difficult to generalize, failing to meet the understandability requirements of business users.
[0053] Based on this, the main solution of this application is as follows: In response to a prediction result interpretation request, input samples and user background data are obtained according to the prediction result interpretation request, wherein the input samples include multiple sample features and their corresponding feature values; the input samples are input into a pre-trained prediction mini-model to obtain prediction results; the local contribution of each sample feature to the prediction results is determined, and key sample features are selected based on the local contribution; prompt words are generated based on the key sample features and the user background data, and the prompt words and the prediction results are input into a large language model to output natural language interpretation text for the prediction results.
[0054] This application effectively solves the core problem of abstract and difficult-to-understand explanations in existing technologies by integrating user background data and the capabilities of a large language model, transforming traditional feature contribution into natural language explanations that align with user cognition. Specifically, after obtaining input samples (such as customer credit data) and user background data (such as user role as "branch risk control specialist" and knowledge tags containing "regional economic indicators"), key sample features that contribute significantly to the prediction results (such as income volatility and the number of credit inquiries) are first extracted. Then, semantic prompts are generated based on these key sample features and user background data. By introducing user background data, key sample features are parsed within the user's specific business context, providing a semantic context for parsing. This allows features to move beyond abstract numerical representations and anchor to the user's specific business entity space, dynamically driving the large language model to generate natural language explanation text that aligns with the user's cognitive framework. This enables users to more intuitively understand the model's decision-making basis, thereby improving the understandability of the interpretable expression of the model's prediction results.
[0055] It should be noted that the execution subject of the various embodiments of the interpretability prediction method of this application can be a computing service device with data processing, network communication and program running functions, such as a server, tablet computer, personal computer, mobile phone, etc., or an interpretability prediction device capable of realizing the above functions. The various embodiments of the interpretability prediction method of this application do not impose specific limitations on this.
[0056] Based on this, this application proposes an interpretability prediction method in the first embodiment, referring to... Figure 1 As shown, the interpretability prediction method includes the following steps S10 to S40:
[0057] Step S10: In response to the prediction result interpretation request, obtain the input sample and user background data according to the prediction result interpretation request, wherein the input sample includes multiple sample features and their corresponding feature values;
[0058] This prediction result interpretation request can be initiated by a user who needs a detailed interpretation of the model's output prediction results to verify the rationality of the model's decisions, meet compliance requirements, or explain the basis of the decisions to the client. Upon receiving this prediction result interpretation request, the corresponding input sample and user background data are obtained. Specifically, the input sample can be the sample data entered by the user. The input sample contains multiple sample features and their corresponding feature values. For example, in the financial field, the input sample may include the applicant's credit score, income level, debt ratio, and other features and their specific values.
[0059] User background data forms the cognitive context base for the generated explanation. This data may include user roles (e.g., loan approvers, account managers), historical indicators (e.g., past loan records, credit history), business scenario descriptions (e.g., "auto finance loans" or "CT image diagnosis"), and domain knowledge bases (e.g., financial industry regulations, industry standards). This embodiment does not impose specific limitations on these. However, in a preferred embodiment, user background data includes user roles, historical indicators, and business scenario descriptions. User roles control the professional depth and target of the explanation; historical indicators reflect user operational preferences to optimize the explanation focus; and business scenario descriptions bind domain-specific semantic parsing logic. Through the synergistic effect of user roles, historical indicators, and business scenario descriptions, deep personalization of the explanation text is achieved while minimizing data fields. This breaks down cross-functional cognitive barriers through role adaptation, optimizes the focus on business priorities based on historical indicators, and ensures traceability of decision-making basis through scenario rule binding. Ultimately, this makes the natural language explanation targeted, understandable, and compliant and auditable, better meeting the needs of different users in different business scenarios.
[0060] Step S20: Input the input sample into the pre-trained prediction mini-model to obtain the prediction result;
[0061] The predictive mini-model can be a machine learning model trained on a large amount of labeled data, such as an ensemble tree model or a neural network model. These models learn patterns and rules from the labeled data, enabling them to quickly and accurately predict new input samples. Specifically, the predictive mini-model receives multiple sample features and their corresponding feature values from the input sample, and obtains a prediction result through internal calculations and inference. This prediction result can be, but is not limited to, an intermediate output of the model, such as a predicted probability vector. For example, in the financial field, a predicted probability vector could represent the probability distribution of an applicant under different credit risk levels, such as [0.1, 0.7, 0.2], indicating that the probability of the applicant belonging to the "low risk" level is 10%, the probability of the "medium risk" level is 70%, and the probability of the "high risk" level is 20%. This intermediate output provides a richer information foundation for subsequent interpretation generation.
[0062] Step S30: Determine the local contribution of each sample feature to the prediction result, and select key sample features based on the local contribution.
[0063] After obtaining the prediction results, the feature contribution analysis stage begins. Specific algorithms, such as SHAP value calculation or LIME local interpretation methods, can be used to calculate the local contribution of each sample feature to the prediction result, thus identifying which features have a significant impact on the prediction outcome. Subsequently, based on these local contributions, a predetermined number of sample features with the largest local contributions are selected as key sample features, i.e., the key sample features that have the greatest impact on the prediction result are identified. For example, in credit risk assessment, it may be found that credit score and debt ratio have relatively large local contributions, while other features have relatively small local contributions; therefore, credit score and debt ratio can be selected as key sample features.
[0064] By filtering key sample features, we can focus on the features that are most explanatory to the prediction results, thereby providing support for the subsequent generation of concise and effective explanatory text.
[0065] Step S40: Generate prompt words based on the key sample features and the user background data, input the prompt words and the prediction results into the large language model, and output a natural language explanation text for the prediction results.
[0066] These key sample features and background data are integrated to generate a natural language prompt. This prompt acts as a guiding framework, providing a clear direction for the large language model's generation and ensuring that the generated explanatory text aligns with the user's actual needs and business scenarios.
[0067] Subsequently, the generated prompts, along with the predictions obtained from the small prediction model, are input into the large language model. Leveraging its powerful natural language processing capabilities, the large language model generates a natural language explanation based on the context set by the prompts and the specific content of the predictions. This explanation, written in plain language, details the reasons behind the model's predictions. For example, it might explain that the applicant's low credit score indicates a possible history of late payments, increasing the risk of default; simultaneously, a high debt ratio signifies a heavy current financial burden, further impacting the model's assessment of credit risk. Therefore, the model predicts a high probability that the applicant is a high-risk customer (the model predicts the applicant is a high-risk customer). This explanation not only clearly demonstrates the basis of the model's decision but also incorporates user background data, making the explanation more targeted and understandable. This effectively solves the problem of abstract and difficult-to-understand explanations in existing technologies, meeting the needs of business personnel and users for model interpretability.
[0068] This embodiment effectively solves the core problem of abstract and difficult-to-understand explanations in existing technologies by integrating user background data and the capabilities of a large language model, transforming traditional feature contribution values into natural language explanations that align with user cognition. Specifically, after obtaining input samples (such as customer credit data) and user background data (such as user role as "branch risk control specialist" and knowledge tags containing "regional economic indicators"), key sample features that contribute significantly to the prediction results (such as income volatility and the number of credit inquiries) are first extracted. Then, semantic prompts are generated based on these key sample features and user background data. By introducing user background data, key sample features are parsed within the user's specific business context, providing a semantic context for parsing. This allows features to move beyond abstract numerical representations and anchor to the user's specific business entity space, dynamically driving the large language model to generate natural language explanation text that aligns with the user's cognitive framework. This enables users to more intuitively understand the model's decision-making basis, thereby improving the understandability of the interpretable expression of the model's prediction results.
[0069] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, the step of generating prompt words based on the key sample features and the user background data includes:
[0070] Step A10: Obtain the feature values, global importance values, and feature sensitivity values of the key sample features;
[0071] Feature values are extracted from the selected key sample features. These feature values are the specific numerical values of each feature in the input sample. For example, in the financial field, credit score might be a key sample feature with a feature value of 750.
[0072] For global importance and feature sensitivity values, the global importance and feature sensitivity values of each sample feature can be pre-calculated and stored based on a large number of samples. This allows for the retrieval of global importance and feature sensitivity values of key sample features when needed. Global importance refers to the contribution of a feature to the output result throughout the model prediction process, and can be calculated using methods such as feature ablation or gradient-weighted activation mapping. Feature sensitivity refers to the impact of small changes in feature values on the model prediction result, and can be evaluated by calculating the change in model output before and after the feature value change, for example, using local gradient analysis.
[0073] Step A20: Concatenate the key sample features, the feature values of the key sample features, the local contribution of the key sample features, the global importance value, the feature sensitivity value, the user background data, and the predefined template guidance information to obtain prompt words. The predefined template guidance information includes one or more of the following: explanatory tone, depth of detail, and domain terminology specifications.
[0074] After obtaining the feature values, global importance values, and feature sensitivity values of key sample features, this information is concatenated with user background data and predefined template guidance information to generate a prompt word. The predefined template guidance information may include the tone of explanation, depth of detail, and / or domain terminology specifications. For example, the tone of explanation can be formal or informal, the depth of detail can be detailed or concise, and the domain terminology specifications can be financial industry-specific terminology or plain, easily understood language. In this way, a comprehensive prompt word is generated to guide the large language model in generating natural language explanation text.
[0075] This embodiment deeply integrates global importance values and feature sensitivity values into the prompt word generation chain. Based on the global importance value, it reveals the stable contribution of features to the model's overall decision-making, enabling the explanatory text to overcome the sample-specific limitations of local explanations. Simultaneously, the feature sensitivity value explicitly reveals the fragility of the decision boundary. More importantly, by coupling these two values with predefined template guidance information, natural language explanation is elevated from a simple result description to a decision diagnostic tool. It uses the global importance value to convey historical pattern summaries and the feature sensitivity value to embed future risk projections, ultimately achieving a unification of credible decision-making and proactive risk control in high-value business scenarios.
[0076] In one possible implementation, refer to Figure 3As shown, after the step of inputting the input sample into a pre-trained prediction model to obtain the prediction result, the method further includes:
[0077] Step B10: Check if a target feature vector matching the prediction result exists in the preset cache interpretation database;
[0078] After obtaining the prediction result, a search is performed in a pre-defined cached explanation database to determine if a target feature vector matching the prediction result exists. The cached explanation database is a database used to store the prediction results of historically processed input samples and their corresponding natural language explanation texts output by the large language model.
[0079] It should be noted that if the prediction result is not a vector, the prediction result can be vectorized, and the target feature vector can be searched in the preset cache interpretation database to see if it exists.
[0080] The existence of a matching target feature vector can be determined by calculating the similarity between the predicted result and the feature vectors stored in the database (e.g., using cosine similarity or Euclidean distance). If a matching target feature vector is found, it means that similar samples have been processed before and corresponding explanatory text already exists. The explanatory text corresponding to this vector can be directly retrieved without repeatedly calling the large language model. Specifically, if the similarity between the predicted result and a certain feature vector stored in the database is greater than a preset threshold, then this feature vector is determined to be a target feature vector matching the predicted result.
[0081] Step B20: If it exists, then determine that the explanation text corresponding to the target feature vector in the preset cache explanation database is the explanation text for the prediction result;
[0082] If a target feature vector matching the prediction result is found in the cached explanation database, the explanation text corresponding to that target feature vector will be directly used as the explanation text for the current prediction result. This caching mechanism can significantly improve the system's efficiency and avoid repeatedly performing complex explanation generation processes for similar samples. For example, if the input sample is highly similar to a previously processed sample in terms of feature vectors, the previously generated explanation text can be returned directly without recalculating feature contributions and calling a large language model to generate new explanation text.
[0083] Step B30: If it does not exist, then perform the step of determining the local contribution of each of the sample features to the prediction result and filtering key sample features based on the local contribution.
[0084] If no target feature vector matching the prediction result is found in the cached interpretation database, proceed to the next step, which is to determine the local contribution of each sample feature to the prediction result and select key sample features based on these contributions.
[0085] This embodiment sets up a cached explanation database and uses vector retrieval techniques (such as calculating the cosine similarity of feature vectors or nearest neighbor search) to match feature vectors that are similar to the explanation requests in the historical prediction results, thereby achieving rapid retrieval of similar explanations and improving system response efficiency and resource utilization.
[0086] Based on the first and / or second embodiments of this application, in the third embodiment of this application, the content that is the same as or similar to the first and second embodiments described above can be referred to the above description and will not be repeated hereafter. Based on this, after the step of inputting the prompt word and the prediction result into the large language model and outputting a natural language explanation text for the prediction result, the method further includes:
[0087] Step C10: Input the natural language interpretation text into the pre-trained compliance detector and output the compliance detection result;
[0088] After the large language model generates natural language explanation text, this text is input into a pre-trained compliance detector. The compliance detector can be a specially designed model to check whether the explanation text meets specific compliance requirements. These requirements may include, but are not limited to, language accuracy, information completeness, whether it contains sensitive content, or compliance with industry standards. The compliance detector analyzes the content of the explanation text and outputs a compliance detection result, which indicates whether the explanation text meets preset compliance standards.
[0089] Step C20: If the compliance detection result indicates that the natural language interpretation text is non-compliant, then the preset security instruction prompt word is input into the large language model so that the large language model can regenerate the natural language interpretation text.
[0090] If the compliance check indicates that the natural language interpretation text is non-compliant, corrective measures will be taken. Specifically, a pre-set security instruction prompt will be input into the large language model. This security instruction prompt can be a pre-set guiding text used to instruct the large language model to generate a new interpretation text that complies with compliance requirements. For example, the security instruction prompt might explicitly instruct the large language model to avoid using sensitive words, ensure the accuracy and completeness of information, etc. The large language model will regenerate the natural language interpretation text based on this new prompt, and then perform compliance checks again until the generated text complies with compliance requirements.
[0091] Step C30: If the compliance detection result indicates that the natural language interpretation text is compliant, then the interpretability prediction process ends.
[0092] If the compliance check results indicate that the natural language interpretation text is compliant, the interpretation text will be confirmed as usable, and the current interpretability prediction process will end. This means that the generated interpretation text has passed the compliance check and can be accepted and used by business personnel or users. After the process ends, the compliant interpretation text and the feature vector of the prediction result can be stored in a preset cache database for possible future queries or audits. It can also be used as cached data for subsequent similar prediction interpretation requests, further improving system efficiency.
[0093] This embodiment introduces a compliance detector to automatically identify and block explanatory texts with compliance risks, and regenerates the content using preset security warning words to ensure the compliance and security of the explanatory results.
[0094] In one possible implementation, before the step of inputting the natural language interpreted text into a pre-trained compliance detector and outputting a compliance detection result, the method further includes:
[0095] Step D10: Obtain the compliance detector to be trained and the original training samples;
[0096] A compliance detector is a model used to detect whether generated natural language interpretation text complies with specific compliance requirements, such as whether it contains sensitive words or conforms to industry standards. The original training samples refer to pre-labeled natural language text samples that are either compliant or non-compliant. These samples can be selected from historical data or generated through manual annotation, and are used to train the compliance detector to identify features that distinguish between compliant and non-compliant text.
[0097] Step D20: Perturb the original training samples to obtain adversarial training samples, wherein the perturbation includes one or more of sensitive word replacement, insertion of sensitive word templates, and semantic perturbation;
[0098] To enhance the robustness of the compliance detector, the original training samples are perturbed to generate adversarial training samples. The purpose of perturbation is to introduce subtle changes to the original samples that may cause the model to output incorrect results, thus simulating adversarial attacks. Perturbation can include one or more of the following: sensitive word replacement, insertion of sensitive word templates, and semantic perturbation. For example, sensitive word replacement replaces certain words in the original text with words that have similar semantics but may raise compliance issues; inserting sensitive word templates involves inserting predefined sensitive word templates into the text; and semantic perturbation makes minor adjustments to the semantics of the text, making it slightly different from the original text, but enough to affect the model's judgment.
[0099] Step D30: Adversarially train the compliance detector based on the adversarial training samples to obtain the trained compliance detector.
[0100] After generating adversarial training samples, these samples are used to adversarially train the compliance detector. The purpose of adversarial training is to enable the compliance detector to not only correctly classify the original samples but also correctly identify compliance issues in the adversarial samples. In this way, the compliance detector can maintain high accuracy and robustness when facing adversarial attacks. Adversarial training typically involves training with a mix of original and adversarial samples, and the model needs to learn how to make correct judgments on these samples. After training, the compliance detector will be able to more effectively detect whether natural language interpretation text complies with compliance requirements.
[0101] Based on the first, second, and / or third embodiments of this application, in the fourth embodiment of this application, the content that is the same as or similar to the above-described embodiments one, two, and three can be referred to the above description and will not be repeated hereafter. On this basis, after the step of inputting the prompt word and the prediction result into a large language model and outputting a natural language explanation text for the prediction result, the method further includes:
[0102] Step E10: In response to the unsatisfactory explanation feedback request for the natural language explanation text, obtain the feedback content based on the unsatisfactory explanation feedback request;
[0103] After the large language model generates natural language explanation text, it receives feedback from users or business personnel regarding this explanation text. If users are dissatisfied with the generated explanation text, they can initiate an "Explanation Dissatisfaction Feedback Request." The system will respond to this request and then retrieve the specific feedback content through the user interface or API (Application Programming Interface). The feedback content may include specific points of dissatisfaction from the user regarding the explanation text, such as insufficient explanation, inclusion of errors, or non-standard formatting.
[0104] Step E20: Based on the feedback content, incrementally train or fine-tune the large language model using reinforcement learning.
[0105] After obtaining feedback, this feedback will be used to further train or fine-tune the large language model. Specific methods can include incremental training or reinforcement learning fine-tuning. Incremental training involves continuing to train the model using new feedback data, updating the model's weights to better adapt to new data. Reinforcement learning fine-tuning optimizes the model's output through a reward mechanism, enabling the model to generate text that better matches user preferences. For example, Human Feedback Reinforcement Learning (RLHF) can be used to train a reward model using manually labeled preference data, and then use this reward model to guide the fine-tuning of the large language model. In this way, the model can learn which generated text is more favored by users, thus avoiding similar unsatisfactory situations in future interpretation generation.
[0106] For example, to aid in understanding the technical concept or principle of the interpretability prediction method combined with the first, second, and third embodiments described above, a specific embodiment is now provided. In this specific embodiment, reference is made to... Figure 2 As shown, the interpretability prediction process includes:
[0107] 1. After a user requests an explanation of the prediction results, a pre-trained small model (such as an ensemble tree model or a neural network) is invoked to predict the input sample and obtain the prediction results.
[0108] 2. Based on the training dataset, use global statistics or model-inherent metrics (such as feature weights of ensemble tree models) and sensitivity analysis methods (adding a dynamometer to each feature value to observe changes in the output results) to calculate the global importance value and feature sensitivity value of each sample feature. At the same time, use SHA P or similar methods to calculate the local feature contribution of the input samples.
[0109] Third, select the top N most critical key sample features and their local contributions from the SHAP values of the current samples. These features are considered to have the greatest impact on the prediction result. At the same time, collect individual background data related to the explanation request (i.e., user background data, such as user roles, historical indicators, business scenario descriptions, etc.) and predefined template guidance information configured with explanation granularity, explanation templates, etc. (such as explanation tone, depth of detail, domain terminology specifications, etc.).
[0110] Fourth, the key features and feature values, global importance values, SHAP scores, and feature sensitivity values of the selected Top N samples are combined with user background data and predefined template guidance information to form a prompt for the large language model. The large language model is then called via API to request it to generate a natural language explanation of the prediction result. This explanation includes an explanation of the role of the top-ranking key sample features and also embeds descriptions relevant to the user scenario, ensuring that the content conforms to both model logic and business context.
[0111] Furthermore, to improve efficiency, a caching and similar request retrieval mechanism was designed. Specifically, the caching unit uses a triple of "explanation request - model output - feature vector" as the key and stores the generated natural language explanation as the value. When a new explanation request arrives, the system calculates the feature vector of the prediction result output by the small model after its input sample is input, and then performs cosine similarity matching with the historical feature vectors in the cache. If the similarity exceeds a preset threshold, the corresponding explanation text already generated in the cache is directly returned, avoiding repeated calls to the large language model, thereby speeding up the response and saving resources. For requests with similarity below the threshold, the large language model is still called to generate an explanation, and the new result is stored in the cache.
[0112] After the large model generates explanatory text, a compliance detection mechanism is further introduced to prevent the generated content from containing sensitive information. To achieve this, a compliance detector based on adversarial generation is designed, and its specific working mechanism is as follows:
[0113] First, the explanatory text generated by the large language model is input into a trained compliance detector to identify whether the text violates industry standards, contains sensitive words, or has other problems.
[0114] The compliance detector is trained using an adversarial approach, which involves inputting "edge cases" that have been perturbed or constructed during the training phase, forcing the model to learn to identify potential compliance risks. This mechanism enhances the robustness and security of the interpreted text.
[0115] Once an anomaly is detected, the interception and repair process will begin, intercepting the current output and preventing it from being directly returned to the user; and based on the anomaly fragment and its context, calling the preset security instruction prompt words to re-request the large model to generate a new compliant interpretation.
[0116] Furthermore, business personnel can provide satisfaction feedback (e.g., satisfied / dissatisfied or offering suggestions for improvement) after reviewing the explanations. Explanations deemed "unsatisfactory" are marked and, along with the original request and feature data, used as training samples, are input into the large language model's online fine-tuning process. Incremental training or reinforcement learning fine-tuning is then performed using this labeled data, allowing the large language model to gradually learn and generate explanation strategies that better meet user expectations.
[0117] It should be noted that the above examples are only used to help understand this embodiment and do not constitute a limitation on the interpretability prediction process of this embodiment. Any simple modifications based on this technical concept are within the protection scope of this application.
[0118] Furthermore, embodiments of this application also propose an interpretability prediction system, referring to... Figure 4 As shown, the interpretability prediction system includes:
[0119] The response module 10 is used to respond to a prediction result interpretation request and obtain input samples and user background data according to the prediction result interpretation request, wherein the input samples include multiple sample features and their corresponding feature values;
[0120] Prediction module 20 is used to input the input sample into a pre-trained prediction mini-model to obtain the prediction result;
[0121] The filtering module 30 is used to determine the local contribution of each sample feature to the prediction result, and to filter key sample features based on the local contribution.
[0122] The explanation module 40 is used to generate prompt words based on the key sample features and the user background data, input the prompt words and the prediction results into the large language model, and output natural language explanation text for the prediction results.
[0123] Furthermore, embodiments of this application also propose an interpretability prediction device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the interpretability prediction method as described above.
[0124] refer to Figure 5The diagram illustrates a structural schematic suitable for implementing the interpretability prediction device of the embodiments of this application. The interpretability prediction device in the embodiments of this application may also include, but is not limited to, mobile terminals such as mobile phones, servers, laptops, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The interpretability prediction device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0125] like Figure 5 As shown, the interpretability prediction device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the interpretability prediction device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the interpretability prediction device to communicate wirelessly or wiredly with other devices to exchange data. While interpretability prediction devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0126] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0127] The interpretability prediction device provided in this application, employing the interpretability prediction method described in the above embodiments, can solve the technical problem that existing model interpretability methods are technically demanding and difficult to meet the requirements of understandability. Compared with the prior art, the beneficial effects of the interpretability prediction device provided in this application are the same as those of the interpretability prediction method provided in the above embodiments, and other technical features in this interpretability prediction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0128] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0130] In addition, to achieve the above objectives, embodiments of this application also provide a readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the interpretability prediction method in the above embodiments.
[0131] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0132] The aforementioned computer-readable storage medium may be included in the interpretability prediction device; or it may exist independently and not assembled into the interpretability prediction device.
[0133] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an interpretable prediction device, cause the interpretable prediction device to: respond to a prediction result interpretation request, acquire input samples and user background data according to the prediction result interpretation request, wherein the input samples include multiple sample features and their corresponding feature values; input the input samples into a pre-trained prediction mini-model to obtain a prediction result; determine the local contribution of each of the sample features to the prediction result, and select key sample features based on the local contribution; generate prompt words based on the key sample features and the user background data, input the prompt words and the prediction result into a large language model, and output a natural language explanation text for the prediction result.
[0134] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0136] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the modules themselves.
[0137] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described interpretability prediction method. This addresses the technical problem that existing model interpretability methods are technically complex and struggle to meet understandability requirements. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the interpretability prediction method provided in the above embodiments, and will not be elaborated upon here.
[0138] Furthermore, embodiments of this application also propose a computer program product, including a computer program that, when executed by a processor, implements the steps of the interpretability prediction method as described above.
[0139] The specific implementation of the computer program product in this application is basically the same as the embodiments of the interpretability prediction method described above, and will not be repeated here.
[0140] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0141] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0142] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software sensor. This computer software sensor is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause an interpretable prediction device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0143] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An interpretable prediction method, characterized in that, The interpretability prediction method includes the following steps: In response to a prediction result interpretation request, input samples and user background data are obtained based on the prediction result interpretation request, wherein the input samples include multiple sample features and their corresponding feature values; The input sample is fed into a pre-trained prediction mini-model to obtain the prediction result; Determine the local contribution of each of the sample features to the prediction result, and select key sample features based on the local contribution. Based on the key sample features and the user background data, prompt words are generated. The prompt words and the prediction results are input into a large language model, and a natural language explanation text for the prediction results is output.
2. The interpretability prediction method as described in claim 1, characterized in that, The step of generating prompt words based on the key sample features and the user background data includes: Obtain the feature values, global importance values, and feature sensitivity values of the key sample features; By concatenating the key sample features, the feature values of the key sample features, the local contribution of the key sample features, the global importance value, the feature sensitivity value, the user background data, and the predefined template guidance information, prompt words are obtained. The predefined template guidance information includes one or more of the following: explanatory tone, depth of detail, and domain terminology specifications.
3. The interpretability prediction method as described in claim 1, characterized in that, After the step of inputting the input sample into a pre-trained prediction model to obtain the prediction result, the method further includes: Check if a target feature vector matching the prediction result exists in the preset cache interpretation database; If it exists, then the explanatory text corresponding to the target feature vector in the preset cache explanation database is determined to be the explanatory text for the prediction result; If not, then perform the step of determining the local contribution of each of the sample features to the prediction result, and filtering key sample features based on the local contribution.
4. The interpretability prediction method as described in claim 1, characterized in that, After the step of inputting the prompt word and the prediction result into a large language model and outputting a natural language explanation text for the prediction result, the method further includes: The natural language interpretation text is input into a pre-trained compliance detector, and the compliance detection result is output. If the compliance detection result indicates that the natural language interpretation text is non-compliant, then the preset security instruction prompt word is input into the large language model so that the large language model can regenerate the natural language interpretation text; If the compliance test result indicates that the natural language interpretation text is compliant, then the interpretability prediction process ends.
5. The interpretability prediction method as described in claim 4, characterized in that, Before the step of inputting the natural language interpreted text into the pre-trained compliance detector and outputting the compliance detection result, the method further includes: Obtain the compliance detector to be trained and the original training samples; The original training samples are perturbed to obtain adversarial training samples, wherein the perturbation includes one or more of the following: sensitive word replacement, insertion of sensitive word templates, and semantic perturbation. The compliance detector is trained adversarially based on the adversarial training samples to obtain the trained compliance detector.
6. The interpretability prediction method as described in claim 1, characterized in that, After the step of inputting the prompt word and the prediction result into a large language model and outputting a natural language explanation text for the prediction result, the method further includes: In response to a request for unsatisfactory interpretation of the natural language interpreted text, feedback content is obtained based on the request for unsatisfactory interpretation. The large language model is incrementally trained or fine-tuned using reinforcement learning based on the feedback content.
7. The interpretability prediction method according to any one of claims 1 to 6, characterized in that, The local contribution is the SHAP value, and the step of screening key sample features based on the local contribution includes: The SHAP values are sorted in descending order to obtain a sorted queue; Select the sample features corresponding to the first preset number of SHAP values in the sorting queue as key sample features.
8. An interpretable prediction system, characterized in that, The interpretability prediction system includes: The response module is used to respond to a prediction result interpretation request and obtain input samples and user background data according to the prediction result interpretation request, wherein the input samples include multiple sample features and their corresponding feature values; The prediction module is used to input the input samples into a pre-trained prediction mini-model to obtain prediction results; A filtering module is used to determine the local contribution of each of the sample features to the prediction result, and to filter key sample features based on the local contribution. The explanation module is used to generate prompt words based on the key sample features and the user background data, input the prompt words and the prediction results into the large language model, and output natural language explanation text for the prediction results.
9. An interpretability prediction device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the interpretability prediction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an interpretability prediction program, which, when executed by a processor, implements the steps of the interpretability prediction method as described in any one of claims 1 to 7.
Citation Information
Cited By
Battery short circuit detection system and device fusing impedance spectroscopy and interpretable ensemble learning
CN121385665A