Data processing method and device, electronic equipment and readable storage medium

By using decision tree models and natural language generation techniques, decision paths and similar sample sets are extracted to generate interpretable target explanation data, which solves the problem of weak interpretation methods in machine learning models and improves the understandability and credibility of prediction results.

CN121189435APending Publication Date: 2025-12-23BEIJING JIZHI DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511107982.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing machine learning models suffer from weak interpretation methods and low correlation with internal decision-making logic, resulting in poor interpretability and low business credibility of prediction results.

Method used

The decision tree model is used to process the samples to be predicted, extract decision path data and similar sample sets, and then combined with natural language generation technology to generate interpretable target explanation data.

Benefits of technology

It improves the interpretability and credibility of prediction results, enhances the relevance and focus of explanations, reduces the difficulty of understanding explanatory texts, and increases users' trust in the model's prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121189435A_ABST
    Figure CN121189435A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a data processing method and device, electronic equipment and a readable storage medium. The method comprises the following steps: processing a to-be-predicted sample through a decision tree model to obtain a predicted value corresponding to the to-be-predicted sample; performing decision path extraction processing on the to-be-predicted sample based on the predicted value to obtain decision path data and leaf nodes corresponding to the to-be-predicted sample; performing normalization screening processing on a historical sample set based on the leaf nodes to obtain a similar sample set, the historical sample set being used for constructing a decision tree model; and performing natural language generation processing on the decision path data, the similar sample set, the predicted value and the feature description text corresponding to the to-be-predicted sample to obtain target explanation data corresponding to the predicted value, thereby enhancing the credibility of an output result, ensuring the consistency of the sample and decision logic, improving the correlation of the similar sample, avoiding information overload, and improving the accuracy of the result. And the interpretation focusing performance and effectiveness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a data processing method and device, electronic equipment and readable storage medium. BACKGROUND

[0002] Machine learning models are widely used in business prediction, but the explainability of the prediction results is limited. The explanations generated by the conventional method are too technical, and non-technical users cannot understand the association between the decision logic and the business meaning. The similar samples found are inconsistent with the current prediction samples in the decision logic of the model, which leads to the disconnection between the explanation and the actual judgment process of the model, the ambiguity of the similarity definition, and the provision of only abstract technical indicators by the decision rule display or feature importance score, resulting in insufficient trust of the prediction results of the users and restricting the landing effect of the model in the actual business.

[0003] As can be seen, the prior art has the problem of poor understandability of the prediction results and low business credibility due to weak association between the explanation method and the internal decision logic of the model, low sample relevance, and no obvious focus point of the explanation text. SUMMARY

[0004] Therefore, the embodiments of the present disclosure provide a data processing method and device, electronic equipment and readable storage medium to solve the problem of poor understandability of the prediction results and low business credibility due to weak association between the explanation method and the internal decision logic of the model, low sample relevance, and no obvious focus point of the explanation text in the prior art.

[0005] In a first aspect, the embodiments of the present disclosure provide a data processing method, comprising: processing a to-be-predicted sample by a decision tree model to obtain a prediction value corresponding to the to-be-predicted sample; performing decision path extraction processing on the to-be-predicted sample based on the prediction value to obtain decision path data and a leaf node corresponding to the to-be-predicted sample; performing normalization filtering processing on a historical sample set based on the leaf node to obtain a similar sample set, wherein the historical sample set is used to construct the decision tree model; and performing natural language generation processing on the decision path data, the similar sample set, the prediction value, and a feature description text corresponding to the to-be-predicted sample to obtain target explanation data corresponding to the prediction value.

[0006] In some embodiments, the normalization filtering processing on the historical sample set based on the leaf node to obtain the similar sample set comprises: performing similarity filtering processing on the historical sample set based on the leaf node to obtain a candidate similar sample subset; and performing similar sample refining processing on the candidate similar sample subset to obtain the similar sample set.

[0007] In some embodiments, the similar sample refining processing is performed on the candidate similar sample subset to obtain a similar sample set, including: determining a number of candidate similar samples included in the candidate similar sample subset; in a case where the number of candidate similar samples is less than or equal to a preset number threshold, determining the candidate similar sample subset as the similar sample set; in a case where the number of candidate similar samples is greater than the preset number threshold, performing a normalized neighbor screening processing on all candidate similar samples to obtain the similar sample set.

[0008] In some embodiments, the decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the to-be-predicted sample are subjected to natural language generation processing to obtain target explanation data corresponding to the predicted value, including: based on the decision path data, the similar sample set, the predicted value, and the feature description text, constructing text generation prompt engineering data; based on the text generation prompt engineering data, generating a natural language explanation text; and integrating the natural language explanation text and the predicted value to obtain the target explanation data.

[0009] In some embodiments, the decision path data and the leaf node corresponding to the to-be-predicted sample are obtained by performing decision path extraction processing on the to-be-predicted sample based on the predicted value, including: performing conditional judgment processing on the to-be-predicted sample based on the predicted value to obtain the decision path data; and performing end node identification processing on the decision path data to obtain the leaf node.

[0010] In some embodiments, the candidate similar sample subset is obtained by performing similarity screening processing on the historical sample set based on the leaf node, including: performing attribution matching processing on the historical sample set based on the leaf node to obtain candidate similar samples; and performing real label collection processing on the candidate similar samples to obtain the candidate similar sample subset.

[0011] In some embodiments, before the to-be-predicted sample is processed by the decision tree model to obtain a predicted value corresponding to the to-be-predicted sample, the method further includes: obtaining a historical sample set; performing root node initialization processing on the historical sample set to obtain an initial root node; performing recursive splitting processing on the historical sample set based on the initial root node to obtain an initial leaf node; and performing cost complexity pruning processing on the initial root node and the initial leaf node to obtain the decision tree model.

[0012] In a second aspect, the embodiment of the present disclosure provides a data processing apparatus, comprising: a first processing module configured to process a to-be-predicted sample through a decision tree model to obtain a prediction value corresponding to the to-be-predicted sample; a second processing module configured to perform decision path extraction processing on the to-be-predicted sample based on the prediction value to obtain decision path data and a leaf node corresponding to the to-be-predicted sample; a third processing module configured to perform normalization screening processing on a historical sample set based on the leaf node to obtain a similar sample set, wherein the historical sample set is used to construct the decision tree model; and a fourth processing module configured to perform natural language generation processing on the decision path data, the similar sample set, the prediction value and a feature description text corresponding to the to-be-predicted sample to obtain target explanation data corresponding to the prediction value.

[0013] In a third aspect, the embodiment of the present disclosure provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.

[0014] In a fourth aspect, the embodiment of the present disclosure provides a readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the steps of the above method.

[0015] Compared with the prior art, the embodiment of the present disclosure has the beneficial effects that: the to-be-predicted sample is processed through the decision tree model to obtain the prediction value corresponding to the to-be-predicted sample; the decision path extraction processing is performed on the to-be-predicted sample based on the prediction value to obtain the decision path data and the leaf node corresponding to the to-be-predicted sample; the normalization screening processing is performed on the historical sample set based on the leaf node to obtain the similar sample set; and the natural language generation processing is performed on the decision path data, the similar sample set, the prediction value and the feature description text corresponding to the to-be-predicted sample to obtain the target explanation data corresponding to the prediction value. In this way, the prediction value is obtained through the decision tree model, the explainability of the prediction process is ensured, and the computing efficiency is improved; the decision path extraction clearly defines the reasoning track, the normalization screening improves the relevance and credibility of the explanation, the natural language generation is performed by fusing information and features, the understanding difficulty of the explanation text is reduced, the focus and effectiveness of the explanation are improved, and information overload is avoided. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0017] Figure 1 is a scene schematic diagram of an application scenario of the embodiment of the present disclosure;

[0018] Figure 2 is a flowchart of a data processing method according to an embodiment of the present disclosure;

[0019] Figure 3 is a flowchart of another data processing method according to an embodiment of the present disclosure;

[0020] Figure 4 is a flowchart of still another data processing method according to an embodiment of the present disclosure;

[0021] Figure 5 is a structural diagram of a data processing apparatus according to an embodiment of the present disclosure;

[0022] Figure 6 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] In the following description, specific details are set forth, such as a particular system architecture, techniques, etc., in order to provide a thorough understanding of the present embodiments. However, persons skilled in the art will understand that the present disclosure can be practiced without these specific details. In other instances, well-known structures, devices, circuits, and methods have not been described in detail in order to avoid obscuring the present disclosure.

[0024] It should be noted that the user information (including but not limited to terminal device information, user personal information, etc.) and data (including but not limited to data for display, analyzed data, etc.) involved in the present disclosure are all information and data authorized by the user or authorized by all parties.

[0025] A data processing method and apparatus according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0026] Figure 1 is a scenario diagram of an application scenario of an embodiment of the present disclosure. The application scenario can include terminal devices 1, 2, and 3, a server 4, and a network 5.

[0027] The terminal devices 1, 2 and 3 can be hardware or software. When the terminal devices 1, 2 and 3 are hardware, they can be various electronic devices with display screens and supporting communication with the server 4, including but not limited to smart phones, tablet computers, laptop computers and desktop computers, etc.; when the terminal devices 1, 2 and 3 are software, they can be installed in the electronic devices as above. The terminal devices 1, 2 and 3 can be implemented as multiple software or software modules, or as a single software or software module, and the embodiments of the present disclosure do not make any limitation in this regard. Further, the terminal devices 1, 2 and 3 can be installed with various applications, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0028] The server 4 can be a server providing various services, for example, a background server receiving requests sent by terminal devices establishing communication connections therewith. The background server can receive and analyze the requests sent by the terminal devices, etc., and generate processing results. The server 4 can be a single server, a server cluster composed of several servers, or a cloud computing service center, and the embodiments of the present disclosure do not make any limitation in this regard.

[0029] It should be noted that the server 4 can be hardware or software. When the server 4 is hardware, it can be various electronic devices providing various services for the terminal devices 1, 2 and 3. When the server 4 is software, it can be multiple software or software modules providing various services for the terminal devices 1, 2 and 3, or a single software or software module providing various services for the terminal devices 1, 2 and 3, and the embodiments of the present disclosure do not make any limitation in this regard.

[0030] The network 5 can be a wired network connected by coaxial cables, twisted pairs and optical fibers, or a wireless network realizing interconnection of various communication devices without wiring, for example, Bluetooth, Near Field Communication (NFC), Infrared, etc., and the embodiments of the present disclosure do not make any limitation in this regard.

[0031] The user can establish a communication connection with the server 4 via the network 5 through the terminal devices 1, 2 and 3 to receive or send information, etc. Specifically, the server 4 can acquire a to-be-predicted sample via the terminal devices 1, 2 and 3, process the to-be-predicted sample through a decision tree model to obtain a prediction value corresponding to the to-be-predicted sample, perform decision path extraction processing on the to-be-predicted sample based on the prediction value to obtain decision path data and a leaf node corresponding to the to-be-predicted sample, perform normalization filtering processing on a historical sample set based on the leaf node to obtain a similar sample set, and perform natural language generation processing on the decision path data, the similar sample set, the prediction value and a feature description text corresponding to the to-be-predicted sample to obtain target explanation data corresponding to the prediction value.

[0032] It should be noted that the specific types, numbers and combinations of the terminal devices 1, 2 and 3, the server 4 and the network 5 can be adjusted according to actual needs of an application scenario, and the embodiments of the present disclosure do not limit them.

[0033] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present disclosure. Figure 2 The data processing method of the embodiment of the present disclosure can be executed by the server. Figure 1 As shown in the figure, the data processing method includes the following steps. Figure 2

[0034] S201, processing the to-be-predicted sample through a decision tree model to obtain a prediction value corresponding to the to-be-predicted sample.

[0035] Specifically, the to-be-predicted sample can be input into the decision tree model, starting from the root node of the decision tree model, and based on the specific feature value of the to-be-predicted sample, the decision condition set by each internal node, such as whether the feature value is less than, greater than or equal to a certain threshold, the branches are selected layer by layer to traverse the decision tree downward until the leaf node is reached; the prediction value calculated or stored in the leaf node is taken as the prediction result corresponding to the to-be-predicted sample. In this way, the new sample is automatically predicted through the decision tree model, which enhances the automation degree of prediction and improves the prediction efficiency.

[0036] ​The decision tree model can be a tree structure machine learning model composed of nodes and directed edges. The topology of the decision tree model can be composed of a root node, multiple internal nodes, and leaf nodes. The root node and the internal nodes respectively store a feature selector and a corresponding threshold. The leaf nodes store a prediction value. The decision tree model can be constructed by recursively splitting the feature space of historical samples or training data. The splitting criterion can be based on information gain, Gini index, or Mean Absolute Error (MSE), without limitation. Through the training process, the decision tree model divides the feature space into multiple disjoint subspaces, each corresponding to a leaf node, and solidifies the statistical results of historical samples or training data in the subspace into the leaf node, thereby completing the construction of the mapping function from input features to output prediction values.

[0037] The to-be-predicted sample can be an unknown output observation that has not yet output a result. The to-be-predicted sample can be a feature vector, where each component corresponds to the value of a defined feature. The feature vector can be numerical or categorical, without limitation. The values of the feature vector can be obtained from a business system or a sensor that is isomorphic to the historical sample or is collected in real time. The to-be-predicted sample can also be used to represent a single data record that is inferred by the decision tree model to obtain a prediction result.

[0038] The prediction value corresponding to the to-be-predicted sample can be an output value stored or calculated in real time by the leaf node to which the inference process of the decision tree model converges. For a regression task, the output value can be the arithmetic mean or weighted mean of the target variable of the historical samples in the leaf node. For a classification task, the output value can be the mode, probability distribution, or logistic regression coefficient of the class label of the historical samples in the leaf node, without limitation. The prediction value can also be used to represent the probability of occurrence of a future event, the numerical estimate of a continuous variable, or the attribution result of the class to which the sample belongs.

[0039] For example, in real-time risk control of commercial bank credit cards, a decision tree model can be trained based on historical transaction and user behavior data, and each leaf node stores the probability of risk for each transaction. When a new online transaction request arrives, the feature vector of the transaction, including transaction amount, merchant category, user historical consumption frequency, and / or device fingerprint, can be input into the decision tree model as a prediction sample. Starting from the root node, the decision tree model traverses down layer by layer based on internal node judgment conditions such as "transaction amount greater than 5000 yuan" and "user historical 90-day out-of-town transaction frequency less than or equal to 2", and falls into a certain leaf node. The risk probability 0.87 stored in the leaf node can be used as the prediction value for this transaction.

[0040] In S202, a decision path extraction process is performed on the prediction sample based on the prediction value, and decision path data and a leaf node corresponding to the prediction sample are obtained.

[0041] Specifically, the feature values of the prediction sample can be matched with the branch conditions in the decision tree model starting from the root node, and the node sequence and corresponding judgment rules traversed are recorded. When the sample reaches the end node without branches, the unique identifier (ID) of the leaf node is recorded. The judgment conditions of all nodes in the traversal process can be integrated into decision path data according to the order, and the form of the decision path data can be an ordered list associated with the leaf node ID. In this way, the traceability of the decision-making process of the decision tree model is enhanced, and the logical consistency constraint for subsequent similar sample screening is provided by explicitly extracting the path rule and the end position. The accuracy of business interpretation is improved, the understandability of complex decision trees is improved, abstract prediction is converted into a parseable rule chain, and the trust of business personnel in the prediction result is supported.

[0042] The decision path data can be an explicit structured representation of all judgment conditions and logical order followed by the prediction sample starting from the root node and passing through each internal node. The form of the decision path data can be an ordered list or a path vector, where each element contains a node number, a judgment feature name, an operator, and a threshold. The decision path data is associated with the leaf node ID to form a traceable rule chain. The decision path data can be generated in real time by online traversal of the decision tree topology. The decision path data can also be used to represent all intermediate logic when making predictions for the current prediction sample, which can be considered as a "model reasoning log" or "explanation evidence chain", providing formal basis for subsequent sample similarity retrieval, rule conflict detection, and business interpretation.

[0043] The leaf node corresponding to the to-be-predicted sample can refer to a terminal node reached after layer-by-layer determination according to a feature vector of the to-be-predicted sample; the leaf node has no child node in the tree structure, and the leaf node can also be used to represent the "home region" of the to-be-predicted sample, and the ID of the leaf node and the prediction value jointly constitute the final positioning of the to-be-predicted sample in the tree structure.

[0044] For example, in the price sensitivity prediction of a retail e-commerce platform, when a new user initiates a product detail page visit, the features of the new user (such as the number of visits in the last 7 days, the number of collected categories, the coupon usage rate, etc.) can be input into the decision tree model as to-be-predicted samples, and the node determination conditions such as "the number of visits in the last 7 days is greater than 15", "the number of collected categories is less than or equal to 3", and "the coupon usage rate is equal to 0.8" are used to determine the path, and the leaf node with ID L_482 is reached.

[0045] In S203, the historical sample set is normalized and filtered based on the leaf node, and a similar sample set is obtained.

[0046] Specifically, candidate similar samples falling into the leaf node to which the to-be-predicted sample belongs in the decision tree model can be filtered from the historical sample set. When the number of candidate similar samples is greater than a preset number threshold, the key features of the candidate samples and the to-be-predicted sample can be normalized to eliminate the dimensional differences between the features. In the normalized feature space, the similarity between the to-be-predicted sample and each candidate similar sample can be calculated based on a distance measurement algorithm, and the preset number of samples with the highest similarity can be selected as the output similar sample set. When the number of candidate similar samples is less than or equal to the preset number threshold, all candidate similar samples can be used as the similar sample set. In this way, the relevance of the candidate similar samples and the decision logic is enhanced, ensuring that the candidate similar samples obtained by filtering follow the same decision rules as the to-be-predicted sample. The accuracy of numerical comparison between samples is improved, and the credibility and business persuasiveness of the explanation are improved.

[0047] The historical sample set can be a collection of all labeled samples applied in the decision tree model training phase and persistently stored, where each historical sample is a multi-dimensional feature vector. The historical sample set can be obtained from a business system, a sensor, or a log through offline batch processing or streaming acquisition, and stored in a training data warehouse after feature engineering, cleaning, and labeling processing. In the decision tree model construction process, the historical sample set can obtain the statistical distribution of each node of the decision tree through recursive division and be indexed to the leaf node.

[0048] The similar sample set can be a set of historical samples highly close to the to-be-predicted sample in feature distribution, obtained from the historical sample set after preliminary selection by the "shared leaf node" rule and further refined by a distance measurement algorithm in the normalized feature space; the similar sample set can also be used to represent a historical instance that follows a consistent decision path with the to-be-predicted sample and is most adjacent in the numerical feature space.

[0049] For example, in the insurance claim risk scenario, historical claim records can be mapped to multiple leaf nodes by the trained decision tree model; when a new claim record is input, the leaf node L_237 can be located according to the feature vector of the record, and then 120 historical claim cases that also fall into L_237 can be retrieved from the historical sample set as candidate similar samples. Since the number exceeds the preset number threshold 100, the key features such as the amount, the time interval of the accident, and the hospital level can be normalized, and the similarity can be calculated by the Euclidean distance to obtain the top 80 historical claim cases in the order of similarity from high to low to constitute the similar sample set.

[0050] S204, the decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the to-be-predicted sample are subjected to natural language generation processing to obtain target explanation data corresponding to the predicted value.

[0051] Specifically, the decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the to-be-predicted sample can be structured and integrated, and input to a large language model (LLM). The LLM is guided by a pre-defined prompt template (Prompt) to integrate information in natural language, and a complete explanation text containing prediction basis, rule interpretation, and instance reference, i.e., readable explanation data, is obtained. In this way, the relevance of the explanation content and the decision logic is enhanced, and it is ensured that the generated basis originates from the actual reasoning process; the credibility and persuasiveness of the explanation are improved, and the similar sample with the same decision logic provides empirical support; the understanding efficiency of business users for complex model prediction is improved, and the natural language generation technology is used to convert technical data into intuitive business scenario description.

[0052] The feature description text corresponding to the to-be-predicted sample can be a set of unambiguous natural language segments obtained by mapping the original feature vector of the to-be-predicted sample through an interpretable feature encoder; the feature description text can be obtained by concatenating the feature name, the feature value, and the semantic label according to a fixed template, and the feature description text can also be used to provide the large language model with input-side context with semantic consistency.

[0053] The target explanation data corresponding to the prediction value can be a business scenario-oriented explainable text output by the large language model in the natural language generation stage; the target explanation data can be obtained by structurally injecting the decision path data, the similar sample set, the prediction value, and the feature description text through a pre-constructed prompt template, and performing conditional generation by the LLM; and the target explanation data can also be used to represent the causal chain explanation of the prediction result.

[0054] For example, in a real-time credit granting scenario of consumer finance, when a prediction value of "default probability 0.87" is given for an application, the decision path data "more than 3 times of overdue in the past 6 months and income-debt ratio greater than 0.6", 3 historical cases in the similar sample set that were rejected for loan due to high overdue and high debt, and the feature description text "age: 27 years old, monthly income: 8000 yuan, current debt: 50000 yuan" are input into the large language model together, and the target explanation data generated can be: "the customer is classified into a high default risk node due to 4 times of overdue in the past half year and an income-debt ratio as high as 6.25; the default rate of similar historical customers is 82%, so it is recommended to reject the current credit granting application."

[0055] According to the technical scheme provided by the embodiments of the present disclosure, the decision tree model is used to process the to-be-predicted sample to obtain the prediction value, the decision path data and the leaf node corresponding to the to-be-predicted sample are extracted, the historical sample set is filtered based on the leaf node to obtain the normalized similar sample set, and the decision path data, the similar sample set, the prediction value, and the feature description text are input into the large language model to generate the target explanation data. In this way, the prediction automation efficiency and the decision traceability are improved, and the similarity of the similar samples and the explanation credibility are enhanced.

[0056] In some embodiments, the historical sample set is normalized and filtered based on the leaf node to obtain the similar sample set, including: performing similarity filtering processing on the historical sample set based on the leaf node to obtain a candidate similar sample subset; and performing similar sample refining processing on the candidate similar sample subset to obtain the similar sample set.

[0057] Specifically, the historical sample set can be traversed based on the leaf node ID to which the to-be-predicted sample belongs in the decision tree model, and all historical samples falling into the same leaf node are filtered to obtain a candidate similar sample subset; whether the number of samples in the candidate similar sample subset is greater than a preset number threshold is determined. If it is less than or equal to the preset number threshold, the candidate similar sample subset can be taken as the similar sample set; if it is greater than the preset number threshold, refining processing can be performed to obtain the similar sample set.

[0058] The historical sample set can be a complete sample set that is applied in the decision tree model training stage, is completely labeled, and shares a feature space and business semantics with the to-be-predicted sample. The historical sample set can be stored in a training data warehouse through an Extract-Transform-Load (ETL) process from a business system, a log, or a sensor. In the decision tree model construction process, the historical sample set can form a statistical distribution of each leaf node of the decision tree through recursive division, and can be stored by reverse indexing with the leaf node ID as the key, to realize fast positioning in the online inference stage. The historical sample set can also be used to obtain a historical instance set consistent with the decision path of the to-be-predicted sample according to the shared leaf node ID in the inference stage.

[0059] The candidate similar sample subset can be a set of all historical samples falling in the same leaf node extracted from the historical sample set. The candidate similar sample subset shares the same decision rule chain with the to-be-predicted sample in the feature distribution. The candidate similar sample subset can also be used to represent a historical sample set that is consistent with the to-be-predicted sample in the internal logic of the model.

[0060] For example, in a commercial bank credit card risk control system, the decision tree model can be trained according to historical transactions and user portraits. The to-be-predicted sample can be located to a leaf node through inference. A leaf node ID inverted index query can be performed on the historical sample set to obtain 150 historical transaction records falling in the same leaf node as the candidate similar sample subset. Since the number exceeds the preset number threshold 100, the 100 or 50 similar samples can be obtained through refinement processing as the similar sample set.

[0061] According to the technical scheme provided in the embodiments of the present disclosure, the leaf node ID to which the to-be-predicted sample belongs in the decision tree model is traversed to filter the candidate similar sample subset from the historical sample set. It is determined whether the number of the candidate similar sample subset is greater than the preset number threshold. If the number is less than or equal to the preset number threshold, the candidate similar sample subset can be used as the similar sample set. If the number is greater than the preset number threshold, the similar sample set can be obtained through refinement processing. In this way, the dual accuracy of the similar sample in decision logic consistency and numerical similarity is improved, the efficiency of locating the sample with the same decision path in the large-scale historical sample is improved, and the controllability and explanation credibility of the refinement are improved.

[0062] In some embodiments, the candidate similar sample subset is subjected to a similar sample refining process to obtain a similar sample set, including: determining the number of candidate similar samples included in the candidate similar sample subset; when the number of candidate similar samples is less than or equal to a preset number threshold, determining the candidate similar sample subset as the similar sample set; when the number of candidate similar samples is greater than the preset number threshold, performing a normalized neighbor screening process on all candidate similar samples to obtain the similar sample set.

[0063] Specifically, the total number of samples included in the candidate similar sample subset can be counted, and the preset number threshold can be a predefined fixed value or a dynamically configured parameter, which is used to control the triggering condition of the refining operation; if the number of candidate samples is less than or equal to the preset number threshold, it indicates that the size of the candidate similar samples does not need to be compressed, and the candidate similar sample subset can be used as the similar sample set; if the number of candidate samples is greater than the preset number threshold, a normalized neighbor screening can be performed, the relevant feature vectors of all samples in the candidate similar sample subset are subjected to a normalization process to eliminate the dimension difference of the feature quantity, and in the normalized feature space, the distance measure between the to-be-predicted sample and each candidate similar sample is calculated, including but not limited to Euclidean distance or cosine similarity, the samples closest to the to-be-predicted sample can be sorted and selected according to the distance in ascending order, and the original features and labels of the selected samples can be used to form the similar sample set.

[0064] The number of candidate similar samples can be the total number of historical samples included in the candidate similar sample subset obtained by completing the leaf node-based screening; the number of candidate similar samples can be obtained by one aggregation query or counting operation in the inverted index structure corresponding to the leaf node, and the number of candidate similar samples can be used to represent the size of historical samples sharing the decision path with the to-be-predicted sample, and can also be used to represent whether the sample size needs to be further compressed to avoid redundancy in the subsequent calculation or interpretation link, thereby serving as a decision variable for triggering the normalized neighbor screening.

[0065] The preset number threshold can be an integer hyperparameter predefined or dynamically calculated by the configuration layer, and the preset number threshold can be used to define the maximum capacity of the candidate similar sample subset that can be directly output without refining; the preset number threshold can be obtained by offline cross-validation, business rules or online adaptive strategy, which is not limited here.

[0066] For example, in online advertisement click conversion rate prediction, 350 historical click records contained in the leaf node can be obtained as a candidate similar sample subset through inverted index statistics; since the configuration center sets the preset number threshold of this scenario to 100, the maximum and minimum normalization can be performed on the key features such as “exposure frequency in the last 7 days”, “device type” and “page dwell time” of the 350 candidate samples, and the 100 samples closest to the current user can be obtained based on the Euclidean distance screening to form a similar sample set.

[0067] According to the technical scheme provided by the embodiments of the present disclosure, by counting the total number of samples contained in the candidate similar sample subset, if the number of candidate samples is less than or equal to the preset number threshold, the candidate similar sample subset can be used as the similar sample set; if the number of candidate samples is greater than the preset number threshold, normalization neighbor screening can be performed, the distance measure between the to-be-predicted sample and each candidate similar sample can be calculated, the preset number of samples closest in distance can be sorted and screened according to the ascending order of distance, and the original features and labels of the screened samples can be used to form the similar sample set, thereby improving the automation degree of sample refinement, enhancing the flexibility of sparse control on large-scale candidate sets, and improving the numerical closeness of similar samples in the feature space.

[0068] In some embodiments, the decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the to-be-predicted sample are subjected to natural language generation processing to obtain target explanation data corresponding to the predicted value, including: based on the decision path data, the similar sample set, the predicted value, and the feature description text, constructing text generation prompt engineering data; based on the text generation prompt engineering data, generating a natural language explanation text; and integrating the natural language explanation text and the predicted value to obtain the target explanation data.

[0069] Specifically, the hierarchical judgment rules in the decision path data, the feature values and the corresponding actual results of each sample in the similar sample set, the predicted value of the to-be-predicted sample, and the feature description text describing the business meaning of each feature can be structured and integrated according to a predefined template, which can include background description, prediction result, decision rule, similar case data, and / or explicit natural language generation instructions, etc., without limitation here, to obtain prompt engineering data that can be processed by LLM. The constructed prompt engineering data can be input into LLM, and based on the instructions and data in the prompt, a readable explanation text for non-technical users, i.e., a natural language explanation text, which combines decision logic and similar case evidence, can be generated. The natural language explanation text can be associated and combined with the predicted value corresponding to the to-be-predicted sample to obtain the target explanation data.

[0070] The text generation prompt engineering data can be an input carrier with structured context and instruction information constructed to drive the LLM to generate natural language explanations that meet the needs of business scenarios. The content of the text generation prompt engineering data can be a composite prompt containing background description, prediction result, decision rule, similar case, and natural language generation instruction obtained by field-level splicing and formatting of the hierarchical decision rule in the decision path data, the feature values and actual results of each sample in the similar sample set, the predicted value of the to-be-predicted sample, and the feature description text through a pre-defined template. The text generation prompt engineering data can be encoded in the form of key-value pairs or tokenized sequences. The text generation prompt engineering data can also be used to provide complete context and task instructions to the LLM.

[0071] The natural language explanation text can be a readable paragraph for non-technical users output by the LLM through a self-recurrent generation mechanism after receiving and analyzing the text generation prompt engineering data. The generation process of the natural language explanation text follows the instruction constraints in the prompt, integrates the logical chain of the decision path, the supporting cases of the similar samples, and the business meaning of the predicted value in sequence, and obtains a coherent, unambiguous, and persuasive explanatory narrative. The natural language explanation text can exist in the form of pure natural language. The natural language explanation text can also be used to represent the complete explanation of the causal chain, case support, and business scenario of the prediction result.

[0072] For example, in the real-time credit granting scenario of online consumer finance, after giving a prediction value of “default probability 0.87” to a new applicant user, the decision path “more than 3 times of overdue in the past 6 months and income-debt ratio greater than 0.6”, three historical records of similar samples in the similar sample set that were also rejected for loans, and the feature description text “age: 27 years old, monthly income, current debt” can be integrated into the text generation prompt engineering data according to the template and input into the LLM, and then the natural language explanation text “the customer has been overdue for 4 times in the past half year and the income-debt ratio is as high as 6.25, consistent with the similar case of 82% default rate, and it is recommended to reject the credit” can be generated. Then the natural language explanation text and the prediction value 0.87 can be combined to form the target explanation data.

[0073] According to the technical scheme provided by the embodiment of the present disclosure, the hierarchical judgment rule in the decision path data, the feature value of each sample in the similar sample set and the corresponding actual result, the prediction value of the to-be-predicted sample, and the feature description text describing the meaning of each feature service are structured and integrated according to a predefined template to obtain prompt engineering data that can be processed by the LLM. Then, the constructed prompt engineering data can be input into the LLM to generate a natural language explanation text according to the instructions and data in the prompt. The natural language explanation text can be associated and combined with the prediction value corresponding to the to-be-predicted sample to obtain target explanation data. In this way, the business readability of the prediction result of the complex model is improved, the consistency between the explanation content and the real decision logic is enhanced, and the understanding efficiency and trust of the non-technical user for the model conclusion are improved.

[0074] In some embodiments, the decision path extraction processing is performed on the to-be-predicted sample based on the prediction value to obtain the decision path data and the leaf node corresponding to the to-be-predicted sample, including: performing conditional judgment processing on the to-be-predicted sample based on the prediction value to obtain the decision path data; and performing end node identification processing on the decision path data to obtain the leaf node.

[0075] Specifically, the to-be-predicted sample can be input into the trained decision tree model, and the conditional judgment is performed layer by layer from the root node according to the predefined node judgment rule, and the traversal path of the to-be-predicted sample in the tree structure is tracked at the same time. The judgment conditions corresponding to each non-leaf node passed by the to-be-predicted sample are recorded, including the feature name, comparison operator and threshold value of the judgment basis, which are not limited here, and an ordered condition sequence is formed. The condition sequence can be used to represent the decision logic chain of the to-be-predicted sample from the root node to the end node, and can be output as structured decision path data. Then, the node recorded at the end of the path can be identified, and the end node can be a leaf node in the decision tree that does not contain further judgment conditions.

[0076] For example, in the real-time transaction scenario of a credit card, the to-be-predicted transaction can be input into the trained decision tree model, and the traversal can be performed layer by layer downward from the root node according to the judgment conditions such as “transaction amount greater than 5000 yuan”, “number of remote transactions in the past 30 days greater than 3”, “device fingerprint risk score greater than 0.8”, and the like. The feature name, comparison operator and threshold value of each non-leaf node passed can be recorded as structured decision path data according to the order of appearance, and the leaf node without child nodes at the end of the path can be identified. Thus, the decision path data and the leaf node corresponding to the to-be-predicted transaction can be obtained.

[0077] According to the technical scheme provided by the embodiment of the present disclosure, by inputting the to-be-predicted sample into the trained decision tree model, starting from the root node and performing conditional judgment layer by layer according to the predefined node judgment rule, and simultaneously tracking the traversal path of the to-be-predicted sample in the tree structure, the judgment conditions corresponding to each non-leaf node passed by the to-be-predicted sample are recorded to form decision path data, and then the node recorded at the end of the path, that is, the leaf node, can be identified, thereby improving the automation degree of decision path data extraction, enhancing the traceability of the decision process, improving the integrity and accuracy of the rule chain extracted from the tree structure, and enhancing the interpretability and credibility of the prediction result.

[0078] In some embodiments, the leaf node is used for similarity screening processing on the historical sample set to obtain a candidate similar sample subset, including: performing attribution matching processing on the historical sample set based on the leaf node to obtain a candidate similar sample; and performing real label collection processing on the candidate similar sample to obtain a candidate similar sample subset.

[0079] Specifically, the ID of the leaf node can be taken as a matching basis to traverse all historical samples in the historical sample set. For each historical sample, the leaf node ID to which the historical sample finally belongs can be determined by the decision tree model. If the ID is consistent with the target leaf node ID, the historical sample can be included in the candidate similar sample set. The actual result value corresponding to the candidate similar sample can be extracted from the annotation information of the historical sample set. The actual result value can be a real observation value or a verified label of the historical sample in the business scenario. The candidate similar sample and the corresponding actual result value can be integrated into a structured data set to obtain a candidate similar sample subset. In the attribution matching processing, the node mapping logic of the decision tree model can be used to ensure that the candidate similar sample and the target sample are consistent in the model decision logic. In the real label collection processing, the association between the sample identifier and the historical sample set annotation field can be used to ensure the accuracy and traceability of the collected label.

[0080] The candidate similar sample can be obtained by the node mapping logic of the decision tree model from the historical sample set, and the candidate similar sample can be obtained by matching the leaf node ID. The candidate similar sample can also be used to represent a set of historical instances that are consistent with the to-be-predicted sample in the internal decision path of the decision tree model and have credible and real labels.

[0081] For example, in an insurance claim settlement scenario, when the claim settlement record falls into a leaf node L after being inferred by the decision tree model, L can be used as a retrieval key to perform a leaf node attribution matching on the historical claim database, and 200 historical claim records that also belong to L are extracted as candidate similar samples. Then, by associating the claim number with the "risk involved" label field in the historical database, the true risk label of each record is obtained, forming a candidate similar sample subset containing sample features and true labels.

[0082] According to the technical scheme provided by the embodiments of the present disclosure, the ID of the leaf node is used as the matching basis to traverse all historical samples in the historical sample set. For each historical sample, the leaf node ID of its final attribution can be determined by the decision tree model. The actual result value corresponding to the candidate similar sample can be extracted from the labeled information of the historical sample set. The candidate similar sample and its corresponding actual result value can be integrated into a structured data set to obtain a candidate similar sample subset. In this way, the retrieval efficiency of the same decision path sample is improved, the accuracy and traceability of the candidate sample label are enhanced, and the consistency of the selected sample in the model macro decision logic is ensured.

[0083] In some embodiments, before the predicted sample is processed by the decision tree model to obtain the predicted value corresponding to the predicted sample, the method further comprises: obtaining a historical sample set; performing root node initialization processing on the historical sample set to obtain an initial root node; performing recursive splitting processing on the historical sample set based on the initial root node to obtain an initial leaf node; and performing cost complexity pruning processing on the initial root node and the initial leaf node to obtain the decision tree model.

[0084] Specifically, a historical sample set containing feature data and its corresponding target value can be received from a data storage or interface as a training data source. The historical sample set can be assigned to the root node of the decision tree, which can contain all historical samples, i.e., the initial root node. Then, based on a pre-set splitting criterion, the splitting quality indicators of all features can be calculated at the current node from the initial root node. The optimal feature and its splitting threshold can be selected to generate a child node. The historical samples of the current node can be assigned to the child node according to the pre-set splitting rule. Each child node can be recursively iterated until the stop condition is met to obtain an unpruned decision tree containing the initial leaf node. Then, the subtree cost can be calculated. The initial root node can be traced back from the initial leaf node to evaluate whether to prune layer by layer: if the overall cost complexity of the leaf node after replacing a subtree is less than or equal to the original subtree, the pruning operation is performed to obtain the decision tree model.

[0085] The historical sample set can be all observation data sets used in the decision tree model training stage, which are complete and consistent with the business scenario, and can be obtained from the business database, logs or sensors through the ETL process, cleaned, processed for missing values, encoded and verified for consistency, and then stored in the training data warehouse.

[0086] The initial root node can be a single node created at the beginning of the decision tree construction and containing the historical sample set. The initial root node can be obtained by loading the historical sample set into the memory and encapsulating it as a node data structure. The attributes of the initial root node can include a sample index set, a feature matrix, a target vector and a node depth. The initial root node can also be used to represent the starting division unit of the feature space.

[0087] The initial leaf node can be a set of all terminal nodes located at the end of the tree structure, which no longer have child nodes and meet the preset stopping condition after the recursive splitting of the unpruned decision tree is completed. The initial leaf node can be obtained by recursively calling the splitting algorithm until the splitting quality cannot be improved or the stopping criterion is triggered. The initial leaf node can also be used to represent the final division area of the feature space by the decision tree model in the fully grown state.

[0088] For example, in the consumer financial risk control scenario, 500,000 labeled loan application records in the past two years can be pulled from the data warehouse as the historical sample set. The historical sample set can be encapsulated as the initial root node. Then, the optimal features such as "income debt ratio" and "number of overdue times in the past 6 months" can be selected based on the Gini index for recursive splitting to obtain a fully grown tree containing 800 initial leaf nodes. Then, the redundant sub-trees can be replaced with leaf nodes through cost complexity pruning to obtain a decision tree model containing 120 leaf nodes.

[0089] According to the technical scheme provided by the embodiments of the present disclosure, the historical sample set containing feature data and its corresponding target value can be received from the data storage or interface as the training data source. The historical sample set can be allocated to the root node of the decision tree. Then, the splitting quality indicators of all features can be calculated at the current node based on the preset splitting criteria. The optimal feature and its splitting threshold can be selected to generate a child node. The historical samples of the current node can be allocated to the child node according to the preset splitting rule. Each child node can be recursively iterated until the stopping condition is met to obtain an unpruned decision tree containing the initial leaf node. Then, the subtree cost can be calculated to obtain the decision tree model. In this way, the generalization ability of the decision tree model is improved, the overfitting risk is reduced, and the stability and interpretability are improved.

[0090] All the optional technical solutions described above can be combined to form optional embodiments of the present disclosure, which will not be described one by one here.

[0091] Figure 3 is a flow diagram of another data processing method provided by an embodiment of the present disclosure. As shown in Figure 3 , the data processing method comprises:

[0092] The functional architecture and data flow relationship of the present data processing method can be explained by Figure 3 . The data processing method mainly comprises the following functional modules:

[0093] The basic resource input includes the following components:

[0094] 1. To be predicted sample: new data instance that needs to be predicted and explained.

[0095] 2. Training data set: contains the features and known target values of historical samples, used for decision tree model training and similar sample searching.

[0096] 3. Trained decision tree model: a decision tree model pre-built based on the training data set (historical sample set).

[0097] 4. Feature information: text description of the meaning of each feature in the historical sample set, for reference when generating explanations later.

[0098] Model prediction module: receives the to-be-predicted sample, calculates it through the trained decision tree model, and outputs the predicted value of the to-be-predicted sample.

[0099] Decision path extraction module: receives the to-be-predicted sample and the trained decision tree model, analyzes the decision-making process of the to-be-predicted sample in the decision tree model, and extracts the specific decision path (i.e. a series of judgment conditions from the root node to the leaf node) that leads to the current predicted value and the leaf node ID to which the to-be-predicted sample belongs finally.

[0100] Similar sample searching module: receives the to-be-predicted sample, the trained decision tree model, the training data set, and the leaf node ID provided by the decision path extraction module.

[0101] The main functions of the similar sample searching module include:

[0102] 1. Same leaf node screening: in the training data set, identify all historical samples that fall into the same leaf node ID as the to-be-predicted sample, forming a candidate similar sample subset. This step ensures that the selected samples are consistent with the to-be-predicted sample in the macro decision logic of the model.

[0103] 2. Quantity judgment and K-Nearest Neighbors (KNN) refinement:

[0104] Determine whether the number of candidate similar samples exceeds a preset threshold (a preset number threshold) N (for example, N = 3).

[0105] If N is not exceeded, all candidate samples (candidate similar samples) are directly taken as the final similar sample set.

[0106] If N is exceeded, a refinement step is performed: the relevant features of the to-be-predicted sample and the candidate similar samples can be normalized to eliminate the influence of different feature scales; the KNN algorithm can be applied in the normalized feature space to find the N closest candidate samples to the to-be-predicted sample; and then the original features and target values of the N most similar samples can be obtained to constitute the final similar sample set. The refinement step further improves the numerical closeness of similar samples on the basis of ensuring logical consistency.

[0107] The module finally outputs a similar sample set containing N or less than N highly relevant historical samples.

[0108] The explanation content generation module receives the decision path (decision path data), the similar sample set, the prediction value, and the feature information. The task of the explanation content generation module can be to integrate and format structured information to build a text prompt (Prompt) containing background, prediction details, decision rules, similar cases, and clear instructions, ready to be submitted to a large language model.

[0109] The large language model (LLM) module receives the Prompt prepared by the explanation content generation module. Through natural language understanding and generation, a fluent, understandable, and natural language explanation text that combines decision logic and instance evidence is generated according to the instructions and data in the Prompt.

[0110] The explanation integration and output module receives the original prediction value and the natural language explanation generated by the LLM module. The two are integrated to obtain the final output result.

[0111] The final output: shows the user complete information containing the prediction value and its detailed natural language explanation.

[0112] According to the technical scheme provided by the embodiments of the present disclosure, through model prediction, path extraction, similar sample finding based on same leaf nodes and optional KNN refinement, LLM Prompt construction, LLM calling, and explanation output modules, a complete technical solution is formed to improve model credibility and business usability.

[0113] Figure 4 is a flowchart of another data processing method provided by the embodiments of the present disclosure. As shown in Figure 4As shown, the data processing method includes:

[0114] Training the decision tree model: A decision tree model can be trained through a training dataset containing historical data and known results. This step is preparatory and can be completed before the interpretation process is executed.

[0115] Receiving the sample to be predicted: When a new data instance needs to be predicted and interpreted, the sample to be predicted is received as input.

[0116] Executing model prediction: The sample to be predicted is input into the pre-trained decision tree model, and its predicted value is calculated and obtained.

[0117] Extracting decision path and leaf node ID: Analyze the specific decision-making process of the sample to be predicted in the decision tree model. Determine the complete decision path from the root node to the leaf node through a series of condition judgments based on feature values. Record the complete decision path and the ID of the leaf node.

[0118] Finding training samples with the same leaf node: Access the training dataset and traverse all samples to obtain historical samples that also belong to the above leaf node ID in the model. The obtained historical samples and their corresponding true target values can be collected to obtain a candidate similar sample subset.

[0119] Judging the number of samples and refining: Check the number of candidate similar samples found in the previous step.

[0120] If the number does not exceed the preset threshold N: All candidate similar samples are selected as the final similar sample set.

[0121] If the number exceeds N: Start the refining process, which can normalize the relevant features of the sample to be predicted and all candidate similar samples (e.g., features involved in the decision path or all features). Then, through the KNN algorithm, the N closest candidate similar samples to the sample to be predicted can be calculated and obtained in the normalized space. These N most similar samples (through corresponding original feature values and target values) can be determined as the final similar sample set.

[0122] Building LLMPrompt: Integrating key information to guide the large language model, the predicted value, the extracted decision path data, the final determined similar sample set (containing features and actual results), and feature information (feature meaning description) can be organized into a text Prompt according to a pre-defined structure. The Prompt needs to clearly state the background, task requirements, and all necessary data.

[0123] Calling LLM to generate explanation: sending the constructed Prompt to a large language model service; the LLM can parse the content in the Prompt and generate a natural language explanation; the natural language explanation can represent the decision path information of the predicted value, and can be supported by actual cases in the similar sample set to justify the rationality of the prediction.

[0124] Integrating output: combining the predicted value obtained by the decision tree model and the natural language explanation generated by the LLM to obtain the target output result and present it to the user (such as a business personnel).

[0125] According to the technical scheme provided by the embodiments of the present disclosure, the decision logic (specific path) inside the decision tree model and the historical data instances (similar samples of the same leaf node) closely related to the logic are integrated to provide a more comprehensive and more convincing explanation basis for machine learning prediction. The method of using "sample falling into the same leaf node of the decision tree" as a strong constraint condition for screening similar training samples is proposed to ensure the consistency of the selected samples with the sample to be explained in terms of macro decision logic and improve the relevance of similar cases. When the number of samples in the same leaf node is too large, the step of first normalizing the features and then applying the KNN algorithm to find the most similar N samples is introduced to improve the closeness of similar samples in specific feature values while ensuring logical consistency, avoid information overload, and improve the focus and effectiveness of the explanation. The extracted decision path, the screened similar sample set, and other context information (such as feature meaning) are structuredly input into the LLM, and through the natural language processing capability of the LLM, a prediction explanation for non-technical business personnel is automatically generated, which is smooth, easy to understand, and combines logic and instances.

[0126] The following is an apparatus embodiment of the present disclosure, which can be used to execute the method embodiments of the present disclosure. For details not disclosed in the apparatus embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.

[0127] Figure 5 is a schematic diagram of a data processing apparatus provided by an embodiment of the present disclosure. As shown in Figure 5 , the data processing apparatus comprises:

[0128] The first processing module 501 is configured to process the to-be-predicted sample by using the decision tree model to obtain a predicted value corresponding to the to-be-predicted sample.

[0129] The second processing module 502 is configured to perform decision path extraction processing on the to-be-predicted sample based on the predicted value to obtain decision path data and a leaf node corresponding to the to-be-predicted sample.

[0130] The third processing module 503 is configured to perform normalization screening processing on the historical sample set based on the leaf node, to obtain a similar sample set, wherein the historical sample set is used to construct the decision tree model.

[0131] The fourth processing module 504 is configured to perform natural language generation processing on the decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the to-be-predicted sample, to obtain target explanation data corresponding to the predicted value.

[0132] According to the technical scheme provided by the embodiments of the present disclosure, the decision tree model is used to process the to-be-predicted sample to obtain the predicted value, the decision path data and the leaf node corresponding to the to-be-predicted sample can be extracted, the historical sample set is screened based on the leaf node to obtain the normalized similar sample set, and the decision path data, the similar sample set, the predicted value, and the feature description text are input into the large language model to generate the target explanation data. In this way, the prediction automation efficiency and the decision traceability are improved, and the similarity of the similar samples and the explanation credibility are enhanced.

[0133] In some embodiments, the third processing module 503 is specifically configured to perform similarity screening processing on the historical sample set based on the leaf node, to obtain a candidate similar sample subset; and perform similar sample refining processing on the candidate similar sample subset, to obtain the similar sample set.

[0134] In some embodiments, the similar sample refining processing on the candidate similar sample subset to obtain the similar sample set is specifically configured to determine the number of candidate similar samples included in the candidate similar sample subset; when the number of candidate similar samples is less than or equal to a preset number threshold, the candidate similar sample subset is determined as the similar sample set; and when the number of candidate similar samples is greater than the preset number threshold, normalized nearest neighbor screening processing is performed on all candidate similar samples, to obtain the similar sample set.

[0135] In some embodiments, the fourth processing module 504 is specifically configured to construct text generation prompt engineering data based on the decision path data, the similar sample set, the predicted value, and the feature description text; generate a natural language explanation text based on the text generation prompt engineering data; and integrate the natural language explanation text and the predicted value, to obtain the target explanation data.

[0136] In some embodiments, the second processing module 502 is specifically configured to perform conditional judgment processing on the to-be-predicted sample based on the predicted value, to obtain the decision path data; and perform end node identification processing on the decision path data, to obtain the leaf node.

[0137] In some embodiments, the process of performing similarity screening on the historical sample set based on leaf nodes to obtain a subset of candidate similar samples is specifically used as follows: performing attribution matching on the historical sample set based on leaf nodes to obtain candidate similar samples; and performing real label collection on the candidate similar samples to obtain a subset of candidate similar samples.

[0138] In some embodiments, the data processing apparatus is further configured to: acquire a historical sample set; perform root node initialization processing on the historical sample set to obtain an initial root node; perform recursive splitting processing on the historical sample set based on the initial root node to obtain an initial leaf node; and perform cost complexity pruning processing on the initial root node and the initial leaf node to obtain a decision tree model.

[0139] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure.

[0140] Figure 6 This is a schematic diagram of the electronic device 6 provided in an embodiment of this disclosure. Figure 6 As shown, the electronic device 6 of this embodiment includes a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, it implements the steps in the various method embodiments described above. Alternatively, when the processor 601 executes the computer program 603, it implements the functions of each module / unit in the various device embodiments described above.

[0141] Electronic device 6 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 6 may include, but is not limited to, processor 601 and memory 602. Those skilled in the art will understand that... Figure 6 This is merely an example of electronic device 6 and does not constitute a limitation on electronic device 6. It may include more or fewer components than shown, or different components.

[0142] The processor 601 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0143] The memory 602 can be an internal storage unit of the electronic device 6, for example, a hard disk or a memory of the electronic device 6. The memory 602 can also be an external storage device of the electronic device 6, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 6. The memory 602 can also include both the internal storage unit and the external storage device of the electronic device 6. The memory 602 is used to store computer programs and other programs and data required by the electronic device.

[0144] It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0145] The integrated module / unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a readable storage medium (such as a computer readable storage medium). Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each method embodiment described above can be implemented. The computer program can include computer program code, which can be in the form of source code, object code, executable file or some intermediate form, etc. The computer readable storage medium can include any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier wave signal, telecommunication signal and software distribution medium, etc.

[0146] The above examples are only used to illustrate the technical solutions of the present disclosure, rather than limit the same; although the present disclosure is described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that the technical solutions recorded in the foregoing examples can be modified, or some technical features thereof can be replaced by equivalents; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the protection scope of the present disclosure.

Claims

1. A data processing method, characterized in that, include: The predicted value of the sample to be predicted is obtained by processing the sample using a decision tree model. Based on the predicted value, the decision path extraction process is performed on the sample to be predicted to obtain decision path data and the leaf node corresponding to the sample to be predicted. Based on the leaf nodes, the historical sample set is normalized and filtered to obtain a similar sample set, wherein the historical sample set is used to construct the decision tree model; The decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the sample to be predicted are processed by natural language generation to obtain the target explanation data corresponding to the predicted value.

2. The data processing method according to claim 1, characterized in that, The process of normalizing and filtering the historical sample set based on the leaf nodes to obtain a similar sample set includes: Based on the leaf nodes, a similarity screening process is performed on the historical sample set to obtain a subset of candidate similar samples; The candidate similar sample subset is refined to obtain the similar sample set.

3. The data processing method according to claim 2, characterized in that, The process of refining the candidate similar sample subset to obtain the similar sample set includes: Determine the number of candidate similar samples included in the subset of candidate similar samples; When the number of candidate similar samples is less than or equal to a preset threshold, the subset of candidate similar samples is determined as the similar sample set; When the number of candidate similar samples is greater than the preset threshold, normalized nearest neighbor filtering is performed on all candidate similar samples to obtain the similar sample set.

4. The data processing method according to claim 1, characterized in that, The step of performing natural language generation processing on the decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the sample to be predicted to obtain the target explanation data corresponding to the predicted value includes: Based on the decision path data, the similar sample set, the predicted value, and the feature description text, construct text generation prompt engineering data; Based on the aforementioned text generation prompt engineering data, a natural language explanation text is generated; The natural language explanation text and the predicted value are integrated to obtain the target explanation data.

5. The data processing method according to claim 1, characterized in that, The step of extracting decision paths from the sample to be predicted based on the predicted values ​​to obtain decision path data and the leaf nodes corresponding to the sample to be predicted includes: Based on the predicted value, the sample to be predicted is subjected to conditional judgment processing to obtain the decision path data; The decision path data is processed to identify the endpoint node, thus obtaining the leaf node.

6. The data processing method according to claim 2, characterized in that, The similarity screening process based on the leaf nodes of the historical sample set to obtain a subset of candidate similar samples includes: Based on the leaf nodes, the historical sample set is subjected to attribution matching to obtain candidate similar samples; The candidate similar samples are processed by collecting real labels to obtain a subset of the candidate similar samples.

7. The data processing method according to claim 1, characterized in that, Before processing the sample to be predicted using the decision tree model to obtain the predicted value corresponding to the sample to be predicted, the method further includes: Obtain historical sample sets; The historical sample set is initialized with root nodes to obtain initial root nodes; Based on the initial root node, the historical sample set is recursively split to obtain the initial leaf node; The initial root node and the initial leaf node are pruned to reduce their cost complexity, thus obtaining the decision tree model.

8. A data processing apparatus, characterized in that, include: The first processing module is used to process the sample to be predicted through a decision tree model to obtain the predicted value corresponding to the sample to be predicted. The second processing module is used to perform decision path extraction processing on the sample to be predicted based on the predicted value, so as to obtain decision path data and the leaf node corresponding to the sample to be predicted. The third processing module is used to perform normalization filtering on the historical sample set based on the leaf nodes to obtain a similar sample set, wherein the historical sample set is used to construct the decision tree model. The fourth processing module is used to perform natural language generation processing on the decision path data, the similar sample set, the predicted value, and the feature description text corresponding to the sample to be predicted, to obtain the target explanation data corresponding to the predicted value.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.