Visualization System for Interpretability-Aware Feature Discretization in Logistic Regression
By providing a visualization system for discretization of interpretability-aware features for logistic regression, the problem that logistic regression models are difficult to capture nonlinear patterns when processing complex data sets is solved, improving the prediction performance and interpretability of the model, and reducing the technical threshold.
Patent Information
- Application Number
- CN202510258873.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-06
AI Technical Summary
Logistic regression models are difficult to capture nonlinear patterns when processing complex data sets, and they face relative data sparsity and long-tail distribution problems in the process of transparent feature discretization, resulting in unstable positive rate estimation, affecting the accurate judgment of nonlinear trends.
It provides a visualization system for interpretable perceptual feature discretization for logistic regression, including model evolution module, feature evaluation module, nonlinear mode evaluation module and control module. Through the visual view of these modules, users can independently explore and adjust discretization schemes, identify nonlinear modes and optimize model performance.
It lowers the technical threshold for data utilization, significantly reduces the workload of data scientists in feature selection, model construction and management, improves the prediction performance and interpretability of logistic regression models, and avoids the risks brought by relative data sparseness and long-tail distribution.
Smart Images

Figure CN119760667B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and visualization, and particularly relates to a visualization system for interpretable perception feature discretization for logistic regression. Background Art
[0002] In recent years, with the continuous progress of artificial intelligence and machine learning technologies, various black-box models have been successively developed, showing strong generalization performance in practice and being able to achieve high-precision prediction results on complex data sets. However, despite their many technical advantages, due to the lack of convincing interpretability, these black-box models have sparked widespread controversy and attention in some highly sensitive industries with extremely high requirements for transparency, and the decision-making processes and bases of the models often need to be clearly explained and understood.
[0003] Logistic Regression (LR) still occupies a central position in the industry due to its fast calculation speed, ability to output probabilities, and high level of interpretability. The LR model predicts the probability of the target variable by constructing a linear relationship, and its coefficients directly reflect the influence direction and degree of the independent variable on the dependent variable. This intuitive interpretation method gives the LR model unique advantages in decision support and risk management. However, as a linear model, LR also has certain limitations. It cannot effectively capture the non-linear patterns in the data, which limits its prediction ability and application scope when dealing with complex data sets. To overcome this limitation, researchers have begun to explore improving the performance of the LR model through Transparent Feature Engineering (TFE).
[0004] TFE aims to transform the original data into a form more suitable for processing by the LR model through a series of technical means while maintaining the interpretability of the model. Among them, Transparent Feature Discretization (TFD), as one of the core links of TFE, plays a crucial role in enabling the LR model to capture non-linear features without sacrificing its interpretability. The main tasks of the TFD process include selecting appropriate features, evenly binning the features, calculating the positive example rate within each bin, estimating non-linearity, discretizing the features, training LR instances, and comparing and managing these instances. Through this process, the continuous features in the original data can be transformed into discrete categorical variables, thereby simplifying the model complexity, improving the calculation efficiency and interpretability. At the same time, by reasonably dividing the bins and calculating the positive example rate, the non-linear relationship between the features and the target variable can be captured, thereby enhancing the prediction ability of the LR model.
[0005] However, in actual operation, the TFD process also faces many challenges. Among them, the problems of relative data sparsity (RDS) and long-tailed distribution (LTD) are the most prominent. RDS refers to the instability of the estimated positive rate due to insufficient sample size in each partition during the binning process, which in turn affects the accurate judgment of the non-linear trend. LTD refers to the imbalance of the data distribution, that is, a few partitions occupy a large number of samples, while most partitions have only a small number of samples, which also interferes with the calculation of the positive rate and the identification of non-linear relationships. The non-intrinsic disorder of the positive rate caused by these problems not only increases the difficulty of judgment for data scientists, but also seriously affects the effectiveness and efficiency of the TFD cycle. Currently, there is no visualization system that can effectively assist the TFD process, and further development and application are needed. Summary of the Invention
[0006] In view of the above, the object of the present invention is to provide a visualization system for interpretable feature discretization in logistic regression, focusing on interpretable feature discretization, reducing the technical threshold of data utilization while ensuring the effectiveness of data analysis, supporting users to conduct independent exploration based on the visualization view of the model evolution module, realizing continuous comparison and management of multiple logistic regression model instances during the construction process of the logistic regression model instance, intuitively displaying information based on the visualization views of the feature evaluation module and the non-linear pattern evaluation module, facilitating users to capture feature non-linear patterns, and being able to explore multiple de-discretization regions and use them to adjust the discretization scheme, thereby improving the prediction performance and interpretability of the logistic regression model.
[0007] To achieve the above object of the invention, the technical solutions provided by the present invention are as follows:
[0008] A visualization system for interpretable feature discretization in logistic regression provided by an embodiment of the present invention includes:
[0009] A model evolution module, which constructs a corresponding logistic regression model instance as a basic model according to the features of the samples input by the user, modifies the basic model or its generated parent model to generate a new logistic regression model instance to establish the evolution and optimization process of the logistic regression model instance and visualizes it as a tree diagram, uses the logistic regression model instance as a node, and constructs a composite diagram for each node in the tree diagram to evaluate the performance of each logistic regression model instance to compare the performance changes of adjacent instances;
[0010] A control module, based on a certain logistic regression model instance constructed in the model evolution module, displays and adjusts the relevant parameters of the current instance and the feature overview of the current input sample in the control panel, and supports users to adjust parameters and select features;
[0011] The feature evaluation module, based on a certain logistic regression model instance constructed in the model evolution module, partitions the sample data according to the discretization scheme and input features set by the control panel, calculates the sample size and positive rate information within all feature partitions, and visualizes them as a line chart area and a histogram area for capturing non-linear feature patterns.
[0012] The non-linear pattern evaluation module calculates the non-linear region based on the positive rate information through a clustering algorithm and visualizes it as a scatter plot. Combined with the feature evaluation module, it provides a new perspective on the non-linear feature pattern for evaluating and adjusting the discretization scheme.
[0013] Preferably, in the model evolution module, the tree diagram of the evolution and optimization process of the logistic regression model instance is presented as a binary tree diagram. It has a total of one root node. Each parent node supports generating one or two child nodes after it. Each node of the binary tree diagram is visualized as a composite diagram representing the logistic regression model instance. When a new logistic regression model instance is constructed, the tree diagram will grow with the corresponding composite diagram node.
[0014] Among them, the composite diagram of the root node includes an interactive text element, an outer ring, and a radar chart. The composite diagram of the child node includes an interactive text element, an outer ring, and two radar charts. The interactive text element includes the naming of the logistic regression example. Click on the interactive text element to select the logistic regression model instance represented by the root node or child node, and support zooming in on the corresponding composite diagram separately. Each segment in the outer ring represents a corresponding input feature. The size of each segment represents the relative contribution of the feature to the logistic regression model instance, and the color depth of each segment represents the absolute intensity of the feature effect. The performance of the logistic regression model instance of the root node is shown through a radar chart in the composite diagram of the root node. The performances of the logistic regression model instances of the child node and its parent node are shown respectively through the two radar charts in the composite diagram of the child node.
[0015] Preferably, the performance of the logistic regression model instance includes: accuracy, precision, recall rate, F1-score, and normalized entropy.
[0016] Preferably, during the generation of the tree diagram, a curve is drawn from the right edge of the composite diagram corresponding to the selected node, indicating that the composite diagram of the next logistic regression model instance will be drawn at the end of this curve. When the composite diagram corresponding to the next logistic regression model instance will exceed the preset view range or overlap, a horizontal line will be drawn automatically instead of a curve to avoid possible view misalignment. If the composite diagram of the next logistic regression model instance fails to be drawn as expected, the corresponding curve or line will be removed.
[0017] Preferably, in the feature evaluation module, the line chart area includes a graph composed of two curves of different colors and a graph composed of two types of color blocks, gray and colored, superimposed. The Y-axis represents the magnitude of the positive rate, and the X-axis represents the number of the feature partitions. The two curves of different colors respectively represent the original positive rate data and the positive rate data processed by the SG filter. The gray-colored blocks are used as the background to indicate that the sample size in this feature partition is insufficient and the confidence level of the positive rate expression is low. The colored blocks are drawn based on the accuracy filtering and their clustering labels. Their width represents the range of the corresponding area, and their height represents the range of the positive rate in the feature partition. Different colors represent different clustering labels.
[0018] Preferably, inside each color block, the area is divided according to the maximum value, upper quartile, median, lower quartile, and minimum value of the ordinate. The greater the height of the area between the upper quartile and the lower quartile indicates the greater the interquartile range of this area, which means the samples in this area are more dispersed and this area is linear. On the contrary, it means the samples in this area are more concentrated and this area is non-linear.
[0019] Preferably, in the feature evaluation module, the histogram area includes a histogram, a composite diagram above the histogram, and an interactive background color block. The Y-axis of the histogram represents the sample size, and the X-axis represents the number of the feature partitions. The height of the histogram represents the sample size in this feature partition, and the color of the histogram represents the confidence level of the positive rate in this feature partition. The histogram is encoded with different colors according to whether the sample size meets the standard. The composite diagram above the histogram index line consists of a straight line and a point. The length of the straight line represents the magnitude of the positive rate fluctuation value, that is, the first-order difference of the positive rates in each filtered interval, and different colors and directions are used to represent whether the fluctuation value is negative or positive. The interactive background color block alternates between gray and transparent. By touching the mouse, a dotted line around the color block is displayed. The width of the color block represents the currently discretized area, and interaction is supported between the color blocks. The user can directly drag the color block to control its width and update the discretization scheme created by the decision tree.
[0020] Preferably, in the non-linear mode evaluation module, the scatter plot shares the Y-axis with the line chart area in the feature evaluation module. The X-axis of the scatter plot shows the positive rate fluctuation value. By taking the positive rate fluctuation value and the SG filtered value of the positive rate in each feature partition as two-dimensional features and inputting them into the k-means clustering model, the clustered point clusters are obtained. The point clusters are presented in the form of a scatter plot and used to observe multiple non-linear areas of the data.
[0021] Preferably, in the non-linear mode evaluation module, there is also a control area that supports the user to modify the number of clusters in the clustering algorithm through the control area, supports clicking the recoloring button to update the number of colors of the color blocks, and supports clicking the save button to save the current discretization scheme and upload it to the control module.
[0022] Preferably, in the control module, it is supported to display and adjust the relevant parameters of the selected logistic regression model instance in the tree diagram of the current model evolution module through the control panel. The relevant parameters include interval, window size, rank, tree depth adjustment, discretization, and prediction. Among them, the interval represents the duration in which continuous data should be segmented into pseudo-stationary periods, and the size of the period is called the interval. The window size represents the sliding window size applicable to the SG filter. The rank represents the rank of the SG filter polynomial. The tree depth adjustment represents adjusting the depth of the tree, which is applicable to the training of the CART algorithm in discretization. Clicking on discretization triggers the discretization command, and clicking on prediction starts the system operation. It is supported to display the feature overview for different features, support selecting the range of the current feature and clicking the input button to input the feature, and support clicking the button to obtain detailed information to get the detailed information of the feature.
[0023] Compared with the prior art, the beneficial effects of the present invention at least include:
[0024] Through the intuitive, efficient, and intelligent visual system interface function design of the present invention, the workload of data scientists in the cycle of selecting features, constructing, comparing, and managing logistic regression model instances is significantly reduced. And through the non-linear identification function design in visualization, it allows data scientists to identify non-linearity based on the SG filter and k-means algorithm, supports interactive adjustment of the discretization scheme, and avoids the risks brought by relative data sparsity and long-tail distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 It is a schematic diagram of the main interface of the visualization system for interpretable-aware feature discretization for logistic regression provided by the embodiment of the present invention;
[0027] Figure 2 It is a schematic diagram of the control panel provided by the embodiment of the present invention;
[0028] Figure 3 It is a schematic diagram of the model evolution view provided by the embodiment of the present invention;
[0029] Figure 4 It is a schematic diagram of the composite diagram representing the logistic regression model instance in the model evolution view provided by the embodiment of the present invention;
[0030] Figure 5 It is a schematic diagram of the feature evaluation view provided by an embodiment of the present invention;
[0031] Figure 6 It is a schematic diagram for the detailed description of the feature evaluation view provided by an embodiment of the present invention;
[0032] Figure 7 It is a schematic diagram of the non - linear mode evaluation view provided by an embodiment of the present invention. Detailed implementation manners
[0033] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific implementation manners described herein are only used to explain the present invention and do not limit the protection scope of the present invention.
[0034] The inventive concept of the present invention is as follows: Aiming at the problem that there is a lack of an efficient, intuitive and fully functional visualization system in the prior art to assist the transparent feature discretization (TFD) process, an embodiment of the present invention provides a visualization system for interpretable - aware feature discretization for logistic regression, which is composed of four main components: A model evolution module is used to manage the instance lineage and represent and compare instance performance and details; A feature evaluation module is used to visually display non - linearity based on the SG filter; A non - linear mode evaluation module is used to explore multiple de - discretization regions based on the k - means algorithm; A control module allows users to identify and input features, adjust controllable parameters and construct instances. These components work together around the core link of re - defining the discretization region in the TFD process. Since the re - definition of the discretization region is often accompanied by complex subsequent changes and many unknown problems, this system liberates data scientists from heavy work through visualization means and assists them to more effectively carry out the TFD process in an intuitive way.
[0035] Figure 1 It is a schematic diagram of the main interface of the visualization system for interpretable - aware feature discretization for logistic regression provided by an embodiment of the present invention. As Figure 1 shown, the embodiment provides a visualization system for interpretable - aware feature discretization for logistic regression, including: A model evolution module (its visualization view is as shown in A in Figure 1 , that is, the model evolution view), a feature evaluation module (its visualization view is as shown in B in Figure 1 , that is, the feature evaluation view), a non - linear mode evaluation module (its visualization view is as shown in C in Figure 1 , that is, the non - linear mode evaluation view) and a control module (its visualization view is as shown in D in Figure 1 , that is, the control panel).
[0036] 1. Control panel.
[0037] In the control module, an integrated interface for input and control is provided to the user through a visual control panel, specifically as Figure 2 shown. The control module will implicitly execute a series of complex code engineering, including positive rate calculation and filtering, tree-based feature discretization, data cleaning, one-hot encoding, model training and testing. Among them, as Figure 2 shown by D1 in
[0038] 2. Model evolution view.
[0039] In the model evolution view of the model evolution module, the evolution and optimization process of the logistic regression model instance is constructed in the form of a tree diagram, supporting users for interpretable feature selection, and used to manage and compare multiple logistic regression model instances constructed by different feature input schemes, specifically as Figure 3 shown. As Figure 3 shown by A1 in
[0040] In the tree diagram of the model evolution view, the evolution of logistic regression is visualized as a binary tree diagram, with a total of one root node (representing the base model), and each parent node supports generating one or two child nodes after it. Each node of the binary tree diagram is visualized as a composite diagram representing a logistic regression model instance. When constructing a new logistic regression model instance, the tree diagram will grow with the corresponding composite diagram node.
[0041] Among them, the composite diagram of the root node includes an interactive text element, an outer ring, and a radar chart, and the composite diagram of the child node includes an interactive text element, an outer ring, and two radar charts.The key to determining the node growth position is the interactive text element. The interactive text element includes the naming of the logistic regression example. When the interactive text element is clicked, the font will be bolded, indicating that the node has been selected. The corresponding composite graph of the node will be immediately copied, enlarged, and placed in the enlarged block diagram (as shown by A2 in Figure 3 ). At the same time, a curve is drawn from the right edge of the composite graph corresponding to the selected node, indicating that another composite graph of the next logistic regression model instance will be drawn at the end of the curve. When the composite graph corresponding to the next logistic regression model instance will exceed the preset view range or overlap will occur, a horizontal line will be automatically drawn instead of a curve to avoid possible view misalignment. When the bolded interactive text element is clicked again, the system will confirm whether the expected composite graph has been drawn. If so, the existence of the child node will be recorded in the binary tree data object, and its font will be restored to the original style. If not, the corresponding curve or line will be removed accordingly. It is worth mentioning that when the node is selected again, the curve will be drawn in the opposite direction. For a node in the binary tree diagram, if its left child node is not generated as expected, it will expect a right child node to be generated next time. The left child node is drawn downwards, and the right child node is drawn upwards, allowing the user to decide whether to draw above or below. Such an arrangement allows the user to adjust the structure of the graph according to the situation, making the overall layout more compact.
[0042] The outer ring is the core component of the composite graph. Each arc in the outer ring is a visual encoding of the corresponding input feature. The size of each arc represents the relative contribution of the feature to the logistic regression model instance. The positive correlation between the feature and the classification result is encoded in blue, and the negative correlation is represented in red, reflecting its importance. The color depth of each arc represents the absolute intensity of the feature effect. That is, if a certain corresponding feature is blue, the larger the feature value, the greater the possibility that the classification result is positive. Red means the opposite. As shown in Figure 4 , dark red indicates a newly added feature, dark blue indicates a relatively important feature, light blue indicates a feature with a positive contribution, light red indicates a relatively unimportant feature, and the lightest part of red indicates a feature with a negative contribution. In addition, when visualizing any LR instance that is not the parent node, the features added to the input scheme of the parent node will be highlighted in a protruding form. Through the design of this component, data scientists can intuitively observe the relative importance and effect intensity structure of features in each logistic regression model instance. In addition, data scientists can very naturally select those features that have a significant impact on the classification result for further analysis.
[0043] The radar chart encodes the performance of the logistic regression model instance, as shown in Figure 4As shown in the figure, it includes: accuracy, precision, recall rate, F1-score, and normalized entropy. Each instance of the logistic regression model that is not a parent node will carry the performance data of its parent node during visualization, and they will be plotted together as a radar chart. Therefore, in the composite chart of the root node, a radar chart is used to show the performance of the logistic regression model instance of the root node. In the composite chart of the child node, two radar charts are used to show the performance of the logistic regression model instances of the child node and its parent node respectively. The radar chart of the current node is orange, and the radar chart of the parent node is gray. To highlight the performance differences between instances, we use the performance of the instance corresponding to the root node as a benchmark.
[0044] 3. Feature evaluation view.
[0045] The feature evaluation view of the feature evaluation module is used to capture the non-linear patterns of features, helping data scientists solve the problem of difficultly confirming non-linear patterns due to the fluctuating line charts caused by RDS and LTD. The SG filter is one of the best methods for dealing with unnecessary positive rate fluctuations. On the other hand, when there are multiple non-linear regions in the data, some visualization algorithms and techniques are needed to mine them. The feature evaluation view is an interactive and continuously updated view composed of a line chart area, a histogram area, and color blocks in the shared x-coordinate area. The feature evaluation view executes the content of the feature evaluation controller, and the feature evaluation is based on the SG filter and the decision tree discretization method.
[0046] As Figure 5 shown, in the line chart area of the feature evaluation view, it includes a graph composed of two curves of different colors and a graph composed of color blocks of two types, gray and colored, superimposed, aiming to identify the positive rate pattern, where the Y-axis represents the size of the positive rate data, and the X-axis represents the number of feature partitions. As Figure 6 shown in (a) and (b) in it, the blue curve represents the original positive rate data, and the orange curve represents the positive rate data processed by the SG filter. The orange curve is used to eliminate data noise and show the essential non-linear pattern of the data. Users can adjust the parameters of the SG filter to make the orange line closer to or farther from the blue line. As Figure 6 shown in (b2) in it, the gray type of color block is used as the background to indicate that the sample size in this area is insufficient and the confidence level of the positive rate expression is low. As Figure 6 shown in (b1) in it, the colored type of color block is drawn based on the correct rate filtering and its clustering labels. Its width represents the range of the corresponding area, and its height represents the range of the positive rate within the partition. Different colors represent different clustering labels. Each color block is divided into regions according to the maximum value, upper quartile, median, lower quartile, and minimum value of the ordinate, as Figure 6 shown in (a1), (a2), and (a3) in it. As Figure 6As shown in (a) therein, the greater the height of the region between the upper quartile and the lower quartile indicates the greater the interval between the upper and lower quartiles of the region, indicating that the samples in this region are more dispersed and the region is linear. On the contrary, it indicates that the samples in this region are more concentrated and the region is non-linear. Combining the above design, this region allows users to quickly visually identify non-linear patterns.
[0047] In the histogram region of the feature evaluation view, including the histogram ( Figure 6 the lower parts of (c) and (d) therein), the composite diagram above the histogram index line ( Figure 6 the upper parts of (c) and (d) therein) and the interactive background color blocks ( Figure 6 the (c1) and (d1) therein), express the sample size and the fluctuation trend of each partition. The Y-axis of the histogram represents the sample size, the X-axis represents the number of feature partitions, the height of the histogram represents the sample size in this area, and the color of the histogram is encoded into two different colors according to whether the sample size meets the standard. Blue indicates that the sample size meets the standard, while red indicates that the sample size does not meet the standard. Therefore, the color of the histogram represents the confidence level of the positive rate in the partition. The composite diagram above the histogram index line consists of a straight line and a point. The length of the straight line represents the magnitude of the positive rate fluctuation value, that is, the first-order difference of the positive rates in each interval after filtering, and different colors and directions are used to represent whether the fluctuation value is negative or positive. Red represents negative values, and blue represents positive values. The interactive background color blocks alternate between gray and transparent. By touching the mouse, a dotted line around the color block is displayed. The width of the color block represents the currently discretized area, and interaction is supported between the color blocks. Users can directly drag the color block to control its width and update the discretization scheme created by the decision tree. Combined with the histogram, it is possible to intuitively understand the impact of the sample size on the positive rate fluctuation.
[0048] 4. Non-linear pattern evaluation view.
[0049] Based on the feature evaluation view, users can perceive multiple non-linear regions, which are difficult to directly evaluate with a line chart. To identify these regions, the positive rate fluctuation value and the SG filtering value of the positive rate in each feature partition are used as two-dimensional features and input into the k-means clustering model, which is executed by the non-linear explorer in the non-linear pattern evaluation module. By using the clustering method, users can observe multiple non-linear regions of the data in the form of a scatter plot in the non-linear pattern evaluation view.
[0050] In the scatter plot of the non-linear pattern evaluation view, the scatter plot shares the Y-axis with the feature evaluation view, as Figure 7As shown in C1 in [reference], the X-axis of the scatter plot shows the positive rate fluctuations. The clustered point clusters are presented in the form of a scatter plot and are used to observe multiple non-linear regions of the data. In this view, fluctuations less than 0 are shaded, and the non-linear patterns are displayed in a more intuitive way.
[0051] In the control area of the non-linear mode evaluation view, as Figure 7 shown in C2 in [reference], it supports the user to modify the number of clusters of k-means clustering through the control area. It supports clicking the recoloring button to update the number of colors of the color blocks. In addition, the text element to the left of the save button (Current scheme: loan amount - 8) represents the currently selected feature and its discrete scheme. Clicking the save button will upload the scheme and be reflected in the control panel. At the same time, the system will save the current page. Feature switching will not cause the feature evaluation and non-linear mode view obtained by the current adjustment to be cleared.
[0052] In summary, for the visualization system of interpretability-aware feature discretization for logistic regression provided by the embodiments of the present invention, the interaction between the views includes the following steps:
[0053] (1) Input multi-dimensional data containing multiple features into the control panel and adjust various parameters.
[0054] (2) Construct a basic model through the logistic regression model instance builder and display it in the model evolution view. Features in different situations in the model (i.e., the logistic regression model instance) will be color-coded differently and displayed in the outer ring of the composite graph. Through the model evolution view, modify based on the basic model or the parent model, and a child node will be drawn. The inner circle part of the child node is the comparison with its parent model.
[0055] (3) According to the model obtained in step (2), through the feature evaluation controller, calculate the positive rate and the positive rate after SG filter filtering, and visually display them to the feature evaluation view. In this view, the data can be fed back to the control panel by adjusting the interval size to achieve the effect of step (2).
[0056] (4) According to the model obtained in step (2), through the non-linear explorer, visualize it in the non-linear mode evaluation view by means of k-means clustering. The number of clusters can be modified in the view and uploaded to the control panel through the save button.
[0057] The above-described specific embodiments have elaborated on the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, supplements, equivalent replacements, etc. made within the scope of the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A visualization system for interpretability-aware feature discretization of logistic regression, characterized in that include: The model evolution module constructs a corresponding logistic regression model instance as a basic model according to the characteristics of the sample input by the user, modifies the basic model or its generated parent model to generate a new logistic regression model instance to establish the evolution and optimization process of the logistic regression model instance and visualize it as a tree diagram, takes the logistic regression model instance as a node, and constructs a composite graph of each node in the tree diagram to evaluate the performance of each logistic regression model instance to compare the performance changes of adjacent instances; The control module displays and adjusts the relevant parameters of the current instance and the feature overview of the current input sample in the control panel based on a logistic regression model instance built in the model evolution module, and supports users to adjust parameters and select features; The feature evaluation module, based on a logistic regression model instance built in the model evolution module, performs feature partitioning on the sample data according to the discretization scheme and input features set by the control panel, calculates the sample size and positive rate information in all feature partitions, and visualizes them as line graph areas and histogram areas to capture feature nonlinear patterns; The nonlinear pattern evaluation module calculates the nonlinear area through a clustering algorithm based on the positive rate information and visualizes it as a scatter plot. It is combined with the feature evaluation module to evaluate and adjust the discretization scheme.
2. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 1, characterized in that: In the model evolution module, the tree diagram of the evolution and optimization process of the logistic regression model instance is presented as a binary tree diagram, with a common root node. Each parent node supports the generation of one or two child nodes after it. Each node of the binary tree diagram is visualized as a composite graph representing the logistic regression model instance. When a new logistic regression model instance is constructed, the tree diagram will grow with the corresponding composite graph node; Among them, the composite graph of the root node includes an interactive text element, an outer ring and a radar chart, and the composite graph of the child node includes an interactive text element, an outer ring and two radar charts; the interactive text element includes the name of the logistic regression example, and the interactive text element is clicked to select the logistic regression model instance represented by the root node or the child node, and supports separate zooming in to display the corresponding composite graph; each segment in the outer ring represents a corresponding feature of the input, the size of each segment represents the relative contribution of the feature to the logistic regression model instance, and the color depth of each segment represents the absolute strength of the feature effect; the performance of the logistic regression model instance of the root node is displayed through a radar chart in the composite graph of the root node; the performance of the logistic regression model instance of the child node and its parent node are respectively displayed through two radar charts in the composite graph of the child node.
3. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 1 or 2, characterized in that: The performance of the logistic regression model instance includes: accuracy, precision, recall, F1-score, and normalized entropy.
4. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 2, characterized in that: During the generation of the tree diagram, a curve is drawn from the right edge of the composite graph corresponding to the selected node, indicating that another composite graph of the next logistic regression model instance will be drawn at the end of the curve. When the composite graph corresponding to the next logistic regression model instance will exceed the preset view range or will overlap, a horizontal line is automatically drawn instead of a curve to avoid possible view misalignment. If the next logistic regression model instance fails to draw the expected composite graph, the corresponding curve or straight line will be removed.
5. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 1, characterized in that: In the feature evaluation module, the line graph area includes a graph composed of two curves of different colors and a graph composed of two types of color blocks, gray and color. The Y-axis represents the size of the positive rate, and the X-axis represents the number of the feature partition; the two curves of different colors represent the original positive rate data and the positive rate data processed by the SG filter respectively; the gray type of color block is used as the background to indicate that the sample size in the feature partition is insufficient and the confidence level of the positive rate expression is low; the color type of color block is drawn based on the accuracy filtering and its clustering label, and its width represents the range of the corresponding area, and the height represents the range of the positive rate in the feature partition. Different colors represent different clustering labels.
6. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 5, characterized in that: Each color block is divided into regions according to the maximum value, upper quartile, median, lower quartile and minimum value of the vertical coordinate. The greater the height of the area between the upper quartile and the lower quartile, the larger the interval between the upper and lower quartiles of the area, which means that the samples in the area are more dispersed and the area is linear. Conversely, the more concentrated the samples in the area are, the more nonlinear the area is.
7. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 1, characterized in that: In the feature evaluation module, the histogram area includes a histogram, a composite diagram on the top of the histogram, and an interactive background color block; the Y-axis of the histogram represents the sample size, the X-axis represents the number of the feature partition, the height of the histogram represents the sample size in the feature partition, and the color of the histogram represents the confidence level of the positive rate in the feature partition. The histogram is encoded in different colors according to whether the sample size meets the standard; the composite diagram on the top of the histogram indicator line consists of a straight line and a point. The length of the straight line represents the size of the positive rate fluctuation value, that is, the first-order difference of the positive rate in each interval after filtering, and different colors and directions are used to represent whether the fluctuation value is negative or positive; the interactive background color block alternates between gray and transparent, and the dotted line around the color block is displayed by touching the mouse. The width of the color block represents the currently discretized area, and the color blocks support interaction. The user controls its width by directly dragging the color block and updates the discretization scheme created by the decision tree.
8. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 1, characterized in that: In the nonlinear pattern evaluation module, the scatter plot shares the Y-axis with the line graph area in the feature evaluation module, and the X-axis of the scatter plot displays the positive rate fluctuation value. The positive rate fluctuation value and the SG filtered value of the positive rate in each feature partition are input as two-dimensional features into the k-means clustering model to obtain the clustered point clusters. The point clusters are presented in the form of a scatter plot and used to observe multiple nonlinear areas of the data.
9. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 1, characterized in that: The nonlinear pattern evaluation module also includes a control area, which supports the user to modify the number of clusters in the clustering algorithm through the control area, supports clicking the recoloring button to update the number of colors of the color block, and supports clicking the save button to save the current discretization scheme and upload it to the control module.
10. The visualization system for interpretable perceptual feature discretization of logistic regression according to claim 1, characterized in that: In the control module, it supports displaying and adjusting the relevant parameters of the logistic regression model instance currently selected in the tree diagram of the model evolution module through the control panel. The relevant parameters include interval, window size, level, tree depth adjustment, discretization and prediction. Among them, the interval indicates that the continuous data should be divided into pseudo-stationary durations, and the size of the duration is called the interval. The window size indicates the sliding window size applicable to the SG filter, the level indicates the level of the SG filter polynomial, and the tree depth adjustment indicates the depth of the tree, which is applicable to the training of the CART algorithm in discretization. Clicking discretization triggers the discretization command, and clicking prediction starts the system operation; it supports displaying a feature overview for different features, supports selecting the range of the current feature and clicking the input button to enter the feature, and supports clicking the get details button to obtain feature detailed information.
Citation Information
Patent Citations
Machine learning algorithm model training method and device and storage medium
CN114219096A
Global trade transfer influence analysis method and visualization system
CN118154224A