Data exploration analysis method and system based on financial risk control data visualization
Through the data exploration and analysis method of financial risk control data visualization, the problem of insufficient adaptability of the risk control system is solved, rapid analysis of data and timely response to risk signals are achieved, and the accuracy and efficiency of risk control decisions are improved.
Patent Information
- Application Number
- CN202510929245.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-21
AI Technical Summary
The existing financial risk control system is unable to adaptively adjust strategies and lacks deep data interaction and real-time feedback mechanisms, resulting in inaccurate risk assessment and low decision-making efficiency.
Design a data exploration and analysis method based on the visualization of financial risk control data, including data collection, preprocessing and cleaning, visualization, interactive analysis and feedback. Through reinforcement learning, optimize risk control strategies to achieve rapid data analysis and timely response to risk signals.
It improves the decision-making accuracy and execution efficiency of risk control data, realizes dynamic optimization and real-time adjustment of risk control strategies, and enhances user interaction experience.
Smart Images

Figure CN120821750A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data visualization management, and specifically relates to a data exploration and analysis method and system based on financial risk control data visualization. Background Art
[0002] Current financial risk control systems mostly rely on traditional static models and rule-based engines. However, with the ever-changing financial market environment and the continuous emergence of massive amounts of data, risk control systems based on traditional static models or fixed rules are becoming increasingly limited. Traditional financial risk control systems utilize historical data to construct static risk assessment models, using pre-set rules to issue risk warnings but unable to adaptively adjust strategies. Some data visualization platforms can display financial data in graphical and chart form, but lack deep data interaction and real-time feedback mechanisms. Some risk control systems have attempted to incorporate reinforcement learning to optimize risk control strategies, but their data processing and display modules are not integrated with the visualization platform, resulting in insufficient overall interactivity.
[0003] Therefore, it is necessary to design a financial risk control data analysis method and system to achieve rapid dynamic analysis of risk control data, data exploration interaction and timely response to risk signals, so as to improve the accuracy of decision-making and execution efficiency. Summary of the Invention
[0004] In order to solve the above problems, this application designs a data exploration and analysis method and system based on financial risk control data visualization, which realizes rapid analysis of risk control data and timely response to risk signals through data collection, data preprocessing and cleaning, data visualization display, data exploration and interactive analysis feedback.
[0005] The purpose of the present invention can be achieved through the following technical solutions:
[0006] A data exploration and analysis method based on financial risk control data visualization includes the following steps:
[0007] Step S1, collecting data;
[0008] Step S2: preprocessing and cleaning the data;
[0009] Step S3: Visualize the data;
[0010] Step S4: interactively explore and analyze the data;
[0011] Step S5: Optimize the risk control strategy based on the feedback data;
[0012] Step S6: interactively feedback data and generate a report.
[0013] Preferably, in step S1, collecting data specifically includes the following steps:
[0014] Step S11: collecting raw data from multiple financial data sources;
[0015] Step S12: Initialize the data interface, automatically call each data interface, and support automatic capture and synchronization of structured, semi-structured and unstructured data.
[0016] Preferably, in step S2, preprocessing and cleaning the data includes the following steps:
[0017] Step S21: Clean the data to filter out noise, duplicate data, and erroneous data to ensure high quality of the data.
[0018] Step S22: Process missing values in the data so that the missing values in the data do not cause deviation or error in subsequent data analysis or model training;
[0019] Step S23: Eliminate abnormal data by identifying outliers in the data and eliminating data records that are significantly deviated from the normal range;
[0020] Step S24: Store the data in a unified format to ensure that the data type is consistent and meets expectations during subsequent processing.
[0021] Preferably, in step S21, data cleaning includes the following steps:
[0022] Step S211: Detect and remove duplicate data by using database query or deduplication function in programming language to detect identical records and determine whether duplicate data exists based on key fields. If duplicate data exists, retain one copy and delete the other records.
[0023] Step S212: Identify and correct erroneous data, check whether there are illegal characters, data type mismatches, or logical errors in the data; correct logical errors, such as replacing erroneous characters with spaces or specific marks, or inferring correct values based on contextual data.
[0024] Preferably, in step S22, processing missing values in the data includes the following steps:
[0025] Step S221: Use statistical analysis or programming tools to detect missing values in each column;
[0026] Step S222: Select an appropriate filling method to fill missing values based on data attributes and business scenarios. For numerical data, use the mean, median, or mode. For skewed data, use the quantile filling method. For data with strong continuity, use interpolation. For categorical data, use the mode or a custom missing marker.
[0027] Step S223: Set a threshold. For example, when the missing rate of a column exceeds 30%, the column is directly discarded, or when the number of missing fields in a record exceeds a certain ratio, the record is discarded.
[0028] Preferably, in step S23, removing abnormal data includes the following steps:
[0029] Step S231: Detect outliers in numerical data using a statistical distribution-based method. The Z-score method is used, which calculates the difference between each data point and the mean. If the Z-score exceeds a set threshold (assuming it is 3), it is considered an outlier. The IQR method is used, which calculates the quartiles of the data, with a lower limit of Q1-1.5*IQR and an upper limit of Q3+1.5*IQR. Data outside this range is considered an outlier.
[0030] Step S232: For the detected abnormal data, select to remove, correct or mark it according to the business scenario.
[0031] Preferably, in step S24, storing the data in a unified format includes the following steps:
[0032] Step S241: Standardize the date and time format, converting various date and time information into a standard format (e.g., YYYY-MM-DD HH:MM:SS); use a programming language (e.g., the datetime module in Python) or a database built-in date function to perform the conversion;
[0033] Step S242: convert the data format, convert all numerical data into the same data type, and perform scaling processing on the numerical values that need to be normalized or standardized;
[0034] Step S243: uniformly encode the text data and convert the text data into UTF-8 encoding to avoid data parsing errors due to inconsistent encoding; remove or replace unnecessary spaces and special symbols in the text to ensure consistent data format.
[0035] Preferably, in step S3, visually displaying the data includes the following steps:
[0036] Step S31: Environment setup and library configuration, setting up the front-end framework and graphics engine, installing dependency management tools and initializing the icon container;
[0037] Step S32: Obtain processed data from the backend via API or WebSocket, and update the chart data periodically or in real time. Use WebSocket to establish a persistent connection, and immediately push new data from the server to the frontend. After obtaining the data, call the chart's setOption method to reload the data, setting notMerge:false in the parameter to achieve a smooth transition update.
[0038] Step S33: Design configuration files for various charts such as bar charts, line charts, heat maps, and relationship charts, and dynamically generate charts using Echarts or D3.js;
[0039] Step S34: In the code and interface display, all English abbreviations are annotated with Chinese to ensure that the system is easy to understand and maintain.
[0040] Preferably, in step S4, interactively exploring and analyzing the data includes the following steps:
[0041] Step S41: Select different data dimensions through the graphical interface for interactive exploration;
[0042] Step S42: Through the data linkage mechanism, click, drag and drop operations to refresh relevant data information in real time;
[0043] Step S43: Integrate data mining algorithms to provide multiple analysis methods such as clustering and association rules to reveal the implicit association risks between data.
[0044] Preferably, in step S42, refreshing relevant data information in real time by operations such as clicking and dragging through a data linkage mechanism includes the following steps:
[0045] Step S421: In the chart component, use the event binding interface provided by the graphics engine to capture user interaction behaviors such as clicks and drags;
[0046] Step S422: Using the front-end event bus or the global state management library, the screening parameters generated by the user operation are transmitted to other components or data processing modules;
[0047] Step S423: In the front-end data filtering module, the local data is filtered, aggregated, and re-counted according to the passed parameters to obtain a new data set. If the data volume is large or the latest data needs to be obtained, after the front-end captures the interaction event, the back-end API is called to request the updated data.
[0048] Step S424: After receiving the filtered data, call the chart engine's setOption method to update the chart display and implement real-time refresh. When the data filtering results or interaction parameters are updated, use the chart engine's setOption method to update the data and configuration, immediately triggering a redraw of the chart. If a smooth transition is required, set notMerge:false and lazyUpdate:true in the parameters to implement a partial update. Update the chart configuration items based on the filtered data and call the refresh method of the chart instance.
[0049] Step S425: All chart components that subscribe to the interactive information respond simultaneously, achieving multi-chart linkage display and synchronous update effects, enhancing the user interaction experience. Through the event bus or global state management, the filter conditions are passed to all relevant components, causing them to call their respective update methods to re-render.
[0050] Preferably, in step S43, integrating data mining algorithms to provide multiple analysis methods such as clustering and association rules to reveal the implicit association risks between data includes the following steps:
[0051] Step S431: Before integrated data mining, the original data is fully preprocessed and feature constructed to ensure that the data input of subsequent algorithms has unified standards and high quality.
[0052] Step S432: Clustering rules can use methods such as K-Means, Hierarchical Clustering, DBSCAN, etc., and select an appropriate algorithm for cluster analysis based on the data characteristics;
[0053] Step S433: Association rule algorithms include the Apriori algorithm and the FP-Growth algorithm. Select an algorithm suitable for the data scale and transaction density to mine frequent item sets and their association rules. From the preprocessed data, scan each transaction or user feature set to generate a basic item set. Based on the set minimum support threshold, filter out frequent item sets that meet the conditions. From the frequent item sets, mine association rules that meet the minimum confidence requirement and output indicators such as support and confidence for each rule. The generated association rules describe the relationships between different attributes or risk characteristics, providing a quantitative basis for interpreting risk correlations.
[0054] Step S434: Collaboration and integration of multiple data mining algorithms. First, the preprocessing module provides input data in the same format for clustering and association rule mining, ensuring that the two algorithms operate based on a unified feature space. The results of cluster analysis (such as cluster labels) can be passed to the association rule module as filtering conditions to achieve localized rule mining. Secondly, in terms of system architecture, a distributed computing framework (such as Spark or Hadoop) is used for parallel computing. Clustering and association rule algorithms can run simultaneously and process different data partitions or clusters respectively. A unified risk scoring system is established by cross-validating clustering results and association rules. Based on historical feedback data, an adaptive adjustment mechanism is established to dynamically adjust the weights of clustering and association rules in the overall risk assessment to achieve multi-objective collaborative optimization. In the visualization display module, graphical display of clustering results and association rules can be achieved through interactive linkage. Users can click on a cluster area to view the key association rules in that area. At the same time, statistical charts and network diagrams of association rules can intuitively display the connections between risk factors.
[0055] Preferably, in step S5, optimizing the risk control strategy includes the following steps:
[0056] Step S51: Based on real-time feedback data, dynamically optimize the risk control strategy using reinforcement learning;
[0057] Step S52: Construct a reward and punishment mechanism to evaluate the decision-making effect in real time and automatically adjust the model parameters;
[0058] Step S53: Output the decision result and feed it back to the back-end execution system through the interface.
[0059] Preferably, in step S51, dynamically optimizing the risk control strategy using reinforcement learning based on real-time feedback data includes the following steps:
[0060] Step S511: Environment construction, including designing states, actions, and reward functions;
[0061] Step S512: Dynamically optimize the risk control strategy using the Q-Learning method;
[0062] Step S513: Dynamically optimize the risk control strategy using the Deep Reinforcement Learning method;
[0063] Preferably, in step S52, building a reward and punishment mechanism, evaluating decision-making effects in real time, and automatically adjusting model parameters include the following steps:
[0064] Step S521: When building a reward and penalty mechanism, determine core business indicators and goals, clarify the main business goals and key indicators of the risk control system, which directly reflect the effectiveness of decision-making. After determining the indicators, set a target value or expected range for each indicator and assign weights based on business needs;
[0065] Step S522: When designing the reward and penalty function mechanism, when constructing the reward function, positive effects (such as high accuracy and low response latency) are mapped as positive rewards, and negative effects (such as high misjudgment rate and risk loss) are mapped as penalties;
[0066] Step S523: Real-time evaluation of decision-making effects, that is, in each decision cycle, the actual risk control effect is quantified into a reward or penalty signal by monitoring key indicators. The latest financial data, transaction records, user feedback, etc. are obtained through the ETL process to ensure the real-time nature of the data. After each decision cycle, the real-time monitoring module is used to collect statistics on the effectiveness of the current risk control decision. Set a calculation to automatically trigger at the end of each fixed time window or each decision cycle to score the performance of the current strategy. The reward value and related indicators of each cycle are stored in the database to facilitate subsequent statistics, analysis, and dynamic adjustment of model parameters.
[0067] Step S524: Automatically adjust model parameters to achieve dynamic optimization;
[0068] Step S525: Comprehensive feedback loop and system integration. From data collection, indicator monitoring, reward calculation to model update, a complete feedback loop is formed. When a new decision is made, the reward calculation module is triggered, and the new reward signal is fed back to the reinforcement learning module, thereby updating the strategy parameters. A real-time monitoring dashboard is established to track cumulative rewards, key indicators, and model performance. If continuous negative rewards or abnormal fluctuations are found, the system automatically triggers an alarm or rollback mechanism to ensure the stable operation of the risk control system. By recording the changes in reward values, decision effects, and feedback data online, grid search, Bayesian optimization, and other methods are used to automatically adjust key parameters in the model (such as learning rate, discount factor, number of network layers, etc.) to achieve dynamic optimization. An online learning mechanism is constructed to allow the system to continuously update the model during actual operation to adapt to changes in the market environment.
[0069] A data exploration and analysis system based on the visualization of financial risk control data, including:
[0070] The data acquisition module is used to collect raw data from various financial data sources. It has a built-in data interface and supports automatic capture and synchronization of structured, semi-structured, and unstructured data.
[0071] The data preprocessing and cleaning module is used to clean the collected data, handle missing values, eliminate abnormal data, and standardize the format. It uses extraction, conversion, and loading technologies to ensure data consistency and timeliness.
[0072] The data visualization display module is used to generate and update data charts in real time. It supports multiple display formats and provides interactive operations. All English abbreviations are annotated in Chinese, such as RL for reinforcement learning and ETL for extraction, transformation, and loading technology.
[0073] The data exploration and interactive analysis module enables users to interactively explore different data dimensions through a graphical interface. The module implements a data linkage mechanism, allowing users to refresh the relevant data display in real time through operations such as clicking and dragging. It also integrates data mining algorithms and provides various analysis methods such as clustering and association rules to reveal implicit correlation risks between data.
[0074] The risk control decision-making and optimization module is used to dynamically optimize risk control strategies, build reward and penalty mechanisms, evaluate decision effects in real time, automatically adjust model parameters, output decision results, and feed them back to the back-end execution system through an interface;
[0075] The interactive feedback and report generation module is used to display the data and system responses after user operations through real-time charts and reports. The system also supports customized report generation and automatic archiving for subsequent statistical analysis and historical data comparison.
[0076] The advantages and effects of this application are as follows:
[0077] This application designs a data exploration and analysis method and system based on the visualization of financial risk control data. The method collects data, pre-processes and cleans the data; then visualizes the processed data, generates charts and updates the data in real time; secondly, it conducts data exploration and interactive analysis, performs interactive data screening and display, and performs data mining; thirdly, it dynamically optimizes risk control decisions through reinforcement learning strategies; finally, it feeds back interactive data, generates data analysis reports and historical records.
[0078] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application so that it can be implemented in accordance with the contents of the specification, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following is a detailed description of the preferred embodiment of the present application in conjunction with the accompanying drawings.
[0079] Based on the detailed description of the specific embodiments of the present application in conjunction with the accompanying drawings below, those skilled in the art will become more aware of the above and other objects, advantages and features of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without inventive work. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.
[0081] Figure 1 A flowchart of a data exploration and analysis method and system based on financial risk control data visualization designed for this application;
[0082] Figure 2 A data processing and visualization flowchart of a data exploration and analysis method and system based on financial risk control data visualization designed for this application;
[0083] Figure 3 A data exploration and analysis method based on financial risk control data visualization and a flow chart of the reinforcement learning optimization module of the system designed for this application. DETAILED DESCRIPTION
[0084] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. In the following description, specific details such as specific configurations and components are provided only to help fully understand the embodiments of the present application. Therefore, it should be clear to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. In addition, for clarity and brevity, the description of known functions and structures has been omitted in the embodiments.
[0085] It should be understood that references throughout this specification to "one embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the appearance of "one embodiment" or "this embodiment" throughout this specification does not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0086] In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or settings discussed.
[0087] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" in this article describes another type of association object relationship, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0088] The term "at least one" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, at least one of A and B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0089] It should also be noted that, in this document, relational terms such as first and second are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include," "comprises," or any other variations thereof are intended to cover non-exclusive inclusion.
[0090] Example 1
[0091] Please refer to Figures 1 to 3 This embodiment mainly introduces a data exploration and analysis method based on financial risk control data visualization, including the following steps:
[0092] Step S1, collecting data, specifically includes the following steps:
[0093] Step S11: Collecting raw data from various financial data sources (such as transaction data, credit records, market conditions, etc.);
[0094] Step S12: Initialize the data interface and automatically call various data interfaces (such as REST API, database direct connection, etc.) to support automatic capture and synchronization of structured, semi-structured and unstructured data.
[0095] Step S2, preprocessing and cleaning the data, includes the following steps:
[0096] Step S21: Clean the data to filter out noise data, duplicate data, and erroneous data to ensure high-quality data. This step specifically includes the following steps:
[0097] Step S211: Duplicate data detection and deduplication: Use database query or deduplication function (such as drop_duplicates()) in programming language (such as Python Pandas) to detect identical records, and determine whether duplicate data exists based on key fields (such as unique identifier, timestamp). If duplicate data exists, retain one copy of the data and delete the other records;
[0098] Step S212: Identify and correct erroneous data, check whether there are illegal characters, data type mismatches, or logical errors in the data; correct logical errors, such as replacing erroneous characters with spaces or specific marks, or inferring correct values based on contextual data.
[0099] Step S22, processing missing values in the data, includes the following steps:
[0100] Step S221: Use statistical analysis or programming tools (such as the isnull() function in Pandas) to detect missing values in each column;
[0101] Step S222: Select an appropriate filling method to fill missing values based on data attributes and business scenarios. For numerical data, use the mean, median, or mode. If the data distribution is skewed, use the quantile filling method. For data with strong continuity, use interpolation (such as linear interpolation). For categorical data, use the mode or a custom missing marker (such as "unknown" or "null").
[0102] Step S223: Set a threshold. For example, when the missing rate of a column exceeds 30%, the column is directly discarded, or when the number of missing fields in a record exceeds a certain ratio, the record is discarded.
[0103] Step S23, removing abnormal data, includes the following steps:
[0104] Step S231: Detect outliers in numerical data using a statistical distribution-based method. The Z-score method is used, which calculates the difference between each data point and the mean. If the Z-score exceeds a set threshold (assuming it is 3), it is considered an outlier. The IQR method is used, which calculates the quartiles of the data, with a lower limit of Q1-1.5*IQR and an upper limit of Q3+1.5*IQR. Data outside this range is considered an outlier.
[0105] Step S232: For detected abnormal data, select whether to remove, modify, or mark it based on the business scenario. For example, if a transaction amount is significantly outside the normal range (perhaps due to a data entry error or system error), the record can be deleted; or the abnormal data can be subject to additional review before deciding whether to remove it.
[0106] Step S24, storing the data in a unified format, includes the following steps:
[0107] Step S241: Standardize the date and time format, converting various date and time information into a standard format (e.g., YYYY-MM-DD HH:MM:SS); use a programming language (e.g., the datetime module in Python) or a database built-in date function to perform the conversion;
[0108] Step S242: convert the data format, convert all numerical data into the same data type (such as float or int), and perform scaling processing on the values that need to be normalized or standardized;
[0109] Step S243: uniformly encode the text data and convert the text data into UTF-8 encoding to avoid data parsing errors due to inconsistent encoding; remove or replace unnecessary spaces and special symbols in the text to ensure consistent data format.
[0110] Step S3: After the data is preprocessed and cleaned, visual display of the data includes the following steps:
[0111] Step S31: Environment setup and library configuration, set up the front-end framework and graphics engine, install dependency management tools, and initialize the icon container. The front-end framework can use Vue.js, React, or Angular to build the overall data page; choose Echarts or D3.js as the graphics engine, use NPM or Yarn for dependency management, and use Webpack or other packaging tools to build the project;
[0112] Step S32: Obtain processed data from the backend via API or WebSocket, and update the chart data periodically or in real time. Use WebSocket to establish a persistent connection, and immediately push new data from the server to the frontend. After obtaining the data, call the chart's setOption method to reload the data. The parameter can be set to notMerge:false to achieve a smooth transition update.
[0113] Step S33: Design configuration files for various charts such as bar charts, line charts, heat maps, and relationship charts, and dynamically generate charts using Echarts or D3.js;
[0114] Step S34: In the code and interface display, all English abbreviations (such as RL, ETL) are annotated with Chinese to ensure that the system is easy to understand and maintain.
[0115] Step S4, interactively exploring and analyzing the data, includes the following steps:
[0116] Step S41: Select different data dimensions through the graphical interface for interactive exploration;
[0117] Step S42, through the data linkage mechanism, click, drag and other operations to refresh relevant data information in real time, includes the following steps:
[0118] Step S421: In the chart component, use the event binding interface provided by the graphics engine (such as Echarts or D3.js) to capture user interaction behaviors such as clicks and drags. For example, for click events, in a bar chart, when a user clicks a bar, the category or value corresponding to the bar is captured; for drag events, in a scatter plot or heat map, when a user drags to select an area, the coordinates or data range of the area are captured;
[0119] Step S422: Use a front-end event bus or global state management library (such as Vuex or Redux) to transmit the filtering parameters (such as category, time, and value range) generated by user operations to other components or data processing modules. In simple projects, establish an empty Vue instance as the event transmission center and implement message transmission between components through the $emit and $on methods; in large projects, use a state management library to store interaction parameters in the global state, and each component subscribes to related state changes and automatically refreshes the display;
[0120] Step S423: In the front-end data filtering module, the local data is filtered, aggregated, and re-counted according to the passed parameters to obtain a new data set. If the data volume is large or the latest data needs to be obtained, after the front-end captures the interaction event, the back-end API is called to request the updated data.
[0121] Step S424: After receiving the filtered data, call the chart engine's setOption method to update the chart display and implement real-time refresh. When the data filtering results or interaction parameters are updated, use the chart engine's setOption method to update the data and configuration, immediately triggering a redraw of the chart. If a smooth transition is required, set notMerge:false and lazyUpdate:true in the parameters to implement a partial update. Update the chart configuration items based on the filtered data and call the refresh method of the chart instance.
[0122] Step S425: All chart components that subscribe to this interactive information respond simultaneously, achieving a multi-chart linkage display and synchronous update effect, enhancing the user interaction experience. Through the event bus or global state management, the filter conditions are passed to all relevant components, causing them to call their respective update methods to re-render. For example, chart component A captures the click event, updates local data, and notifies chart component B via the event bus to recalculate and refresh the display. The two chart components display data of different dimensions respectively, share common filter parameters, and achieve chart linkage.
[0123] Step S43: Integrate data mining algorithms to provide multiple analysis methods such as clustering and association rules to reveal the implicit association risks between data. The steps include:
[0124] Step S431: Before integrated data mining, the raw data is fully pre-processed and feature-constructed to ensure that the data input to each subsequent algorithm has a unified standard and high quality. During data cleaning, missing data, outliers, and duplicate records are removed, and the data format is standardized; key attributes (such as transaction frequency, credit score, behavioral characteristics, etc.) are extracted from the raw data based on the business scenario, and multi-dimensional feature vectors are constructed. The data is processed through standardization and normalization techniques to ensure the consistency of the data used by different algorithms; the data is preliminarily divided according to business requirements, such as grouping by time, region, or user type, to provide a basis for subsequent localized analysis (such as mining association rules within different groups);
[0125] Step S432: Clustering rules can use methods such as K-Means, Hierarchical Clustering, and DBSCAN, and select appropriate algorithms for cluster analysis based on data characteristics. For example, by clustering user or transaction data, high-risk groups and low-risk groups can be quickly located. Initialize the number of clusters or parameter settings (such as K value, distance measurement method), and set the initial parameters based on pre-investigation or heuristic methods; use the processed feature vector as input data to perform clustering operations; evaluate the clustering quality through indicators such as silhouette coefficient and intra-cluster error sum of squares, and adjust and optimize the clustering parameters; each data sample corresponds to a cluster label, which serves as an important basis for subsequent mining of association rules and risk assessment;
[0126] Step S433: Common association rule algorithms include the Apriori algorithm and the FP-Growth algorithm. Select an algorithm suitable for the data scale and transaction density to mine frequent item sets and their association rules. From the preprocessed data, scan each transaction or user feature set to generate a basic item set; based on the set minimum support threshold, filter out frequent item sets that meet the conditions; mine association rules that meet the minimum confidence requirements from the frequent item sets, and output indicators such as the support and confidence of each rule; the generated association rules describe the relationship between different attributes (or risk characteristics), providing a quantitative basis for interpreting risk correlations;
[0127] Step S434: Collaboration and integration of multiple data mining algorithms. First, the preprocessing module provides input data in the same format for clustering and association rule mining, ensuring that both algorithms operate within a unified feature space. The results of cluster analysis (e.g., cluster labels) can be passed as filtering criteria to the association rule module, enabling localized rule mining. For example, performing association rule analysis within high-risk groups can uncover more targeted risk combination patterns. Second, within the system architecture, a distributed computing framework (e.g., Spark or Hadoop) is employed for parallel computing, allowing clustering and association rule algorithms to run simultaneously and process different data partitions or clusters. By cross-validating clustering results and association rules, a unified risk scoring system is established. For example, if multiple high-frequency association rules (with high support and confidence) are found in a cluster, the risk warning signal within that cluster is stronger. Using historical feedback data, an adaptive adjustment mechanism is established to dynamically adjust the weights of clustering and association rules in the overall risk assessment, achieving multi-objective collaborative optimization. In the visualization module, graphical presentation of clustering results and association rules can be achieved through interactive linkage. Users can click on a cluster area to view the key association rules in the area. At the same time, the statistical graphs and network graphs of the association rules can intuitively display the connections between risk factors.
[0128] Step S5: Optimizing the risk control strategy based on the feedback data includes the following steps:
[0129] Step S51, based on real-time feedback data, dynamically optimizes the risk control strategy using reinforcement learning (e.g., Q-Learning, deep reinforcement learning), including the following steps:
[0130] Step S511, environment construction, including the design of state, action, and reward function. In state design, the state vector is defined to reflect the current market conditions and risk characteristics. The state space can be discrete (simpler, lower dimension) or continuous (high dimension, continuous characteristics); in action design, the action can be defined as adjusting certain parameters in the decision model, adjusting the risk threshold, or selecting different risk control strategy combinations. For example, action A1 means lowering the risk warning line, A2 means improving the credit approval standard, etc.; in reward function design, the reward function is designed based on risk loss, risk prediction accuracy, cost-benefit and other indicators. The reward function design takes into account both short-term (immediate risk warning accuracy) and long-term (historical cumulative risk and benefit balance) factors. For example, the reward function R = α × (prediction accuracy) - β × (default loss) - γ × (adjustment cost) is designed. The weights of parameters α, β, γ can be adjusted according to actual data and business strategies;
[0131] Step S512: Dynamically optimize the risk control strategy using the Q-Learning method. Construct a Q-table that records the expected long-term return of taking a certain action under a certain state. The update rule is based on the Bellman equation:
[0132]
[0133] Where α is the learning rate and γ is the discount factor.
[0134] For a given discrete state space, initialize the Q-table, where all state-action pairs can be set to zero or small random values. Use an ε-greedy strategy to balance exploration (randomly selecting actions) and exploitation (selecting the action with the largest Q value according to the Q-table). Gradually reduce the ε value during training to increase the exploitation rate. At each step, select action a from the current state s, observe the next state s' and the reward R obtained, and update Q(s,a) according to the formula. Set the number of iterations or monitor the change in Q-value. When the change in Q-value falls below a certain threshold, the model is considered converged.
[0135] Step S513: Use the Deep Reinforcement Learning method to dynamically optimize the risk control strategy. Normalize and standardize the original state vector of the data to form the input data of the network; design a network structure suitable for the current business scenario: for example, a 3-5 layer fully connected neural network, and the activation function can use ReLU. The number of nodes in the output layer is consistent with the total number of actions, and the Q value corresponding to each action is output; each (s, a, R, s') is stored in the experience replay buffer, and small batch samples are randomly selected for training to break the data correlation; the target network is used to regularly synchronize the main network parameters to ensure training stability; the mean square error (MSE) is used to compare the predicted Q value with the target Q value:
[0136]
[0137] Use the ε-greedy strategy to balance exploration and exploitation during training; use the validation set or simulation environment to evaluate model performance, monitor the convergence speed and accumulated rewards, and adjust the network structure, learning rate, experience replay capacity and ε strategy based on the results.
[0138] Step S52, building a reward and penalty mechanism, real-time evaluation of decision-making effects and automatic adjustment of model parameters, includes the following steps:
[0139] Step S521: When building a reward and penalty mechanism, determine core business indicators and goals, and clarify the main business goals and key indicators of the risk control system. These indicators directly reflect the effectiveness of decision-making. For example, the risk prediction accuracy rate is the proportion of correctly predicted risk events, the misjudgment rate is the ratio of incorrectly misjudging low risks as high risks or vice versa, the response delay is the time required from data input to decision output, and the risk loss is the actual economic loss caused by decision-making errors. After determining the indicators, set a target value or expected range for each indicator and assign weights based on business needs;
[0140] Step S522: When designing the reward and penalty function mechanism, when constructing the reward function, positive effects (such as high accuracy and low response latency) are mapped to positive rewards, and negative effects (such as high misjudgment rate and risk loss) are mapped to penalties. The mathematical expression can be selected as follows:
[0141] R = α·Accuracy - β·False positive rate - γ·Risk loss - δ·Response delay
[0142] in:
[0143] α, β, γ, and δ are adjustment coefficients, and the best effect is obtained through experimental tuning; each indicator is normalized to ensure reasonable comparison between different dimensions.
[0144] Step S523, real-time evaluation of decision-making effects, that is, in each decision cycle, the actual risk control effect is quantified into a reward (or punishment) signal by monitoring key indicators. The latest financial data, transaction records, user feedback, etc. are obtained through the ETL process to ensure the real-time nature of the data. After the end of each decision cycle, the real-time monitoring module is used to count the effects of the current risk control decision, for example: statistical prediction accuracy, the number of misjudgments, and the occurrence of real risk events; record decision delays and processing time, calculate the average response delay; collect actual economic loss data caused by decision-making errors, and these data are summarized in real time through system logs, online databases or message queues. Perform real-time calculations on numerical values, design an independent reward calculation module, and after receiving the data of each key indicator, calculate a reward value R according to the pre-designed reward function formula. Set a calculation to be automatically triggered at the end of each fixed time window or each decision cycle to score the performance of the current strategy. Store the reward value and related indicators of each cycle in the database to facilitate subsequent statistics, analysis and dynamic adjustment of parameters;
[0145] Step S524: Automatically adjust model parameters to achieve dynamic optimization. For discrete state spaces, automatically adjust model parameters based on Q-Learning: Initialize Q values for all possible state-action pairs (s, a)(s, a)(s, a), usually set to 0 or a random value. When the system selects action aaa in state sss and observes the new state s′s′s′ and reward RRR, update the Q value according to the Bellman equation:
[0146]
[0147] Where: η is the learning rate, which controls the impact of new information on the Q value; γ is the discount factor, which balances the immediate reward with the future reward; R is the reward value calculated previously.
[0148] The action to take in the current state is determined using an ε-greedy strategy, utilizing the current optimal strategy (greedily selecting the action with the maximum Q value) while also retaining the random exploration ratio. The Q table is updated after each decision cycle to continuously move the strategy toward the optimal solution. For high-dimensional or continuous state spaces, model parameters are automatically adjusted based on deep reinforcement learning. A multi-layer fully connected neural network (or convolutional / recurrent network, depending on the data type) is designed, with the current state vector as input and the Q values corresponding to all possible actions as output. An experience replay buffer is set up to store each (s, a, R, s') sample to break correlations between data and stabilize training. A target network is introduced to stabilize the training process, and the main network parameters are periodically copied to the target network to prevent oscillations during training. The mean squared error (MSE) is used as the loss function to compare the current predicted Q value with the target Q value:
[0149]
[0150] The gradient descent algorithm is used to update the neural network parameters, so that the loss function gradually decreases and the Q value output by the model becomes more accurate. In actual operation, the system continuously collects real-time data and reward feedback, and through online network training or fine-tuning, the model parameters automatically adapt to market changes.
[0151] Step S525: Comprehensive feedback loop and system integration. From data collection, indicator monitoring, reward calculation to model update, a complete feedback loop is formed. When a new decision is made, the reward calculation module is triggered, and the new reward signal is fed back to the reinforcement learning module, thereby updating the strategy parameters. A real-time monitoring dashboard is established to track cumulative rewards, key indicators, and model performance. If continuous negative rewards or abnormal fluctuations are found, the system automatically triggers an alarm or rollback mechanism to ensure the stable operation of the risk control system. By recording the changes in reward values, decision effects, and feedback data online, grid search, Bayesian optimization, and other methods are used to automatically adjust key parameters in the model (such as learning rate, discount factor, number of network layers, etc.) to achieve dynamic optimization. An online learning mechanism is constructed to allow the system to continuously update the model during actual operation to adapt to changes in the market environment.
[0152] Step S53: Output the decision result and feed it back to the back-end execution system through the interface.
[0153] Step S6: interactively feedback data and generate a report.
[0154] The above description is merely a preferred embodiment of the present invention and does not limit the scope of protection of the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any variation, modification, replacement, integration, or parameter change to these embodiments, which is within the spirit and principles of the present invention and which achieves the same functionality through conventional substitutions, without departing from the principles and spirit of the present invention, falls within the scope of the present invention.
Claims
1. A data exploration and analysis method based on the visualization of financial risk control data, characterized in that: The following steps are involved: Step S1, collecting data; Step S2: preprocessing and cleaning the data; Step S3: Visualize the data; Step S4: interactively explore and analyze the data; Step S5: Optimize the risk control strategy based on the feedback data; Step S6: interactively feedback data and generate a report.
2. The data exploration and analysis method based on visualization of financial risk control data according to claim 1, characterized in that: In step S1, collecting data includes the following steps: Step S11: collecting raw data from multiple financial data sources; Step S12: Initialize the data interface, automatically call each data interface, and support automatic capture and synchronization of structured, semi-structured and unstructured data.
3. The data exploration and analysis method based on financial risk control data visualization according to claim 1 is characterized in that: In step S2, preprocessing and cleaning the data includes the following steps: Step S21: Clean the data to filter out noise, duplicate data, and erroneous data to ensure high quality of the data. Step S22: Process missing values in the data so that the missing values in the data do not cause deviation or error in subsequent data analysis or model training; Step S23: Eliminate abnormal data by identifying outliers in the data and eliminating data records that are significantly deviated from the normal range; Step S24: Store the data in a unified format to ensure that the data type is consistent and meets expectations during subsequent processing; In step S22, processing missing values in the data includes the following steps: Step S221: Use statistical analysis or programming tools to detect missing values in each column; Step S222: Select an appropriate filling method to fill missing values based on data attributes and business scenarios. For numerical data, use the mean, median, or mode. For skewed data, use the quantile filling method. For data with strong continuity, use interpolation. For categorical data, use the mode or a custom missing marker. Step S223: Set a threshold. For example, if the missing rate of a column exceeds 30%, the column is discarded directly, or if the number of missing fields in a record exceeds a certain ratio, the record is discarded. In step S23, the abnormal data removal includes the following steps: Step S231: Detect outliers in numerical data using a statistical distribution-based method. The Z-score method is used, which calculates the difference between each data point and the mean. If the Z-score exceeds a set threshold, it is considered an outlier. The IQR method is used, which calculates the quartiles of the data, with a lower limit of Q1-1.5*IQR and an upper limit of Q3+1.5*IQR. Data outside this range is considered an outlier. Step S232: For the detected abnormal data, select to remove, correct or mark it according to the business scenario.
4. The data exploration and analysis method based on visualization of financial risk control data according to claim 3 is characterized in that: In step S3, visualizing the data includes the following steps: Step S31: Environment setup and library configuration, setting up the front-end framework and graphics engine, installing dependency management tools and initializing the icon container; Step S32: Obtain processed data from the backend via API or WebSocket, and update the chart data periodically or in real time. Use WebSocket to establish a persistent connection, and immediately push new data from the server to the frontend. After obtaining the data, call the chart's setOption method to reload the data, setting notMerge:false in the parameter to achieve a smooth transition update. Step S33: Design configuration files for various charts such as bar charts, line charts, heat maps, and relationship charts, and dynamically generate charts using Echarts or D3.js; Step S34: In the code and interface display, all English abbreviations are annotated with Chinese to ensure that the system is easy to understand and maintain.
5. The data exploration and analysis method based on visualization of financial risk control data according to claim 4, characterized in that: In step S4, interactive exploration and analysis of data includes the following steps: Step S41: Select different data dimensions through the graphical interface for interactive exploration; Step S42: Through the data linkage mechanism, click, drag and drop operations to refresh relevant data information in real time; Step S43: Integrate data mining algorithms to provide various analysis methods such as clustering and association rules to reveal the implicit association risks between data; In step S42, the real-time refreshing of relevant data information by operations such as clicking and dragging through the data linkage mechanism includes the following steps: Step S421: In the chart component, use the event binding interface provided by the graphics engine to capture user interaction behaviors such as clicks and drags; Step S422: Using the front-end event bus or the global state management library, the screening parameters generated by the user operation are transmitted to other components or data processing modules; Step S423: In the front-end data filtering module, the local data is filtered, aggregated, and re-counted according to the passed parameters to obtain a new data set. If the data volume is large or the latest data needs to be obtained, after the front-end captures the interaction event, the back-end API is called to request the updated data. Step S424: After receiving the filtered data, call the chart engine's setOption method to update the chart display and implement real-time refresh. When the data filtering results or interaction parameters are updated, use the chart engine's setOption method to update the data and configuration, immediately triggering a redraw of the chart. If a smooth transition is required, set notMerge:false and lazyUpdate:true in the parameters to implement a partial update. Update the chart configuration items based on the filtered data and call the refresh method of the chart instance. Step S425: All chart components that subscribe to the interactive information respond simultaneously, achieving multi-chart linkage display and synchronous update effects, enhancing the user interaction experience. Through the event bus or global state management, the filter conditions are passed to all relevant components, causing them to call their respective update methods to re-render.
6. The data exploration and analysis method based on visualization of financial risk control data according to claim 5, characterized in that: In step S43, the data mining algorithm is integrated to provide a variety of analysis methods such as clustering and association rules to reveal the implicit association risks between data, including the following steps: Step S431: Before integrated data mining, perform comprehensive preprocessing and feature construction on the original data to ensure that the data input of subsequent algorithms has unified standards and high quality; Step S432: Clustering rules can use K-Means, hierarchical clustering, DBSCAN and other methods, and select appropriate algorithms for cluster analysis according to data characteristics; Step S433: Association rule algorithms include the Apriori algorithm and the FP-Growth algorithm. Select an algorithm suitable for the data scale and transaction density to mine frequent item sets and their association rules. From the preprocessed data, scan each transaction or user feature set to generate a basic item set. Based on the set minimum support threshold, filter out frequent item sets that meet the conditions. From the frequent item sets, mine association rules that meet the minimum confidence requirement and output indicators such as support and confidence for each rule. The generated association rules describe the relationships between different attributes or risk characteristics, providing a quantitative basis for interpreting risk correlations. Step S434: Collaboration and integration of multiple data mining algorithms. First, the preprocessing module provides input data in the same format for clustering and association rule mining, ensuring that the two algorithms operate based on a unified feature space. The results of cluster analysis can be passed to the association rule module as filtering conditions to achieve localized rule mining. Secondly, in the system architecture, a distributed computing framework is used for parallel computing, allowing clustering and association rule algorithms to run simultaneously and process different data partitions or clusters respectively. A unified risk scoring system is established by cross-validating clustering results and association rules. Using historical feedback data, an adaptive adjustment mechanism is established to dynamically adjust the weights of clustering and association rules in the overall risk assessment, achieving multi-objective collaborative optimization. In the visualization module, graphical display of clustering results and association rules can be achieved through interactive linkage. Users can click on a cluster area to view the key association rules within that area. At the same time, statistical charts and network diagrams of association rules can intuitively display the connections between risk factors.
7. The data exploration and analysis method based on visualization of financial risk control data according to claim 6, characterized in that: In step S5, optimizing the risk control strategy includes the following steps: Step S51: Based on real-time feedback data, dynamically optimize the risk control strategy using reinforcement learning; Step S52: Construct a reward and punishment mechanism to evaluate the decision-making effect in real time and automatically adjust the model parameters; Step S53: Output the decision result and feed it back to the back-end execution system through the interface.
8. The data exploration and analysis method based on visualization of financial risk control data according to claim 7, characterized in that: In step S51, based on real-time feedback data, the risk control strategy is dynamically optimized using reinforcement learning, including the following steps: Step S511: Environment construction, including designing states, actions, and reward functions; Step S512: Dynamically optimize the risk control strategy using the Q-Learning method; Step S513: Use the Deep Reinforcement Learning method to dynamically optimize the risk control strategy.
9. The data exploration and analysis method based on visualization of financial risk control data according to claim 8, characterized in that: In step S52, building a reward and punishment mechanism, evaluating decision-making results in real time, and automatically adjusting model parameters include the following steps: Step S521: When building a reward and penalty mechanism, determine core business indicators and goals, clarify the main business goals and key indicators of the risk control system, which directly reflect the effectiveness of decision-making. After determining the indicators, set a target value or expected range for each indicator and assign weights based on business needs; Step S522: When designing the reward and penalty function mechanism, when constructing the reward function, positive effects are mapped to positive rewards, and negative effects are mapped to penalties; Step S523: Real-time evaluation of decision-making effects, that is, in each decision cycle, the actual risk control effect is quantified into a reward or penalty signal by monitoring key indicators. The latest financial data, transaction records, user feedback, etc. are obtained through the ETL process to ensure the real-time nature of the data. After each decision cycle, the real-time monitoring module is used to collect statistics on the effectiveness of the current risk control decision. Set a calculation to automatically trigger at the end of each fixed time window or each decision cycle to score the performance of the current strategy. The reward value and related indicators of each cycle are stored in the database to facilitate subsequent statistics, analysis, and dynamic adjustment of model parameters. Step S524: Automatically adjust model parameters to achieve dynamic optimization; Step S525: Comprehensive feedback loop and system integration. A complete feedback loop is formed, encompassing data collection, indicator monitoring, reward calculation, and model updating. When a new decision is made, the reward calculation module is triggered, and the new reward signal is fed back to the reinforcement learning module, which in turn updates the strategy parameters. A real-time monitoring dashboard is established to track cumulative rewards, key indicators, and model performance. If continuous negative rewards or abnormal fluctuations are detected, the system automatically triggers an alarm or rollback mechanism to ensure stable operation of the risk control system. By online recording of reward value changes, decision-making results, and feedback data, key model parameters are automatically adjusted using methods such as grid search and Bayesian optimization to achieve dynamic optimization. An online learning mechanism is established to allow the system to continuously update the model during actual operation to adapt to changes in the market environment.
10. A data exploration and analysis system based on the visualization of financial risk control data, characterized in that: include: The data acquisition module is used to collect raw data from various financial data sources. It has a built-in data interface and supports automatic capture and synchronization of structured, semi-structured, and unstructured data. The data preprocessing and cleaning module is used to clean the collected data, handle missing values, eliminate abnormal data, and standardize the format. It uses extraction, conversion, and loading technologies to ensure data consistency and timeliness. The data visualization display module is used to generate and update data charts in real time. It supports multiple display formats and provides interactive operations. All English abbreviations are annotated in Chinese, such as RL for reinforcement learning and ETL for extraction, transformation, and loading technology. The data exploration and interactive analysis module enables users to interactively explore different data dimensions through a graphical interface. The module implements a data linkage mechanism, allowing users to refresh the relevant data display in real time through operations such as clicking and dragging. It also integrates data mining algorithms and provides various analysis methods such as clustering and association rules to reveal implicit correlation risks between data. The risk control decision-making and optimization module is used to dynamically optimize risk control strategies, build reward and penalty mechanisms, evaluate decision effects in real time, automatically adjust model parameters, output decision results, and feed them back to the back-end execution system through an interface; The interactive feedback and report generation module is used to display the data and system responses after user operations through real-time charts and reports. The system also supports customized report generation and automatic archiving for subsequent statistical analysis and historical data comparison.