Intelligent decision-making system based on reinforcement learning
Through an intelligent decision-making system based on reinforcement learning, dynamic risk thresholds are generated and Lagrange multipliers are adjusted in real time to optimize strategy network parameters. This solves the problem of inaccurate risk assessment in complex market environments caused by traditional credit decision-making methods, realizes intelligent and precise credit decision-making, and improves business efficiency and security.
Patent Information
- Application Number
- CN202510981183.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional credit decision-making methods are based on fixed rules and static risk assessment models, which are difficult to adapt to complex and changing market environments and customer characteristics, resulting in inaccurate credit risk assessment and low efficiency.
An intelligent decision-making system based on reinforcement learning is adopted. Through the collaborative work of multiple modules, dynamic risk thresholds are generated in real time, Lagrange multipliers are dynamically adjusted, strategy network parameters are optimized, and early warnings are triggered when risks exceed limits, thus realizing intelligent, dynamic and precise credit decision-making.
It improves the accuracy of credit risk assessment, balances multiple risk constraints, enhances business efficiency, and ensures the security of credit business.
Smart Images

Figure CN120634274A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial technology, and in particular to an intelligent decision-making system based on reinforcement learning. Background Art
[0002] In the credit business, intelligent and accurate decision-making systems are crucial for controlling risk and improving business efficiency. Traditional credit decision-making methods, often based on fixed rules and static risk assessment models, struggle to adapt to complex and changing market environments and customer characteristics. With the continuous development and innovation of the financial market, credit operations face increasing uncertainty and risk factors, such as industry fluctuations and changes in customer credit status.
[0003] Therefore, it is necessary to provide an intelligent decision-making system based on reinforcement learning to solve the above technical problems. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides an intelligent decision-making system based on reinforcement learning. Through multi-module collaboration, it realizes the intelligence, dynamism and precision of credit decision-making, effectively reduces credit risks and improves business efficiency.
[0005] The present invention provides an intelligent decision-making system based on reinforcement learning, the system comprising:
[0006] A threshold generation module is used to extract customer feature vectors and environment state vectors in real time in response to the current credit application flow, and generate a dynamic risk threshold set based on the environment state vectors;
[0007] A decision action determination module, configured to input the customer feature vector and the environment state vector into a pre-trained policy network and determine a decision action based on an output result of the policy network;
[0008] A risk function value calculation module, configured to execute the decision action and calculate the function values of multiple risk constraints in real time based on the execution results;
[0009] a Lagrange multiplier adjustment module, configured to determine deviations by correspondingly comparing the function values of the plurality of risk constraints with the dynamic risk threshold set, and dynamically adjust the Lagrange multipliers associated with each risk constraint according to the deviations;
[0010] A policy network parameter updating module, configured to update the parameters of the policy network using a primal-dual gradient descent method based on a real-time reward signal fed back after the decision action is executed and a dynamically adjusted Lagrange multiplier;
[0011] The risk warning module is used to output the decision action to the risk control execution end, and trigger a real-time risk warning when the function value of any risk constraint reaches the preset proportional coefficient of the corresponding dynamic risk threshold.
[0012] Preferably, the threshold generation module is specifically used to:
[0013] Extracting real-time risk factors of preset dimensions from the environmental state vector to form a risk factor vector, wherein the real-time risk factors include the real-time default rate of the industry and the change in capital adequacy ratio;
[0014] Recalling a pre-stored weight matrix and bias vector, wherein the weight matrix and bias vector are generated by learning the historical credit data used for training the policy network, and each column vector in the weight matrix corresponds to a risk type of a plurality of risk constraints;
[0015] The weight matrix and the bias vector are combined to generate a dynamic risk threshold set corresponding to the multiple risk constraints, wherein the generation formula of the dynamic risk threshold set is:
[0016]
[0017] Among them, σ represents the Sigmoid activation function, T j represents the dynamic risk threshold, b j is the j-th component in the bias vector b, F risk represents the risk factor vector; Indicates that the column vector W j Transposed to a row vector, W j represents the column vector corresponding to the j-th risk constraint in the weight matrix W, and · represents the vector dot product operation.
[0018] Preferably, the decision action determination module is specifically used to:
[0019] Concatenate the customer feature vector and the environment state vector along the feature dimension to generate a joint feature vector;
[0020] Inputting the joint feature vector into the policy network and outputting a decision action probability distribution, wherein the dimensions of the decision action probability distribution correspond one-to-one to the discrete options of the decision action;
[0021] Based on the decision action probability distribution, a decision action is determined based on a maximizing probability selection strategy.
[0022] Preferably, the risk function value calculation module is specifically used to:
[0023] Executing the decision action determined based on the maximum probability selection strategy, and obtaining credit business execution result data after the decision action is executed;
[0024] extracting characteristic indicators related to each risk constraint from the credit business execution result data;
[0025] Based on the characteristic indicators, according to the preset calculation rules corresponding to each risk constraint, the function values of multiple risk constraints are calculated in real time.
[0026] Preferably, the Lagrange multiplier adjustment module is specifically used to:
[0027] For each risk constraint, calculate its function value C j With dynamic risk threshold T j The difference between the two generates a sub-deviation, where the sub-deviation δ j The generation formula is:
[0028] δ j =C j -T j (j=1,2,…,m);
[0029] For all sub-deviations, weighted norm aggregation is performed according to the preset weights to generate deviations;
[0030] Based on the sub-deviations corresponding to each risk constraint, the Lagrange multipliers associated with each risk constraint are dynamically adjusted according to a preset Lagrange multiplier adjustment rule. The preset Lagrange multiplier adjustment rule is to increase or decrease the Lagrange multiplier according to a preset proportional coefficient based on the size and positive or negative sign of the sub-deviation.
[0031] Preferably, the policy network parameter updating module is specifically used to:
[0032] Obtaining a real-time reward signal fed back after the decision action is executed, and calculating an original optimization target gradient based on the real-time reward signal and current parameters of the policy network;
[0033] For each risk constraint, the gradient of the constraint penalty term is calculated by combining the updated Lagrange multiplier and the function value;
[0034] The constraint penalty term gradient and the original optimization target gradient are combined to update the parameters of the policy network through the primal-dual gradient descent method.
[0035] Preferably, the risk warning module is specifically used to:
[0036] The function value C of each risk constraint calculated j Corresponding dynamic risk threshold T j , calculate the proportion of the target value T j ·K j , where K j Indicates the preset scale factor;
[0037] When any function value C j Satisfy C j ≥T j ·K j When , a warning trigger signal for the risk constraint is generated;
[0038] The following actions are performed according to the early warning trigger signal:
[0039] Transmit the determined decision action to the risk control execution end;
[0040] The warning trigger signal is synchronously sent to the risk monitoring end.
[0041] Compared with related technologies, the intelligent decision-making system based on reinforcement learning provided by the present invention has the following beneficial effects:
[0042] The threshold generation module of the present invention can generate a dynamic risk threshold set in real time based on the environmental state vector, allowing the system to flexibly adapt to changing market and customer conditions, thereby improving the accuracy of risk assessment. Secondly, the Lagrange multiplier adjustment module determines the deviation by comparing the risk function value and the dynamic risk threshold, and dynamically adjusts the multiplier accordingly, effectively balancing multiple risk constraints and avoiding neglecting other risks due to excessive focus on a single risk. Furthermore, the policy network parameter update module combines the real-time reward signal and the adjusted multiplier, optimizes the parameters through the primal-dual gradient descent method, and improves the decision-making performance of the policy network. Finally, the risk warning module triggers a warning in a timely manner when the value of any risk constraint function reaches a preset ratio, thereby ensuring the security of the credit business. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of the module structure of an intelligent decision-making system based on reinforcement learning provided by the present invention. DETAILED DESCRIPTION
[0044] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all of the structures. Furthermore, the embodiments of the present invention and the features of the embodiments may be combined with one another unless there is a conflict.
[0045] It should also be noted that, for ease of description, only the part relevant to the present invention, rather than all of the content, is shown in the accompanying drawings. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processing or methods depicted as flow charts. Although flow charts describe various operations (or steps) as sequential processing, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of various operations can be rearranged. When its operation is completed, the processing can be terminated, but can also have additional steps not included in the accompanying drawings. The processing can correspond to methods, functions, procedures, subroutines, subprograms, etc.
[0046] The present invention provides an intelligent decision-making system based on reinforcement learning. Figure 1 As shown, the system includes:
[0047] The threshold generation module 100 is used to extract customer feature vectors and environment state vectors in real time in response to the current credit application flow, and generate a dynamic risk threshold set according to the environment state vectors.
[0048] Specifically, the threshold generation module 100 is specifically configured to:
[0049] Real-time risk factors of preset dimensions are extracted from the environmental state vector to form a risk factor vector, wherein the real-time risk factors include the real-time default rate of the industry and the change in capital adequacy ratio.
[0050] In this embodiment, a real-time environment state vector (data structure: [timestamp, macroeconomic index, industry default rate, capital adequacy ratio, market volatility, policy adjustment coefficient]) is received. The vector is parsed by a preset factor positioning engine:
[0051] Real-time industry default rate: The ratio of the total loan default amount to the total loan amount for the entire industry on that day is obtained from the API of the central bank's credit reporting system, smoothed using a sliding window (window size = 7 days), and retained to 4 decimal places.
[0052] Change in capital adequacy ratio: Calculates the deviation between the current commercial bank's core tier 1 capital adequacy ratio and the average of the previous five days. Data is sourced from the regulatory reporting system of the State Financial Supervision and Administration Bureau. Any deviation outside the regulatory range of [8%, 15%] will be automatically marked as abnormal.
[0053] Supplementary factor processing: The annualized market volatility is calculated using the GARCH (1, 1) model, and the policy adjustment coefficient is quantified based on the semantic analysis of the central bank's monetary policy announcement (loose = 0.8 / neutral = 0.5 / tightening = 0.2).
[0054] Finally, the standardized risk factor vector F is output. risk .
[0055] Calling pre-stored weight matrix and bias vector, wherein the weight matrix and bias vector are generated by learning the historical credit data used for training the policy network, and each column vector in the weight matrix corresponds to a risk type of multiple risk constraints.
[0056] In this embodiment, the weight matrix and bias vector are stored in a financial-grade encrypted database (including but not limited to Oracle Exadata) using three copies. The parameter structure is:
[0057]
[0058] Among them, the rows correspond to three types of constraints: credit risk, market risk, and operational risk, and the columns correspond to four risk factors.
[0059] The weight matrix and the bias vector are combined to generate a dynamic risk threshold set corresponding to the multiple risk constraints, wherein the generation formula of the dynamic risk threshold set is:
[0060]
[0061] Among them, σ represents the Sigmoid activation function, T j represents the dynamic risk threshold, b j is the j-th component in the bias vector b, F risk represents the risk factor vector; Indicates that the column vector W j Transposed to a row vector, W j represents the column vector corresponding to the j-th risk constraint in the weight matrix W, and · represents the vector dot product operation.
[0062] In this embodiment, this formula is the core calculation model of the dynamic risk threshold generation mechanism. Its technical function is to generate dynamic regulatory boundaries that match each type of risk constraint in real time through the fusion of multi-source risk factors and nonlinear probability conversion. The specific technical logic is as follows:
[0063] Input layer processing: based on the risk factor vector F risk Integrate real-time risk indicators: industry default rate (degree of deterioration of credit assets), capital adequacy ratio (institutional risk resistance), market volatility (systemic risk pressure) and policy adjustment index (changes in the regulatory environment).
[0064] Linear transformation layer processing: weight matrix column vector W j Mark risk sensitivity. A positive weight indicates that an increase in capital adequacy ratio suppresses the risk threshold, while a negative weight indicates that an increase in industry default rate pushes up the risk threshold.
[0065] Bias term b j Provides benchmark calibration: automatically increases b during economic recessionsj Relax risk tolerance and reduce b when the market is overheated j Tighten regulatory standards and use dot product operations to quantify risk exposure.
[0066] The probabilistic output layer Sigmoid function σ(·) performs nonlinear mapping, which compresses the linear combination result to the (0, 1) interval, and the output value T j This is an example of the probability of triggering risk intervention.
[0067] The decision action determination module 200 is used to input the customer feature vector and the environment state vector into a pre-trained policy network, and determine a decision action according to an output result of the policy network.
[0068] Specifically, the decision action determination module 200 is specifically configured to:
[0069] The customer feature vector and the environment state vector are concatenated along the feature dimension to generate a joint feature vector.
[0070] In this embodiment, real-time input of customer feature vectors and environmental state vectors is received. The customer feature vector includes key indicators such as credit score, historical overdue payments, and debt-to-income ratio (typically 15 dimensions), while the environmental state vector covers macroeconomic factors such as economic volatility index and industry risk coefficient (typically 6 dimensions). The input data is first normalized: missing values are filled using a 7-day moving average for continuous features and a historical mode for categorical features. For outliers, credit scores outside the range of [300, 850] are automatically truncated to the boundary value.
[0071] Then, a concatenation operation is performed along the feature dimensions to generate a joint feature vector whose dimensions are the sum of the customer feature dimension and the environment state dimension (21 dimensions in this example). This concatenated vector is then Z-score normalized, with the normalization parameters (mean and standard deviation) derived from the training dataset statistics. The final output is the normalized joint feature vector.
[0072] The joint feature vector is input into the policy network, and a decision action probability distribution is output, wherein the dimensions of the decision action probability distribution correspond one-to-one to the discrete options of the decision action.
[0073] In this example, the pre-trained policy network uses a three-layer fully connected architecture: the number of input layer nodes matches the dimension of the joint feature vector (21 nodes), the hidden layer has 64 nodes and uses the ReLU activation function f(x) = max(0, x), and the number of output layer nodes corresponds to the number of decision actions (5 nodes) and uses the Softmax activation function. Financial-grade security measures are employed during network deployment: model parameters are stored with AES-256 encryption, and inference is performed within the Intel SGX trusted execution environment.
[0074] The inference execution process involves layer-by-layer forward computation: After the input layer receives the joint feature vector, the hidden layer performs a linear transformation and ReLU activation, and the output layer generates a probability distribution for the decision action. A typical output is [Reject: 0.2, Credit Level 1: 0.5, Credit Level 2: 0.3], indicating the probability of selecting each action.
[0075] Based on the decision action probability distribution, a decision action is determined based on a maximizing probability selection strategy.
[0076] In this embodiment, differentiated action selection strategies are adopted according to the system operation mode: in training mode, the ε-greedy strategy is implemented, and actions are randomly selected with probability ε to ensure exploration (the initial value of ε is 0.3 and decays in each round), and the highest probability action is selected with probability 1-ε; in production mode, the maximum probability strategy is strictly adopted, and the action with the highest Softmax output probability is directly selected.
[0077] The discrete action space is predefined as: {0: "application rejected", 1: "credit limit 10,000", 2: "credit limit 50,000", 3: "credit limit 100,000", 4: "manual review"}. After execution, a structured output instruction is generated, including fields such as timestamp, decision action, confidence level, and action ID.
[0078] For example, when a customer's credit score is 725, debt ratio is 35%, there is one past payment, environmental economic index is 0.6, and industry risk is 0.4: feature concatenation generates a standardized joint vector [1.2, -0.8, 0.3, 0.4, -1.1]; policy network inference outputs a probability distribution [reject: 0.1, 10,000: 0.2, 50,000: 0.6, 100,000: 0.1]; in production mode, the action with the highest probability, ID = 2, is selected, and a decision to grant a credit of 50,000 is output. The entire process completes within 22ms, meeting the real-time decision-making requirements of financial systems.
[0079] The risk function value calculation module 300 is used to execute the decision action and calculate the function values of multiple risk constraints in real time according to the execution results.
[0080] Specifically, the risk function value calculation module 300 is specifically used to:
[0081] Execute the decision action determined based on the maximum probability selection strategy, and obtain credit business execution result data after the decision action is executed.
[0082] In this embodiment, the system receives the structured instructions output by the decision action determination module 200 and executes the decision action through the bank core system API: when action_id=0, the loan rejection interface is called; when action_id=1-3, the credit approval interface is called; when action_id=4, the manual review process is triggered.
[0083] Before execution, a double security check (decision signature verification and credit limit compliance check) is performed. Upon passing the check, credit business data is collected in real time from multiple source systems: basic customer information is obtained through a JDBC direct connection to the customer relationship management system, credit results are captured from the Kafka message queue of the core credit system, and repayment forecast data is generated by calling the gRPC interface of the risk measurement engine. The collected data is integrated into structured records, for example, including key fields such as timestamp, customer ID, decision action, credit limit, probability of default (PD), loss given default (LGD), and industry type.
[0084] Extract characteristic indicators related to each risk constraint from the credit business execution result data.
[0085] In this embodiment, feature mapping rules are designed for each type of risk constraint:
[0086] Correlation characteristics of single expected loss function:
[0087] Probability of default (PD) → directly read the pd_score field of the data record;
[0088] Loss Given Default (LGD) → read the lgd_rate field;
[0089] Risk exposure (EAD) → equal to the credit limit loan_amount;
[0090] Industry concentration function correlation characteristics:
[0091] Industry type → parse the industry field;
[0092] Current credit limit → directly read loan_amount;
[0093] Total institution risk exposure → Real-time SQL query credit system (SELECT SUM (loan_balance) FROM loan_portfolio);
[0094] To ensure query efficiency, a cache optimization mechanism is implemented: if the total exposure cache time related to the industry concentration calculation does not exceed 5 minutes and there are no large loan events (single loan > 1 million), the cache value will be used directly; otherwise, the database will be re-queried and the cache will be updated.
[0095] Based on the characteristic indicators, according to the preset calculation rules corresponding to each risk constraint, the function values of multiple risk constraints are calculated in real time.
[0096] In this embodiment, the single expected loss function value is calculated using the formula C1 = PD × LGD × EAD. The input example pd_score = 0.032, lgd_rate = 0.65, loan_amount = 50000 outputs 0.032 × 0.65 × 50000 = 1040 yuan (precision is retained to an integer). The industry concentration function value calculation executes a step-by-step process: first query the current exposure of the manufacturing industry (such as 38 million), combine it with the total exposure cache value of the institution (4.2 billion), and then calculate it according to the formula Calculated (Percentages are rounded to three decimal places.) A single loss exceeding 1 million triggers a red alert and interrupts the process. Industry concentration exceeding 15% indicates a high-risk status. A calculation timeout of 200ms automatically switches to the historical mean approximation algorithm.
[0097] The Lagrange multiplier adjustment module 400 is configured to determine deviations by correspondingly comparing the function values of the plurality of risk constraints with the dynamic risk threshold set, and dynamically adjust the Lagrange multipliers associated with each risk constraint according to the deviations.
[0098] Specifically, the Lagrange multiplier adjustment module 400 is specifically used to:
[0099] For each risk constraint, calculate its function value C j With dynamic risk threshold T j The difference between the two generates a sub-deviation, where the sub-deviation δ j The generation formula is:
[0100] δ j =C j -T j (j=1,2,…,m).
[0101] In this embodiment, the system receives the risk constraint function value set (example format: {credit risk: 0.085, market risk: 0.182, operational risk: 0.043}) output by the risk function value calculation module 300 and the dynamic risk threshold value set (example: {credit risk: 0.07, market risk: 0.15, operational risk: 0.05}) generated by the threshold generation module (100). The input data is strictly verified: when the function value is missing, the cached data within 5 seconds is called to supplement it; when the threshold exceeds the range of (0, 1), it is automatically truncated to the boundary value of 0.01 or 0.99. The sub-deviation is calculated by constraint-by-constraint subtraction according to the above formula, where j in the above formula represents credit, market, and operation in this embodiment.
[0102] For all sub-deviations, weighted norm aggregation is performed according to the preset weights to generate deviations.
[0103] In this example, risk weights based on regulatory rules are preset: a credit risk weight of 0.50 (based on Basel III core capital requirements), a market risk weight of 0.35 (based on historical backtesting of the VaR model), and an operational risk weight of 0.15 (set based on the annual loss ratio). Deviation aggregation is performed using the L2 norm.
[0104] Based on the sub-deviations corresponding to each risk constraint, the Lagrange multipliers associated with each risk constraint are dynamically adjusted according to a preset Lagrange multiplier adjustment rule. The preset Lagrange multiplier adjustment rule is to increase or decrease the Lagrange multiplier according to a preset proportional coefficient based on the size and positive or negative sign of the sub-deviation.
[0105] In this embodiment, a rule-driven multiplier update engine is constructed, specifically including:
[0106] Rule base configuration: positive deviation (δ j >0) Scenario: where α j is the magnification factor (credit risk 0.12 / market risk 0.10 / operational risk 0.08), where is the Lagrange multiplier at the current moment (the penalty intensity of the risk constraint), when δ j When α > 0, it means the risk exceeds the limit and the penalty needs to be strengthened; j Equivalent to risk price regulator: α j The larger the value, the greater the multiplier increase when the unit risk exceeds the limit;
[0107] Negative deviation (δ j <0)Scene: 0.05 is the risk release attenuation coefficient, when δ j When <0, it means that the risk is controllable and the penalty can be weakened; the risk release attenuation coefficient is designed to be less than α j : Embody the principle of asymmetric regulation (rapid response when risks accumulate, slow exit when risks are released);
[0108] Update execution process: Enter the current multiplier value (for example: credit risk: 0.40, market risk: 0.35, operational risk: 0.25);
[0109] Credit risk (δ = +0.015): 0.40 + 0.12 × 0.015 = 0.4018;
[0110] Market risk (δ = +0.032): 0.35 + 0.10 × 0.032 = 0.3532;
[0111] Operational risk (δ = -0.007): max(0, 0.25 - 0.05 × 0.007) = 0.24965;
[0112] Output the updated multiplier subset: [0.4018, 0.3532, 0.24965], and transmit it to the policy network parameter update module 500.
[0113] The policy network parameter update module 500 is used to update the parameters of the policy network by the primal-dual gradient descent method based on the real-time reward signal fed back after the execution of the decision-making action and the dynamically adjusted Lagrange multiplier.
[0114] Specifically, the policy network parameter update module 500 is specifically used for:
[0115] Obtain the real-time reward signal fed back after the execution of the decision-making action, and calculate the gradient of the original optimization objective based on the real-time reward signal and the current parameters of the policy network.
[0116] In this embodiment, the system receives the real-time reward signal r fed back by the credit core system t , and the data structure includes a timestamp, a customer ID, a reward type, and a numerical reward value (range [-1, 1]). The reward calculation follows strict economic rules: normal repayment reward +1 × (interest income - capital cost), default loan penalty -k × (principal loss + recovery cost) (k > 1 reflects risk aversion), and opportunity cost deduction -c (c << k) for rejected loans. Based on the current policy network parameters θ [[ID=S]] (t) and the state-action pair (s t , a t ), calculate the policy gradient:
[0117]
[0118] where Q π (s t , a t ) is the state-action value function, and the future cumulative discounted reward is estimated through the Critic network (the discount factor γ is 0.95). When implemented, the PyTorch automatic differentiation mechanism is adopted. First, the advantage function is evaluated through the Critic network, and then the gradient is calculated by combining the logarithm of the action probability output by the policy network.
[0119] For each risk constraint, calculate the gradient of the constraint penalty term by combining the updated Lagrange multiplier and the function value.
[0120] In this embodiment, input the multiplier subset updated by the Lagrange multiplier adjustment module 400 (Example: credit risk 0.4018, market risk 0.3532) and the C output by the risk function value calculation module 300 j (Example: Single loss 1040 yuan, industry concentration 0.904%). Calculate the gradient of the constraint penalty term:
[0121]
[0122] This is achieved by chain derivation: first calculate the risk function C j For action a t The partial derivative of , combined with the gradient of the action probability of the policy network to the parameter θ. j As the weight coefficient, construct the loss function ∑λ j ·C j Then call the automatic differentiation engine (such as torch.autograd.grad) to obtain the gradient tensor.
[0123] The constraint penalty term gradient and the original optimization target gradient are combined to update the parameters of the policy network through the primal-dual gradient descent method.
[0124] In this embodiment, the policy network parameters are updated by the primal-dual gradient descent method:
[0125] θ (t+1) =θ (t) -η·(g policy -g constraint );
[0126] The learning rate η is dynamically adjusted using the adaptive Adam optimizer, and gradient clipping is performed before updating to ensure numerical stability. The updated parameter θ (t+1) It is encrypted with AES-256 and stored in the parameter server, and synchronized in real time to the policy network instance of the decision action determination module 200.
[0127] The risk warning module 600 is used to output the decision action to the risk control execution end, and trigger a real-time risk warning when the function value of any risk constraint reaches a preset proportional coefficient of the corresponding dynamic risk threshold.
[0128] Specifically, the risk warning module 600 is specifically used to:
[0129] The function value C of each risk constraint calculated j Corresponding dynamic risk threshold T j , calculate the proportion of the target value T j ·K j , where K j Indicates the preset scale factor.
[0130] In this embodiment, the system receives the risk constraint function value set (including indicators such as single loss value and industry concentration) output by the risk function value calculation module 300 and the dynamic risk threshold set (such as credit risk threshold 0.467 and market risk threshold 0.582) generated by the threshold generation module 100.
[0131] Load the preset scaling factor configuration table: Credit risk scaling factor of 0.90 (based on Basel III capital buffer requirements), Market risk factor of 0.85 (based on stress testing VAR breach probability), and Operational risk factor of 0.95 (based on historical loss event statistics). Calculate the target scaling value by risk type multiplication: multiply each risk type threshold by its corresponding scaling factor, retaining the result to five decimal places. For example: Multiplying the market risk threshold of 0.582 by the scaling factor of 0.85 yields a target scaling value of 0.4947.
[0132] When any function value C j Satisfy C j ≥T j ·K j When , an early warning trigger signal for the risk constraint is generated.
[0133] In this embodiment, a multi-level early warning judgment engine is established: each risk function value is compared with the corresponding proportion standard value in real time. When the function value ≥ the proportion standard value, a warning signal is triggered, and three levels of response are divided according to the degree of violation:
[0134] Yellow warning: The function value is between 90% and 100% of the standard value;
[0135] Orange warning: The function value is between 100% and 110% of the standard value;
[0136] Red warning: the function value exceeds the proportion of the standard value by 110%;
[0137] The judgment result generates a structured warning signal, which includes timestamp, risk type, warning level, current value and threshold benchmark.
[0138] The following actions are performed according to the early warning trigger signal:
[0139] Transmit the determined decision action to the risk control execution end;
[0140] The warning trigger signal is synchronously sent to the risk monitoring end.
[0141] In this embodiment, a dual-channel early warning output mechanism is implemented:
[0142] Risk control execution end transmission:
[0143] Send the original decision action via HTTPS protocol (two-way certificate authentication);
[0144] Add risk tags to data packets (e.g., "MarketRisk-Alert");
[0145] Example: A risk tag is embedded in a credit decision instruction of RMB 50,000, triggering the approval system to automatically reduce the credit limit by 20%.
[0146] Risk monitoring terminal transmission:
[0147] Push warning details to the monitoring screen in real time via WebSocket;
[0148] Includes heat map generation instructions and disposal suggestions (such as "suspending credit approval for the manufacturing industry");
[0149] A closed-loop feedback loop was simultaneously activated: warning signals were injected into the environment state vector. Credit risk warnings increased the credit risk index by 0.15, and market risk warnings increased market volatility by 0.10. Three consecutive red warnings triggered the model's automatic rollback mechanism.
[0150] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0151] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, including a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data.
[0152] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
Claims
1. An intelligent decision-making system based on reinforcement learning, characterized in that: The system comprises: A threshold generation module is used to extract customer feature vectors and environment state vectors in real time in response to the current credit application flow, and generate a dynamic risk threshold set based on the environment state vectors; A decision action determination module, configured to input the customer feature vector and the environment state vector into a pre-trained policy network and determine a decision action based on an output result of the policy network; A risk function value calculation module, configured to execute the decision action and calculate the function values of multiple risk constraints in real time based on the execution results; a Lagrange multiplier adjustment module, configured to determine deviations by correspondingly comparing the function values of the plurality of risk constraints with the dynamic risk threshold set, and dynamically adjust the Lagrange multipliers associated with each risk constraint according to the deviations; A policy network parameter updating module, configured to update the parameters of the policy network using a primal-dual gradient descent method based on a real-time reward signal fed back after the decision action is executed and a dynamically adjusted Lagrange multiplier; The risk warning module is used to output the decision action to the risk control execution end, and trigger a real-time risk warning when the function value of any risk constraint reaches the preset proportional coefficient of the corresponding dynamic risk threshold.
2. The intelligent decision-making system based on reinforcement learning according to claim 1, characterized in that: The threshold generation module is specifically used to: Extracting real-time risk factors of preset dimensions from the environmental state vector to form a risk factor vector, wherein the real-time risk factors include the real-time default rate of the industry and the change in capital adequacy ratio; Recalling a pre-stored weight matrix and bias vector, wherein the weight matrix and bias vector are generated by learning the historical credit data used for training the policy network, and each column vector in the weight matrix corresponds to a risk type of a plurality of risk constraints; The weight matrix and the bias vector are combined to generate a dynamic risk threshold set corresponding to the multiple risk constraints, wherein the generation formula of the dynamic risk threshold set is: Among them, σ represents the Sigmoid activation function, T j represents the dynamic risk threshold, b j is the j-term component in the bias vector b, F risk represents the risk factor vector; Indicates that the column vector W j Transposed to a row vector, W j represents the column vector corresponding to the j-th risk constraint in the weight matrix W, and · represents the vector dot product operation.
3. The intelligent decision-making system based on reinforcement learning according to claim 2, characterized in that: The decision action determination module is specifically used to: Concatenate the customer feature vector and the environment state vector along the feature dimension to generate a joint feature vector; Inputting the joint feature vector into the policy network and outputting a decision action probability distribution, wherein the dimensions of the decision action probability distribution correspond one-to-one to the discrete options of the decision action; Based on the decision action probability distribution, a decision action is determined based on a maximizing probability selection strategy.
4. The intelligent decision-making system based on reinforcement learning according to claim 3, characterized in that: The risk function value calculation module is specifically used to: Executing the decision action determined based on the maximum probability selection strategy, and obtaining credit business execution result data after the decision action is executed; extracting characteristic indicators related to each risk constraint from the credit business execution result data; Based on the characteristic indicators, according to the preset calculation rules corresponding to each risk constraint, the function values of multiple risk constraints are calculated in real time.
5. The intelligent decision-making system based on reinforcement learning according to claim 4, characterized in that: The Lagrange multiplier adjustment module is specifically used to: For each risk constraint, calculate its function value C j With dynamic risk threshold T j The difference between the two generates a sub-deviation, where the sub-deviation δ j The generation formula is: δ j =C j -T j (j=1,2,…,m); For all sub-deviations, weighted norm aggregation is performed according to the preset weights to generate deviations; Based on the sub-deviations corresponding to each risk constraint, the Lagrange multipliers associated with each risk constraint are dynamically adjusted according to a preset Lagrange multiplier adjustment rule. The preset Lagrange multiplier adjustment rule is to increase or decrease the Lagrange multiplier according to a preset proportional coefficient based on the size and positive or negative sign of the sub-deviation.
6. The intelligent decision-making system based on reinforcement learning according to claim 5, characterized in that: The policy network parameter updating module is specifically used to: Obtaining a real-time reward signal fed back after the decision action is executed, and calculating an original optimization target gradient based on the real-time reward signal and current parameters of the policy network; For each risk constraint, the gradient of the constraint penalty term is calculated by combining the updated Lagrange multiplier and the function value; The constraint penalty term gradient and the original optimization target gradient are combined to update the parameters of the policy network through the primal-dual gradient descent method.
7. The intelligent decision-making system based on reinforcement learning according to claim 6, characterized in that: The risk warning module is specifically used to: The function value C of each risk constraint calculated j Corresponding dynamic risk threshold T j , calculate the proportion of the target value T j ·K j , where K j Indicates the preset scale factor; When any function value C j Satisfy C j ≥T j ·K j When , a warning trigger signal for the risk constraint is generated; The following actions are performed according to the early warning trigger signal: Transmit the determined decision action to the risk control execution end; The warning trigger signal is synchronously sent to the risk monitoring end.
Citation Information
Cited By
Credit risk intelligent management and control method based on deep learning and massive data
CN122549942A