A method, system and device for identifying and responding to risks of a treasury of an enterprise
By preprocessing and feature extraction of multi-source time-series data in corporate treasury scenarios, combined with clustering and Markov state transition matrix simulation, and using reinforcement learning to generate real-time risk response strategies, the problems of static and isolated risk identification and lagging response decisions in existing technologies are solved, realizing the foresight and autonomy of risk management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSPUR GENERSOFT CO LTD
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-05
AI Technical Summary
Existing corporate treasury risk management technologies suffer from problems such as static and isolated risk identification, lack of foresight, delayed risk response decisions, and reliance on manual labor, making it difficult to meet the needs of enterprises for efficient risk management during the digital transformation phase.
By preprocessing multi-source time-series data, risk state vectors with global context information are extracted using an encoder. Risk simulation is then performed using clustering algorithms and Markov state transition matrices to generate risk evolution paths. Finally, real-time risk response strategies are generated through a reinforcement learning model.
It enables proactive risk identification and autonomous response, allowing for advance risk simulation, generation of optimal strategies, significant reduction in response time, and enhancement of corporate risk resistance capabilities.
Smart Images

Figure CN121120290B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, and in particular relates to a method, system and equipment for identifying and responding to corporate treasury risks. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] In the field of corporate financial management, treasury risk management, as a core component of ensuring corporate capital security and maintaining stable financial operations, primarily focuses on the identification, early warning, and response to various risks that may be encountered in corporate financial activities, such as operational risks, liquidity risks, and market risks. With the expansion of corporate operations and the increasing complexity of the market environment, the need for more refined and forward-looking treasury risk management is becoming increasingly prominent. Although existing technologies have built a rule-based risk monitoring and threshold early warning system, providing basic support for treasury risk management, there are still many limitations in practical applications that urgently need to be addressed, making it difficult to meet the needs of enterprises for efficient risk management during their digital transformation.
[0004] Existing treasury risk management technologies generally rely on rule-based or experience-based risk thresholds for risk identification and early warning. For example, in market risk management, when exchange rate fluctuations exceed a preset percentage, or in liquidity risk management, when a company's cash flow balance falls below a set safety line, the system can automatically trigger an alarm. While such mechanisms can achieve preliminary monitoring of some clearly defined risks, they suffer from the core problem of static and isolated risk identification mechanisms lacking foresight. On the one hand, treasury systems accumulate a large amount of multi-source information during operation, covering financial reports, transaction data, investment and financing records, historical risk records, etc. However, existing technologies fail to effectively mine and utilize the value of this information. They cannot improve risk identification efficiency through multi-source data correlation analysis, nor can they cope with the dynamic correlation and coupling effects between different risk factors, such as the mutual influence between exchange rate fluctuations, interest rate adjustments, and credit risk. On the other hand, facing the ever-changing risk scenarios in complex market environments, existing static threshold mechanisms cannot achieve risk simulation and adaptation. They can only identify single risks that are clearly defined and easily quantifiable, making it difficult to predict potential risks that may be triggered by changes in the macroeconomic environment or extreme market scenarios in advance, resulting in risk identification always being in a passive response state.
[0005] Meanwhile, existing technologies exhibit significant lag in risk response decision-making. Their core function is limited to risk reporting, meaning they can only transmit risk warning signals but cannot provide managers with timely and actionable risk response strategies. After a warning signal is issued, the entire process, from in-depth analysis of the risk's causes and development of a response plan based on the company's actual financial situation, to the final implementation of the strategy, heavily relies on manual operation by managers. This requires not only manually integrating multi-dimensional financial data to assess the scope of the risk's impact but also relying on experience to select the optimal response path, resulting in a significantly prolonged decision-making cycle and low response efficiency. In today's rapidly changing market environment, key financial indicators such as exchange rates and interest rates can fluctuate dramatically in a short period, and the window for risk management is often fleeting. The lag in existing technologies makes it difficult for companies to seize the best response opportunity, potentially exacerbating financial losses caused by risks and adversely affecting the stability of the company's cash flow and long-term investment and financing plans.
[0006] In summary, the problems existing in current corporate treasury risk management technologies, such as static and isolated risk identification, lack of foresight, delayed risk response decisions, and reliance on manual labor, are issues that urgently need to be addressed. Summary of the Invention
[0007] To overcome the shortcomings of the prior art, the present invention provides a method, system and equipment for corporate treasury risk identification and response, realizing forward-looking risk perception and autonomous risk response.
[0008] To achieve the above objectives, the present invention adopts the following technical solution:
[0009] In a first aspect, the present invention provides a method for identifying and responding to corporate treasury risks, including:
[0010] The multi-source time-series data acquired in the enterprise treasury scenario are preprocessed to obtain a multivariate time-series matrix;
[0011] The encoder is used to extract features from the multivariate time series matrix to obtain a risk state vector that incorporates global context information;
[0012] Clustering algorithms are used to perform cluster analysis on the risk state vector to obtain the risk scenario label to which the risk state vector belongs;
[0013] Using historical risk scenario labels as the discrete state space, a Markov state transition matrix is generated based on the statistical analysis of historical risk scenario sequences. Risk simulation is performed starting from the initial risk state, generating multiple risk evolution paths.
[0014] The reinforcement learning model is pre-trained based on the generated risk evolution path, and the pre-trained reinforcement learning model is used to generate risk response strategies for real-time risk states.
[0015] Secondly, the present invention provides a corporate treasury risk identification and response system, comprising:
[0016] The preprocessing module is configured to preprocess the acquired multi-source time-series data in the enterprise treasury scenario to obtain a multivariate time series matrix.
[0017] The extraction module is configured to: use the encoder to extract features from the multivariate time series matrix to obtain a risk state vector that incorporates global context information;
[0018] The clustering module is configured to: perform cluster analysis on the risk state vector using a clustering algorithm to obtain the risk scenario label to which the risk state vector belongs;
[0019] The simulation module is configured to: use historical risk scenario labels as the discrete state space, generate Markov state transition matrices based on the historical risk scenario sequence statistics, perform risk simulation starting from the initial risk state, and generate multiple risk evolution paths.
[0020] The response module is configured to: pre-train the reinforcement learning model based on the generated risk evolution path, and use the pre-trained reinforcement learning model to generate risk response strategies for real-time risk states.
[0021] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0022] The above one or more technical solutions have the following beneficial effects:
[0023] In this invention, multi-source time-series data acquired in a corporate treasury scenario are preprocessed to obtain a multivariate time-series matrix. An encoder is used to extract features from the multivariate time-series matrix, resulting in a risk state vector that integrates global contextual information. Deep fusion of multi-source data addresses the problems of fragmented multi-source information and the inability to identify risk coupling effects inherent in traditional technologies. Clustering algorithms are used for cluster analysis to match the risk state vector corresponding to new data to the relevant risk scenario, upgrading from single-point alerts to high-dimensional scenario recognition and covering complex risk patterns that traditional technologies cannot reach. Based on Markov state transition matrices and path simulation, multiple risk evolution paths are generated, allowing for the pre-simulation of various possible outcomes of the current scenario before the risk occurs, solving the problem that traditional technologies can only provide passive alerts and cannot make forward-looking predictions. Reinforcement learning enables autonomous and precise risk response. This invention pre-simulates risk paths, shifting response nodes from after the risk occurs to before it occurs, and generating optimal strategies, significantly shortening response time, adapting to multiple risk scenarios, and enhancing the corporate treasury's ability to withstand market fluctuations.
[0024] In this invention, a fractal temporal correlation attention layer is introduced into the encoder. This layer can effectively capture the multi-scale self-similarity characteristics existing in multi-source historical time-series data and generate a fractal prior matrix accordingly. By introducing a gating fusion mechanism, the attention score and the fractal prior matrix are adaptively fused to generate the final attention weight matrix. By introducing fractal priors, it is possible to capture recurring multi-scale self-similarity patterns in historical data, thereby more accurately identifying and characterizing different types of typical risk scenarios.
[0025] In this invention, a clustering method based on a weighted mechanism of maximum correlation entropy improves robustness to noise and outliers.
[0026] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0027] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0028] Figure 1 This is a flowchart of a corporate treasury risk identification and response method according to Embodiment 1 of the present invention. Detailed Implementation
[0029] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0030] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0031] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0032] Example 1
[0033] This embodiment discloses a method for identifying and responding to corporate treasury risks, including:
[0034] Step 1: Preprocess the multi-source time-series data obtained from the enterprise treasury scenario to obtain a multivariate time series matrix;
[0035] Step 2: Use the encoder to extract features from the multivariate time series matrix to obtain a risk state vector that incorporates global context information;
[0036] Step 3: Use a clustering algorithm to perform cluster analysis on the risk state vector to obtain the risk scenario label to which the risk state vector belongs;
[0037] Step 4: Using historical risk scenario labels as the discrete state space, generate a Markov state transition matrix based on the historical risk scenario sequence statistics, and perform risk simulation starting from the initial risk state to generate multiple risk evolution paths.
[0038] Step 5: Pre-train the reinforcement learning model based on the generated risk evolution path, and use the pre-trained reinforcement learning model to generate risk response strategies for real-time risk states.
[0039] In this embodiment, the multi-source time-series data acquired in the enterprise treasury scenario is preprocessed to obtain a multivariate time-series matrix. An encoder is used to extract features from the multivariate time-series matrix to obtain a risk state vector that integrates global context information. Deep fusion of multi-source data solves the problems of fragmented multi-source information and inability to identify risk coupling effects in traditional technologies. Clustering algorithms are used for cluster analysis to match the risk state vector corresponding to new data to the relevant risk scenario, achieving an upgrade from single-point alarm to high-dimensional scenario recognition, covering complex risk patterns that traditional technologies cannot reach. Based on Markov state transition matrices and path simulation, multiple risk evolution paths are generated, pre-simulating various possible outcomes of the current scenario before the risk occurs, solving the problem that traditional technologies can only passively alarm and cannot proactively predict. Reinforcement learning enables autonomous and precise risk response. This embodiment pre-simulates risk paths, shifting response nodes from after the risk occurs to before it occurs, and generating optimal strategies, significantly shortening response time, adapting to multiple risk scenarios, and improving the enterprise treasury's ability to withstand market fluctuations.
[0040] The following is combined Figure 1 This embodiment provides a detailed explanation of a corporate treasury risk identification and response method.
[0041] Step 1: Preprocess the multi-source time series data obtained from the enterprise treasury scenario to obtain a multivariate time series matrix.
[0042] In this embodiment, the multi-source time-series data in the corporate treasury scenario includes, but is not limited to: internal cash flow series, investment and financing record series, changes in credit spreads of major counterparties, daily spot exchange rates such as USD / CNY, interest rates such as SHIBOR and LIBOR, and commodity prices such as Brent crude oil.
[0043] Preprocessing of multi-source time-series data in corporate treasury scenarios includes missing value imputation, standardization, and unified timestamps to form a multivariate time series matrix. Where T is the time step and N is the dimension of historical data.
[0044] Step 2: Use the encoder to extract features from the multivariate time series matrix to obtain a risk state vector that incorporates global context information.
[0045] In this embodiment, an encoder incorporating a fractal temporal correlation attention layer is used to extract features from a multivariate time series matrix to obtain a risk state vector that integrates global contextual information, specifically:
[0046] Multiscale fractal features of multivariate time series matrices are extracted using the multiscale Hearst exponent.
[0047] Based on the correlation strength between the extracted multi-scale fractal features, a fractal prior matrix is constructed.
[0048] The attention score and dynamic gating value are calculated for the multivariate time series matrix. The attention score, dynamic gating value and fractal prior matrix are fused through a gating fusion mechanism to obtain the dynamic attention weight.
[0049] By using dynamic attention weights for weighted summation, a risk state vector that incorporates global contextual information is obtained.
[0050] As a specific implementation, the encoder employs an improved Transformer encoder, which replaces the self-attention layer of the Transformer encoder with a fractal temporal correlation attention layer. The improved Transformer encoder is used to describe multi-scale self-similarity in multi-source historical time-series data, dynamically calculates the correlation weights between data indicators at different time sections, and transforms the dynamically calculated correlation weights into a fixed-size risk state vector that incorporates global contextual information.
[0051] Specifically, for multivariate time series matrices Here, T represents the time dimension, and N represents the data feature dimension. To achieve multi-scale extraction, a set of sliding windows covering different time granularities needs to be defined first, i.e., the time scale set S = [s1, s2, ..., s]. E For example, S = [4, 8, 16, 32] represents sliding time windows of different lengths, such as 4 days, 8 days, 16 days, and 32 days. The time scale is determined by combining the actual cyclical characteristics of the company's treasury risk. For example, short-term transaction risks correspond to a small scale, while long-term investment and financing risks correspond to a large scale, avoiding scales that are too dense or too sparse.
[0052] For multivariate time series matrices For each feature dimension, at each preset time scale, extract all continuous subsequences that satisfy the window length, specifically as follows:
[0053] For matrix Each of the N feature dimensions is labeled as feature dimension d;
[0054] For each feature dimension d, iterate through each scale s in the time scale set S. e For time-series data with feature dimension d, according to scale s e The window length is determined by sliding from the starting point of the time sequence to extract all continuous subsequences.
[0055] The Hearst exponent is a core fractal characteristic indicator for measuring the long-term memory and trend persistence of time series data. The Hearst exponent needs to be calculated for each extracted continuous subsequence. The Hearst exponent can be estimated by using rescaled range analysis (R / S Analysis) or detrended fluctuation analysis (DFA).
[0056] For each feature dimension d and each scale s e Extract all sequences of length s from the sequence of this feature dimension. e For each continuous subsequence, the Hearst exponent is estimated by applying rescaled range analysis (R / S Analysis) or other computational methods such as detrended volatility analysis (DFA). For each feature dimension, a fractal feature vector is obtained. All feature dimensions constitute a fractal feature matrix. N represents the feature dimension, and E represents the time scale.
[0057] The above-mentioned multi-scale fractal features are transformed into priors of the fundamental correlation strength between features, specifically including:
[0058] Reshaping the projection: transforming the fractal feature matrix The fractal representation is obtained by feeding the data into a lightweight fractal encoding network, such as a two-layer MLP, and projecting it onto a new P-dimensional feature space. .
[0059] Calculate the prior association: for a feature pair, i.e., fractal feature i and fractal feature j, the strength of the fractal prior association between them. It is obtained by calculating the similarity of their fractal representations. This embodiment uses cosine similarity or dot product followed by scaling: or The superscript T indicates transpose. fractal representation fractal characteristics i, fractal representation The fractal characteristic j.
[0060] Perform this operation on all feature pairs to obtain a fractal prior matrix. Each element in the fractal prior matrix B This represents the intrinsic fundamental correlation strength between fractal feature i and fractal feature j, i.e., the similarity between fractal feature i and fractal feature j.
[0061] Multivariate time series matrix The query is obtained through a learnable linear transformation. ), Key ( Value ); Calculate the standard scaled dot product attention score matrix This indicates that the attention weights focus on the relationships between different time steps. Indicates the dimension of the query and key. The dimension representing the value.
[0062] To adaptively balance the contributions of fractal priors and data features, a gating mechanism is introduced to generate a dynamically gated scalar g∈ [0, 1]. This is applied to multivariate time series matrices. A context vector is obtained by performing global average pooling. The context vector c is passed through a fully connected layer with a sigmoid activation function to generate a gated scalar g.
[0063] The final attention matrix A is calculated using the following formula:
[0064]
[0065] The gating factor g is dynamic. If the current sequence pattern closely follows fractal rules, such as a stable trend, g may approach 0, and the model relies more on the fractal prior matrix B. If the sequence exhibits unprecedented abnormal fluctuations, i.e., the fractal rules are broken, g may approach 1, and the model relies more on the immediate correlations S learned from the data.
[0066] The output of the fractal temporal correlation attention layer is obtained by weighting Value(V) using the fused attention weight matrix A, which is the same as the attention of the standard Transformer block: The output of this Transformer block is the weighted average of the associated weights of the data indicators over the entire time window, which is the risk state vector within that time window.
[0067] Step 3: Use clustering algorithms to perform cluster analysis on the risk state vectors to obtain the risk scenario labels to which the risk state vectors belong.
[0068] As a specific implementation method, the k-means clustering method based on the weighted mechanism of maximum correlation entropy is used for cluster analysis. The specific steps are as follows:
[0069] Step 1: Set the number of clusters From the set of risk state vectors, k historical risk state vectors are randomly selected as initial cluster centers, and for each historical risk state vector... Initialize weights This indicates that all sample points have the same importance in the initial state.
[0070] Step 2: Calculate the weighted Euclidean distance from each historical risk state vector to each cluster center:
[0071]
[0072] in, Indicates the first j Cluster centers, This represents the vector of the i-th historical risk state. This represents the weight corresponding to the i-th historical risk state vector.
[0073] Each historical risk state vector is assigned to the cluster to which the cluster center with the nearest weighted distance belongs.
[0074] Step 3: For each cluster, recalculate the cluster centers using a weighted average:
[0075]
[0076] At this point, the influence of outlier points (with small weights) on the new center point location decreases.
[0077] Update the weights of each historical risk state vector according to the maximum correlation entropy criterion:
[0078]
[0079] Where σ is the kernel width parameter, which controls the sensitivity to outliers; ) represents the historical risk state vector The center vector of the current cluster. This formula assigns smaller weights to outliers (far from the center) and larger weights to normal points.
[0080] Step 4: Repeat steps 2-3 until convergence meets one of the following convergence conditions: 1. The positions of the cluster centers no longer change significantly, i.e., the movement distance of all center points is less than a preset threshold; 2. The preset maximum number of iterations is reached; 3. The cluster assignment relationship no longer changes. After clustering, the clusters are semantically defined, and each cluster is defined as a typical risk scenario pattern. The center vector of the cluster is the corresponding risk scenario label, and the risk scenario labels corresponding to multiple clusters constitute a risk scenario label library.
[0081] For newly accessed real-time multi-source time series data, the improved Transformer encoder is first used to generate a real-time risk state vector. The cosine similarity or Euclidean distance between the real-time risk state vector and all labels in the risk scenario label library, i.e., the target cluster center vector, is calculated. The risk scenario label with the highest similarity or the closest distance is selected as the label to which the current real-time risk state vector belongs, and the matching confidence is output.
[0082] Traditional k-means clustering uses Euclidean distance, which is highly sensitive to outliers and noise. Risk data often contains outliers caused by extreme events such as flash crashes or sudden policy changes. These outliers distort the location of cluster centers, leading to inaccurate risk scenario labels that fail to represent most common scenarios. This embodiment employs a weighted k-means clustering method based on maximum relevance entropy, which improves robustness to noise and outliers.
[0083] As an optional implementation, for any historical moment, after calculating the historical risk state vector, if there is a risk event record at the current moment, such as over-budget payment, then the risk scenario label corresponding to the risk state vector is matched with the risk event record.
[0084] Specifically, the overall risk metric is calculated based on the risk state vector, which includes a deep learning binary classification model. The input is the risk state vector, and the output is a classification confidence score between 0 and 1. The closer the value is to 1, the higher the risk; the closer the value is to 0, the lower the risk, i.e., the higher the health level. The training data is the risk state vector, and the output is whether the risk scenario label to which the risk state vector belongs matches a risk event record. If the risk scenario label has a risk event record, the output is 1; otherwise, it is 0.
[0085] In this embodiment, an improved clustering algorithm is used to identify and segment recurring typical risk scenarios from historical risk state vectors. Finally, a risk scenario label is assigned to each risk vector. The risk level is quantitatively evaluated using a deep learning binary classification model. That is, based on the correlation between the risk scenario label output by the clustering algorithm and historical risk event records, a risk metric value in the range of 0-1 is calculated for each risk vector.
[0086] Step 4: Using historical risk scenario labels as the discrete state space, generate a Markov state transition matrix based on the historical risk scenario sequence statistics, and perform risk simulation starting from the initial risk state to generate multiple risk evolution paths.
[0087] In this embodiment, each historical risk scenario label is defined as a discrete state of a Markov chain, constructing a discrete state space; the historical risk scenario sequence is traversed, and the states are statistically analyzed. i Transition to state j The frequency and state i The total frequency of occurrence is used to calculate the number of occurrences from the state. i to state j The transition probabilities are used to construct a Markov state transition matrix; starting from the initial risk state, multiple risk evolution paths are generated through Monte Carlo simulation.
[0088] In this embodiment, starting from the initial risk state, multiple risk evolution paths are generated through Monte Carlo simulation. Specifically, starting from the initial risk state, the transition probability distribution corresponding to the current risk state is queried; based on the transition probability distribution corresponding to the initial risk state, random sampling is performed to determine the next risk state; the next risk state is recorded as the first step state in the path, and the next risk state is taken as the new current risk state. The steps of querying the transition probability distribution and random sampling are repeated until the multi-step simulation is completed, resulting in a risk evolution path; the process of generating a single risk evolution path is repeated multiple times to obtain multiple risk evolution paths.
[0089] As a specific implementation method, each risk scenario label is defined as a discrete state of a Markov chain. Assuming five typical risk scenario labels are obtained through clustering: X1 - mild imported inflation; X2 - monetary policy tightening; X3 - industry-wide credit crisis; X4 - stable corporate financial condition; X5 - commodity-driven volatility, the discrete state space can be represented as M = (X1, X2, X3, X4, X5). The matching results of historical risk state vectors and risk scenario labels are extracted and arranged chronologically to form a historical risk scenario sequence, for example: [X4, X4, X1, X1, X2, ...]. The entire historical risk scenario sequence is traversed, and the frequency of transitions from risk state i to risk state j is counted. For example, the number of times "X1 - mild imported inflation" occurs, and the number of times the state immediately following the transition to "X2 - monetary policy tightening" is counted. Based on the above statistical results, the frequency is converted into Markov state transition probabilities. Each element in the Markov state transition matrix... The calculation formula is ,in: Let i be the number of times a state transitions from state i to state j. Let i be the total number of times state i occurs.
[0090] For example, the Markov state transition matrix is shown in Table 1.
[0091] Table 1:
[0092]
[0093] Starting from the initial risk state, multiple risk evolution paths are generated through Monte Carlo simulation, specifically:
[0094] Initial risk status The number of simulation steps is T, for example, simulating 30 days; the number of simulation paths is N, for example, 1000 paths, to cover various low-probability but high-impact "tail risks".
[0095] From the initial state Initially, for step t, query the current state in the Markov state transition matrix. The corresponding row probability distribution, based on which the next state is randomly sampled. ;Will Record the new state on the path and use it as the new current state; repeat the above process, t from 1 to T, until T steps of simulation are completed, and a complete risk evolution path is obtained.
[0096] Repeat the single-path simulation process described above independently N times to obtain N possible time series paths. For example, if N=3, the following randomized simulation paths may be obtained: Path 1: [X1, X1, X1, X2, X2, X2, X3, ...]; Path 2: [X1, X1, X4, X4, X4, X1, X1, ...]; Path 3: [X1, X2, X2, X5, X3, X3, X3, ...].
[0097] This embodiment does not simply define risk using rules or fixed thresholds, but rather summarizes and generalizes core risk scenarios through intelligent algorithms. When new market data emerges, the system no longer checks whether a single indicator exceeds a threshold, but determines which risk scenario the current financial situation of the enterprise falls into, thereby achieving high-dimensional and dynamic risk perception. The scenario simulation in this method is not based on economic forecasts, but rather on scenario-based simulations of actual data and risk records in the enterprise's treasury. Intelligent algorithms quantify potential risk paths, serving as the training basis for the following proactive intelligent strategies.
[0098] Step 5: Pre-train the reinforcement learning model based on the generated risk evolution path, and use the pre-trained reinforcement learning model to generate risk response strategies for real-time risk states.
[0099] In this embodiment, risk scenario labels are randomly selected from a risk scenario label library, and the typical risk state vectors corresponding to the selected risk scenario labels are used as the initial state for each training round. Starting from the initial state, a training sample set covering multiple risk scenarios is constructed based on the generated risk evolution path. A deep reinforcement learning algorithm is used to enable the reinforcement learning agent to learn through trial and error in the simulated scenario corresponding to the risk evolution path, completing the offline pre-training of the reinforcement learning model. Starting from the current risk state vector, online training samples are constructed by combining the action execution results of real-time environmental feedback. The offline pre-trained reinforcement learning model is fine-tuned online using the online training samples. The current risk state vector is input into the online fine-tuned reinforcement learning model to obtain the risk response strategy for the current risk state.
[0100] As a specific implementation method, a policy network that can automatically output optimal hedging instructions based on real-time risk scenarios is continuously trained by allowing the agent to learn and experiment in a simulated environment. The specific technical solution includes the following framework and parameter settings:
[0101] State (S): Represented by a risk state vector. At time step t, the risk state vector is a high-dimensional vector, directly derived from the output of the improved Transformer encoding, representing a representation of the current complex risk situation in the market, for example, a 256-dimensional vector.
[0102] Action (A): Defined in a high-dimensional continuous action space, it represents the position adjustment amount on various hedging instruments, i.e., the intelligent hedging strategy. For example, the output action [0.5, -0.3] might represent increasing a 50% long position in foreign exchange forward contracts and decreasing an interest rate swap position by 30%. Actions are hedging strategies continuously trained by the intelligent strategy engine, intelligently responding to the input risk state.
[0103] Reward function (Reward, R): The reward for policy training is set as follows:
[0104]
[0105] Wherein, ΔC represents the reduction in the overall risk metric after hedging, with risk reduction being a positive reward. The overall risk metric is obtained through a deep learning binary classification model. Cost represents transaction costs such as commissions and impact costs. Penalty represents a significant penalty for violating the company's custom risk control rules, such as excessive leverage. λ and μ are hyperparameters of the system, used to adjust the company's sensitivity to transaction costs and risk control rule penalties. Larger parameters lead the agent to be more inclined to choose lower-cost and milder trading strategies, while avoiding dangerous actions that may lead to violations, thus ensuring strategy compliance.
[0106] The reinforcement learning training architecture consists of two parts: pre-training and online learning. The specific implementation method is as follows:
[0107] Offline pre-training: In order to enhance the generalization ability of reinforcement learning models and avoid overfitting to a certain market scenario that has already been experienced, in the offline pre-training stage, the initial state of each training episode is not from real-time data, but a scenario risk label is randomly selected from the risk scenario label library, and then the typical risk state vector corresponding to the label is matched as the starting point of the initial state. In each subsequent step, the training is carried out by simulating the risk evolution path, so that the agent is trained through various possible market situation scenarios.
[0108] Online fine-tuning: The pre-trained reinforcement learning model is deployed to a real-time production environment for online learning, connected to a real-time data stream. Starting from the current risk state vector, online training samples are constructed by combining the action execution results fed back from the real-time environment. An incremental update strategy is adopted. After the agent outputs an action, the risk changes after the action are monitored in real time to see if they are compliant. Real-time rewards are calculated, and the policy network is fine-tuned based on the real reward feedback, so that it can continuously adapt to the latest market structure that has not appeared in the pre-training, thereby achieving capability evolution.
[0109] This embodiment improves the Transformer encoder by using a fractal temporal correlation attention layer. It can extract multi-scale self-similar features from multi-source time-series data, including internal cash flow and investment and financing records, as well as external exchange rates (USD / CNY), interest rates (SHIBOR), and commodity prices. Compared with traditional isolated indicator monitoring, it can identify complex risks with multiple coupled factors and avoid risk omissions due to data fragmentation. For common extreme outliers in risk data, such as policy changes and market flash crashes, it adopts k-means clustering based on maximum correlation entropy weighting. By dynamically updating the weights of data points, it avoids outliers from distorting the cluster centers and ensures that each risk scenario label can truly represent a typical risk pattern. Compared with traditional k-means, the clustering results are more robust, and the accuracy of subsequent real-time scenario matching is higher. Based on Markov chains, a risk scenario transition matrix is constructed. Combined with Monte Carlo simulation, thousands of risk evolution paths are generated, which can quantify low-probability but high-impact tail risks. This breaks through the limitation of traditional warnings that only warn of risks that have already occurred, allowing enterprises to anticipate the possibility of risks before they actually occur and prepare contingency plans in advance.
[0110] This embodiment of the reinforcement learning framework is divided into two parts: offline pre-training and online fine-tuning. In the offline stage, the model is trained with a full range of risk scenarios, covering historical and simulated scenarios to ensure the model's generalization ability. In the online stage, the model is connected to real-time data streams and continuously fine-tunes the strategy based on real market feedback. Even if the market structure changes, such as the new policy of interest rate liberalization, the model can adapt and iterate to avoid the strategy becoming outdated and ineffective.
[0111] Example 2
[0112] The purpose of this embodiment is to provide a corporate treasury risk identification and response system, including:
[0113] The preprocessing module is configured to preprocess the acquired multi-source time-series data in the enterprise treasury scenario to obtain a multivariate time series matrix.
[0114] The extraction module is configured to: use the encoder to extract features from the multivariate time series matrix to obtain a risk state vector that incorporates global context information;
[0115] The clustering module is configured to: perform cluster analysis on the risk state vector using a clustering algorithm to obtain the risk scenario label to which the risk state vector belongs;
[0116] The simulation module is configured to: use historical risk scenario labels as the discrete state space, generate Markov state transition matrices based on the historical risk scenario sequence statistics, perform risk simulation starting from the initial risk state, and generate multiple risk evolution paths.
[0117] The response module is configured to: pre-train the reinforcement learning model based on the generated risk evolution path, and use the pre-trained reinforcement learning model to generate risk response strategies for real-time risk states.
[0118] In further embodiments, the following is also provided:
[0119] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0120] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0121] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0122] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0123] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0124] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0125] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for identifying and responding to corporate treasury risks, characterized in that, include: The multi-source time-series data acquired in the enterprise treasury scenario are preprocessed to obtain a multivariate time-series matrix; An encoder incorporating a fractal temporal correlation attention layer is used to extract features from a multivariate time series matrix, resulting in a risk state vector incorporating global contextual information. Specifically, multi-scale fractal features are extracted from the multivariate time series matrix using a multi-scale Hearst exponent. A fractal prior matrix is constructed based on the correlation strength between the extracted multi-scale fractal features. Attention scores and dynamic gating values are calculated for the multivariate time series matrix, and the attention scores, dynamic gating values, and fractal prior matrix are fused through a gating fusion mechanism to obtain dynamic attention weights. We use dynamic attention weights to perform weighted summation to obtain a risk state vector that incorporates global contextual information; Clustering algorithms are used to perform cluster analysis on the risk state vectors to obtain the risk scenario labels to which the risk state vectors belong, specifically: Step 1: Set the number of clusters, randomly select k initial cluster centers from the historical risk state vector set, and initialize the weights of each historical risk state vector in the historical risk state vector set; Step 2: Calculate the distance from each historical risk state vector in the historical risk state vector set to each cluster center, and assign each historical risk state vector to the cluster to which the nearest cluster center belongs; Step 3: For each cluster, recalculate the cluster centers using a weighted average, and update the weights of each historical risk state vector according to the maximum relevance entropy criterion. Where σ is the kernel width parameter, which controls the sensitivity to outliers; ) represents the historical risk state vector The center vector of the current cluster; Step 4: Repeat steps 2-3 until convergence is achieved and one of the following convergence conditions is met:
1. The positions of the cluster centers no longer change significantly, i.e., the movement distance of all center points is less than a preset threshold; 2. The preset maximum number of iterations is reached; 3. The cluster assignment relationship no longer changes; After clustering, the semantic definition of the clusters is performed, defining each cluster as a typical risk scenario pattern. The center vector of the cluster is the corresponding risk scenario label, and the risk scenario labels corresponding to multiple clusters constitute a risk scenario label library; Using historical risk scenario labels as the discrete state space, a Markov state transition matrix is statistically generated based on the historical risk scenario sequence. Risk simulation is performed starting from the initial risk state, generating multiple risk evolution paths. Specifically, each historical risk scenario label is defined as a discrete state of a Markov chain, constructing a discrete state space; the historical risk scenario sequence is traversed, and the risk states are statistically analyzed. i Shift to a risky state j Frequency and risk status i The total frequency of occurrence is used to calculate the risk status. i to risk status j The transition probabilities are used to construct a Markov state transition matrix; starting from the initial risk state, multiple risk evolution paths are generated through Monte Carlo simulation; The reinforcement learning model is pre-trained based on the generated risk evolution path, and the pre-trained reinforcement learning model is used to generate risk response strategies for real-time risk states.
2. The method for identifying and responding to corporate treasury risks as described in claim 1, characterized in that, Starting from the initial risk state, multiple risk evolution paths are generated through Monte Carlo simulation, specifically: Starting from the initial risk state, query the transition probability distribution corresponding to the current risk state; Random sampling is performed based on the transition probability distribution corresponding to the initial risk state to determine the next risk state; The next risk state is recorded as the first step state in the path, and the next risk state is the new current risk state. The steps of querying the transition probability distribution and random sampling are repeated until the multi-step simulation is completed and a risk evolution path is obtained. The process of generating a single risk evolution path is repeated multiple times to obtain multiple risk evolution paths.
3. The method for identifying and responding to corporate treasury risks as described in claim 1, characterized in that, The reinforcement learning model is pre-trained based on the generated risk evolution path. The pre-trained model is then used to generate risk response strategies for the current risk state. Specifically: Risk scenario labels are randomly selected from the risk scenario label library, and the typical risk state vector corresponding to the selected risk scenario label is used as the initial state for each training round. Starting from the initial state, a training sample set covering multiple risk scenarios is constructed based on the generated risk evolution path; Deep reinforcement learning algorithms are used to enable reinforcement learning agents to learn through trial and error in simulated scenarios corresponding to risk evolution paths, thereby completing the offline pre-training of reinforcement learning models. A risk response strategy for the current risk state is generated using a pre-trained reinforcement learning model.
4. A method for identifying and responding to corporate treasury risks as described in claim 1 or 3, characterized in that, The risk response strategy for the current risk state is generated using a pre-trained reinforcement learning model, specifically as follows: Starting with the current risk state vector, online training samples are constructed by combining the action execution results of real-time environmental feedback. The online training samples are then used to fine-tune the offline pre-trained reinforcement learning model online. By inputting the real-time risk state vector into the online fine-tuned reinforcement learning model, a risk response strategy for the current risk state can be obtained.
5. The method for identifying and responding to corporate treasury risks as described in claim 1, characterized in that, Multi-scale fractal feature extraction of multivariate time series matrices is performed using the multi-scale Hearst exponent, specifically as follows: Pre-set a set of time scales based on the actual cyclical characteristics of corporate treasury risks; For each feature dimension of the multivariate time series matrix, at each preset time scale, extract all continuous subsequences that satisfy the window length; The Hearst exponent is calculated for each extracted continuous subsequence, and the Hearst exponents of each feature dimension at all time scales are integrated to form the fractal feature vector of the feature dimension. By integrating the fractal feature vectors of all feature dimensions, multi-scale fractal features are obtained.
6. A corporate treasury risk identification and response system, characterized in that, include: The preprocessing module is configured to preprocess the acquired multi-source time-series data in the enterprise treasury scenario to obtain a multivariate time series matrix. The extraction module is configured to: extract features from a multivariate time series matrix using an encoder that incorporates a fractal temporal correlation attention layer to obtain a risk state vector that integrates global context information; specifically: extract multi-scale fractal features from the multivariate time series matrix using a multi-scale Hearst exponent; construct a fractal prior matrix based on the correlation strength between the extracted multi-scale fractal features; calculate attention scores and dynamic gating values for the multivariate time series matrix; and fuse the attention scores, dynamic gating values, and fractal prior matrix through a gating fusion mechanism to obtain dynamic attention weights. We use dynamic attention weights to perform weighted summation to obtain a risk state vector that incorporates global contextual information; The clustering module is configured to: perform cluster analysis on the risk state vector using a clustering algorithm to obtain the risk scenario label to which the risk state vector belongs, specifically: Step 1: Set the number of clusters, randomly select k initial cluster centers from the historical risk state vector set, and initialize the weights of each historical risk state vector in the historical risk state vector set; Step 2: Calculate the distance from each historical risk state vector in the historical risk state vector set to each cluster center, and assign each historical risk state vector to the cluster to which the nearest cluster center belongs; Step 3: For each cluster, recalculate the cluster centers using a weighted average, and update the weights of each historical risk state vector according to the maximum relevance entropy criterion. Where σ is the kernel width parameter, which controls the sensitivity to outliers; ) represents the historical risk state vector The center vector of the current cluster; Step 4: Repeat steps 2-3 until convergence is achieved and one of the following convergence conditions is met:
1. The positions of the cluster centers no longer change significantly, i.e., the movement distance of all center points is less than a preset threshold; 2. The preset maximum number of iterations is reached; 3. The cluster assignment relationship no longer changes; After clustering, the semantic definition of the clusters is performed, defining each cluster as a typical risk scenario pattern. The center vector of the cluster is the corresponding risk scenario label, and the risk scenario labels corresponding to multiple clusters constitute a risk scenario label library; The simulation module is configured to: use historical risk scenario labels as a discrete state space, statistically generate a Markov state transition matrix based on the historical risk scenario sequence, simulate risks starting from the initial risk state, and generate multiple risk evolution paths. Specifically, each historical risk scenario label is defined as a discrete state of a Markov chain, constructing a discrete state space; the historical risk scenario sequence is traversed, and the risk states are statistically analyzed. i Shift to a risky state j Frequency and risk status i The total frequency of occurrence is used to calculate the risk status. i to risk status j The transition probabilities are used to construct a Markov state transition matrix; starting from the initial risk state, multiple risk evolution paths are generated through Monte Carlo simulation; The response module is configured to: pre-train the reinforcement learning model based on the generated risk evolution path, and use the pre-trained reinforcement learning model to generate risk response strategies for real-time risk states.
7. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-5.
Citation Information
Patent Citations
Multivariable time sequence prediction method based on channel grouping and prediction correction
CN120806249A
Credit risk dynamic feature extraction method based on hierarchical reinforcement learning
CN120894124A