Auditable asset allocation system and method based on time-varying risk preference estimation

By using a state-space model and sequential Monte Carlo method for online Bayesian estimation, combined with dynamic asset allocation optimization and compliance auditability modules, the static lag in risk preference assessment and the unexplainable nature of the decision-making process are resolved. This achieves endogenous linkage between risk preference and asset allocation and full-process auditability, meeting financial regulatory requirements.

CN122048546APending Publication Date: 2026-05-15ZHEJIANG FINANCIAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610117701.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-28
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies suffer from static lag in risk preference assessment, a disconnect between asset allocation and real-time risk preferences, and a lack of auditable and explainable mechanisms in the decision-making process, making it difficult to meet the transparent audit requirements of financial regulation.

Method used

Online Bayesian estimation is performed using a state-space model and a sequential Monte Carlo method. Combined with a dynamic asset allocation optimization module and a compliance and auditability module, the system achieves real-time dynamic tracking of investor risk preferences and end-to-end auditability. The system's adaptability and compliance are improved by introducing a robust optimization and constrained reinforcement learning framework.

Benefits of technology

It enables real-time, dynamic tracking of investor risk preferences, enhances the personalization and adaptability of asset allocation, meets the transparency and traceability requirements of financial regulation, and ensures the robustness and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122048546A_ABST
    Figure CN122048546A_ABST
Patent Text Reader

Abstract

The invention discloses an auditable asset allocation system and method based on time-varying risk preference estimation. The system comprises a data acquisition and preprocessing module; the time-varying risk preference estimation module is used for carrying out online Bayesian estimation on the risk preference time-varying hidden state by adopting a sequential Monte Carlo method based on the state space model; the dynamic asset allocation optimization module is used for taking the estimated value as an endogenous variable to be embedded into an optimization model to solve the optimal weight; and the compliance and auditing module is used for recording the full-link log in a tamper-proof manner and generating an interpretable report. According to the method, closed-loop optimization from dynamic risk preference estimation to asset allocation decision is realized, and auditing performance of the whole process is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of financial technology and artificial intelligence, and in particular to an auditable asset allocation system and method based on time-varying risk preference estimation. Background Technology

[0002] With the development of fintech, robo-advisors and automated asset allocation systems have become important tools in the securities, fund, and bank wealth management sectors. Their core lies in accurately assessing investors' risk preferences and generating personalized asset allocation plans accordingly. However, existing mainstream technologies face significant technical bottlenecks in achieving this goal, particularly in terms of dynamism, interconnectivity, and auditability.

[0003] At the risk preference assessment level, existing solutions are mostly based on static models. For example, they may linearly weight questionnaire answers based on preset weights, or match customer characteristics with fixed asset allocation templates. The results of these methods remain unchanged long-term once generated, failing to reflect the dynamic process of investor psychology changing with market fluctuations and personal financial situations. Although some technologies attempt to use machine learning to analyze customer trading behavior, these models are often trained on static datasets, making it difficult to effectively model the temporal evolution of risk preferences, resulting in assessment lags.

[0004] At the asset allocation decision-making level, even with the introduction of dynamic market factors or target volatility control mechanisms, the optimization models typically treat the client's risk preference as a fixed parameter of external input. This severs the intrinsic link between risk assessment and allocation optimization, failing to establish a feedback channel that automatically and in a closed loop adjusts allocation weights based on the investor's real-time risk status, thus limiting the personalization and adaptability of the strategy.

[0005] Regarding system transparency and compliance, existing solutions generally suffer from a "black box" problem. For example, the decision-making logic of deep learning models used for prediction or recommendation is complex and unexplainable. Furthermore, these systems generally lack a complete and tamper-proof logging mechanism for model inputs, versions, parameters, and outputs, and lack standardized audit traceability interfaces. This makes it difficult for existing systems to meet the increasingly stringent financial regulatory requirements for traceable business processes, explainable decision-making logic, and end-to-end auditability.

[0006] In summary, current technologies have not yet been able to achieve real-time dynamic estimation of risk appetite, nor have they been able to endogenously optimize this estimation in conjunction with asset allocation, nor have they provided a transparent audit framework to support effective financial regulation. Summary of the Invention

[0007] To address the problems of static lag in risk preference assessment, disconnect between asset allocation and real-time risk preference, and lack of auditable and explainable mechanisms in the decision-making process in existing technologies, this invention proposes an auditable asset allocation system and method based on time-varying risk preference estimation.

[0008] The specific technical solution is as follows: An auditable asset allocation system based on time-varying risk preference estimation, comprising:

[0009] The data acquisition and preprocessing module is used to acquire and process investors' static attributes, dynamic behavior data, and market environment data.

[0010] The time-varying risk preference estimation module is connected to the data acquisition and preprocessing module. It is used to model the investor's risk preference parameters as time-varying hidden states based on the state-space model, and to perform online Bayesian estimation of the hidden states using the sequential Monte Carlo method, and output the estimated value of the risk preference parameters and their uncertainty.

[0011] The dynamic asset allocation optimization module is connected to the time-varying risk preference estimation module. It is used to embed the estimated risk preference parameters as endogenous variables into the objective function or reward function of the asset allocation optimization model to solve for the optimal asset allocation weight.

[0012] The compliance and auditability module records the entire system's decision logs in an immutable manner and generates auditable reports containing decision attribution analysis based on model interpretability technology. By constructing a complete technical closed loop of data collection, dynamic estimation, optimized decision-making, and audit evidence storage, it achieves real-time and continuous tracking of investor risk preferences, ensuring that asset allocation strategies can be adjusted in sync with investors' latest psychological states. At the same time, standardized technical means guarantee the traceability and auditability of the entire decision-making process, fundamentally solving the pain points of existing systems such as static evaluation, black-box decision-making, and weak compliance.

[0013] Furthermore, the state-space model includes:

[0014] The state transition equation is used to describe the dynamic evolution of the hidden state vector from the previous time step to the current time step. This evolution process is affected by the state at the previous time step and the current market environment data.

[0015] The observation equation is used to map the current hidden state vector and market environment data into observable investor behavior indicators.

[0016] The latent state vector contains at least a logarithmic form of a parameter representing the investor's risk aversion. Employing a state-space model allows psychological preferences, which are inherently difficult to observe directly, to be estimated indirectly and quantitatively through observable behavioral data. In particular, modeling the risk aversion coefficient in its logarithmic form effectively ensures a positive value constraint for this parameter during the estimation process, enhancing the model's rationality and numerical stability.

[0017] Furthermore, the online Bayesian estimation using the sequential Monte Carlo method is performed according to the following recursive steps:

[0018] Based on the state transition equation, state prediction is performed on the set of particles representing the posterior distribution;

[0019] Obtain the observed value of the investor behavior index at the current time t, and calculate the likelihood of each particle and update its normalized weight based on the observation equation and the observed value.

[0020] When the number of effective particles representing the particle weight dispersion is lower than a preset threshold, system resampling is performed to generate a new set of equally weighted particles.

[0021] Based on the particle set and its weights obtained after resampling, the minimum mean square error estimate and estimated covariance of the risk preference parameters are calculated. The sequential Monte Carlo method is employed, which can efficiently and flexibly handle the nonlinearity and non-Gaussianity in the state-space model, achieving an approximation of the posterior distribution of the hidden states. This provides a dynamic estimation result with uncertainty measurement for risk preference, enhancing the model's adaptability to complex real-world situations and improving estimation reliability.

[0022] Furthermore, the time-varying risk preference estimation module also includes:

[0023] A jump detection unit is used to continuously monitor the investor behavior data sequence through a cumulative sum control graph algorithm to detect sudden changes in its statistical characteristics;

[0024] An adaptive adjustment unit is used to temporarily increase the covariance matrix of the process noise in the state transition equation when a sudden change is detected, thereby enabling the particle set to quickly adapt to structural changes in investor risk preferences. A jump detection mechanism based on statistical process control is introduced, allowing the system to sensitively capture structural changes in investor behavior patterns caused by drastic market fluctuations or significant personal decisions. By adaptively adjusting the process noise, the model converges to a new preference state more quickly, significantly improving the system's response speed and estimation accuracy to sudden changes in preferences.

[0025] Furthermore, the dynamic asset allocation optimization module is configured to: construct an objective function, which incorporates at least the following two items: an expected utility term with the estimated risk preference parameter as the core parameter; and a conditional value-at-risk penalty term or constraint for controlling tail risk; wherein the objective function is used to solve for the optimal asset allocation weights. By directly embedding the dynamically estimated risk preference parameter as a core variable into the optimization objective, a native linkage and endogenous driving force between risk preference and asset allocation decisions is achieved. Simultaneously, the explicit introduction of CVaR risk constraints into the objective enables the optimization model to effectively control tail extreme losses in the portfolio while pursuing maximum utility, achieving a refined balance between return and risk.

[0026] Furthermore, the dynamic asset allocation optimization module is configured to perform the following robust optimization process:

[0027] Construct a set of parameters to characterize the uncertainty in estimating the expected rate of return and the return covariance matrix of an asset;

[0028] For each possible set of parameters in the parameter set, calculate the value of the objective function;

[0029] The asset allocation weights that maximize the objective function under the worst-case parameters on the parameter set are selected as the optimal asset allocation weights. By considering the uncertainty set of expected return and covariance, and seeking the optimal strategy under the worst-case parameters, the robustness and stability of the asset allocation scheme are greatly enhanced. This ensures that the final portfolio weights maintain relatively robust performance in the face of market parameter estimation errors or adverse changes in the market environment, reducing the model's over-reliance on the accuracy of parameter estimation.

[0030] Furthermore, the expected utility term adopts the form of an exponential utility function, specifically:

[0031] ;

[0032] in, R is the risk aversion coefficient estimated at the current moment, and R is the stochastic return of the portfolio.

[0033] Furthermore, the dynamic asset allocation optimization module employs a constrained reinforcement learning framework, which is configured as follows:

[0034] Using a policy network, the state vector containing the estimated risk preference parameters is mapped to an adjustment action for asset weights;

[0035] The policy network is trained by interacting with a simulated environment, wherein the evaluation function used for training includes at least two of the following: one positively correlated with the risk-adjusted return on investment; and the other a penalty term for violations of investment constraints. By leveraging the ability of a reinforcement learning agent to autonomously learn the optimal policy through interaction with the simulated environment, it is possible to handle high-dimensional, continuous state and action spaces. Furthermore, by designing a penalty term in the reward function, it is ensured that the learned policy automatically satisfies various investment constraints, achieving adaptive and intelligent decision-making for seeking long-term optimal configurations under complex constraints.

[0036] Furthermore, the compliance and auditability module specifically includes:

[0037] The structured log unit is used to associate and store the input data fingerprint, model version, key parameters, output results and timestamp of each decision in the form of a hash chain to ensure data integrity and traceability.

[0038] The interpretable reporting unit is used to automatically generate attribution reports after each asset allocation adjustment, and to clarify the main driving factors of the changes in risk preference estimates and asset weight adjustments through feature contribution analysis.

[0039] An auditable asset allocation method based on time-varying risk preference estimation includes the following steps:

[0040] Collect and preprocess multi-dimensional investor data;

[0041] A state-space model is constructed, and the investor's risk preference parameter is defined as a time-varying hidden state;

[0042] By applying the sequential Monte Carlo method, new behavioral data and market data are recursively fused to estimate the latent state online and output the current risk preference parameter estimate.

[0043] Using the estimated risk preference parameters as the core optimization variables, an asset allocation optimization model that integrates expected utility, risk control, and transaction costs is constructed and solved, and the optimal allocation weights are determined.

[0044] The rebalancing operation is performed according to the optimal configuration weights, and the entire chain data of this decision is recorded in the tamper-proof audit log and a decision attribution report is generated.

[0045] The above technical solution has the following advantages or technical effects:

[0046] 1. This invention achieves dynamic, real-time, and quantitative tracking of investor risk preferences through state-space modeling and online particle filtering, solving the problem of lag in traditional static assessment.

[0047] 2. This invention greatly improves personalization and adaptability by directly embedding dynamic risk preference estimates as endogenous variables into the asset allocation optimization model.

[0048] 3. This invention provides a standardized, transparent, and traceable technical solution for intelligent investment advisory systems by constructing a compliance audit module that includes end-to-end tamper-proof logs and interpretable report generation functions, effectively meeting the compliance requirements of financial supervision.

[0049] 4. This invention introduces two advanced optimization frameworks, robust optimization and constrained reinforcement learning, to ensure the robustness, stability and intelligence of asset allocation strategies in the face of market uncertainty and complex constraints. Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the overall architecture of the system of the present invention;

[0051] Figure 2 This is an overall flowchart of the method of the present invention. Detailed Implementation

[0052] To make the technical solution of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0053] Example 1

[0054] like Figure 1 As shown, an auditable asset allocation system based on time-varying risk preference estimation includes: a data acquisition and preprocessing module, a time-varying risk preference estimation module, a dynamic asset allocation optimization module, a compliance and auditability module, and an interaction and display module.

[0055] Data Acquisition and Preprocessing Module: Acquires and processes investors' static attributes, dynamic behavioral data, and market environment data. The collected data includes:

[0056] Customer static information, such as age, occupation, financial status, and historical risk questionnaire results;

[0057] Customer dynamic behavior data, such as transaction frequency, subscription and redemption amount, changes in holdings, order cancellation rate, page dwell time, and stop-loss / average-down behavior;

[0058] Market environment data, such as market volatility indices (e.g., VIX), macroeconomic factors (e.g., Consumer Price Index (CPI), Purchasing Managers' Index (PMI), market sentiment indices, and liquidity indicators;

[0059] Regulatory parameters, such as investor risk classification standards, product risk levels, and regulatory thresholds.

[0060] The data processing workflow includes: desensitizing and encrypting all sensitive data before transmission and storage. Data acquisition adopts a combined event-triggered and periodic sampling model: acquisition is triggered immediately when significant market fluctuations or abnormal customer transactions are detected; otherwise, updates are performed at fixed intervals (e.g., daily). After cleaning, missing value handling, and normalization, the raw data is transformed into standardized vectors through feature engineering. For example, transaction behavior sequences are transformed into statistical features (such as recent means and variances), and market data is transformed into factor exposures.

[0061] Time-varying risk preference estimation module: Based on the state-space model, the risk preference parameters of investors are modeled as time-varying hidden states, and the sequential Monte Carlo method is used to perform online Bayesian estimation of the hidden states, outputting the estimated values ​​of risk preference parameters and their uncertainties.

[0062] State Space Model (SSM) Construction: Define the hidden state vector at time t as... ,in, Arrow-Pratt relative risk aversion coefficient The natural logarithmic form; This is a parameter representing investors' preference for the expected return of their investment portfolio. This represents the investor's sensitivity to portfolio volatility. The constructed state-space model is as follows:

[0063] State transition equation: ;

[0064] The observation equation is as follows: ;

[0065] in: Let t be the external market state feature vector (such as volatility, market return, sentiment index, etc.). As a nonlinear state transition function, it can be adopted in a smooth form with a forgetting factor to characterize the inertial evolution of preferences; For process noise, Let its covariance matrix be ; Let t be the vector of investor behavior indicators observed at time t (such as standardized trading volume, redemption amount, and order cancellation frequency). The observation function maps the implicit risk preference state to observable behavioral expectations. To observe the noise, Let it be its covariance matrix.

[0066] This model treats risk preference parameters as latent time-series variables, capturing their continuous evolution over time and market conditions through state equations, and establishing a mapping between risk preference and observable behavior through observation equations. This structure assumes that risk preference is a continuously changing latent state over time, influenced by market conditions. The combined influence of historical conditions, and through investor behavior. Indirect manifestation.

[0067] Online estimation based on particle filtering: Particle filtering (PF, a sequential Monte Carlo method) is used to perform online Bayesian inference on the above state-space model. The specific steps are as follows:

[0068] a) Initialization: At the initial time t=0, start from the prior distribution Extract N particles from and assign initial weights The number of particles N is dynamically configured based on computing resources and accuracy requirements, typically ranging from 500 to 2000.

[0069] b) Prediction: For t ≥ 1, each particle i is sampled according to the state transition equation:

[0070] ;

[0071] c) Update: When new observation data is received Then, the likelihood value for each particle is calculated. and update the weights The weights are then normalized: .

[0072] d) Resampling: Calculate the effective number of particles .like (The threshold is usually set to N / 2), then the system is resampled, low-weight particles are discarded, high-weight particles are copied, and a new set of equally weighted particles is obtained. .

[0073] e) Output estimation: Calculate the minimum mean square error estimate of the state as the output: ,in, This is the logarithmic estimate of the risk aversion coefficient. It also outputs the estimation uncertainty (covariance): .

[0074] Jump Detection and Adaptive Adjustment: To further improve the model's response speed to sudden changes in preferences, this module integrates a jump detection mechanism. This mechanism continuously monitors the behavioral observation sequence through a Cumulative Sum Control Chart (CUSUM) algorithm. When a significant abrupt change in statistical characteristics is detected (such as a sharp increase in the redemption rate), the system determines that a structural shift in investor risk appetite may have occurred. At this point, the adaptive adjustment unit temporarily increases the covariance matrix of the process noise. This allows particles to explore a wider range of states, thereby accelerating the model's convergence to new preferred states.

[0075] Output and Standardization: The final output of the module includes:

[0076] Risk aversion coefficient estimate: and its variance ;

[0077] Risk level mapping: based on The numerical range is mapped to the risk level required by regulatory requirements. Example mapping: This mapping relationship can be configured;

[0078] Uncertainty measure: It serves as a quantitative indicator for estimating confidence levels and is used for compliant disclosure.

[0079] The dynamic asset allocation optimization module receives risk preference estimates from the time-varying risk preference estimation module. It then embeds these risk preference parameter estimates as endogenous variables into the objective or reward function of the asset allocation optimization model to obtain the optimal asset allocation weights. The specific implementation method is as follows:

[0080] Market parameter prediction: Based on the cleaned market data provided by the data acquisition and preprocessing module, the expected return vector μ and return covariance matrix Σ of each asset in the future investment period are predicted using methods such as factor model and time series analysis.

[0081] Optimization Problem Definition and Solution: Construct a multi-objective optimization problem whose objective function integrates expected utility, risk control, and transaction costs.

[0082] ;

[0083] in, Let be the asset allocation weight vector to be solved. This represents the current portfolio weighting; Including budget constraints and regulatory and product constraints feasible domain, and These are the lower and upper limits of asset weights, respectively, set according to regulatory requirements, product contracts, and internal risk control rules. Let r be the future random return of the portfolio, and r be the asset return vector. The utility function is used; an exponential utility function is employed. And the time-varying risk preference estimation module estimates in real time Substitute directly. The conditional value of risk at a confidence level α (e.g., 95%) is used to control for extreme losses. From position Adjust to The resulting transaction cost function; The pre-defined tradeoff coefficients control the degree of aversion to tail risk and transaction costs, respectively.

[0084] This module provides two optional solution schemes to address different market environments and complexity requirements:

[0085] Option 1: Robust optimization implementation. This approach addresses the estimation error of the market parameters (μ, Σ) by introducing a set of parameter uncertainties, Ξ. The original problem is transformed into a minimization / maximization problem seeking the optimal strategy in the worst-case scenario. The uncertainty set Ξ is constructed as follows:

[0086] ;

[0087] in For estimating the revenue covariance matrix, and These are the lower and upper bounds of the expected return vector, respectively, which can be set based on historical volatility, for example... , Let k be the standard error vector of the return estimate, and k be the scaling factor (e.g., k=2). The preset minimum risk structure matrix, This indicates that the estimation bias of the covariance matrix is ​​limited to the baseline estimate. Within the Frobenius norm sphere centered at ρ, the size of radius ρ is configured according to the degree of fluctuation of the estimation error or the risk tolerance.

[0088] Based on this, the robust optimization problem can be formulated as follows:

[0089] ;

[0090] This problem can be solved using robust second-order cone programming or semidefinite programming algorithms, which significantly improves the robustness and stability of asset allocation strategies. This ensures that even in the worst-case scenario where market parameters exhibit unfavorable estimation biases, the portfolio can still maintain superior performance, reducing the model's over-reliance on prediction accuracy.

[0091] Option 2: Constrained reinforcement learning implementation, employing algorithms such as Proximal Policy Optimization (PPO) or Deep Deterministic Policy Gradient (DDPG) to train the agent in a simulation environment constructed from historical data. Implementation is as follows:

[0092] state: This includes current risk appetite, market characteristics, previous portfolio performance, and weightings. This is the current risk aversion coefficient estimate output by the time-varying risk preference estimation module; This represents the current market feature vector. and These represent the portfolio's return and volatility in the previous period, respectively. This is the asset weight vector at the end of the previous period.

[0093] action: This represents the adjustment vector for asset weights. It is output through an action network and processed via projection or a Softmax / Tanh layer to ensure that it meets the following requirements. And the adjusted weights satisfy Constraints.

[0094] Reward function: ,in, t represents the Sharpe ratio (or risk-adjusted return) of the portfolio at time t. Adjust the weight vector The L1 norm is used to penalize the costs incurred by frequent or large transactions; For indicator functions, Indicates the action This leads to the next option being overweight. Violation of budget constraints ) or weighted boundary constraints ( The value is 1 if any of the conditions in the equation are met, otherwise it is 0; the coefficient k is set to a positive number much greater than 1 to ensure that the constraints are strictly satisfied.

[0095] Training: The agent learns through interaction with a historical market simulation environment, using algorithms such as Proximal Policy Optimization (PPO) for training, ultimately resulting in an agent capable of adapting to different risk appetite states. and market environment The optimal configuration strategy network.

[0096] Dynamic rebalancing triggering mechanism: To reduce unnecessary transaction costs, this system adopts rule-based, event-triggered rebalancing instead of fixed-period adjustments. Triggering conditions include:

[0097] The estimated risk aversion coefficient changes by more than the threshold δ: , where d is the last adjustment interval;

[0098] Market volatility indices (such as VIX) break through a preset threshold range.

[0099] After rebalancing is triggered, the system calls the optimizer to calculate the target weights. The system generates rebalancing instructions. The trading execution module uses an order splitting algorithm (such as TWAP or VWAP) to break down large orders into multiple smaller orders, which are then executed within a specified time period to minimize market impact costs.

[0100] Compliance and Auditability Module: Records the entire system's decision logs in an immutable manner and generates auditable reports including decision attribution analysis based on model interpretability technology. Specifically, it includes:

[0101] Structured log unit: Automatically records the entire chain of data for each system decision in an immutable manner (such as by constructing a Merkle tree hash chain), including the hash fingerprint of the input data, the model version number used, all parameters, output results, timestamps, and a text summary of the decision logic.

[0102] Explainable Reporting Unit: Utilizing Explainable Artificial Intelligence (XAI) technologies, such as SHAP values ​​or LIME, this unit performs attribution analysis on each change in risk preference estimates and adjustment of asset allocation weights, automatically generating natural language reports that clarify the main driving factors (e.g., "The increase in the risk aversion coefficient this time is mainly due to the client's three consecutive stop-loss transactions in recent days").

[0103] Audit Interface Unit: Provides standardized API interfaces for regulatory agencies, supporting the "replay" verification of model decisions at any historical point in time, and querying inputs, outputs, and intermediate inference processes.

[0104] Customer Confirmation Module: When the risk level dynamically assessed by the system differs from the customer's most recent questionnaire result by more than a certain range, or when the single portfolio adjustment is too large, the confirmation process is automatically triggered through the client and the customer's confirmation operation (time, method, and result) is fully recorded to meet the compliance requirements of "seller's due diligence".

[0105] Interaction and Display Module: This module provides differentiated graphical interfaces for different users.

[0106] Customer interface: Displays individual risk preference curves in chart format. The changes over time, the current asset allocation radar chart, historical rebalancing records, and the simulated future return and risk distribution.

[0107] Compliance personnel interface: Provides log query and model "replay" functions, supports filtering by time, customer, and decision type, and allows tracing the complete decision chain.

[0108] Administrator Interface: Provides a strategy backtesting platform, model performance monitoring dashboard (such as estimation error, strategy Sharpe ratio) and parameter tuning interface, and supports A / B testing and hot model switching.

[0109] Example 2

[0110] like Figure 2 As shown, an auditable asset allocation method based on time-varying risk preference estimation, whose steps correspond to the execution flow of the above system embodiment, includes the following steps:

[0111] S1: Data Acquisition and Feature Extraction. Multi-dimensional raw data is acquired from customer databases, trading systems, and market data sources, and then cleaned, encrypted, and feature-engineered to convert it into standardized input vectors.

[0112] S2: Dynamic modeling of risk preference. Construct a state-space model and define the hidden state vectors as described in Example 1. The state transition equation and observation equation are derived, and the parameters and noise covariance matrix are initialized.

[0113] S3: Particle Filter Estimation and Jump Detection. Performs the recursive steps of particle filtering (prediction, update, resampling) and outputs the current time-to-time risk preference parameter estimate. Its uncertainty. The CUSUM algorithm is run in parallel to adaptively adjust process noise when behavioral abrupt changes are detected.

[0114] S4: Robust optimization / reinforcement learning configuration solution. The result obtained in step S3... As the core input, a multi-objective optimization model is constructed that integrates expected utility, CVaR risk constraints, and transaction costs. The optimal asset allocation weights are obtained by solving the model using robust optimization or constrained reinforcement learning algorithms.

[0115] S5: Configuration, Execution, and Recording. Generate and execute transaction instructions based on the output of S4. Simultaneously, record the entire data chain (inputs, model, parameters, outputs) from S1 to S4 in a hash chain into an immutable audit log.

[0116] S6: Explanatory Analysis and Audit Report Generation. Based on the data from this decision-making process, an attribution analysis report is generated using XAI technology to explain the main reasons for changes in risk appetite and asset adjustments, supporting inquiries from clients and regulators.

[0117] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. An auditable asset allocation system based on time-varying risk preference estimation, characterized in that, include: The data acquisition and preprocessing module is used to acquire and process investors' static attributes, dynamic behavior data, and market environment data. The time-varying risk preference estimation module is connected to the data acquisition and preprocessing module. It is used to model the investor's risk preference parameters as time-varying hidden states based on the state-space model, and to perform online Bayesian estimation of the hidden states using the sequential Monte Carlo method, and output the estimated value of the risk preference parameters and their uncertainty. The dynamic asset allocation optimization module is connected to the time-varying risk preference estimation module. It is used to embed the estimated risk preference parameters as endogenous variables into the objective function or reward function of the asset allocation optimization model to solve for the optimal asset allocation weight. The compliance and auditability module is used to record the system's end-to-end decision logs in an immutable manner and generate auditable reports containing decision attribution analysis based on model interpretability technology.

2. The auditable asset allocation system based on time-varying risk preference estimation according to claim 1, characterized in that, The state-space model includes: The state transition equation is used to describe the dynamic evolution of the hidden state vector from the previous time step to the current time step. This evolution process is affected by the state at the previous time step and the current market environment data. The observation equation is used to map the current hidden state vector and market environment data into observable investor behavior indicators. The hidden state vector contains at least a logarithmic form of a parameter representing the investor's risk aversion level.

3. The auditable asset allocation system based on time-varying risk preference estimation according to claim 2, characterized in that, The online Bayesian estimation using the sequential Monte Carlo method is performed according to the following recursive steps: Based on the state transition equation, state prediction is performed on the set of particles representing the posterior distribution; Obtain the observed value of the investor behavior index at the current time t, and calculate the likelihood of each particle and update its normalized weight based on the observation equation and the observed value. When the number of effective particles representing the particle weight dispersion is lower than a preset threshold, system resampling is performed to generate a new set of equally weighted particles. Based on the particle set and its weights obtained after resampling, the minimum mean square error estimate and estimated covariance of the risk preference parameter are calculated.

4. The auditable asset allocation system based on time-varying risk preference estimation according to claim 3, characterized in that, The time-varying risk preference estimation module also includes: A jump detection unit is used to continuously monitor the investor behavior data sequence through a cumulative sum control graph algorithm to detect sudden changes in its statistical characteristics; An adaptive adjustment unit is used to temporarily increase the covariance matrix of the process noise in the state transition equation when a sudden change is detected, so that the particle set can quickly adapt to structural changes in investor risk preferences.

5. An auditable asset allocation system based on time-varying risk preference estimation according to claim 1, characterized in that, The dynamic asset allocation optimization module is configured to: construct an objective function, which incorporates at least the following two items: an expected utility term with the estimated risk preference parameter as the core parameter; and a conditional value at risk penalty term or constraint for controlling tail risk; wherein the objective function is used to solve for the optimal asset allocation weights.

6. An auditable asset allocation system based on time-varying risk preference estimation according to claim 5, characterized in that, The dynamic asset allocation optimization module is configured to perform the following robust optimization process: Construct a set of parameters to characterize the uncertainty in estimating the expected rate of return and the return covariance matrix of an asset; For each possible set of parameters in the parameter set, calculate the value of the objective function; The asset allocation weight that maximizes the objective function under the worst-case parameters on the parameter set is selected as the optimal asset allocation weight.

7. An auditable asset allocation system based on time-varying risk preference estimation according to claim 6, characterized in that, The expected utility term adopts the form of an exponential utility function, specifically: ; in, R is the risk aversion coefficient estimated at the current moment, and R is the stochastic return of the portfolio.

8. An auditable asset allocation system based on time-varying risk preference estimation according to claim 5, characterized in that, The dynamic asset allocation optimization module adopts a constrained reinforcement learning framework, which is configured as follows: Using a policy network, the state vector containing the estimated risk preference parameters is mapped to an adjustment action for asset weights; The policy network is trained by interacting with an environmental simulation, wherein the evaluation function used for training includes at least two of the following: one is positively correlated with risk-adjusted investment returns; and the other is a penalty for violations of investment constraints.

9. An auditable asset allocation system based on time-varying risk preference estimation according to claim 1, characterized in that, The compliance and auditability module specifically includes: The structured log unit is used to associate and store the input data fingerprint, model version, key parameters, output results and timestamp of each decision in the form of a hash chain to ensure data integrity and traceability. The interpretable reporting unit is used to automatically generate attribution reports after each asset allocation adjustment, and to clarify the main driving factors of the changes in risk preference estimates and asset weight adjustments through feature contribution analysis.

10. An auditable asset allocation method based on time-varying risk preference estimation, characterized in that, Includes the following steps: Collect and preprocess multi-dimensional investor data; A state-space model is constructed, and the investor's risk preference parameter is defined as a time-varying hidden state; By applying the sequential Monte Carlo method, new behavioral data and market data are recursively fused to estimate the latent state online and output the current risk preference parameter estimate. Using the estimated risk preference parameters as the core optimization variables, an asset allocation optimization model that integrates expected utility, risk control, and transaction costs is constructed and solved, and the optimal allocation weights are determined. The rebalancing operation is performed according to the optimal configuration weights, and the entire chain data of this decision is recorded in the tamper-proof audit log and a decision attribution report is generated.