User financial risk assessment method and system based on composite features
By constructing composite feature vectors, combining dynamic behavioral features and topological network features, and utilizing phased ensemble learning and knowledge graph reasoning, the problem of insufficient comprehensiveness in user financial risk assessment in existing technologies is solved, achieving more accurate and secure risk assessment and meeting the privacy protection and compliance requirements of financial data.
Patent Information
- Application Number
- CN202511808493.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-01-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies lack effective capture and quantification of users' dynamic behavioral responses in user financial risk assessment, resulting in insufficient comprehensiveness and predictability of risk assessment. They mainly rely on static topological relationships, which cannot fully reflect user behavior patterns.
By acquiring multiple heterogeneous data sources, performing data fusion and preprocessing, constructing composite feature vectors, and combining graph neural networks and graph theory algorithms, the system captures users' dynamic behavioral characteristics and topological network characteristics. It also utilizes a phased ensemble learning model and a knowledge graph inference engine for risk assessment, and finally protects privacy through blockchain notarization and federated learning.
It improves the comprehensiveness and accuracy of risk assessment, enhances the security and compliance of decision-making, achieves a comprehensive characterization of user risks and privacy protection, and provides a traceable decision-making process.
Smart Images

Figure CN121280141A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of financial data processing technology, and in particular to a user financial risk assessment method and system based on composite features. Background Technology
[0002] In the fintech field, risk management is central to ensuring financial security. Existing technologies, such as graph neural networks, are beginning to utilize topological relationships by constructing graphs of relationships between entities, such as transaction graphs between users and merchants, to identify anomalous patterns for fraud prevention or credit assessment. These methods assess risk by analyzing a user's structural position within the financial network, representing a significant improvement over traditional statistical methods.
[0003] However, the aforementioned existing technologies primarily focus on analyzing relatively static or slowly changing topological relationships, i.e., "who the user is associated with." They lack the ability to effectively capture and quantify dynamic changes in user behavior, especially the immediate behavioral response patterns of users when facing specific stressful events, such as sudden changes in their financial situation. Therefore, existing technologies, when characterizing user risk, still suffer from the shortcomings of having a single dimension and failing to fully reflect dynamic user behavior patterns, resulting in insufficient comprehensiveness and predictability in risk assessment. Summary of the Invention
[0004] The purpose of this application is to provide a user financial risk assessment method and system based on composite features, in order to solve the technical problem that existing technologies mainly rely on static topological relationships for risk assessment, while ignoring the dynamic behavioral responses of users to specific events, resulting in an incomplete and inaccurate risk profile.
[0005] In a first aspect, embodiments of this application provide a user financial risk assessment method based on composite features, comprising the following steps: acquiring multiple heterogeneous data sources associated with a user, wherein the heterogeneous data sources include at least user behavior data reflecting individual user behavior and macro-environmental data reflecting the external environment; performing fusion preprocessing on the multiple heterogeneous data sources to generate a unified paradigm dataset, wherein the fusion preprocessing includes standardizing the format and fields of data from different sources and repairing missing or non-standard data in the dataset based on user context information; constructing a composite feature vector representing user financial behavior based on the unified paradigm dataset, wherein the composite feature vector includes at least dynamic behavioral features reflecting dynamic changes in user behavior and topological network features reflecting the relationship between the user and other entities; inputting the composite feature vector into one or more prediction models for analysis to generate a risk assessment result representing the user's risk level; and executing predefined risk disposal actions when the risk assessment result meets preset risk conditions.
[0006] Optionally, constructing a composite feature vector representing a user's financial behavior specifically includes: The dynamic behavioral characteristics are generated by monitoring preset trigger events in the user data stream and capturing user behavior data within a preset time window after the trigger event occurs. Furthermore, the topological network features are generated by constructing user-related graph structures and using graph theory algorithms or graph neural networks to compute graph metrics.
[0007] Optionally, the preset triggering event includes: The user's monthly income changes by more than a preset fluctuation threshold; Alternatively, the user can log in from a new geographical location or using a new device.
[0008] Optionally, the user-related graph structure includes: The user-merchant bipartite diagram illustrates the transaction relationship between users and merchants. Alternatively, it could be a user-to-user directed graph representing the fund transfer relationships between users.
[0009] Optionally, the step of repairing missing or non-standard data in the dataset based on user context information specifically includes: Using users' historical behavior patterns or occupational types as conditions, generative models are used to fill in missing values in the dataset.
[0010] Optionally, the prediction model is a staged ensemble learning model, and its analysis process includes: In the first stage, the first model is used to perform feature filtering on the composite feature vector; In the second stage, the filtered features are input into a second model used to capture temporal dependencies in order to generate intermediate risk predictions. In the third stage, a knowledge graph reasoning engine is introduced to match user behavior with the rule base, so as to output the risk assessment result based on the intermediate risk prediction.
[0011] Optionally, the predefined risk management action includes at least one of the following: Push early warning information to the risk control system; Restrict user account permissions; Generate visualized risk reports; The risk assessment results and decision-making basis will be stored on the blockchain.
[0012] Optionally, the step of inputting the composite feature vector into one or more prediction models for analysis to generate risk assessment results characterizing the user's risk level is performed within a federated learning framework, specifically including: The model update amount is calculated locally by multiple participants using their respective user data; The central server aggregates the encrypted model update data uploaded by each participant without accessing the original data, in order to generate the global model update data. Each participant updates its local model based on the global model update amount.
[0013] Optionally, when the predefined risk management action includes storing the risk assessment results and decision-making basis on the blockchain, the storage specifically includes: Calculate the hash value for the composite feature vector or a subset thereof; The hash value, the decision path of the prediction model, and the risk assessment result are written into the blockchain via a smart contract. Secondly, embodiments of this application also provide a user financial risk assessment system, which includes: The data acquisition module is used to acquire multiple heterogeneous data sources associated with the user. The heterogeneous data sources include at least user behavior data reflecting individual user behavior and macro-environmental data reflecting the external environment. The data fusion module is used to perform fusion preprocessing on the various heterogeneous data sources to generate a dataset with a unified paradigm. The fusion preprocessing includes standardizing the format and fields of data from different sources and repairing missing or non-standard data in the dataset based on user context information. The feature construction module is used to construct a composite feature vector representing user financial behavior based on the dataset of the unified paradigm. The composite feature vector includes at least dynamic behavioral features that reflect the dynamic changes of user behavior and topological network features that reflect the relationship between the user and other entities. The risk assessment module is used to input the composite feature vector into one or more prediction models for analysis, so as to generate risk assessment results that characterize the user's risk level; The risk management module is used to execute predefined risk management actions when the risk assessment results meet preset risk conditions.
[0014] Compared with the prior art, this application has the following beneficial effects: 1. Enhance the comprehensiveness and accuracy of assessment: By combining topological network features reflecting static relationships with dynamic behavioral features reflecting dynamic changes, the composite feature vector constructed by this invention can more comprehensively and profoundly depict user behavior patterns, significantly overcoming the shortcomings of existing technologies with single feature dimensions, thereby greatly improving the comprehensiveness, accuracy and predictability of risk assessment.
[0015] 2. Enhance the security and compliance of decision-making: By preferentially adopting privacy-preserving computing technologies such as federated learning, joint modeling and analysis can be carried out without exposing the original data of each participant, effectively protecting user privacy, meeting financial data security and compliance requirements, and solving the data silo problem.
[0016] 3. Achieve traceability and transparency in the decision-making process: By preferentially adopting the blockchain evidence storage mechanism, key links in risk control decision-making are generated into tamper-proof on-chain records, providing a highly credible traceability basis for post-audit, making the complex risk control decision-making process easier to understand and review. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. Furthermore, these drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments.
[0018] Figure 1 A flowchart illustrating an overall method for assessing user financial risk based on composite features, provided in this application embodiment; Figure 2 A schematic diagram illustrating the construction of composite feature vectors provided in an embodiment of this application; Figure 3 This application provides an overall structural diagram of a user financial risk assessment system as an embodiment of the present application. Figure 4 Signaling interaction timing diagram of the federated learning analysis process provided for embodiments of this application; Figure 5 This is a schematic diagram of the internal structure of the blockchain evidence storage module provided in the embodiments of this application. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0020] Example 1 This embodiment provides a basic implementation scheme for a user financial risk assessment method and system based on composite features. It fully demonstrates the entire process from multi-source heterogeneous data acquisition, data fusion preprocessing, composite feature construction, phased model evaluation, to final risk disposal and blockchain evidence storage.
[0021] Please see Figure 1 This illustrates the overall flow of a user financial risk assessment method based on composite features provided in this application embodiment. The method can be executed by one or more data processing devices (e.g., a server cluster) and specifically includes the following steps: S101: Data Acquisition.
[0022] In one embodiment of this application, the system first performs data acquisition, aiming to comprehensively collect multi-dimensional information related to the target user. This information comes from a wide range of sources and is diverse in nature, collectively forming the data foundation for risk assessment. Specifically, it acquires multiple heterogeneous data sources associated with the user, including at least user behavior data reflecting individual user behavior and macro-environmental data reflecting the external environment.
[0023] The acquisition of user behavior data is contingent upon obtaining explicit user authorization. For example, users can authorize the system to securely access their data across different financial institutions or platforms via an interface adhering to the Open License 2.0 protocol. The types of user behavior data acquired are diverse, including at least: 1. Structured transaction data, such as a user's bank account statements, which include fields such as counterparty account number, transaction amount, transaction time, currency type, and transaction remarks; and a user's securities transaction records, including the securities codes bought and sold, transaction price, transaction quantity, and operation time. This also includes data sources such as APP operation logs, IoT sensor and time-series behavioral data including page dwell time, and biometric authentication records. This type of data uses financial ontology mapping rules to unify multi-source fields.
[0024] 2. Unstructured and semi-structured data, such as user operation logs within financial service applications, which record users' click flows, feature usage preferences, etc.; and user comments related to financial products posted on public social media or product review sections, which may contain signals of users' emotions and specific life events.
[0025] 3. Time-series behavioral data, such as fine-grained behaviors when users interact with financial applications, such as the duration of time spent on key information input pages, and the response time and number of attempts when performing biometric authentication such as facial recognition or fingerprint authentication.
[0026] The data sources include transaction data from the bank's core system, third-party payment platforms, tax declaration interfaces, and IoT devices. The OAuth 2.0 protocol is used for dynamic token verification. For unstructured text, sentiment and event entities are extracted using BERT-BiLSTM. Time-series data is dynamically normalized to eliminate noise. Cross-cycle behavioral features, capital closed-loop topology features, and stress response features are innovatively constructed.
[0027] The acquisition of macroeconomic data is primarily achieved through web crawling technology. Understandably, the system can deploy a distributed crawler cluster based on the Scrapy framework to periodically scrape macroeconomic indicators from authoritative data sources such as the National Bureau of Statistics and major economic information websites. In this embodiment, the system focuses on collecting recent Consumer Price Index (CPI) and Purchasing Managers' Index (PMI), as these data serve as important factors in measuring the overall economic climate and provide a macroeconomic context for understanding user behavior.
[0028] S102: Data fusion and preprocessing.
[0029] Because the raw data obtained comes from diverse sources and has different formats, and may contain missing data or outliers, data fusion and preprocessing are required to generate a clean, standardized, and uniform dataset, laying the foundation for subsequent feature construction.
[0030] Specifically, multiple heterogeneous data sources are merged and preprocessed to generate a dataset with a unified paradigm. The fusion and preprocessing includes standardizing the format and fields of data from different sources, as well as repairing missing or non-standard data in the dataset based on user context information.
[0031] This step specifically includes: 1. Standardized Mapping: The system first applies a set of predefined financial ontology mapping rules to unify fields from different data sources. For example, the "credit amount" field in one banking system and the "income" field in another payment platform will both be mapped to the "credit_amount" standard field in the unified dataset. Correspondingly, transaction types with similar semantics but different names, such as "payment," "transfer," and "consumption," will be uniformly classified into standard transaction event codes, thereby ensuring semantic consistency of the data.
[0032] 2. Unstructured Data Processing: For unstructured data such as text comments obtained from social media, the system employs natural language processing techniques for analysis. Specifically, a deep learning model combining a BERT pre-trained language model and a bidirectional long short-term memory network can be used. This model can deeply understand the contextual semantics of the text and perform two tasks: first, sentiment polarity analysis, classifying the text into positive, negative, or neutral categories and quantifying their sentiment intensity; second, key entity identification, extracting predefined key event entities from the text, such as "unemployment," "overdue payment," "complaint," and "loan rejection." This extracted structured information will be added to the dataset as new features.
[0033] In addition, the mean behavior is calculated using a preset period as a window to eliminate collection errors. A conditional generative adversarial network is used to address missing data. Specifically, this includes: window weight decays over time to eliminate device jitter errors; for sudden external events, an adaptive window expansion mechanism is activated, automatically enlarging the window to smooth out abnormal spikes over 30 minutes to ensure the stability of behavioral trends; the generator uses user occupation type and historical behavior patterns as input conditions to synthesize filler data that matches their characteristics; the discriminator forces the synthesized data to align with the real distribution; and after filling, missing flags are added as derived features input to the subsequent processing model.
[0034] 3. Data Imputation and Repair: To address potential missing values in the dataset, this embodiment employs an advanced generative model for data imputation, rather than simple mean or median filling. Specifically, it repairs missing or non-standard data based on user context information, including using the user's historical behavior patterns or occupation type as conditions to fill in missing values in the dataset using a generative model.
[0035] Specifically, a conditional generative adversarial network (GAN) is employed. This network uses stable and important contextual information about the user (such as the user's "occupation type," "age group," and "historical spending level") as generation conditions. Guided by these conditions, the generator network generates reasonable imputation values that are highly consistent with the user profile. For example, for a "high-income IT professional," the model will generate spending data that matches their identity to fill in missing transaction records. It should be noted that, to ensure data rigor, all imputed data points are given a Boolean flag "is_imputed," which itself can also be used as a derived feature for subsequent models to evaluate the impact of the imputed data on risk prediction.
[0036] S103: Construction of composite features.
[0037] This step aims to construct a composite feature vector that can comprehensively and deeply characterize a user's financial risk. Specifically, based on a unified paradigm dataset, a composite feature vector representing a user's financial behavior is constructed. This composite feature vector includes at least dynamic behavioral features reflecting the dynamic changes in user behavior, and topological network features reflecting the user's relationships with other entities.
[0038] More specifically, constructing a composite feature vector representing a user's financial behavior includes: generating the dynamic behavioral features by monitoring preset trigger events in the user data stream and capturing user behavior data within a preset time window after the trigger event occurs; and generating the topological network features by constructing a user-related graph structure and using graph theory algorithms or graph neural networks to calculate graph indicators.
[0039] Please see Figure 2 This is a schematic diagram of the construction of a composite feature vector. As shown in the figure, the composite feature vector 60 is composed of topological network feature 30 and dynamic behavior feature 50.
[0040] In this embodiment, the system constructs a composite feature vector 60 with a total of 326 dimensions, which consists of: traditional statistical features, deep learning embedded features, and dynamic event markers. Traditional statistical features include, for example, the Herfindahl-Hirschman Index of transaction amount over the past 30 days to quantify consumption concentration and cross-platform fund flow closed-loop detection, which are automatically generated through a sliding window; the deep learning embedded features are output by the GNN node embedding engine, which constructs a user-merchant bipartite graph, calculates the merchant clustering coefficient, and generates a 128-dimensional vector, while dynamically capturing event markers to form a real-time behavioral fingerprint; the event markers use one-hot encoding, and the numerical features are normalized to the [0, 1] interval using Min-Max to output a 326-dimensional vector, which is compressed and transmitted to the federated learning node through the Protobuf protocol to support real-time risk prediction and compliance analysis.
[0041] More specifically, the process of constructing a composite feature vector is as follows: 1. Construction of Topological Network Feature 30: This feature aims to quantify a user's structural position and influence in the financial relationship network. The system first constructs a user-merchant bipartite graph based on the user's transaction data. In this graph, nodes are divided into two categories: user nodes 10 representing users and merchant nodes 20 representing merchants. If a user and a merchant have transacted, an edge is established between the corresponding user node 10 and merchant node 20. Based on this graph structure, the system uses graph neural network technology to learn the node representations. Specifically, the GraphSAGE model can be used to iteratively aggregate the feature information of the node's neighbors to generate a low-dimensional, dense embedding vector for each user node 10. This embedding vector can encode the local network topology information of the user node. In this embodiment, a 128-dimensional embedding vector is generated as part of the topological network feature 30. Furthermore, some traditional graph theory metrics, such as the clustering coefficient of the group of merchants transacting with the user, can also be calculated to measure the concentration of their trading partners.
[0042] 2. Construction of Dynamic Behavioral Feature 50: This feature aims to capture changes in user behavior patterns when coping with specific stressful events, representing a significant supplement to traditional static features. The system continuously monitors the user's data stream to detect preset trigger events 41. In this embodiment, a key trigger event 41 is defined as: a change in the user's monthly income exceeding a preset fluctuation threshold (i.e., a sharp fluctuation in the user's monthly income), for example, when the total monthly income decreases by more than 20% compared to the moving average of the past six months. Once this event is detected, the system immediately initiates a 72-hour response window 42. Within this response window 42, the system closely records all of the user's high-risk operational behaviors, such as large transfers (a single transaction exceeding 50,000 RMB), cross-border transactions (the counterparty is located overseas), and nighttime transactions (occurring between 1 AM and 5 AM). Subsequently, the system calculates a series of quantitative indicators based on these records, such as the total frequency, total amount, and proportion of high-risk operations to the total transaction amount. These indicators collectively constitute the dynamic behavioral feature 50.
[0043] 3. Additional Features: In addition to the two core features mentioned above, the composite feature vector 60 also includes other types of features. For example, composite indicators calculated based on macroeconomic data (such as the Consumer Price Index and Purchasing Managers' Index) and user transaction data, such as the Herfindahl-Hirschman Index, which measures the changing trend of user spending over the past 30 and 90 days, are used to measure the changing trend of consumption concentration. All these features are combined to form the final 326-dimensional composite feature vector 60.
[0044] S104: Risk assessment.
[0045] The constructed composite feature vector 60 is input into one or more carefully designed staged ensemble learning models (predictive models) for analysis to generate the final risk assessment result representing the user's risk level. The model is designed to combine the advantages of different models, achieving a progressive process from feature selection and deep analysis to rule validation.
[0046] In some embodiments, the prediction model is a staged ensemble learning model, and its analysis process includes: 1. First Stage: Feature filtering of the composite feature vector using the first model. Specifically, the complete 326-dimensional feature vector is input into a gradient boosting decision tree model (the first model). This model not only has powerful predictive capabilities but also provides a reliable ranking of feature importance. By analyzing the gain or coverage of each feature after model training, the system can filter out the top 50 features that contribute the most to risk prediction. This step effectively reduces the computational complexity of subsequent models and eliminates noisy features.
[0047] 2. Second Stage: The selected features are input into a second model to capture temporal dependencies, generating an intermediate risk prediction. Specifically, the 50 selected features, especially those with temporal attributes (such as recent transaction sequences), are input into a Long Short-Term Memory (LSTM) network with an attention mechanism. LTM networks excel at capturing long-term dependencies in sequential data, while the attention mechanism allows the model to dynamically assign different weights to inputs at different time points during sequence analysis, thus focusing on historical behaviors most critical to the current risk assessment. The output of this stage is an intermediate risk prediction, such as a user's short-term liquidity risk score.
[0048] 3. Third Stage: Introducing a knowledge graph inference engine to match user behavior with a rule base, outputting the risk assessment result based on the intermediate risk prediction. Specifically, this stage introduces a knowledge graph inference engine. This engine receives the intermediate risk prediction results from the second stage and incorporates an expert rule base in the field of financial risk control (e.g., typical pattern rules for anti-money laundering and anti-fraud). The inference engine performs subgraph matching or logical reasoning between the user's behavioral characteristics (from the filtered feature set) and the rules in the rule base. For example, a rule can be defined as: "If a user's short-term liquidity risk score is higher than 0.9 and they have recently engaged in multiple cross-border small-amount high-frequency transactions without clear commercial justification, then their compliance risk level is high." In this way, the model combines data-driven prediction with expert knowledge, ultimately outputting a comprehensive risk assessment result, such as a compliance score ranging from 0 to 1.
[0049] S105: Risk Management.
[0050] The system determines whether the risk assessment result generated in the previous step meets preset risk conditions. When the risk assessment result meets the preset risk conditions, a predefined risk handling action is executed. In this embodiment, the risk condition is "compliance score higher than 0.85". If this condition is met, the system will automatically trigger and execute one or more predefined risk handling actions. Specifically, these include: 1. Generate a Visualized Risk Report: The system utilizes a geographic information visualization engine to display the user's recent high-risk transaction behavior in a spatiotemporal manner on a map. For example, it overlays the user's consumption and investment activities according to geographic hotspots and timelines, and uses density clustering algorithms (such as DBSCAN) to automatically mark abnormal activity clusters, such as a dense cluster of cross-border transfers occurring in a non-residential area in the early morning. This visualized report is presented intuitively to risk control analysts. This function can be developed by... Figure 3 The decision visualization module 500 shown is implemented.
[0051] 2. Blockchain storage of risk assessment results and decision-making basis: To ensure the transparency, immutability, and traceability of the decision-making process, the system performs a storage operation. This storage specifically includes: calculating the hash value of the composite feature vector or a subset thereof; and writing the hash value, the decision path of the prediction model, and the risk assessment results to the blockchain via a smart contract. Please refer to [link to relevant documentation]. Figure 5 This is a schematic diagram of the internal structure of the blockchain evidence storage module 600. First, the data digest unit 610 calculates the SHA-256 hash value of the 326-dimensional composite feature vector 60, which serves as the core basis for decision-making, generating a unique digital fingerprint. Then, the encryption-to-chain unit 620 calls a smart contract pre-deployed on a consortium blockchain (e.g., a consortium blockchain based on Hyperledger Fabric) to write the hash value, the key decision path based on the model in the third stage (e.g., the triggered rule ID), and the final compliance score, as a transaction into the blockchain ledger. In the future, the audit verification unit 630 can respond to audit requests to verify the integrity and consistency of the records on the chain.
[0052] 3. Push early warning information to the risk control system and restrict user account permissions: The system through... Figure 3 The Real-Time Feedback Engine 700 shown pushes high-risk warning information, links to visual reports, and suggested handling measures to the financial institution's core risk control system in real time via message queues (such as Kafka).
[0053] 4. Restrict user account permissions: After receiving an alert, the risk control system can automatically implement intervention measures according to preset strategies, such as temporarily freezing the user and its associated accounts for 72 hours, pending further manual review.
[0054] The entire process in this embodiment, from data acquisition to risk management, constitutes an automated data processing and decision-making pipeline, which can be operated by... Figure 3 The user financial risk assessment system shown is implemented. The system includes: The data acquisition module 100 is used to acquire multiple heterogeneous data sources associated with the user. The heterogeneous data sources include at least user behavior data reflecting individual user behavior and macro-environmental data reflecting the external environment. The data fusion module 200 is used to perform fusion preprocessing on the multiple heterogeneous data sources to generate a dataset with a unified paradigm. The fusion preprocessing includes standardizing the format and fields of data from different sources and repairing missing or non-standard data in the dataset based on user context information. The feature construction module 300 is used to construct a composite feature vector representing user financial behavior based on the dataset of the unified paradigm. The composite feature vector includes at least dynamic behavioral features that reflect the dynamic changes of user behavior and topological network features that reflect the relationship between the user and other entities. The risk assessment module is used to input the composite feature vector into one or more prediction models for analysis, in order to generate risk assessment results that characterize the user's risk level (specifically...). Figure 3 Federated Learning Analytics Module 400). The risk handling module is used to execute predefined risk handling actions (specifically including) when the risk assessment result meets preset risk conditions. Figure 3 The system includes a decision visualization module 500, a blockchain evidence storage module 600, and a real-time feedback engine 700 to achieve the corresponding functions. These modules work together to complete end-to-end, high-precision assessment and handling of users' financial risks.
[0055] Example 2 As an optional implementation, this embodiment is a variation of embodiment 1, mainly demonstrating another extended construction method of dynamic behavioral feature 50, aiming to enhance the ability to identify specific types of risks (such as identity theft and account hijacking). Most of the steps in this embodiment are the same as those in embodiment 1, the main difference being the generation logic of dynamic behavioral feature 50 in step S103.
[0056] In this embodiment, when constructing dynamic behavior features 50 in step S103, the monitored preset trigger event 41 is expanded to include the following two situations: the user logs in at a new geographical location or using a new device. Specifically: 1. High-Risk Geographic Location Login Event: The system maintains a dynamically updated list of high-risk countries or regions. When a user's account is detected logging in for the first time from a location on this list, this login action is defined as a trigger event 41. This detection is achieved by parsing the IP address of the login request and comparing it with the geographic location database and historical login locations.
[0057] 2. Device Abnormal Change Event: The system monitors login device information (such as device ID, device fingerprint) associated with the user account. When it is detected that a user frequently changes login devices within a very short time window (e.g., within 1 hour) (e.g., more than 3 new, previously unseen device IDs appear), this pattern is also defined as a trigger event 41.
[0058] When any of the aforementioned triggering events 41 occurs, the system will open a preset response window 42, for example, 24 hours after the event. Within this window, the system captures all of the user's trading activities and calculates a series of targeted quantitative indicators as dynamic behavioral characteristics 50, such as: Total transaction amount during the window period; Counterparty dispersion: Calculates the unique number of counterparties within the trading window. Excessive dispersion may indicate abnormal fund dispersion behavior. High-frequency small-amount transactions: Transaction patterns in which the amount of a single transaction is less than a certain threshold (e.g., 100 yuan) but the total number of transactions exceeds another threshold (e.g., 10 times) within a statistical window period. This may be a signal of money laundering or testing the effectiveness of stolen accounts.
[0059] These newly generated dynamic behavioral features 50 will be incorporated into the composite feature vector 60, and together with other features such as the topological network feature 30 described in Example 1, will be input into the subsequent risk assessment model. In this way, the method of this embodiment can more sensitively capture abnormal signals directly related to account security, making the risk assessment model more comprehensive and robust.
[0060] Example 3 This solution is another variation of Embodiment 1, mainly demonstrating an alternative implementation of the topological network feature 30 and the risk assessment model (corresponding to step S104). This solution aims to simplify the model structure and focuses on identifying risk scenarios with strong network correlations, such as gang fraud or money laundering networks.
[0061] In this embodiment, the construction method of the topology network feature 30 in step S103 is different: Instead of constructing a user-merchant bipartite graph, the system builds a directed graph of fund flows between users based on their transfer records. In this graph, each node represents a user. If user A transfers money to user B, there exists a directed edge from node A to node B. The weight of the edge can be set to the total transfer amount or the number of transfers. Based on this fund flow graph, the system calculates the following indicators as topological network features 30: 1. PageRank Centrality Score: The PageRank algorithm is used to calculate the score of each node in the graph. This score measures a user's importance or influence in the overall money laundering network. Nodes with abnormally high scores may be core nodes playing a role in consolidating funds in the money laundering network.
[0062] 2. Community Tightness Score: First, a community discovery algorithm (such as the Louvain algorithm) is used to divide the entire graph into several financial communities (i.e., user groups with close financial transactions). Then, the internal tightness (such as the density of edges within the community) of each user is calculated. If a user is in an unusually tight financial community, it may mean that they are involved in group fraud activities.
[0063] In the risk assessment of step S104, this embodiment adopts a different model architecture than Embodiment 1: the system no longer uses a phased integration model, but instead employs an end-to-end graph convolutional network model. The input to this model includes two parts: first, the adjacency matrix of the entire user-user fund flow graph, which describes the network topology; second, the initial feature matrix of the nodes, where each row represents a user, and the features include the user's basic attributes and dynamic behavioral features 50 generated according to the methods of Embodiment 1 or Embodiment 2. The graph convolutional network model, through its multi-layer graph convolutional operations, can automatically learn and fuse the node's own features with the features of its neighboring nodes. At each layer, a node's representation vector is updated by aggregating the representation vectors of its first-order neighbors. After multiple layers of propagation, the final representation vector of each node contains rich information from its higher-order neighborhood, that is, it simultaneously fuses topological information and the user's own dynamic behavioral information. The last layer of the model is typically a classifier (such as a Softmax layer), which directly outputs the fraud risk probability of each user based on the node's final representation vector.
[0064] This end-to-end model architecture simplifies complex feature engineering and multi-stage modeling processes, improves analytical efficiency, and excels in identifying risk patterns that rely on complex network structures (such as money laundering chains and fraud rings). Its risk handling actions can adopt more granular, hierarchical strategies based on the output fraud probability value. For example, for users with a fraud probability between 0.7 and 0.9, the system can automatically trigger secondary verification (such as SMS verification codes or human video calls) instead of directly freezing accounts, thus achieving a better balance between risk control and user experience.
[0065] Example 4 This embodiment details how to securely execute a risk assessment process by applying privacy-preserving computation technology in scenarios involving multiple data holders (e.g., multiple banks jointly conducting risk assessments), particularly a deeper implementation of step S104 in Embodiment 1. The core of this solution lies in utilizing a federated learning framework, combined with homomorphic encryption and differential privacy technologies, to achieve cross-institutional collaborative modeling while ensuring that the original data of each party does not leave their local machine, thus meeting stringent data security and privacy compliance requirements.
[0066] Specifically, the composite feature vector is input into one or more prediction models for analysis to generate risk assessment results that characterize the user's risk level. The analysis process is executed under a federated learning framework and includes: multiple participants calculating model update amounts locally using their respective user data; a central server aggregating the encrypted model update amounts uploaded by each participant without accessing the original data to generate a global model update amount; and each participant updating its local model based on the global model update amount.
[0067] Please see Figure 4 This is a signaling interaction sequence diagram of the federated learning analysis process provided in this application embodiment. It is assumed that two parties, financial institution A (participant A) and financial institution B (participant B), are involved, and the training process is coordinated by a neutral central server S. The training and application process of the risk assessment model is carried out under the federated learning framework, as follows: 1. Initialization phase: The central server S initializes the parameters of a global risk assessment model (e.g., the graph convolutional network model described in Example 3) and distributes them to all participants A and B.
[0068] 2. Local Computation and Update: In each round of training, participant A and participant B each use their own isolated user data on their local servers. They first perform data acquisition (S101), data fusion preprocessing (S102), and composite feature construction (S103) as described in Example 1 to generate composite feature vectors for their respective users. Then, they use a batch of local data to perform forward and backward propagation computations on the current local model to obtain the gradient of the model parameters, which indicates how the model parameters should be adjusted to reduce prediction errors on the local data.
[0069] 3. Gradient Protection and Upload: Before uploading the calculated gradients to the central server S, each participating party performs critical privacy protection operations. These privacy protection operations include: Differential Privacy Noise Addition: To prevent gradient information from leaking sensitive information of individual users, participants first apply differential privacy techniques to the gradient vector. Specifically, a Laplace mechanism can be used, based on a Laplace distribution. Sampling noise is applied to each component of the gradient vector. The scale parameter is... , It refers to the global sensitivity of the function. This refers to a privacy budget. In this embodiment, a smaller privacy budget can be set. (e.g., 0.3) to provide strong privacy protection and ensure that the impact of individual user data on the final aggregation gradient is effectively obscured.
[0070] Homomorphic encryption: After adding noise, the participants use a public-key encryption algorithm with additive homomorphism (such as the Paillier algorithm) to encrypt the gradient vector. The characteristic of this type of encryption scheme is that performing a specific operation on two ciphertexts results in a decrypted sum equal to the sum of the two original plaintexts. .
[0071] Encrypted upload: Each participant sends the encrypted and noise-added gradients to the central server S via a secure transport layer protocol channel using the upload_encrypted_gradient (upload encrypted gradient data) signaling.
[0072] 4. Ciphertext Aggregation: The central server S receives encrypted gradients from all participants. Because the data is encrypted, the server cannot see the actual gradient values of any party, nor can it access the original user data. The server utilizes the additive homomorphism of the encryption algorithm to aggregate all received encrypted gradients in the ciphertext state (e.g., by multiplying all ciphertexts to add the plaintext), and can perform a weighted average to obtain the aggregated global encrypted gradient. This process corresponds to... Figure 4The aggregate_gradients (ciphertext aggregation) operation in the code.
[0073] 5. Model Update and Distribution: The central server S distributes the aggregated global encrypted gradient back to all participants A and B via the `distribute_global_model` signaling. Each participant decrypts the gradient locally using its own private key to obtain the noisy global average gradient, and uses this gradient to update its local model parameters.
[0074] Steps 2 through 5 constitute a complete federated learning iteration round. This process is repeated multiple times until the performance of the global model converges on the validation set. Ultimately, all participants will have a high-performance global model that integrates the data and knowledge from all parties, and no party's original data is shared or leaked throughout the entire process.
[0075] This embodiment effectively solves the "data silo" problem that is prevalent in the financial field by deeply integrating federated learning with technologies such as differential privacy and homomorphic encryption. This makes cross-institutional joint risk control possible, greatly increases the scale and diversity of data that the model can utilize, thereby enhancing the accuracy and coverage of risk assessment. At the same time, it provides a solid compliance and security foundation for the commercial implementation of the entire solution.
[0076] It is understood that the same or similar parts in the above embodiments can be referred to each other, and the contents not described in detail in some embodiments can be referred to the same or similar contents in other embodiments.
[0077] It should be noted that in the description of this invention, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this invention, unless otherwise stated, "a plurality of" means at least two.
[0078] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.
[0079] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0080] Those skilled in the art will understand that all or part of the steps of the methods implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0081] Furthermore, the functional units in the various embodiments of this invention can be integrated into a single processing module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The aforementioned storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0082] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0083] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A user financial risk assessment method based on composite features, characterized in that, The method comprises the following steps: obtaining a plurality of heterogeneous data sources associated with a user, the heterogeneous data sources at least including user behavior data reflecting individual behavior of the user and macro environment data reflecting an external environment; performing fusion preprocessing on the plurality of heterogeneous data sources to generate a uniform paradigm data set, the fusion preprocessing including standardizing mapping of formats and fields of data from different sources and repairing missing or non-standard data in the data set based on user context information; based on the uniform paradigm data set, constructing a composite feature vector representing user financial behavior, the composite feature vector at least including dynamic behavior features reflecting dynamic changes in user behavior and topological network features reflecting association relationships between the user and other entities; inputting the composite feature vector into one or more prediction models for analysis to generate a risk assessment result for representing a user risk level; when the risk assessment result meets a preset risk condition, performing a predefined risk handling action.
2. The method of claim 1, wherein, The method of constructing a composite feature vector representing user financial behavior specifically comprises: generating the dynamic behavior features by monitoring a preset trigger event in a user data stream and capturing user behavior data within a preset time window after the trigger event occurs; and generating the topological network features by constructing a user-related graph structure and calculating graph indicators using graph theory algorithms or graph neural networks.
3. The method of claim 2, wherein, The preset trigger event includes: a change in the user's monthly income exceeding a preset fluctuation threshold; or, the user logging in at a new geographic location or using a new device.
4. The method of claim 2, wherein, The user-related graph structure includes: a user-merchant bipartite graph formed by transaction relationships between the user and a merchant; or, a user-user directed graph formed by fund flow relationships between users.
5. The method of claim 1, wherein, The repairing missing or non-standard data in the data set based on user context information specifically includes: using a generative model to fill in missing values in the data set based on historical behavior patterns or occupation types of the user.
6. The method of claim 1, wherein, The prediction model is a phased ensemble learning model, and the analysis process includes: in a first phase, using a first model to perform feature selection on the composite feature vector; in a second phase, inputting the selected features into a second model for capturing temporal dependencies to generate an intermediate risk prediction; in a third phase, introducing a knowledge graph reasoning engine to match user behavior with a rule base to output the risk assessment result based on the intermediate risk prediction.
7. The method of claim 1, wherein, The predefined risk handling action includes at least one of the following: pushing warning information to a risk control system; limiting user account permissions; generating a visual risk report; blockchain notarization of risk assessment results and decision basis.
8. The method of claim 1, wherein, In the analysis process of inputting the composite feature vector into one or more prediction models to generate a risk assessment result for representing a user risk level, the analysis process is performed under a federated learning framework, specifically including: computing model updates by multiple participants locally using their own user data; The central server aggregates the encrypted model update uploaded by each participant to generate a global model update without accessing the original data; Each participant updates the local model according to the global model update.
9. The method according to claim 1 or 7, characterized in that, When the predefined risk treatment action includes block chain notarization of the risk assessment result and the basis for decision-making, the notarization specifically includes: calculating a hash value of the composite feature vector or a subset thereof; writing the hash value, the decision path of the prediction model and the risk assessment result into the block chain through the smart contract.
10. A user financial risk assessment system, characterized by, Comprise: a data acquisition module configured to acquire a plurality of heterogeneous data sources associated with a user, the heterogeneous data sources comprising at least user behavior data reflecting individual behavior of the user and macro environment data reflecting an external environment; a data fusion module configured to perform fusion preprocessing on the plurality of heterogeneous data sources to generate a unified paradigm data set, the fusion preprocessing comprising standardizing mapping of formats and fields of data from different sources, and repairing missing or non-standard data in the data set based on user context information; a feature construction module configured to construct a composite feature vector representing financial behavior of the user based on the unified paradigm data set, the composite feature vector comprising at least dynamic behavior features reflecting dynamic changes in user behavior and topological network features reflecting association relationships between the user and other entities; a risk assessment module configured to input the composite feature vector into one or more prediction models for analysis to generate a risk assessment result representing a risk level of the user; a risk treatment module configured to perform a predefined risk treatment action when the risk assessment result meets a preset risk condition.