An investor behavior risk portrait construction method based on a knowledge graph

By integrating diverse and heterogeneous data, constructing knowledge graphs, and modeling hybrid models, the problems of single data and delayed response in traditional investor behavior assessment methods have been solved, achieving efficient, secure, and interpretable risk assessment and improving the accuracy and timeliness of the assessment.

CN120782568BActive Publication Date: 2026-02-17股掌柜证券投资咨询有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511067063.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2026-02-17
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional investor behavior risk assessment methods suffer from problems such as limited data sources, single assessment dimensions, delayed response, insufficient data security, and inability to dynamically reflect market changes, making it difficult to guarantee the accuracy and timeliness of risk assessment results.

Method used

It employs multi-source heterogeneous data collection and fusion, privacy protection and data governance, construction of a three-layer entity system knowledge graph, multimodal risk feature extraction and hybrid model risk modeling, combined with a three-level early warning mechanism, and achieves risk assessment through a knowledge graph-neural network joint model.

Benefits of technology

It improves the accuracy and timeliness of risk assessment, solves the problems of single assessment dimensions and delayed response in traditional methods, realizes accurate assessment and effective management of investor behavior, and ensures data security and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120782568B_ABST
    Figure CN120782568B_ABST
Patent Text Reader

Abstract

This invention relates to the interdisciplinary field of fintech and artificial intelligence, and particularly to a method for constructing investor behavior risk profiles based on knowledge graphs. The method includes: synchronously collecting investor behavior data, public opinion data, and related market data via the PTP protocol; processing investor identity information using k-anonymization; constructing a knowledge graph comprising a three-layer entity system, establishing behavioral relationship networks and risk transmission relationship networks; fusing multimodal risk features to generate knowledge graph embedding vectors; constructing a knowledge graph-neural network joint model, achieving feature interaction through cross-compression units, and outputting four risk levels; generating a risk heatmap and executing a three-level early warning mechanism. This invention can improve the accuracy of high-risk account identification, reduce response latency, and can be applied to financial regulatory scenarios such as securities and banking to achieve real-time risk warnings.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of financial technology and artificial intelligence, and in particular to a method for constructing investor behavior risk profiles based on knowledge graphs. Background Technology

[0002] In the financial sector, accurately assessing investor behavior risk is crucial for maintaining market stability and protecting investor interests. However, traditional methods for assessing investor behavior risk have numerous limitations. Early assessment methods primarily relied on single financial data points or simple questionnaires, resulting in limited data sources and an inability to fully reflect the complexity of investor behavior. Furthermore, the lack of effective privacy protection and data governance measures fails to ensure data security and reliability. Moreover, these traditional methods cannot dynamically reflect market changes and are ill-suited to the rapidly evolving needs of the financial market.

[0003] With the diversification of financial markets, the value of massive amounts of unstructured public opinion data and real-time market fluctuation data is becoming increasingly prominent. However, traditional models, unable to effectively integrate diverse and heterogeneous data and capture dynamic risk characteristics, are gradually revealing shortcomings such as single assessment dimensions and delayed response. Although knowledge graph technology has been introduced into the financial field to model entity relationships, existing methods mostly remain at the level of static graph construction, lacking deep integration of temporal features and graph structure features; at the same time, traditional machine learning models struggle to achieve high-order interactions of multimodal features, resulting in insufficient accuracy in risk scoring.

[0004] While knowledge graphs and machine learning technologies are currently used for risk assessment, several problems remain. Data fusion is difficult, particularly integrating unstructured public opinion data with structured investor behavior data; knowledge graph construction and updates lag behind, failing to dynamically adjust to real-time market fluctuations; the accuracy of risk feature extraction and modeling is insufficient, making it difficult to accurately identify high-risk investors and potential risk events; visualization and early warning mechanisms are inadequate, and the risk response record traceability chain is incomplete, failing to meet financial regulators' requirements for interpretability and compliance. These issues make it difficult to guarantee the accuracy and timeliness of risk assessment results, hindering financial institutions' needs for precise assessment and effective management of investor behavior risks.

[0005] Therefore, this invention proposes a method for constructing investor behavior risk profiles based on knowledge graphs. Summary of the Invention

[0006] The purpose of this invention is to propose a knowledge graph-based method for constructing investor behavior risk profiles to address the problems of real-time performance, interpretability, data heterogeneity, and model dynamism in existing technologies.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for constructing investor behavior risk profiles based on knowledge graphs, comprising the following steps:

[0008] Step S1: Collect and integrate diverse and heterogeneous data. Simultaneously collect structured investor behavior data, unstructured public opinion data, and related market data, and establish a millisecond-level time synchronization benchmark through the PTP protocol.

[0009] Step S2, privacy protection and data governance, uses k-anonymization to process investor identity information and uses the isolated forest algorithm to detect abnormal investor behavior;

[0010] Step S3: Construct a knowledge graph of a three-tiered entity system including investors, financial products, and risk events, and establish a behavioral relationship network and a risk transmission relationship network;

[0011] Step S4: Multimodal risk feature extraction, fusing node attribute features, graph structure features and temporal features, and using the TransE algorithm to generate knowledge graph embedding vectors;

[0012] Step S5: Hybrid model risk modeling, constructing a knowledge graph-neural network joint model, realizing feature interaction through cross compression units, and outputting four risk levels;

[0013] Step S6: Generate a risk heatmap based on the knowledge graph and risk score R, execute a three-level early warning mechanism, and write risk response records into the knowledge graph.

[0014] The beneficial effects of the technical solution provided by this invention include at least the following:

[0015] This invention solves the problem that traditional static models cannot capture cross-layer risk transmission paths by constructing a three-layer knowledge graph of investors, financial products, and risk events, and a dynamic weight network.

[0016] This invention solves the problem of delayed risk response and improves the response speed to market volatility events by using PTP protocol time synchronization and Kafka streaming processing.

[0017] This invention integrates node attributes, graph structure, and temporal features through cross-compression units and introduces dynamic weights α and β to address the problem of insufficient feature interaction in traditional models, thereby significantly improving the accuracy of risk prediction.

[0018] This invention addresses the issues of inexplicable risk transmission paths and lack of closed-loop traceability of handling results through heat map visualization and a three-level early warning closed loop.

[0019] This invention addresses the issues of privacy protection and compliant application of financial data through k-anonymization and differential privacy technology. Attached Figure Description

[0020] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating a method for constructing an investor behavior risk profile based on a knowledge graph, as provided in an embodiment of the present invention. Detailed Implementation

[0022] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a knowledge graph-based investor behavior risk profiling method proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0024] The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.

[0025] The following description, in conjunction with the accompanying drawings, details a specific scheme for constructing an investor behavior risk profile based on a knowledge graph, as provided by this invention.

[0026] Please see Figure 1 The diagram illustrates a flowchart of a method for constructing an investor behavior risk profile based on a knowledge graph, as provided in an embodiment of the present invention. The method includes the following steps:

[0027] Step S1: Collect and integrate diverse and heterogeneous data. Simultaneously collect structured investor behavior data, unstructured public opinion data, and related market data, and establish a millisecond-level time synchronization benchmark through the PTP protocol.

[0028] Step S1 further includes the following sub-steps:

[0029] S1-1 acquires structured investor behavior data through a front-end tracking system authorized by investors. This data includes risk preference tags, activity indicators, operational entropy values, and risk assessment scores. Risk preference tags are derived from the risk level distribution of financial products investors focus on, the historical volatility of simulated trading portfolios, and their preference for setting risk warning thresholds. Activity indicators are calculated by weighting the number of market data views, research report reads, risk warning settings, and cross-sector page jumps within a unit of time. The weights are 0.4, 0.3, 0.2, and 0.1, respectively.

[0030] Operational entropy is calculated based on the information entropy of the diversity of investor behavior types, including market information acquisition, investment research, and simulated trading; risk assessment scores are obtained through standardized questionnaires authorized by investors.

[0031] S1-2, using distributed crawlers to collect unstructured public opinion data, and converting the news text in the unstructured public opinion data into attributes of risk event entities through sentiment analysis;

[0032] S1-3: Real-time acquisition of relevant market data through financial data interface. Relevant market data includes financial product risk level and market volatility events. Financial product risk level is used to define the attributes of financial product entities. Market volatility event data is used to trigger time series feature calculation.

[0033] It should be noted that the integration of structured investor behavior data, unstructured public opinion data, and related market data aims to construct a three-dimensional feature system covering behavior, sentiment, and environment, thereby addressing the one-sidedness of traditional methods that rely solely on transaction data.

[0034] PTP stands for Precision Time Protocol, which is used to achieve millisecond-level time synchronization in distributed systems, providing a foundation for subsequent time series feature analysis and avoiding risky misjudgments caused by time misalignment.

[0035] A front-end tracking system is a technical system that legally collects page interaction data triggered by investors by embedding monitoring code in the front-end code of web pages / apps.

[0036] Risk preference labels are classification labels that reflect an investor's risk tolerance and subjective inclination.

[0037] The activity index is a comprehensive indicator used to quantify the frequency of investor participation in the market, and it is calculated using a weighted summation.

[0038] Operational entropy is a measure of behavioral diversity based on information entropy theory.

[0039] The public opinion data uses the Scrapy framework to build a distributed crawler to selectively scrape news text from websites such as Reuters and Sina Finance. It then uses the FinBERT model to perform sentiment analysis and converts the results into attribute values ​​for risk event entities.

[0040] Public opinion data reflects market participants' views, sentiments, and expectations regarding the financial market, and may influence investor behavior and market trends.

[0041] It connects market data to financial data interfaces, monitors market fluctuations in real time, and stores the risk level of financial products directly as a graph node attribute.

[0042] The risk assessment score is derived from a standardized questionnaire completed by investors and is mapped to risk preference labels of four risk levels: A, B, C, and D.

[0043] Step S2, privacy protection and data governance, uses k-anonymization to process investor identity information and uses the isolated forest algorithm to detect abnormal investor behavior;

[0044] Step S2 further includes the following sub-steps:

[0045] S2-1, k-anonymize investor identity information, with anonymized group size k≥10;

[0046] S2-2 detects abnormal investor behavior based on the isolated forest algorithm, with the judgment condition being that the activity index or operational entropy value deviates from the mean by ±3σ.

[0047] It should be noted that k-anonymization is a data privacy protection technology that aims to prevent privacy leaks by processing sensitive data to ensure that no individual record in a dataset can be uniquely identified.

[0048] k-anonymization, investors are grouped by asset size and occupation type, with ≥10 people in each group.

[0049] The Isolation Forest algorithm can use the anomaly detection results of investor behavior data as a preliminary screening indicator for risk scoring, reducing noise in subsequent modeling.

[0050] Step S3: Construct a knowledge graph of a three-tiered entity system including investors, financial products, and risk events, and establish a behavioral relationship network and a risk transmission relationship network;

[0051] Step S3 further includes the following sub-steps:

[0052] S3-1 defines the attributes of the investor entity, including investor ID, risk preference label, and operational entropy value;

[0053] S3-2 defines the attributes of a financial product entity, including product type, historical volatility, and credit rating.

[0054] S3-3 defines the attributes of a risk event entity, including event type, scope of impact, and timestamp;

[0055] S3-4, Establish two types of relationship networks, including:

[0056] The behavioral relationship network reflects the strength of behavioral interactions between investors and financial products. The behavioral relationship weights are positively correlated with activity indicators. The benchmark value of the behavioral relationship weights is dynamically linked to the diversity of products investors focus on and the industry average, specifically manifested as follows:

[0057] When the concentration of investors' attention on a certain type of product is higher than the industry average, the benchmark weighting coefficient increases;

[0058] When investors prioritize product diversity above the industry average, the benchmark weighting coefficient is lowered.

[0059] Risk transmission relationship network is used to reflect the indirect propagation of risk events. The strength of risk transmission relationship is quantitatively assessed by the number of common risk events.

[0060] The knowledge graph is updated incrementally when the market volatility exceeds 20%, with an update delay of ≤1 second. The Apache Kafka streaming platform is used to monitor market volatility events.

[0061] The knowledge graph is stored using the Neo4j graph database, and the Cypher language is supported for querying risk transmission paths within 3 hops.

[0062] It should be noted that,

[0063] Neo4j's efficient query support for risk transmission paths within 3 hops enables risk control personnel to quickly locate the source of risk and potentially affected groups.

[0064] The three-tier entity system supports cross-tier risk transmission analysis from individual behavior to product risk to market events, and can capture more complex risk relationships compared to the traditional two-dimensional model.

[0065] Kafka's second-level update mechanism enables knowledge graphs to respond to market changes in real time, avoiding the failure of risk assessment caused by the lag of static graphs.

[0066] Step S4: Multimodal risk feature extraction, fusing node attribute features, graph structure features and temporal features, and using the TransE algorithm to generate knowledge graph embedding vectors;

[0067] Step S4 further includes the following sub-steps:

[0068] S4-1, node attribute characteristics include activity index, operational entropy value and risk assessment score;

[0069] S4-2, the graph structure features include betweenness centrality and risk transmission efficiency; betweenness centrality is calculated through the behavioral relationship network; risk transmission efficiency is calculated by dividing the risk transmission relationship strength by the shortest path hop count, where the shortest path hop count is calculated using Neo4j's shortestPath function;

[0070] S4-3, the time-series characteristics include the rate of change in activity within 24 hours before and after market fluctuations; the benchmark value of the rate of change in activity is taken from the point when the market fluctuation event was triggered;

[0071] S4-4 uses the TransE algorithm to generate 128-dimensional knowledge graph embedding vectors to represent the semantic relationships between investors, financial products, and risk events; the embedding vectors are part of the graph structure feature matrix.

[0072] It should be noted that node attribute features reflect investors' static risk propensity, graph structure features depict their pivotal role in the behavioral network, and time series features capture dynamic market responses. The fusion of these three types of features addresses the shortcomings of traditional models that "emphasize history over real-time" and "emphasize individuals over the network."

[0073] The embedding vectors generated by the TransE algorithm can be input into neural networks to transform the symbolic knowledge of knowledge graphs into numerical features, thus achieving a combination of symbolic AI and numerical AI.

[0074] Betweenness centrality is a feature of graph structure that reflects the pivotal role of investors in behavioral relationship networks. A higher value indicates that the investor is more likely to become a key node in risk transmission and is used for node size mapping in risk heatmaps.

[0075] Time-series features refer to the rate of change in activity within 24 hours before and after a market volatility event. They are used to capture investors' real-time responses to market dynamics and are an important component of multimodal feature fusion.

[0076] The generated embedding vectors, as one of the inputs to the knowledge graph-neural network joint model, together with node attributes and temporal features, constitute a multi-dimensional feature matrix, providing a data foundation for higher-order interactions of cross-compression units.

[0077] Step S5: Hybrid model risk modeling, constructing a knowledge graph-neural network joint model, realizing feature interaction through cross compression units, and outputting four risk levels;

[0078] Step S5 further includes the following sub-steps:

[0079] S5-1, the knowledge graph-neural network joint model consists of an embedding layer, a cross-compression unit, and a classification layer;

[0080] S5-2, the input to the embedding layer includes node attribute feature vectors, graph structure feature matrices, and temporal feature sequences;

[0081] S5-3, the cross-compression unit generates a 128-dimensional high-order feature representation h, which is used to calculate the risk score R;

[0082] S5-4, the classification layer outputs four risk levels, including A, B, C and D;

[0083] The cross-compression unit includes:

[0084] The dynamic weight fusion module fuses basic scoring terms using trainable parameters α and β. With logarithmic transformation term The logarithmic transformation term is a nonlinear term obtained by performing a logarithmic transformation on the higher-order characteristic representation h. The specific formula is as follows:

[0085]

[0086] in, ϵ is the logarithmic transformation term, h is the higher-order feature representation, W is the trainable weight matrix, and ϵ is the smoothing factor.

[0087] The parameter constraint module uses the gradient projection method to maintain α+β≡1 during backpropagation.

[0088] The standardized output module outputs a 128-dimensional high-order feature representation h normalized by LayerNorm.

[0089] The formula for calculating the risk score R is:

[0090]

[0091] Where R is the risk score; α and β are weight parameters obtained through model training, satisfying α+β=1; Here, h represents the logarithmic transformation term, and h is the higher-order feature representation generated by the cross-compression unit.

[0092] The calculation formula is:

[0093]

[0094] in, The calculation is based on node attribute features, and the specific formula is as follows:

[0095] = ⋅Standardized value of activity index + ⋅Standardized value of operational entropy + Standardized values ​​of risk assessment scores

[0096] Among them, weight It is determined through training with historical data;

[0097] Based on graph structure features, the specific formula is as follows:

[0098] = ⋅Standardized value of intermediary centrality+ Standardized value of risk transmission path length

[0099] Among them, weight This was determined through knowledge graph analysis;

[0100] Based on time-series features, the specific formula is as follows:

[0101] =Standardized value of the rate of change in market activity 24 hours before and after market fluctuations.

[0102] It should be noted that in the knowledge graph-neural network joint model, the knowledge graph provides prior risk associations, while the neural network mines nonlinear feature interactions, thus improving prediction accuracy compared to a single model.

[0103] The cross-compression unit adaptively adjusts the contributions of the linear basic score and the nonlinear logarithmic transformation term through dynamic weights α and β, achieving a flexible response to different risk scenarios. For example, during periods of sharp market fluctuations, the β value automatically increases, enhancing the logarithmic transformation term's ability to capture sudden risks; while during periods of stability, the α value dominates, maintaining model stability.

[0104] The parameter constraint module ensures that α+β≡1, avoiding overfitting caused by weight shift.

[0105] LayerNorm normalization ensures that the distribution of high-order feature representation h remains stable during training, solving the gradient vanishing / exploding problem in deep neural networks and accelerating model convergence.

[0106] The output risk levels, from A to D, directly correspond to the three-tiered early warning mechanism in step S6. Level A investors can enjoy regular services, Level B investors receive email alerts, Level C investors are restricted from high-risk transactions, and Level D investors have their accounts frozen, forming a closed-loop management system of risk assessment and tiered response.

[0107] Step S6: Generate a risk heatmap based on the knowledge graph and risk score R, execute a three-level early warning mechanism, and write risk response records into the knowledge graph.

[0108] Step S6 further includes the following sub-steps:

[0109] S6-1 generates a risk heatmap based on a knowledge graph and risk score R, where:

[0110] The size of a node is positively correlated with the centrality of the investor entity, reflecting its role as a risk transmission hub in the behavioral network;

[0111] The depth of edge color is positively correlated with the strength of risk transmission, highlighting the risk transmission path;

[0112] The node color mapping risk score R has four risk levels: A is green, B is yellow, C is orange, and D is red.

[0113] S6-2, implements a three-level early warning mechanism:

[0114] Level B, risk score 60≤R<75, daily email reminders will be sent;

[0115] Grade C, risk score 75≤R<90, restricted leveraged trading privileges;

[0116] Level D, risk score R≥90, freeze the account and generate a risk report containing the risk transmission path; call Neo4j's shortestPath function to generate a risk transmission path within 3 hops;

[0117] S6-3, based on the three-level early warning mechanism, risk response records are automatically written into the risk event entity in the knowledge graph, forming a closed-loop traceability chain, including:

[0118] Warning level, trigger time, and risk score R value;

[0119] Risk transmission path ID and processing result.

[0120] It should be noted that risk heatmaps use visualization techniques to transform abstract risk quantification results into intuitive graphs, reducing the decision-making costs for risk control personnel.

[0121] The early warning mechanism employs a tiered intervention strategy to balance risk control and investor experience. Level A involves no early warning response; Level B involves email alerts to notify investors of risks in a non-intrusive manner; Level C restricts leveraged trading to reduce potential losses; and Level D involves account freezing to protect investor assets in extreme risk scenarios.

[0122] Risk response records are written into a knowledge graph to form a complete closed loop of assessment, early warning, handling, and feedback.

[0123] The combination of risk heat maps and a three-level early warning mechanism enables risk control decisions to shift from experience-driven to data visualization-driven.

[0124] In this way, a method for constructing investor behavior risk profiles based on knowledge graphs can be achieved.

[0125] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for constructing an investor behavioral risk portrait based on a knowledge graph, characterized in that, The method includes: Step S1: Collect and integrate diverse and heterogeneous data. Simultaneously collect structured investor behavior data, unstructured public opinion data, and related market data, and establish a millisecond-level time synchronization benchmark through the PTP protocol. Step S2, privacy protection and data governance, uses k-anonymization to process investor identity information and uses the isolated forest algorithm to detect abnormal investor behavior; Step S3: Construct a knowledge graph of a three-tiered entity system including investors, financial products, and risk events, and establish a behavioral relationship network and a risk transmission relationship network; Step S4: Multimodal risk feature extraction, fusing node attribute features, graph structure features and temporal features, and using the TransE algorithm to generate knowledge graph embedding vectors; Step S5: Hybrid model risk modeling, constructing a knowledge graph-neural network joint model, realizing feature interaction through cross compression units, and outputting four risk levels; Step S6: Generate a risk heat map based on the knowledge graph and risk score R, execute a three-level early warning mechanism, and write risk response records into the knowledge graph; Step S5 further includes the following sub-steps: S5-1, the knowledge graph-neural network joint model consists of an embedding layer, a cross-compression unit, and a classification layer; S5-2, the input to the embedding layer includes node attribute feature vectors, graph structure feature matrices, and temporal feature sequences; S5-3, the cross-compression unit generates a 128-dimensional high-order feature representation h, which is used to calculate the risk score R; S5-4, the classification layer outputs four risk levels, including level A, level B, level C and level D; The cross-compression unit includes: The dynamic weight fusion module fuses the basic score items through trainable parameters a and β and a logarithmic transformation term The logarithmic transformation term is a nonlinear term obtained by logarithmic transformation on the high-order feature representation h, and the specific formula is: wherein, is a log-transformed term, h is a high-order feature representation, W is a trainable weight matrix, and e is a smoothing factor. The parameter constraint module uses the gradient projection method to maintain α+β≡1 during backpropagation. The standardized output module outputs a 128-dimensional high-order feature representation h normalized by LayerNorm.

2. The method for constructing an investor behavior risk profile based on knowledge graphs according to claim 1, characterized in that: Step S1 further includes the following sub-steps: S1-1, obtain structured investor behavior data through a front-end tracking system authorized by the investor, the investor behavior data including risk preference tags, activity indicators, operational entropy values ​​and risk assessment scores; The risk preference label is generated by inferring the investor’s preference for the risk level distribution of financial products they pay attention to, the historical volatility of simulated trading portfolios, and risk warning threshold settings. The activity index is calculated by weighting the number of times market data is viewed, the number of times research reports are read, the number of times risk warnings are set, and the number of cross-sector page jumps within a unit of time. The weights used in the weighting calculation are 0.4, 0.3, 0.2, and 0.1, respectively. The operational entropy value is calculated based on the information entropy of the diversity of investor behavior types, including market information acquisition behavior, investment research behavior, and simulated trading behavior; the risk assessment score is obtained through a standardized questionnaire authorized by the investor. S1-2, Distributed crawlers are used to collect unstructured public opinion data, and the news text in the unstructured public opinion data is transformed into the attributes of risk event entities through sentiment analysis; S1-3, real-time acquisition of relevant market data through a financial data interface, the relevant market data including the risk level of financial products and market volatility events; the risk level of financial products is used to define the attributes of financial product entities; the market volatility event data is used to trigger time-series feature calculations.

3. The method for constructing an investor behavior risk profile based on knowledge graphs according to claim 1, characterized in that: Step S2 further includes the following sub-steps: S2-1, k-anonymize investor identity information, with anonymized group size k≥10; S2-2 detects abnormal investor behavior based on the isolated forest algorithm, with the judgment condition being that the activity index or operational entropy value deviates from the mean by ±3σ.

4. The method for constructing an investor behavior risk profile based on knowledge graphs according to claim 1, characterized in that: Step S3 further includes the following sub-steps: S3-1 defines the attributes of the investor entity, including investor ID, risk preference label, and operational entropy value; S3-2 defines the attributes of a financial product entity, including product type, historical volatility, and credit rating. S3-3 defines the attributes of a risk event entity, including event type, scope of impact, and timestamp; S3-4, Establish two types of relationship networks, including: The behavioral relationship network reflects the intensity of behavioral interactions between investors and financial products. The behavioral relationship weights are positively correlated with activity indicators. The benchmark values ​​of these behavioral relationship weights are dynamically linked to the diversity of products investors focus on and the industry average, specifically manifested as follows: When the concentration of investors' attention on a certain type of product is higher than the industry average, the benchmark weighting coefficient increases; When investors prioritize product diversity above the industry average, the benchmark weighting coefficient is lowered. Risk transmission relationship network is used to reflect the indirect propagation of risk events. The strength of risk transmission relationship is quantitatively assessed by the number of common risk events.

5. The method for constructing an investor behavior risk profile based on knowledge graphs according to claim 4, characterized in that: The knowledge graph is updated incrementally when the market volatility exceeds 20%, with an update delay of ≤1 second. The Apache Kafka streaming platform is used to monitor market volatility events. The knowledge graph is stored using the Neo4j graph database, and the Cypher language is supported for querying risk transmission paths within 3 hops.

6. The method for constructing an investor behavior risk profile based on knowledge graphs according to claim 1, characterized in that: Step S4 further includes the following sub-steps: S4-1, the node attribute features include activity index, operational entropy value and risk assessment score; S4-2, the graph structure features include betweenness centrality and risk transmission efficiency; the betweenness centrality is calculated through a behavioral relationship network; the risk transmission efficiency is calculated by dividing the risk transmission relationship strength by the number of hops on the shortest path, wherein the number of hops on the shortest path is calculated using the shortestPath function of Neo4j; S4-3, the time-series feature includes the rate of change of activity within 24 hours before and after the market fluctuation; the benchmark value of the rate of change of activity is taken from the trigger point of the market fluctuation event; S4-4 uses the TransE algorithm to generate 128-dimensional knowledge graph embedding vectors to represent the semantic relationships between investors, financial products and risk events; the embedding vectors are part of the graph structure feature matrix.

7. The method for constructing an investor behavior risk profile based on knowledge graphs according to claim 1, characterized in that: The formula for calculating the risk score R is: Wherein, R is a risk score; a, β are weight parameters obtained through model training, and a+β=1 is satisfied; is a log-transformed term, and h is a high-order feature representation generated by the cross compression unit. The calculation formula is: wherein, Based on the node attribute feature calculation, the specific formula is: = ⋅ Activity Index Normalized Value ⋅ Operational Entropy Normalized Value ⋅ Risk Assessment Score Normalized Value Among them, weight It is determined through training with historical data; The Based on the aforementioned graph structure features, the specific formula is as follows: = ⋅Standardized value of intermediary centrality+ Standardized value of risk transmission path length Among them, weight This was determined through knowledge graph analysis; The Based on the aforementioned time-series characteristics, the specific formula is as follows: =Standardized value of the rate of change in market activity 24 hours before and after market fluctuations.

8. The method for constructing an investor behavior risk profile based on knowledge graphs according to claim 1, characterized in that: Step S6 further includes the following sub-steps: S6-1, generates a risk heatmap based on knowledge graph and risk score R, where: The size of a node is positively correlated with the centrality of the investor entity, reflecting its role as a risk transmission hub in the behavioral network; The depth of edge color is positively correlated with the strength of risk transmission, highlighting the risk transmission path; The node color mapping risk score R has four risk levels: A is green, B is yellow, C is orange, and D is red. S6-2, implements a three-level early warning mechanism: Level B, risk score 60≤R<75, daily email reminders will be sent; Grade C, risk score 75≤R<90, restricted leveraged trading privileges; Level D, risk score R≥90, freeze the account and generate a risk report containing the risk transmission path; call Neo4j's shortestPath function to generate the risk transmission path within 3 hops; S6-3, The risk response records generated based on the three-level early warning mechanism are automatically written into the risk event entities in the knowledge graph, forming a closed-loop traceability chain, including: Warning level, trigger time, and risk score R value; Risk transmission path ID and processing result.

Citation Information

Patent Citations

  • Collaborative knowledge perception enhanced network recommendation method

    CN113420233A

  • User portrait determination method and device based on knowledge graph

    CN119829829A

  • Multi-source heterogeneous data fusion knowledge graph method and system

    CN120181198A