Financial product recommendation method and computer equipment based on probabilistic knowledge graph
By identifying target user groups through probabilistic knowledge graphs and recommending financial products based on the probability distribution of users and products, we can solve the problems of high development and maintenance costs and poor interpretability of existing systems, and achieve efficient and personalized financial product recommendations.
Patent Information
- Application Number
- CN202410785778.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-18
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2044-06-18
AI Technical Summary
Existing financial product recommendation systems rely on large amounts of labeled data and complex models, resulting in high development and maintenance costs and a lack of explainability.
A method based on probabilistic knowledge graph is adopted to generate target product nodes by obtaining target product information, determine the product node center, and identify the target user node center based on probability distribution, send recommendation information to them, and combine the user's purchase history data and cluster analysis to improve the accuracy and personalization of recommendations.
It reduces dependence on hardware configuration, improves the accuracy and personalization of recommendations, reduces development and maintenance costs, and enhances the interpretability and adaptability of the system.
Smart Images

Figure CN118710358B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular to a financial product recommendation method and computer device based on a probabilistic knowledge graph. Background Art
[0002] With the rapid development of financial technology, financial product recommendation systems have become a crucial tool for improving customer service and market competitiveness for banks and financial institutions. Existing systems often rely on large amounts of labeled data and complex models, resulting in high development and maintenance costs and a lack of interpretability. Summary of the Invention
[0003] The purpose of this disclosure is to provide a financial product recommendation method and computer device based on probabilistic knowledge graphs to reduce data requirements, reduce dependence on hardware configuration, and improve the interpretability of recommendations.
[0004] In a first aspect, a method for recommending financial products based on a probabilistic knowledge graph is provided, the method comprising:
[0005] Obtaining target product information, where the target product information is used to describe a target financial product;
[0006] Generate a corresponding target product node according to the target product information;
[0007] Determining a target product node center to which the target product node belongs based on a predetermined probabilistic knowledge graph;
[0008] Determine the probability distribution of the target product node center with respect to each user node center;
[0009] Determine the target user node center corresponding to the highest probability value in the probability distribution;
[0010] Send recommendation information for the target financial product to a target user corresponding to at least one user node under the target user node center.
[0011] Optionally, the method comprises:
[0012] Determine a first user node center among the user node centers that has the highest average purchase amount of the financial product;
[0013] The recommendation information is sent to a target user corresponding to at least one user node under the first user node center.
[0014] Optionally, determining the probabilistic knowledge graph includes:
[0015] Obtaining purchase history data of a user for a financial product, the purchase history data including user information describing the user and product information describing the financial product, wherein the number of the user and the number of the financial product are both greater than or equal to 1;
[0016] Constructing an initial probabilistic knowledge graph based on the user information and the product information, the initial probabilistic knowledge graph including a first knowledge graph and a second knowledge graph, the first knowledge graph including at least one user node, the second knowledge graph including at least one product node, and a fully connected relationship between the user node and the product node;
[0017] Extracting probability features for each of the user nodes and each of the product nodes from the purchase history data;
[0018] Fitting the probability features to obtain parameterized description information for each of the user nodes and each of the product nodes;
[0019] Establishing a multidimensional feature library according to the parameterized description information;
[0020] The probabilistic knowledge graph is determined based on the multidimensional feature library.
[0021] Optionally, determining the probabilistic knowledge graph according to the multidimensional feature library includes:
[0022] Clustering the user nodes and the product nodes according to the parameterized description information of each user node and each product node in the multidimensional feature library to obtain at least one product node center and at least one user node center;
[0023] The probabilistic knowledge graph is generated based on the at least one product node center and the at least one user node center.
[0024] Optionally, clustering the user nodes and the product nodes includes at least one of the following:
[0025] Performing density-based clustering on the user nodes and the product nodes;
[0026] Performing hierarchical clustering on the user nodes and the product nodes.
[0027] Optionally, obtaining the user's purchase history data of financial products includes:
[0028] Acquiring a plurality of first historical data of the user for the financial product in a plurality of different time periods;
[0029] The determination of the probabilistic knowledge graph further includes:
[0030] Determining a plurality of probabilistic knowledge graphs based on the plurality of first historical data;
[0031] The determining, based on a predetermined probabilistic knowledge graph, a target product node center to which the target product node belongs, includes:
[0032] Based on any one or more probabilistic knowledge graphs among the multiple probabilistic knowledge graphs, determine one or more target product node centers to which the target product node belongs.
[0033] Optionally, the multiple different time periods include a most recent time period corresponding to the current time, and the method includes:
[0034] Performing similarity analysis on the multiple probabilistic knowledge graphs to determine the similarity between any two of the probabilistic knowledge graphs;
[0035] The determining, based on a predetermined probabilistic knowledge graph, a target product node center to which the target product node belongs, includes:
[0036] Based on the first probabilistic knowledge graph corresponding to the most recent time period and the second probabilistic knowledge graph that is most similar to the first probabilistic knowledge graph, determine the first product node center and the second product node center to which the target product node belongs. The first product node center is the target product node center corresponding to the first probabilistic knowledge graph, and the second product node center is the target product node center corresponding to the second probabilistic knowledge graph.
[0037] Optionally, the method comprises:
[0038] The purchase history data is preprocessed, and the preprocessing includes at least one of the following: abnormal value processing, repeated value processing, and encoding processing.
[0039] Optionally, the method comprises:
[0040] Obtain target user information, where the first user information is used to describe the user to be recommended;
[0041] Generate a target user node according to the target user information;
[0042] Determining, based on the probabilistic knowledge graph, a second user node center to which the target user node belongs;
[0043] Determine a first probability distribution of the second user node center with respect to each product node center;
[0044] Determine the third product node center corresponding to the highest probability value in the first probability distribution;
[0045] Send recommendation information of the financial product corresponding to at least one product node under the third product node center to the user to be recommended.
[0046] In a second aspect, a computer device is provided, comprising a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the financial product recommendation method based on the probabilistic knowledge graph described in the first aspect.
[0047] Compared with the existing technology, the beneficial effects provided by the present disclosure include: by obtaining target product information that describes the characteristics of the target financial product, and generating a target product node based on the target product information, the product node center to which the target product node belongs in a predetermined probabilistic knowledge graph can be determined, and the probability distribution of the target product node center for each user node center can be further determined, from which the user node center with the highest probability of purchasing the target financial product can be determined, and corresponding recommendation information can be sent to each user under the user node center. The target user group can be identified through the probabilistic knowledge graph, and specific financial products can be recommended to these users, which can effectively improve the accuracy and personalization of the recommendation. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be considered as limiting the scope. Those skilled in the art can also derive other relevant drawings based on these drawings without inventive effort.
[0049] Figure 1 A flowchart of a method for recommending financial products based on a probabilistic knowledge graph according to an embodiment of the present disclosure;
[0050] Figure 2 A schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0051] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, not all of them. Generally, the components of the embodiments of the present disclosure described and shown in the drawings herein can be arranged and designed in a variety of different configurations.
[0052] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0053] In order to solve the technical problems in the above background technology, Figure 1 The following is a flow chart of a method for recommending financial products based on a probabilistic knowledge graph, which is provided in an embodiment of the present disclosure. The method for recommending financial products based on a probabilistic knowledge graph is described in detail. The method may include the following steps:
[0054] Step S101: Acquire target product information, where the target product information is used to describe a target financial product.
[0055] The target product information may be a description of the characteristics of the target financial product.
[0056] In this step, detailed information about the target financial product is collected, including but not limited to product name, interest rate, investment period, risk level, product features, etc. This information will be used to generate the target product node and serve as the starting point for the recommendation process.
[0057] Step S102: Generate a corresponding target product node according to the target product information.
[0058] The target product node may include a multi-dimensional feature vector corresponding to the target product information.
[0059] Step S103: determining the target product node center to which the target product node belongs based on a predetermined probabilistic knowledge graph.
[0060] It is understandable that the product node center may be obtained by clustering multiple product nodes, and the product node center may be the product node at the most central position in a cluster.
[0061] Optionally, the target product node center to which the target product node belongs may be determined based on the similarity between the target product node and a plurality of product center nodes.
[0062] Step S104: determine the probability distribution of the target product node center with respect to each user node center.
[0063] The probability distribution can be used to indicate the probability that the user corresponding to each user node center will purchase the target financial product. This probability value can be calculated based on the parameterized description information of each user node center in the probabilistic knowledge graph and the parameterized description information corresponding to the target product node center.
[0064] Step S105: determining the target user node center corresponding to the highest probability value in the probability distribution.
[0065] Step S106: Send recommendation information for the target financial product to the target user corresponding to at least one user node under the target user node center.
[0066] The recommendation information for the target financial product may be sent to the target user via SMS, email, phone call, etc., which is not limited in the embodiments of the present disclosure.
[0067] In an embodiment of the present disclosure, target product information that describes the characteristics of a target financial product can be obtained, and a target product node can be generated based on the target product information, thereby determining the product node center to which the target product node belongs in a predetermined probabilistic knowledge graph. The probability distribution of the target product node center with respect to each user node center can be further determined, from which the user node center with the highest probability of purchasing the target financial product can be determined, and corresponding recommendation information can be sent to each user under the user node center. The target user group can be identified through the probabilistic knowledge graph, and specific financial products can be recommended to these users, which can effectively improve the accuracy and personalization of the recommendations.
[0068] In some embodiments, the method comprises:
[0069] Determine a first user node center among the user node centers that has the highest average purchase amount of the financial product;
[0070] The recommendation information is sent to a target user corresponding to at least one user node under the first user node center.
[0071] With the above scheme, the average purchase amount of users is taken into account in the recommendation process, and the user group with the strongest purchasing power is selected for recommendation, thereby improving the economic benefits of the recommendation.
[0072] In some embodiments, determining the probabilistic knowledge graph includes:
[0073] Obtaining purchase history data of a user for a financial product, the purchase history data including user information describing the user and product information describing the financial product, wherein the number of the user and the number of the financial product are both greater than or equal to 1;
[0074] Constructing an initial probabilistic knowledge graph based on the user information and the product information, the initial probabilistic knowledge graph including a first knowledge graph and a second knowledge graph, the first knowledge graph including at least one user node, the second knowledge graph including at least one product node, and a fully connected relationship between the user node and the product node;
[0075] Extracting probability features for each of the user nodes and each of the product nodes from the purchase history data;
[0076] Fitting the probability features to obtain parameterized description information for each of the user nodes and each of the product nodes;
[0077] Establishing a multidimensional feature library according to the parameterized description information;
[0078] The probabilistic knowledge graph is determined based on the multidimensional feature library.
[0079] It is understood that the process of determining a probabilistic knowledge graph can include data collection, initial graph construction, feature extraction, parameterized description, and the establishment of a multidimensional feature library. The multidimensional feature library can be established based on the parameterized description information to create a multidimensional feature vector for each user node and product node. These feature vectors will be used for subsequent cluster analysis.
[0080] An initial probabilistic knowledge graph can be constructed based on the preprocessed data. This graph consists of user nodes and product nodes, with a fully connected relationship between them. Probabilistic features are extracted from the purchase history data, such as the probability of a user purchasing a specific product and the probability of a product being purchased by a specific user. These probabilistic features are then fitted using a parameterized probabilistic model (e.g., a Gaussian mixture model) to obtain parameterized descriptions of each node.
[0081] In some embodiments, determining the probabilistic knowledge graph based on the multidimensional feature library includes:
[0082] Clustering the user nodes and the product nodes according to the parameterized description information of each user node and each product node in the multidimensional feature library to obtain at least one product center node and at least one user center node;
[0083] The probabilistic knowledge graph is generated based on the at least one product center node and the at least one user center node.
[0084] It further explains how to determine the probabilistic knowledge graph based on the multidimensional feature library and generate central nodes by clustering user nodes and product nodes. This helps to identify and build more accurate user and product groups, thereby improving the relevance of recommendations.
[0085] During the implementation process, different clustering algorithms, parameterized probability models, and recommendation strategies can be selected according to specific needs and data characteristics to adapt to different application scenarios.
[0086] In some embodiments, clustering the user nodes and the product nodes includes at least one of the following:
[0087] Performing density-based clustering on the user nodes and the product nodes;
[0088] Performing hierarchical clustering on the user nodes and the product nodes.
[0089] For example, density-based clustering can be, for example, BDSCAN (Density-Based Spatial Clustering of Applications with Noise), which is a clustering algorithm for identifying clusters of arbitrary shapes and has good robustness to noise points. It identifies clusters by measuring the density distribution of data points, dividing areas with higher density into clusters, and treating areas with lower density as noise or points separated from other clusters. Hierarchical clustering can be, for example, AGNES (A Gathering of Nearest Neighbors-based Ensemble System), which belongs to the agglomerative hierarchical clustering method. It calculates the distance between data points and gradually merges the nearest points or clusters until a predetermined number of clusters or other termination conditions are reached. The AGNES algorithm is simple, fast, and suitable for large-scale data sets.
[0090] The two algorithms mentioned above each have their own advantages and are suitable for different data characteristics and application scenarios. DBSCAN is suitable for datasets with irregular or noisy clusters, while AGNES is suitable for data analysis tasks that require generating cluster hierarchies. In financial product recommendation systems, the appropriate clustering algorithm can be selected based on user and product characteristics to optimize recommendation results.
[0091] In some embodiments, obtaining the user's purchase history data for financial products includes:
[0092] Acquiring a plurality of first historical data of the user for the financial product in a plurality of different time periods;
[0093] The determination of the probabilistic knowledge graph further includes:
[0094] Determining a plurality of probabilistic knowledge graphs based on the plurality of first historical data;
[0095] The determining, based on a predetermined probabilistic knowledge graph, a target product node center to which the target product node belongs, includes:
[0096] Based on any one or more probabilistic knowledge graphs among the multiple probabilistic knowledge graphs, determine one or more target product node centers to which the target product node belongs.
[0097] Among them, by considering data from multiple time periods, the recommendation system can adapt to market changes and improve the timeliness of recommendations.
[0098] In some embodiments, the multiple different time periods include a most recent time period corresponding to the current time, and the method includes:
[0099] Performing similarity analysis on the multiple probabilistic knowledge graphs to determine the similarity between any two of the probabilistic knowledge graphs;
[0100] The determining, based on a predetermined probabilistic knowledge graph, a target product node center to which the target product node belongs, includes:
[0101] Based on the first probabilistic knowledge graph corresponding to the most recent time period and the second probabilistic knowledge graph that is most similar to the first probabilistic knowledge graph, determine the first product center node and the second product center node to which the target product node belongs, the first product center node being the target product center node corresponding to the first probabilistic knowledge graph, and the second product center node being the target product center node corresponding to the second probabilistic knowledge graph.
[0102] In some embodiments, similarity analysis of probabilistic knowledge graphs can be used for the following: Cross-time period comparison: By calculating the similarity between PKGs trained in different time periods, the system can identify patterns and trends in customer purchasing behavior at different points in time. This helps understand market dynamics and customer needs that change over time. Product recommendation relevance analysis: By comparing PKGs from the current time period with those from previous time periods, the system can identify historical periods that are most similar to current market conditions. This helps the recommendation system learn from successful historical recommendation models, thereby improving the relevance and accuracy of recommendations. Market trend prediction: Similarity calculation can reveal market trends and cyclicalities, helping financial institutions predict future market trends and prepare corresponding financial products or marketing strategies in advance. User behavior pattern recognition: By analyzing the similarity between different PKGs, the system can identify specific user behavior patterns, such as whether certain products are more popular during a specific time period or whether certain user groups are more likely to purchase specific financial products under certain conditions. Product strategy adjustment: Financial institutions can adjust financial product design and promotion strategies based on the results of PKG similarity calculations, such as adding new financial products that are similar to successful products from the period of high-similarity PKGs. Risk management: By analyzing the similarities between different PKGs, financial institutions can identify potential market risks. For example, if the current PKG is similar to a period in history with higher risks, the institution can take measures in advance to avoid risks. Enhance the generalization ability of the model: By considering the similarities of different time periods, the recommendation system can improve its generalization ability, not only limited to performance on a specific data set, but also able to make reasonable recommendations under different market conditions. Optimize resource allocation: Financial institutions can optimize the allocation of marketing and product design resources based on the results of PKG similarity calculations, and invest more resources in products that are highly similar to current market conditions and have performed well in history. Using the above solution, similarity analysis is introduced, and the period in historical data that is most similar to the current market conditions is used to guide recommendations, thereby enhancing the adaptability and accuracy of the recommendation system.
[0103] This approach expands the scope of data acquisition to include user purchase history data from multiple time periods, and uses this data to generate multiple probabilistic knowledge graphs. This allows the system to consider the impact of time, improving the timeliness and adaptability of recommendations.
[0104] In some embodiments, the method comprises:
[0105] The purchase history data is preprocessed, and the preprocessing includes at least one of the following: abnormal value processing, repeated value processing, and encoding processing.
[0106] By adopting the above solution and preprocessing the purchase history data, the data quality and model reliability are improved, thereby improving the overall performance of the recommendation system.
[0107] It is understandable that not only can the product be recommended to the user based on the probabilistic knowledge graph according to the product information of the product, but the corresponding product can also be recommended to the user according to the user information of the user.
[0108] In some embodiments, the method comprises:
[0109] Obtain target user information, where the first user information is used to describe the user to be recommended;
[0110] Generate a target user node according to the target user information;
[0111] Determining, based on the probabilistic knowledge graph, a second user node center to which the target user node belongs;
[0112] Determine a first probability distribution of the second user node center with respect to each product node center;
[0113] Determine the third product node center corresponding to the highest probability value in the first probability distribution;
[0114] Send recommendation information of the financial product corresponding to at least one product node under the third product node center to the user to be recommended.
[0115] In the above scheme, a reverse recommendation strategy is provided, that is, recommending financial products based on target user information, which increases the flexibility of the recommendation system and allows product recommendations from the user's perspective.
[0116] In order to enable those skilled in the art to better understand the overall solution provided by the present disclosure, the present disclosure also provides the following embodiments.
[0117] In this implementation, a probabilistic grammar model can be used to capture the underlying patterns in bank customers' purchasing behavior for financial products, cluster financial product purchasing customer groups, and recommend target customer groups for specific products. This solution can include a model training phase and a model usage phase.
[0118] During the model training phase, the probabilistic grammar model can be trained with the user's financial product purchase data represented by <user, user asset level, user consumption level, user credit level, user hidden asset level, user income level, product, product interest rate, product installment>.
[0119] Given a certain amount of data over a period of time, this paper uses a probabilistic grammar model defined by a probabilistic knowledge graph (PKG). The graph can be divided into two layers, one for financial products and the other for purchasing users, with a fully connected relationship between the nodes in the two layers.
[0120] The model training phase mainly involves statistical and parameterized modeling of various probabilities. Common probabilities include but are not limited to:
[0121] 1) The probability that node Ui is connected to Pj is P0(Pj|Ui), where the total number of edges between U and P is equal to the total number of valid training data, which can be used for statistics. For example, P0(Pj|Ui) =<Ui,…,Pj,…,…> The number of times the sample appears / the total number of samples in the training set.
[0122] 2) The prior probability of a node state occurring, for example, the probability of a state occurring at time t, P1(U_t), where U_t represents the state at time t, which can be s = {the number of financial products purchased by the user per unit time (unit: per month)}.
[0123] 3) The state transition probability of a node from time t1 to time t2 is P12(U_t2|U_t1).
[0124] 4) The probability of a node state change, for example, P2(△s_t), where △s = {the change in the amount of financial product purchases per unit time (unit: units / month)}, or △s = {…, -20, -10, -5, 0, 5, 10, 20, …}.
[0125] A negative number represents a decrease in the purchase volume of financial products, a positive number represents an increase in the purchase volume of financial products, and 0 represents an unchanged purchase volume of financial products.
[0126] 5) The probability that node Pi is connected to Uj is P0(Uj|Pi), where the total number of edges between P and U is equal to the total number of valid training data, which can be used for statistics. For example, P3(Uj|Pi) =<Ui,…,Pj,…,…> The number of times the sample appears / the total number of samples in the training set.
[0127] 6) The asset level probability of a product node, for example: P4(property_levelk|Pi) = the number of times asset level k is used by node Pi / the total number of behaviors of node Pi.
[0128] 7) The consumption level probability of a product node, for example: P5(consumption_levelk|Pi) = the number of times consumption level k is used by node Pi / the total number of behaviors of node Pi.
[0129] 8) The credit rating probability of a product node, for example: P6(credit_ratinglk|Pi) = the number of times credit rating k is used by node Pi / the total number of actions of node Pi.
[0130] By quantifying the variable values and performing statistics on the defined probabilities, a discrete probability distribution can be obtained. Then, an appropriate parameterized probability model (such as a normal distribution model, a multinomial distribution model, a Gaussian mixture model, etc.) is used to fit the statistically obtained probability distribution, and a parameterized description of each node can be obtained, which is recorded as ...and so on, where μ0, μ1, μ 12 etc. represent probability parameters. Once the parameterized descriptions of all nodes are obtained, the training of the probabilistic grammar model PKG is completed.
[0131] During the model implementation phase, users' financial product purchasing behavior is analyzed based on the trained PKG, enabling product clustering and customer group analysis. The specific process includes feature construction, node similarity calculation, and target customer group recommendation.
[0132] Given data from a product node of interest over a period of time (referred to as Data1), the main purpose of this step is to analyze the consumer customer base of the product node of interest during this period and cluster the products. Assuming that PKG1, PKG2, ..., PKGn are n PKGs trained using product data from different time periods, in order to analyze the potential customer base of the new product, the following three steps can be used:
[0133] 1) Feature construction: Given a PKG, the node probability parameters are arranged together to obtain a structured feature vector TZ = {μ0,μ1,…,μ k Then, according to the similarity between node feature vectors, group and cluster user nodes and product nodes respectively to obtain Nd product node cluster centers. and Ns user node cluster centers Thus, Ns+Nd eigenvectors about PKG are obtained.
[0134] By using the above feature construction method for PKG1, PKG2, ..., PKGn, n groups of PKG feature vectors can be obtained, each group including Ns+Nd cluster centers.
[0135] 2) Node similarity calculation: Using Data1 as training data and the model training method of the present invention, a new probabilistic knowledge graph PKG_new can be obtained. For each user node in PKG_new, its similarity (or distance) with the Ns cluster centers of PKG1 is calculated. If the similarity (or distance) with a certain center is i * If the similarity is greater than the threshold (or the distance is less than the threshold), the node is considered to belong to S in PKG1. i * Similarly, if for each product node in PKG_new, calculate its similarity with the Nd cluster centers of PKG1, if it is similar to a certain center If the similarity is greater than the threshold (or the distance is less than the threshold), the node is considered to belong to PKG1. If a node in PKG_new does not belong to any category in PKG1, the node is considered an outlier.
[0136] Based on the above calculation, the similarity between PKG1 and PKG_new can be expressed as:
[0137]
[0138] Where Sim(A,B) represents the similarity between node A and cluster center B. Possible calculation methods include Gaussian similarity metric Cosine similarity measure Wait, N in1 N is the number of source nodes in PKG_new that can be regarded as a certain type in PKG1. in2 N is the number of destination nodes in PKG_new that can be regarded as a certain type in PKG1. all is the number of all nodes in PKG_new, α, β, γ are adjustable weights.
[0139] Similarly, the similarity between PKG_new and PKG1, PKG2, ..., PKGn can be calculated.
[0140] 3) Cluster center modeling: This step builds a probabilistic knowledge graph based on the cluster center, and clusters to obtain Nd product node cluster centers. and Ns user node cluster centers Count the probabilities of user node centers under different product node centers, that is, calculate:
[0141] where i∈[1,N s ],j∈[1,N d ].
[0142] 4) Target customer group recommendation: For the data of the specified product P*, the probabilistic knowledge graph PKG_new can be calculated according to the above method. According to the similarity between the node feature vectors, the user nodes and product nodes are grouped and clustered respectively to obtain Nd product node cluster centers and Ns user node cluster centers. First, find the cluster center to which P* belongs. Computing Product Node Center Under , the probability of different user node centers: i∈[1,N s ];
[0143] At the same time, calculate the node center of the product Under each user node center The average amount of products in the. You can find the user center with the highest probability or the user center with the highest expected return The user center includes the user collection This user set is used as the recommended users for the product P*.
[0144] In a specific embodiment:
[0145] 1. Model training stage:
[0146] 1) Data Acquisition and Preprocessing: We selected one year of financial product sales data as our dataset. We divided the data into four quarters: January to March, April to June, July to September, and October to December. This data set consists of four datasets, Data1, Data2, Data3, and Data4.
[0147] The sample data of each data set is preprocessed, mainly including:
[0148] Outlier handling: Identifying and handling outliers in the data. This typically involves statistical analysis and visualization of the data to determine which values are unreasonable and whether to remove them or replace them using some method (such as median filling).
[0149] Duplicate value processing: remove or merge duplicate values in data.
[0150] Encoding: For some feature data, such as asset level and consumption level, one-hot encoding is used to convert categorical variables into an understandable format.
[0151] 2) Feature extraction: PKG contains user nodes Ui and product nodes Pi.
[0152] 3) Establishment of a multi-dimensional feature library: After obtaining the discrete probability distributions of the two features of user node Ui and product node Pi respectively, we can use an appropriate parameterized probability model (here, a Gaussian mixture model) to fit the statistically obtained probability distributions, and then we can obtain a parametric description of each node. For the probability P(Pj|Ui) of user node Ui, a 4-component one-dimensional Gaussian mixture model is used for fitting. Therefore, the parameterized model contains 4 means, 4 standard deviations, and 4 component weights, for a total of 12 parameters. For user node Ui, P(rating|Ui) is fitted using a 4-component one-dimensional Gaussian mixture model. Similarly, the parameterized model contains 4 means, 4 covariances, and 4 component weights, for a total of 12 parameters. After flattening the two features of Ui, a 24-dimensional feature vector of Si can be obtained. Similarly, a 24-dimensional feature vector can be obtained for Pi.
[0153] 2. Model usage stage:
[0154] 1) Feature Clustering: Data2 contains 6950 user nodes Ui and 114 product nodes Pi. Clustering the 6950 user nodes Ui according to the 24-dimensional vectors used in model construction yields 15 24-dimensional user class centers. Clustering the 114 product nodes Pi according to the 24-dimensional vectors used in model construction yields 8 36-dimensional product class centers. Therefore, by arranging the 15 class centers of all user nodes Ui and the 8 class centers of all product nodes together, we obtain a structured feature vector for PKG2 containing 23 elements, where each element is a 24-dimensional vector.
[0155] For the other datasets, Data1, Data3, and Data4, the discrete probability distributions of node features were fitted to parameterized models with a dimension of 24. However, during clustering, the number of cluster centers varied: 28, 20, and 35, respectively. Therefore, the following probabilistic grammar models (PKGs) were obtained.
[0156]
[0157] 2) PKG similarity calculation: PKG similarity calculation can be implemented using the formula: In the formula, α, β, and γ are 0.4, 0.4, and 0.2 respectively. The similarity is measured using cosine similarity. The similarity between the above four data sets was calculated separately, and it was found that the similarity between Data2 and Data4 was higher, which was 0.79.
[0158] 3) Cluster center modeling: This step uses Data2 as sample data and the probabilistic grammar model PKG established above as the basis to establish the cluster center PKG model. The cluster centers of Data2 are the 8 product node cluster centers and the 15 user node cluster centers. Statistical clustering is used to obtain the probability of the 15 user node centers under the 8 product node centers, that is, to calculate: Where i∈[1,15],j∈[1,8].
[0159] 4) Product target customer group recommendation: In this embodiment, product 2 is selected, and according to the probability knowledge graph PKG_new calculated in step 3), and the probability grammar model of the cluster center, first find the cluster center to which product 2 belongs. Computing Product Node Center Next, the probability of 15 user node centers is: i∈[1,15].
[0160] The probability of 15 user node centers, each user node center The average amount of all products purchased in the app is as follows:
[0161]
[0162] Select the user center with the highest expected average amount Or the user center with the highest probability The user collections included Recommend product P* to these users.
[0163] In the above embodiment, compared with the related technologies, it has the following significant advantages: (1) Low deployment cost: The system does not require high hardware configuration, is easy to deploy, and has low operation and maintenance costs, making it suitable for resource-constrained environments. (2) Strong interpretability: The probabilistic grammar model provides an intuitive way to understand customers' financial product purchasing behavior. (3) Strong generalization ability: The model has a high tolerance for data diversity and low data requirements, and can achieve good performance without a large amount of labeled data and customer personal information. (4) Low migration cost: The system can automatically adapt to different financial business scenarios, reducing migration costs.
[0164] On the other hand, an embodiment of the present disclosure provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned financial product recommendation method based on probabilistic knowledge graph. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present disclosure. The computer device 100 includes a memory 111 , a processor 112 , and a communication unit 113 .
[0165] To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other, directly or indirectly. For example, these components can be electrically connected via one or more communication buses or signal lines. Computer device 100 includes at least one software function module that can be stored in the form of software or firmware in memory 111 or embedded in the operating system (OS) of computer device 100. Processor 112 is used to execute the financial product recommendation method based on the probabilistic knowledge graph stored in memory 111.
[0166] An embodiment of the present disclosure provides a readable storage medium, which includes a computer program. When the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the aforementioned financial product recommendation method based on probabilistic knowledge graph.
[0167] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.
Claims
1. A financial product recommendation method based on probabilistic knowledge graph, characterized in that: The method comprises: Obtaining target product information, where the target product information is used to describe a target financial product; Generate a corresponding target product node according to the target product information; Determining a target product node center to which the target product node belongs based on a predetermined probabilistic knowledge graph; Determine a probability distribution of the target product node center with respect to each user node center; wherein the probability distribution indicates a probability value of the user corresponding to each user node center purchasing the target financial product, and the probability value is calculated based on parameterized description information of each user node center in the probabilistic knowledge graph and parameterized description information corresponding to the target product node center; Determine the target user node center corresponding to the highest probability value in the probability distribution; Sending recommendation information for the target financial product to a target user; the target user is a user corresponding to a user node under the target user node center, and the user nodes under the target user node center include at least one user node in a cluster corresponding to the target user node center; The product node center is obtained by clustering multiple product nodes in the probabilistic knowledge graph. Each product node center is the product node at the most central position in a cluster. The target product node center is determined based on the similarity between the target product node and multiple product center nodes. The user node center is obtained by clustering multiple user nodes in the probabilistic knowledge graph, and each user node center is a user node at the most central position in a cluster; Determining the probabilistic knowledge graph includes: Obtaining purchase history data of a user for a financial product, the purchase history data including user information describing the user and product information describing the financial product, wherein the number of the user and the number of the financial product are both greater than or equal to 1; Constructing an initial probabilistic knowledge graph based on the user information and the product information, the initial probabilistic knowledge graph including a first knowledge graph and a second knowledge graph, the first knowledge graph including at least one user node, the second knowledge graph including at least one product node, and a fully connected relationship between the user node and the product node; Extracting probability features for each of the user nodes and each of the product nodes from the purchase history data; Fitting the probability features to obtain parameterized description information for each of the user nodes and each of the product nodes; Establishing a multidimensional feature library according to the parameterized description information; The probabilistic knowledge graph is determined based on the multidimensional feature library.
2. The method according to claim 1, characterized in that The method comprises: Determine a first user node center among the user node centers that has the highest average purchase amount of the financial product; The recommendation information is sent to a target user corresponding to at least one user node under the first user node center.
3. The method according to claim 1, characterized in that Determining the probabilistic knowledge graph based on the multidimensional feature library includes: Clustering the user nodes and the product nodes according to the parameterized description information of each user node and each product node in the multidimensional feature library to obtain at least one product node center and at least one user node center; The probabilistic knowledge graph is generated based on the at least one product node center and the at least one user node center.
4. The method according to claim 3, characterized in that The clustering of the user nodes and the product nodes includes at least one of the following: Performing density-based clustering on the user nodes and the product nodes; The user nodes and the product nodes are clustered based on hierarchy.
5. The method according to claim 1, wherein The obtaining of the user's purchase history data of financial products includes: Acquire a plurality of first historical data of the user for the financial product in a plurality of different time periods; The determination of the probabilistic knowledge graph further includes: Determining a plurality of probabilistic knowledge graphs based on the plurality of first historical data; The determining, based on a predetermined probabilistic knowledge graph, a target product node center to which the target product node belongs, includes: Based on any one or more probabilistic knowledge graphs among the multiple probabilistic knowledge graphs, determine one or more target product node centers to which the target product node belongs.
6. The method according to claim 5, characterized in that The multiple different time periods include a most recent time period corresponding to the current time, and the method includes: Performing similarity analysis on the multiple probabilistic knowledge graphs to determine the similarity between any two of the probabilistic knowledge graphs; The determining, based on a predetermined probabilistic knowledge graph, a target product node center to which the target product node belongs, includes: Based on the first probabilistic knowledge graph corresponding to the most recent time period and the second probabilistic knowledge graph that is most similar to the first probabilistic knowledge graph, determine the first product node center and the second product node center to which the target product node belongs. The first product node center is the target product node center corresponding to the first probabilistic knowledge graph, and the second product node center is the target product node center corresponding to the second probabilistic knowledge graph.
7. The method according to claim 1, characterized in that The method comprises: The purchase history data is preprocessed, and the preprocessing includes at least one of the following: abnormal value processing, repeated value processing, and encoding processing.
8. The method according to claim 1, characterized in that The method comprises: Obtain target user information, where the first user information is used to describe the user to be recommended; Generate a target user node according to the target user information; Determining, based on the probabilistic knowledge graph, a second user node center to which the target user node belongs; Determine a first probability distribution of the second user node center with respect to each product node center; Determine the third product node center corresponding to the highest probability value in the first probability distribution; Send recommendation information of the financial product corresponding to at least one product node under the third product node center to the user to be recommended.
9. A computer device, characterized in that: The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the financial product recommendation method based on probabilistic knowledge graph described in any one of claims 1 to 7.
Citation Information
Patent Citations
Network resource configuration automatic adjustment method based on probability grammar model
CN118353777A
Financial product recommendation method and device
CN118469679A