A method for recommending financial products
Through the combination of perfect data processing and multiple recommendation algorithms, financial product information is dynamically updated, and user feedback is used to optimize the recommendation algorithm, the problems of imperfect data processing and inaccurate recommendation algorithms in the existing financial product recommendation methods are solved, and personalized and accurate financial product recommendations are achieved.
Patent Information
- Application Number
- CN202510241728.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing financial product recommendation methods have problems such as incomplete data processing, inaccurate recommendation algorithms, untimely update of product information, and inability to effectively utilize user feedback optimization algorithms.
Through a complete data processing process, including abnormal data detection and missing data filling, combining multiple recommendation algorithms such as neural networks, content-based filtering and collaborative filtering, financial product information is dynamically updated, and user feedback is used to optimize the recommendation algorithm.
It realizes a comprehensive and accurate analysis of user characteristics, generates a personalized financial product recommendation list, improves the accuracy and effectiveness of recommendations, and meets the diverse financial product needs of users.
Smart Images

Figure CN119722254B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of financial technology, and in particular to a method for recommending financial products. Background Art
[0002] In today's financial market, there are many types of financial products, and users face many problems when choosing financial products that suit them. Traditional financial product recommendation methods have great limitations. On the one hand, the data processing method is not perfect. When collecting user information, there is a lack of effective processing methods for abnormal data and missing data, resulting in inaccurate user portraits built based on these data. For example, some abnormal historical financial behavior data may interfere with the judgment of the user's true investment preferences, and missing data will cause information gaps in the portrait and fail to fully reflect the user's characteristics.
[0003] On the other hand, the recommendation algorithm is not accurate and intelligent enough. The simple content filtering algorithm only makes recommendations based on the feature matching between the product and the user portrait, ignoring the similarity between users and group behavior; although the collaborative filtering algorithm takes into account the similarity between users, it is insufficient in processing the complex features of user portraits and cannot fully combine key factors such as the user's risk tolerance to make personalized recommendations. Moreover, in the process of recommendation, traditional methods rarely dynamically update the key information of financial products, such as the yield fluctuation range and risk level, and cannot adapt to the dynamic changes of the market.
[0004] In addition, traditional recommendation methods lack effective use of user feedback and do not optimize the recommendation algorithm based on user feedback on the recommendation results, resulting in a low degree of fit between the recommended financial products and the actual needs of users.
[0005] Therefore, a method for recommending financial products has become an urgent problem to be solved. Summary of the invention
[0006] The purpose of the present invention is to provide a method for recommending financial products, which solves the technical problems in the prior art such as imperfect data processing, inaccurate recommendation algorithms, untimely product information updates, and inability to effectively utilize user feedback to optimize algorithms. Through a complete data processing flow, personalized recommendations that integrate multiple algorithms, and algorithm optimization based on user feedback, user characteristics are comprehensively and accurately analyzed, and a personalized recommendation list is generated for the user in combination with financial product information. The recommendation algorithm is continuously optimized through user feedback, thereby improving the accuracy and effectiveness of recommendations and meeting users' diverse financial product needs.
[0007] To achieve the above object, the present invention provides a technical solution: a method for recommending financial products, comprising the following steps:
[0008] S1. Collect the user's basic information, historical financial behavior data, and risk tolerance questionnaire filled in by the user and generate data samples. Use the Isolation Forest algorithm to calculate abnormal samples and remove them. Use the linear regression model to predict and fill in missing data to form a user profile.
[0009] S2. Analyze user portraits using neural network algorithms to extract user preference characteristics and risk tolerance quantification values;
[0010] S3. Establish a database containing detailed information of various financial products, including the product's rate of return, risk level, and investment period;
[0011] S4. Generate a personalized list of financial product recommendations for users based on key information in the user portrait and product information in the financial product database by combining content-based filtering algorithm with collaborative filtering algorithm;
[0012] S5. Collect user feedback on recommendation results, generate comprehensive feedback data, and continuously optimize the recommendation algorithm for financial products based on the comprehensive feedback data.
[0013] Further, in step S1, the basic information includes age, gender, occupation, and income level; the historical financial behavior data includes historical investment product types, investment amounts, transaction times, and transaction frequencies;
[0014] The collected data is processed using the data cleaning algorithm. Assume that the collected historical financial behavior data constitutes the data set D. For each sample x in the data set i ∈D, i=1,2,…,N, where N is the number of samples; the IsolationForest algorithm is used to calculate the anomaly score The average path length correction factor
[0015] Where n = N, H(i) = ln(i) + 0.5772156649, 0.5772156649 is the Euler constant, For sample x i Average path length among multiple isolated trees;
[0016] Set the threshold τ, if s(x i )<τ, then determine x i It is an abnormal sample and is removed;
[0017] For missing data, linear regression models are used to predict and fill in the missing data, ultimately forming complete and accurate user portrait data.
[0018] Further, feature extraction step: In step S2, the neural network model is used to conduct in-depth analysis of the user portrait; let the neural network input layer feature vector be x=x 1 ,x 2 ,…,x n0 , the lth layer has n l neurons where l=1,2,…,L, L is the number of network layers; input z (l) =W (l) a (l-1) +b (l) , where W (l) is the weight matrix of the lth layer, with dimension n l ×n l-1 , b (l) is the bias vector of the lth layer, with dimension n l ×1; the activation function σ adopts the ReLU function, that is, a (l) =σ(z (l) )=max(0,z (l) );
[0019] The output layer L outputs y = a (L) =σ(z (L) ), the output layer has n L neurons, corresponding to the extracted preference features P f and the risk tolerance quantification value R t ; Update the weight matrix W through the back propagation algorithm and gradient descent method (l) and the bias vector b (l) , to minimize the mean square error loss function Where M is the number of training samples, y m is the true value, is the predicted value. Furthermore, in the training process of the neural network model, a regularization method is used to regularize the weight matrix W (l) Constraints are imposed to prevent the model from overfitting.
[0020] Further, in step S3, the yield of the product includes the expected annualized yield and the fluctuation range of the actual yield; the risk level is divided into low, medium and high risk levels, represented by the values 1, 2 and 3 respectively; the investment period includes short-term, medium-term and long-term, the short-term is within one year, the medium-term is one to three years, and the long-term is more than three years;
[0021] The ARIMA (p, d, q) model is used to analyze the historical yield time series of financial products. Combined with market macroeconomic indicators, the actual yield fluctuation range is updated through a multivariate linear regression model.
[0022] Furthermore, for the update of risk level, in addition to considering the product’s own factors, the risk level is re-evaluated through a logistic regression model in combination with the market risk index.
[0023] Further, in step S4, based on the content filtering algorithm, let the user portrait feature vector u = (u 1 ,u 2 ,…,u n ,) by the preference feature P f and the risk tolerance quantification value R t The financial product feature vector p = (p 1 ,p 2 ,…,p n ,) to calculate the cosine similarity Get the score vector S based on content filtering c =(s c1 ,s c2 ,…,s ci ,…,s cm ), where m is the number of financial products;
[0024] Based on the collaborative filtering algorithm, let the user-product matrix be R, where R vw represents the rating of user v on product w; using the risk tolerance quantification value R in the user portrait t and preference characteristics P f Cluster users and calculate the cosine similarity between user v and user k of the same category in and are the average ratings of user v and user k respectively;
[0025] Predict user v’s rating for product w Where N v is a set of users of the same category and similar to user v, and the score vector S based on collaborative filtering is obtained cf =(s cf1 ,s cf2 ,…,s cfi ,…,s cfm );
[0026] Assume weight α (0≤α≤1), and α is quantified according to the risk tolerance value R t Dynamic adjustment, fused score vector S f =(s f1 ,s f2 ,…,s fi ,…,s fm ), where S fi =αs ci +(1-α)s cfi;
[0027] According to S f Sort the financial products and select the top K products with relatively high scores to generate a recommendation list RL, where 5≤K≤20.
[0028] Furthermore, the DBSCAN clustering algorithm is used to cluster financial products. The DBSCAN algorithm calculates the distance between samples and divides the samples whose distance is within a certain threshold and the number of neighborhood samples is greater than the minimum number of samples into the same category, thereby ensuring that the products in the recommendation list come from different categories and increasing product diversity.
[0029] Furthermore, in step S5, the feedback includes the user's click behavior on the recommended product, whether to purchase the recommended product, and the evaluation of the recommended product, and a user feedback matrix F is constructed, where F vw represents the comprehensive feedback score of user v on the recommended product w;
[0030] Collect user browsing behavior data on recommended pages, including dwell time t s , browse order O s ; Obtain the comprehensive feedback value C through weighted calculation f =w 1 F vw +w 2 t s +w 3 O s , where w 1 ,w 2 ,w 3 is the weight, which is determined according to the actual business importance;
[0031] Use the Q-Learning algorithm to optimize the recommendation algorithm. Suppose the state space S consists of the user portrait and the current recommendation list RL, and the action space A is different recommendation algorithm parameters.
[0032] After taking action a∈A in state s∈S, it moves to the new state s′ and obtains reward r, which is calculated based on the comprehensive feedback value C f Calculation, that is, r = C f ;
[0033] The Q function Q(s,a) represents the expectation of the long-term cumulative reward of taking action a in state s. The update formula is:
[0034] Q(s,a)=Q(s,a)+α Q [r+γmax a′ Q(s′,a′)-Q(s,a)], where α Q is the learning rate, γ is the discount factor;
[0035] By continuously iteratively updating the Q function, the action that maximizes the Q value is selected, that is, the optimal recommendation algorithm parameters;
[0036] In the A / B test, users are randomly divided into two groups, and different versions of the recommendation algorithm are used for recommendation. The reward r is calculated based on the feedback data of the two groups of users, and then the Q function is updated to determine the optimal algorithm version.
[0037] The advantages of the present invention over the prior art are that the present invention constructs an accurate and complete user portrait through multi-dimensional data collection, data cleaning and preprocessing algorithms, laying a solid foundation for personalized recommendations.
[0038] The present invention utilizes deep learning neural network and regularization technology to deeply mine user features and improve the accuracy and reliability of feature extraction.
[0039] The present invention combines multiple algorithms to dynamically update the financial product database, ensuring the timeliness and accuracy of product information and reflecting dynamic changes in the market.
[0040] The present invention integrates multiple recommendation algorithms and considers multiple factors to generate a recommendation list that better meets the user's personalized needs and market dynamics.
[0041] The present invention continuously optimizes the recommendation algorithm through multiple optimization algorithms and user feedback, improves the accuracy and effectiveness of recommendations, and enhances user experience and market competitiveness. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flow chart of a method for recommending financial products of the present invention.
[0043] Figure 2 It is a flow chart of the steps of constructing a user portrait of a financial product recommendation method of the present invention.
[0044] Figure 3 It is a flow chart of the feature extraction steps of a method for recommending financial products of the present invention.
[0045] Figure 4 It is a flow chart of the database establishment steps of a financial product recommendation method of the present invention.
[0046] Figure 5 It is a flow chart of the steps of generating a recommendation list of a method for recommending financial products of the present invention.
[0047] Figure 6 It is a flow chart of algorithm optimization steps of a financial product recommendation method of the present invention. DETAILED DESCRIPTION
[0048] Various exemplary embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention.
[0049] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0050] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.
[0051] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0052] Combined with Figure 1-6 , a method for recommending a financial product of the present invention is further described in detail.
[0053] A method for recommending financial products, comprising the following steps:
[0054] Steps to construct a user portrait: Collect basic information about users, which should at least include age, gender, occupation, and income level; collect historical financial behavior data about users, which should include but not be limited to historical investment product types, investment amounts, transaction times, and transaction frequencies; and obtain risk tolerance questionnaires filled out by users, and combine the above information to form a user portrait.
[0055] The collected data is processed using the data cleaning algorithm. Assume that the collected historical financial behavior data constitutes the data set D. For each sample x in the data set i ∈D, i = 1, 2, ..., N (N is the number of samples), the IsolationForest algorithm is used to calculate the anomaly score The average path length correction factor (n=N,H(i)=ln(i)+0.5772156649,0.5772156649 is Euler's constant), For sample x i The average path length in multiple isolated trees. Set the threshold τ, if s(x i )<τ, then determine x i The samples are considered abnormal and removed.
[0056] For missing data, assume that the missing investment amount data A mmiss , using the linear regression model A m=β 0 +β 1 T+β 2 F+∈ performs prediction filling and estimates the regression coefficient β by the least squares method 0 ,β 1 ,β 2 , get the predicted value Fill in missing values and finally form complete and accurate user portrait data U.
[0057] Feature extraction step: Use deep learning models such as neural networks to conduct in-depth analysis of user profile U. Suppose the neural network input layer feature vector is x = x 1 ,x 2 ,…,x n0 , the lth layer (l=1,2,…,L, L is the number of network layers) has n l neurons, input z (l) =W (l) a (l-1) +b (l) , where W (l) is the weight matrix of the lth layer, with dimension n l ×n l-1 , b (l) is the bias vector of the lth layer, with dimension n l ×1, the activation function σ adopts the ReLU function, that is, a (l) =σ(z (l) )=max(0,z (l) ). The output layer L outputs y = a (L) =σ(z (L) ), assuming that the output layer has n L neurons, corresponding to the extracted preference features P f and the risk tolerance quantification value R t . Update the weight matrix W through the back propagation algorithm and gradient descent method (l) and the bias vector b (l) , to minimize the mean square error loss function Where M is the number of training samples, y m is the true value, is the predicted value.
[0058] Database establishment steps: Establish a database containing detailed information of various financial products, which records the yield of the products in detail, including the expected annualized yield R ea , actual yield fluctuation range ΔR; risk level R g According to industry standards or custom standards, it is divided into low, medium and high risk levels, represented by values 1, 2 and 3 respectively; investment period T i, including short-term (less than one year), medium-term (one to three years), and long-term (more than three years), represented by the values 1, 2, and 3 respectively.
[0059] Use the ARIMA (p, d, q) model to analyze the historical yield time series y of financial products. t , t=1,2,…,T (T is the length of the time series) for analysis, and the original time series is differentiated by d order to obtain a stationary series in is the first-order difference. The ARIMA model represents Where B is the lag operator, B k y t =y t-k , is the autoregressive part, is the moving average part, ∈ t is a white noise sequence with mean 0 and variance σ 2 . Estimate the model parameter φ by using methods such as maximum likelihood estimation i and θ j , predict the future rate of return y T+h , update the expected annualized rate of return R ea At the same time, combined with the market macroeconomic indicators E (such as GDP growth rate, inflation rate, etc.), through the multivariate linear regression model ΔR = α 0 +α 1 E+∈ updates the actual yield fluctuation range ΔR and estimates the regression coefficient α by the least squares method 0 ,α 1 .
[0060] Recommendation list generation steps: Based on key information in the user profile, such as the risk tolerance quantification value R t , preference feature P f , and the product information in the financial product library, and use the recommendation algorithm to generate a personalized financial product recommendation list for the user. The recommendation algorithm uses a combination of content filtering algorithm and collaborative filtering algorithm. Based on the content filtering algorithm, let the user portrait feature vector u = (u 1 ,u 2 ,…,u n ,) by the preference feature P f and the risk tolerance quantification value R t The financial product feature vector p = (p 1 ,p 2 ,…,p n ,) to calculate the cosine similarity Get the score vector S based on content filtering c =(s c1 ,s c2,…,s ci ,…,s cm ), where m is the number of financial products.
[0061] Based on the collaborative filtering algorithm, let the user-product matrix be R, where R vw Represents the user v's rating of product w (or purchase behavior, which can be expressed as 0 / 1). Use the risk tolerance quantification value R in the user portrait t and preference characteristics P f Cluster users and calculate the cosine similarity between user v and user k of the same category in and are the average ratings of user v and user k respectively. Predict the rating of user v for product w Where N v is a set of users of the same category and similar to user v, and the score vector S based on collaborative filtering is obtained cf =(s cf1 ,s cf2 ,…,s cfi ,…,s cfm ). Assume weight α (0≤α≤1), and α is quantified according to the risk tolerance value R t Dynamic adjustment, when R t When it is higher, increase α appropriately; when R t When it is low, reduce α appropriately. The fused score vector S f =(s f1 ,s f2 ,…,s fi ,…,s fm ), where S fi =αs ci +(1-α)s cfi According to S f Sort the financial products and select the top K products with higher scores to generate a recommendation list RL, where 5≤K≤20.
[0062] Algorithm optimization steps: Collect user feedback on recommendation results, including user click behavior on recommended products, purchase behavior, evaluation of recommended products, etc., and construct user feedback matrix F, where F vw represents the comprehensive feedback score of user v on the recommended product w. At the same time, the browsing behavior data of users on the recommended page is collected, such as the stay time t s , browse order O s Etc., and the comprehensive feedback value is obtained through weighted calculation:
[0063] C f =w 1 F vw +w 2 ts +w 3 O s ; where w 1 ,w 2 ,w 3 is the weight, which is determined according to the actual business importance. Using the Q-Learning algorithm to optimize the recommendation algorithm, the state space S is composed of user portraits (including the risk tolerance quantification value R t and preference characteristics P f ) and the current recommendation list RL. The action space A is composed of different recommendation algorithm parameters (such as different values of weight α based on content filtering and collaborative filtering). After taking action a∈A in state s∈S, it transfers to the new state s′ and obtains reward r. The reward r is calculated based on the comprehensive feedback value C f Calculation, for example, r = C f The Q function Q(s,a) represents the expected long-term cumulative reward for taking action a in state s, and the update formula is Q(s,a)=Q(s,a)+α Q [r+γmax a′ Q(s′,a′)-Q(s,a)], where α Q is the learning rate, and γ is the discount factor. By continuously iteratively updating the Q function, the action that maximizes the Q value (i.e., the optimal recommendation algorithm parameters) is selected. In the A / B test, users are randomly divided into two groups, and different versions of the recommendation algorithm (corresponding to different actions a) are used for recommendation. The reward r is calculated based on the feedback data of the two groups of users, and then the Q function is updated to determine the optimal algorithm version.
[0064] In the user portrait construction step, data cleaning and preprocessing also include one-hot encoding of categorical data (such as gender, occupation, product type, etc.) to meet the input requirements of subsequent machine learning algorithms.
[0065] In the feature extraction step, during the training of the neural network model, a regularization method (such as L1 or L2 regularization) is used to regularize the weight matrix W. (l) Constraints are imposed to prevent the model from overfitting. The L2 regularization term is Where λ is the regularization parameter.
[0066] In the database establishment step, in addition to considering the product's own factors, the risk level is updated by combining the market risk index M through a logistic regression model. Corresponding to low, medium and high risk levels respectively, X is a related variable including product characteristics and market risk index M) re-evaluate the risk level R g .
[0067] In the step of generating the recommendation list, the diversity algorithm is based on the clustering method and uses the DBSCAN clustering algorithm to cluster financial products. The DBSCAN algorithm calculates the distance d(x i ,x j )(such as Euclidean distance The samples whose distance is within a certain threshold ∈ and whose number of neighborhood samples is greater than the minimum number of samples MinPTs are divided into the same category, thereby ensuring that the products in the recommendation list come from different categories and increasing product diversity.
[0068] In the algorithm optimization step, during the A / B test, statistical hypothesis tests (such as t-tests) are used to determine whether the performance differences of different versions of recommendation algorithms are significant. Assume that the null hypothesis H 0 :There is no difference in performance between the two algorithms. The alternative hypothesis H 1 :The performance of the two algorithms is different, by calculating the t value (in are the means of the two groups of user feedback indicators, are the variances of the two groups of user feedback indicators, n 1 ,n 2 are the number of users in two groups respectively), and compared with the critical value. ( is the significance level), the null hypothesis is rejected and it is believed that the performance of the two algorithms is significantly different, thus more accurately determining the optimal algorithm version.
[0069] The specific implementation process of a financial product recommendation method of the present invention is as follows:
[0070] 1. User portrait construction steps
[0071] (1) Data collection:
[0072] The basic information of a user was collected: age 35, male, occupation as middle-level manager in a company, and annual salary of 500,000 yuan.
[0073] Historical financial behavior data: In the past year, he has invested in stock funds (investment amount of 100,000 yuan, trading time is X month of XXX year, and trading frequency is once a month), bonds (investment amount of 200,000 yuan, trading time is X month of XXXX year, and trading frequency is once a quarter).
[0074] Risk tolerance questionnaire: This user self-assessed that his / her risk tolerance is medium.
[0075] (2) Data cleaning and preprocessing:
[0076] Outlier detection: Assume that the number of samples of historical financial behavior data collected is N = 100 (including the data of this user and other users). For each sample x i, the Isolation Forest algorithm is used to calculate the anomaly score in n=N,
[0077] H(i)=ln(i)+0.5772156649, For sample x i The average path length in multiple isolated trees. Set the threshold τ = 0.2, if s(x i )<τ, then determine x i The samples are considered abnormal and removed.
[0078] Missing value filling: The user's investment amount data is complete and there are no missing values.
[0079] Categorized data processing: One-hot encoding is performed on gender (male is coded as [1,0], female is coded as [0,1]), occupation (middle-level management is coded as [0,0,1,0,…], there are 10 occupational categories in total), and product type (stock funds are coded as [1,0,0], bonds are coded as [0,1,0], etc.), ultimately forming complete and accurate user portrait data U.
[0080] 2. Feature extraction steps
[0081] (1) Neural network model construction:
[0082] Neural network input layer feature vector x = x 1 ,x 2 ,…,x n0 , n0=10 (including age, income level, gender, occupation and other information after one-hot encoding).
[0083] The number of network layers is L = 3, and the first layer has n 1 = 20 neurons, the second layer has n 2 = 15 neurons, the output layer L has n L = 2 neurons, corresponding to the extracted preference features P f and the risk tolerance quantification value R t .
[0084] Enter z (l) =W (l) a (l-1) +b (l) , the activation function uses the ReLU function, that is, a (l) =σ(z (l) )=max(0,z (l) ).
[0085] Output layer y = a (L) =σ(z (L) ).
[0086] (2) Model training:
[0087] The weight matrix W is regularized using L2 regularization method. (l) Constrained, the L2 regularization term is Assume λ=0.01.
[0088] Update the weight matrix W through the back-propagation algorithm and gradient descent method (l) and the bias vector b (l) , to minimize the mean square error loss function The number of training samples is M = 1000. Finally, the user's preference characteristics and risk tolerance quantification value are obtained, R t =0.6 (quantified value, the higher the value, the stronger the risk tolerance).
[0089] 3. Database establishment steps
[0090] (1) Financial product information entry:
[0091] There is a product called “Steady Growth Fund” in the database, with an expected annualized rate of return R ea =6%, actual rate of return fluctuation range ΔR = ±2%, risk level R g =2 (medium risk), investment period T i =2 (mid-term, 1-3 years).
[0092] (2) Update of yield and risk level:
[0093] Yield prediction: Use the ARIMA (p, d, q) model to predict the fund's historical yield time series y t (T = 100 time points) for analysis, and the original time series is differentiated by d = 1 to obtain a stationary series ARIMA model representation Estimate the model parameters φ by using methods such as maximum likelihood estimation i and θ j , predict the future rate of return y T+h , update the expected annualized rate of return R after prediction ea It is 6.5%.
[0094] Update of actual yield fluctuation range: Combined with market macroeconomic indicators E (such as GDP growth rate, inflation rate, etc.), through the multivariate linear regression model ΔR=α 0 +α 1 E+∈ updates the actual yield fluctuation range ΔR and estimates the regression coefficient α by the least squares method 0 =0.01,α 1 =0.5, after update ΔR=±2.5%.
[0095] Risk level update: Combined with the market risk index M, through the logistic regression model (k=1,2,3 correspond to low, medium and high risk levels respectively, X is a related variable including product characteristics and market risk index M) Re-evaluate the risk level R g , the risk level after evaluation is still 2.
[0096] 4. Recommendation list generation steps
[0097] (1) Based on content filtering algorithm:
[0098] User portrait feature vector u=(u 1 ,u 2 ,…,u n ,) by the preference feature P f and the risk tolerance quantification value R t Construct, u = [0.3, 0.6] (here simplified to a two-dimensional vector).
[0099] The characteristic vector of the financial product “Steady Growth Fund” p=(p 1 ,p 2 ,…,p n ,), p = [0.2, 0.5] (corresponding to characteristics such as risk preference and expected return).
[0100] Calculate cosine similarity Get the score S based on content filtering c ≈0.9967.
[0101] (2) Based on collaborative filtering algorithm:
[0102] The user-product matrix R, using the risk tolerance quantification value R in the user profile t and preference characteristics P f Cluster users. The cosine similarity between user v and user k of the same category Predict user v’s rating for “Steady Growth Fund” Get the score S based on collaborative filtering cf =0.7.
[0103] (3) Fusion and sorting:
[0104] Because R t =0.6 is relatively high, so increase α appropriately and set the weight α=0.6.
[0105] The fusion score S f =αs c +(1-α)s cf =0.6×0.9967+(1-0.6)×0.7=0.87802.
[0106] The above calculation is performed on all financial products and they are sorted, and the top K=10 products with higher scores are selected to generate a recommendation list RL. The “Steady Growth Fund” is included in the recommendation list due to its higher score.
[0107] (4) Increased diversity:
[0108] The DBSCAN clustering algorithm is used to cluster financial products based on the Euclidean distance. Calculate the distance between samples, set the threshold ∈=0.5, the minimum number of samples MinPTs=5, and classify the samples whose distance is within the threshold and whose number of neighborhood samples is greater than the minimum number of samples into the same category, ensuring that the products in the recommendation list come from different categories and increasing product diversity.
[0109] 5. Algorithm optimization steps
[0110] (1) User feedback collection:
[0111] The user's click behavior score for the recommended product "Steady Growth Fund" is 3 (out of 5 points), the non-purchase behavior score is 0, and the evaluation score is 4. The user feedback matrix F is constructed. vw =3+0+4=7 (comprehensive score).
[0112] At the same time, the browsing behavior data of users on the recommended pages and the stay time t s = 60 seconds, browsing order O s =3 (the third product viewed in order).
[0113] The comprehensive feedback value C is obtained through weighted calculation f =w 1 F vw +w 2 t s +w 3 O s , where w 1 =0.5, w 2 =0.3, w 3 =0.2, then C f =0.5×7+0.3×60+0.2×3=22.1.
[0114] (2) Q-Learning algorithm optimization:
[0115] The state space S consists of user profiles (including the risk tolerance quantification value R t and preference characteristics P f ) and the current recommendation list RL, and the action space A is composed of different recommendation algorithm parameters (such as different values of the weight α based on content filtering and collaborative filtering).
[0116] After taking action a∈A in state s∈S, transfer to new state s′ and obtain reward r, where r=C f =22.1.
[0117] The Q function Q(s,a) represents the expected long-term cumulative reward of taking action a in state s. The update formula is Q(s,a)=Q(s,a)+α Q [r+γmax a′ Q(s′,a′)-Q(s,a)], where the learning rate α Q =0.1, discount factor γ=0.9.
[0118] (3) A / B testing:
[0119] The users are randomly divided into two groups, and different versions of the recommendation algorithm (corresponding to different actions a, assuming that α=0.6 for one group and α=0.5 for the other group) are used for recommendation.
[0120] The number of users in the first group n 1 =100, mean value of feedback index variance The number of users in the second group n 2 =120, mean value of feedback index variance
[0121] Calculate the t value,
[0122] Defining the significance level Critical value Because |t|=7.84>1.96, the null hypothesis is rejected, and it is believed that there is a significant difference in the performance of the two algorithms, thereby more accurately determining the optimal algorithm version.
[0123] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. A method for recommending financial products, characterized in that: The following steps are included: S1. Collect the user's basic information, historical financial behavior data, and risk tolerance questionnaire filled in by the user and generate data samples. Use the Isolation Forest algorithm to calculate abnormal samples and remove them. Use the linear regression model to predict and fill in missing data to form a user profile. S2. Analyze user portraits using neural network algorithms to extract user preference characteristics and risk tolerance quantification values; S3. Establish a database containing detailed information of various financial products, wherein the database contains the product's rate of return, risk level, and investment period; S4. Generate a personalized list of financial product recommendations for users based on key information in the user portrait and product information in the financial product database by combining content-based filtering algorithm with collaborative filtering algorithm; Based on the content filtering algorithm: Let the user portrait feature vector u = (u1, u2, ..., u n ,) by the preference feature P f and the risk tolerance quantification value R t The financial product feature vector p = (p1, p2, ..., p n ,) to calculate the cosine similarity Get the score vector S based on content filtering c =(s c1 ,s c2 ,...,s ci ,…,s cm ), where m is the number of financial products; Based on collaborative filtering algorithm: Let the user-product matrix be R, where R vw represents the rating of user v on product w; using the risk tolerance quantification value R in the user portrait t and preference characteristics P f Cluster users and calculate the cosine similarity between user v and user k of the same category in and are the average ratings of user v and user k respectively; Predict user v’s rating for product w Where N v is a set of users of the same category and similar to user v, and the score vector S based on collaborative filtering is obtained cf =(s cf1 ,s cf2 ,…,s cfi ,…,s cfm ); Assume weight α (0≤α≤1), and α is quantified according to the risk tolerance value R t Dynamic adjustment, fused score vector S f =(s f1 ,s f2 ,…,s fi ,...,s fm ), where S fi =αs ci +(1-α)s cfi ; According to S f Sort the financial products and select the top K products with high scores to generate a recommendation list RL, where 5≤K≤20; S5. Collect user feedback on recommendation results, generate comprehensive feedback data, and continuously optimize the recommendation algorithm for financial products based on the comprehensive feedback data; The feedback includes the user's click behavior on the recommended product, whether to purchase the recommended product, and the evaluation of the recommended product. The user feedback matrix F is constructed, where F vw represents the comprehensive feedback score of user v on the recommended product w; Collect user browsing behavior data on recommended pages, including dwell time t s , browse order O s ; Get the comprehensive feedback value C through weighted calculation f =w1F vw +w2t s +w3O s , where w1, w2, and w3 are weights, determined according to the actual business importance; Use the Q-Learning algorithm to optimize the recommendation algorithm. Suppose the state space S consists of the user portrait and the current recommendation list RL, and the action space A is different recommendation algorithm parameters. After taking action a∈A in state s∈S, transition to new state s ′ And get reward r, which is based on the comprehensive feedback value C f Calculation, that is, r = C f ; The Q function Q(s,a) represents the expectation of the long-term cumulative reward of taking action a in state s. The update formula is: Q(s,a)=Q(s,a)+α Q [r+γmax a′ Q(s ′ ,a ′ )-Q(s,a)], where α Q is the learning rate, γ is the discount factor; By continuously iteratively updating the Q function, the action that maximizes the Q value is selected, that is, the optimal recommendation algorithm parameters; In the A / B test, users are randomly divided into two groups, and different versions of the recommendation algorithm are used for recommendation. The reward r is calculated based on the feedback data of the two groups of users, and then the Q function is updated to determine the optimal algorithm version.
2. A method for recommending financial products according to claim 1, characterized in that: In step S1, the basic information includes age, gender, occupation, and income level; the historical financial behavior data includes historical investment product types, investment amounts, transaction times, and transaction frequencies; The collected data is processed using the data cleaning algorithm. Assume that the collected historical financial behavior data constitutes the data set D. For each sample x in the data set i ∈D, i=1,2,...,N, where N is the number of samples; the IsolationForest algorithm is used to calculate the anomaly score The average path length correction factor Where n = N, H(i) = ln(i) + 0.5772156649, 0.5772156649 is the Euler constant, For sample x i Average path length among multiple isolated trees; Set the threshold τ, if s(x i )<τ, then determine x i It is an abnormal sample and is removed; For missing data, linear regression models are used to predict and fill in the missing data, ultimately forming complete and accurate user portrait data.
3. A method for recommending financial products according to claim 2, characterized in that: Feature extraction Steps: In step S2, the neural network model is used to conduct in-depth analysis of the user portrait; let the neural network input layer feature vector be x = x1, x2, …, x n0 , the lth layer has n l neurons where l=1,2,…,L, L is the number of network layers; input z (l) =W (l) a (l-1) +b (l) , where W (l) is the weight matrix of the lth layer, with dimension n l ×n l-1 , b (l) is the bias vector of the lth layer, with dimension n l ×1; the activation function σ adopts the ReLU function, that is, a (l) =σ(z (l) )=max(0,z (l) ); The output layer L outputs y = a (L) =σ(z (L) ), the output layer has n L neurons, corresponding to the extracted preference features P f and the risk tolerance quantification value R t ; Update the weight matrix W through the back propagation algorithm and gradient descent method (l) and the bias vector b (l) , to minimize the mean square error loss function Where M is the number of training samples, y m is the true value, is the predicted value.
4. A method for recommending financial products according to claim 3, characterized in that: In the training process of the neural network model, the weight matrix W is regularized by using a regularization method. (l) Constraints are imposed to prevent the model from overfitting.
5. A method for recommending financial products according to claim 4, characterized in that: In step S3, the yield of the product includes the expected annualized yield and the fluctuation range of the actual yield; the risk level is divided into low, medium and high risk levels, represented by the values 1, 2 and 3 respectively; the investment period includes short-term, medium-term and long-term, the short-term is within one year, the medium-term is one to three years, and the long-term is more than three years; The ARIMA (p, d, q) model is used to analyze the historical yield time series of financial products. Combined with market macroeconomic indicators, the actual yield fluctuation range is updated through a multivariate linear regression model.
6. A method for recommending financial products according to claim 5, characterized in that: When updating the risk level, in addition to considering the product’s own factors, the risk level is re-evaluated through a logistic regression model in combination with the market risk index.
7. A method for recommending financial products according to claim 6, characterized in that: The DBSCAN clustering algorithm is used to cluster financial products. By calculating the distance between samples, samples whose distance is within the threshold and whose number of neighborhood samples is greater than the minimum number of samples are classified into the same category, ensuring that the products in the recommendation list are from different categories.
Citation Information
Patent Citations
Financial product recommendation method and device based on data medium table, and electronic equipment
CN118195786A
Personalized recommendation method and system for enterprise financial products based on preference analysis
CN118657592A