A Method for Recommending Potential Users of Financial Products in Multi-Party Semi-Supervised Learning

Through the combination of multi-party semi-supervised learning and vertical federal learning models, the problem of financial products recommendation for financial institutions in the case of data silos is solved, accurate and batch recommendations for potential users are achieved, and the reliability and accuracy of recommendations are improved.

CN116523602BActive Publication Date: 2025-07-01CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310508313.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-08
Publication Date
2025-07-01
Estimated Expiration
2043-05-08

AI Technical Summary

Technical Problem

When recommending new financial products, financial institutions face the challenge of how to quickly and accurately identify potential users, especially in the case of data silos, it is difficult for multi-party data to aggregate and build machine learning models.

Method used

The multi-party semi-supervised learning method is adopted to establish a data set containing user information of financial products and other multi-party user information, pre-process and sample alignment, build positive and unlabeled sample data sets, and use vertical federated learning models for training and prediction, and gradually enhance the positive sample data set to achieve accurate recommendations of potential users.

Benefits of technology

While protecting the security and privacy of multi-party data, the batch recommendation problem with only a small number of positive samples and a large number of unlabeled samples is effectively solved, which improves the reliability and accuracy of recommendations, and realizes the accuracy and batch recommendation of potential users of financial products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116523602B_ABST
    Figure CN116523602B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for recommending potential users of financial products in multi-party semi-supervised learning, belonging to the field of big data recommendation. Aiming at the problem that the provider of financial products only has its own positive sample data and cannot make recommendations, under the condition of protecting the security and privacy of multi-party data, multiple random samplings are performed by combining multi-party unlabeled data to construct a binary classification dataset with balanced positive and negative samples, and a vertical federated learning model based on a base learner is trained. According to its prediction results, reliable positive samples are selected from the unlabeled sample data, and the processes of dataset reconstruction sampling and model training prediction are iterated multiple times to select a batch of reliable positive samples. This method effectively solves the problem of batch recommendation with only a small number of positive samples and a large number of unlabeled samples, improves the reliability of recommendation, and realizes the accurate and batch recommendation of potential users of financial products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of big data recommendation, and relates to a method for recommending potential users of financial products through multi-party semi-supervised learning. Background Art

[0002] With the continuous development of the social economy, more and more financial institutions have been launching different financial products for different user groups with different characteristics. Financial products refer to various financial services and products provided by financial institutions such as banks, securities, insurance, and provident funds, such as loans, credit cards, funds, stocks, provident fund contributions, etc. However, due to the different demands of different users for different financial products, when a specific new financial product is released, there are huge challenges in quickly and accurately recommending it to potential users. By using machine learning and big data technologies, the characteristics of users who have purchased a certain financial product of a certain financial institution can be analyzed, and potential users with the same characteristics can be found from those who have not purchased the financial product, and accurate recommendations can be made to the financial institution to meet the needs of users and financial institutions. However, when a financial institution only has information on a small number of people who have purchased the financial product and does not have information on those who have not purchased the financial product, the financial institution needs to obtain more data from other institutions or channels to assist in the recommendation. However, due to the requirements of data privacy and security protection for each party, the data is isolated from each other, forming data islands, making it very difficult for different participants to aggregate the data together to build a higher-performance machine learning model. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method for recommending potential users of financial products through joint multi-party semi-supervised learning.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] A method for recommending potential users of financial products through multi-party semi-supervised learning includes the following steps:

[0006] S1: Establish a multi-party dataset for recommending potential users of financial products that includes information on users who have purchased financial products and information on other multi-party users, and perform preprocessing and sample alignment on the multi-party dataset for recommending potential users of financial products to construct a positive sample dataset and an unlabeled sample dataset;

[0007] S2: Perform random sampling with replacement in the unlabeled sample dataset to establish a negative sample dataset; use the negative sample dataset and the positive sample dataset to construct a training set, and use the samples in the unlabeled sample dataset that have not been sampled to construct a prediction set; construct a vertical federated learning model based on a base learner, train it on the training set, and make predictions on the prediction set to obtain the prediction scores of each sample in the prediction set;

[0008] S3: Repeat the sampling, training, and prediction processes in step S2 multiple times; calculate the probability that a sample is predicted as a positive sample based on the sum of the prediction scores of each sample in the unlabeled sample dataset and the number of times it appears in the prediction set; sort all the samples in the unlabeled sample dataset in descending order according to the probability of positive samples, select the samples with the top rankings as reliable positive samples based on prior knowledge, add them to the positive sample dataset, and delete them from the unlabeled sample dataset at the same time;

[0009] S4: Repeat steps S2 - S3 until the preset maximum number of iterations is reached; thus, all the reliable positive samples selected from the unlabeled sample dataset are used as potential users and precisely and batch - recommended to the financial product providers.

[0010] Furthermore, the multi - party dataset for financial product potential user recommendation in step S1 includes Party A with information on users who have purchased financial products, and other multi - party data sources Party B and Party C other than Party A; step S1 specifically includes the following steps:

[0011] S11: Perform data pre - processing on the multi - party dataset for financial product potential user recommendation, including redundant data processing, missing value processing, outlier processing, data standardization processing, and label data processing, to obtain the Party A dataset represents the i - th sample feature vector of Party A, and the vector dimension is d A , represents the corresponding label, where i = 1,..., a; the Party B dataset represents the i - th sample feature vector of Party B, and the vector dimension is d B , represents the corresponding label, where i = 1,..., b; the Party C dataset represents the i - th sample feature vector of Party C, and the vector dimension is dc, represents the corresponding label, where i = 1,..., c;

[0012] S12: Encrypt and align the samples of Party B and Party C datasets D B and D C according to their sample IDs, retain the aligned sample data of Party B and Party C, discard the unaligned sample data, and obtain n samples; the aligned Party B dataset is The Party C dataset is represents the corresponding label, denote the corresponding tags, where i = 1,..., n;

[0013] S13: For the dataset D A and and encrypt and align the samples according to their sample IDs, and use the three - party aligned samples as positive samples to form the positive sample dataset P = {(xp i , yp i )}, where respectively represent the i - th sample feature vectors after alignment of Party A, Party B, and Party C, yp i ∈{1} is the positive sample label, |P| represents the number of positive samples, i = 1,..., |P|; the samples that are not aligned among the three parties are used as unlabeled samples to form the unlabeled sample dataset U = {(xu i , yu i )}, where respectively represent the i - th sample feature vectors after alignment of Party B and Party C, yu i ∈{0} is the unlabeled sample label, |U| represents the number of unlabeled samples, i = 1,..., |U|.

[0014] Furthermore, in step S11, the redundant data processing specifically includes: judging whether the data is repeated through some fields in the data record, for the repeated data, only keep one of them, delete the other repeated data from the dataset, and at the same time keep a backup of the original data for retrospective and comparative analysis when needed;

[0015] The missing value processing specifically includes: statistically calculating the missing rate for each feature in the data, and regarding the data with the number of missing features greater than half of the overall sample feature scale as invalid data and eliminating it; for the remaining data, fill it using a specific method according to the data distribution;

[0016] The outlier processing specifically includes: according to the actual situation, confirm the threshold of the outliers; conduct visual analysis on the data to find the data beyond the threshold; process the outliers, and after processing, check the dataset again to ensure that the outliers have been processed and the basic features of the dataset have not changed significantly;

[0017] The data standardization processing specifically includes: classifying the existing features into continuous features and discrete features according to the data type, performing maximum - minimum standardization on the continuous features, and performing one - hot encoding on the discrete features;

[0018] The specific processing of the label data includes: adding a column of label columns to the dataset of Party A, setting its label data to "1", representing positive sample data; adding a column of label columns to the datasets of Party B and Party C respectively, and setting their label data to "0", representing unlabeled data.

[0019] Further, the specific steps of step S2 include: establishing a process for predicting candidate recommendations for unlabeled samples, and executing this process M rounds in a loop. The prediction process in the m-th round is as follows: randomly draw |P| samples from U with replacement, where |P| represents the number of samples in P, and form a negative sample dataset Nm with these |P| unlabeled samples; form a training set from P and Nm The samples in U that are not drawn form a prediction set Construct a vertical federated model with Gradient Boosting Decision Tree (GBDT) as the base learner, and train it on ; input into the trained model for prediction, and obtain the scores of each sample in as the prediction scores of the samples; specifically including the following steps:

[0020] S21: Use the bootstrap sampling method to randomly draw |P| samples from U with replacement, and form the m-th round negative sample dataset with these |P| unlabeled samples where are the feature vectors of Party B and Party C in the i-th negative sample in the m-th round respectively, represents the label of the i-th negative sample in the m-th round, i = 1,..., |P|, |N m | represents the number of samples in the m-th round negative sample dataset N m ; form the corresponding m-th round training set from the positive sample dataset P and the m-th round negative sample dataset N m where where are the feature vectors of Party B and Party C in the i-th sample in the m-th round respectively, represents the label of the i-th sample in the m-th round, i = 1, 2,..., 2|P|; the samples in the unlabeled sample dataset U that are not drawn form the corresponding m-th round prediction set where represents the feature vector of the j-th sample in the m-th round, j = 1, 2,.., |U|-|N m |;

[0021] S22: Use the GBDT algorithm as the base learner to construct a vertical federated model, and through the integration of T decision trees for Train, for the input data For each sample in To predict its output in the m-th round Where T represents the total number of decision trees in the m-th round, Represents the prediction result of the t-th tree in the m-th round, Represents the feature vector of the i-th sample in the m-th round, i = 1, 2,..., 2|P|; According to the prediction output And the true label Between the gradient values of the loss function, establish the first-order derivative and second-order derivative gradient histograms, and jointly find the global best split of the current node for all feature gradient histograms of both sides to construct the optimal decision tree;

[0022] S23: Use the trained GBDT decision model of the m-th round to predict the m-th round test set To obtain the prediction scores of each sample in the m-th round test set For each sample in Traverse all samples in U, and find the sample xu in U with the same ID as And assign the prediction score of i To xu As the m-th round prediction score of the sample xu i Denoted by i Where i = 1,..., |U| is the sample number in the unlabeled dataset U; If the sample xu in U Does not exist in the m-th round in i Then the m-th round prediction score corresponding to the sample xu Is equal to 0. i

[0023] Furthermore, in step S22, first initialize the prediction result of each sample in the m-th round To a random value, and then the specific process of the federated training of the t-th tree in the m-th round is as follows: S221: Start from Party B. First, calculate the first-order gradient

[0024] Of the loss function of the model in the m-th round for each sample of Party B And the second-order gradient Where i = 1, 2,..., 2|P|, Represents the prediction result of aggregating the first t - 1 trees in the m-th round for the sample Represents the true label of the i-th sample in the m-th round, ​​​​is the loss function; using additive homomorphic encryption for and are encrypted to obtain and Party B sends and to Party C;

[0025] S222: For Party C, establish the m-th round gradient histogram based on its own feature data, and send the encrypted gradient histogram to Party B;

[0026] S223: Party B decrypts the m-th round encrypted gradient histogram from Party C, calculates the optimal solution by enumerating each feature gradient histogram according to the splitting gain calculation formula, finds the global optimal splitting point, and returns the splitting information to Party C for parsing;

[0027] S224: Party C determines the threshold of the feature according to the feature number K opt and the threshold number V opt and divides the current sample space; then Party C establishes a lookup table locally, records the threshold of the selected feature, forms a record [record number, feature, threshold], and returns the record number and the sample space on the left side (I L ) after division to Party B;

[0028] S225: Party B divides the current node according to the received [record number, I L and associates the current node with [participant, record number]; Party B synchronizes the division information of the current node with Party C and enters the division of the next node;

[0029] S226: Iterate steps S222 - S225 until the training stop condition or the maximum depth of the tree is reached.

[0030] Furthermore, step S222 specifically includes the following steps:

[0031] S2221: For all Party C features in the current m-th round samples , sort all samples according to the feature values of each feature, and then divide the sorted samples into q categories by bucketing to obtain the corresponding feature thresholds of each category where k represents the feature number, represents the threshold of the q-th category of the feature numbered k in the m-th round;

[0032] S2222: According to the and obtained from Party B in the m-th round, Party C performs encrypted gradient information aggregation to construct the m-th round encrypted gradient histogram, that is

[0033]

[0034] where \(i = 1, 2, \cdots, 2|P|\) and \(v = 1, 2, \cdots, q\);

[0035] S2223: Party C sends the and calculated in the \(m\) - th round to Party B.

[0036] Furthermore, step S223 specifically includes the following steps:

[0037] S2231: Party B aggregates the first - order gradients and second - order gradients of all samples in the current node space, and executes where \(I\) represents all samples of the current node;

[0038] S2232: Party B decrypts the and obtained from Party C in the \(m\) - th round to get the decryption values in the \(m\) - th round and and calculates the following for all categories of all features of Party C in sequence to obtain and

[0039]

[0040]

[0041] where \(I\) L represents the sample space of the left child node after splitting, and \(I\) R represents the sample space of the right child node after splitting,

[0042] represents the sum of the first - order gradients of the loss functions of all samples in the sample space of the left child node in the \(m\) - th round,

[0043] represents the sum of the second - order gradients of the loss functions of all samples in the sample space of the left child node in the \(m\) - th round,

[0044] represents the sum of the first - order gradients of the loss functions of all samples in the sample space of the right child node in the \(m\) - th round,

[0045] represents the sum of the second - order gradients of the loss functions of all samples in the sample space of the right child node in the \(m\) - th round;

[0046] S2233: Calculate the best splitting value of the current node in the \(m\) - th round

[0047]

[0048] where λ is a hyperparameter;

[0049] S2234: For all thresholds of each feature of the sample a value can be obtained, and the largest value is selected, and the feature threshold is determined as the global optimal split for the m-th round. The global optimal split is represented by [participating party, feature number (K opt ), threshold number (V opt )], and the feature number (K opt ) and the threshold number (V opt ) are returned to Party C.

[0050] Furthermore, step S23 specifically includes the following steps:

[0051] S231: Party B queries the [participating party, record number] record associated with the current node; based on this record, Party B sends the sample number to be labeled and the record number to Party C, and asks about the next tree search direction, that is, to the left child node or the right child node;

[0052] S232: After receiving the sample number to be labeled and the record number, Party C compares the value of the corresponding feature in the sample to be labeled with the threshold in the record [record number, feature, threshold] in the local lookup table to obtain the next tree search direction; then, Party C sends the search decision to Party B;

[0053] S233: Party B receives the search decision sent by Party C and goes to the corresponding child node;

[0054] S234: Iterate steps S231 to S233 until a leaf node is reached, obtain the corresponding classification label and the weight of this label, so as to obtain the m-th round prediction score of the sample xu corresponding to the sample in U i

[0055]

[0056] where I represents the sample space of the leaf node, and λ is a hyperparameter.

[0057] Furthermore, in step S3, from the sum of the M-round prediction scores of each sample xu i in U and the sum of the number of times it appears in the M-round prediction set, calculate the sample xu i ​The probability of being predicted as a positive sample; sort all samples in U in descending order according to the probability of positive samples, select the top-ranked samples as reliable positive samples based on prior knowledge, and add them to the positive sample dataset P, while deleting them from U. The specific steps are as follows:

[0058] S31: From the sum of the prediction scores of each sample in U and the sum of the number of times it appears in the M-round prediction set, calculate the probability ρ that the sample is predicted as a positive sample i , and the calculation formula is as follows:

[0059]

[0060] where is an indicator function, indicating that if the sample xu in U i exists in the test set of the m-th round , then I = 1, otherwise I = 0;

[0061] S32: Use its probability ρ for all samples in U j to sort them in descending order, select the top θ samples as reliable positive samples, and add them to the positive sample dataset P, while deleting them from U, where the value of θ is set according to prior knowledge.

[0062] Furthermore, the specific steps of step S4 are as follows: Repeat steps S2 - S3 until the preset maximum number of iterations is reached. In each iteration, for the reliability and accuracy of the recommendation, only a small number of reliable positive samples can be selected each time. Therefore, in order to select a certain amount of reliable positive samples for batch recommendation, multiple iterations are required. In each iteration, step S3 will select some reliable positive samples from the unlabeled dataset U and add them to the positive sample dataset P, while deleting them from the unlabeled dataset U. In this case, the sample size of |P| will increase with the increase of the number of iterations, and the sample size of |U| will decrease with the increase of the number of iterations until the preset maximum number of iterations is reached. Thus, all the reliable positive samples selected from U during multiple iterations can be used as potential purchasers of financial products and be accurately batch-recommended to the financial product owners.

[0063] The beneficial effects of the present invention are as follows: The present invention aims at the problem that the financial product provider only has positive sample data and cannot make recommendations. By combining vertical federated learning and semi-supervised learning methods, while protecting the security and privacy of multi-party data, it jointly trains and predicts the financial product potential user recommendation model with multiple parties, effectively solves the batch recommendation problem with only a small number of positive samples and a large number of unlabeled samples, improves the reliability of the recommendation, and realizes the accurate and batch recommendation of financial product potential users.

[0064] Other advantages, objects and features of the present invention will be set forth in part in the following description, and in part will be obvious to those skilled in the art based on the examination of the following, or can be learned from the practice of the present invention. The objects and other advantages of the present invention can be achieved and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to make the objects, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings, where:

[0066] Figure 1 Flow schematic diagram of a method for recommending potential users of financial products for multi-party semi-supervised learning;

[0067] Figure 2 Schematic diagram of multi-party data sample alignment for vertical federated learning;

[0068] Figure 3 Schematic diagram of randomly sampling to form a training set and a test set;

[0069] Figure 4 Schematic diagram of constructing a gradient histogram;

[0070] Figure 5 Schematic diagram of the training process of a vertical federated GBDT model;

[0071] Figure 6 Schematic diagram of the prediction process of a vertical federated GBDT model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] The following illustrates the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention schematically, and the following embodiments and the features in the embodiments can be combined with each other without conflict.

[0073] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0074] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and cannot be understood as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0075] Taking the recommended application scenario of the provident fund contribution of flexible employees as an example, the relevant information of potential users who have not yet contributed to the provident fund of flexible employees is in multiple parties, and each party has requirements for data security and privacy protection. When vertical federated learning is carried out on the data of multiple parties and sample alignment is performed, Party A has the information of flexible employees who have contributed to the provident fund, and Party B and Party C have the information of flexible employees who have contributed to the provident fund and the information of potential flexible employees who have not contributed to the provident fund.

[0076] In this embodiment: Party A is the provident fund party and has the provident fund contribution data of flexible employees, including personal basic information, provident fund account number, provident fund balance, provident fund contribution period, provident fund contribution ratio, provident fund contribution amount and other information. Optionally, Party B can be the tax data party, including personal basic information, personal income tax, housing management tax and other information. Optionally, Party C can be the social security data party, including personal basic information, medical insurance, old-age insurance, unemployment insurance and other information.

[0077] Please refer to Figures 1 to 6 , a method for recommending potential users of financial products for multi-party semi-supervised learning, the method comprising the following steps:

[0078] S1: Establish a multi-party dataset for the contribution recommendation of flexible employees, including Party A, Party B, and Party C with the provident fund contribution information of flexible employees. Preprocess the multi-party dataset and perform sample alignment to form a positive sample dataset P and an unlabeled sample dataset U. The specific steps are as follows:

[0079] S11: Perform data preprocessing operations on the multi-party dataset for the contribution recommendation of flexible employees, including redundant data processing, missing value processing, outlier processing, data standardization processing, and label data processing, to obtain the dataset of Party A represents the i-th sample feature vector of Party A, and the vector dimension is d A , represents the corresponding label, where i = 1,..., a; the dataset of Party B denote the i-th sample feature vector of Party B, with the vector dimension of d B , denote the corresponding label, where i = 1, …, b; the dataset of Party C denote the i-th sample feature vector of Party C, with the vector dimension of d C , denote the corresponding label, where i = 1, …, c. The specific operations of data preprocessing are as follows:

[0080] The redundant data processing specifically includes: judging whether the data is repeated through certain fields in the data record. For the repeated data, only one of them is retained, and the other repeated data is deleted from the dataset. Meanwhile, a backup of the original data is retained for retrospective and comparative analysis when needed.

[0081] The missing value processing specifically includes: statistically calculating the missing rate for each feature in the data. For the data with the number of missing features greater than half of the overall sample feature scale, it is regarded as invalid data and excluded; for the remaining data, specific methods are used for filling according to the data distribution. Optionally, the method of filling with the median is used for missing value filling.

[0082] The outlier processing specifically includes: according to the actual situation, optionally, using the standard deviation method to confirm the threshold of outliers. Optionally, using the box plot for visual analysis of the data to find the data beyond the threshold. Optionally, the method of replacing outliers with the median is selected to process the outliers. After processing, the dataset is checked again to ensure that the outliers have been processed and the basic features of the dataset have not changed significantly.

[0083] The data standardization processing specifically includes: classifying the existing features into continuous features and discrete features according to the data type. Optionally, the maximum-minimum standardization is used for continuous features, and one-hot encoding is used for discrete features.

[0084] The label data processing specifically includes: adding a label column to the dataset of Party A, and setting its label data to "1", representing positive sample data; adding a label column to the datasets of Party B and Party C respectively, and setting their label data to "0", representing unlabeled data.

[0085] S12: For the datasets D B and D C of Party B and Party C, perform sample encryption alignment according to their sample IDs, retain the aligned sample data of Party B and Party C, discard the unaligned sample data, and obtain n samples. The aligned dataset of Party B is The dataset of Party C is denote the corresponding tag denote the corresponding tag, where i = 1, …, n

[0086] Optionally, the RSA algorithm and hash function can be used for encrypted alignment of samples

[0087] S13: Encrypt and align the dataset D A and and according to their sample IDs, and use the tripartite-aligned samples as positive samples to form the positive sample dataset P = {(xp i , yp i )}, where respectively represent the i-th sample feature vectors after alignment of Party A, Party B, and Party C, yp i ∈ {1} is the positive sample label, |P| represents the number of positive samples, i = 1, …, |P|; the unaligned samples of the three parties are used as unlabeled samples to form the unlabeled sample dataset U = {(xu i , yu i )}, where respectively represent the i-th sample feature vectors after alignment of Party B and Party C, yu i ∈ {0} is the unlabeled sample label, |U| represents the number of unlabeled samples, i = 1, …, |U|

[0088] Optionally, the RSA algorithm and hash function can be used for encrypted alignment of samples

[0089] S2: Establish a process for predicting candidate recommendations for unlabeled samples and execute this process M rounds in a loop. Optionally, set the total number of loop rounds M to 10. The prediction process in the m-th round is as follows: Randomly and with replacement, draw |P| samples from U, where |P| represents the number of samples in P, and form the negative sample dataset N m . Combine P and N m to form the training set The samples not drawn from U form the prediction set Construct a vertical federated model with the Gradient Boosting Decision Tree (GBDT) as the base learner and train it on . Input into the trained model for prediction to obtain the scores of each sample in

[0090] S21: Randomly and with replacement draw |P| samples from U using the bootstrap sampling method, and form the negative sample dataset for the m-th round with these |P| unlabeled samples. where are the feature vectors of Party B and Party C in the i-th negative sample of the m-th round respectively. represents the label of the i-th negative sample in the m-th round, i = 1, …, |P|, |N m | represents the number of samples in the negative sample dataset N m for the m-th round. The positive sample dataset P and the negative sample dataset N m for the m-th round form the corresponding training set for the m-th round. where represent the feature vectors of Party B and Party C in the i-th sample of the m-th round respectively. represents the label of the i-th sample in the m-th round, i = 1, 2, ..., 2|P|; the samples not drawn from the unlabeled sample dataset U form the corresponding prediction set for the m-th round. where represents the feature vector of the j-th sample in the m-th round, j = 1, 2, ..., |U| - |N m |.

[0091] S22: Use the GBDT algorithm as the base learner to construct a vertical federated model, and train through the integration of T decision trees for the input data to predict its output for the m-th round for each sample where T represents the total number of decision trees for the m-th round. Optionally, set the value of the total number of decision trees T to 40. represents the prediction result of the t-th tree for the m-th round. represents the feature vector of the i-th sample for the m-th round, i = 1, 2, ..., 2|P|. According to the prediction output and the true label between them, establish the gradient histograms of the first-order derivative and the second-order derivative, and jointly find the global best split of the current node for all feature gradient histograms of both parties to construct the optimal decision tree.

[0092] First, initialize the prediction result for each sample for the m-th round to a random value. Optionally, the random value of is -1 or 1. Then the specific process of the federated training of the t-th tree for the m-th round is as follows:

[0093] S221: Starting from Party B, first for each sample of Party B Calculate the first-order gradient of the model loss function in the m-th round And the second-order gradient Where i = 1, 2,..., 2|P|, Represents the prediction result of aggregating the first t - 1 trees in the m-th round for the sample Of, Represents the true label of the i-th sample in the m-th round, Is the loss function. Use additive homomorphic encryption for And To encrypt, obtaining And Party B sends And To Party C.

[0094] Optionally, here the loss function uses the mean squared error loss function, and the homomorphic encryption scheme adopted is Paillier homomorphic encryption. Use To represent that data a has been homomorphically encrypted.

[0095] S222: For Party C, build the gradient histogram in the m-th round based on its own feature data, and send the encrypted gradient histogram to Party B. The specific process is as follows:

[0096] S2221: For all Party C features in the current m-th round of samples Sort all samples according to the feature values of each feature, and then divide the sorted samples into q categories by bucketing to obtain the corresponding feature thresholds for each category Where k represents the feature number, Represents the threshold of the q-th category of the feature numbered k in the m-th round. Optionally, here the category q is set to 10.

[0097] S2222: According to the And Obtained from Party B in the m-th round, Party C performs encrypted gradient information aggregation to construct the encrypted gradient histogram in the m-th round, that is

[0098]

[0099] Where i = 1, 2,..., 2|P|, v = 1, 2,..., q.

[0100] S2223: Party C sends the calculated And In the m-th round to Party B.

[0101] S223: Party B decrypts the m-th round of encrypted gradient histograms received from Party C. According to the split gain calculation formula, it enumerates each feature gradient histogram to calculate the optimal solution, finds the global optimal split point, and returns the split information to Party C for parsing. The specific process is as follows:

[0102] S2231: Party B aggregates the first-order gradients and second-order gradients of all samples in the current node space and executes where I represents all samples of the current node.

[0103] S2232: Party B decrypts the m-th round and obtains the decrypted values of the m-th round and and calculates the following for all categories of all features of Party C in sequence to obtain and

[0104]

[0105]

[0106] where I L represents the sample space of the left child node after splitting, and I R represents the sample space of the right child node after splitting.

[0107] represents the sum of the first-order gradients of the loss functions of all samples in the sample space of the left child node in the m-th round,

[0108] represents the sum of the second-order gradients of the loss functions of all samples in the sample space of the left child node in the m-th round,

[0109] represents the sum of the first-order gradients of the loss functions of all samples in the sample space of the right child node in the m-th round,

[0110] represents the sum of the second-order gradients of the loss functions of all samples in the sample space of the right child node in the m-th round.

[0111] S2233: Calculate the best split value of the current node in the m-th round

[0112]

[0113] where λ is a hyperparameter. Optionally, the hyperparameter λ is set to 0.5.

[0114] S2234: For all thresholds of each feature of the samples A value can be obtained for each value, and the largest one is selected as the value. The feature threshold is determined as the global optimal split for the m-th round. The global optimal split can be represented by [participant, feature number (K opt ), threshold number (V opt )], and the feature number (K opt ) and threshold number (V opt ) are returned to Party C.

[0115] S224: Party C determines the threshold of the feature based on the feature number K opt and threshold number V opt sent from Party B, and divides the current sample space. Then, Party C locally creates a lookup table to record the threshold of the selected feature, forming a record [record number, feature, threshold], and returns the record number and the sample space (I L ) on the left side after division to Party B.

[0116] S225: Party B will divide the current node according to the received [record number, I L and associate the current node with [participant, record number]. Party B synchronizes the division information of the current node with Party C and proceeds to the division of the next node.

[0117] S226: Iterate steps S222 - S225 until the training stop condition is reached or the maximum depth of the tree is reached.

[0118] S23: Use the trained GBDT decision model of the m-th round to predict the m-th round test set to obtain the prediction scores of each sample in the m-th round test set . Traverse all samples in U, find the sample xu in U with the same ID as i , and assign the prediction score of to xu i as the m-th round prediction score of the sample xu i , denoted by , where i = 1,..., |U| is the sample number in the unlabeled dataset U. If the sample xu i in the m-th round does not exist in , then the m-th round prediction score i corresponding to the sample xu is equal to 0. The specific process is as follows:

[0119] S231: Party B queries the [participant, record number] record associated with the current node. Based on this record, Party B sends the sample number to be labeled and the record number to Party C, and asks about the next tree search direction, that is, the left child node or the right child node.

[0120] S232: After receiving the sample number to be labeled and the record number, Party C compares the value of the corresponding feature in the sample to be labeled with the threshold in the record [record number, feature, threshold] in the local lookup table to obtain the next tree search direction. Then, Party C sends the search decision to Party B.

[0121] S233: Party B receives the search decision sent by Party C and goes to the corresponding child node.

[0122] S234: Iterate steps S231 - S233 until reaching a leaf node, obtaining the corresponding classification label and the weight of this label, then the m - th round prediction score of the sample corresponding to xu in U i can be obtained.

[0123]

[0124] where I represents the sample space of the leaf node and λ is a hyperparameter. Optionally, the hyperparameter λ is set to 0.5.

[0125] S3: From the sum of the M - round prediction scores of each sample xu in U i and the sum of the number of times it appears in the M - round prediction set, calculate the probability that the sample xu i is predicted as a positive sample. Sort all samples in U according to the probability of being a positive sample from large to small, and select the top - ranked samples as reliable positive samples according to prior knowledge, and add them to the positive sample dataset P, and delete them from U at the same time. Optionally, the total number of rounds M is set to 10.

[0126] S31: From the sum of the prediction scores of each sample in U and the sum of the number of times it appears in the M - round prediction set, calculate the probability ρ i that the sample is predicted as a positive sample. The calculation formula is as follows:

[0127]

[0128] where is the indicator function, indicating that if the sample xu in U i exists in the test set of the m - th round , then I = 1, otherwise I = 0.

[0129] S32: Use its probability ρ for all samples in U iSort in descending order, select the top θ samples as reliable positive samples, add them to the positive sample dataset P, and delete them from U at the same time, where the value of θ is set according to prior knowledge. Optionally, θ is set to 0.1.

[0130] S4: Repeat steps S2 - S3 until the preset maximum number of iterations is reached. Optionally, the maximum number of iterations is set to 5. In each iteration, for the reliability and accuracy of recommendations, only a small number of reliable positive samples can be selected each time. Therefore, in order to select a certain amount of reliable positive samples for batch recommendation, multiple iterations are required. In each iteration, step S3 will select some reliable positive samples from the unlabeled dataset U and add them to the positive sample dataset P, and at the same time delete them from the unlabeled dataset U. In this case, the sample size of |P| will increase with the increase of the number of iterations, and the sample size of |U| will decrease with the increase of the number of iterations until the preset maximum number of iterations is reached. Thus, all the reliable positive samples selected from U during multiple iterations can be used as potential provident fund depositors and recommended to the provident fund side in batches.

[0131] For example, personal basic information, personal income tax, housing management tax, etc. of a batch of customers in Party B's tax data, which are relevant to the application scenario; personal basic information, medical insurance, old-age insurance, unemployment insurance, etc. of this batch of customers in Party C's social security data, which are relevant to the application scenario information. After preprocessing, they are respectively put into the trained vertical federated GBDT model. Each party judges whether there are potential users of flexible employment provident fund deposits among this batch of customers according to the characteristics of this batch of customers, and then recommends the customers judged as potential users to the provident fund side in batches.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for recommending potential users of financial products in multi-party semi-supervised learning, characterized in that: It includes the following steps: S1: Establish a multi-party dataset for recommending potential users of financial products that includes information of users who have purchased financial products and information of other multi-party users, and perform preprocessing and sample alignment on the multi-party dataset for recommending potential users of financial products to construct a positive sample dataset and an unlabeled sample dataset; S2: Conduct random sampling with replacement in the unlabeled sample dataset to establish a negative sample dataset; use the negative sample dataset and the positive sample dataset to construct a training set, and use the samples in the unlabeled sample dataset that have not been sampled to construct a prediction set; construct a vertical federated learning model based on base learners, train it on the training set, and make predictions on the prediction set to obtain the prediction scores of each sample in the prediction set; S3: Repeat the sampling, training, and prediction processes in step S2 multiple times; calculate the probability that a sample is predicted as a positive sample according to the sum of the prediction scores of each sample in the unlabeled sample dataset and the number of times it appears in the prediction set; sort all samples in the unlabeled sample dataset in descending order according to the probability of positive samples, select the top-ranked samples as reliable positive samples according to prior knowledge, add them to the positive sample dataset, and delete them from the unlabeled sample dataset at the same time; S4: Repeat steps S2 - S3 until the preset maximum number of iterations is reached; thus, use all the reliable positive samples selected from the unlabeled sample dataset as potential users and make accurate batch recommendations to the financial product providers.

2. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 1, wherein: The multi-party dataset for recommending potential users of financial products in step S1 includes Party A that has information of users who have purchased financial products, and other multi-party data sources, Party B and Party C, other than Party A; step S1 specifically includes the following steps: S11: Preprocess the multi-party dataset recommended for potential users of financial products, including redundant data processing, missing value processing, outlier processing, data standardization processing, and labeled data processing, to obtain Dataset A represents the feature vector of the i-th sample in Party A, and the vector dimension is d A , represents the corresponding label, where i = 1, …, a; Dataset B represents the feature vector of the i-th sample in Party B, and the vector dimension is d B , represents the corresponding label, where i = 1, …, b; Dataset C represents the feature vector of the i-th sample in Party C, and the vector dimension is d C , represents the corresponding label, where i = 1, …, c; S12: Encrypt and align the samples of datasets D of Party B and Party C B and D C according to their sample IDs, retain the sample data aligned by Party B and Party C, discard the unaligned sample data, and obtain n samples; the aligned dataset of Party B is The dataset of Party C is denote the corresponding label, denote the corresponding label, where i = 1, …, n; S13: For the dataset D A , and , encrypt and align the samples according to their sample IDs, and use the samples aligned by the three parties as positive samples to form the positive sample dataset P = {(xp i , yp i ), where respectively represent the i-th sample feature vectors after alignment by Party A, Party B, and Party C, yp i ∈{1} is the positive sample label, |P| represents the number of positive samples, and i = 1, …, |P|; the samples not aligned by the three parties are used as unlabeled samples to form the unlabeled sample dataset U = {(xu i , yu i ), where respectively represent the i-th sample feature vectors after alignment by Party B and Party C, yu i ∈{0} is the unlabeled sample label, |U| represents the number of unlabeled samples, and i = 1, …, |U|.

3. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 2, wherein: In step S11, the redundant data processing specifically includes: judging whether the data is repeated through certain fields in the data record, for repeated data, only keep one of them, delete other repeated data from the dataset, and at the same time keep a backup of the original data for retrospective and comparative analysis when needed; The missing value processing specifically includes: statistically calculating the missing rate for each feature in the data, regarding the data with the number of missing features greater than half of the overall sample feature scale as invalid data and eliminating it; filling the remaining data using specific methods according to the data distribution; The outlier processing specifically includes: confirming the threshold of outliers according to the actual situation; performing visual analysis on the data to find the data beyond the threshold; processing the outliers, and then checking the dataset again to ensure that the outliers have been processed and the basic features of the dataset have not changed significantly; The data standardization processing specifically includes: classifying the existing features into continuous features and discrete features according to the data type, performing maximum - minimum standardization on the continuous features, and performing one - hot encoding on the discrete features; The label data processing specifically includes: adding a label column to the Party A dataset, setting its label data as "1", representing positive sample data; adding a label column to the Party B and Party C datasets respectively, setting their label data as "0", representing unlabeled data.

4. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 3, wherein: The specific steps of step S2 include: establishing a prediction process for candidate recommendations of unlabeled samples, and looping through this process M times. The prediction process in the m-th round is as follows: randomly and with replacement, draw |P| samples from U, where |P| represents the number of samples in P, and form a negative sample dataset N with these |P| unlabeled samples m ; combine P and N m to form a training set The samples in U that are not drawn form a prediction set Construct a vertical federated model with Gradient Boosting Decision Tree (GBDT) as the base learner, and train it on ; input into the trained model for prediction, and obtain the scores of each sample in, as the prediction scores of the samples; specifically including the following steps: S21: Randomly and with replacement draw |P| samples from U using the bootstrap sampling method, and form the negative sample dataset for the m-th round with these |P| unlabeled samples. where are the feature vectors of Party B and Party C in the i-th negative sample of the m-th round respectively, represents the label of the i-th negative sample in the m-th round, i = 1, …, |P|, |N m | represents the number of samples in the negative sample dataset N m for the m-th round; The positive sample dataset P and the negative sample dataset N m for the m-th round form the corresponding training set for the m-th round where represent the feature vectors of Party B and Party C in the i-th sample of the m-th round respectively, represents the label of the i-th sample in the m-th round, i = 1, 2, …, 2|P|; The samples not drawn from the unlabeled sample dataset U form the corresponding prediction set for the m-th round where represents the feature vector of the j-th sample in the m-th round, j = 1, 2, …, |U| - |N m |; S22: Use the GBDT algorithm as the base learner to construct a vertical federated model, and train it through the integration of T decision trees to predict the m-th round output for each sample in the input data where T represents the total number of decision trees in the m-th round, denotes the prediction result of the t-th tree in the m-th round, denotes the feature vector of the i-th sample in the m-th round, i = 1, 2,..., 2|P|; according to the gradient value of the loss function between the predicted output and the true label establish the first-order and second-order derivative gradient histograms, and jointly find the global best split of the current node for the gradient histograms of all features of both parties to construct the optimal decision tree; S23: Use the trained GBDT decision model of the m-th round to predict the m-th round test set to obtain the prediction scores of each sample in the m-th round test set Traverse all samples in U, and find the sample xu in U with the same ID as Assign the prediction score of to xu i as the prediction score of the m-th round of the sample xu denoted by i where i = 1, …, |U| is the sample number in the unlabeled dataset U; if the sample xu in U i does not exist in the m-th round of then the prediction score of the m-th round corresponding to the sample xu i is equal to 0 i ​​​ 5. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 4, characterized in that: In step S22, first, initialize the prediction result of each sample in the m-th round to a random value. Then, the specific process of the federated training of the t-th tree in the m-th round is as follows: ​ S221: Starting from Party B, first for each sample of Party B calculate the first-order gradient of the model loss function in the m-th round and the second-order gradient where i = 1, 2,..., 2|P|, represents the prediction result of aggregating the first t - 1 trees for the sample in the m-th round, represents the true label of the i-th sample in the m-th round, is the loss function; use additive homomorphic encryption for and to encrypt them, obtaining and Party B sends and to Party C; S222: For Party C, construct the m-th round of gradient histograms based on its own feature data, and send the encrypted gradient histograms to Party B. S223: Party B decrypts the m-th round of encrypted gradient histograms received from Party C, calculates the optimal solution by enumerating each feature gradient histogram according to the split gain calculation formula, finds the global optimal split point, and returns the split information to Party C for parsing. S224: Party C determines the threshold of the feature based on the feature number K sent from Party B opt and the threshold number V opt to determine the threshold of the feature and partition the current sample space; then Party C locally creates a lookup table to record the threshold of the selected feature, forming a record [record number, feature, threshold], and returns the record number and the sample space (I L ) on the left side after partitioning to Party B; S225: Party B divides the current node according to the received [record number, I L , and associates the current node with [participating party, record number]; Party B synchronizes the division information of the current node with Party C and proceeds to divide the next node; S226: Iterate steps S222 - S225 until the training stop condition is reached or the maximum depth of the tree is achieved.

6. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 5, characterized in that: Step S222 specifically includes the following steps: S2221: For the current m-th round of samples For all C-party features in it, sort all samples according to the eigenvalue of each feature, and then divide the sorted samples into q categories by bucketing to obtain the corresponding feature thresholds for each category where k represents the feature number, represents the threshold of the q-th category of the feature numbered k in the m-th round; S2222: According to the and received from Party B in the m-th round, Party C performs encrypted gradient information aggregation to construct the encrypted gradient histogram in the m-th round, namely where i = 1, 2,..., 2|P|, v = 1, 2,..., q; S2223: Party C will send the and calculated in the m-th round to Party B.

7. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 5, characterized in that: Step S223 specifically includes the following steps: S2231: Party B aggregates the first-order gradients and second-order gradients of all samples in the current node space and executes where I represents all samples of the current node; S2232: Party B decrypts the m-th round and obtaining the decryption value of the m-th round and Performs the following calculations on all categories of all features of Party C in sequence to obtain and where I L represents the sample space of the left child node after splitting, and I R represents the sample space of the right child node after splitting, represents the sum of the first-order gradients of the loss functions of all samples in the sample space of the left child node in the m-th round, represents the sum of the second-order gradients of the loss functions of all samples in the sample space of the left child node in the m-th round, represents the sum of the first-order gradients of the loss functions of all samples in the sample space of the right child node in the m-th round, represents the sum of the second-order gradients of the loss functions of all samples in the sample space of the right child node in the m-th round; S2233: Calculate the optimal splitting value of the current node in the m-th round where λ is a hyperparameter; S2234: For all thresholds of each feature of the sample a value can be obtained. Select the largest value, and determine that the feature threshold is the global optimal segmentation in the m-th round. The global optimal segmentation is represented by [participating party, feature number (K opt ), threshold number (V opt )], and return the feature number (K opt ) and the threshold number (V opt ) to Party C.

8. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 4, wherein: Step S23 specifically includes the following steps: S231: Party B queries the [participating party, record number] record associated with the current node; based on this record, Party B sends the sample number to be labeled and the record number to Party C, and asks about the next tree search direction, i.e., to the left child node or the right child node. S232: After receiving the sample number to be labeled and the record number, Party C compares the value of the corresponding feature in the sample to be labeled with the threshold in the record [record number, feature, threshold] in the local lookup table to obtain the next tree search direction; then, Party C sends the search decision to Party B. S233: Party B receives the search decision sent by Party C and goes to the corresponding child node. S234: Iterate through steps S231 - S233 until a leaf node is reached, obtaining the corresponding classification label and the weight of this label, thereby obtaining the sample in U corresponding to the sample xu in U i the predicted score in the m-th round where I represents the sample space of the leaf node, and λ is a hyperparameter.

9. The method for recommending potential users of financial products in multi-party semi-supervised learning according to claim 4, characterized in that: In the step S3, for each sample xu in U i calculate the sum of the prediction scores in M rounds and the sum of the number of occurrences in the M-round prediction set, and calculate the probability that the sample xu i is predicted as a positive sample; sort all the samples in U according to the probability of positive samples from large to small, select the samples with the top rankings as reliable positive samples according to prior knowledge, add them to the positive sample data set P, and delete them from U at the same time. The specific steps are as follows: S31: Calculate the sum of the prediction scores of each sample in U and the total number of times it appears in the M-round prediction set, and calculate the probability ρ that the sample is predicted as a positive sample i , and the calculation formula is as follows: wherein is an indicator function, indicating that if the sample xu in U i exists in the test set in the m-th round then I = 1, otherwise I = 0; S32: Use the probability ρ for all samples in U i Sort them in descending order, select the top θ samples as reliable positive samples, add them to the positive sample dataset P, and delete them from U at the same time, where the value of θ is set according to prior knowledge.