Financial user portrait generation method, system and device and medium
By combining differential privacy noise addition and feature decoupling generation model, the problems of accuracy and privacy protection in the generation of financial user profiles are solved, and high-quality user profile generation is achieved under small data conditions.
Patent Information
- Application Number
- CN202511058669.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies for generating financial user profiles suffer from insufficient accuracy and comprehensiveness, as well as inadequate privacy protection. In particular, data quality is poor in small datasets, and there is a risk of privacy breaches.
The original dataset is processed by differential privacy noise addition to generate a noisy dataset. A feature decoupling generation model is used to generate data. A hierarchical mechanism decouples feature subspaces of different feature attribute types. Combined with adaptive noise addition and dynamic correlation strength, the target dataset is generated. Finally, a user profile is generated through a machine learning model.
While ensuring privacy protection, it improves the accuracy and comprehensiveness of user profiles, avoids incorrect feature binding caused by noisy and ambiguous associations, and enhances data quality and security.
Smart Images

Figure CN120995233A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, system, device, and medium for generating financial user profiles. Background Technology
[0002] In the financial sector, building accurate, comprehensive, and privacy-protected user profiles has become one of the core requirements for banks, insurance companies, and other financial institutions to conduct risk assessments, marketing activities, and other business activities.
[0003] Currently, the relevant technologies are usually based on big data technology to collect and analyze a certain scale of user data, and generate user profiles by directly adding noise to the user data. The accuracy and comprehensiveness of the user profiles generated in this way are not satisfactory.
[0004] Therefore, the problems with the relevant technologies still need to be solved and optimized. Summary of the Invention
[0005] The purpose of this invention is to at least partially solve one of the technical problems existing in the related art.
[0006] Therefore, one objective of this invention is to provide a method, system, device, and medium for generating financial user profiles, wherein the method can improve the accuracy and comprehensiveness of user profile generation.
[0007] To achieve the above-mentioned technical objectives, the technical solutions adopted in the embodiments of this application include:
[0008] In a first aspect, embodiments of this application provide a method for generating financial user profiles, including:
[0009] Obtain the original dataset of the feature decoupling generation model and the target user, wherein the original dataset includes several user data;
[0010] Differential privacy noise is applied to the original dataset to obtain a noisy dataset. The noisy dataset includes several noisy user data sets. The noise intensity added to each noisy user data set is not exactly the same. The noise intensity added is positively correlated with the sensitivity and importance of the user data set.
[0011] The noisy dataset is input into the feature decoupling generation model to generate data, thus obtaining the target dataset.
[0012] User profiles are generated from the target dataset to obtain user profiles of the target users.
[0013] In addition, the method according to the above embodiments of this application may also have the following additional technical features:
[0014] Furthermore, in one embodiment of this application, the step of performing differential privacy noise addition on the original dataset to obtain a noisy dataset includes:
[0015] The original dataset is classified to obtain several classification information, and each classification information corresponds to one user data;
[0016] Based on all the classification information, adaptive noise is added to the original dataset to obtain the noisy dataset.
[0017] Furthermore, in one embodiment of this application, the step of adaptively adding noise to the original dataset based on all the classification information to obtain the noisy dataset includes:
[0018] Obtain the sensitivity rules corresponding to the classification information;
[0019] Based on the sensitivity rules, data sensitivity analysis is performed on the user data to obtain the privacy budget and several field sensitivities of the user data, with each field sensitivity corresponding to a data field in the user data;
[0020] Based on the privacy budget and the sensitivity of all the fields, the user data is noise-added to obtain the noise-added user data.
[0021] Furthermore, in one embodiment of this application, the step of adding noise to the user data based on the privacy budget and the sensitivity of all the fields to obtain the noisy user data includes:
[0022] Perform joint field analysis on the sensitivity of all the aforementioned fields to obtain the target sensitivity;
[0023] Based on the target sensitivity, noise intensity analysis is performed on the privacy budget to obtain the added noise intensity;
[0024] The user data is noise-added based on the added noise intensity to obtain the noise-added user data.
[0025] Furthermore, in one embodiment of this application, the step of inputting the noisy dataset into the feature decoupling generation model for data generation to obtain the target dataset includes:
[0026] Obtain the dynamic correlation strength;
[0027] The noisy dataset is subjected to feature decomposition to obtain a first feature subspace and a second feature subspace, wherein the feature attribute types in the first feature subspace are different from the feature attribute types in the second feature subspace.
[0028] Heterogeneous generation is performed on the first feature subspace and the second feature subspace to obtain a plurality of first target features corresponding to the first feature subspace and a plurality of second target features corresponding to the second feature subspace;
[0029] Based on the dynamic association strength, all the first target features and the second target features are dynamically coupled to obtain the target dataset.
[0030] Furthermore, in one embodiment of this application, the method further includes:
[0031] Obtain the compensation dataset, which is the noisy dataset or the target dataset;
[0032] Distortion detection is performed on the compensation dataset to obtain distortion detection results;
[0033] If the distortion detection result indicates that the compensation dataset is distorted, then the compensation dataset is repaired to obtain the repaired compensation dataset.
[0034] Furthermore, in one embodiment of this application, the step of performing data repair on the compensation dataset to obtain the repaired compensation dataset includes:
[0035] Obtain several original repair rules;
[0036] Distortion type analysis is performed on the compensation dataset to obtain the target distortion type of the compensation dataset;
[0037] Based on the target distortion type, all the original repair rules are filtered to obtain the target repair rule, which is the original repair rule that corresponds to the target distortion type among all the original repair rules.
[0038] According to the target repair rules, the compensation dataset is repaired to obtain the repaired compensation dataset.
[0039] Secondly, embodiments of this application provide a system for generating financial user profiles, including:
[0040] The first processing unit is used to obtain the feature decoupling generation model and the original dataset of the target user, wherein the original dataset includes several user data.
[0041] The second processing unit is used to perform differential privacy noise addition on the original dataset to obtain a noisy dataset. The noisy dataset includes several noisy user data. The noise intensity added to each noisy user data is not exactly the same. The noise intensity added is positively correlated with the sensitivity and importance of the user data.
[0042] The third processing unit is used to input the noisy dataset into the feature decoupling generation model to generate data and obtain the target dataset.
[0043] The fourth processing unit is used to generate user profiles from the target dataset to obtain user profiles of the target users.
[0044] Thirdly, embodiments of this application also provide an electronic device, including:
[0045] At least one processor;
[0046] At least one memory for storing at least one program;
[0047] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.
[0048] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a processor-executable program, which, when executed by the processor, is used to implement the above-described method.
[0049] The advantages and beneficial effects of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application:
[0050] This application discloses a method, system, device, and medium for generating financial user profiles. The method involves acquiring a feature decoupling generation model and a raw dataset of target users, the raw dataset comprising several user data points. Differential privacy noise is applied to the raw dataset to obtain a noisy dataset, which includes several noisy user data points. The noise intensity of each noisy user data point is not entirely the same, and the noise intensity is positively correlated with the sensitivity and importance of the user data. The noisy dataset is input into the feature decoupling generation model for data generation to obtain a target dataset. Finally, user profiles are generated from the target dataset to obtain the user profile of the target user. This method, by applying differential privacy noise to the raw dataset and inputting the noisy raw dataset into the feature decoupling generation model for generation, can effectively improve the data quality of the dataset used to generate user profiles while ensuring privacy protection, thereby improving the accuracy and comprehensiveness of user profile generation. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following description is provided with accompanying drawings of the relevant technical solutions in the embodiments of this application or the prior art. It should be understood that the accompanying drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions in this application. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0052] Figure 1 A flowchart illustrating a method for generating a financial user profile, provided in an embodiment of this application;
[0053] Figure 2 A schematic diagram of the framework of a financial user profile generation system provided in an embodiment of this application;
[0054] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0055] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application. The step numbers in the following embodiments are set only for ease of explanation, and there is no limitation on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0057] In the financial sector, building accurate, comprehensive, and privacy-protected user profiles has become one of the core requirements for banks, insurance companies, and other financial institutions to conduct risk assessments, marketing activities, and other business activities.
[0058] Currently, related technologies typically rely on big data technology to collect and analyze a certain scale of user data, generating user profiles by directly adding noise to the data. However, constructing accurate and comprehensive user profiles requires a large amount of high-quality data, and the available high-quality data in practical applications is often limited. For example, high-quality customer data is scarce in credit card risk assessment, and historical data on rare high-risk events is lacking in insurance business. Furthermore, the noise directly introduced into the data distorts it, reduces data quality, and significantly decreases the accuracy and comprehensiveness of the generated user profiles. For instance, according to experimental data, when the privacy budget for privacy protection is 1, adding noise to credit card user data reduces the accuracy of the generated user profiles applied to risk assessment models from 90% to around 50%.
[0059] Furthermore, some related technologies directly generate synthetic data similar to user data through generative models (such as GAN models) to address data shortages and then use this synthetic data to create user profiles. However, this approach may leak real user privacy data due to the generative model's memory limitations. For example, the generative model might leak user privacy information due to overfitting during training, or it might infer original user information through reverse engineering, posing a risk of privacy breaches and resulting in low data security. Additionally, when processing noisy data, this type of generative model might learn residual, noise-blurred attribute associations, leading to biased data. For instance, records of "high spending amounts" (even with added noise) often appear alongside "dining category" records; the generative model might incorrectly associate high spending with dining categories, even though in the real world, high-spending users might be more inclined towards luxury goods or travel. Moreover, this method generates poor-quality data with small datasets, resulting in inaccurate user profiles. For example, with limited training data, the generative model might produce "illusionary" data, creating data points inconsistent with the actual data distribution, leading to biased user profiles.
[0060] Furthermore, a very small number of related technologies combine differential privacy protection and generative models to generate user profiles. However, this approach often fails to achieve a good balance between strict privacy protection and high profile quality, resulting in poor accuracy and data security. For example, training a generative model directly on noisy data may cause the model to learn noise rather than the true data distribution, leading to poor accuracy in the subsequently generated user profiles. On the other hand, applying differential privacy technology after the generative model output may fail to effectively protect intermediate features, making the subsequently generated user profiles susceptible to privacy leaks and resulting in low data security.
[0061] It should be noted that the aforementioned related technologies are only used to assist in understanding the technical solutions of this application and do not mean that they belong to the publicly disclosed prior art.
[0062] In view of this, embodiments of this application provide a method, system, device, and medium for generating financial user profiles. This method generates user profiles by differentially privacy-preserving noise addition to the original dataset and inputting the noise-added original dataset into a feature decoupling generation model. This can effectively improve the data quality of the dataset used to generate user profiles while ensuring the security of privacy protection, thereby improving the accuracy and comprehensiveness of user profile generation.
[0063] Furthermore, this method decouples and generates data in noisy datasets through a feature decoupling generation model. Specifically, it generates feature subspaces of different feature attribute types through a hierarchical mechanism, and then generates target features corresponding to each feature subspace. This enables the decoupling generation of privacy and non-privacy information and differentiated privacy protection strategies, preventing the generation model from incorrectly binding features due to the blurred correlation of noise in the data. This effectively improves the quality of the synthetic data generated by the generation model, thereby helping to improve the accuracy of generating user profiles.
[0064] Furthermore, this method adds differential privacy noise to the original dataset, specifically by adaptively adding noise to user data of different classification types. This allows for the addition of appropriate noise to user data of each classification type, preserving the statistical characteristics of user data to the greatest extent possible while protecting privacy. This enables the subsequent feature decoupling generation model to generate a target dataset based on the noise-added dataset, achieving strict privacy protection and high data quality, thereby improving the accuracy and data security of the subsequently generated user profiles.
[0065] Reference Figure 1 In this embodiment of the application, a method for generating a financial user profile includes:
[0066] Step 110: Obtain the feature decoupling generation model and the original dataset of the target user, wherein the original dataset includes data of several users;
[0067] In this embodiment of the application, the feature decoupling generation model can be an optimized generative adversarial (GAN) model, and the original dataset of the target user can be a collection of user data such as user ID, age, gender, income, and transaction records of financial users.
[0068] Step 120: Perform differential privacy noise addition on the original dataset to obtain a noisy dataset. The noisy dataset includes several noisy user data. The noise intensity added to each noisy user data is not exactly the same. The noise intensity added is positively correlated with the sensitivity and importance of the user data.
[0069] In this embodiment of the application, random noise can be added to each user data in the original dataset based on the optimized differential privacy technology, wherein the noise intensity of the random noise added to each user data is positively correlated with its sensitivity and importance.
[0070] In some embodiments, performing differential privacy noise addition on the original dataset to obtain a noisy dataset includes:
[0071] The original dataset is classified to obtain several classification information, and each classification information corresponds to one user data;
[0072] In this embodiment of the application, for any user data in the original dataset, the user data can be classified to obtain the classification information of the user data. Specifically, in the financial field, if the user data is income, single transaction amount, account balance, etc., the corresponding classification information can be numerical; or, if the user data is occupation, transaction type, payment method, etc., the corresponding classification information can be categorical; or, if the user data is user login count, user transaction count, user loan application count, etc., the corresponding classification information can be count.
[0073] Based on all the classification information, adaptive noise is added to the original dataset to obtain the noisy dataset.
[0074] Further, the step of adaptively adding noise to the original dataset based on all the classification information to obtain the noisy dataset includes:
[0075] Obtain the sensitivity rules corresponding to the classification information;
[0076] Based on the sensitivity rules, data sensitivity analysis is performed on the user data to obtain the privacy budget and several field sensitivities of the user data, with each field sensitivity corresponding to a data field in the user data;
[0077] In this embodiment of the application, if the classification information is numerical, its sensitivity rule can be defined as the maximum change in query results between two adjacent datasets. This numerical sensitivity rule can be expressed as:
[0078] Δf1=max(f(D)-f(D ′ ))
[0079] Where Δf1 is the numerical sensitivity rule function; max(·) is the maximum value function; and f(D) is the first neighboring dataset D. ′ The function; f(D) ′) is the second adjacent dataset D ′ The function.
[0080] If the classification information is categorical, its sensitivity rule can be defined as the maximum impact a single record (i.e., a data field in user data) may have on the target classification statistics (such as category technique, frequency distribution) in the database. This can be achieved through hierarchical sensitivity. Specifically, many financial categories have a hierarchical structure; for example, transaction types can be divided into "Restaurant" (including subcategories such as "Chinese food" and "Western food"). When calculating sensitivity, sensitivity values can be assigned hierarchically. For top-level categories (such as "Restaurant" vs. "Shopping"), adding or removing a record at most changes the count of one top-level category, so the sensitivity is 1. For subcategories, the sensitivity can be set to 1 divided by the total number of subcategories under that parent category. For example, if "Restaurant" only has two subcategories, "Chinese food" and "Western food," then changing one record in "Chinese food" will have a sensitivity of 1 / 2 on the "Chinese food" count.
[0081] If the classification information is count-based, its sensitivity rule can be defined as the maximum change in count that a single record may cause within a specific time window.
[0082] Understandably, for any user data in the original dataset, data sensitivity analysis can begin by analyzing several data fields based on sensitivity rules to obtain the field sensitivity of each data field. The field sensitivity of all data fields can be used to characterize the sensitivity and importance of the user data. Next, based on the field sensitivity of each data field, a field budget is allocated to each data field. The smaller the field budget, the stronger the privacy protection, but the lower the data availability. For example, a smaller field budget is allocated to a data field with higher sensitivity, while a field privacy budget is allocated to a data field with lower sensitivity. Finally, the final privacy budget is determined by directly summing or weighted summing all field budgets.
[0083] Based on the privacy budget and the sensitivity of all the fields, the user data is noise-added to obtain the noise-added user data.
[0084] Further, the step of adding noise to the user data based on the privacy budget and the sensitivity of all the fields to obtain the noisy user data includes:
[0085] Perform joint field analysis on the sensitivity of all the aforementioned fields to obtain the target sensitivity;
[0086] Based on the target sensitivity, noise intensity analysis is performed on the privacy budget to obtain the added noise intensity;
[0087] The user data is noise-added based on the added noise intensity to obtain the noise-added user data.
[0088] In this embodiment, since financial data fields often exhibit strong correlations (such as income and credit scores), adding noise independently to highly correlated fields could introduce excessive noise and reduce data utility. Therefore, joint field analysis can involve squaring the sensitivity of each field and taking the square root of the sum to account for the correlation between data fields in the user data, thereby obtaining the target sensitivity of the user data. Noise intensity analysis can involve calculating the ratio limit between the target sensitivity and the privacy budget to obtain the added noise intensity. The noise addition process can involve generating random noise that adds noise intensity and introducing this random noise into the user data to obtain the noisy user data.
[0089] For example, if a user's data has two fields with sensitivity, the intensity of added noise can be expressed as:
[0090]
[0091] Where, λ jonit ε represents the noise intensity; ε is the privacy budget; (Δf3) is the sensitivity of the first field of the user data; Δf4 is the sensitivity of the second field of the fuzzy data.
[0092] Step 130: Input the noisy dataset into the feature decoupling generation model to generate data and obtain the target dataset;
[0093] In this embodiment of the application, after adding noise to each user data to obtain a noisy dataset, the noisy dataset can be input into a feature structure generation model. The feature decoupling generation model decouples the features in the noisy dataset user data and generates corresponding synthetic data, thereby obtaining the target dataset.
[0094] In some embodiments, inputting the noisy dataset into the feature decoupling generation model for data generation to obtain the target dataset includes:
[0095] Obtain the dynamic correlation strength;
[0096] The noisy dataset is subjected to feature decomposition to obtain a first feature subspace and a second feature subspace, wherein the feature attribute types in the first feature subspace are different from the feature attribute types in the second feature subspace.
[0097] Heterogeneous generation is performed on the first feature subspace and the second feature subspace to obtain a plurality of first target features corresponding to the first feature subspace and a plurality of second target features corresponding to the second feature subspace;
[0098] Based on the dynamic association strength, all the first target features and the second target features are dynamically coupled to obtain the target dataset.
[0099] In this embodiment, the dynamic association strength can be the association strength value generated by the feature decoupling generation model based on the policy optimization algorithm (such as the proximal policy optimization algorithm PPO, the evolutionary policy optimization algorithm ES, and the group relative policy optimization algorithm GRPO) during the previous data generation. This dynamic association strength is used to dynamically adjust the association strength between the first feature subspace and the second feature subspace to reduce the residual sensitive information contained in the two feature subspaces.
[0100] Understandably, feature decomposition can be achieved by using an improved adversarial training technique to decompose each user's data in a noisy dataset into two independent feature subspaces. Specifically, a frequency domain transformation technique is used to transform the sensitive information in each user's data, such as removing identity features from the sensitive information, thus obtaining a feature subspace with sensitive attribute type, denoted as the first feature subspace. On the other hand, a dynamic weight allocation mechanism is used to generate a corresponding feature space for non-sensitive information in each user's data, thus obtaining a feature subspace with non-sensitive attribute type, denoted as the second feature subspace.
[0101] It should be noted that heterogeneous generation can be achieved through a dual-channel mechanism, acquiring target features from a first feature subspace and a second feature subspace respectively. Specifically, the dual channels can be a sensitive channel and a non-sensitive channel. The sensitive channel can be a simplified recurrent neural network (RNN), which obtains the first target feature by inputting the first feature subspace into the sensitive channel; while the non-sensitive channel can be a long sequence modeling model, specifically any of the Mamba model, Transformer model, etc., which obtains the second target feature by inputting the second feature subspace into the non-sensitive channel.
[0102] It is worth mentioning that dynamic coupling can be based on the dynamic association strength to couple and bind the first target feature and the second target feature, and generate the corresponding target dataset. Specifically, for any first target feature and second target feature, dynamic coupling can be performed by calculating the feature association degree between the first target feature and the second target feature. There are various evaluation metrics for this feature association degree, such as semantic similarity. Then, the first target feature with a feature association strength greater than the dynamic association strength is coupled and bound to the second target feature. The remaining first target features and second target features are treated similarly. Finally, all the coupled and bound first target features and second target features, as well as the remaining uncoupled first target features and second target features, are determined as the target dataset.
[0103] Step 140: Generate user profiles from the target dataset to obtain user profiles of the target users.
[0104] In this embodiment of the application, the target dataset can be input into a trained machine learning model (such as a logistic regression model, a random forest model, a gradient boosting tree model, or a deep learning model) to generate a user profile of the target user.
[0105] In some embodiments, the method further includes:
[0106] Obtain the compensation dataset, which is the noisy dataset or the target dataset;
[0107] Distortion detection is performed on the compensation dataset to obtain distortion detection results;
[0108] In this embodiment, after obtaining the noisy dataset or the target dataset, distortion detection can be performed on the obtained noisy dataset or the target dataset to obtain the distortion detection result. Specifically, taking the supplementary dataset as the noisy dataset as an example, distortion detection can first obtain a preset threshold and calculate the difference values between the noisy dataset and the original dataset in multiple statistical features. Specifically, it can calculate the mean, variance, correlation sparsity, distribution shape, etc. of the noisy dataset. For example, for user data such as "transaction data", it can calculate the difference between the mean of the original transaction amount and the mean of the noisy transaction amount; or, for data such as "transaction frequency", it can compare the difference between the distribution of the original transaction frequency and the distribution of the noisy transaction frequency.
[0109] Understandably, after obtaining the difference between the compensated dataset and the original dataset, the magnitude of this difference can be compared with a preset threshold. Specifically, if the difference is less than or equal to the preset threshold, a distortion detection result indicating that the compensated dataset does not have distortion can be generated. In this case, the process can return to the step of inputting the noisy dataset into the feature decoupling generation model to generate data and obtain the target dataset. Alternatively, if the difference is greater than the preset threshold, a distortion detection result indicating that the compensated dataset has distortion can be generated.
[0110] It should be noted that the compensation dataset is the content of the target dataset, which is similar to the aforementioned compensation dataset being the content of the noisy dataset. This can be easily deduced by analogy, and will not be elaborated further in this application.
[0111] If the distortion detection result indicates that the compensation dataset is distorted, then the compensation dataset is repaired to obtain the repaired compensation dataset.
[0112] Further, the step of performing data repair on the compensation dataset to obtain the repaired compensation dataset includes:
[0113] Obtain several original repair rules;
[0114] Distortion type analysis is performed on the compensation dataset to obtain the target distortion type of the compensation dataset;
[0115] Based on the target distortion type, all the original repair rules are filtered to obtain the target repair rule, which is the original repair rule that corresponds to the target distortion type among all the original repair rules.
[0116] According to the target repair rules, the compensation dataset is repaired to obtain the repaired compensation dataset.
[0117] In this embodiment, the original repair rule can be a repair rule based on Conditional Generative Adversarial Network (CGAN), a repair rule based on wavelet reconstruction technology, or a repair rule based on a normalized flow model, etc. Distortion type analysis can analyze the distortion type of the compensation dataset to obtain the target distortion type of the compensation dataset. Specifically, taking the supplementary dataset as a noisy dataset as an example, if the data features in the compensation dataset (such as the average consumption cycle of the target user) become blurred due to noise, the corresponding target distortion type is a feature distortion type, which corresponds to a repair rule based on Conditional Generative Adversarial Network (CGAN); or, if the temporal features in the compensation dataset are blurred, the corresponding target distortion type is a temporal distortion type, which corresponds to a repair rule based on wavelet reconstruction technology; or, if the statistical correlation of the user data in the compensation dataset is distorted, the corresponding target distortion type is a correlation distortion type, which corresponds to a repair rule based on a normalized flow model.
[0118] Understandably, if the target repair rule is based on Conditional Generative Adversarial Networks (CGAN), rule repair can involve using CGAN to generate more reasonable data features from the remaining data in the compensation dataset (such as consumption amount and consumption category) to obtain the repaired compensation dataset. Alternatively, if the target repair rule is based on wavelet reconstruction techniques, rule repair can involve extracting time patterns from the user data in the compensation dataset. For example, when the time features in the user data "large-scale consumption on weekend nights" are ambiguous, the low-frequency component (i.e., the component containing the consumption trend) and high-frequency component (i.e., the component containing random noise) of the user data can be obtained by decomposing the data. By enhancing the low-frequency component and suppressing the high-frequency noise, the time features in the user data "large-scale consumption on weekend nights" can be repaired to obtain the repaired compensation dataset. Or, if the target repair rule is based on a normalized flow model, rule repair can involve correcting unreasonable correlations between different data fields in each user data set of the compensation dataset. Specifically, this can be done by using a normalized flow model to detect each user data set in the compensation dataset, detecting and correcting unreasonable correlations in the data fields to obtain the repaired compensation dataset.
[0119] The following describes in detail, with reference to the accompanying drawings, a financial user profile generation system proposed according to an embodiment of this application.
[0120] Reference Figure 2 The financial user profile generation system proposed in this application includes:
[0121] The first processing unit 101 is used to obtain the feature decoupling generation model and the original dataset of the target user, wherein the original dataset includes several user data.
[0122] The second processing unit 102 is used to perform differential privacy noise addition on the original dataset to obtain a noisy dataset. The noisy dataset includes several noisy user data. The noise intensity added to each noisy user data is not exactly the same. The noise intensity added is positively correlated with the sensitivity and importance of the user data.
[0123] The third processing unit 103 is used to input the noisy dataset into the feature decoupling generation model to generate data and obtain the target dataset.
[0124] The fourth processing unit 104 is used to generate user profiles from the target dataset to obtain user profiles of the target users.
[0125] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0126] Reference Figure 3 This application also provides an electronic device, including:
[0127] At least one processor 201;
[0128] At least one memory 202 is used to store at least one program;
[0129] When the at least one program is executed by the at least one processor 201, the at least one processor 201 implements the method embodiment described above.
[0130] Similarly, it can be understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0131] This application also provides a computer-readable storage medium storing a program executable by a processor 201, which, when executed by the processor 201, is used to implement the above-described method embodiments.
[0132] Similarly, the content of the above method embodiments is applicable to the present computer-readable storage medium embodiments. The specific functions implemented by the present computer-readable storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0133] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0134] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods.
[0135] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this application are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.
[0136] Furthermore, although this application is described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding this application. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional technology for an engineer. Therefore, those skilled in the art can implement the application set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of this application, which is determined by the full scope of the appended claims and their equivalents.
[0137] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0138] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0139] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0140] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0141] In the foregoing description of this specification, the references to terms such as "one embodiment," "another embodiment," or "some embodiments," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0142] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.
[0143] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.
Claims
1. A method for generating financial user profiles, characterized in that, include: Obtain the original dataset of the feature decoupling generation model and the target user, wherein the original dataset includes several user data; Differential privacy noise is applied to the original dataset to obtain a noisy dataset. The noisy dataset includes several noisy user data sets. The noise intensity added to each noisy user data set is not exactly the same. The noise intensity added is positively correlated with the sensitivity and importance of the user data set. The noisy dataset is input into the feature decoupling generation model to generate data, thus obtaining the target dataset. User profiles are generated from the target dataset to obtain user profiles of the target users.
2. The method according to claim 1, characterized in that, The step of performing differential privacy noise addition on the original dataset to obtain a noisy dataset includes: The original dataset is classified to obtain several classification information, and each classification information corresponds to one user data; Based on all the classification information, adaptive noise is added to the original dataset to obtain the noisy dataset.
3. The method according to claim 2, characterized in that, The step of adaptively adding noise to the original dataset based on all the classification information to obtain the noisy dataset includes: Obtain the sensitivity rules corresponding to the classification information; Based on the sensitivity rules, data sensitivity analysis is performed on the user data to obtain the privacy budget and several field sensitivities of the user data, with each field sensitivity corresponding to a data field in the user data; Based on the privacy budget and the sensitivity of all the fields, the user data is noise-added to obtain the noise-added user data.
4. The method according to claim 3, characterized in that, The step of adding noise to the user data based on the privacy budget and the sensitivity of all the fields to obtain the noisy user data includes: Perform joint field analysis on the sensitivity of all the aforementioned fields to obtain the target sensitivity; Based on the target sensitivity, noise intensity analysis is performed on the privacy budget to obtain the added noise intensity; The user data is noise-added based on the added noise intensity to obtain the noise-added user data.
5. The method according to claim 1, characterized in that, The step of inputting the noisy dataset into the feature decoupling generation model to generate the target dataset includes: Obtain the dynamic correlation strength; The noisy dataset is decomposed into a first feature subspace and a second feature subspace, wherein the feature attribute types in the first feature subspace are different from the feature attribute types in the second feature subspace. Heterogeneous generation is performed on the first feature subspace and the second feature subspace to obtain a plurality of first target features corresponding to the first feature subspace and a plurality of second target features corresponding to the second feature subspace; Based on the dynamic association strength, all the first target features and the second target features are dynamically coupled to obtain the target dataset.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Obtain the compensation dataset, which is the noisy dataset or the target dataset; Distortion detection is performed on the compensation dataset to obtain distortion detection results; If the distortion detection result indicates that the compensation dataset is distorted, then the compensation dataset is repaired to obtain the repaired compensation dataset.
7. The method according to claim 6, characterized in that, The step of performing data repair on the compensation dataset to obtain the repaired compensation dataset includes: Obtain several original repair rules; Distortion type analysis is performed on the compensation dataset to obtain the target distortion type of the compensation dataset; Based on the target distortion type, all the original repair rules are filtered to obtain the target repair rule, which is the original repair rule that corresponds to the target distortion type among all the original repair rules. According to the target repair rules, the compensation dataset is repaired to obtain the repaired compensation dataset.
8. A system for generating financial user profiles, characterized in that, include: The first processing unit is used to obtain the feature decoupling generation model and the original dataset of the target user, wherein the original dataset includes several user data. The second processing unit is used to perform differential privacy noise addition on the original dataset to obtain a noisy dataset. The noisy dataset includes several noisy user data. The noise intensity added to each noisy user data is not exactly the same. The noise intensity added is positively correlated with the sensitivity and importance of the user data. The third processing unit is used to input the noisy dataset into the feature decoupling generation model to generate data and obtain the target dataset. The fourth processing unit is used to generate user profiles from the target dataset to obtain user profiles of the target users.
9. An electronic device, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor performs the method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to implement the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Domain generalization image recognition method based on causal decoupling generation model
CN114863213A
User portrait generation method and device, electronic equipment and storage medium
CN118094590A
Dynamic advertisement putting strategy generation method and system based on user behaviors
CN120298054A