A Personal Credit Data Completion Method, Medium and Device
Through the autoencoder model and sparse processing technology, the problem of missing values of personal credit data is solved, and high-precision data completion and automated processing is achieved, which is suitable for improving the integrity and accuracy of large-scale credit data.
Patent Information
- Application Number
- CN202410260623.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-07
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-03-07
AI Technical Summary
In the prior art, the missing value of personal credit data results in low data analysis and application accuracy, and the inability to effectively process complex nonlinear data.
The data is cleaned and pre-processed by the autoencoder model, and the data is compressed and decoded through the autoencoder model composed of the encoder and the decoder. Combined with sparse processing, the sparse parameters and deep learning algorithms are dynamically adjusted, and the model is iteratively optimized to fill in the missing data.
It achieves high-precision data completion, improves the integrity and accuracy of personal credit data, is suitable for large-scale data processing, has a high degree of automation, and reduces errors caused by manual intervention.
Smart Images

Figure CN118113693B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of credit data, and more specifically, to a method, medium, and device for completing personal credit data. Background Art
[0002] With the continuous development of big data technology, personal credit data is increasingly widely used in fields such as financial risk management. However, due to various reasons, missing values often exist in personal credit data, which can affect data analysis and applications. Currently, common data completion methods include simple filling, interpolation, and statistical-based methods, etc., but these methods often have low accuracy and cannot handle complex non-linear data. Summary of the Invention
[0003] The purpose of the present invention is to provide a method, medium, and device for completing personal credit data.
[0004] The present invention aims to solve the problems existing in the prior art.
[0005] Compared with the prior art, the technical solution and its beneficial effects of the present invention are as follows:
[0006] In the first aspect disclosed by the present invention, a method for completing personal credit data is provided, including: cleaning and preprocessing the original personal credit data; training an autoencoder model using the processed data; the autoencoder model consists of an encoder and a decoder, the encoder compresses the input data into a low-dimensional latent representation, and the decoder decodes the latent representation into complete output data; sparsifying the autoencoder model to obtain the trained autoencoder model; by inputting partially missing personal credit data, performing calculation, filling, and optimization processing through the autoencoder model to obtain the supplemented complete data; evaluating the data quality and integrity of the supplemented complete data; if the evaluation result meets the standard, output the final result; if the evaluation result does not meet the standard, iteratively optimize the autoencoder model, and dynamically adjust the sparse parameters related to the sparsification process and the deep learning association algorithm of the autoencoder model according to the defects and deficiencies in the evaluation result.
[0007] As a further improvement, the autoencoder model includes an input layer, a hidden layer, and an output layer. The input layer is used for data input, the hidden layer is used for the processing of the autoencoder model, and the output layer is used for data output.
[0008] As a further improvement, the cleaning and preprocessing of the original personal credit data include, but are not limited to, data deduplication, outlier processing, and missing value identification.
[0009] As a further improvement, the sparsification process of the autoencoder model includes: calculating the average activation degree, that is, calculating the average activation and utilization degree of a random activation unit for all input unit data. The calculation formula of the average activation degree is as follows:
[0010]
[0011] where ρ j is the average activation degree, is the activation degree of the activated hidden layer unit, and x i is the input unit.
[0012] As a further improvement, the sparsification process of the autoencoder model further includes: calculating the KL divergence; the KL divergence describes the relative relationship between two distributions and describes the association information of another distribution based on the encoding of one distribution. The calculation formula of the KL divergence is as follows:
[0013]
[0014] where p and q are two distributions.
[0015] As a further improvement, if the distribution of the activated units and unactivated units in the hidden layer follows a binomial distribution, its KL divergence, that is, the information entropy between two binomial distributions with one ρ as the mean and one ρ j as the mean, then the KL divergence can be expressed as:
[0016]
[0017] where ρ is the sparsity parameter.
[0018] In the second aspect disclosed by the present invention, a storage medium is provided. The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the personal credit data completion method are implemented.
[0019] In the third aspect disclosed by the present invention, an electronic device is provided, which at least includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program on the memory, the steps of the personal credit data completion method are implemented.
[0020] The beneficial effects of the present invention are as follows:
[0021] The present invention proposes a personal credit data completion method based on deep learning, aiming to automatically identify and predict missing data, and improve the integrity and accuracy of personal credit data. This method uses statistical principles and the sparse autoencoder model in deep learning for data completion, and continuously updates and improves the model through an optimization algorithm to improve the prediction accuracy.
[0022] The new technology applied in the present invention is a comprehensive data completion model based on statistical principles and deep learning techniques. The present invention can effectively make up for the defects of existing data completion techniques in the current market, and while taking into account accuracy and efficiency, complete the improvement, integration and unification processing of personal credit-related data, and break through data barriers and statistical and analysis errors caused by data standard defects and statistical and recording errors.
[0023] The method of the present invention realizes high-precision data completion; through the deep learning method, missing data can be automatically learned and predicted, with higher precision than traditional data completion methods. The method of the present invention has strong non-linear data processing capabilities: compared with traditional single statistical methods, deep learning can automatically extract non-linear features of data and is suitable for processing complex personal credit data. The method of the present invention has strong large-scale processing capabilities: this method can be applied to large-scale personal credit data, and as the data increases, the prediction accuracy can be further improved. The method of the present invention realizes high automation: the whole process can be automated, and most of the data missing and errors caused by manual intervention are excluded as much as possible. The method of the present invention has a large space for iterative optimization. In the process of promoting and continuously optimizing the evaluation method, this evaluation method does not need to redesign the entire model structure (add or delete existing logic and architecture), and only needs to perform retraining of parameters using various algorithms covered by relevant deep learning, saving the cost of optimization iteration. Description of the Drawings
[0024] Figure 1 It is a schematic diagram of the steps of a personal credit data completion method provided by an embodiment of the present invention.
[0025] Figure 2 It is a logical schematic diagram of a personal credit data completion method provided by an embodiment of the present invention.
[0026] Figure 3 It is a schematic diagram of reconstructing and adjusting parameters of a sparse matrix autoencoder model provided by an embodiment of the present invention.
[0027] Figure 4 It is ρ provided by an embodiment of the present invention j And the associated image of the KL divergence of ρ. Detailed Embodiments
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0029] In the description of the present invention, the terms "first" and "second" are used only for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more, unless otherwise specifically defined.
[0030] Referring to Figures 1 to 4 As shown, a method for completing personal credit data includes: cleaning and preprocessing the original personal credit data; training an autoencoder model using the processed data; the autoencoder model consists of an encoder and a decoder, the encoder compresses the input data into a low-dimensional latent representation, and the decoder decodes the latent representation into complete output data; performing sparsification processing on the autoencoder model to obtain the trained autoencoder model; by inputting partially missing personal credit data, performing calculation, filling, and optimization processing through the autoencoder model to obtain the supplemented complete data; evaluating the data quality and integrity of the supplemented complete data; if the evaluation result meets the standard, output the final result; if the evaluation result does not meet the standard, iteratively optimize the autoencoder model, and dynamically adjust the sparsity parameters related to the sparsification processing and the deep learning correlation algorithm of the autoencoder model according to the defects and deficiencies in the evaluation result.
[0031] The autoencoder model includes an input layer, a hidden layer, and an output layer. The input layer is used for data input, the hidden layer is used for the processing of the autoencoder model, and the output layer is used for data output.
[0032] The cleaning and preprocessing of the original personal credit data includes, but is not limited to, data deduplication, outlier processing, and missing value identification.
[0033] The sparsification process of the autoencoder model includes: calculating the average activation degree, that is, calculating the average activation and utilization degree of a random activation unit for all input unit data. The calculation formula of the average activation degree is as follows:
[0034]
[0035] where ρ j is the average activation degree, is the activation degree of the activated hidden layer unit, and x i is the input unit.
[0036] The sparsification process of the autoencoder model also includes: KL divergence, that is, describing the relative relationship between two distributions, and describing the correlation information of another distribution based on the encoding of one distribution; the calculation formula of KL divergence is as follows:
[0037]
[0038] where p and q are two distributions.
[0039] The distributions of the activated and unactivated units in the hidden layer follow a binomial distribution, then its KL divergence, that is, describing the information entropy between two binomial distributions with one ρ as the mean and one ρ j as the mean, then the KL divergence can be expressed as:
[0040]
[0041] where ρ is the sparsity parameter.
[0042] In order to make the autoencoder sparse, then according to the specific situation of personal credit data, a very small ρ can be adjusted so that the activated units ρ in the hidden layer of the autoencoder j also become sparse.
[0043] In an ideal situation, when ρ = ρ j , the KL divergence is 0, and its distribution relationship is affected by the value of the sparsity parameter ρ, that is, the sparsity degree of the hidden layer of the autoencoder can be adjusted by the sparsity parameter ρ.
[0044] Figure 4 is the KL divergence correlation image of ρ j and ρ.
[0045] After constructing the sparse autoencoder model, incomplete personal credit data with some parts missing can be input, and through calculations, filling, and optimization by this model, complete data after supplementation can be obtained. The supplemented data will undergo evaluations of data quality and integrity, and the final result will be output after meeting the relevant standards; otherwise, iterative optimization will be carried out, and according to the defects and deficiencies in the evaluation, relevant sparse parameters and deep learning correlation algorithms will be dynamically adjusted.
[0046] Specifically, this method includes the following steps:
[0047] Step 1: Data preprocessing: Clean and preprocess the original personal credit data, including data deduplication, outlier handling, missing value identification, etc.
[0048] Step 2: Construct an autoencoder model: Use the processed data to train an autoencoder model, which can learn the internal structure and patterns of the data. The autoencoder model consists of two parts, an encoder and a decoder. The encoder compresses the input data into a low-dimensional latent representation, and the decoder decodes the latent representation into complete output data.
[0049] Step 3: Sparse processing of the autoencoder model: For the hidden layer units of the autoencoder model in Step 2, add a sparsity constraint as a regularization term for the feedforward network, so that the model only activates a small number of hidden units at any time, avoiding overfitting and improving the robustness of the model; at the same time, implicit feature selection is carried out, which has high interpretability.
[0050] Step 4: Data completion: Use the trained sparse autoencoder model to predict and complete the missing data. Specifically, the missing data is input into the autoencoder model, and through the combined action of the encoder and the decoder, the predicted complete data is output.
[0051] Step 5: Optimize and iterate the model: According to the continuous supplementation and improvement of the data source, use optimization algorithms to continuously update and improve the sparse autoencoder model to improve the prediction accuracy. Optimization algorithms such as gradient descent and stochastic gradient descent are used.
[0052] The specific implementation of the present invention is as follows: First, construct an autoencoder model, then perform sparsification processing on it, and combine it with the personal credit data pattern for targeted specialization. Then, use the processed personal credit data to train the model. The autoencoder model can utilize the complete part of the data to train the mapping model, learn the potential laws in the data, and accordingly supplement the missing and incorrect parts in the data, thereby improving the data. During the training process, an optimization algorithm is adopted to continuously adjust the parameters of the model to improve the prediction accuracy. After training, the data with a part missing is used as the input and fed into the trained sparse autoencoder model, and the predicted complete data is output through the decoder corresponding to the model. Finally, certain evaluation metrics are used to evaluate and adjust the prediction results.
[0053] The standard deep learning autoencoder model consists of three parts: an input layer (data input), a hidden layer (model processing), and an output layer (data output); the method of the autoencoder model is to perform deep learning based on other complete parts in the original personal credit data, so as to complete the missing data. In view of the particularity of personal credit data, that is, the correlation differences between various data items are relatively large and cannot be simply represented by a linear relationship, so this method uses a sparse matrix to reconstruct and adjust the parameters of the autoencoder model. The schematic diagram is as Figure 3 shown.
[0054] In Figure 3 the sparse autoencoder model, the hidden layer is jointly affected by a sparse matrix, a loss function, and the standard structure of preprocessed personal credit data: selectively activate and deactivate each unit in the hidden layer; moreover, the activation degree of each unit is not the same, but is calculated and updated according to the specific structure and pattern of personal credit data; as the data is improved and completed, the activation degree of each unit will also change accordingly. Even for a single hidden layer unit, its activation degree will change according to the change of the input unit data, which can greatly make up for the defect that the traditional completion method cannot complete non-linear and unstable correlation data well.
[0055] The sparse autoencoder model sparsifies the autoencoder model by using a sparse matrix and KL divergence.
[0056] A consumer finance company compared the user portraits obtained from different channels and found many problems such as inaccurate correspondence and missing data; according to data risk monitoring, a batch of potential high-risk customers was classified. Due to reasons such as missing and incorrect personal data, the uncertainty and potential default risk of this batch of users are very high.
[0057] The format of his personal credit data is as follows, and about 20% of the data is missing, which generates systematic risks: user ID, gender, age, province, city, issue time, expiration time, total principal and interest.
[0058] This batch of desensitized data was processed through data cleaning, completion, and reanalysis, etc. The data completion model included in this patent was used, and then operations such as desensitization, cleaning, and integrity completion analysis were performed on this batch of potential risk data; finally, the scope of high-risk customers was narrowed and confirmed, providing data indicators for precise corresponding strategies and reducing the overdue risk.
[0059]
[0060] This method does not guarantee the absolute accuracy of the prediction of personal specific information, but reasonably infers and evaluates the missing data through the model method, increasing the information density and quality of personal credit data overall, so as to increase the overall effective evaluation and utilization of personal credit data packets.
[0061] A storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the personal credit data completion method are implemented.
[0062] An electronic device includes at least a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program on the memory, the steps of the personal credit data completion method are implemented.
[0063] The above embodiments are only used to explain and illustrate the technical solutions of the present invention, rather than to limit it. Those skilled in the art should understand that any modification and equivalent replacement without departing from the spirit and scope of the present invention should fall within the protection scope of the claims of the present invention.
Claims
1. A method for completing personal credit data, characterized in that, Including: Cleaning and preprocessing the original personal credit data; Training an autoencoder model using the processed data; the autoencoder model consists of an encoder and a decoder, the encoder compresses the input data into a low-dimensional latent representation, and the decoder decodes the latent representation into complete output data; Performing sparsification processing on the autoencoder model to obtain the trained autoencoder model; by inputting personal credit data with partial content missing, performing calculation, filling, and optimization processing through the autoencoder model to obtain the supplemented complete data; evaluating the data quality and integrity of the supplemented complete data; if the evaluation result meets the standard, output the final result; if the evaluation result does not meet the standard, iteratively optimize the autoencoder model, dynamically adjust the sparse parameters related to the sparsification processing according to the defects and deficiencies in the evaluation result, and the deep learning association algorithm of the autoencoder model; The autoencoder model includes an input layer, a hidden layer, and an output layer. The input layer is used for data input, the hidden layer is used for the processing of the autoencoder model, and the output layer is used for data output; The cleaning and preprocessing of the original personal credit data includes, but is not limited to, data deduplication, outlier processing, and missing value identification; Sparsifying the autoencoder model includes: calculating the average activation degree, that is, calculating the average activation and utilization degree of a random activation unit for all input unit data. The calculation formula of the average activation degree is as follows: where ρ j is the average activation degree, is the activation degree of the activated hidden layer unit, and x i is the input unit; The sparsification processing of the autoencoder model further includes: calculating the KL divergence; the KL divergence describes the relative relationship between two distributions, and the correlation information of another distribution is described based on the encoding of one distribution; the calculation formula of the KL divergence is as follows: where p and q are two distributions; If the distribution of the active and inactive units in the hidden layer follows a binomial distribution, then its KL divergence, which describes the information entropy between two binomial distributions with means of ρ and ρj respectively, can be expressed as follows: , where ρ is the sparsity parameter. For the hidden layer units of the autoencoder model, adding a sparsity constraint as a regularization term for the feedforward network enables the model to activate only a small number of hidden units at any given time, avoiding overfitting and enhancing the model's robustness. At the same time, implicit feature selection is performed, making it highly interpretable.
2. A storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for completing personal credit data described in claim 1 above.
3. An electronic device, characterized in that, At least including a memory and a processor, a computer program is stored on the memory, and when the processor executes the computer program on the memory, it implements the steps of the method for completing personal credit data described in claim 1 above.
Citation Information
Patent Citations
Traffic data make-up method
CN104091081A