Method, device, equipment and medium for building a default model based on data migration
By obtaining source and target domain data, using preset diffusion models for data migration, filtering out similar sample data and building a default model, the problem of insufficient sample data in the new credit business scenario is solved, and the accuracy and applicability of the credit model is improved.
Patent Information
- Application Number
- CN202210095000.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-26
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-01-26
AI Technical Summary
In the case of insufficient sample data in the early stages of the new credit business scenario, the existing technology cannot build a suitable credit model, resulting in poor application effect of the model.
By obtaining source domain data and target domain data, using preset diffusion models for data migration, filtering out similar sample data, and building a default model.
It realizes data migration of different credit scenarios in new credit business scenarios, builds a suitable credit model, and improves the accuracy and applicability of the model.
Smart Images

Figure CN114418749B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pre-loan credit approval, and in particular to a method, device, equipment and medium for constructing a default model based on data migration. Background Art
[0002] With the development of internet technology and financial markets, credit information has exploded, and credit business scenarios have become increasingly diverse. New credit business scenarios often lack sufficient sample data in the early stages of the business, while building credit models typically requires a large amount of modeling sample data.
[0003] Currently, credit models from established credit scenarios are typically applied directly to new credit scenarios by adjusting the rejection threshold. Credit models for these scenarios are then developed after sufficient sample data has been accumulated. However, the customer bases of established and new credit scenarios differ significantly. Directly applying adjusted credit models from established scenarios to new ones results in poor performance in the new scenarios, and the new models fail to leverage the sample data from these scenarios.
[0004] At present, due to insufficient accumulation of initial sample data for new credit business scenarios, it is impossible to build a credit model suitable for new credit business scenarios. Summary of the Invention
[0005] The main purpose of the present invention is to propose a method, device, equipment and medium for constructing a default model based on data migration, aiming to use data migration of different credit scenarios to construct a credit model suitable for new loan business scenarios.
[0006] To achieve the above-mentioned object, the present invention provides a method for constructing a default model based on data migration, which comprises the following steps:
[0007] Obtain source domain data and target domain data;
[0008] Based on the source domain data and the target domain data, data migration is performed using a preset diffusion model to determine target data;
[0009] A default model is constructed based on the target data.
[0010] Preferably, the step of acquiring source domain data and target domain data includes:
[0011] Filtering sample data of the first business scenario based on preset business requirements to obtain source domain data;
[0012] The sample data of the second business scenario is filtered based on preset business requirements to obtain target domain data.
[0013] Preferably, before the step of performing data migration based on the source domain data and the target domain data using a preset diffusion model and determining the target data, the method further includes:
[0014] Obtaining distinguishing feature variables of the source domain data and the target domain data;
[0015] Based on the source domain data, the target domain data and the distinguishing feature variables, the initial model is iteratively trained to obtain a preset diffusion model.
[0016] Preferably, the step of performing data migration based on the source domain data and the target domain data by using a preset diffusion model and determining target data includes:
[0017] Evaluate the source domain data and the target domain data using a preset diffusion model to obtain a first evaluation score corresponding to the source domain data and a second evaluation score corresponding to the target domain data;
[0018] setting a preset score according to the second evaluation score;
[0019] Filtering the source domain data corresponding to the first evaluation score according to a preset score to obtain filtered source domain data;
[0020] The filtered source domain data is determined as target data.
[0021] Preferably, the step of constructing a target default model based on the target data includes:
[0022] performing weighted processing on the target data according to the target domain data to obtain processed target data;
[0023] The processed target data is used as a sample, and positive and negative labels are added to the sample to obtain a sample with a positive label and a sample with a negative label;
[0024] A default model is constructed according to the samples with the positive label and the samples with the negative label.
[0025] Preferably, the step of performing weighted processing on the target data according to the target domain data to obtain processed target data includes:
[0026] Acquire a first sample size of the target domain data and a second sample size of the target data;
[0027] The target data is weighted according to the first sample quantity and the second sample quantity to obtain processed target data.
[0028] The present application provides a credit default analysis method, which includes:
[0029] Obtaining data to be analyzed and credit information of the data to be analyzed;
[0030] The data to be analyzed is used as a sample, and positive and negative labels are added to the sample to obtain positive and negative labels of the data to be analyzed;
[0031] Based on the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed, credit default analysis is performed using the default model, wherein the default model is constructed based on target data determined according to source domain data and target domain data.
[0032] In addition, to achieve the above-mentioned purpose, the present invention further provides a device for constructing a default model based on data migration, the device comprising:
[0033] Acquisition module, used to acquire source domain data and target domain data;
[0034] a determination module, configured to perform data migration based on the source domain data and the target domain data through a preset diffusion model to determine target data;
[0035] A construction module is used to construct a default model based on the target data.
[0036] Preferably, the acquisition module is further used for:
[0037] Filtering sample data of the first business scenario based on preset business requirements to obtain source domain data;
[0038] The sample data of the second business scenario is filtered based on preset business requirements to obtain target domain data.
[0039] Preferably, the determination module is further configured to:
[0040] Obtaining distinguishing feature variables of the source domain data and the target domain data;
[0041] Based on the source domain data, the target domain data and the distinguishing feature variables, the initial model is iteratively trained to obtain a preset diffusion model.
[0042] Preferably, the determination module is further configured to:
[0043] Evaluate the source domain data and the target domain data using a preset diffusion model to obtain a first evaluation score corresponding to the source domain data and a second evaluation score corresponding to the target domain data;
[0044] setting a preset score according to the second evaluation score;
[0045] Filtering the source domain data corresponding to the first evaluation score according to a preset score to obtain filtered source domain data;
[0046] The filtered source domain data is determined as target data.
[0047] Preferably, the building block is further configured to:
[0048] performing weighted processing on the target data according to the target domain data to obtain processed target data;
[0049] The processed target data is used as a sample, and positive and negative labels are added to the sample to obtain a sample with a positive label and a sample with a negative label;
[0050] A default model is constructed according to the samples with the positive label and the samples with the negative label.
[0051] Preferably, the building block is further configured to:
[0052] Acquire a first sample size of the target domain data and a second sample size of the target data;
[0053] The target data is weighted according to the first sample quantity and the second sample quantity to obtain processed target data.
[0054] The present application also provides a credit default analysis device, the credit default analysis device comprising:
[0055] An acquisition module is used to acquire the data to be analyzed and the credit information of the data to be analyzed;
[0056] An adding module is used to take the data to be analyzed as a sample and add positive and negative labels to the sample to obtain positive and negative labels of the data to be analyzed;
[0057] An analysis module is configured to perform credit default analysis using the default model based on the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed, wherein the default model is constructed based on target data determined based on source domain data and target domain data.
[0058] In addition, to achieve the above-mentioned purpose, the present invention also provides a device, which is a default model construction device based on data migration, and the default model construction device based on data migration includes: a memory, a processor, and a default model construction program based on data migration stored on the memory and executable on the processor, and when the default model construction program based on data migration is executed by the processor, the steps of the default model construction method based on data migration as described above are implemented.
[0059] In addition, to achieve the above-mentioned purpose, the present invention also provides a device, which is a credit default analysis device, comprising: a memory, a processor, and a credit default analysis program stored in the memory and executable on the processor, wherein the credit default analysis program, when executed by the processor, implements the steps of the credit default analysis method described above.
[0060] In addition, to achieve the above-mentioned purpose, the present invention also provides a medium, which is a computer-readable storage medium, and a default model construction program based on data migration is stored on the computer-readable storage medium. When the default model construction program based on data migration is executed by a processor, the steps of the default model construction method based on data migration as described above are implemented.
[0061] In addition, to achieve the above-mentioned purpose, the present invention also provides a medium, which is a computer-readable storage medium, on which a credit default analysis program is stored. When the credit default analysis program is executed by a processor, the steps of the credit default analysis method described above are implemented.
[0062] The present invention proposes a method, apparatus, device, and medium for constructing a default model based on data migration. These methods involve obtaining source domain data and target domain data; determining target data based on the source and target domain data using a preset diffusion model; and constructing a default model based on the target data. Thus, the present invention utilizes data migration from different credit scenarios to construct a credit model by obtaining source domain data for established business scenarios and target domain data for new business scenarios; utilizing a preset diffusion model to filter sample data from the source domain data for established business scenarios that is similar to the target domain data for new business scenarios, and determining this similar sample data as the target data; and constructing a default model using the target data. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 Schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention;
[0064] Figure 2 This is a flow chart of a first embodiment of a method for constructing a default model based on data migration according to the present invention;
[0065] Figure 3 This is a flow chart of a second embodiment of a method for constructing a default model based on data migration according to the present invention;
[0066] Figure 4 for Figure 2 A schematic diagram of a sub-process of step S20 in the method shown;
[0067] Figure 5 for Figure 2 A schematic diagram of a sub-process of step S30 in the method shown;
[0068] Figure 6 for Figure 2 Schematic diagram of the score distribution of source domain data and target domain data for the preset diffusion model in the illustrated method;
[0069] Figure 7 This is a flow chart of a first embodiment of the credit default analysis method of the present invention;
[0070] Figure 8 This is a functional module diagram of a first embodiment of a method for constructing a default model based on data migration according to the present invention;
[0071] Figure 9 This is a functional module diagram of the first embodiment of the credit default analysis method of the present invention.
[0072] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0073] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0074] like Figure 1 As shown, Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention.
[0075] The device in the embodiment of the present invention may be a mobile terminal or a server device.
[0076] like Figure 1As shown, the device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory, or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0077] Those skilled in the art will understand that Figure 1 The device structure shown in the figure does not constitute a limitation of the device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0078] like Figure 1 As shown, the memory 1005 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a default model building program based on data migration.
[0079] Among them, the operating system is a program that manages and controls the default model construction equipment and software resources based on data migration, and supports the operation of the network communication module, the user interface module, the default model construction program based on data migration, and other programs or software; the network communication module is used to manage and control the network interface 1002; the user interface module is used to manage and control the user interface 1003.
[0080] exist Figure 1 In the device for constructing a default model based on data migration shown, the device for constructing a default model based on data migration calls a page generation program stored in the memory 1005 through the processor 1001, and executes the operations in each embodiment of the following method for constructing a default model based on data migration.
[0081] Based on the above hardware structure, an embodiment of a method for constructing a default model based on data migration of the present invention is proposed.
[0082] Reference Figure 2 , Figure 2 This is a flow chart of a first embodiment of a method for constructing a default model based on data migration according to the present invention. The method includes:
[0083] Step S10, obtaining source domain data and target domain data;
[0084] Step S20: Based on the source domain data and the target domain data, data migration is performed using a preset diffusion model to determine target data;
[0085] Step S30: constructing a default model based on the target data.
[0086] This embodiment obtains source domain data of mature business scenarios and target domain data of new business scenarios; uses a preset diffusion model to filter out sample data similar to the target domain data of new business scenarios from the source domain data of mature business scenarios, and determines the similar sample data as target data; uses the target data to build a default model, thereby realizing the construction of a credit model by using data migration from different credit scenarios.
[0087] The following describes each step in detail:
[0088] Step S10: Acquire source domain data and target domain data.
[0089] In this embodiment, source and target domain data are obtained from different channels. This can be from the system's business database or from different users' clients. These different users include business personnel, clients, and third-party agency personnel. This embodiment does not limit the channels for obtaining source and target domain data. In this embodiment, source and target domain data are obtained for established business scenarios and target domain data for new business scenarios.
[0090] Furthermore, in one embodiment, step S10 includes:
[0091] Step S11, filtering sample data of the first business scenario based on preset business requirements to obtain source domain data;
[0092] In one embodiment, a preset business requirement is first set. This preset business requirement serves as a condition for filtering source and target domain data. The preset business requirement can be the business requirement of a new business scenario. For new business scenarios, there is often insufficient initial sample data accumulated for the new business scenario, and building a default model for the new business scenario requires a large number of modeling samples. Based on the business requirements of the new business scenario, sample data from established business scenarios is filtered to obtain source domain data for the established business scenario.
[0093] Step S12: filtering the sample data of the second business scenario based on preset business requirements to obtain target domain data.
[0094] In one embodiment, a preset business requirement is first set. This preset business requirement serves as a condition for filtering source and target domain data. The preset business requirement can be the business requirement of a new business scenario. For new business scenarios, there is often insufficient initial sample data accumulated for the new business scenario, and building a default model for the new business scenario requires a large number of modeling samples. Based on the business requirement of the new business scenario, the sample data for the new business scenario is filtered to obtain the target domain data for the new business scenario.
[0095] Step S20 : Based on the source domain data and the target domain data, data migration is performed through a preset diffusion model to determine target data.
[0096] In this embodiment, a preset diffusion model is first constructed. This preset diffusion model is used to filter out sample data similar to the target domain data of the new business scenario from the source domain data of the mature business scenario. The preset diffusion model can be a looklike (crowd diffusion algorithm) model. Based on the acquired source domain data of the mature business scenario and the target domain data of the new business scenario, the looklike model is used to select sample data similar to the target domain data of the new business scenario from the source domain data of the mature business scenario. This similar sample data is then matched with the target domain data of the new business scenario, and the matched similar sample data is determined as the target data.
[0097] Step S30: constructing a default model based on the target data.
[0098] In this embodiment, the target data is processed and the processed target data is fit to the target domain data of the new business scenario; the processed target data is used to build a default model suitable for the new business scenario, so as to realize the construction of the default model of the new business scenario by using the source domain data of the mature business scenario and the target domain data of the new business scenario through data migration.
[0099] This embodiment obtains source domain data of mature business scenarios and target domain data of new business scenarios; uses a preset diffusion model to filter out sample data similar to the target domain data of new business scenarios from the source domain data of mature business scenarios, and determines the similar sample data as target data; uses the target data to build a default model, thereby realizing the construction of a credit model by using data migration from different credit scenarios.
[0100] Furthermore, based on the first embodiment of the method for constructing a default model based on data migration of the present invention, a second embodiment of the method for constructing a default model based on data migration of the present invention is proposed.
[0101] The difference between the second embodiment of the method for constructing a default model based on data migration and the first embodiment of the method for constructing a default model based on data migration is that, in step S20, based on the source domain data and the target domain data, data migration is performed using a preset diffusion model. Before determining the target data, the default model is referenced. Figure 3 ,The default model construction based on data migration also includes:
[0102] Step A10, obtaining distinguishing feature variables of the source domain data and the target domain data;
[0103] Step A20: Iteratively train the initial model based on the source domain data, the target domain data, and the distinguishing feature variables to obtain a preset diffusion model.
[0104] In this embodiment, the source domain data of mature business scenarios and the distinguishing characteristic variables of target domain data of new business scenarios are obtained; and the source domain data of mature business scenarios and the target domain data of new business scenarios and the distinguishing characteristic variables are input into an initial model, the initial model is iteratively trained, and the training model with the best model evaluation index is selected as the preset diffusion model, thereby improving the accuracy of subsequent data migration of the preset diffusion model.
[0105] The following describes each step in detail:
[0106] Step A10: Obtain distinguishing feature variables of the source domain data and the target domain data.
[0107] In this embodiment, the distinguishing characteristic variables of the source domain data of the mature business scenario and the target domain data of the new business scenario are obtained. The distinguishing characteristic variables mainly select variable dimensions related to the classification of the source domain data of the mature business scenario and the target domain data of the new business scenario, such as some basic attribute information, such as age, gender, education, region, etc., as well as some dimensional information of the source domain data of the mature business scenario and the target domain data of the new business scenario related to the preset diffusion model to be built. The obtained distinguishing characteristic variables are formed into a candidate variable pool. For example, the distinguishing characteristic variable is selected as the age in the basic information of the user. The users in the source domain data of the mature business scenario may be mostly middle-aged people between the ages of 30 and 50, while the users in the target domain data of the new business scenario may be mostly young people under the age of 30.
[0108] Step A20: Iteratively train the initial model based on the source domain data, the target domain data, and the distinguishing feature variables to obtain a preset diffusion model.
[0109] In one embodiment, the acquired source domain data of mature business scenarios and target domain data of new business scenarios, as well as a variable pool of candidate distinguishing feature variables, are input into an initial model, the initial model is iteratively trained, and the training model with the best AUC (Area Under Curve, model evaluation index) is selected as the preset diffusion model.
[0110] The AUC is defined as the area under the ROC curve (Receiver Operating Characteristic Curve) and the coordinate axes. AUC values range between 0.5 and 1. AUC values closer to 1.0 indicate a strong discriminatory ability for the detection method; those closer to 0.5 indicate a weak discriminatory ability and lack of practical application. The ROC curve is a binary classification method plotted based on different cutoff values or decision thresholds, with the true positive rate as the vertical axis and the false positive rate as the horizontal axis.
[0111] In this embodiment, the distinguishing characteristic variables of the source domain data of the mature business scenario and the target domain data of the new business scenario are obtained; and the source domain data of the mature business scenario and the target domain data of the new business scenario and the distinguishing characteristic variables are input into an initial model, the initial model is iteratively trained, and the training model with the best model evaluation index is selected as the preset diffusion model, thereby improving the accuracy of the subsequent data migration of the preset diffusion model.
[0112] Furthermore, based on the first and second embodiments of the method for constructing a default model based on data migration of the present invention, a third embodiment of the method for constructing a default model based on data migration of the present invention is proposed.
[0113] The difference between the third embodiment of the method for constructing a default model based on data migration and the first and second embodiments of the method for constructing a default model based on data migration is that in step S20, based on the source domain data and the target domain data, data migration is performed through a preset diffusion model to determine the refinement of the target data, and reference is made to the following: Figure 4 , this step specifically includes:
[0114] Step S21: evaluating the source domain data and the target domain data using a preset diffusion model to obtain a first evaluation score corresponding to the source domain data and a second evaluation score corresponding to the target domain data;
[0115] Step S22, setting a preset score according to the second evaluation score;
[0116] Step S23, filtering the source domain data corresponding to the first evaluation score according to a preset score to obtain filtered source domain data;
[0117] Step S24: determining the filtered source domain data as target data.
[0118] In this embodiment, a preset diffusion model is used to distinguish the source domain data of mature business scenarios and the target domain data of new business scenarios by using a scoring method, and a first evaluation score corresponding to the source domain data of mature business scenarios and a second evaluation score corresponding to the target domain data of new business scenarios are obtained; a preset score is set according to the second evaluation score, and the source domain data corresponding to the first evaluation score is filtered according to the preset score to obtain the filtered source domain data; and the filtered source domain data is determined as the target data; thereby improving the accuracy of the target data.
[0119] The following describes each step in detail:
[0120] Step S21 : evaluating the source domain data and the target domain data through a preset diffusion model to obtain a first evaluation score corresponding to the source domain data and a second evaluation score corresponding to the target domain data.
[0121] In this embodiment, in the preset diffusion model, the credit information, basic information and other information in each source domain data and each target domain data are used; and based on this information, the source domain data of mature business scenarios and the target domain data of new business scenarios are scored to obtain a first evaluation score corresponding to the source domain data and a second evaluation score corresponding to the target domain data.
[0122] Step S22: setting a preset score according to the second evaluation score.
[0123] In this embodiment, the preset score is set according to the second evaluation score corresponding to the target domain data. For example, in the target domain data of the new business scenario, the second evaluation scores of 10 target domain data are selected to include [0.5, 0.6, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9, 1]; in the source domain data of the mature business scenario, the first evaluation scores corresponding to 10 source domain data are selected to include [0, 0.1, 0.2, 0.3, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9]. The preset score can be set between the numerical range of (0.6, 0.9). In another embodiment, the actual preset score can be set according to the actual situation.
[0124] Step S23: filtering the source domain data corresponding to the first evaluation score according to a preset score to obtain filtered source domain data.
[0125] In this embodiment, the source domain data is screened according to the preset scores. For example, in the source domain data of the mature business scenario, the first evaluation scores corresponding to 10 source domain data are selected to include [0, 0.1, 0.2, 0.3, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9]; in the target domain data of the new business scenario, the second evaluation scores of 10 target domain data are selected to include [0.5, 0.6, 0.6, 0.7, 0.7, 0.7, 0.8 , 0.8, 0.9, 1]; the preset score set according to the second evaluation score is between the numerical range of (0.6, 0.9); according to the preset score, the 10 source domain data [0, 0.1, 0.2, 0.3, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9] are filtered, and the 6 source domain data after filtering include [0.6, 0.7, 0.7, 0.8, 0.8, 0.9], and these 6 source domain data are the filtered source domain data.
[0126] Step S24: determining the filtered source domain data as target data.
[0127] In this embodiment, the source domain data is screened according to the preset scores. For example, in the source domain data of the mature business scenario, the first evaluation scores corresponding to 10 source domain data are selected to include [0, 0.1, 0.2, 0.3, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9]; in the source domain data of the new business scenario, the second evaluation scores of 10 target domain data are selected to include [0.5, 0.6, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9, 1]; according to the target domain The preset score of the second evaluation score corresponding to the data is set, and the preset score is set between the numerical range of (0.6, 0.9); according to the preset score, 10 source domain data [0, 0.1, 0.2, 0.3, 0.6, 0.7, 0.7, 0.8, 0.8, 0.9] in the mature business scenario are screened, and the 6 source domain data after screening include [0.6, 0.7, 0.7, 0.8, 0.8, 0.9]; and these 6 source domain data are determined as the target data as the filtered source domain data.
[0128] In this embodiment, a preset diffusion model is used to distinguish the source domain data of mature business scenarios and the target domain data of new business scenarios by using a scoring method, and a first evaluation score corresponding to the source domain data of mature business scenarios and a second evaluation score corresponding to the target domain data of new business scenarios are obtained; a preset score is set according to the second evaluation score, and the source domain data corresponding to the first evaluation score is filtered according to the preset score to obtain the filtered source domain data; and the filtered source domain data is determined as the target data; thereby improving the accuracy of the target data.
[0129] Furthermore, based on the first, second and third embodiments of the method for constructing a default model based on data migration of the present invention, a fourth embodiment of the method for constructing a default model based on data migration of the present invention is proposed.
[0130] The fourth embodiment of the method for constructing a default model based on data migration is different from the first, second and third embodiments of the method for constructing a default model based on data migration in that this embodiment is a refinement of step S30, constructing a default model based on the target data, referring to Figure 5 , this step specifically includes:
[0131] Step S31, performing weighted processing on the target data according to the target domain data to obtain processed target data;
[0132] Step S32: taking the processed target data as samples and adding positive and negative labels to the samples to obtain samples with positive labels and samples with negative labels;
[0133] Step S33: constructing a default model based on the samples with positive labels and the samples with negative labels.
[0134] In this embodiment, the target data is weighted according to the target domain data to obtain the processed target data; the processed data is used as samples, and a positive label or a negative label is added to each sample to obtain samples with positive labels and samples with negative labels; and a default model is constructed based on the samples with positive labels and samples with negative labels, thereby further improving the accuracy of the default model.
[0135] The following describes each step in detail:
[0136] Step S31 : performing weighted processing on the target data according to the target domain data to obtain processed target data.
[0137] In this embodiment, all target data are put into a sample set to obtain the number of sample sets; referring to the sample number of the target data, the target data is weighted by using the target domain data of the new business scenario to obtain the processed target data.
[0138] Furthermore, in one embodiment, in step S31, the step of performing weighted processing on the target data according to the target domain data to obtain the processed target data specifically includes:
[0139] Step B10: Obtain a first sample size of the target domain data and a second sample size of the target data.
[0140] In this embodiment, referring to Figure 6 , Figure 6The dashed line in the middle represents the looklike score distribution of the source domain data, the solid line represents the score distribution of the target domain data, and the gray area represents the intersection of the score releases for the source and target domain data. The horizontal axis represents the score, and the vertical axis represents the ratio of the source and target domain scores. The target domain data sample size for the new business scenario is obtained by superimposing each sample data point in the target domain data, and this sample size is used as the first sample size. The target data sample size is obtained by superimposing each sample data point in the target data, and this sample size is used as the first sample size and the second sample size.
[0141] Step B20: performing weighted processing on the target data according to the first sample quantity and the second sample quantity to obtain processed target data.
[0142] In this embodiment, referring to Figure 6 By adjusting the weights of target data with consistent scores, the target domain data distribution can be fitted. By weighting the target data according to the first sample size of the target domain data and the second sample size of the target data, the processed target data is obtained. For example, in the target domain data of the new business scenario, the evaluation scores of 10 target domain data are selected, including [0.1, 0.2, 0.2, 0.3, 0.3, 0.3, 0.4, 0.4, 0.4]; in the target data screened from the source domain data of the mature business scenario, the evaluation scores of the target data include [0.1, 0.1, 0.2, 0.2, 0.3, 0.4]; in the evaluation scores of the new business scenario, 0.1 appears once, 0.2 appears twice, 0.3 appears three times, and 0.4 appears four times. In the evaluation scores of mature business scenarios, 0.1 appears twice, 0.2 appears twice, 0.3 appears once, and 0.4 appears once. The evaluation scores of the target data of mature business scenarios are adjusted by weights to fit the evaluation score distribution of the target domain data of new business scenarios. The fitting result of the scoring scores of the target data of mature business scenarios is that 0.1 appears once, 0.2 appears twice, 0.3 appears three times, and 0.4 appears four times. That is, the fitting result of the scoring scores of the target data is the distribution of the scoring scores of the target domain data. Among them, the corresponding weights of the target data evaluation scores [0.1, 0.1, 0.2, 0.2, 0.3, 0.4] are [0.5, 0.5, 1, 1, 3, 4]. The actual weights can be set according to the actual situation.
[0143] Step S32: taking the processed target data as samples, and adding positive and negative labels to the samples to obtain samples with positive labels and samples with negative labels.
[0144] In this embodiment, by adding a positive label or a negative label to each sample data in the processed target data, samples with positive labels and samples with negative labels are obtained. For example, there are five processed data including [a, b, c, d, e]. By adding positive labels or negative labels to these five processed data, labeled samples are obtained. The positively labeled samples include [a, b, c], and the negatively labeled samples include [d, e].
[0145] Step S33: constructing a default model based on the samples with positive labels and the samples with negative labels.
[0146] In this embodiment, samples with positive labels and samples with negative labels are input into an initial model, and the initial model is iteratively trained, and a training model with the best model evaluation index is selected as the default model.
[0147] In this embodiment, the target data is weighted according to the target domain data to obtain processed target data; the processed data is used as samples, and a positive label or a negative label is added to each sample to obtain samples with positive labels and samples with negative labels; and a default model is constructed based on the target data and the positive label or negative label samples corresponding to the target data, thereby further improving the accuracy of the default model.
[0148] Reference Figure 7 , Figure 7 This is a flow chart of a first embodiment of a credit default analysis method according to the present invention. The credit default analysis method includes:
[0149] Step C10, obtaining the data to be analyzed and the credit information of the data to be analyzed;
[0150] Step C20, taking the data to be analyzed as a sample, and adding positive and negative labels to the sample to obtain positive and negative labels of the data to be analyzed;
[0151] Step C30: performing credit default analysis using the default model based on the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed, wherein the default model is constructed based on target data determined based on the source domain data and the target domain data.
[0152] In this embodiment, the data to be analyzed and the credit information corresponding to the data to be analyzed are obtained; the data to be analyzed is used as a sample, and a positive label or a negative label is added to each sample data in the data to be analyzed to obtain the positive and negative labels of the data to be analyzed; the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed are input into a trained default model, and credit default analysis is performed on the data to be analyzed through the default model, thereby improving the efficiency of credit analysis of the data to be analyzed.
[0153] The following describes each step in detail:
[0154] Step C10, obtaining the data to be analyzed and the credit information of the data to be analyzed;
[0155] In this embodiment, the data to be analyzed and its credit information are obtained from different channels, such as the business database in the system or through different user terminals. Different users include business personnel, customers, and third-party agency personnel. The credit information of the data to be analyzed includes, but is not limited to, basic user information and loan information. Basic user information includes, but is not limited to, name, age, education level, job, and region. Loan information includes, but is not limited to, information about the user's car loan and mortgage.
[0156] Step C20, taking the data to be analyzed as a sample, and adding positive and negative labels to the sample to obtain positive and negative labels of the data to be analyzed;
[0157] In this embodiment, positive or negative labels are obtained by adding positive or negative labels to each sample data in the data to be analyzed. For example, if the data to be analyzed includes [a, d, c, e, f, g], positive or negative labels are added to the data to be analyzed. The samples with positive labels are [a, d, c], and the samples with negative labels are [e, f, g], thus obtaining positive or negative labels for each sample in the data to be analyzed.
[0158] Step C30: performing credit default analysis using the default model based on the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed, wherein the default model is constructed based on target data determined based on the source domain data and the target domain data.
[0159] In this embodiment, the data to be analyzed, its credit information, and its positive and negative labels are input into a trained default model, and the default model is used to perform credit default analysis on the data to be analyzed. This is accomplished by filtering sample data similar to the target data from the source domain data and using these sample data as the target data. The target data is then weighted to fit the target data, thereby constructing a default model suitable for the target data.
[0160] In this embodiment, the data to be analyzed and the credit information corresponding to the data to be analyzed are obtained; the data to be analyzed is used as a sample, and a positive label or a negative label is added to each sample data in the data to be analyzed to obtain the positive and negative labels of the data to be analyzed; the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed are input into a trained default model, and credit default analysis is performed on the data to be analyzed through the default model, thereby improving the efficiency of credit analysis of the data to be analyzed.
[0161] The present invention also provides a device for constructing a default model based on data migration. Figure 8 The present invention provides a device for constructing a default model based on data migration, comprising:
[0162] Acquisition module D10, used to acquire source domain data and target domain data;
[0163] A determination module D20 is configured to perform data migration based on the source domain data and the target domain data using a preset diffusion model to determine target data;
[0164] A construction module D30 is used to construct a default model based on the target data.
[0165] The present invention also provides a credit default analysis device. Figure 9 The credit default analysis device of the present invention comprises:
[0166] An acquisition module E10 is used to acquire the data to be analyzed and the credit information of the data to be analyzed;
[0167] An adding module E20 is configured to take the data to be analyzed as a sample and add positive and negative labels to the sample to obtain positive and negative labels for the data to be analyzed;
[0168] The analysis module E30 is used to perform credit default analysis through the default model based on the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed, wherein the default model is constructed based on the target data determined by the source domain data and the target domain data.
[0169] In addition, the present invention also provides a medium, which is a computer-readable storage medium, on which a default model construction program based on data migration is stored. When the default model construction program based on data migration is executed by a processor, the steps of the default model construction method based on data migration as described above are implemented.
[0170] In addition, the present invention also provides a medium, which is a computer-readable storage medium and stores a credit default analysis program. When the credit default analysis program is executed by a processor, the steps of the credit default analysis method described above are implemented.
[0171] Among them, the method implemented when the data migration-based default model construction program running on the processor is executed can refer to the various embodiments of the data migration-based default model construction method of the present invention, and will not be repeated here.
[0172] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.
[0173] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0174] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0175] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A method for constructing a default model based on data migration, characterized in that: The method for constructing a default model based on data migration includes the following steps: Obtain source domain data and target domain data; Obtaining distinguishing feature variables of the source domain data and the target domain data; Iteratively training the initial model based on the source domain data, the target domain data, and the distinguishing feature variables to obtain a preset diffusion model; Evaluate the source domain data and the target domain data using a preset diffusion model to obtain a first evaluation score corresponding to the source domain data and a second evaluation score corresponding to the target domain data; setting a preset score according to the second evaluation score; Filtering the source domain data corresponding to the first evaluation score according to a preset score to obtain filtered source domain data; determining the filtered source domain data as target data; A default model is constructed based on the target data.
2. The method for constructing a default model based on data migration according to claim 1, wherein: The step of obtaining source domain data and target domain data includes: Filtering sample data of the first business scenario based on preset business requirements to obtain source domain data; The sample data of the second business scenario is filtered based on preset business requirements to obtain target domain data.
3. The method for constructing a default model based on data migration according to claim 1, wherein: The step of constructing a default model based on the target data comprises: performing weighted processing on the target data according to the target domain data to obtain processed target data; The processed target data is used as a sample, and positive and negative labels are added to the sample to obtain a sample with a positive label and a sample with a negative label; A default model is constructed according to the samples with the positive label and the samples with the negative label.
4. The method for constructing a default model based on data migration according to claim 3, wherein: The step of performing weighted processing on the target data according to the target domain data to obtain processed target data comprises: Acquire a first sample size of the target domain data and a second sample size of the target data; The target data is weighted according to the first sample quantity and the second sample quantity to obtain processed target data.
5. A credit default analysis method, characterized in that: The credit default analysis method comprises the following steps: Obtaining data to be analyzed and credit information of the data to be analyzed; The data to be analyzed is used as a sample, and positive and negative labels are added to the sample to obtain positive and negative labels of the data to be analyzed; Based on the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed, credit default analysis is performed using a default model, wherein the default model is The method for constructing a default model based on data migration according to claim 1 is constructed.
6. A device for constructing a default model based on data migration, characterized in that: The default model construction device based on data migration includes: Acquisition module, used to acquire source domain data and target domain data; a determination module configured to obtain distinguishing feature variables of the source domain data and the target domain data; iteratively train an initial model based on the source domain data, the target domain data, and the distinguishing feature variables to obtain a preset diffusion model; evaluate the source domain data and the target domain data using the preset diffusion model to obtain a first evaluation score corresponding to the source domain data and a second evaluation score corresponding to the target domain data; set a preset score based on the second evaluation score; filter the source domain data corresponding to the first evaluation score based on the preset score to obtain filtered source domain data; and determine the filtered source domain data as target data; A construction module is used to construct a default model based on the target data.
7. A credit default analysis device, characterized in that: The credit default analysis device comprises: An acquisition module is used to acquire the data to be analyzed and the credit information of the data to be analyzed; An adding module is used to take the data to be analyzed as a sample and add positive and negative labels to the sample to obtain positive and negative labels of the data to be analyzed; An analysis module is configured to perform credit default analysis using a default model based on the data to be analyzed, the credit information of the data to be analyzed, and the positive and negative labels of the data to be analyzed, wherein the default model is constructed according to the default model construction method based on data migration according to claim 1.
8. A device for constructing a default model based on data migration, characterized in that: The data migration-based default model construction device includes: a memory, a processor, and a data migration-based default model construction program stored in the memory and executable on the processor. When the data migration-based default model construction program is executed by the processor, it implements the data migration-based default model construction method as described in any one of claims 1 to 4 or the credit default analysis method as described in claim 5.
9. A medium, which is a computer-readable storage medium, characterized in that: The computer-readable storage medium stores a default model construction program based on data migration. When the default model construction program based on data migration is executed by the processor, it implements the default model construction method based on data migration as described in any one of claims 1 to 4 or the credit default analysis method as described in claim 5.
Citation Information
Patent Citations
Small, medium and micro-sized enterprise credit evaluation method based on sample transfer learning
CN113159461A