Model training method and device, electronic equipment and storage medium
By cleaning and enhancing the training data set of the risk assessment model, the impact of missing and inaccurate data on the accuracy of model assessment is resolved, and the accuracy of model assessment is improved and its effectiveness is guaranteed.
Patent Information
- Application Number
- CN202510769542.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
AI Technical Summary
In existing technologies, the evaluation accuracy of risk assessment models is easily affected by missing and inaccurate data in the training dataset, causing the model to lose trust and use value.
By performing data cleaning and data enhancement on the processed data set, including missing value processing, outlier processing and data balancing processing, the accuracy of the data set is improved, thereby improving the assessment accuracy of the risk assessment model.
Through data cleaning and enhancement technology, the quality of the training data set is improved, the generalization ability and evaluation accuracy of the risk assessment model are enhanced, and the effectiveness and reliability of the model are ensured.
Smart Images

Figure CN120670844A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training method, device, electronic device and storage medium. Background Art
[0002] With the rapid development of society, in the field of financial lending business, the use of big data and deep learning technology to build risk assessment models has become a common practice in the industry. When an object applies for a financial loan, the constructed risk assessment model will be used to conduct a risk assessment for the object.
[0003] In related technologies, the constructed risk assessment model needs to improve its assessment accuracy through model training, and the assessment accuracy of the risk assessment model depends on the training data set input during the training process. If the data in the training data set is missing or inaccurate, the assessment accuracy of the risk assessment model will be reduced, and the trained risk assessment model will be untrustworthy and lose its use value. Therefore, it is urgent to provide a model training method to solve the above technical problems. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose a model training method, device, electronic device and storage medium, which improves the data accuracy of the data set to be used, including the training data set, by performing data cleaning and data enhancement on the data set to be processed, thereby improving the assessment accuracy of the risk assessment model.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present application proposes a model training method, the method comprising:
[0006] Acquire a data set to be processed, and perform data cleaning on the data set to be processed to obtain a cleaned data set;
[0007] Performing data enhancement on the cleaned data set to obtain a data set to be used;
[0008] Dividing the dataset to be used according to a preset data ratio to obtain a training dataset and a verification dataset;
[0009] Training the risk assessment model to be trained using the training data set to obtain a trained risk assessment model;
[0010] Inputting the validation data set into the trained risk assessment model to obtain a predicted risk result corresponding to each validation data in the output validation data set;
[0011] Based on the actual risk result and the corresponding predicted risk result corresponding to each verification data, verify whether the trained risk assessment model is trained.
[0012] In some embodiments, performing data cleaning on the dataset to be processed to obtain a cleaned training dataset includes:
[0013] Performing missing value processing on each data in the data set to be processed to obtain a data set after missing value processing;
[0014] Performing outlier processing on each data in the data set after missing value processing to obtain an outlier-processed data set;
[0015] Consistency processing is performed on each data in the data set after outlier processing to obtain a cleaned training data set.
[0016] In some embodiments, each data in the dataset to be processed includes the same data field, and performing missing value processing on each data in the dataset to be processed to obtain a dataset after missing value processing includes:
[0017] Obtaining the number of missing field values from each of the data fields;
[0018] Calculating the ratio of the number of missing items corresponding to each data field to the total number of each data item in the data set to be processed, to obtain the missing item ratio of each data field;
[0019] Eliminate the missing field values corresponding to the first data field whose missing percentage is less than or equal to the first missing percentage threshold;
[0020] Fill the missing field value corresponding to the second data field whose missing ratio is greater than or equal to the second missing ratio threshold to obtain a data set after missing value processing, and the second missing ratio threshold is greater than the first missing ratio threshold.
[0021] In some embodiments, performing outlier processing on each data in the missing value processed data set to obtain the outlier processed data set includes:
[0022] Obtaining the field mean and the corresponding standard deviation of each of the data fields;
[0023] Sequentially calculating the difference between the field value of each data field in each data and the corresponding field mean value to obtain the difference value of each data field in each data;
[0024] Calculating the ratio of the difference value of each data field in each data to the standard deviation to obtain a calculation result for each data field in each data;
[0025] Determine as an abnormal data field a data field whose calculation result in each data field is greater than or equal to a preset threshold;
[0026] Process the field value corresponding to the abnormal data field.
[0027] In some embodiments, performing outlier processing on each data in the missing value processed data set to obtain the outlier processed data set includes:
[0028] Sort the field values of each data field in each data to obtain a sorted sequence of each data field;
[0029] Obtaining the first quartile and the third quartile in the sorted sequence of each of the data fields;
[0030] determining an outlier interval for each of the data fields based on the first quartile and the third quartile of the data fields;
[0031] Determine a data field in each of the data whose field value is within a corresponding abnormal value interval as an abnormal data field;
[0032] Process the field value corresponding to the abnormal data field.
[0033] In some embodiments, performing data enhancement on the cleaned dataset to obtain a dataset to be used includes:
[0034] Performing denoising on the cleaned data set to obtain a denoised data set;
[0035] The denoised data set is balanced to obtain a balanced data set.
[0036] In some embodiments, performing balancing on the denoised data set to obtain a balanced data set includes:
[0037] Determine the ratio of the difference between the positive data and the negative data in the denoising data set;
[0038] When the difference ratio is greater than or equal to a preset difference ratio, undersampling the second data with a larger number between the positive data and the negative data is performed, and part of the data is removed from the second data;
[0039] Oversampling is performed on first data with a smaller number among the positive data and the negative data to obtain generated data, and the generated data is allocated to the denoised data set to obtain a balanced data set.
[0040] To achieve the above objectives, a second aspect of an embodiment of the present application provides a model training device, comprising:
[0041] A data cleaning unit is used to obtain a data set to be processed, perform data cleaning on the data set to be processed, and obtain a cleaned data set;
[0042] A data enhancement unit, configured to perform data enhancement on the cleaned data set to obtain a data set to be used;
[0043] A division unit, configured to divide the data set to be used according to a preset data ratio to obtain a training data set and a verification data set;
[0044] A training unit, configured to train the risk assessment model to be trained using the training data set to obtain a trained risk assessment model;
[0045] A prediction unit, configured to input the verification data set into the trained risk assessment model and obtain an outputted predicted risk result corresponding to each verification data in the verification data set;
[0046] The verification unit is used to verify whether the trained risk assessment model is trained based on the actual risk result and the corresponding predicted risk result corresponding to each verification data.
[0047] In some embodiments, the data cleaning unit includes:
[0048] A missing value processing subunit is used to perform missing value processing on each data in the data set to be processed to obtain a data set after missing value processing;
[0049] an outlier processing subunit, configured to perform outlier processing on each data in the data set after missing value processing to obtain a data set after outlier processing;
[0050] The consistency processing subunit is used to perform consistency processing on each data in the data set after the outlier processing to obtain a cleaned training data set.
[0051] In some embodiments, each data in the to-be-processed data set includes the same data field, and the missing value processing subunit is configured to:
[0052] Obtaining the number of missing field values from each of the data fields;
[0053] Calculating the ratio of the number of missing items corresponding to each data field to the total number of each data item in the data set to be processed, to obtain the missing item ratio of each data field;
[0054] Eliminate the missing field values corresponding to the first data field whose missing percentage is less than or equal to the first missing percentage threshold;
[0055] Fill the missing field value corresponding to the second data field whose missing ratio is greater than or equal to the second missing ratio threshold to obtain a data set after missing value processing, and the second missing ratio threshold is greater than the first missing ratio threshold.
[0056] In some embodiments, the outlier processing subunit is configured to:
[0057] Obtaining the field mean and the corresponding standard deviation of each of the data fields;
[0058] Sequentially calculating the difference between the field value of each data field in each data and the corresponding field mean value to obtain the difference value of each data field in each data;
[0059] Calculating the ratio of the difference value of each data field in each data to the standard deviation to obtain a calculation result for each data field in each data;
[0060] Determine as an abnormal data field a data field whose calculation result in each data field is greater than or equal to a preset threshold;
[0061] Process the field value corresponding to the abnormal data field.
[0062] In some embodiments, the outlier processing subunit is configured to:
[0063] Sort the field values of each data field in each data to obtain a sorted sequence of each data field;
[0064] Obtaining the first quartile and the third quartile in the sorted sequence of each of the data fields;
[0065] determining an outlier interval for each of the data fields based on the first quartile and the third quartile of the data fields;
[0066] Determine a data field in each of the data whose field value is within a corresponding abnormal value interval as an abnormal data field;
[0067] Process the field value corresponding to the abnormal data field.
[0068] In some embodiments, the data enhancement unit includes:
[0069] a denoising subunit, configured to perform denoising on the cleaned data set to obtain a denoised data set;
[0070] The balancing subunit is configured to perform balancing on the denoised data set to obtain a balanced data set.
[0071] In some embodiments, the balancing processing subunit is configured to:
[0072] Determine the ratio of the difference between the positive data and the negative data in the denoising data set;
[0073] When the difference ratio is greater than or equal to a preset difference ratio, undersampling the second data with a larger number between the positive data and the negative data is performed, and part of the data is removed from the second data;
[0074] Oversampling is performed on first data with a smaller number among the positive data and the negative data to obtain generated data, and the generated data is allocated to the denoised data set to obtain a balanced data set.
[0075] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.
[0076] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.
[0077] The model training method, device, electronic device, and storage medium proposed in this application obtain a dataset to be processed, clean the dataset to be processed, and obtain a cleaned dataset; perform data enhancement on the cleaned dataset to obtain a dataset to be used; divide the dataset to be used according to a preset data ratio to obtain a training dataset and a verification dataset; train a risk assessment model to be trained using the training dataset to obtain a trained risk assessment model; input the verification dataset into the trained risk assessment model to obtain a predicted risk result corresponding to each verification data in the output verification dataset; and verify whether the trained risk assessment model is trained based on the actual risk result and the corresponding predicted risk result corresponding to each verification data. In this way, the data accuracy of the dataset to be used, including the training dataset, is improved by performing data cleaning and data enhancement on the dataset to be processed, thereby improving the assessment accuracy of the risk assessment model. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 This is a flow chart of the model training method provided in the embodiment of the present application;
[0079] Figure 2 yes Figure 1 Flowchart of step S101 in FIG.
[0080] Figure 3 yes Figure 2 Flowchart of step S201 in FIG.
[0081] Figure 4 yes Figure 2 Flowchart of step S202 in FIG.
[0082] Figure 5 yes Figure 3 Flowchart of step S202 in FIG.
[0083] Figure 6 yes Figure 1 Flowchart of step S103 in FIG.
[0084] Figure 7 yes Figure 6 Flowchart of step S405 in FIG.
[0085] Figure 8 Schematic diagram of the structure of the model training device provided in the embodiment of the present application;
[0086] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0087] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0088] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0089] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0090] First, let’s analyze some of the terms used in this application:
[0091] Model training: It is the core link in machine learning and deep learning. It refers to the process of using algorithms to enable the model to learn patterns and extract features from data, and ultimately have the ability to predict or make decisions about unknown data.
[0092] The core goal of model training is to enable the model to establish a mapping relationship between input data (features) and output results (labels) by learning from the training data, so that it can make reasonable predictions about new data.
[0093] The key elements of model training include:
[0094] 1. Data. Training data: The dataset used for model learning, which must include features (such as image pixels, text words, user behavior data, etc.) and labels (such as classification results and regression target values). Validation data: Data used during training to evaluate model performance and adjust hyperparameters to avoid overfitting. Test data: An independent dataset used to ultimately evaluate the model's generalization ability after training.
[0095] 2. Algorithm (model): Choose an appropriate model architecture, such as: For classification tasks, use logistic regression, random forest, or convolutional neural network (CNN). For regression tasks, use linear regression, gradient boosted tree (GBDT), or recurrent neural network (RNN). For complex tasks, use Transformer (used for natural language processing, image generation, etc.).
[0096] 3. Loss function, which measures the gap between the model's predicted value and the true value, is the goal of model optimization.
[0097] Example: Classification task: Cross-Entropy Loss. Regression task: Mean Squared Error (MSE).
[0098] 4. Optimizer, an algorithm that adjusts model parameters to minimize the loss function, such as stochastic gradient descent (SGD), Adam, RMSprop, etc.
[0099] Based on this, the embodiments of the present application provide a model training method, device, electronic device and storage medium, which aim to improve the data accuracy of the data set to be used, including the training data set, by performing data cleaning and data enhancement on the data set to be processed, thereby improving the assessment accuracy of the risk assessment model.
[0100] The model training method, device, electronic device and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the model training method in the embodiments of the present application is described.
[0101] The model training method provided in the embodiment of the present application relates to the field of business processing technology. The model training method provided in the embodiment of the present application can be applied to an electronic device as a server. The server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers. It can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the model training method, etc., but is not limited to the above forms.
[0102] The present application can be used in many general or special computer system environments or configurations. For example: server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0103] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0104] Figure 1 This is an optional flowchart of the model training method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S106.
[0105] Step S101: obtaining a data set to be processed, and performing data cleaning on the data set to be processed to obtain a cleaned data set.
[0106] The dataset to be processed is a collection of raw data obtained from the business system. For example, in a financial lending system, the dataset to be processed includes raw fields such as customer information and loan records (e.g., age, income, loan amount, and repayment status). Data cleaning is a core step in data preprocessing. It refers to the process of identifying and correcting erroneous, incomplete, inaccurate, or inconsistent records in a dataset to improve data quality. Its goal is to ensure the accuracy, completeness, consistency, and reliability of the data, providing high-quality input for subsequent analysis and modeling.
[0107] In some embodiments, performing data cleaning on the dataset to be processed to obtain a cleaned training dataset includes:
[0108] S201, performing missing value processing on each data in the data set to be processed to obtain a data set after missing value processing;
[0109] S202, performing outlier processing on each data in the data set after missing value processing to obtain an outlier-processed data set;
[0110] S203 , performing consistency processing on each data in the data set after outlier processing to obtain a cleaned training data set.
[0111] Among them, see Figure 2 , Figure 2 yes Figure 1 Flowchart of step S101 in the process. A missing value is a missing value in a data field of a data point in the dataset to be processed. An outlier is a data point whose value in a data field significantly deviates from the overall distribution of the dataset (e.g., a customer's age is 200 years old). Consistency processing unifies the data format and eliminates contradictory or redundant data. By performing missing value processing, outlier processing, and consistency processing on the data, the dataset to be processed is cleaned, resulting in a cleaned training dataset.
[0112] Specifically, missing value processing, outlier processing and consistency processing are processed in sequence, that is, first, missing value processing is performed on each data in the data set to be processed to obtain a data set after missing value processing; then, outlier processing is performed on each data in the data set after missing value processing to obtain a data set after outlier processing; finally, consistency processing is performed on each data in the data set after outlier processing to obtain a cleaned training data set.
[0113] Consistency processing can include format standardization, category merging, and deduplication. Format standardization, for example, unifies date formats (e.g., "2025-05-15") and currency units (e.g., converting "$100" to the value 100). Category merging, for example, unifies different representations of the same category (e.g., unifying "Male / Female" and "M / F" into "M / F"). Deduplication, for example, deletes exact duplicate records or merges duplicate data based on primary keys.
[0114] This approach helps avoid analytical bias caused by missing data through missing values. For example, filling missing values in customer revenue prevents model learning errors, corrects unreasonable values, ensures that data truly reflects business characteristics, and provides a reliable data foundation for model training. Missing value processing ensures that no key information is missing from data records through filling or deletion. Deduplication in consistency processing avoids interference from duplicate data, ensuring that every record in the dataset is valuable, ensuring that the dataset fully covers business scenarios, and improving data availability.
[0115] In some embodiments, each data in the dataset to be processed includes the same data field, and performing missing value processing on each data in the dataset to be processed to obtain a dataset after missing value processing includes:
[0116] S301, obtaining the number of missing field values from each of the data fields;
[0117] S302, calculating the ratio of the number of missing items corresponding to each data field to the total number of each data item in the data set to be processed, to obtain a missing item ratio for each data field;
[0118] S303: Eliminate missing field values corresponding to the first data field whose missing percentage is less than or equal to a first missing percentage threshold;
[0119] S304: Fill in the missing field value corresponding to the second data field whose missing ratio is greater than or equal to a second missing ratio threshold to obtain a data set after missing value processing, where the second missing ratio threshold is greater than the first missing ratio threshold.
[0120] Among them, see Figure 3 , Figure 3 yes Figure 2 Flowchart of step S201 in
[15] . Each data point in the dataset to be processed includes the same data fields. For example, each data point includes the three data fields of "age," "income," and "loan amount." Missing field values are unfilled or invalid values in a data field. Data analysis tools (such as Python's pandas library) can be used to traverse each data field and count the number of missing values in each field.
[0121] For example, using the data.isnull().sum() function, we can calculate the number of missing values column by column to get the number of missing values for each field (e.g., 5 missing values for the "age" field and 20 missing values for the "income" field). If there are 1000 records in total, and 5 missing values for the "age" field, the missing percentage is 5 / 1000 = 0.5%; if 20 missing values for the "income" field, the missing percentage is 20 / 1000 = 2%.
[0122] Specifically, the missing value ratio is the ratio of the number of missing values in each data field to the total number of records, reflecting the severity of missing data. There are two different treatment options for missing values. The first is to directly delete the missing values for the first data field whose missing value ratio is less than or equal to the first missing value ratio threshold. Since the missing value ratio is low, the missing values can be directly deleted. The second is to fill in the missing values for the second data field whose missing value ratio is greater than or equal to the second missing value ratio threshold to make up for the missing parts. The second missing value ratio threshold is greater than the first missing value ratio threshold.
[0123] For example, the first missing ratio threshold is 5%, the second missing ratio threshold is 20%, the missing ratio of the data field "age" is 0.5%, the missing ratio of the data field "loan amount" is 1%, and the missing ratio of the data field "income" is 20%. Then the data fields of "age" and "loan amount" are the first data fields, and the missing values in the data fields of "age" and "loan amount" can be directly eliminated; the data field of "income" is the second data field, and the missing values of the data field of "income" are filled.
[0124] Specifically, for missing values that are numerical, the filling method can be through the use of mean, median or through prediction by machine learning models (such as random forest); for missing values that are categorical, the filling method can be through the use of mode (the category with the highest frequency) or the most likely category (such as prediction by Bayesian algorithm), and finally a data set with missing values processed is obtained.
[0125] This allows us to directly delete a small number of missing records for fields with low missingness rates, preventing the discarding of entire columns of features due to missing individual data points. For fields with high missingness rates, we fill them in so they can continue to be used for model training, preventing them from being discarded due to high missingness rates. Thresholding prevents fields with low missingness rates from being entirely deleted due to individual missing values, ensuring the integrity of the model input dimensions.
[0126] In some embodiments, performing outlier processing on each data in the missing value processed data set to obtain the outlier processed data set includes:
[0127] S401, obtaining the field mean and corresponding standard deviation of each data field;
[0128] S402, sequentially calculating the difference between the field value of each data field in each data and the corresponding field mean value, to obtain the difference value of each data field in each data;
[0129] S403, calculating the ratio of the difference value of each data field in each data to the standard deviation, and obtaining a calculation result for each data field in each data;
[0130] S404, determining a data field in each of the data whose calculation result is greater than or equal to a preset threshold as an abnormal data field;
[0131] S405: Process the field value corresponding to the abnormal data field.
[0132] Among them, see Figure 4 , Figure 4 yes Figure 2 Flowchart of step S202 in . The field mean is the arithmetic mean of all non-missing values in the data field, reflecting the central tendency of the data; the standard deviation is a measure of the degree of dispersion of all non-missing values in the data field, and the calculation formula is the square root of the variance. The difference value is the difference between the field value of a single data and the mean of the field, that is, the difference value = field value - field mean; the calculation result is the ratio of the difference value to the standard deviation (Z-score), that is, the calculation result = standard deviation / difference value, which is used to measure the degree to which the field value deviates from the mean. The preset threshold is the critical value for judging anomalies (such as +3 and -3), and the field value whose Z-score exceeds the threshold is considered abnormal. The abnormal data field is a data field whose calculation result ≥ the preset threshold, and its field value is an abnormal value. The processing of abnormal data fields is achieved by correcting, replacing or deleting the abnormal values.
[0133] For example, there is a loan dataset containing 1,000 customer data, where the value distribution of the "loan amount" field is: [5000, 8000, 10000, ..., 200000, 300000, 500000], the field mean is 50000, and the standard deviation is 30000. For each customer's loan amount x i , calculate the difference from the mean value x i -50000, divide the difference value by the standard deviation to get the Z-score, Z-score = (x i -50000) / 30000. The preset threshold is set to 3. For the current data field, data fields with Z-score ≥ 3 are screened from each data and determined to be abnormal data fields, so as to be processed.
[0134] Specifically, the abnormal data fields can be processed in the following ways: 1. If there is an obvious logical contradiction in the abnormal value, it can be corrected through business rules or external data. For example, in a data set, the age field of a patient in a certain hospital contains values such as -5 and 200. Processing: Correct -5 to a reasonable value (such as confirming that the actual age is 5 through supplementary medical records). Correct 200 to the logical maximum value (such as correcting it to 120 based on demographic data). 2. When the abnormal value cannot be deleted or corrected and the data integrity needs to be retained, refer to the missing value filling method. That is, for numerical types: fill with the mean, median or model prediction value (such as random forest). For categorical types: fill with the mode or the most likely category.
[0135] By calculating the field mean, standard deviation, and Z-score, we define outliers as data that deviates from the mean (default threshold), eliminating the subjectivity of manually set thresholds. Outliers with obvious logical errors (such as age -5 years) are directly corrected to ensure data authenticity. For true extreme values (such as large corporate loans), winsorization or logarithmic transformation is used to preserve business characteristics while reducing statistical noise.
[0136] In some embodiments, performing outlier processing on each data in the missing value processed data set to obtain the outlier processed data set includes:
[0137] S501, sorting the field values of each data field in each data to obtain a sorting sequence of each data field;
[0138] S502, obtaining the first quartile and the third quartile in the sorting sequence of each of the data fields;
[0139] S503, determining an outlier interval of each of the data fields based on the first quartile and the third quartile of the data field;
[0140] S504, determining each data field in the data whose field value is in the corresponding abnormal value interval as an abnormal data field;
[0141] S505: Process the field value corresponding to the abnormal data field.
[0142] Among them, see Figure 5 , Figure 5 yes Figure 2Another flowchart of step S202 in . The sorting sequence is a sequence formed by arranging the field values of the data field in ascending order; the first quartile (Q1) is the value at the 25th percentile position in the sorting sequence, i.e., the lower quartile, marking the dividing point of the first 25% of the data; the third quartile (Q3) is the value at the 75th percentile position in the sorting sequence, i.e., the upper quartile, marking the dividing point of the last 25% of the data; the outlier interval is the outlier judgment range calculated by Q1, Q3, and the interquartile range (IQR), including the lower limit and the upper limit; an abnormal data field is a data field whose field value falls outside the outlier interval, and its field value is an outlier.
[0143] For example, there is a "consumption amount" field containing 10 customer data, and the values are as follows (unit: yuan): [500, 800, 1000, 1200, 1500, 2000, 2500, 3000, 5000, 10000]. The following are the specific steps to use the IQR method to deal with outliers:
[0144] Sort "consumption amount" from small to large and get the sorting sequence: [500, 800, 1000, 1200, 1500, 2000, 2500, 3000, 5000, 10000]; data length: n = 10, Q1 position: (n+1) / 4 = 2.75, that is, 25% between the second value (800) and the third value (1000), Q1 = 800 + 0.75 × (1000-800) = 800 + 150 = 950; Q3 position: 3(n+1) / 4 = 8.25, that is, 25% between the eighth value (3000) and the ninth value (5000): Q3 = 3000 + 0.25 × (5000-3000) = 3000 + 500 = 3500.
[0145] Interquartile range (IQR): IQR = Q3 - Q1 = 3500 - 950 = 2550. The lower limit of the outlier interval is: Q1 - 1.5 × IQR = 950 - 1.5 × 2550 = 950 - 3825 = -2875.
[0146] The upper limit is: Q3 + 1.5 × IQR = 3500 + 1.5 × 2550 = 3500 + 3825 = 7325, so the outlier range is [-2875, 7325]. Values outside this range are considered outliers. Traverse the original data and determine whether each field value in the "Consumption Amount" data field is outside the outlier range: 5000: within the range (5000 < 7325), normal; 10000: outside the upper limit (10000 > 7325), determined to be an outlier data field, and then process the corresponding field value of the outlier data field.
[0147] For the abnormal data field, its corresponding abnormal field value is determined, and the value closest to the abnormal field value in the abnormal value interval is determined as the corrected field value of the abnormal data field.
[0148] For example, 10000 is closest to 7325 in [-2875,7325], so 7325 is used as the correction field value.
[0149] Therefore, compared to the Z-score method, the IQR method is based on quantile calculations and does not require data to follow a normal distribution. It is more friendly to skewed data (such as the right-skewed distribution of consumption amounts) or data with long tails. Q1 and Q3 are insensitive to extreme values, and even if outliers are present, the quantile calculation will not be significantly affected, ensuring the stability of the outlier interval.
[0150] Step S102 : performing data enhancement on the cleaned data set to obtain a data set to be used.
[0151] Data augmentation is a common technique in machine learning and deep learning. It involves expanding the size, diversity, and quality of a dataset by performing controllable transformations on existing data or generating new samples without directly increasing the number of original data samples, thereby improving the generalization, robustness, and training effectiveness of the model. By performing data augmentation on the cleaned dataset, the dataset to be used is obtained.
[0152] In some embodiments, performing data enhancement on the cleaned dataset to obtain a dataset to be used includes:
[0153] S601, performing denoising processing on the cleaned data set to obtain a denoised data set;
[0154] S602: Perform balancing processing on the denoised data set to obtain a balanced data set.
[0155] Among them, see Figure 6 , Figure 6 yes Figure 1 Flowchart of step S102 in FIG. Since data entry may be subject to noise interference, the cleaned dataset must first be denoised to obtain a denoised dataset. Then, due to the imbalance of positive and negative data (for example, 90% positive data and only 10% negative data), the model will be biased towards the majority class, so balancing is required to obtain a balanced dataset with balanced positive and negative data.
[0156] Specifically, the denoising process includes: 1. Smoothing process: Using smoothing techniques such as moving average method, weighted average method, etc. to reduce the random noise of data. 2. Denoising algorithm: Applying data denoising algorithms, such as wavelet transform, etc., to clean the noise in the data.
[0157] In some embodiments, the balanced processing of the denoised dataset to obtain a balanced dataset includes:
[0158] S701, determining the proportion of the difference in the number of positive data and negative data in the denoised dataset;
[0159] S702, when the proportion of the difference in the number is greater than or equal to the preset proportion of the difference in the number, performing undersampling on the second data with a larger number among the positive data and the negative data, and removing some data from the second data;
[0160] S703, performing oversampling on the first data with a smaller number among the positive data and the negative data to obtain generated data, and allocating the generated data to the denoised dataset to obtain a balanced dataset.
[0161] Among them, please refer to Figure 7 , Figure 7 is Figure 6 the flowchart of step S602 in
[0162] Specifically, if the difference quantity ratio is greater than or equal to the preset difference quantity ratio, it means that the quantity gap between the positive data and the negative data is large, and balancing processing is required. The specific method of balancing processing is to undersample the second data with a larger quantity in the positive data and the negative data to eliminate part of the data from the second data (that is, reduce the second data with a larger quantity); oversample the first data with a smaller quantity in the positive data and the negative data to obtain generated data (that is, increase the first data with a larger quantity), and distribute the generated data to the data set after the denoising processing to obtain the data set after balanced processing.
[0163] In this way, the degree of data imbalance can be accurately identified through the proportion of the number of differences to avoid subjective misjudgment. The dominance of the majority class can be reduced through undersampling, and the representativeness of the minority class can be enhanced through oversampling to alleviate the model's bias towards the majority class. The processing logic can be dynamically triggered through preset thresholds to adapt to the data balance needs of different businesses.
[0164] Furthermore, data can be filtered to identify features that contribute most to the prediction objective (such as risk assessment results) while eliminating redundant or irrelevant features. For example, in customer credit risk assessment, features such as income stability and credit history are highly correlated with risk, while features such as the customer's name and preferences are less relevant. Through methods such as correlation coefficient calculation, chi-square tests, and recursive feature elimination, feature importance can be quantified, reducing model computational effort and overfitting risk, thereby improving training efficiency and generalization capabilities.
[0165] Convert non-numeric data (such as text and categories) into numerical form for easier model processing. Common methods include one-hot encoding, which converts each category into a binary vector. For example, in the "gender" field, "male" is encoded as [1,0] and "female" is encoded as [0,1]. Another method is label encoding, which assigns a unique number to each category, such as encoding "low," "medium," and "high" risk as 0, 1, and 2, respectively. This encoding allows the model to understand and learn categorical information, avoiding training errors caused by data formatting issues.
[0166] Step S103 : dividing the dataset to be used according to a preset data ratio to obtain a training dataset and a verification dataset.
[0167] The dataset to be used is divided according to a preset data ratio (for example, 7:3) to obtain a training dataset and a validation dataset, and the data distribution of the two must be kept consistent.
[0168] Specifically, the partitioning methods can be: 1. Stratified sampling: Stratify the data by the distribution ratio of the target variable (default), ensuring that the proportion of defaulting customers in the training and validation sets is consistent with the original data. 2. Random partitioning: Split the data using random indices to avoid order bias that affects model generalization.
[0169] Step S104: training the risk assessment model to be trained using the training data set to obtain a trained risk assessment model.
[0170] The risk assessment model to be trained can be constructed using the Sequential model to construct an LSTM model. First, add an LSTM layer with a certain number of units, for example, 50 units, and possibly subsequent LSTM layers. The input shape is determined to be (1,7) based on the training set data. Next, add a Dropout layer to randomly drop 20% of the neuron connections to prevent model overfitting. Add a second LSTM layer, also with 50 units, and then add another Dropout layer. Finally, add an output layer using a sigmoid activation function with an output dimension of 1 to predict whether the customer has defaulted (the output value is a probability between 0 and 1). Select the Adam optimizer to adjust the model weights to minimize the loss function. Select binary_crossentropy as the loss function because this is a binary classification problem. At the same time, use accuracy as one of the indicators to evaluate model performance.
[0171] Specifically, setting epochs = 10 means that the model will learn the training set data 10 times. Batch size batch_size = 32 means that the model will process 32 samples at a time during each training.
[0172] During each training epoch, the model reads the training data in batches. For example, in the first batch of the first training epoch, the model extracts samples 1-32, extracts their corresponding labels, and then feeds this data into the model for forward propagation, calculating the predicted results. Next, the loss is calculated using binary_crossentropy based on the predicted results and the true labels. Backpropagation is then used to calculate the gradient and update the model weights. After training a batch, the model continues with the next batch until all 3840 samples in the training set have been traversed. This constitutes one training epoch. After each epoch, the model is evaluated on the test set, and the loss and accuracy on the test set are calculated and recorded. During training, the loss and accuracy of the training and test sets can be observed to change as the number of training epochs increases. For example, in the first few epochs, the loss may decrease and the accuracy may increase for both the training and test sets, indicating that the model is continuously learning and optimizing. However, if the number of training rounds is too large, the loss value of the training set may continue to decrease, while the loss value of the test set begins to increase and the accuracy decreases. This is a manifestation of overfitting.
[0173] Step S105 , inputting the verification data set into the trained risk assessment model to obtain the predicted risk result corresponding to each verification data in the output verification data set.
[0174] After training, the model is evaluated using the validation dataset. The model uses the trained weights to predict the test set data, then calculates the loss and accuracy between the predicted results and the true labels. For example, if the model's accuracy on the test set is 0.85, this means the model correctly predicts whether 85% of the test set customers will default.
[0175] Step S106: Verify whether the trained risk assessment model is completed based on the actual risk result and the corresponding predicted risk result corresponding to each verification data.
[0176] Among them, the loss value (such as binary cross entropy), accuracy, recall rate, F1 score, etc. on the validation set are calculated to evaluate the overall performance of the model. The confusion matrix is drawn to analyze the classification effect of the model on defaulting customers (minority class) and non-defaulting customers (majority class). Compare the loss values of the training set and the validation set. If the loss of the validation set is significantly higher than that of the training set and the accuracy no longer increases, it indicates that the model is overfitting, and it is necessary to adjust the hyperparameters (such as increasing the Dropout ratio, reducing the number of LSTM layers) or terminate the training early. When the loss of the validation set decreases steadily, the accuracy reaches the preset threshold (such as 85%) and there is no obvious overfitting, the model training is determined to be complete; otherwise, return to adjust the model parameters or re-enhance the data until the requirements are met.
[0177] A threshold is set based on the predicted probability output by the model, assuming it is set at 0.5. If a customer's predicted default probability is greater than or equal to 0.5, they are classified as high-risk; if it is less than 0.5, they are classified as low-risk. A risk report is generated based on the classification results, listing high-risk customers and assigning corresponding risk scores. For example, for high-risk customers with higher risk scores, it is recommended to strengthen the review, further investigate their financial situation, or increase the loan interest rate to compensate for potential risks. For low-risk customers with lower risk scores, it is possible to consider simplifying the approval process and offering more favorable loan terms to attract high-quality customers.
[0178] Finally, once the trained risk assessment model is obtained, we can use the model's predictions to assess the risk of lending businesses and generate corresponding risk reports. We set thresholds based on the model's predicted probabilities, classify customers as high-risk or low-risk, and generate corresponding risk reports.
[0179] In the steps S101 to S106 shown in the embodiment of the present application, by obtaining a data set to be processed, performing data cleaning on the data set to be processed to obtain a cleaned data set; performing data enhancement on the cleaned data set to obtain a data set to be used; dividing the data set to be used according to a preset data ratio to obtain a training data set and a verification data set; training the risk assessment model to be trained with the training data set to obtain a trained risk assessment model; inputting the verification data set into the trained risk assessment model to obtain a predicted risk result corresponding to each verification data in the output verification data set; and verifying whether the trained risk assessment model is trained based on the actual risk result and the corresponding predicted risk result corresponding to each verification data. In this way, the data accuracy of the data set to be used including the training data set is improved by performing data cleaning and data enhancement on the data set to be processed, thereby improving the assessment accuracy of the risk assessment model.
[0180] See also Figure 8 The present application also provides a model training device that can implement the above-mentioned model training method. The device includes:
[0181] A data cleaning unit is used to obtain a data set to be processed, perform data cleaning on the data set to be processed, and obtain a cleaned data set;
[0182] A data enhancement unit, configured to perform data enhancement on the cleaned data set to obtain a data set to be used;
[0183] A division unit, configured to divide the data set to be used according to a preset data ratio to obtain a training data set and a verification data set;
[0184] A training unit, configured to train the risk assessment model to be trained using the training data set to obtain a trained risk assessment model;
[0185] A prediction unit, configured to input the verification data set into the trained risk assessment model and obtain an outputted predicted risk result corresponding to each verification data in the verification data set;
[0186] The verification unit is used to verify whether the trained risk assessment model is trained based on the actual risk result and the corresponding predicted risk result corresponding to each verification data.
[0187] In some embodiments, the data cleaning unit includes:
[0188] A missing value processing subunit is used to perform missing value processing on each data in the data set to be processed to obtain a data set after missing value processing;
[0189] an outlier processing subunit, configured to perform outlier processing on each data in the data set after missing value processing to obtain a data set after outlier processing;
[0190] The consistency processing subunit is used to perform consistency processing on each data in the data set after the outlier processing to obtain a cleaned training data set.
[0191] In some embodiments, each data in the to-be-processed data set includes the same data field, and the missing value processing subunit is configured to:
[0192] Obtaining the number of missing field values from each of the data fields;
[0193] Calculating the ratio of the number of missing items corresponding to each data field to the total number of each data item in the data set to be processed, to obtain the missing item ratio of each data field;
[0194] Eliminate the missing field values corresponding to the first data field whose missing percentage is less than or equal to the first missing percentage threshold;
[0195] Fill the missing field value corresponding to the second data field whose missing ratio is greater than or equal to the second missing ratio threshold to obtain a data set after missing value processing, and the second missing ratio threshold is greater than the first missing ratio threshold.
[0196] In some embodiments, the outlier processing subunit is configured to:
[0197] Obtaining the field mean and the corresponding standard deviation of each of the data fields;
[0198] Sequentially calculating the difference between the field value of each data field in each data and the corresponding field mean value to obtain the difference value of each data field in each data;
[0199] Calculating the ratio of the difference value of each data field in each data to the standard deviation to obtain a calculation result for each data field in each data;
[0200] Determine as an abnormal data field a data field whose calculation result in each data field is greater than or equal to a preset threshold;
[0201] Process the field value corresponding to the abnormal data field.
[0202] In some embodiments, the outlier processing subunit is configured to:
[0203] Sort the field values of each data field in each data to obtain a sorted sequence of each data field;
[0204] Obtaining the first quartile and the third quartile in the sorted sequence of each of the data fields;
[0205] determining an outlier interval for each of the data fields based on the first quartile and the third quartile of the data fields;
[0206] Determine a data field in each of the data whose field value is within a corresponding abnormal value interval as an abnormal data field;
[0207] Process the field value corresponding to the abnormal data field.
[0208] In some embodiments, the data enhancement unit includes:
[0209] a denoising subunit, configured to perform denoising on the cleaned data set to obtain a denoised data set;
[0210] The balancing subunit is configured to perform balancing on the denoised data set to obtain a balanced data set.
[0211] In some embodiments, the balancing processing subunit is configured to:
[0212] Determine the ratio of the difference between the positive data and the negative data in the denoising data set;
[0213] When the difference ratio is greater than or equal to a preset difference ratio, undersampling the second data with a larger number between the positive data and the negative data is performed, and part of the data is removed from the second data;
[0214] Oversampling is performed on first data with a smaller number among the positive data and the negative data to obtain generated data, and the generated data is allocated to the denoised data set to obtain a balanced data set.
[0215] The model training device provided in the embodiment of the present application obtains a data set to be processed through a data cleaning unit, performs data cleaning on the data set to be processed, and obtains a cleaned data set; a data enhancement unit performs data enhancement on the cleaned data set to obtain a data set to be used; a division unit divides the data set to be used according to a preset data ratio to obtain a training data set and a verification data set; the training unit trains the risk assessment model to be trained through the training data set to obtain a trained risk assessment model; the prediction unit inputs the verification data set into the trained risk assessment model to obtain a predicted risk result corresponding to each verification data in the output verification data set; the verification unit verifies whether the trained risk assessment model is trained based on the actual risk result corresponding to each verification data and the corresponding predicted risk result. In this way, the data accuracy of the data set to be used including the training data set is improved by performing data cleaning and data enhancement on the data set to be processed, thereby improving the assessment accuracy of the risk assessment model.
[0216] The specific implementation of the model training device is basically the same as the specific embodiment of the above-mentioned model training method, and will not be repeated here.
[0217] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-described model training method when executing the computer program. The electronic device can be any smart terminal, such as a tablet computer or an in-vehicle computer.
[0218] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:
[0219] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0220] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the model training method of the embodiment of the present application;
[0221] Input / output interface 903, used to implement information input and output;
[0222] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0223] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );
[0224] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .
[0225] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned model training method is implemented.
[0226] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0227] The model training method, model training device, electronic device and storage medium provided in the embodiment of the present application obtain a data set to be processed, perform data cleaning on the data set to be processed to obtain a cleaned data set; perform data enhancement on the cleaned data set to obtain a data set to be used; divide the data set to be used according to a preset data ratio to obtain a training data set and a verification data set; train the risk assessment model to be trained with the training data set to obtain a trained risk assessment model; input the verification data set into the trained risk assessment model to obtain a predicted risk result corresponding to each verification data in the output verification data set; based on the actual risk result corresponding to each verification data and the corresponding predicted risk result, verify whether the trained risk assessment model is trained. In this way, the data accuracy of the data set to be used including the training data set is improved by performing data cleaning and data enhancement on the data set to be processed, thereby improving the assessment accuracy of the risk assessment model.
[0228] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0229] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0230] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0231] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0232] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0233] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0234] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0235] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0236] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0237] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0238] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A model training method, characterized in that: include: Acquire a data set to be processed, and perform data cleaning on the data set to be processed to obtain a cleaned data set; Performing data enhancement on the cleaned data set to obtain a data set to be used; Dividing the dataset to be used according to a preset data ratio to obtain a training dataset and a verification dataset; Training the risk assessment model to be trained using the training data set to obtain a trained risk assessment model; Inputting the validation data set into the trained risk assessment model to obtain a predicted risk result corresponding to each validation data in the output validation data set; Based on the actual risk result and the corresponding predicted risk result corresponding to each verification data, verify whether the trained risk assessment model is trained.
2. The model training method according to claim 1, characterized in that The step of performing data cleaning on the data set to be processed to obtain a cleaned training data set includes: Performing missing value processing on each data in the data set to be processed to obtain a data set after missing value processing; Performing outlier processing on each data in the data set after missing value processing to obtain an outlier-processed data set; Consistency processing is performed on each data in the data set after outlier processing to obtain a cleaned training data set.
3. The model training method according to claim 2, characterized in that Each data in the data set to be processed includes the same data field, and performing missing value processing on each data in the data set to be processed to obtain a data set after missing value processing includes: Obtaining the number of missing field values from each of the data fields; Calculating the ratio of the number of missing items corresponding to each data field to the total number of each data item in the data set to be processed, to obtain the missing item ratio of each data field; Eliminate the missing field values corresponding to the first data field whose missing percentage is less than or equal to the first missing percentage threshold; Fill the missing field value corresponding to the second data field whose missing ratio is greater than or equal to the second missing ratio threshold to obtain a data set after missing value processing, and the second missing ratio threshold is greater than the first missing ratio threshold.
4. The model training method according to claim 3, characterized in that The step of performing outlier processing on each data in the data set after missing value processing to obtain the data set after outlier processing includes: Obtaining the field mean and the corresponding standard deviation of each of the data fields; Sequentially calculating the difference between the field value of each data field in each data and the corresponding field mean value to obtain the difference value of each data field in each data; Calculating the ratio of the difference value of each data field in each data to the standard deviation to obtain a calculation result for each data field in each data; Determine as an abnormal data field a data field whose calculation result in each data field is greater than or equal to a preset threshold; Process the field value corresponding to the abnormal data field.
5. The model training method according to claim 3, characterized in that: The step of performing outlier processing on each data in the data set after missing value processing to obtain the data set after outlier processing includes: Sort the field values of each data field in each data to obtain a sorted sequence of each data field; Obtaining the first quartile and the third quartile in the sorted sequence of each of the data fields; determining an outlier interval for each of the data fields based on the first quartile and the third quartile of the data fields; Determine a data field in each of the data whose field value is within a corresponding abnormal value interval as an abnormal data field; Process the field value corresponding to the abnormal data field.
6. The model training method according to claim 1, characterized in that The step of performing data enhancement on the cleaned data set to obtain a data set to be used includes: Performing denoising on the cleaned data set to obtain a denoised data set; The denoised data set is balanced to obtain a balanced data set.
7. The model training method according to claim 6, characterized in that The performing balancing processing on the denoised data set to obtain a balanced data set includes: Determine the ratio of the difference between the positive data and the negative data in the denoising data set; When the difference ratio is greater than or equal to a preset difference ratio, undersampling the second data with a larger number between the positive data and the negative data is performed, and part of the data is removed from the second data; Oversampling is performed on first data with a smaller number among the positive data and the negative data to obtain generated data, and the generated data is allocated to the denoised data set to obtain a balanced data set.
8. A model training device, characterized in that: include: A data cleaning unit is used to obtain a data set to be processed, perform data cleaning on the data set to be processed, and obtain a cleaned data set; A data enhancement unit, configured to perform data enhancement on the cleaned data set to obtain a data set to be used; A division unit, configured to divide the data set to be used according to a preset data ratio to obtain a training data set and a verification data set; A training unit, configured to train the risk assessment model to be trained using the training data set to obtain a trained risk assessment model; A prediction unit, configured to input the verification data set into the trained risk assessment model and obtain an outputted predicted risk result corresponding to each verification data in the verification data set; The verification unit is used to verify whether the trained risk assessment model is trained based on the actual risk result and the corresponding predicted risk result corresponding to each verification data.
9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the model training method described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the model training method according to any one of claims 1 to 7 is implemented.