Account classification method and device, storage medium and electronic equipment
By oversampling and fusion processing of multiple classification models, the problem of low account classification accuracy was solved, enabling accurate identification and optimized resource management of low-frequency accounts, thereby improving the accuracy of account classification and service efficiency.
Patent Information
- Application Number
- CN202510895713.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-21
AI Technical Summary
In existing technologies, the accuracy of account classification is low, making it difficult to effectively identify whether user accounts will reduce their frequency of using resource exchange services in the future.
Multiple first-classification models are used to oversample and predict the service data of the target account to obtain a confidence set. The confidence set is then fused using a second-classification model to determine that the target account is a low-frequency account. After classification, the frequency of using resource replacement services for low-frequency accounts will decrease.
It improves the accuracy of account classification, helps financial service platforms to more accurately identify and manage low-frequency accounts, optimize resource allocation, improve service quality and efficiency, and reduce costs.
Smart Images

Figure CN120822175A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of financial technology, and specifically, to an account classification method and device, a storage medium, and an electronic device. Background Art
[0002] Some financial service platforms often offer resource swap services, such as credit-based resource swaps. Specifically, these allow users to pre-qualify resources provided by the financial service platform based on their credit rating and future compensation. These resource swap services typically require user account classification to determine whether their account will decrease its frequency of use of the resource swap service in the future, thereby optimizing resource allocation.
[0003] In related technologies, account classification typically relies on a single machine learning model. However, this approach often struggles to capture important feature relationships and patterns when processing data, leading to low classification accuracy.
[0004] Currently, no effective solution has been proposed to address the problem of low accuracy in account classification in related technologies. Summary of the Invention
[0005] The main purpose of this application is to provide an account classification method to solve the problem of low accuracy of account classification in related technologies.
[0006] To achieve the above-mentioned objectives, according to one aspect of the present application, an account classification method is provided. The method comprises: in each first classification model of the account classification model, predicting the type of the target account based on service data generated when the target account uses a resource replacement service, thereby obtaining a confidence set, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold, and the first classification model is trained based on sample data after oversampling the original sample data; in a second classification model of the account classification model, fusing the confidence set to obtain a fused confidence, wherein the second classification model is trained based on the training output results obtained by processing the sample data by the first classification model; and when the fused confidence is greater than the second threshold, determining that the target account is a low-frequency account, wherein the frequency of the low-frequency account using the resource replacement service in a target time period after classification will decrease until it is lower than the target frequency.
[0007] To achieve the above-mentioned purpose, according to one aspect of the present application, a method for training an account classification model is provided. The method comprises: training each initial first classification model based on sample data obtained by oversampling original sample data to obtain each first classification model, wherein the original sample data includes service data generated when a sample account uses a resource replacement service, and the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold; predicting the type of the sample account based on the sample data in each first classification model to obtain a sample confidence set, wherein the sample confidence is used to characterize the probability that the sample account is a low-frequency account, and the frequency of the low-frequency account using the resource replacement service in a target time period will decrease until it is lower than the target frequency; using the sample confidence set to train the initial second classification model to obtain a second classification model, wherein the account classification model includes each first classification model and the second classification model.
[0008] To achieve the above-mentioned purpose, according to another aspect of the present application, an account classification device is provided. The device includes: a prediction unit, configured to predict the type of a target account based on service data generated when the target account uses a resource replacement service in each first classification model of the account classification model, and obtain a confidence set, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold, and the first classification model is trained based on sample data after oversampling the original sample data; a fusion unit, configured to fuse the confidence set in a second classification model of the account classification model, and obtain a fusion confidence, wherein the second classification model is trained based on the training output result obtained by processing the sample data by the first classification model; and a determination unit, configured to determine that the target account is a low-frequency account when the fusion confidence is greater than a second threshold, wherein the frequency of the low-frequency account using the resource replacement service in the target time period after classification will decrease until it is lower than the target frequency.
[0009] Optionally, in the account classification device provided in the embodiment of the present application, the above-mentioned prediction unit includes: a preprocessing module, used to preprocess the service data to obtain preprocessed service data; a prediction module, used to predict the type of the target account based on the preprocessed service data in each first classification model, and obtain the confidence output by each first classification model; an adding module, used to add the confidence output by each first classification model to the confidence set.
[0010] Optionally, in the account classification device provided in the embodiment of the present application, the above-mentioned fusion unit includes: a first fusion module, used to perform weighted fusion processing on each confidence in the confidence set in the second classification model to obtain a fused confidence; a second fusion module, used to perform nonlinear fusion processing on each confidence in the confidence set in the second classification model to obtain a fused confidence.
[0011] Optionally, the account classification device provided in the embodiment of the present application further includes: a first determination unit, used to determine first sample data and second sample data from the original sample data, wherein the sample account corresponding to the first sample data is marked as a sample high-frequency account, and the sample account corresponding to the second sample data is marked as a sample low-frequency account, the frequency of the sample low-frequency account using the resource replacement service is lower than the target frequency, and the frequency of the sample high-frequency account using the resource replacement service is higher than or equal to the target frequency; a second determination unit, used to determine the sample data whose number of sample data is less than the target number in the first sample data and the second sample data as the data to be sampled, and to determine the sample data whose number of sample data is greater than or equal to the target number as the first target data; an oversampling unit, used to perform oversampling processing on the data to be sampled to obtain the second target data; a third determination unit, used to determine the first target data and the second target data as sample data.
[0012] Optionally, the account classification device provided in the embodiment of the present application further includes: a first prediction unit, used to predict the type of the current sample account corresponding to the current sample data based on the current sample data in the current initial first classification model in each initial first classification model, and obtain the current sample confidence, wherein the sample data includes the current sample data; a fourth determination unit, used to determine the current first loss based on the current sample confidence and the type label of the current sample account, wherein the type label is used to characterize the account type to which the sample account belongs; a fifth determination unit, used to determine the current initial first classification model as the current first classification model that has reached convergence when the current first loss reaches the first loss threshold.
[0013] Optionally, the account classification device provided in the embodiment of the present application further includes: a second prediction unit, used to predict the type of the current sample account based on the current sample data in each first classification model, and obtain a current sample confidence set; a first fusion unit, used to fuse each current sample confidence in the current sample confidence set in the initial second classification model, and obtain a current fused sample confidence; a sixth determination unit, used to determine the current second loss based on the current fused sample confidence and the type label of the current sample account; a seventh determination unit, used to determine the initial second classification model as a converged second classification model when the current second loss reaches the second loss threshold.
[0014] To achieve the above-mentioned purpose, according to another aspect of the present application, a training device for an account classification model is provided. The device includes: a first training unit, configured to train each initial first classification model based on sample data obtained by oversampling original sample data to obtain each first classification model, wherein the original sample data includes service data generated when a sample account uses a resource replacement service, and the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold; a prediction unit, configured to predict the type of the sample account based on the sample data in each first classification model, to obtain a sample confidence set, wherein the sample confidence is used to characterize the probability that the sample account is a low-frequency account, and the frequency of the low-frequency account using the resource replacement service in a target time period will decrease until it is lower than the target frequency; a second training unit, configured to train the initial second classification model using the sample confidence set, to obtain a second classification model, wherein the account classification model includes each first classification model and the second classification model.
[0015] In an embodiment of the present application, in each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and a confidence set is obtained, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold, and the first classification model is trained based on sample data after oversampling the original sample data; in the second classification model of the account classification model, the confidence set is fused to obtain a fused confidence, wherein the second classification model is trained based on the training output result obtained by processing the sample data by the first classification model; when the fused confidence is greater than the second threshold, it is determined that the target account belongs to a low-frequency account, wherein the frequency of the low-frequency account using the resource replacement service in the target time period after classification will decrease until it is lower than the target frequency. In other words, the embodiment of the present application solves the technical problem of low account classification accuracy and achieves the technical effect of improving account classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0017] Figure 1 A hardware structure block diagram of a computer terminal for implementing an account classification method is shown;
[0018] Figure 2 This is a flowchart of an account classification method provided according to an embodiment of the present application;
[0019] Figure 3 This is a flowchart of an account classification method provided according to an embodiment of the present application;
[0020] Figure 4 This is a flowchart of a method for training an account classification model according to an embodiment of the present application;
[0021] Figure 5 is a flowchart of another account classification model training method provided according to an embodiment of the present application;
[0022] Figure 6 is a schematic diagram of an account classification device provided according to an embodiment of the present application;
[0023] Figure 7 is a schematic diagram of a training device for an account classification model provided according to an embodiment of the present application;
[0024] Figure 8 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0027] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0028] Synthetic Minority Over-sampling Technique (SMOTE), the SMOTE algorithm is an oversampling technique used to deal with data imbalance problems. It is mainly used to increase the number of minority class samples, so that the number of samples of different categories in the data set is more balanced.
[0029] Stacked Generalization (Stacking) is an advanced ensemble learning method used to combine the predictions of multiple base learners to improve the accuracy and robustness of the overall prediction.
[0030] Model training refers to the process of adjusting model parameters using training set data. During this phase, data is fed into the model, which attempts to learn patterns and regularities within the data. The training set is the primary source of information for model learning. Based on the inputs and outputs in the training set, the algorithm optimizes the model parameters to minimize the loss function, which measures the difference between predicted and actual values. Common training methods include gradient descent and stochastic gradient descent. After training, the model should be able to make good predictions about the data seen in the training set.
[0031] Model validation uses a validation set to evaluate the model's performance on unseen data to prevent overfitting. During model training, the validation set is used to adjust the model's hyperparameters (such as learning rate, regularization strength, etc.), evaluate the model's generalization ability, and compare and select between different model architectures. The validation set can be considered a simulated test set, allowing developers to gain a preliminary understanding of and adjust the model's performance before the final evaluation. By observing the model's performance on the validation set, the training process can be adjusted, such as adding more regularization, early stopping, and selecting the optimal model version.
[0032] It should be noted that the collected information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, analysis, service data, etc.) involved in this application are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation portals for users to choose to authorize or refuse. For example, an interface is set up between this system and relevant users or institutions to provide users with corresponding operation portals for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.
[0033] Example 1
[0034] According to an embodiment of the present application, an embodiment of an account classification method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0035] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an account classification method. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0036] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0037] Memory 104 can be used to store software programs and modules for application software, such as the program instructions / data storage device corresponding to the account classification method in the embodiments of the present application. Processor 102 executes the software programs and modules stored in memory 104 to perform various functional applications and data processing, thereby implementing the aforementioned account classification method. Memory 104 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory 104 may further include memory remotely located relative to processor 102, and such remote memory may be connected to computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0038] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0039] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0040] Under the above operating environment, this application provides Figure 2 The account classification method shown. Figure 2 This is a flowchart of the account classification method according to Example 1 of the present application.
[0041] Step S101: In each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service to obtain a confidence set, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold. The first classification model is trained based on sample data after oversampling of the original sample data.
[0042] It should be noted that the above-mentioned account classification method can be applied to, but is not limited to, the financial service platform in the classification scenario of user accounts. Specifically, the above-mentioned account classification method is used to determine whether the user account is a low-frequency account that will no longer use or reduce the use of the resource replacement service provided by the financial service platform in the future.
[0043] Optionally, the account classification model is a machine learning model for classifying account types, and may include, but is not limited to, multiple first classification models and one second classification model. The first classification model may be, but is not limited to, a base learner, which is used to predict input data and output each confidence level for the sample category. The multiple first classification models in the account classification model may, but are not limited to, use different types of base learners, such as random forests, extreme gradient boosting learners, support vector machines, logistic regression, K-nearest neighbor algorithms, or other base learners with similar functions, and this is not limited in this embodiment.
[0044] Furthermore, the second classification model is used to integrate the prediction results of the first classification model to improve classification accuracy. The second classification model can be, but is not limited to, logistic regression, linear regression, multi-layer perceptron, support vector machine, or other models with similar functions, which are not limited in this embodiment.
[0045] It should be noted that the resource exchange service described above refers to a service provided to accounts that meet certain credit resource requirements. Credit resources may, but are not limited to, indicate a user's accumulated credit value or credit score within a system or platform. The resource exchange service may pre-provision a corresponding number of exchange resources based on the number of credit resources held by an account. However, the account must commit to returning the exchanged resources to the service platform providing the resource exchange service within a predetermined timeframe. The amount of the returned resources is greater than the amount of the exchanged resources.
[0046] Optionally, the service data includes the account's usage records for resource replacement services, such as frequency of use, quantity of replacement resources obtained, time of obtaining replacement resources, etc. In addition, the service data may, but is not limited to, include information about the object holding the account, and this is not limited to any restrictions in this embodiment.
[0047] It should be noted that the above confidence set includes the confidence output by each first classification model for the type prediction of the target account.
[0048] Step S102: In the second classification model of the account classification model, the confidence set is fused to obtain fused confidence, wherein the second classification model is trained based on the training output result obtained by processing the sample data by the first classification model.
[0049] Optionally, in the second classification model of the account classification model, the confidence set is fused to obtain a fused confidence, which may include but is not limited to: performing weighted fusion processing on each confidence in the confidence set in the second classification model to obtain a fused confidence; or performing nonlinear fusion processing on each confidence in the confidence set in the second classification model to obtain a fused confidence.
[0050] Step S103: When the fusion confidence is greater than the second threshold, the target account is determined to be a low-frequency account, wherein the frequency of using the resource replacement service by the low-frequency account in the target time period after classification will decrease until it is lower than the target frequency.
[0051] As an optional example, assuming that the target account is an account that uses the resource exchange service provided by the financial service platform, the above steps can be explained based on, but not limited to, the following examples:
[0052] S1. In each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service to obtain a confidence set, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold. The first classification model is trained based on sample data after oversampling of the original sample data.
[0053] S2, in a second classification model of the account classification model, performing fusion processing on the confidence set to obtain fused confidence, wherein the second classification model is trained based on the processing results of the sample data by the first classification model;
[0054] S3. When the fusion confidence is greater than the second threshold, it is determined that the target account is a low-frequency account, and then a processing strategy is adopted for the target account, such as providing customized services, that is, providing resource accumulation guidelines and credit resource management suggestions to help the account use its credit resources more effectively, and providing incentives, that is, providing additional rewards for low-frequency accounts to encourage users to increase the frequency of using resource exchange services, etc. This is not limited in this embodiment.
[0055] According to the embodiment of the present application, in each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and a confidence set is obtained, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold, and the first classification model is trained based on sample data after oversampling the original sample data; in the second classification model of the account classification model, the confidence set is fused to obtain a fused confidence, wherein the second classification model is trained based on the training output result obtained by processing the sample data by the first classification model; when the fused confidence is greater than the second threshold, it is determined that the target account belongs to a low-frequency account, wherein the frequency of the low-frequency account using the resource replacement service in the target time period after classification will decrease until it is lower than the target frequency. In other words, according to the embodiment of the present application, on the one hand, when training the model, the oversampling technology is used to balance the number of different types of sample data, thereby making the accuracy of the account classification model obtained based on the sample data training higher, thereby making the type of the target account predicted based on the account classification model more accurate. On the other hand, the output of the first classification model group (i.e., the prediction confidence set for the target account type) is fused through the second classification model to obtain a more stable fusion confidence. This integrated learning method can combine the advantages of multiple models, reduce the deviation between models, improve the generalization ability and robustness of predictions, and further improve the accuracy of account classification. On the other hand, through this method, the relevant service platform can more accurately identify and manage low-frequency accounts, especially in the scenario of resource replacement services, and then adjust the resource replacement service strategy, optimize resource allocation, improve service quality and efficiency, and reduce unnecessary cost losses. In summary, the embodiment of the present application solves the technical problem of low account classification accuracy and achieves the technical effect of improving account classification accuracy.
[0056] Optionally, in the account classification method provided in an embodiment of the present application, in each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and the obtained confidence set includes:
[0057] S1, preprocessing the service data to obtain preprocessed service data.
[0058] Optionally, the above preprocessing may include, but is not limited to, a series of data preprocessing operations such as one-hot encoding and normalization of the service data, which is not limited in this embodiment.
[0059] S2. In each first classification model, the type of the target account is predicted based on the preprocessed service data to obtain the confidence level of the output of each first classification model.
[0060] S3: Add the confidence output by each first classification model to the confidence set.
[0061] In this embodiment of the present application, service data is preprocessed to obtain preprocessed service data; in each first classification model, the type of the target account is predicted based on the preprocessed service data to obtain the confidence level output by each first classification model; and the confidence level output by each first classification model is added to the confidence level set. In other words, in this embodiment of the present application, service data is preprocessed before the first classification model makes predictions, which has the beneficial effect of improving data quality and ensuring the accuracy of model training and predictions.
[0062] Optionally, in the account classification method provided in the embodiment of the present application, in the second classification model of the account classification model, the confidence set is fused to obtain the fused confidence including:
[0063] In the second classification model, each confidence in the confidence set is weighted and fused to obtain a fused confidence.
[0064] Optionally, the weighted fusion processing may include, but is not limited to, weighted summation processing, weighted average processing, weighted multiplication processing, etc., which is not limited in this embodiment.
[0065] Alternatively, nonlinear fusion processing is performed on each confidence level in the confidence level set in the second classification model to obtain a fused confidence level.
[0066] Optionally, the nonlinear fusion process may include, but is not limited to, using a nonlinear function to fuse the confidence values in the confidence set. Nonlinear fusion can capture the complex relationship between confidence values, thereby potentially generating more accurate fused confidence.
[0067] In an embodiment of the present application, a weighted fusion process is performed on each confidence level in the confidence level set in the second classification model to obtain a fused confidence level; or a nonlinear fusion process is performed on each confidence level in the confidence level set in the second classification model to obtain a fused confidence level. In other words, using an embodiment of the present application, weighted fusion can consider the performance (such as accuracy and recall rate) of each first classification model to assign different weights, while nonlinear fusion uses a model (such as a neural network) to capture the complex relationship between confidence levels. This method effectively enhances the generalization ability of the model, reduces the uncertainty of single model predictions, and improves the accuracy of the final classification decision.
[0068] Optionally, in the account classification method provided in the embodiment of the present application, in each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and before obtaining the confidence set, the method further includes:
[0069] S1, the first sample data and the second sample data are determined from the original sample data, wherein the sample account corresponding to the first sample data is marked as a sample high-frequency account, and the sample account corresponding to the second sample data is marked as a sample low-frequency account, the frequency of the sample low-frequency account using the resource replacement service is lower than the target frequency, and the frequency of the sample high-frequency account using the resource replacement service is higher than or equal to the target frequency.
[0070] Optionally, before determining the first sample data and the second sample data from the original sample data, the process may further include, but is not limited to: preprocessing the initial sample data to obtain the original sample data, wherein each item of original sample data includes service data of the sample account corresponding to the original sample data.
[0071] It should be noted that the above-mentioned initial first classification model is an original model that has not been fully trained. It can use, but is not limited to, any machine learning classifier, such as logistic regression, random forest, neural network, etc. For details, please refer to the introduction to the first classification model above, and will not be repeated here.
[0072] S2, determining the sample data whose number of sample data is less than the target number in the first sample data and the second sample data as the data to be sampled, and determining the sample data whose number of sample data is greater than or equal to the target number as the first target data.
[0073] It should be noted that the first and second sample data are two sets of data divided from the original sample data by account type (high-frequency or low-frequency). The first sample data corresponds to data from high-frequency accounts, while the second sample data corresponds to data from low-frequency accounts. The frequency of use of the resource exchange service by high-frequency accounts is greater than or equal to the target threshold.
[0074] Furthermore, the target number is a value set to ensure the balance of sample data and is used to guide oversampling or undersampling operations. This target number is set to ensure that the data volumes of high-frequency and low-frequency accounts are as close as possible, thereby eliminating the impact of class imbalance on model training. In some examples, for the first sample data, the target number can be set to the number of the second sample data, and for the second sample data, the target number can be set to the number of the first sample data.
[0075] Optionally, the data to be sampled may be, but is not limited to, used to indicate a type of data with a smaller quantity in the sample data, which needs to be oversampled to increase its quantity.
[0076] S3, performing oversampling processing on the data to be sampled to obtain second target data.
[0077] Optionally, the to-be-sampled data may be oversampled by, but not limited to, simple replication (replicating minority class samples), SMOTE (Synthetic Minority Oversampling Technology), and the like, which is not limited in this embodiment.
[0078] S4, determining the first target data and the second target data as sample data.
[0079] In an embodiment of the present application, first sample data and second sample data are determined from the original sample data, wherein the sample account corresponding to the first sample data is marked as a sample high-frequency account, and the sample account corresponding to the second sample data is marked as a sample low-frequency account, the frequency of the sample low-frequency account using the resource replacement service is lower than the target frequency, and the frequency of the sample high-frequency account using the resource replacement service is greater than or equal to the target frequency; the sample data whose number of sample data is less than the target number in the first sample data and the second sample data is determined as the data to be sampled, and the sample data whose number of sample data is greater than or equal to the target number is determined as the first target data; the data to be sampled is oversampled to obtain the second target data; the first target data and the second target data are determined as the sample data. In other words, using the embodiment of the present application, the sample data with an unbalanced number is oversampled, and the beneficial effect is to balance the number of high-frequency accounts and low-frequency accounts in the data set. This data balancing strategy improves the representativeness of low-frequency accounts, so that the model can more comprehensively learn the characteristics of various accounts during training, thereby more accurately distinguishing low-frequency accounts from high-frequency accounts during prediction, reducing the bias of the model and improving classification accuracy.
[0080] Optionally, in the account classification method provided in the embodiment of the present application, in each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and before obtaining the confidence set, the method further includes:
[0081] S1. In the current initial first classification model among the initial first classification models, the type of the current sample account corresponding to the current sample data is predicted based on the current sample data to obtain the current sample confidence, wherein the sample data includes the current sample data.
[0082] It should be noted that the sample data can be divided into training data and validation data, but is not limited to the above. The training data is used to train the first and second classification models to adjust the model parameters in the models, and the validation data is used to validate the first and second classification models to evaluate the generalization ability of the models, that is, the performance of the models on unseen data. This helps to determine whether the models are overfitting.
[0083] Optionally, the current sample data is a subset of data used during model training to evaluate model performance and adjust parameters. This data is randomly sampled from the total sample dataset and may vary between iterations. The current sample confidence level is the confidence level of the initial first classification model's prediction of the account type in the current sample data at a particular iteration. Furthermore, the type tag is the actual type label of the sample account, used to compare and verify the accuracy of the model's prediction results.
[0084] S2. Determine the current first loss based on the current sample confidence and the type tag of the current sample account, wherein the type tag is used to characterize the account type to which the sample account belongs.
[0085] It should be noted that the above-mentioned current first loss is the error between the confidence of the model predicting the current sample data and the actual type label of the sample account, which is usually calculated using a loss function (such as cross entropy loss, mean square error).
[0086] S3: When the current first loss reaches the first loss threshold, determine the current initial first classification model as the current first classification model that has reached convergence.
[0087] Optionally, after determining the current first loss based on the current sample confidence and the type label of the current sample account, it can, but is not limited to, also include: when the current first loss reaches the first loss threshold, adjusting the above-mentioned current initial first classification model to obtain the adjusted current initial first classification model, and obtaining the next sample data to continue training the adjusted current initial first classification model.
[0088] It should be noted that the sample data used in the above steps can be but is not limited to training data. After the first classification model is obtained through training data, the first classification model needs to be verified through verification data to finally obtain a first classification model that can be used in practical applications.
[0089] In an embodiment of the present application, in the current initial first classification model among the various initial first classification models, the type of the current sample account corresponding to the current sample data is predicted based on the current sample data to obtain the current sample confidence, wherein the sample data includes the current sample data; based on the current sample confidence and the type label of the current sample account, the current first loss is determined, wherein the type label is used to characterize the account type to which the sample account belongs; when the current first loss reaches the first loss threshold, the current initial first classification model is determined as the current first classification model that has reached convergence. In other words, by using the embodiment of the present application, the degree of fit to the training data is improved while preventing overfitting.
[0090] Optionally, in the account classification method provided in the embodiment of the present application, in each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and before obtaining the confidence set, the method further includes:
[0091] S1. In each first classification model, based on the current sample data, predict the type of the current sample account and obtain the current sample confidence set.
[0092] S2: In the initial second classification model, each current sample confidence in the current sample confidence set is fused to obtain a current fused sample confidence.
[0093] It should be noted that the above-mentioned initial second classification model is a second classification model that has not been fully trained.
[0094] S3: Determine a current second loss based on the current fused sample confidence and the type label of the current sample account. Optionally, the current second loss is a loss value calculated based on the current fused sample confidence and the actual type label (true label) of the sample account, and is used to evaluate the prediction accuracy of the second classification model in the current training round.
[0095] S4. When the current second loss reaches the second loss threshold, determine the initial second classification model as the converged second classification model.
[0096] Optionally, after determining the current second loss based on the current fusion sample confidence and the type label of the current sample account, it can include but is not limited to: when the current second loss does not reach the second loss threshold, adjusting the model parameters of the initial second classification model to obtain an adjusted initial second classification model, and continuing to refer to the above steps S1 to S4 to train the adjusted initial second classification model.
[0097] It should be noted that the sample data used in the above steps can be but is not limited to training data. After the second classification model is obtained through training data, the second classification model needs to be verified through training data to finally obtain a second classification model that can be used in practical applications.
[0098] In the embodiment of the present application, in each first classification model, the type of the current sample account is predicted based on the current sample data to obtain the current sample confidence set; in the initial second classification model, the current sample confidences in the current sample confidence set are fused to obtain the current fused sample confidence; based on the current fused sample confidence and the type label of the current sample account, the current second loss is determined; when the current second loss reaches the second loss threshold, the initial second classification model is determined as the converged second classification model. In other words, the use of the embodiment of the present application ensures that the second classification model can reach the optimal state when the prediction results of the first classification model are integrated. The convergence training of the second classification model enhances the model's ability to integrate the prediction results of the first classification model, thereby improving the accuracy of the classification of the target account type.
[0099] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0100] Example 2
[0101] This application provides Figure 3 The training method of the account classification model shown. Figure 3 This is a flowchart of the training method of the account classification model according to Example 2 of the present application.
[0102] Step S201: Based on the sample data obtained by oversampling the original sample data, each initial first classification model is trained to obtain each first classification model, wherein the original sample data includes service data generated when the sample account uses the resource replacement service, and the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold.
[0103] It should be noted that the above-mentioned sample data obtained by oversampling the original sample data is used to train each initial first classification model, and each first classification model obtained may include but is not limited to: in the current initial first classification model in each initial first classification model, the type of the current sample account corresponding to the current sample data is predicted based on the current sample data to obtain the current sample confidence, wherein the sample data includes the current sample data; based on the current sample confidence and the type label of the current sample account, the current first loss is determined, wherein the type label is used to characterize the account type to which the sample account belongs; when the current first loss reaches the first loss threshold, the current initial first classification model is determined as the current first classification model that has reached convergence.
[0104] In step S202, in each first classification model, the type of the sample account is predicted based on the sample data to obtain a sample confidence set, wherein the sample confidence is used to characterize the probability that the sample account is a low-frequency account. The frequency of the low-frequency account using the resource exchange service in the target time period will decrease until it is lower than the target frequency.
[0105] Optionally, in each first classification model, based on the sample data, the type of the sample account is predicted to obtain a sample confidence set, which may include but is not limited to: in each first classification model, based on the current sample data, the type of the current sample account is predicted to obtain a current sample confidence set.
[0106] Step S203: Use the sample confidence set to train the initial second classification model to obtain a second classification model, wherein the account classification model includes each first classification model and the second classification model.
[0107] It should be noted that, in this embodiment, the target time period may be, but is not limited to, used to indicate a period of time after the above confidence level is determined.
[0108] Optionally, the above-mentioned use of the sample confidence set to train the initial second classification model to obtain the second classification model includes: in the initial second classification model, fusing the current sample confidences in the current sample confidence set to obtain the current fused sample confidence; determining the current second loss based on the current fused sample confidence and the type label of the current sample account; when the current second loss reaches the second loss threshold, determining the initial second classification model as the second classification model that has reached convergence.
[0109] As an optional example, it is possible but not limited to Figure 4 The following example illustrates the above steps:
[0110] Step S301, obtaining original sample data;
[0111] Step S302, performing oversampling processing on the original sample data to obtain sample data;
[0112] Step S303, dividing the sample data into multiple training sets and one validation set;
[0113] Step S304: Using the training set, train the base learner (used to represent the first classification model) and the meta learner (used to represent the second classification model) in the account classification model, and then use the validation set to validate the base learner and the meta learner in the account classification model.
[0114] Step S305: classify the accounts using the trained and verified account classification model.
[0115] As an optional example, taking the example of dividing the sample data into 6 data sets including training set 1, training set 2, training set 3, training set 4, training set 5 and validation set 6, it can be but not limited to the following method: Figure 5 The following example illustrates the above steps:
[0116] After performing a series of data preprocessing operations such as one-hot encoding and normalization on the initial data, the original sample data is obtained.
[0117] Then, in order to make the subsequently constructed model have better prediction effect, the SMOTE oversampling algorithm is used to balance the imbalanced data set to obtain sample data.
[0118] Next, 70% of the sample data is divided out for training the fusion model (used to represent the account classification model) as the training set (including training set 1, training set 2, training set 3, training set 4, and training set 5), and 30% of the data is used to evaluate the performance of the model as the validation set (i.e., validation set 6).
[0119] Then, we select the optimal hyperparameters for the ensemble model. We use a grid search plus cross-validation method to select the optimal hyperparameters for the first-layer base classifier (used to represent the first classification model).
[0120] Next, the five-fold cross-validation method is used to train each base classifier (i.e., base classifier 1, base classifier 2, base classifier 3, and base classifier 4) on the divided training set. The prediction results of the individual classifiers are spliced into a column, denoted as Ai (i = 1, 2, 3, 4), and then the prediction results of the spliced individual base classifiers are integrated together, denoted as (A1, A2, A3, A4), and the integrated results are used as new training samples for the second-layer classifier (used to represent the second classification model). Then, each trained base classifier is predicted on the test set. For a single base classifier, its multiple test results are averaged to obtain a column of values, denoted as Bi (i = 1, 2, 3, 4). The individual test results are integrated together, denoted as (B1, B2, B3, B4), and used as the test sample of the second-layer learner.
[0121] Furthermore, the new training samples (A1, A2, A3, A4) are used to train the meta-learner of the second layer to obtain the account classification model, and then the newly generated test samples (B1, B2, B3, B4) of the first layer are input into the model to obtain the prediction results of the newly constructed model.
[0122] It should be noted that, for other specific implementations of the above steps S201 to S203 in this embodiment, please refer to the implementations and examples provided in the above account classification method, and will not be described in detail in this embodiment.
[0123] In an embodiment of the present application, based on the sample data obtained by oversampling the original sample data, each initial first classification model is trained to obtain each first classification model, wherein the original sample data includes service data generated when the sample account uses the resource replacement service, and the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold; in each first classification model, based on the sample data, the type of the sample account is predicted to obtain a sample confidence set, wherein the sample confidence is used to characterize the probability that the sample account is a low-frequency account, and the frequency of the low-frequency account using the resource replacement service within the target time period will decrease until it is lower than the target frequency; the initial second classification model is trained using the sample confidence set to obtain a second classification model, wherein the account classification model includes each first classification model and the second classification model. In other words, the embodiment of the present application combines the advantages of sample data balance and ensemble learning by adopting a two-layer model training strategy, which not only improves the model's ability to identify low-frequency accounts, but also enhances the model's comprehensive decision-making ability through the fusion processing of the second classification model. The resulting account classification model can maintain high classification accuracy when processing unbalanced data sets.
[0124] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0125] Example 3
[0126] The embodiment of the present application further provides an account classification device. It should be noted that the account classification device of the embodiment of the present application can be used to execute the account classification method provided in the embodiment of the present application. The following introduces the account classification device provided in the embodiment of the present application.
[0127] According to an embodiment of the present application, a device for implementing the above account classification method is also provided, such as Figure 6 As shown, the device includes:
[0128] Prediction unit 602 is configured to predict the type of the target account in each first classification model of the account classification model based on service data generated when the target account uses the resource replacement service, thereby obtaining a confidence set. The resource replacement service is configured to provide replacement resources to accounts whose credit resources reach a first threshold. The first classification model is trained based on sample data that has been oversampled from the original sample data.
[0129] a fusion unit 604 configured to fuse the confidence set in a second classification model of the account classification model to obtain a fused confidence, wherein the second classification model is trained based on a training output result obtained by processing the sample data by the first classification model;
[0130] The determination unit 606 is configured to determine that the target account is a low-frequency account when the fusion confidence is greater than a second threshold, wherein the frequency of the low-frequency account using the resource replacement service in the target time period after classification will decrease until it is lower than the target frequency.
[0131] The account classification device provided in the embodiment of the present application, on the one hand, uses oversampling technology to balance the number of different types of sample data when training the model, thereby making the accuracy of the account classification model obtained based on the sample data training higher, thereby making the type predicted for the target account based on the account classification model more accurate. On the other hand, the output of the first classification model group (i.e., the prediction confidence set for the target account type) is fused by the second classification model to obtain a more stable fusion confidence. This integrated learning method can combine the advantages of multiple models, reduce the deviation between models, improve the generalization ability and robustness of the prediction, and further improve the accuracy of account classification. On the other hand, through this method, the relevant service platform can more accurately identify and manage low-frequency accounts, especially in the scenario of resource replacement services, and then adjust the resource replacement service strategy, optimize resource allocation, improve service quality and efficiency, and reduce unnecessary cost losses. In summary, the embodiment of the present application solves the technical problem of low account classification accuracy and achieves the technical effect of improving account classification accuracy.
[0132] Optionally, in the account classification device provided in the embodiment of the present application, the above-mentioned prediction unit includes: a preprocessing module, used to preprocess the service data to obtain preprocessed service data; a prediction module, used to predict the type of the target account based on the preprocessed service data in each first classification model, and obtain the confidence output by each first classification model; an adding module, used to add the confidence output by each first classification model to the confidence set.
[0133] Optionally, in the account classification device provided in the embodiment of the present application, the above-mentioned fusion unit includes: a first fusion module, used to perform weighted fusion processing on each confidence in the confidence set in the second classification model to obtain a fused confidence; a second fusion module, used to perform nonlinear fusion processing on each confidence in the confidence set in the second classification model to obtain a fused confidence.
[0134] Optionally, the account classification device provided in the embodiment of the present application further includes: a first determination unit, used to determine first sample data and second sample data from the original sample data, wherein the sample account corresponding to the first sample data is marked as a sample high-frequency account, and the sample account corresponding to the second sample data is marked as a sample low-frequency account, the frequency of the sample low-frequency account using the resource replacement service is lower than the target frequency, and the frequency of the sample high-frequency account using the resource replacement service is higher than or equal to the target frequency; a second determination unit, used to determine the sample data whose number of sample data is less than the target number in the first sample data and the second sample data as the data to be sampled, and to determine the sample data whose number of sample data is greater than or equal to the target number as the first target data; an oversampling unit, used to perform oversampling processing on the data to be sampled to obtain the second target data; a third determination unit, used to determine the first target data and the second target data as sample data.
[0135] Optionally, the account classification device provided in the embodiment of the present application further includes: a first prediction unit, used to predict the type of the current sample account corresponding to the current sample data based on the current sample data in the current initial first classification model in each initial first classification model, and obtain the current sample confidence, wherein the sample data includes the current sample data; a fourth determination unit, used to determine the current first loss based on the current sample confidence and the type label of the current sample account, wherein the type label is used to characterize the account type to which the sample account belongs; a fifth determination unit, used to determine the current initial first classification model as the current first classification model that has reached convergence when the current first loss reaches the first loss threshold.
[0136] Optionally, the account classification device provided in the embodiment of the present application further includes: a second prediction unit, used to predict the type of the current sample account based on the current sample data in each first classification model, and obtain a current sample confidence set; a first fusion unit, used to fuse each current sample confidence in the current sample confidence set in the initial second classification model, and obtain a current fused sample confidence; a sixth determination unit, used to determine the current second loss based on the current fused sample confidence and the type label of the current sample account; a seventh determination unit, used to determine the initial second classification model as a converged second classification model when the current second loss reaches the second loss threshold.
[0137] It should be noted that the prediction unit 602, the fusion unit 604, and the determination unit 606 correspond to steps S101 to S103 in Example 1. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned modules can also be run as part of the device in the computer terminal 10 provided in Example 1.
[0138] Example 4
[0139] The present application also provides an account classification model training device. It should be noted that the account classification model training device of the present application can be used to execute the account classification model training method provided in the present application. The following describes the account classification model training device provided in the present application.
[0140] According to an embodiment of the present application, a device for implementing the above-mentioned account classification model training method is also provided. Figure 7 As shown, the device includes:
[0141] A first training unit 702 is configured to train each initial first classification model based on sample data obtained by oversampling the original sample data to obtain each first classification model, wherein the original sample data includes service data generated when a sample account uses a resource replacement service, and the resource replacement service is configured to provide replacement resources for accounts whose credit resources reach a first threshold;
[0142] Prediction unit 704 is configured to predict the type of the sample account based on the sample data in each first classification model, and obtain a sample confidence set, wherein the sample confidence is used to represent the probability that the sample account is a low-frequency account. The frequency of use of the resource exchange service by the low-frequency account will decrease within the target time period until it falls below the target frequency.
[0143] The second training unit 706 is used to train the initial second classification model using the sample confidence set to obtain a second classification model, wherein the account classification model includes each first classification model and the second classification model.
[0144] The account classification model training device provided in the embodiments of this application utilizes a two-layer model training strategy that combines the advantages of sample data balance and ensemble learning. This not only improves the model's ability to identify low-frequency accounts, but also enhances the model's comprehensive decision-making capabilities through the fusion of a second classification model. The resulting account classification model maintains high classification accuracy when processing unbalanced datasets.
[0145] It should be noted that the first training unit 702, the prediction unit 704, and the second training unit 706 correspond to steps S201 to S203 in Example 2. The examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above-mentioned modules can also be run as part of the device in the computer terminal 10 provided in Example 1.
[0146] Example 5
[0147] An embodiment of the present application may provide an electronic device, Figure 8 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 8 As shown, the electronic device may include: one or more ( Figure 8 Only one is shown) processor 802, memory 804, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0148] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the above-mentioned method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network and a combination thereof.
[0149] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: in each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service to obtain a confidence set, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold, and the first classification model is trained based on the sample data after oversampling the original sample data; in the second classification model of the account classification model, the confidence set is fused to obtain a fused confidence, wherein the second classification model is trained based on the training output result obtained by processing the sample data by the first classification model; when the fused confidence is greater than the second threshold, it is determined that the target account belongs to a low-frequency account, wherein the frequency of the low-frequency account using the resource replacement service in the target time period after classification will decrease until it is lower than the target frequency.
[0150] The processor can also call the information and applications stored in the memory through the transmission device to perform the following steps: preprocess the service data to obtain preprocessed service data; in each first classification model, predict the type of the target account based on the preprocessed service data to obtain the confidence output of each first classification model; add the confidence output of each first classification model to the confidence set.
[0151] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: performing weighted fusion processing on each confidence level in the confidence level set in the second classification model to obtain a fused confidence level; or performing nonlinear fusion processing on each confidence level in the confidence level set in the second classification model to obtain a fused confidence level.
[0152] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: determine the first sample data and the second sample data from the original sample data, wherein the sample account corresponding to the first sample data is marked as a sample high-frequency account, and the sample account corresponding to the second sample data is marked as a sample low-frequency account, the frequency of the sample low-frequency account using the resource replacement service is lower than the target frequency, and the frequency of the sample high-frequency account using the resource replacement service is higher than or equal to the target frequency; determine the sample data in the first sample data and the second sample data whose number of sample data is less than the target number as the data to be sampled, and determine the sample data whose number of sample data is greater than or equal to the target number as the first target data; perform oversampling processing on the data to be sampled to obtain the second target data; determine the first target data and the second target data as the sample data.
[0153] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: in the current initial first classification model in each initial first classification model, predict the type of the current sample account corresponding to the current sample data based on the current sample data to obtain the current sample confidence, wherein the sample data includes the current sample data; based on the current sample confidence and the type label of the current sample account, determine the current first loss, wherein the type label is used to characterize the account type to which the sample account belongs; when the current first loss reaches the first loss threshold, determine the current initial first classification model as the current first classification model that has reached convergence.
[0154] The processor can also call the information and application stored in the memory through the transmission device to perform the following steps: in each first classification model, based on the current sample data, predict the type of the current sample account to obtain the current sample confidence set; in the initial second classification model, fuse the individual current sample confidences in the current sample confidence set to obtain the current fused sample confidence; determine the current second loss based on the current fused sample confidence and the type label of the current sample account; when the current second loss reaches the second loss threshold, determine the initial second classification model as the second classification model that has reached convergence.
[0155] The processor can also call the information and application programs stored in the memory through the transmission device to perform the following steps: based on the sample data obtained by oversampling the original sample data, each initial first classification model is trained to obtain each first classification model, wherein the original sample data includes service data generated when the sample account uses the resource replacement service, and the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold; in each first classification model, based on the sample data, the type of the sample account is predicted to obtain a sample confidence set, wherein the sample confidence is used to characterize the probability that the sample account is a low-frequency account, and the frequency of the low-frequency account using the resource replacement service in the target time period will decrease until it is lower than the target frequency; the initial second classification model is trained using the sample confidence set to obtain a second classification model, wherein the account classification model includes each first classification model and the second classification model.
[0156] By using the embodiment of the present application, a solution for account classification is provided. On the one hand, when training the model, the oversampling technology is used to balance the number of different types of sample data, thereby making the accuracy of the account classification model obtained based on the sample data training higher, thereby making the type predicted for the target account based on the account classification model more accurate. On the other hand, the output of the first classification model group (i.e., the prediction confidence set for the target account type) is fused by the second classification model to obtain a more stable fusion confidence. This integrated learning method can combine the advantages of multiple models, reduce the deviation between models, improve the generalization ability and robustness of the prediction, and further improve the accuracy of account classification. On the other hand, through this method, the relevant service platform can more accurately identify and manage low-frequency accounts, especially in the scenario of resource replacement services, and then adjust the resource replacement service strategy, optimize resource allocation, improve service quality and efficiency, and reduce unnecessary cost losses. In summary, by using the embodiment of the present application, the technical problem of low account classification accuracy is solved, and the technical effect of improving account classification accuracy is achieved.
[0157] The present invention also provides a training scheme for an account classification model. By employing a two-layer model training strategy that combines the advantages of sample data balance and ensemble learning, the model not only improves its ability to identify low-frequency accounts but also enhances its comprehensive decision-making capabilities through the integration of a second classification model. The resulting account classification model maintains high classification accuracy when processing unbalanced datasets.
[0158] It can be understood by those skilled in the art that Figure 8 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 8 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 8 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 8 Different configurations shown.
[0159] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0160] Example 6
[0161] The embodiment of the present application further provides a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the account classification method provided in the first embodiment.
[0162] Optionally, in this embodiment, the above-mentioned storage medium can also be used to store the program code executed by the training method of the account classification model provided in the above-mentioned embodiment 2.
[0163] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0164] The present application also provides a computer program product that, when executed on a data processing device, is suitable for executing the steps of the account classification method and the steps of the account classification model training method.
[0165] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0166] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0167] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0168] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0169] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0170] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0171] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. An account classification method, characterized in that: include: In each first classification model of the account classification model, a type of the target account is predicted based on service data generated when the target account uses a resource replacement service to obtain a confidence set, wherein the resource replacement service is used to provide replacement resources for accounts whose credit resources reach a first threshold, and the first classification model is trained based on sample data that has been oversampled from the original sample data; In a second classification model of the account classification model, the confidence set is fused to obtain a fused confidence, wherein the second classification model is trained based on a training output result obtained by processing the sample data by the first classification model; When the fusion confidence is greater than the second threshold, the target account is determined to be a low-frequency account, wherein the frequency of the low-frequency account using the resource exchange service in the target time period after classification will decrease until it is lower than the target frequency.
2. The method according to claim 1, characterized in that In each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and the confidence set obtained includes: Preprocessing the service data to obtain preprocessed service data; In each of the first classification models, predicting the type of the target account based on the preprocessed service data, and obtaining a confidence level output by each of the first classification models; The confidences output by the respective first classification models are added to the confidence set.
3. The method according to claim 1, characterized in that In the second classification model of the account classification model, the confidence set is fused to obtain the fused confidence, including: performing weighted fusion processing on each confidence level in the confidence level set in the second classification model to obtain the fused confidence level; or In the second classification model, nonlinear fusion processing is performed on each confidence in the confidence set to obtain the fused confidence.
4. The method according to any one of claims 1 to 3, characterized in that In each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and before obtaining the confidence set, the method further includes: First sample data and second sample data determined from the original sample data, wherein the sample account corresponding to the first sample data is marked as a sample high-frequency account, and the sample account corresponding to the second sample data is marked as a sample low-frequency account, the frequency of using the resource exchange service by the sample low-frequency account is lower than the target frequency, and the frequency of using the resource exchange service by the sample high-frequency account is higher than or equal to the target frequency; Determine, among the first sample data and the second sample data, sample data whose number of sample data is less than the target number as data to be sampled, and determine sample data whose number of sample data is greater than or equal to the target number as first target data; Performing oversampling processing on the data to be sampled to obtain second target data; The first target data and the second target data are determined as the sample data.
5. The method according to any one of claims 1 to 3, characterized in that In each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and before obtaining the confidence set, the method further includes: In a current initial first classification model among the initial first classification models, predicting the type of a current sample account corresponding to the current sample data based on the current sample data to obtain a current sample confidence, wherein the sample data includes the current sample data; Determining a current first loss based on the current sample confidence and a type tag of the current sample account, wherein the type tag is used to characterize the account type to which the sample account belongs; In a case where the current first loss reaches a first loss threshold, the current initial first classification model is determined as a current first classification model that has reached convergence.
6. The method according to claim 5, characterized in that In each first classification model of the account classification model, the type of the target account is predicted based on the service data generated when the target account uses the resource replacement service, and before obtaining the confidence set, the method further includes: Predicting the type of the current sample account based on the current sample data in each of the first classification models to obtain a current sample confidence set; In the initial second classification model, each current sample confidence in the current sample confidence set is fused to obtain a current fused sample confidence; Determining a current second loss based on the current fusion sample confidence and the type label of the current sample account; When the current second loss reaches a second loss threshold, the initial second classification model is determined as a converged second classification model.
7. A training method for an account classification model, characterized in that: include: Training each initial first classification model based on sample data obtained by oversampling original sample data to obtain each first classification model, wherein the original sample data includes service data generated when a sample account uses a resource replacement service, the resource replacement service being used to provide replacement resources for an account whose credit resources reach a first threshold; In each of the first classification models, based on the sample data, the type of the sample account is predicted to obtain a sample confidence set, wherein the sample confidence is used to represent the probability that the sample account is a low-frequency account. The frequency of the low-frequency account using the resource exchange service within a target time period will decrease until it falls below a target frequency. The initial second classification model is trained using the sample confidence set to obtain a second classification model, wherein the account classification model includes the first classification models and the second classification model.
8. An account classification device, characterized in that: include: a prediction unit configured to predict, in each first classification model of the account classification model, the type of the target account based on service data generated when the target account uses a resource replacement service, to obtain a confidence set, wherein the resource replacement service is configured to provide replacement resources to an account whose credit resources reach a first threshold, and the first classification model is trained based on sample data that has been oversampled from the original sample data; a fusion unit, configured to perform fusion processing on the confidence set in a second classification model of the account classification model to obtain a fusion confidence, wherein the second classification model is trained based on a training output result obtained by processing the sample data by the first classification model; A determination unit is used to determine that the target account is a low-frequency account when the fusion confidence is greater than a second threshold, wherein the frequency of the low-frequency account using the resource replacement service in the target time period after classification will decrease until it is lower than the target frequency.
9. A training method for an account classification model, characterized in that: include: a first training unit configured to train each initial first classification model based on sample data obtained by oversampling original sample data to obtain each first classification model, wherein the original sample data includes service data generated when a sample account uses a resource replacement service, wherein the resource replacement service is configured to provide replacement resources for an account whose credit resources reach a first threshold; a prediction unit, configured to predict the type of the sample account based on the sample data in each of the first classification models, and obtain a sample confidence set, wherein the sample confidence is used to represent the probability that the sample account is a low-frequency account, and the frequency of the low-frequency account using the resource exchange service within a target time period will decrease until it falls below a target frequency; The second training unit is used to train the initial second classification model using the sample confidence set to obtain a second classification model, wherein the account classification model includes the first classification models and the second classification model.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
11. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 7 when running.
12. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.