Data processing method and related apparatus
By adding confidence parameters to the sample data, filtering out high confidence samples and generating a new training sample set, the problem that the risk control model is affected by noise and interference data during the training process is solved, and the goal of improving the model prediction effect is achieved.
Patent Information
- Application Number
- PCT/CN2024/124240
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-10-11
- Publication Date
- 2025-05-08
AI Technical Summary
In the prior art, risk control models are susceptible to noise and interference data during training, resulting in poor prediction results. How to improve the prediction effect of risk control models is a technical problem that technicians are studying.
By adding confidence parameters to the sample data, black and white samples with high confidence are selected, and a new training sample set is generated based on this to reduce the negative impact of noise and interference data on model training.
It improves the accuracy and recall rate of the model, reduces the problem of inaccurate model identification, and improves the overall performance of the risk control system.
Smart Images

Figure CN2024124240_08052025_PF_FP_ABST
Abstract
Description
A data processing method and related device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on October 31, 2023, with application number 202311436106.1, and invention name “A data processing method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a data processing method and related devices. Background Art
[0003] Intelligent risk control refers to the process of identifying, assessing, predicting, and controlling financial risks using artificial intelligence technologies, such as machine learning, deep learning, and natural language processing. Intelligent risk control can improve the efficiency and accuracy of risk management, reduce the likelihood of human intervention and misjudgment, and enhance the competitiveness and risk resilience of financial institutions. Risk control models are the core of intelligent risk control. As shown in Figure 1, the development and application process of a risk control model consists of three phases: risk control data preparation 101, risk control model training 102, and risk control model deployment 103. The risk control data preparation phase involves processing various data types, including raw transaction events, customer data, and behavioral data, identifying fraudulent and legitimate samples, and developing fraud signatures for targeted fraud scenarios such as account takeover (ATO) and telecommunications fraud. The risk control model training phase involves model design, training, evaluation, and validation, with continuous iteration until a risk control model with acceptable key metrics is achieved. The risk control model deployment phase involves deploying the model on the risk control platform, implementing its operation, and conducting online evaluations. This phase allows for real-time risk assessment of key business operations, while also continuously monitoring the model's effectiveness and optimizing it in a timely manner. Automated training and iteration of risk control models based on automatic machine learning (AutoML) is a key technical direction for the current evolution of risk control systems towards automation and intelligence. It enables the automated cyclical evolution of risk control system AI models from "training -> deployment -> launch -> execution -> optimization -> training." This changes the existing risk control model development and application process, which relies on manual participation in fraud feature selection, model training and parameter adjustment, and evaluation and feedback during model training and usage. This significantly reduces reliance on expert manual tuning and has become a key focus for mainstream vendors in the intelligent risk control field. Figure 2 illustrates a typical solution for applying AutoML to risk control business scenarios. It proposes a method for automatically updating risk control models, primarily for updating risk control models. As shown in Figure 2, risk control model updates include three scenarios: automatic model refitting 201, automatic model retraining 202, and incremental learning 203. The automatic model refitting scenario is mainly used for the automated update of models that have been launched, and is generally triggered by problems discovered through model performance monitoring and operations; the automatic model retraining method can complete the automated training of models for new risks and new businesses; incremental learning emphasizes the reuse of trained models and the continuous incremental update of models based on knowledge bases and online data.
[0004] The specific execution process of the risk control business scenario shown in Figure 2 is shown in Figure 3, which specifically includes the following steps:
[0005] Step S301: input data.
[0006] Step S302: obtaining an original data set through data input, the data set including transaction records (ie events) and a label indicating whether the transaction is fraudulent.
[0007] Step S303: Generate data features from multiple dimensions such as event attributes, event accumulation, event sequence, and relationship topology for all labeled data as candidate features.
[0008] Step S304: Utilize preset feature selection methods, such as feature subset search, feature subset evaluation, etc., to filter data features and select high-value features to obtain a model with higher discrimination.
[0009] Step S305: The system uses the selected high-value features to train the model based on the labeled data. The training process uses methods such as network search and Bayesian optimization, combined with the indicator data of model verification and evaluation to iterate and automatically adjust the model hyperparameters. After multiple rounds of iterations, the model with the optimal indicators can be obtained.
[0010] Step S306: Output the model trained based on the latest data to complete the model update.
[0011] The method for automatically updating the risk control model based on AutoML shown in Figure 3 has high noise and a lot of interference data, which makes the prediction effect of the risk control model finally trained poor. How to improve the prediction effect of the risk control model is a technical problem that technicians in this field are studying.
[0012] Summary of the Invention
[0013] The embodiments of the present application provide a data processing method and related devices, which can improve the quality of the training sample set, thereby improving the performance indicators of the trained model.
[0014] In a first aspect, an embodiment of the present application provides a data processing method, the method comprising:
[0015] Acquire multiple sample data about a business system, wherein each sample data includes feature data, label data, and a confidence level, the feature data including business data during the operation of the business system, the label data being used to characterize whether the sample data is a black sample or a white sample, and the confidence level being used to characterize the credibility of the sample data being the sample type identified by the label data;
[0016] generating second sample data of a black sample type according to a plurality of first sample data, wherein the first sample data includes sample data of which confidence level is higher than a preset threshold value among the plurality of sample data;
[0017] A new training sample set is generated based on the historical data training sample set and the second sample data, wherein the training sample set is used to train a business model, and the business model is used to predict sample types based on business data during the operation of the business system.
[0018] In the above method, the sample data is screened by adding a confidence parameter to the sample data to obtain high-confidence black samples and white samples, and subsequent training sample sets are generated based on this. When such a training sample set is used for model training, it can reduce the negative impact of noise and interference data on model training, thereby improving the model (such as the fraud detection classification model of the risk control system) The key indicators such as precision and recall rate.
[0019] In conjunction with the first aspect, in an optional implementation of the first aspect, the method further includes:
[0020] Screening the sample data in the historical training sample set by a sample elimination algorithm to obtain a plurality of third sample data;
[0021] The step of generating second sample data whose sample type is a black sample according to the plurality of first sample data includes:
[0022] generating second sample data of a black sample type according to the plurality of third sample data and the plurality of first sample data;
[0023] Generating a new training sample set based on the historical data training sample set and the second sample data includes:
[0024] A new training sample set is generated according to the plurality of third sample data and the second sample data.
[0025] In this implementation, inappropriate samples in the historical training sample set are eliminated to avoid the existence of historical backward samples that do not conform to the current black and white user behavior patterns. The eliminated training sample set is used together with the above-mentioned second sample data to generate a new training sample set, which can further improve the performance of the final trained model and avoid the problem of "inaccurate" model recognition.
[0026] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, obtaining multiple sample data about the business system includes:
[0027] The label data and confidence corresponding to the accumulated multiple business data are re-evaluated through a confidence evaluation algorithm to obtain multiple sample data, wherein each sample data includes a business data, and the label data and confidence corresponding to the one business data, and each business data in the accumulated multiple business data corresponds to old label data and old confidence before the re-evaluation.
[0028] In this implementation, new label data and confidence levels are generated for business data to overwrite the old ones. This ensures that the acquired sample data is up to date, thereby improving sample quality. Furthermore, the differences between the newly determined label data and confidence levels and the old ones can reveal potential characteristics or changes in the business data. This information can also be used to select training samples or train models, thereby improving the performance of the final model.
[0029] In combination with the first aspect, or any one of the above-mentioned possible implementations of the first aspect, in another possible implementation, the accumulated multiple business data include newly added business data and historical business data, and the business data in the first sample data all belong to the newly added business data.
[0030] In this implementation, the selected first sample data targets newly added business data, which can better highlight the impact of the newly added business data on generating new second sample data. Therefore, from a time dimension, the generated second sample data has more reference value.
[0031] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in another possible implementation, re-evaluating the label data and confidences corresponding to the accumulated multiple business data using a confidence assessment algorithm includes:
[0032] For the first business data among the accumulated multiple business data, the corresponding label data and confidence are evaluated according to the target information, wherein the target information includes one or more of feedback information, occurrence time and clustering characteristics, and the feedback information includes supplementary label data and / or confidence information indicating the first business data. The occurrence time is the time when the first business data occurs, and the clustering characteristics are used to characterize the category to which the first business data belongs in at least two clustering results formed by the multiple business data. The first business data is any one of the accumulated multiple business data.
[0033] In this implementation, feedback information, occurrence time, clustering characteristics, and other information are used to evaluate label data and confidence, which can improve the accuracy and stability of label data and confidence.
[0034] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, the target information includes the feedback information, the occurrence time, and the clustering characteristics; and evaluating, for each of the accumulated multiple business data, the corresponding label data and confidence level based on the target information, includes:
[0035] evaluating first label data and a first confidence level corresponding to the first service data according to the feedback information;
[0036] evaluating a second confidence level corresponding to the first business data according to the occurrence time;
[0037] evaluating a third confidence level corresponding to the first business data according to the clustering characteristics;
[0038] Determine label data corresponding to the first business data based on the first label data, and determine a confidence level corresponding to the first business data based on the first confidence level, the second confidence level, and the second confidence level.
[0039] In this implementation, the first label data and the first confidence level, especially the first confidence level, are obtained based on the fusion of feedback information, occurrence time, and clustering characteristics. This can take into account the impact of user experience, time changes, and clustering characteristics on the results. Therefore, the obtained first confidence level has higher accuracy and better stability.
[0040] In conjunction with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, the first label data is a sample type indicated by the feedback information, and if the first label data after evaluation based on the feedback information is unchanged compared to the label data of the first service data before evaluation based on the feedback information, then the first confidence level is greater than the confidence level corresponding to the first service data before evaluation based on the feedback information. It will be understood that the unchanged label data indicates, to a certain extent, good stability, and therefore a higher confidence level may be assigned.
[0041] In conjunction with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, if the first label data corresponding to the first business data does not change after evaluation using the feedback information compared to the label data corresponding to the first business data before evaluation, the longer the first business data occurred, the greater the second confidence level generated for the first business data. It will be appreciated that a longer period of time during which the label data remains unchanged indicates greater stability, and thus a higher confidence level is assigned.
[0042] In combination with the first aspect, or any one of the above-mentioned possible implementation methods of the first aspect, in another possible implementation method, there are multiple clustering results among the accumulated multiple business data, and the third confidence corresponding to the forward business data that is closer to the cluster center in each clustering result is higher, and the third confidence corresponding to the reverse business data that is closer to the cluster center in each clustering result is lower. The forward business data is business data whose sample type is consistent with the sample type represented by the clustering result to which it belongs, and the reverse business data is business data whose sample type is inconsistent with the sample type represented by the clustering result to which it belongs. The sample type represented by the clustering result is the sample type corresponding to more business data in the clustering result.
[0043] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, the method further includes:
[0044] Determining second business data, wherein the second business data is business data whose label data changes after the re-evaluation compared to before the re-evaluation or whose confidence level changes by more than a reference threshold;
[0045] The method of filtering the sample data in the historical training sample set by the sample elimination algorithm to obtain a plurality of third sample data includes:
[0046] A target deletion operation is performed on the sample data in the historical training sample set, and the target deletion operation includes: deleting the fourth sample data containing the second business data, and / or deleting the fifth sample data in the sample data cluster to which it belongs and for which no new sample data has been added for more than a preset time period; wherein, the sample data remaining in the historical training sample set after the target deletion operation is performed is the third sample data.
[0047] In this implementation, deleting the fourth sample data and / or the fifth sample data can further improve the quality of the remaining samples in the training data set and reduce the noise impact caused by unstable sample data and sample data that is no longer representative.
[0048] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, the generated second sample data carries an association identifier, where the association identifier is used to identify other sample data used to generate the second sample data; and the target deletion operation further includes:
[0049] The sample data associated with the fourth sample data and / or the fifth sample data is deleted.
[0050] It can be understood that further deleting sample data associated with the fourth sample data and / or the fifth sample data can further improve the quality of the remaining samples in the training data set and reduce the noise impact caused by unstable sample data and sample data that is no longer representative.
[0051] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, generating second sample data whose service type is a black sample based on the multiple third sample data and the multiple first sample data includes:
[0052] Oversampling the sample data of which the sample type is a black sample in the reference data set to generate second sample data, wherein the reference data set includes the plurality of third sample data and the plurality of first sample data.
[0053] In this way, by oversampling, the proportion of black samples in the new training sample set is increased, which can significantly improve the imbalance of training samples caused by the small number of black samples, thereby eliminating the "bias" problem of the model caused by the imbalance of sample classification.
[0054] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, oversampling the sample data of the reference dataset whose sample type is a black sample to generate the second sample data includes:
[0055] Selecting k-nearest neighbors of sixth sample data from the reference data set, where the sixth sample data is any one of the plurality of first sample data whose sample type is a black sample, or the sixth sample data is any one of the reference data set whose sample type is a black sample;
[0056] Determine the sample data with the highest confidence among the k nearest neighbors;
[0057] The second sample data associated with the sixth sample data is generated according to the sample data with the highest confidence among the k nearest neighbors.
[0058] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, generating the new training sample set based on the plurality of third sample data and the second sample data includes:
[0059] performing denoising processing on sample data in a prepared data set, wherein the prepared data set includes the plurality of third sample data and the second sample data;
[0060] Under-sampling is performed on the sample data of the white sample type in the denoised prepared data set to obtain the new training sample set.
[0061] It can be understood that denoising can further improve the quality of training samples, and undersampling can further improve the imbalance of black and white samples.
[0062] In combination with the first aspect, or any one of the foregoing possible implementations of the first aspect, in yet another possible implementation, the prepared data set further includes the plurality of first sample data.
[0063] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, the method further includes:
[0064] The business model is trained using the newly generated training sample set.
[0065] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in yet another possible implementation, the method further includes:
[0066] Sample type prediction is performed based on business data during the operation of the business system.
[0067] In combination with the first aspect, or any one of the above-mentioned possible implementations of the first aspect, in another possible implementation, the business data includes transaction data, and when the label data corresponding to the transaction data is a black sample, it indicates that there is fraudulent behavior or transaction risk in the transaction data.
[0068] In a second aspect, an embodiment of the present application provides a data processing device, the device comprising:
[0069] an acquisition unit, configured to acquire a plurality of sample data regarding a business system, wherein each sample data includes feature data, label data, and a confidence level, wherein the feature data includes business data during the operation of the business system, the label data is used to characterize whether the sample data is a black sample or a white sample, and the confidence level is used to characterize the credibility of the sample data being the sample type identified by the label data;
[0070] A first generating unit is configured to generate second sample data of a black sample type according to a plurality of first sample data, wherein the first sample data includes sample data of which confidence level is higher than a preset threshold value among the plurality of sample data;
[0071] The second generating unit is used to generate a new training sample set based on the historical data training sample set and the second sample data, wherein the training sample set is used to train the business model, and the business model is used to predict the sample type based on the business data during the operation of the business system.
[0072] In the above method, the sample data is screened by adding a confidence parameter to the sample data to obtain high-confidence black samples and white samples, and subsequent training sample sets are generated based on this. When such a training sample set is used for model training, it can reduce the negative impact of noise and interference data on model training, thereby improving the model (such as the fraud detection classification model of the risk control system) The key indicators such as precision and recall rate.
[0073] In conjunction with the second aspect, in an optional implementation of the second aspect, the apparatus further includes:
[0074] a screening unit, configured to screen the sample data in the historical training sample set by using a sample elimination algorithm to obtain a plurality of third sample data;
[0075] In terms of generating second sample data of a black sample type according to a plurality of first sample data, the first generating unit is specifically configured to:
[0076] generating second sample data of a black sample type according to the plurality of third sample data and the plurality of first sample data;
[0077] In terms of generating a new training sample set according to the existing data training sample set and the second sample data, the second generating unit is specifically configured to:
[0078] A new training sample set is generated according to the plurality of third sample data and the second sample data.
[0079] In this implementation, inappropriate samples in the historical training sample set are eliminated to avoid the existence of historical backward samples that do not conform to the current black and white user behavior patterns. The eliminated training sample set is used together with the above-mentioned second sample data to generate a new training sample set, which can further improve the performance of the final trained model and avoid the problem of "inaccurate" model recognition.
[0080] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, in terms of acquiring multiple sample data about the business system, the acquiring unit is specifically configured to:
[0081] The label data and confidence corresponding to the accumulated multiple business data are re-evaluated through a confidence evaluation algorithm to obtain multiple sample data, wherein each sample data includes a business data, and the label data and confidence corresponding to the one business data, and each business data in the accumulated multiple business data corresponds to old label data and old confidence before the re-evaluation.
[0082] In this implementation, new label data and confidence levels are generated for business data to overwrite the old ones. This ensures that the acquired sample data is up to date, thereby improving sample quality. Furthermore, the differences between the newly determined label data and confidence levels and the old ones can reveal potential characteristics or changes in the business data. This information can also be used to select training samples or train models, thereby improving the performance of the final model.
[0083] In combination with the second aspect, or any one of the above-mentioned possible implementations of the second aspect, in another possible implementation, the accumulated multiple business data include newly added business data and historical business data, and the business data in the first sample data all belong to the newly added business data.
[0084] In this implementation, the selected first sample data targets newly added business data, which can better highlight the impact of the newly added business data on generating new second sample data. Therefore, from a time dimension, the generated second sample data has more reference value.
[0085] In combination with the second aspect, or any of the above possible implementations of the second aspect, in another possible implementation, the label data and confidence levels corresponding to the accumulated multiple business data are re-evaluated using a confidence evaluation algorithm, and the acquisition unit is specifically configured to:
[0086] For the first business data among the accumulated multiple business data, the corresponding label data and confidence are evaluated according to the target information, wherein the target information includes one or more of feedback information, occurrence time and clustering characteristics, and the feedback information includes supplementary label data and / or confidence information indicating the first business data. The occurrence time is the time when the first business data occurs, and the clustering characteristics are used to characterize the category to which the first business data belongs in at least two clustering results formed by the multiple business data. The first business data is any one of the accumulated multiple business data.
[0087] In this implementation, feedback information, occurrence time, clustering characteristics, and other information are used to evaluate label data and confidence, which can improve the accuracy and stability of label data and confidence.
[0088] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in another possible implementation, the target information includes the feedback information, the occurrence time, and the clustering characteristics; and for each of the accumulated multiple business data, the corresponding label data and confidence level are evaluated based on the target information, and the acquisition unit is specifically configured to:
[0089] evaluating first label data and a first confidence level corresponding to the first service data according to the feedback information;
[0090] evaluating a second confidence level corresponding to the first business data according to the occurrence time;
[0091] evaluating a third confidence level corresponding to the first business data according to the clustering characteristics;
[0092] Determine label data corresponding to the first business data based on the first label data, and determine a confidence level corresponding to the first business data based on the first confidence level, the second confidence level, and the second confidence level.
[0093] In this implementation, the first label data and the first confidence level, especially the first confidence level, are obtained based on the fusion of feedback information, occurrence time, and clustering characteristics. This can take into account the impact of user experience, time changes, and clustering characteristics on the results. Therefore, the obtained first confidence level has higher accuracy and better stability.
[0094] In combination with the second aspect, or any one of the above-mentioned possible implementations of the second aspect, in another possible implementation, the first label data is the sample type indicated by the feedback information. If the first label data after evaluation according to the feedback information is not changed compared to the label of the first business data before evaluation according to the feedback information, then the first confidence is greater than the confidence corresponding to the first business data before evaluation according to the feedback information.
[0095] In conjunction with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, if the first label data corresponding to the first business data does not change after evaluation using the feedback information compared to the label data corresponding to the first business data before evaluation, the longer the first business data occurred, the greater the second confidence level generated for the first business data. It will be appreciated that a longer period of time during which the label data remains unchanged indicates greater stability, and thus a higher confidence level is assigned.
[0096] In combination with the second aspect, or any of the above-mentioned possible implementation methods of the second aspect, in another possible implementation method, there are multiple clustering results among the accumulated multiple business data, and the third confidence corresponding to the forward business data that is closer to the cluster center in each clustering result is higher, and the third confidence corresponding to the reverse business data that is closer to the cluster center in each clustering result is lower. The forward business data is business data whose sample type is consistent with the sample type represented by the clustering result to which it belongs, and the reverse business data is business data whose sample type is inconsistent with the sample type represented by the clustering result to which it belongs. The sample type represented by the clustering result is the sample type corresponding to more business data in the clustering result.
[0097] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, the apparatus further includes:
[0098] a determining unit, configured to determine second business data, wherein the second business data is business data whose label data changes after the re-evaluation compared with before the re-evaluation or whose confidence level changes by an amount exceeding a reference threshold;
[0099] In the aspect of obtaining a plurality of third sample data by screening the sample data in the historical training sample set through the sample elimination algorithm, the screening unit is specifically configured to:
[0100] A target deletion operation is performed on the sample data in the historical training sample set, and the target deletion operation includes: deleting the fourth sample data containing the second business data, and / or deleting the fifth sample data in the sample data cluster to which it belongs and for which no new sample data has been added for more than a preset time period; wherein, the sample data remaining in the historical training sample set after the target deletion operation is performed is the third sample data.
[0101] In this implementation, deleting the fourth sample data and / or the fifth sample data can further improve the quality of the remaining samples in the training data set and reduce the noise impact caused by unstable sample data and sample data that is no longer representative.
[0102] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, the generated second sample data carries an association identifier, where the association identifier is used to identify other sample data used to generate the second sample data; and the target deletion operation further includes:
[0103] The sample data associated with the fourth sample data and / or the fifth sample data is deleted.
[0104] It can be understood that further deleting sample data associated with the fourth sample data and / or the fifth sample data can further improve the quality of the remaining samples in the training data set and reduce the noise impact caused by unstable sample data and sample data that is no longer representative.
[0105] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, in the aspect of generating the second sample data with a business type of black sample based on the multiple third sample data and the multiple first sample data, the first generating unit is specifically configured to:
[0106] Oversampling the sample data of which the sample type is a black sample in the reference data set to generate second sample data, wherein the reference data set includes the plurality of third sample data and the plurality of first sample data.
[0107] In this way, by oversampling, the proportion of black samples in the new training sample set is increased, which can significantly improve the imbalance of training samples caused by the small number of black samples, thereby eliminating the "bias" problem of the model caused by the imbalance of sample classification.
[0108] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, in the aspect of oversampling the sample data whose sample type is a black sample in the reference dataset to generate the second sample data, the first generating unit is specifically configured to:
[0109] Selecting k-nearest neighbors of sixth sample data from the reference data set, where the sixth sample data is any one of the plurality of first sample data whose sample type is a black sample, or the sixth sample data is any one of the reference data set whose sample type is a black sample;
[0110] Determine the sample data with the highest confidence among the k nearest neighbors;
[0111] The second sample data associated with the sixth sample data is generated according to the sample data with the highest confidence among the k nearest neighbors.
[0112] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, in the aspect of generating the new training sample set based on the multiple third sample data and the second sample data, the second generating unit is specifically configured to:
[0113] performing denoising processing on sample data in a prepared data set, wherein the prepared data set includes the plurality of third sample data and the second sample data;
[0114] Under-sampling is performed on the sample data of the white sample type in the denoised prepared data set to obtain the new training sample set.
[0115] It can be understood that denoising can further improve the quality of training samples, and undersampling can further improve the imbalance of black and white samples.
[0116] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, the prepared data set further includes the plurality of first sample data.
[0117] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, the apparatus further includes:
[0118] A training unit is used to train the business model using the newly generated training sample set.
[0119] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in yet another possible implementation, the apparatus further includes:
[0120] The prediction unit is used to predict the sample type based on the business data during the operation of the business system.
[0121] In combination with the second aspect, or any of the above-mentioned possible implementations of the second aspect, in another possible implementation, the business data includes transaction data, and when the label data corresponding to the transaction data is a black sample, it indicates that there is fraudulent behavior or transaction risk in the transaction data.
[0122] In a third aspect, an embodiment of the present application provides a data processing device comprising at least one processor and at least one memory, wherein the at least one memory stores a computer program; when the computer program is executed by the processor, the method described in the first aspect or any possible implementation of the first aspect is implemented.
[0123] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium runs on a processor, it implements the method described in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0124] FIG1 is a schematic diagram of a scenario of a risk control model in the prior art;
[0125] FIG2 is a schematic diagram of applying AutoML to a risk control business scenario provided by an embodiment of the present application;
[0126] FIG3 is a schematic diagram of a training process of a risk control model provided in an embodiment of the present application;
[0127] FIG4 is a schematic diagram of a scenario of a risk control model provided in an embodiment of the present application;
[0128] FIG5 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;
[0129] FIG6 is a schematic structural diagram of another data processing device provided in an embodiment of the present application;
[0130] FIG7 is a flow chart of a data processing method provided in an embodiment of the present application;
[0131] FIG8 is a schematic diagram of a confidence assessment process provided in an embodiment of the present application;
[0132] FIG9 is a schematic diagram of a confidence assessment scenario provided by an embodiment of the present application;
[0133] FIG10 is a schematic diagram of a scenario of sample data deletion provided in an embodiment of the present application;
[0134] FIG11 is a schematic diagram of an oversampling scenario provided in an embodiment of the present application;
[0135] FIG12 is a schematic structural diagram of a data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0136] The embodiments of the present invention are described below with reference to the accompanying drawings.
[0137] The data processing method in the embodiments of the present application can be applied to various scenarios involving model training, such as data type prediction models, risk control models, object recognition models, etc. Taking the risk control model as an example, as shown in Figure 4, it can be used in the following links of the risk control model:
[0138] 1. Online automatic refitting of risk control models: When the risk control model of an online application is automatically updated, the embodiments of the present application can be used to obtain high-quality training samples to perform higher-quality model fitting.
[0139] 2. Offline automatic retraining of risk control models: Offline training of risk control models is a common scenario in the process of building risk control models. The embodiments of the present application can provide high-quality training samples during the offline automatic training of risk control models, thereby improving the performance of the final risk control model.
[0140] 3. Offline / online incremental learning of risk control models: For incremental learning scenarios of risk control models, incremental high-quality training samples are provided for automatic training of risk control models.
[0141] The focus of the embodiments of this application is how to obtain high-quality training samples.
[0142] Please refer to Figure 5, which shows a data processing device provided in an embodiment of the present application. The data processing device can also be referred to as a data processing system. In essence, it can be a single device (such as a server) or a device cluster consisting of multiple devices. The data processing device can be deployed locally or in the cloud. The data processing device can be a physical hardware device or a virtual device, such as a virtual machine. In short, any device with a certain computing capability can be used as a data processing device in the embodiments of the present application.
[0143] The data processing device 50 shown in FIG5 includes one or more processors 501, one or more memories 502, and optionally, one or more communication interfaces 503, wherein the processor 501, memory 502 and communication interface 503 are interconnected via a bus or other means.
[0144] The memory 502 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM). The memory 502 is used for related computer programs and data, such as for storing computer programs for sample data processing and for storing sample data.
[0145] The processor 501 can be single-core or multi-core. The processor can be various forms of integrated circuits, such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), and of course other forms of processors, which will not be listed here one by one.
[0146] The processor 501 calls the computer program stored in the memory 502 to execute the data processing method shown in Figure 7. After the data processing device 50 obtains a training sample set by executing the data processing method through the processor 501, the training sample set can be cached, temporarily stored, or stored for a long time in the memory 502. The data processing device 50 can also send the training sample set (or part of the samples therein) to other devices through the communication interface 503. The other device can be a device dedicated to storing data or a device for performing model training.
[0147] In the embodiment of the present application, after the data processing device 50 generates the final training sample set, it can perform model training based on the training sample set itself, or send the training sample set to other devices for training. Furthermore, the trained model can be used by the data processing device 50 itself to perform data prediction or classification, or the model can be sent to other devices for prediction or classification.
[0148] In the embodiment of the present application, the process of the data processing device 50 generating a training sample set through the processor 501 requires a series of operations to realize the corresponding functions. From the perspective of software functions, the data processing device 50 includes multiple functional systems. For example, as shown in FIG6 , the data processing device includes a confidence assessment system 601, an elimination system 602, an oversampling system 603, a denoising system 604, an undersampling system 605, an original data set 606, and a training sample set 607. Among them, the business system (such as the risk control system) can submit the business data generated during operation to the data processing device 50, and this part of the data can be added to the original data set 606. In addition, the user can submit feedback information for the business data, or report information, which is the user's evaluation result or judgment result of the business data; the confidence assessment system 601 evaluates the label data and confidence of the input business data based on the feedback information, occurrence time, clustering characteristics and other dimensions. The elimination system 602 maintains the historical training sample set, mainly to eliminate the black samples (such as fraud samples) with reduced confidence and long-occurring time in the historical training data. The oversampling system 603, denoising system 604, and undersampling system 605 work together to address the imbalance between black and white samples. Specifically, the oversampling system 603 generates new high-confidence black samples, the denoising system 604 filters the newly generated black samples and deletes erroneous samples, and the undersampling system 605 samples the filtered samples to reduce the proportion of white samples. The resulting sample set is the new training sample set 607, which can be used by the model training system to train the model. The specific working principles of the various functional systems shown in Figure 6 will be explained later in conjunction with the method shown in Figure 7.
[0149] It should be noted that some of the functional systems shown in Figure 6 are optional, such as the elimination system 602, the denoising system 604, the undersampling system 605, etc., which may not exist. The structure shown in Figure 6 is only an example. When one or several of the functional systems do not exist, the remaining functional systems can still realize the corresponding functions and ultimately contribute to the addition of more high-quality black samples and / or white samples to the newly generated training sample set. Therefore, compared with the existing technology solutions, it can still achieve better results, and all functional systems sometimes have effects that promote each other.
[0150] Please refer to FIG. 7 , which is a flow chart of a data processing method provided in an embodiment of the present application. The method can be implemented based on the architecture shown in FIG. 5 , the architecture shown in FIG. 6 , or other architectures. The method includes but is not limited to the following steps:
[0151] Step S701: Acquire an original data set for the business system.
[0152] Specifically, the data business generated during the operation of the business system can be actively collected, or the business system can actively report the business data generated during the operation. The business data may include the time when the business occurred, the operating status of the business system, the attributes of the business system, etc. For example, if the business system is a risk control system, the business data can be transaction data, and the transaction data may include the time of occurrence, the transaction object, the transaction target, etc. In the embodiment of the present application, the collection of these acquired business data is called the original data set. Each business data in the original data set corresponds to a label data and a confidence level. Each business data in the original data set, together with its corresponding label data and confidence level, constitutes an original sample data. In the original sample data, the label data is used to characterize the sample type of the sample data as a black sample or a white sample, and the confidence level is used to characterize the credibility of the sample data as the sample type identified by the label data. The label data and confidence level initially corresponding to each business data can be default or generated by a pre-set algorithm, which is not limited here. If it is set by default, the confidence level initially corresponding to the latest generated business data is usually relatively low, for example, it can be set to 0%.
[0153] Optionally, the original data set may also include user feedback or report information on the above-mentioned business data. One piece of feedback information (or report information) may be for one business data or for multiple business data. In addition, multiple pieces of feedback information (or report information) may be for one business data. The feedback information may include supplementary (such as user feedback or report) label data indicating the business data and / or confidence information. "Confidence" in the embodiments of this application may also be referred to as reliability, probability, etc.
[0154] Step S702: Acquire multiple sample data about the business system.
[0155] What is actually needed here is to obtain new feature data and new confidence for each business data in the original data set, so as to form a new sample. There are many ways to obtain new multiple sample data, such as sending by other devices, or generating by themselves. Taking self-generation as an example, the label data and confidence corresponding to the multiple business data accumulated in the original data set can be re-evaluated through the confidence evaluation algorithm to obtain multiple sample data, where each sample data includes a business data, and the label data and confidence corresponding to the one business data.
[0156] In an optional scheme, the re-evaluation of the label data and confidence corresponding to the multiple accumulated business data through the confidence evaluation algorithm includes: for the first business data among the multiple accumulated business data, evaluating the corresponding label data and confidence according to the target information, wherein the target information includes one or more of feedback information, occurrence time and clustering characteristics, the feedback information includes supplementary (such as user feedback or reporting) information indicating the label data and / or confidence of the first business data, the occurrence time is the time when the first business data occurs, and the clustering characteristics are used to characterize the category to which the first business data belongs in at least two clustering results formed by the multiple business data, and the first business data is any one of the multiple accumulated business data, that is, each business data in the multiple accumulated business data has the characteristics of the first business data described here.
[0157] The label data and / or confidence level may be determined based on one of the feedback information, the occurrence time, and the clustering characteristics alone, or based on two or three of the information. As shown in FIG8 , the following description is given using the method of determining the label data and / or confidence level based on the above three information levels as an example, i.e., the target information includes the feedback information, the occurrence time, and the clustering characteristics; and for each of the accumulated multiple business data, evaluating the corresponding label data and confidence level based on the target information may include:
[0158] Step 1: Evaluate first label data and first confidence level corresponding to first service data based on the feedback information.
[0159] For example, the initial label data corresponding to the first business data is 0, and the initial corresponding confidence is 0%, where the label data of 0 represents a white sample, and the label data of 1 represents a black sample. Then, after receiving the user's feedback information, the feedback information indicates the data label 1, then this link can determine the first label data as 1 and the first confidence as 60% (or other values, depending on the specific rules). Of course, the feedback information may also indicate a confidence level, then the confidence level can also be used as the basis for determining the first confidence level. Generally, the higher the confidence level, the higher the corresponding first confidence level. The confidence level can be directly used as the first confidence level, or the value calculated by inputting the confidence level into the preset algorithm can be used as the first confidence level, depending on the specific rules.
[0160] For another example, if the first label data after evaluation based on the feedback information is unchanged from the label data of the first business data before evaluation based on the feedback information, then the first confidence level is greater than the confidence level corresponding to the first business data before evaluation based on the feedback information. For example, if the label data initially corresponding to the first business data is 0 and the initial corresponding confidence level is 0%, then upon receiving user feedback information indicating data label 0, it indicates that the user approves of the initial sample type of the first business data. Therefore, the confidence level initially corresponding to the first business data, 0%, is increased to obtain the first confidence level, for example, increased to 65%.
[0161] In general, compared with the label data and confidence initially corresponding to the first business data, the first label data and first confidence corresponding to the first business data are closer to the label data and confidence indicated in the feedback information (if any).
[0162] Step 2: Evaluate a second confidence level corresponding to the first business data according to the occurrence time.
[0163] For example, if the first label data corresponding to the first business data does not change after the evaluation of the feedback information compared to the label data corresponding to the first business data before the evaluation, the longer the occurrence time of the first business data, the greater the second confidence level generated for the first business data. Taking the first business data as transaction data as an example, when a fraudulent transaction (corresponding to a black sample) occurs, the victim generally chooses to report the case, so the system will receive feedback information on the fraud case. The reporting time usually occurs within a period of time after the fraudulent transaction occurs. If a first business data with label data of 0 is received, and no feedback information including label data 1 is received, the longer the occurrence time of the first business data, the lower the probability that the first business data is a black sample (i.e., a fraudulent transaction), and the higher the probability that it is a white sample, that is, the higher the second confidence level generated for the first business data with label data of 0.
[0164] Step 3: Evaluate a third confidence level corresponding to the first business data based on the clustering characteristics.
[0165] Specifically, there are multiple clustering results in the accumulated multiple business data. The closer the positive business data is to the cluster center in each clustering result, the higher the third confidence level is. The closer the negative business data is to the cluster center in each clustering result, the lower the third confidence level is. The positive business data is business data whose sample type is consistent with the sample type represented by the clustering result to which it belongs. The negative business data is business data whose sample type is inconsistent with the sample type represented by the clustering result to which it belongs. The sample type represented by the clustering result is the sample type corresponding to more business data in the clustering result. For example, a clustering result includes 10 items. Business data, 8 of which have label data of 0 (i.e., white samples), and the other 2 have label data of 1 (i.e., black samples), then the sample type represented by the clustering result is a white sample. Then, all white samples in this clustering result (i.e., business data with label data of 0) are positive business data, and all black samples in this clustering result (i.e., business data with label data of 1) are negative business data; based on the idea that there are multiple clustering results in the accumulated multiple business data, the positive business data closer to the cluster center in each clustering result has a higher third confidence level, the corresponding algorithm can be used to calculate the third confidence level.
[0166] It should be noted that clustering is a key characteristic for detecting the behavior of black sample generators. This refers to the fact that the behavioral characteristics of black sample generators are relatively similar, but significantly different from those of normal users. This characteristic is used to perform cluster analysis on the entire data set. Samples can be clustered based on the distance between them, with samples with close distances forming a cluster result (also called a sample cluster). It can generally be assumed that within a cluster result, the label data of the sample data is more likely to be consistent. Therefore, the clustering characteristics can be used to evaluate the confidence level of the label data (i.e., the sample type) corresponding to the sample data.
[0167] Step 4: Determine the label data corresponding to the first business data based on the first label data, and determine the confidence level corresponding to the first business data based on the first confidence level, the second confidence level, and the second confidence level.
[0168] There are many ways to determine the confidence level corresponding to the first business data based on the first confidence level, the second confidence level, and the second confidence level. For example, one confidence level can be selected from the three confidence levels according to a certain algorithm as the confidence level corresponding to the first business data, or these three confidence levels can be input into the corresponding algorithm to generate a new confidence level as the confidence level corresponding to the first business data.
[0169] Optionally, a weighted approach is used to calculate the confidence corresponding to the first business data for the first confidence C1, the second confidence C2, and the third confidence C3, for example, confidence = MAX(C0+αC1+βC2+γC3, 100), where C0 is the confidence initially corresponding to the first business data, and α, β, and γ are the weights corresponding to the first confidence C1, the second confidence C2, and the third confidence C3, respectively.
[0170] Figure 9 illustrates how label data and confidence levels are generated, using specific data. The "Number" column represents different pieces of first business data, with five pieces shown here. When performing a weighted calculation, the first confidence level C1, second confidence level C2, and third confidence level C3 corresponding to each piece of first business data with the same number are weighted to obtain the overall confidence level for that piece of first business data. Therefore, in the example shown in Figure 9, the overall confidence level corresponding to each of the five pieces of first business data can be obtained.
[0171] Step S703: Determine a plurality of first sample data from the plurality of sample data.
[0172] The first sample data includes sample data with a confidence level higher than a preset threshold value among the plurality of sample data. The preset threshold value can be configured as needed, for example, if it is configured to 70%, then among the plurality of sample data, black samples with a confidence level higher than 70% will be selected as the first sample data, and white samples with a confidence level higher than 70% will also be selected as the first sample data.
[0173] Step S704: Filter the sample data in the historical training sample set by a sample elimination algorithm to obtain a plurality of third sample data.
[0174] It should be noted that the spatial distribution of sample features will change over time, as shown in part a of Figure 10. We need to find a way to eliminate and delete samples with such changes or large changes. There are many ways to eliminate and delete, such as deleting samples that are too far away from the time of occurrence; for example, deleting samples that are too far away from the cluster center according to clustering characteristics, and so on; the specific elimination method is not limited here, and can be determined according to the specific needs of the scene, as long as the samples in the training sample set that do not meet the needs of the current scene can be eliminated to a certain extent. For specific information on how to determine which samples need to be deleted due to feature offset, please refer to the examples in parts b and c of Figure 10. For ease of understanding, the following is an example of an optional elimination method:
[0175] First, determine the second business data, wherein the second business data is business data whose label data changes or the confidence changes by more than the reference threshold after the re-evaluation compared to before the re-evaluation (corresponding to case 1 shown in part b of Figure 10); for example, a certain business data, whose initial corresponding data label is 0 (i.e., white sample) and the confidence is 60%, after the previous re-evaluation, the new corresponding label data of the business data is 1 (i.e., black sample) and the confidence is 70%. In this case, the label data has changed, so the business data can be determined as the second business data. For another example, if the above-mentioned reference threshold is 30%, the initial corresponding data label of a certain business data is 0 (i.e., white sample) and the confidence is 60%, after the previous re-evaluation, the new corresponding label data of the business data is still 0 (i.e., white sample) but the confidence is 10%. In this case, the confidence change exceeds the reference threshold, so the business data can be determined as the second business data. Optionally, the change exceeding the reference threshold can include the change from large to small exceeding the reference threshold.
[0176] Then, a target deletion operation is performed on the sample data in the historical training sample set, and the target deletion operation includes: deleting the fourth sample data containing the second business data (corresponding to situation 1 shown in part b of Figure 10), and / or deleting the fifth sample data in the sample data cluster to which it belongs that has no new sample data for more than a preset time (corresponding to situation 2 shown in part c of Figure 10). It can be understood that if a cluster that mainly represents black samples has no black samples added for a long time, the representativeness of the cluster for black samples is weakened, and the behavioral characteristics reflected by the business data in this cluster are no longer adopted by fraudsters, so it needs to be deleted; similarly, if a cluster that mainly represents white samples has no white samples added for a long time, the representativeness of the cluster for white samples is weakened, and the behavioral characteristics reflected by the business data in this cluster are no longer adopted by normal users, so it needs to be deleted.
[0177] Optionally, the historical training sample set also includes sample data generated based on existing samples, such as black samples generated by oversampling. For this type of generated sample data, if the above-mentioned fourth sample data or fifth sample data is used in the generation process, then this type of generated sample data will carry an association identifier, which is used to characterize the fourth sample data or fifth sample data used to generate the sample data; therefore, when deleting the fourth sample data and / or fifth sample data, the sample data generated by this type can also be found according to the association identifier, and then deleted, that is, the sample data associated with the fourth sample data and / or fifth sample data with the association identifier is deleted.
[0178] In the embodiment of the present application, the sample data remaining after the above-mentioned target deletion operation is performed on the historical training sample set is the third sample data.
[0179] Step S705: Generate second sample data whose sample type is a black sample according to the plurality of third sample data and the plurality of first sample data.
[0180] It can be understood that the third sample data and the first sample data obtained in the above manner are both sample data with relatively high sample quality. Therefore, using these sample data as a reference, black samples with relatively high sample quality can also be generated, which are referred to as second sample data for ease of description.
[0181] There are many ways to generate the second sample data based on the plurality of third sample data and the plurality of first sample data. For ease of understanding, the following example illustrates: oversampling the sample data of a black sample type in a reference dataset to generate the second sample data, wherein the reference dataset includes the plurality of third sample data and the plurality of first sample data. For example, the following detailed process may be included:
[0182] Step 1: Select the k nearest neighbors of the sixth sample data from the reference data set. There are many implementation methods for the sixth sample data, such as:
[0183] The sixth sample data in mode 1 is any sample data whose sample type is a black sample in the reference data set. That is to say, every sample data whose sample type is a black sample in the reference data set can be executed in the same way as the sixth sample data and finally generate the second sample data.
[0184] Method 2: The sixth sample data is any sample data whose sample type is a black sample among the multiple first sample data. That is to say, each sample data whose sample type is a black sample among the multiple first sample data can be executed in accordance with the sixth sample data and finally generate the second sample data. For example, the business data in the first sample data does not have corresponding sample data in the training data set.
[0185] The k-nearest neighbor in the embodiment of the present application is an algorithm, also called a k-nearest neighbor system. For example, as shown in Figure 11, the triangle ▲ represents the sixth sample data. Through the k-nearest neighbor system, based on the distance calculation formula SMOTE-ENC, k-nearest neighbor black samples can be obtained, that is, 4 black solid circles connected to the triangle ▲.
[0186] Step 2: Determine the sample data with the highest confidence among the k nearest neighbors. Optionally, a black sample can be randomly selected from the k nearest neighbor black samples, or it can be selected from them according to certain rules. This step takes the selection of the one with the highest confidence as an example, such as the top one of the four nearest neighbor black samples shown in Figure 11, which is connected by a solid line.
[0187] Step 3: Generate the second sample data associated with the sixth sample data based on the sample data with the highest confidence among the k nearest neighbors. For example, as shown in FIG11 , the confidence of the selected sample data with the highest confidence is Cb, and the confidence of the sixth sample data is Ca. Then, a sample data, i.e., the second sample data, is generated in an interval between the sample data with the highest confidence and the sixth sample data. For example, the second sample data is generated in the interval indicated by “}” in FIG11 . The generated second sample data is indicated by a black ★, wherein the distance between the sample data with the highest confidence and the sixth sample data is L. In this case, the length of the interval “}” satisfies the following relationship:
[0188] ①If Ca>Cb, then the length of the interval “}” is equal to L*Cb / Ca.
[0189] ②If Ca<Cb, then the length of the interval “}” is equal to L*Ca / Cb.
[0190] ③If Ca=Cb, then the length of the interval “}” is equal to L.
[0191] It should be noted that the case shown in FIG11 is only an example, and the second sample data can actually be generated in other ways.
[0192] Optionally, the generated second sample data carries an association identifier, and the association identifier is used to represent other sample data used to generate the second sample data, that is, the association identifier can point to the original sample to facilitate subsequent execution of some association operations.
[0193] In an optional solution, the second sample data can be generated without the third sample data, that is, the third sample data can be generated based on multiple first sample data. Since the first sample data is sample data of relatively high quality, the first sample data can be used as a reference to generate some second sample data that are relatively similar to the black samples in the first sample data. There are many specific implementation methods, and the k-nearest neighbor algorithm mentioned above can be used as an optional solution.
[0194] Step S706: Generate a new training sample set according to the plurality of third sample data and the second sample data.
[0195] In one solution, the multiple third sample data and the second sample data are used to generate the new training sample set. The specific use is not limited here. For example, the union of the two can be used as a new training sample set, or the union of the two can be sampled to generate a new training sample set. The union of the two can also be subjected to relevant deletion or generation operations to obtain a new training sample set. The specific implementation method can be determined according to the needs of the scenario. For ease of understanding, the following example illustrates an optional implementation method:
[0196] First, denoising is performed on sample data in a prepared dataset, wherein the prepared dataset includes the plurality of third sample data and the second sample data, i.e., the union of the plurality of third sample data and the second sample data; optionally, the prepared dataset may also include the plurality of first sample data. Denoising may remove sample data that is clearly inappropriate or clearly defective. The specific data samples to be removed by denoising may be determined by a preset algorithm or rule, which is not specifically limited herein.
[0197] Then, the sample data of the sample type of white samples in the denoised preliminary data set are undersampled to obtain the new training sample set. This process reduces the proportion of white samples and increases the proportion of black samples, which is conducive to the balance of sample data.
[0198] In another scheme, the historical training sample set and the second sample data are used to generate the new training sample set. For example, the second sample data is added to the historical training sample set to obtain a new training sample set; for another example, the second sample data is added to the historical training set and the old sample data in the historical training set that does not meet the preset elimination algorithm is eliminated, and the remaining samples are used as the training sample set; of course, it is also possible to perform specific processing on the second sample data before adding it to the historical training sample set, which is not specifically limited here.
[0199] Optionally, after obtaining a training sample set, the new training sample set can be output to other devices for training, or the data processing device can itself train the aforementioned business model using the newly generated training sample set. If the data processing device itself has trained the business model, it can also predict the input business data based on the business model to determine the sample type, that is, determine whether the business data is a black sample (such as fraudulent data) or a white sample (such as non-fraudulent data). Of course, the business model can also be sent to other devices for prediction or classification.
[0200] In the embodiments of the present application, as mentioned above, there are many forms of business data, and the business data may be different in different application scenarios. For example, the business data includes transaction data. In this case, when the label data corresponding to the transaction data is a black sample, it indicates that the transaction data contains fraudulent behavior or transaction risks. The same applies to other scenarios.
[0201] In the method described in Figure 7, the sample data is screened by adding a confidence parameter to the sample data to obtain high-confidence black samples and white samples, and subsequent training sample sets are generated based on this. When such training sample sets are used for model training, the negative impact of noise and interference data on model training can be reduced, thereby improving key indicators such as the precision and recall rate of the model (such as the fraud detection classification model of the risk control system).
[0202] Furthermore, by eliminating inappropriate samples from the historical training sample set, it is possible to avoid the problem of "inaccurate" model recognition caused by the existence of historical backward samples that do not conform to the current black and white user behavior patterns.
[0203] Furthermore, the high-confidence black and white samples are used to generate new black samples (i.e., oversampling) from the eliminated training sample set to increase the proportion of black samples in the new training sample set. This can significantly improve the imbalance of training samples caused by the lack of black samples, thereby reducing the "bias" problem of the model caused by imbalanced sample classification. In addition, denoising can further improve the quality of training samples, and undersampling can further improve the imbalance of black and white samples.
[0204] The above describes in detail the method according to the embodiment of the present invention. The following provides an apparatus according to the embodiment of the present invention.
[0205] Please refer to Figure 12, which is a structural diagram of a data processing device provided by an embodiment of the present invention. The data processing device can be the data processing equipment mentioned above or a device (or module) in the data processing device. The data processing device 120 may include an acquisition unit 1201, a first generation unit 1202, and a second generation unit 1203, wherein each unit is described in detail as follows.
[0206] The acquisition unit 1201 is used to acquire multiple sample data about the business system, wherein each sample data includes feature data, label data and confidence, the feature data includes business data during the operation of the business system, the label data is used to characterize the sample type of the sample data as a black sample or a white sample, and the confidence is used to characterize the credibility of the sample data as the sample type identified by the label data; the acquisition unit 1201 here can be the confidence assessment system 601 mentioned above.
[0207] The first generating unit 1202 is configured to generate second sample data of a black sample type based on a plurality of first sample data, wherein the first sample data includes sample data of which confidence level is higher than a preset threshold value among the plurality of sample data; the first generating unit 1202 may be the aforementioned oversampling system 603.
[0208] The second generating unit 1203 is used to generate a new training sample set based on the historical data training sample set and the second sample data, wherein the training sample set is used to train the business model, and the business model is used to predict the sample type based on the business data during the operation of the business system.
[0209] In the above method, the sample data is screened by adding a confidence parameter to the sample data to obtain high-confidence black samples and white samples, and subsequent training sample sets are generated based on this. When such a training sample set is used for model training, it can reduce the negative impact of noise and interference data on model training, thereby improving the model (such as the fraud detection classification model of the risk control system) The key indicators such as precision and recall rate.
[0210] In an optional implementation, the apparatus 120 further includes:
[0211] A screening unit is configured to screen the sample data in the historical training sample set by using a sample elimination algorithm to obtain a plurality of third sample data; the screening unit may be the elimination system 602 mentioned above.
[0212] In terms of generating second sample data of a black sample type according to a plurality of first sample data, the first generating unit is specifically configured to:
[0213] generating second sample data of a black sample type according to the plurality of third sample data and the plurality of first sample data;
[0214] In terms of generating a new training sample set according to the existing data training sample set and the second sample data, the second generating unit is specifically configured to:
[0215] A new training sample set is generated according to the plurality of third sample data and the second sample data.
[0216] In this implementation, inappropriate samples in the historical training sample set are eliminated to avoid the existence of historical backward samples that do not conform to the current black and white user behavior patterns. The eliminated training sample set is used together with the above-mentioned second sample data to generate a new training sample set, which can further improve the performance of the final trained model and avoid the problem of "inaccurate" model recognition.
[0217] In yet another possible implementation, in terms of acquiring a plurality of sample data related to the business system, the acquiring unit is specifically configured to:
[0218] The label data and confidence corresponding to the accumulated multiple business data are re-evaluated through a confidence evaluation algorithm to obtain multiple sample data, wherein each sample data includes a business data, and the label data and confidence corresponding to the one business data, and each business data in the accumulated multiple business data corresponds to old label data and old confidence before the re-evaluation.
[0219] In this implementation, new label data and confidence levels are generated for business data to overwrite the old ones. This ensures that the acquired sample data is up to date, thereby improving sample quality. Furthermore, the differences between the newly determined label data and confidence levels and the old ones can reveal potential characteristics or changes in the business data. This information can also be used to select training samples or train models, thereby improving the performance of the final model.
[0220] In another possible implementation, the accumulated multiple business data include newly added business data and historical business data, and the business data in the first sample data all belong to the newly added business data.
[0221] In this implementation, the selected first sample data targets newly added business data, which can better highlight the impact of the newly added business data on generating new second sample data. Therefore, from a time dimension, the generated second sample data has more reference value.
[0222] In another possible implementation, the re-evaluating the label data and confidence levels corresponding to the accumulated plurality of business data using a confidence evaluation algorithm, the acquiring unit is specifically configured to:
[0223] For the first business data among the accumulated multiple business data, the corresponding label data and confidence are evaluated according to the target information, wherein the target information includes one or more of feedback information, occurrence time and clustering characteristics, and the feedback information includes supplementary label data and / or confidence information indicating the first business data. The occurrence time is the time when the first business data occurs, and the clustering characteristics are used to characterize the category to which the first business data belongs in at least two clustering results formed by the multiple business data. The first business data is any one of the accumulated multiple business data.
[0224] In this implementation, feedback information, occurrence time, clustering characteristics, and other information are used to evaluate label data and confidence, which can improve the accuracy and stability of label data and confidence.
[0225] In another possible implementation, the target information includes the feedback information, the occurrence time, and the clustering characteristics; for each of the accumulated multiple business data, the corresponding label data and confidence level are evaluated according to the target information, and the acquisition unit is specifically configured to:
[0226] evaluating first label data and a first confidence level corresponding to the first service data according to the feedback information;
[0227] evaluating a second confidence level corresponding to the first business data according to the occurrence time;
[0228] evaluating a third confidence level corresponding to the first business data according to the clustering characteristics;
[0229] Determine label data corresponding to the first business data based on the first label data, and determine a confidence level corresponding to the first business data based on the first confidence level, the second confidence level, and the second confidence level.
[0230] In this implementation, the first label data and the first confidence level, especially the first confidence level, are obtained based on the fusion of feedback information, occurrence time, and clustering characteristics. This can take into account the impact of user experience, time changes, and clustering characteristics on the results. Therefore, the obtained first confidence level has higher accuracy and better stability.
[0231] In another possible implementation, the first label data is the sample type indicated by the feedback information. If the label of the first label data after evaluation according to the feedback information is not changed compared to the label of the first business data before evaluation according to the feedback information, then the first confidence level is greater than the confidence level corresponding to the first business data before evaluation according to the feedback information.
[0232] In another possible implementation, if the first label data corresponding to the first service data does not change after evaluation using the feedback information compared to the label data corresponding to the first service data before evaluation, the longer the first service data occurred, the greater the second confidence level generated for the first service data. It will be appreciated that a longer period of unchanged label data indicates greater stability, and thus a higher confidence level is assigned.
[0233] In another possible implementation, there are multiple clustering results among the accumulated multiple business data. The closer the forward business data is to the cluster center in each clustering result, the higher the third confidence corresponding to the data. The closer the reverse business data is to the cluster center in each clustering result, the lower the third confidence corresponding to the data. The forward business data is business data whose sample type is consistent with the sample type represented by the clustering result to which it belongs. The reverse business data is business data whose sample type is inconsistent with the sample type represented by the clustering result to which it belongs. The sample type represented by the clustering result is the sample type corresponding to more business data in the clustering result.
[0234] In yet another possible implementation, the apparatus further includes:
[0235] A determination unit is used to determine second business data, wherein the second business data is business data whose label data changes or the confidence level changes by more than a reference threshold after the re-evaluation compared to before the re-evaluation; the operation performed by the determination unit can also be completed by the confidence evaluation system 601 or the elimination system 602 mentioned above.
[0236] In the aspect of obtaining a plurality of third sample data by screening the sample data in the historical training sample set through the sample elimination algorithm, the screening unit is specifically configured to:
[0237] A target deletion operation is performed on the sample data in the historical training sample set, and the target deletion operation includes: deleting the fourth sample data containing the second business data, and / or deleting the fifth sample data in the sample data cluster to which it belongs and for which no new sample data has been added for more than a preset time period; wherein, the sample data remaining in the historical training sample set after the target deletion operation is performed is the third sample data.
[0238] In this implementation, deleting the fourth sample data and / or the fifth sample data can further improve the quality of the remaining samples in the training data set and reduce the noise impact caused by unstable sample data and sample data that is no longer representative.
[0239] In another possible implementation, the generated second sample data carries an association identifier, where the association identifier is used to identify other sample data used to generate the second sample data; and the target deletion operation further includes:
[0240] The sample data associated with the fourth sample data and / or the fifth sample data is deleted.
[0241] It can be understood that further deleting sample data associated with the fourth sample data and / or the fifth sample data can further improve the quality of the remaining samples in the training data set and reduce the noise impact caused by unstable sample data and sample data that is no longer representative.
[0242] In yet another possible implementation, in terms of generating the second sample data with a business type of black sample according to the plurality of third sample data and the plurality of first sample data, the first generating unit is specifically configured to:
[0243] Oversampling the sample data of which the sample type is a black sample in the reference data set to generate second sample data, wherein the reference data set includes the plurality of third sample data and the plurality of first sample data.
[0244] In this way, by oversampling, the proportion of black samples in the new training sample set is increased, which can significantly improve the imbalance of training samples caused by the small number of black samples, thereby eliminating the "bias" problem of the model caused by the imbalance of sample classification.
[0245] In yet another possible implementation, in terms of oversampling the sample data of the reference data set whose sample type is a black sample to generate the second sample data, the first generating unit is specifically configured to:
[0246] Selecting k-nearest neighbors of sixth sample data from the reference data set, where the sixth sample data is any one of the plurality of first sample data whose sample type is a black sample, or the sixth sample data is any one of the reference data set whose sample type is a black sample;
[0247] Determine the sample data with the highest confidence among the k nearest neighbors;
[0248] The second sample data associated with the sixth sample data is generated according to the sample data with the highest confidence among the k nearest neighbors.
[0249] In yet another possible implementation, in terms of generating the new training sample set according to the plurality of third sample data and the second sample data, the second generating unit is specifically configured to:
[0250] De-noising is performed on the sample data in the prepared data set, wherein the prepared data set includes the plurality of third sample data and the second sample data; for example, this operation can be completed by the aforementioned de-noising system 604.
[0251] Under-sampling is performed on the sample data of the white sample type in the denoised preliminary data set to obtain the new training sample set. For example, this operation can be completed by the under-sampling system 605 mentioned above.
[0252] It can be understood that denoising can further improve the quality of training samples, and undersampling can further improve the imbalance of black and white samples.
[0253] In yet another possible implementation, the prepared data set further includes the plurality of first sample data.
[0254] In yet another possible implementation, the apparatus further includes:
[0255] A training unit is used to train the business model using the newly generated training sample set.
[0256] In yet another possible implementation, the apparatus further includes:
[0257] The prediction unit is used to predict the sample type based on the business data during the operation of the business system.
[0258] In another possible implementation, the business data includes transaction data, and when the label data corresponding to the transaction data is a black sample, it indicates that the transaction data contains fraudulent behavior or transaction risks.
[0259] It should be noted that the implementation and beneficial effects of each unit may also correspond to the corresponding description of the method embodiment shown in FIG7 .
[0260] An embodiment of the present invention also provides a chip system, which includes at least one processor, a memory and an interface circuit, wherein the memory, the transceiver and the at least one processor are interconnected through lines, and a computer program is stored in the at least one memory; when the computer program is executed by the processor, the method flow shown in Figure 7 is implemented.
[0261] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is executed on a processor, the method flow shown in FIG. 7 is implemented.
[0262] An embodiment of the present invention further provides a computer program product, which, when executed on a processor, implements the method flow shown in FIG. 7 .
[0263] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program or computer program-related hardware. The computer program can be stored in a computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing computer program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A data processing method, characterized in that: include: Acquire multiple sample data about the business system, wherein each sample data includes feature data, label data and confidence, the feature data includes business data during the operation of the business system, the label data is used to characterize whether the sample type of the sample data is a black sample or a white sample, and the confidence is used to characterize the credibility of the sample data as the sample type identified by the label data; Generate second sample data whose sample type is a black sample according to a plurality of first sample data, wherein the first sample data includes sample data whose confidence level is higher than a preset threshold value among the plurality of sample data; A new training sample set is generated based on the historical data training sample set and the second sample data, wherein the training sample set is used to train a business model, and the business model is used to predict sample types based on business data during the operation of the business system.
2. The method according to claim 1, characterized in that The method further comprises: The sample data in the historical training sample set are screened by a sample elimination algorithm to obtain a plurality of third sample data; The step of generating second sample data whose sample type is a black sample according to the plurality of first sample data comprises: Generate second sample data whose sample type is a black sample according to the plurality of third sample data and the plurality of first sample data; The step of generating a new training sample set according to the historical data training sample set and the second sample data includes: A new training sample set is generated according to the plurality of third sample data and the second sample data.
3. The method according to claim 2, characterized in that The obtaining of a plurality of sample data about the business system includes: Re-evaluate the label data and confidence corresponding to the accumulated multiple business data through a confidence evaluation algorithm to obtain multiple sample data, wherein each sample data includes a business data, and the label data and confidence corresponding to the one business data, and each business data in the accumulated multiple business data corresponds to old label data and old confidence before the re-evaluation.
4. The method according to claim 3, characterized in that: The accumulated multiple business data include newly added business data and historical business data, and the business data in the first sample data all belong to the newly added business data.
5. The method according to claim 3 or 4, characterized in that: The re-evaluating the label data and confidence corresponding to the accumulated multiple business data by using the confidence evaluation algorithm includes: For the first business data among the accumulated multiple business data, the corresponding label data and confidence are evaluated according to the target information, wherein the target information includes one or more of feedback information, occurrence time and clustering characteristics, and the feedback information includes supplementary information indicating the label data and / or confidence of the first business data. The occurrence time is the time when the first business data occurs, and the clustering characteristics are used to characterize the category to which the first business data belongs in at least two clustering results formed by the multiple business data, and the first business data is any one of the accumulated multiple business data.
6. The method according to claim 5, characterized in that The target information includes the feedback information, the occurrence time and the clustering characteristics; for each of the accumulated multiple business data, evaluating the corresponding label data and confidence according to the target information includes: evaluating first label data and a first confidence level corresponding to the first service data according to the feedback information; evaluating a second confidence level corresponding to the first business data according to the occurrence time; evaluating a third confidence level corresponding to the first business data according to the clustering characteristics; The label data corresponding to the first business data is determined according to the first label data, and the confidence level corresponding to the first business data is determined according to the first confidence level, the second confidence level and the second confidence level.
7. The method according to claim 6, characterized in that The first label data is the sample type indicated by the feedback information. If the first label data after evaluation according to the feedback information is not changed compared to the label data of the first business data before evaluation according to the feedback information, then the first confidence is greater than the confidence corresponding to the first business data before evaluation according to the feedback information.
8. The method according to claim 6 or 7, characterized in that: If the first label data corresponding to the first business data has not changed after evaluation through the feedback information compared to the label data corresponding to the first business data before evaluation, the longer the occurrence time of the first business data, the greater the second confidence corresponding to the generated first business data.
9. The method according to any one of claims 6 to 8, characterized in that: There are multiple clustering results among the accumulated multiple business data. The third confidence corresponding to the forward business data that is closer to the cluster center in each clustering result is higher, and the third confidence corresponding to the reverse business data that is closer to the cluster center in each clustering result is lower. The forward business data is business data whose sample type is consistent with the sample type represented by the clustering result to which it belongs, and the reverse business data is business data whose sample type is inconsistent with the sample type represented by the clustering result to which it belongs. The sample type represented by the clustering result is the sample type corresponding to more business data in the clustering result.
10. The method according to any one of claims 3 to 9, characterized in that: The method further comprises: Determine second business data, wherein the second business data is business data whose label data changes or whose confidence changes by a magnitude exceeding a reference threshold after the re-evaluation compared with before the re-evaluation; The method of screening the sample data in the historical training sample set by using the sample elimination algorithm to obtain a plurality of third sample data includes: A target deletion operation is performed on the sample data in the historical training sample set, and the target deletion operation includes: deleting the fourth sample data containing the second business data, and / or deleting the fifth sample data in the sample data cluster to which it belongs that has no new sample data for more than a preset time period; wherein the sample data remaining in the historical training sample set after performing the target deletion operation is the third sample data.
11. The method according to claim 10, characterized in that The generated second sample data carries an association identifier, where the association identifier is used to represent other sample data used to generate the second sample data; The target deletion operation further includes: The sample data associated with the fourth sample data and / or the fifth sample data is deleted.
12. The method according to any one of claims 2 to 11, characterized in that: The generating second sample data with a business type of black sample according to the plurality of third sample data and the plurality of first sample data comprises: The sample data whose sample type is a black sample in the reference data set is oversampled to generate second sample data, wherein the reference data set includes the plurality of third sample data and the plurality of first sample data.
13. The method according to claim 12, characterized in that The oversampling of the sample data of the reference data set whose sample type is a black sample to generate the second sample data comprises: Selecting k nearest neighbors of sixth sample data from the reference data set, wherein the sixth sample data is any sample data of the plurality of first sample data whose sample type is a black sample, or the sixth sample data is any sample data of the reference data set whose sample type is a black sample; Determine the sample data with the highest confidence among the k nearest neighbors; The second sample data associated with the sixth sample data is generated according to the sample data with the highest confidence among the k nearest neighbors.
14. The method according to any one of claims 2 to 13, characterized in that: The generating the new training sample set according to the plurality of third sample data and the second sample data comprises: Performing denoising processing on sample data in a prepared data set, wherein the prepared data set includes the plurality of third sample data and the second sample data; Under-sampling is performed on sample data whose sample type is white samples in the denoised prepared data set to obtain the new training sample set.
15. The method according to claim 14, characterized in that The prepared data set also includes the plurality of first sample data.
16. The method according to any one of claims 2 to 15, characterized in that: The method further comprises: The business model is trained using the newly generated training sample set.
17. The method according to claim 16, characterized in that The method further comprises: Sample type prediction is performed based on business data during the operation of the business system.
18. The method according to any one of claims 1 to 17, characterized in that: The business data includes transaction data. When the label data corresponding to the transaction data is a black sample, it indicates that the transaction data contains fraudulent behavior or transaction risks.
19. A data processing device, characterized in that: include: an acquisition unit, configured to acquire a plurality of sample data about a business system, wherein each sample data includes feature data, label data, and confidence, wherein the feature data includes business data during the operation of the business system, the label data is used to characterize whether the sample type of the sample data is a black sample or a white sample, and the confidence is used to characterize the credibility of the sample data as the sample type identified by the label data; A first generating unit, configured to generate second sample data whose sample type is a black sample according to a plurality of first sample data, wherein the first sample data includes sample data whose confidence level is higher than a preset threshold value among the plurality of sample data; The second generating unit is used to generate a new training sample set based on the historical data training sample set and the second sample data, wherein the training sample set is used to train the business model, and the business model is used to predict the sample type according to the business data during the operation of the business system.
20. The device according to claim 19, characterized in that The device also includes: A screening unit, configured to screen the sample data in the historical training sample set by using a sample elimination algorithm to obtain a plurality of third sample data; In the aspect of generating second sample data whose sample type is a black sample according to a plurality of first sample data, the first generating unit is specifically used for: Generate second sample data whose sample type is a black sample according to the plurality of third sample data and the plurality of first sample data; In terms of generating a new training sample set according to the existing data training sample set and the second sample data, the second generating unit is specifically used for: A new training sample set is generated according to the plurality of third sample data and the second sample data.
21. The device according to claim 20, characterized in that In terms of acquiring a plurality of sample data about the business system, the acquiring unit is specifically used to: Re-evaluate the label data and confidence corresponding to the accumulated multiple business data through a confidence evaluation algorithm to obtain multiple sample data, wherein each sample data includes a business data, and the label data and confidence corresponding to the one business data, and each business data in the accumulated multiple business data corresponds to old label data and old confidence before the re-evaluation.
22. The device according to claim 20, characterized in that The accumulated multiple business data include newly added business data and historical business data, and the business data in the first sample data all belong to the newly added business data.
23. The device according to claim 21 or 22, characterized in that The re-evaluating the label data and confidences corresponding to the accumulated multiple business data by the confidence evaluation algorithm, the acquisition unit is specifically used for: For the first business data among the accumulated multiple business data, the corresponding label data and confidence are evaluated according to the target information, wherein the target information includes one or more of feedback information, occurrence time and clustering characteristics, and the feedback information includes supplementary information indicating the label data and / or confidence of the first business data. The occurrence time is the time when the first business data occurs, and the clustering characteristics are used to characterize the category to which the first business data belongs in at least two clustering results formed by the multiple business data, and the first business data is any one of the accumulated multiple business data.
24. The device according to claim 23, characterized in that The target information includes the feedback information, the occurrence time and the clustering characteristics; for each of the accumulated multiple business data, the corresponding label data and confidence level are evaluated according to the target information, and the acquisition unit is specifically used to: evaluating first label data and a first confidence level corresponding to the first service data according to the feedback information; evaluating a second confidence level corresponding to the first business data according to the occurrence time; Evaluating a third confidence level corresponding to the first business data according to the clustering characteristics; The label data corresponding to the first business data is determined according to the first label data, and the confidence level corresponding to the first business data is determined according to the first confidence level, the second confidence level and the second confidence level.
25. The device according to claim 24, characterized in that The first label data is the sample type indicated by the feedback information. If the label of the first label data after evaluation according to the feedback information is not changed compared to the label of the first business data before evaluation according to the feedback information, then the first confidence is greater than the confidence corresponding to the first business data before evaluation according to the feedback information.
26. The device according to claim 24 or 25, characterized in that If the first label data corresponding to the first business data has not changed after evaluation through the feedback information compared to the label data corresponding to the first business data before evaluation, the longer the occurrence time of the first business data, the greater the second confidence corresponding to the generated first business data.
27. The device according to any one of claims 24 to 26, characterized in that There are multiple clustering results among the accumulated multiple business data. The third confidence corresponding to the forward business data that is closer to the cluster center in each clustering result is higher, and the third confidence corresponding to the reverse business data that is closer to the cluster center in each clustering result is lower. The forward business data is business data whose sample type is consistent with the sample type represented by the clustering result to which it belongs, and the reverse business data is business data whose sample type is inconsistent with the sample type represented by the clustering result to which it belongs. The sample type represented by the clustering result is the sample type corresponding to more business data in the clustering result.
28. The device according to any one of claims 21 to 27, characterized in that The device also includes: A determination unit, configured to determine second business data, wherein the second business data is business data whose label data changes after the re-evaluation compared with before the re-evaluation or whose confidence changes by an amount exceeding a reference threshold; In the aspect of obtaining a plurality of third sample data by screening the sample data in the historical training sample set through the sample elimination algorithm, the screening unit is specifically used for: A target deletion operation is performed on the sample data in the historical training sample set, and the target deletion operation includes: deleting the fourth sample data containing the second business data, and / or deleting the fifth sample data in the sample data cluster to which it belongs that has no new sample data for more than a preset time period; wherein the sample data remaining in the historical training sample set after performing the target deletion operation is the third sample data.
29. A data processing device, characterized in that: The device comprises at least one processor and at least one memory, wherein the at least one memory stores a computer program; when the computer program is executed by the processor, the device according to any one of claims 1 to 18 is implemented.
30. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed on a processor, implements the apparatus described in any one of claims 1 to 18.
Citation Information
Patent Citations
Training method and device of business prediction model
CN113537630A
Model training method and related device
CN114359635A
Illegal commercial tenant identification model construction method and device and illegal commercial tenant identification method
CN115330401A
Model training method and device, electronic equipment and storage medium
CN116166959A
Systems and Methods for Semi-Supervised Active Learning
US20220391765A1
Cited By
Intelligent equipment visual data management and AI model development platform and method
CN120356038A