Data processing method and device, storage medium and processor
By processing sample label data based on preset rules in credit scenarios and training the target model, the problem of low accuracy in credit scenarios is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202210130563.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-11
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-02-11
AI Technical Summary
In existing technologies, the classification performance of credit scenario models deteriorates due to differences in sample labels, resulting in low accuracy in predicting credit scenarios.
The target model is trained based on these labels by obtaining first sample label data and second sample label data from the original data according to the first preset rule, and selecting third sample label data that meets the second preset rule from the second sample label data. The data type is confirmed by processing multiple preset rules.
This improved the accuracy of predicting credit scenarios and ensured the accuracy of the model training data labels, thereby enhancing the model's classification performance.
Smart Images

Figure CN114462541B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data processing, in particular to a data processing method and device, a storage medium and a processor. BACKGROUND
[0002] Currently, in the suspicious case model training process, whether to report is used as the label of the suspicious case to train the model or judge the model effect. However, due to the difference between the sample labels, the classification effect of the model on the sample is poor, thereby there is a technical problem of low accuracy in the credit scene.
[0003] At present, no effective solution has been proposed for the above technical problem of low accuracy in predicting the credit scene. SUMMARY
[0004] The embodiments of the present application provide a data processing method, device, storage medium and processor to at least solve the technical problem of low accuracy in predicting the credit scene
[0005] According to one aspect of the embodiments of the present application, a data processing method is provided, comprising: obtaining first sample label data and second sample label data in first original data, wherein the first sample label data is sample data satisfying a first preset rule, and the second sample label data is sample data not satisfying the first preset rule; obtaining third sample label data in the second sample label data, wherein the third sample label data is sample data satisfying a second preset rule selected from the second sample label data; and determining a target sample label based on the first sample label data and the third sample label data, wherein the target sample label is used to train a target model.
[0006] Optionally, first approval data of the first original data is obtained, wherein the first approval data is data of the last approval of the first original data; and the first sample label data is obtained in the first original data, comprising: marking the first original data satisfying the first preset rule of the first approval data to obtain the first sample label data, wherein the first original data satisfying the first preset rule of the first approval data is suspicious data.
[0007] Optionally, the second sample label data is obtained in the first original data, comprising: marking the first original data not satisfying the first preset rule of the first approval data to obtain the second sample label data, wherein the first original data not satisfying the first preset rule of the first approval data is trusted data.
[0008] Optionally, a keyword of historical data and a data type of the historical data are determined, wherein the historical data is data in a database and comprises the first original data; and the keyword and the data type are fitted and iterated to obtain the second preset rule.
[0009] Optionally, the data types are different, and the corresponding keywords are different.
[0010] Optionally, based on the first sample label data and the third sample label data, the target sample label is determined, wherein the target sample label is used to train the target model, and the data processing method further includes: the sample label of the first sample label data is the same as the sample label of the third sample label data, and the target sample label is an actual label of the sample; and the sub-model is trained based on the sample label to obtain the target model.
[0011] Optionally, the prediction label of the first original data is determined; and the sub-model is trained based on the prediction label of the first original data and the target sample label to obtain the target model.
[0012] Optionally, the prediction label of the first original data is determined, and the method further includes: the prediction label of the first original data is determined based on the feature data of the first original data.
[0013] Optionally, in the first original data, the fourth sample label data and the fifth sample label data are obtained, wherein the fourth sample label data is sample data that meets at least one of the following conditions: the first preset rule, the second preset rule, and the third preset rule; and the fifth sample label data is sample data that meets at least one of the following conditions: the first preset rule, the second preset rule, and the third preset rule.
[0014] According to another aspect of the embodiments of the present application, a data processing apparatus is further provided. The apparatus includes: a first obtaining unit, configured to obtain, in first original data, first sample label data and second sample label data, wherein the first sample label data is sample data that meets a first preset rule, and the second sample label data is sample data that does not meet the first preset rule; a second obtaining unit, configured to obtain, in the second sample label data, third sample label data, wherein the third sample label data is sample data that meets a second preset rule and is selected from the second sample label data; and a determining unit, configured to determine, based on the first sample label data and the third sample label data, a target sample label, wherein the target sample label is used to train a target model.
[0015] According to another aspect of the embodiments of the present application, a computer readable storage medium is further provided. The computer readable storage medium includes a stored program, wherein the program, when running, controls a device where the computer readable storage medium is located to perform the data processing method of the embodiments of the present application.
[0016] According to another aspect of the embodiments of the present application, a processor is further provided. The processor is used to run a program, wherein the program, when running, performs the data processing method of the embodiments of the present application.
[0017] In the embodiment of the present application, in the first original data, the first sample label data and the second sample label data are obtained, wherein the first sample label data is sample data satisfying the first preset rule, and the second sample label data is sample data not satisfying the first preset rule; the third sample label data is obtained from the second sample label data, wherein the third sample label data is sample data satisfying the second preset rule selected from the second sample label data; and the target sample label is determined based on the first sample label data and the third sample label data, wherein the target sample label is used to train the target model. That is, the first sample label data and the second sample label data are obtained by processing the first original data in the first approval data based on the first preset rule, the third sample label data is obtained by processing the second sample label data based on the second preset rule, and the first original data type is accurately confirmed through multiple preset rule processing, thereby realizing the technical effect of improving the prediction credit scene accuracy and solving the technical problem of low prediction credit scene accuracy. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are included to provide a further understanding of the present application and constitute a part of this application, illustrate embodiments of the present application and together with the description serve to explain the present application. In the drawings:
[0019] Figure 1 is a flowchart of a data processing method according to an embodiment of the present application;
[0020] Figure 2 is a flowchart of a data flow according to a related art;
[0021] Figure 3 is a flowchart of a data flow of total case counting keywords and extraction rules according to an embodiment of the present application;
[0022] Figure 4 is a flowchart of a case data flow of original label 0 according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of a data processing device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part but not all of the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work should belong to the protection scope of the present application.
[0025] It should be noted that the terms "first", "second" and the like in the description and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a list of steps or units need not be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to such processes, methods, products or devices.
[0026] Embodiment 1
[0027] According to an embodiment of the present application, a method embodiment of data processing is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0028] Figure 1 is a flowchart of a method of data processing according to an embodiment of the present application. As shown in Figure 1 , the method comprises the following steps:
[0029] Step S102, in the first original data, the first sample label data and the second sample label data are obtained, wherein the first sample label data is the sample data satisfying the first preset rule, and the second sample label data is the sample data not satisfying the first preset rule.
[0030] In the technical solution provided in the step S102 of the present application, the first original data is processed based on the first preset rule, the first original data meeting the first preset rule is marked to obtain first sample label data, and the first original data not meeting the first preset rule is marked to obtain second sample label data, wherein the first original data can be case data in a period of time, the first sample label data can be positive sample label data, i.e., suspicious case data with a label of 1, and the second sample label data can be negative sample label data, i.e., trusted case data with a label of 0.
[0031] Optionally, the first preset rule can be a rule set according to actual needs, i.e., a case reporting rule. The case data meeting the first preset rule (the case reporting rule) can be regarded as positive sample data, and the case label is marked as 1. The case data not meeting the first preset rule can be regarded as negative sample data, and the case label is marked as 0.
[0032] In the step S104, third sample label data is obtained from the second sample label data, wherein the third sample label data is sample data meeting a second preset rule selected from the second sample label data.
[0033] In the technical solution provided in the step S104 of the present application, the second sample label data meeting the second preset rule is selected from the second sample label data based on the second preset rule to obtain the third sample label data. The label of the second sample label data meeting the second preset rule can be modified to obtain the third sample label data, wherein the third sample label data can be sample data with a sample label of 1. The second preset rule can be a suspicious unreported case rule, which can be used to determine whether the second sample label data is positive sample data, wherein the positive sample data can be used to represent suspicious case data.
[0034] Optionally, the second sample label data is processed based on the second preset rule to obtain third sample label data meeting the second preset rule. For example, when the second sample label data meets the second preset rule, it can be indicated that the second sample label data meets the suspicious unreported case rule. The label of the second sample label data can be modified to 1 to obtain third sample label data with a sample label of 1.
[0035] In the step S106, a target sample label is determined based on the first sample label data and the third sample label data, wherein the target sample label is used to train a target model.
[0036] In the technical solution provided in the step S106 of the present application, in the first original data, the first sample label data is obtained based on the first preset rule, in the second sample label data, the third sample label data is obtained based on the second preset rule, and the target sample label is determined based on the first sample label data and the second sample label data. The target label is used to train the sub-model to obtain the target model. The target sample label can be a positive sample label with a label of 1, or a negative sample label with a label of 0. It can be a sample label obtained after processing the original data, that is, it can be the actual sample label of the first original data.
[0037] Optionally, in the supervised model, the model evaluation index can be calculated based on the target sample label to evaluate the model. The model is adjusted based on the evaluation result to obtain the target model.
[0038] The steps S102 to S106 of the present application, in the first original data, the first sample label data and the second sample label data are obtained. The first sample label data is the sample data satisfying the first preset rule, and the second sample label data is the sample data not satisfying the first preset rule. The third sample label data is obtained in the second sample label data. The third sample label data is the sample data satisfying the second preset rule selected from the second sample label data. The target sample label is determined based on the first sample label data and the third sample label data. The target sample label is used to train the target model. That is, the first original data is processed based on the first preset rule to obtain the first sample label data and the second sample label data. The second sample label data is processed based on the second preset rule to obtain the third sample label data. After multiple preset rule processing, the type of the first original data is accurately confirmed, and the technical effect of improving the prediction accuracy of the credit scenario is realized, and the technical problem of low prediction accuracy of the credit scenario is solved.
[0039] The above method of the embodiment will be further introduced below.
[0040] As an optional implementation, the first approval data of the first original data is obtained, wherein the first approval data is the data of the last approval of the first original data. In the first original data, the first sample label data is obtained, including: marking the first original data satisfying the first preset rule of the first approval data to obtain the first sample label data, wherein the first original data satisfying the first preset rule of the first approval data is suspicious data.
[0041] In this embodiment, first approval data of the first original data is obtained, first original data satisfying the first preset rule is marked based on the first preset rule, and first sample label data is obtained. The first original data satisfying the first preset rule can be suspicious data, the suspicious data can be reported case data, and the suspicious data can be represented by label 1. The first approval data can be the last approval data of the first original data. Due to process return and other reasons, there can be multiple approval data of the first original data. The last approval data can be selected according to the date of generating the approval data.
[0042] Optionally, in the multiple approval data of the first original data, the last approval data is selected according to the date of generating the approval data to obtain the first approval data. The first approval data of the first original data satisfying the first preset rule is marked based on the first preset rule, and the first sample label data allowed to be reported can be obtained, that is, the sample label data marked as 1.
[0043] As an optional implementation, in the first original data, the second sample label data is obtained, including: marking the first original data of the first approval data not satisfying the first preset rule to obtain the second sample label data, wherein the first original data of the first approval data not satisfying the first preset rule is trusted data.
[0044] In this embodiment, the first original data of the first approval data not satisfying the first preset rule is marked based on the first preset rule to obtain the second sample label data. The first original data of the first approval data not satisfying the first preset rule can be trusted data. The trusted data can be non-suspicious sample label data. The second sample label data can be unreported data, which can be sample label data marked as 0.
[0045] As an optional implementation, the keyword of the historical data and the data type of the historical data are determined, wherein the historical data is data in a database and includes the first original data; the keyword and the data type are fitted and iterated to obtain the second preset rule.
[0046] In this embodiment, the keyword of the historical data and the data type of the historical data can be the keyword extracted from the suspicious unreported case and the type of the historical data. The fuzzy matching rule is constructed, the fuzzy matching results of the keyword of the historical data and the data type are fitted and iterated, and the second preset rule is obtained. The data type of the historical data can be the case type of the historical data. The second preset rule can be used to determine whether the second sample label data is positive sample data. The historical data can include suspicious cases and trusted cases.
[0047] Optionally, the data type corresponding to the keyword of "confirming reporting" is a positive sample, for example, the keyword of "confirming reporting" exists in the historical data, the label of the corresponding case is modified to 1; the historical data exists "no money laundering risk", which means that the historical data is a trusted case, and the label of the corresponding case is modified to 0.
[0048] As an optional implementation, the corresponding data type is different, and the corresponding keyword is different.
[0049] In this embodiment, according to the data type of the case, the corresponding keyword can be extracted, and the corresponding keyword is different when the corresponding data type is different, for example, the keywords corresponding to the reporting case and the non-reporting case cannot be "confirming reporting".
[0050] As an optional implementation, based on the first sample label data and the third sample label data, the target sample label is determined, wherein the target sample label is used to train the target model, and the method further comprises: the sample label of the first sample label data and the sample label of the third sample label data are the same, and the target sample label is the actual label of the sample; and the sub-model is trained based on the sample label to obtain the target model.
[0051] In this embodiment, the first sample label data is obtained based on the first preset rule, the third sample label data is obtained based on the second preset rule, the sample label of the first sample label data and the sample label of the third sample label data are the same, and the sample label data with label 1 is obtained based on the sample label, and the sub-model is trained based on the sample label to obtain the target model, wherein the target sample label is the actual label of the sample, which can include a positive sample label with label 1 or a negative sample label with label 0, and the sub-model is trained based on the sample label to obtain the target model.
[0052] As an optional implementation, the prediction label of the first original data is determined; and the sub-model is trained based on the prediction label of the first original data and the target sample label to obtain the target model.
[0053] In this embodiment, the target sample label of the first original data is obtained based on the first sample label data and the second sample label data, the actual label of the first original data is compared with the target sample label, if they are consistent, the sub-model is trained based on the sample label data to obtain the target model, wherein the actual label of the first original data can be a positive sample label with label 1 or a negative sample label with label 0.
[0054] As an optional implementation, the prediction label of the first original data is determined, and the method further comprises: determining the prediction label of the first original data based on the feature data of the first original data.
[0055] In this embodiment, feature data of the first raw data is determined, and a predicted label of the first raw data is determined based on the feature data of the first raw data, wherein the feature data can be user age, gender, cumulative transaction amount in a period of time, number of transactions, and the like.
[0056] Optionally, the first raw data is processed, feature data in the first raw data is extracted, and a predicted label of the first raw data is determined based on the feature data in the first raw data, for example, when the cumulative transaction amount in a period of time in the feature data exceeds a set threshold value, it can be judged that the first raw data is suspicious data, that is, a positive sample data with a label of 1.
[0057] Optionally, the target model can be a supervised model, for the supervised model, before the model is trained, based on fitting and iteration of the keyword and the data type, a second preset rule is obtained, the second preset rule can be input into the system in the form of code in the terminal, the first raw data is judged based on the first preset rule and the second preset rule, and a target sample label, that is, an actual label of the sample, is obtained; when the model is trained, the first raw data is input into the sub-model, a predicted label of the first raw data is obtained based on the feature data; when the model is evaluated, the target sample label and the predicted label are used together to calculate a model evaluation index, evaluate the model, and optimize the model to obtain the target model.
[0058] Optionally, the target model can also be an unsupervised model, for the unsupervised model, the target sample label and the predicted label are only used to evaluate the model, and do not participate in the model training process.
[0059] As an optional implementation, in the first raw data, fourth sample label data and fifth sample label data are obtained, wherein the fourth sample label data is at least one of the following sample data: sample data satisfying the first preset rule, sample data satisfying the second preset rule, and sample data not satisfying the third preset rule; and the fifth sample label data is at least one of the following sample data: sample data not satisfying the first preset rule, sample data not satisfying the second preset rule, and sample data satisfying the third preset rule.
[0060] In this embodiment, the sample data satisfying at least one of the following rules is obtained in the first original data, which can be sample data satisfying a first preset rule, sample data satisfying a second preset rule, sample data not satisfying a third preset rule, the obtained sample data is taken as fourth sample label data, sample data satisfying at least one of the following rules is obtained, which can be sample data satisfying the third preset rule, sample data not satisfying the first preset rule, sample data not satisfying the second preset rule, and the obtained sample data is taken as fifth sample label data, wherein the fourth sample label data can be positive sample label data, can be a reported case, and can be a suspicious case; the fifth sample label data can be negative sample label data, can be an unreported case, and can be a trusted case.
[0061] Optionally, in the first original data, the fourth sample label data can be obtained by marking at least one of the following rules in the first original data: sample data satisfying a first preset rule, sample data satisfying a second preset rule, and sample data not satisfying a third preset rule, to obtain the fourth sample label data.
[0062] Optionally, in the first original data, the fifth sample label data can be obtained by marking at least one of the following rules in the first original data: sample data not satisfying the first preset rule, sample data not satisfying the second preset rule, and sample data satisfying the third preset rule, to obtain the fifth sample label data.
[0063] Optionally, the fourth sample label data and the fifth sample label data can be obtained in the first original data by processing the first original data based on the first preset rule, the second preset rule, and the third preset rule, or by processing the first original data based on the first preset rule first, then based on the second preset rule, and finally based on the third preset rule. It should be noted that this is only an example and is not limited in detail.
[0064] Optionally, the fourth sample label data and the fifth sample label data can be obtained in the first original data by matching the first original data based on the first preset rule, if the matching fails, matching the first original data based on the second preset rule, if the matching succeeds, no further matching is needed, and the fourth sample label data and / or the fifth sample label data marked with the target sample label are obtained.
[0065] Optionally, the fourth sample label data and the fifth sample label data can be obtained in the first original data by matching the first original data based on the second preset rule, if the matching fails, matching the first original data based on the third preset rule, if the matching succeeds, no further matching is needed, and the fourth sample label data and / or the fifth sample label data marked with the target sample label are obtained.
[0066] The embodiment processes the first approval data in the first original data based on a first preset rule to obtain first sample label data and second sample label data, processes the second sample label data based on a second preset rule to obtain third sample label data, and accurately confirms the type of the first original data through multiple preset rule processing, thereby achieving the technical effect of improving the prediction credit scene accuracy and solving the technical problem of low prediction credit scene accuracy.
[0067] Embodiment 2
[0068] The technical solutions of the embodiments of the application will be described below in conjunction with preferred embodiments.
[0069] In the existing anti-money laundering suspicious case model training process, whether supervised or unsupervised model, in the process of training the model, the training data is the suspicious case generated by anti-money laundering, and whether to report is taken as the label of the suspicious case, and then the model is trained or the model effect is judged.
[0070] In the related art, because the sample labels are different, the classification effect of the model on the samples is poor. Figure 2 It is a flowchart according to a data flow in the related art, as shown in Figure 2 The data flow can include:
[0071] Step S201, obtaining database information.
[0072] The database information is extracted in the storage device or server, wherein the database information includes case information to be analyzed.
[0073] Step S202, obtaining case data information.
[0074] The case information is obtained from the database, and the case information is further analyzed.
[0075] Step S207, judging whether to report.
[0076] According to whether the case rule is reported, the obtained case data is processed, the reported case is marked as a positive sample label, marked as 1, and the unreported case is marked as a negative sample label, marked as 0.
[0077] Step S208, obtaining a label.
[0078] After the case data is marked, the label corresponding to the case data is extracted to obtain the actual label of the case.
[0079] Step S203, obtaining transaction data information.
[0080] Obtain transaction data of the case from the database, wherein the transaction data information can include age of a person, transaction details in a period of time, such as a transaction time point, a transaction amount, a balance after the transaction, a transaction mode, and the like, which are not limited herein.
[0081] Step S204: analyze and extract features.
[0082] The transaction data information is analyzed to extract keywords and identify factors that can determine whether the case is suspicious.
[0083] Optionally, the feature data information can include age of a person, cumulative transaction amount in a period of time, and number of transactions, which are only used as examples and are not limited herein.
[0084] Step S206: obtain feature data.
[0085] Based on the analysis result of the transaction data, the feature data is obtained, wherein the feature data should be a static attribute of the transaction or the customer.
[0086] Optionally, the feature data can be a customer feature, such as age, gender, and the like, which are not limited herein.
[0087] Step S209: model training.
[0088] The predicted label of the obtained case data is compared with the actual label to evaluate the model effect, and if the effect is poor, that is, if the predicted label and the actual label are too different, the model parameters are adjusted and trained.
[0089] In the related art steps S201 to S209, although some cases are actually suspicious, they are not reported because a customer does not report repeatedly in a period of time, and according to the above method, the case is determined as a negative sample.
[0090] Optionally, in the model parameter adjustment and training stage, the model effect is judged according to the difference between the predicted label of the model for the verification data and the actual label of the verification data after each training is completed, to determine whether the current model can well perform prediction. If there is a deviation in the actual label of the training data, a good model will be judged, and the result is not accurate when the new data is predicted.
[0091] To address the aforementioned issues, this invention proposes a method for optimizing training data for a suspicious case model. By organizing the actual handling opinions of all cases, common fields in the handling opinions of truly suspicious cases are identified, such as "confirmed suspicious" and "reported." During model training, initial labels are generated based on whether a report was submitted. Then, the handling opinions of suspicious cases are retrieved to update the labels of cases in the training data. The new labels are used to train the model or evaluate its performance. This ensures that the labels in the training data are correct, resulting in better model discrimination.
[0092] According to this embodiment of the present invention, a method for optimizing training data for a suspected case model is proposed, which may include the following.
[0093] Step 1: Acquiring case data.
[0094] Retrieve the case number, approval comments, and whether it has been reported from the database. The case number must be a unique value, and it should only be retrieved once for the same case. Only the final approval comment should be retrieved. For cases with multiple approval comments due to process reversals or other reasons, the last approval comment should be selected based on the date the approval comment was generated.
[0095] Step 2: Generate initial tags.
[0096] Determine whether the last approval opinion was to be submitted. Cases with the last approval opinion of submission are treated as positive samples and labeled as 1, while cases with the last approval opinion of non-submission are treated as negative samples and labeled as 0, thus obtaining the initial label for the case.
[0097] Optionally, the cases can be: reported cases with an initial label of 1; suspected unreported cases with an initial label of 0 that were not reported due to the reason of not reporting repeatedly; and credible cases with an initial label of 0, that is, cases that are confirmed to be credible and have not been reported.
[0098] Step 3: Analyze the case and proceed with the next steps.
[0099] When analyzing case approval opinions and extracting keywords, the following should be noted: keywords should not be repeated between different types of cases; rules for reported and unreported cases should be separated, and the corresponding rules should be executed according to the initial case label.
[0100] The analysis of the obtained approval opinions can include the following two situations.
[0101] The first approach applies to all cases.
[0102] like Figure 3 As shown, Figure 3 This is a flowchart of the data flow of total keywords and extraction rules for all cases according to an embodiment of the present invention, which may include...
[0103] Step S301, obtain case data.
[0104] The case data can be approval information or case-related information of a case used for model training.
[0105] Step S302, process the case data according to the reported case rule, the suspicious unreported case rule, and the trusted case rule.
[0106] The reported case rule, the suspicious unreported case rule, and the trusted case rule can be obtained by analyzing the approval opinion, using keywords to build fuzzy matching rules, and then extracting case keywords and matching the case data based on the reported case rule, the suspicious unreported case rule, and the trusted case rule. Alternatively, when processing the case, the three reported case rules can be processed simultaneously or sequentially. For example, the case can be matched based on the reported case rule first, and if the matching fails, the case can be matched based on the suspicious unreported case rule, and if the matching fails, the case can be matched based on the trusted unreported case rule, so as to achieve the purpose of screening suspicious unreported cases.
[0107] Alternatively, building fuzzy matching rules can include building fuzzy matching rules based on the extracted keywords and the type of the case. For example, if the approval opinion contains the keyword "confirm reporting", the type of the corresponding case is a reported case; if the approval opinion contains the keyword "no money laundering risk", the type of the corresponding case is an unreported case.
[0108] Alternatively, the keywords cannot be repeated between different types of cases, and the rules of reported and unreported cases need to be separated, and the corresponding rules are executed according to the initial label of the case.
[0109] Alternatively, this method requires analysis of all cases, which is a large amount of analysis, and each type of case needs to summarize keywords, and attention should be paid to whether the keywords are repeated.
[0110] Step S303, output sample labels.
[0111] The matching result of the case based on the case reporting rule outputs the corresponding sample label, wherein the sample label can include a positive sample label, which can be marked as 1, representing a suspicious case, and a negative sample label, which can be marked as 0, representing a trusted case.
[0112] Optionally, cases can be matched based on the rules for reported cases. If the match is successful, the sample is a positive sample and the case is a suspicious case. If the match fails, the sample is a negative sample and the case is a trustworthy case. Trustworthy cases can be matched based on the rules for suspicious unreported cases. If the match is successful, the sample is a positive sample, i.e., the case is a suspicious case, and the case label is changed to a positive sample label. If the match fails, the sample is a negative sample and the case is a trustworthy case. Cases can also be matched based on the rules for trustworthy unreported cases. If the match is successful, the sample is a negative sample and the case is a trustworthy case. If the match fails, the case is a suspicious case, and the case label is determined to be a positive sample label.
[0113] Optionally, the reporting case rules involved in the technical solution of the present invention are only used to modify sample labels and are not used to process cases.
[0114] Optionally, all data in this invention is processed offline without modifying the original database.
[0115] The second type is for cases that have not been reported.
[0116] like Figure 4 As shown, Figure 4 This is a flowchart illustrating the data flow of a case with an initial label of 0 according to an embodiment of the present invention. The data flow may include:
[0117] Step S401: The initial label is a negative sample.
[0118] Initial labels are assigned based on whether the original approval opinion was reported to the case information, and negative samples are obtained based on the initial labels.
[0119] Step S402, Suspicious Unreported Case Rules. Cases with an initial negative label are processed based on the suspicious unreported case rules. For example, case information can be matched with the suspicious unreported case rules, and the actual label of the case can be determined based on the matching result.
[0120] Step S403: Determine if it is suspicious.
[0121] Based on the rule of suspicious unreported cases, cases with negative initial labels are processed to determine whether the cases are suspicious.
[0122] For example, the case information is matched with the rules for suspicious unreported cases to determine whether the match is successful. If the match is successful, the actual label of the case is a positive sample label, and the initial label of the case is changed to 1, that is, the case is a suspicious case. If the match fails, the case is not modified, and the actual label is a negative sample label, that is, the case is a trustworthy case.
[0123] Step S404: Obtain the final label.
[0124] Based on the suspicious unreported case rule, the case with the initial label of negative sample is judged. If it is credible, the initial label is not modified. If it is suspicious, the initial label is modified, that is, the negative sample is modified to a positive sample label, and the final label is obtained.
[0125] In this embodiment, through steps S401 to S404, only the case with the original label of 0 (unreported case) is analyzed, and the keyword of the suspicious but unreported case is extracted. The credible and unreported case does not need to extract the keyword.
[0126] Optionally, this method only needs to analyze most of the cases, and the amount of analysis is reduced compared with the first method; only the case that is actually suspicious but not reported needs to be focused on, and only the keyword of this case needs to be extracted, so it is easier to generate the keyword, and the problem of keyword duplication does not need to be considered.
[0127] For the unsupervised model, the modified label is only used together with the actual label and the model prediction label to calculate the model evaluation index when the model is evaluated, to evaluate the model.
[0128] For the supervised model, the actual label needs to be used when training the model using the training set data. The label and feature data obtained based on the keyword and data type matching are transmitted into the model, and the model is fitted and iterated according to the label and feature data. When the model is evaluated, the label and feature data are used together to calculate the model evaluation index, to obtain the target sample label, and the model is evaluated based on the target sample label (actual label) and the prediction label.
[0129] Optionally, in this embodiment, the suspicious unreported case can be separately regarded as a category, and the model can be set as a multi-classification model during the training of the model, that is, the model can divide the case into a reported case, an unreported case and a suspicious unreported case. For example, the model can distinguish the cases by marking them with labels, such as marking the reported case as 1, marking the unreported case as 0, and marking the suspicious unreported case as 2.
[0130] In this embodiment, by sorting the actual handling opinions of all cases, the same fields (such as: confirming suspiciousness, reporting, etc.) in the truly suspicious case handling opinions are found out. After the initial label is generated according to whether the case is reported, the handling opinion of the suspicious case is taken, fuzzy matching and the fields in the truly suspicious case handling opinion are used for matching, the label of the case in the training data is updated, and the model is trained using the new label or the effect of the model is judged, thereby solving the technical problem of low accuracy in the prediction credit scene and achieving the technical effect of improving the accuracy in the prediction credit scene.
[0131] Embodiment 3
[0132] According to an embodiment of the present application, a data processing apparatus is also provided. It should be noted that the data processing apparatus can be used to execute the method for data processing in embodiment 1.
[0133] Figure 5 is a schematic diagram of a data processing apparatus according to an embodiment of the present application. As shown in Figure 5 the data processing apparatus 500 can include a first obtaining unit 501, a second obtaining unit 502 and a determining unit 503.
[0134] The first obtaining unit 501 is configured to obtain first sample label data and second sample label data from the first original data, wherein the first sample label data is sample data satisfying a first preset rule, and the second sample label data is sample data not satisfying the first preset rule.
[0135] The second obtaining unit 502 is configured to obtain third sample label data from the second sample label data, wherein the third sample label data is sample data satisfying a second preset rule selected from the second sample label data.
[0136] The determining unit 503 is configured to determine a target sample label based on the first sample label data and the third sample label data, wherein the target sample label is used to train a target model.
[0137] Optionally, the apparatus can further include a third obtaining unit configured to obtain first approval data of the first original data, wherein the first approval data is data of the last approval of the first original data.
[0138] Optionally, the first obtaining unit 501 includes a first obtaining module configured to mark the first original data satisfying the first preset rule of the first approval data to obtain the first sample label data, wherein the first original data satisfying the first preset rule of the first approval data is suspicious data.
[0139] Optionally, the apparatus can further include a fifth obtaining unit configured to mark the first original data not satisfying the first preset rule of the first approval data to obtain the second sample label data, wherein the first original data not satisfying the first preset rule of the first approval data is trusted data.
[0140] Optionally, the apparatus can further include a first determining unit configured to determine a keyword of historical data and a data type of the historical data, wherein the historical data is data in a database and includes the first original data; and perform fitting iteration on the keyword and the data type to obtain the second preset rule.
[0141] Optionally, the determining unit 503 comprises a training module configured to: when the sample label of the first sample label data and the sample label of the third sample label data are the same, the target sample label is the actual label of the sample; and train the sub-model based on the sample label to obtain the target model.
[0142] Optionally, the training module comprises a second determining sub-module configured to determine the predicted label of the first original data; and train the sub-model based on the predicted label of the first original data and the target sample label to obtain the target model.
[0143] Optionally, the training module comprises a third determining sub-module configured to determine the predicted label of the first original data based on the feature data of the first original data.
[0144] Optionally, the apparatus further comprises a sixth obtaining unit configured to obtain, in the first original data, fourth sample label data and fifth sample label data, wherein the fourth sample label data is at least one of the following sample data: sample data satisfying the first preset rule, sample data satisfying the second preset rule, and sample data not satisfying the third preset rule; and the fifth sample label data is at least one of the following sample data: sample data not satisfying the first preset rule, sample data not satisfying the second preset rule, and sample data satisfying the third preset rule.
[0145] In the data processing apparatus of this embodiment, the first obtaining unit obtains, in the first original data, the first sample label data and the second sample label data, the second obtaining unit obtains, in the second sample label data, the third sample label data, and the determining unit determines the target sample label based on the first sample label data and the third sample label data, thereby achieving the technical effect of improving the prediction accuracy of the credit scenario and solving the technical problem of low prediction accuracy of the credit scenario.
[0146] Embodiment 4
[0147] According to the embodiments of the present application, a storage medium is also provided, which comprises a stored program, wherein the program, when executed by a processor, controls the device in which the computer readable storage medium is located to perform the method of data processing in Embodiment 1 of the present application.
[0148] Embodiment 5
[0149] According to the embodiments of the present application, a processor is also provided, which is used to run a program, wherein the program, when executed, performs the method of data processing described in Embodiment 1.
[0150] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0151] In the above-mentioned embodiments of the present application, the description of each embodiment is focused on, and the part not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0152] In several embodiments provided in the present application, it should be understood that the disclosed technical contents can be implemented by other ways. Among them, the above-mentioned device embodiments are only schematic, for example, the division of the units can be a logical function division, and in actual implementation, there can be another division way, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units or modules shown or discussed can be indirect coupling or communication connection through some interfaces, and can be electrical or other forms.
[0153] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed to a plurality of units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0154] In addition, each functional unit in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0155] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art that contributes or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0156] The above-mentioned is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be regarded as the protection scope of the present application.
Claims
1. A data processing method, characterized by, The method comprises: In the first original data, first sample label data and second sample label data are obtained, wherein the first sample label data is sample data satisfying a first preset rule, and the second sample label data is sample data not satisfying the first preset rule, and the first preset rule is a reported case rule; In the second sample label data, third sample label data is obtained, wherein the third sample label data is sample data satisfying a second preset rule selected from the second sample label data; Based on the first sample label data and the third sample label data, a target sample label is determined, wherein the target sample label is used to train a target model; Based on the target sample label, the target model is determined; The target model is used to determine the data type of the first original data in a credit scenario; The method further comprises: determining a keyword of historical data and a data type of the historical data, wherein the historical data is data in a database and comprises the first original data; based on the keyword and the data type, a fuzzy matching result between the keyword and the data type is determined; the fuzzy matching result is fitted and iterated to obtain the second preset rule, wherein the second preset rule is a suspicious unreported case rule and is used to determine whether the second sample label data is suspicious case data.
2. The method of claim 1, wherein, The method further comprises: First approval data of the first original data is obtained, wherein the first approval data is data of the last approval of the first original data; In the first original data, first sample label data is obtained, comprising: marking the first original data satisfying the first preset rule of the first approval data to obtain the first sample label data, wherein the first original data satisfying the first preset rule of the first approval data is suspicious data.
3. The method of claim 2, wherein, In the first original data, second sample label data is obtained, comprising: Marking the first original data not satisfying the first preset rule of the first approval data to obtain the second sample label data, wherein the first original data not satisfying the first preset rule of the first approval data is reliable data.
4. The method of claim 1, wherein, If the data types are different, the corresponding keywords are different.
5. The method of claim 1, wherein, Based on the first sample label data and the third sample label data, a target sample label is determined, wherein the target sample label is used to train the target model, and the method further comprises: The sample label of the first sample label data and the sample label of the third sample label data are the same, and the target sample label is the actual label of the sample; Based on the target sample label, a sub-model is trained to obtain the target model.
6. The method of claim 5, wherein the method further comprises: determining a predicted label of the first original data; Based on the predicted label of the first original data and the target sample label, the sub-model is trained to obtain the target model. 7. The method of claim 6, wherein, The determining the predicted label of the first original data comprises: determining the predicted label of the first original data based on the feature data of the first original data.
8. The method of claim 1, wherein, The method further comprises: In the first original data, fourth sample label data and fifth sample label data are obtained, wherein the fourth sample label data is sample data of at least one of the following: sample data satisfying the first preset rule, sample data satisfying the second preset rule, and sample data not satisfying the third preset rule; and the fifth sample label data is sample data of at least one of the following: sample data not satisfying the first preset rule, sample data not satisfying the second preset rule, and sample data satisfying the third preset rule.
9. A data processing apparatus, characterized by, Comprise: The first obtaining unit is configured to obtain first sample label data and second sample label data in first original data, wherein the first sample label data is sample data satisfying a first preset rule, and the second sample label data is sample data not satisfying the first preset rule, and the first preset rule is a reported case rule; The second obtaining unit is configured to obtain third sample label data in the second sample label data, wherein the third sample label data is sample data satisfying a second preset rule selected from the second sample label data; The determining unit is configured to determine a target sample label based on the first sample label data and the third sample label data, wherein the target sample label is used to train a target model; The first determining unit is configured to determine the target model based on the target sample label; The second determining unit is configured to determine a data type of the first original data in a credit scenario by using the target model; The device is further configured to: determine a keyword of historical data and a data type of the historical data, wherein the historical data is data in a database and comprises the first original data; determine a fuzzy matching result between the keyword and the data type based on the keyword and the data type; and perform fitting iteration on the fuzzy matching result to obtain the second preset rule, wherein the second preset rule is a suspicious unreported case rule and is used to determine whether the second sample label data is suspicious case data.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium comprises a stored program, wherein the program controls the device in which the computer-readable storage medium is located to perform the data processing method of any one of claims 1 to 8 when the program is run.
11. A processor, comprising: The processor is configured to run a program, wherein the program performs the data processing method of any one of claims 1 to 8 when the program is run by the processor.
Citation Information
Patent Citations
Information pushing method, device and system
CN108205766A