Data classification method and device, computer readable storage medium, and electronic device

By performing target aggregation and black sample marking on the sample data, the classification model is trained to determine the black and white samples in the sample data, which solves the problem of low classification accuracy of sample data in the prior art, and achieves higher classification accuracy and model stability.

CN114742152BActive Publication Date: 2025-05-16HANGZHOU FRAUDMETRIX TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210364684.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-05-16
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

In the prior art, manual analysis and judgment methods rely on expert experience and are highly subjective. Multi-channel risk summary methods introduce a large amount of noise in model training, resulting in poor accuracy in sample data classification.

Method used

By obtaining sample data, determining the training data based on the target aggregate data, labeling and model training of the black sample data, obtaining the classification model, and then determining the black sample data in the unlabeled sample data, using the target black sample and white sample data for model training, obtaining the target classification model.

Benefits of technology

It improves the accuracy of sample data, reduces noise during model training, and enhances the classification accuracy of the target classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114742152B_ABST
    Figure CN114742152B_ABST
Patent Text Reader

Abstract

The present disclosure is about a data classification method, a data classification device, a computer-readable storage medium and an electronic device, and relates to the field of computer technology. The method comprises: acquiring sample data, obtaining target aggregate data based on the sample data, and determining the training data included in the sample data according to the target aggregate data; marking the black sample data in the training data to obtain black sample data and unlabeled sample data, and using the black sample data to perform model training to obtain a classification model; determining the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data; using the target black sample data and the target white sample data to perform model training to obtain a target classification model, and classifying the sample data through the target classification model. The present disclosure improves the accuracy of sample data classification.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the emerging big data risk control system, each participant can make risk assessments on the users or enterprises included in the sample data based on the sample data they own. In related technologies, the sample data of the participants are usually classified using manual analysis methods or multi-channel risk aggregation methods.

[0003] In the manual analysis method, experts in related fields classify sample data based on their own experience. This method mainly relies on the experience of experts and is highly subjective. In addition, if experts only analyze the black sample data in the sample data, a series of statistical problems will be introduced, that is, the common problems of black sample data may be common problems of all sample data, resulting in poor accuracy in the classification of sample results. In multi-channel risk aggregation, each participant uses the black sample data and the remaining sample data to build a model, that is, it is considered that all the sample data in the remaining sample data are white sample data, which brings a lot of noise to the model training and reduces the accuracy of model classification.

[0004] Therefore, a new data classification method needs to be provided.

[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention

[0006] The purpose of the present invention is to provide a data classification method, a data classification device, a computer-readable storage medium and an electronic device, thereby at least to a certain extent overcoming the problem of low accuracy in classifying sample data due to the limitations and defects of related technologies.

[0007] According to one aspect of the present disclosure, a data classification method is provided, comprising:

[0008] Acquire sample data, obtain target aggregate data based on the sample data, and determine training data included in the sample data according to the target aggregate data;

[0009] Marking the black sample data in the training data to obtain black sample data and unlabeled sample data, and using the black sample data to perform model training to obtain a classification model;

[0010] Determine the black sample data included in the unlabeled sample data by using the classification model to obtain target black sample data and target white sample data;

[0011] The target black sample data and the target white sample data are used to perform model training to obtain a target classification model, and the sample data are classified using the target classification model.

[0012] In an exemplary embodiment of the present disclosure, acquiring sample data, and obtaining target aggregated data based on the sample data includes:

[0013] Acquire sample data, and perform homomorphic encryption on the sample data to obtain homomorphically encrypted sample data;

[0014] The target aggregate data obtained by performing sample alignment and aggregation on the homomorphically encrypted sample data of each participant is decrypted to obtain the target aggregate data.

[0015] In an exemplary embodiment of the present disclosure, marking the black sample data in the training data to obtain the black sample data and the unlabeled sample data includes:

[0016] Acquire black sample data included in the training data, mark the black sample data, and obtain black sample data;

[0017] The unlabeled sample data is obtained through the sample data other than the black sample data.

[0018] In an exemplary embodiment of the present disclosure, model training is performed using data corresponding to the training features included in the black sample data to obtain a classification model, including:

[0019] Randomly marking the sample data included in the black sample data as pseudo-white sample data;

[0020] Adding the pseudo-white sample data to the unlabeled sample data to obtain first sample data corresponding to the black sample and second sample data corresponding to the unlabeled sample;

[0021] Model training is performed using the first sample data and the second sample data to obtain a classification model.

[0022] In an exemplary embodiment of the present disclosure, determining the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data includes:

[0023] Performing multiple rounds of random marking of pseudo-white sample data in the black sample data, inputting each pseudo-white sample data into the classification model in turn, and obtaining a classification result of the pseudo-white sample data;

[0024] Determining a distribution threshold according to the classification result of the pseudo-white sample data;

[0025] The classification result of the unlabeled sample data is obtained through the classification model, and the target black sample data and the target white sample data are obtained according to the distribution threshold and the classification result of the unlabeled sample data.

[0026] In an exemplary embodiment of the present disclosure, obtaining the target black sample data and the target white sample data according to the distribution threshold and the classification result of the unlabeled sample data includes:

[0027] Obtaining a classification result of the unlabeled sample data, and comparing the classification result of the unlabeled sample data with the distribution threshold;

[0028] When it is determined that the classification result of the unlabeled sample data is less than the distribution threshold, marking the unlabeled sample data as white sample data;

[0029] When it is determined that the classification result of the unlabeled sample data is not less than the distribution threshold, marking the unlabeled sample data as black sample data;

[0030] The target black sample data and the target white sample data are obtained through the black sample data in the sample data and the black sample data and the white sample data included in the unlabeled sample data.

[0031] In an exemplary embodiment of the present disclosure, before using the target black sample data and the target white sample data for model training, the data classification method further includes:

[0032] Acquire the number of times each sample data in the unlabeled sample data is marked as black sample data, and acquire the sample data whose number of times being marked as black sample data is greater than a preset number;

[0033] According to the application scenario, the features included in the sample data that are marked as black sample data more than a preset number of times are analyzed to obtain features with low correlation with the application scenario, and the features with low correlation are deleted from the sample data to obtain the target features.

[0034] In an exemplary embodiment of the present disclosure, the target black sample data and the target white sample data are used to perform model training to obtain a target classification model, including:

[0035] Acquire third sample data corresponding to the target feature and included in the target black sample data, and fourth sample data corresponding to the target feature and included in the target white sample data;

[0036] The federal model is trained using the third sample data and the fourth sample data to obtain the target classification model.

[0037] According to one aspect of the present disclosure, there is provided a data classification device, comprising:

[0038] A training feature determination module, used to obtain sample data, obtain target aggregate data based on the sample data, and determine training data included in the sample data according to the target aggregate data;

[0039] A classification model training module, used for marking black sample data in the training data to obtain black sample data and unlabeled sample data, and performing model training using data corresponding to the training features included in the black sample data to obtain a classification model;

[0040] A sample data determination module, configured to determine the black sample data included in the unlabeled sample data by using the classification model, and obtain target black sample data and target white sample data;

[0041] The data classification module is used to perform model training using the target black sample data and the target white sample data to obtain a target classification model, and classify the sample data using the target classification model.

[0042] According to one aspect of the present disclosure, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the data classification method described in any one of the above is implemented.

[0043] According to one aspect of the present disclosure, there is provided an electronic device, including:

[0044] Processor; and

[0045] A memory, configured to store executable instructions of the processor;

[0046] The processor is configured to perform any one of the above-mentioned data classification methods by executing the executable instructions.

[0047] A data classification method provided by an embodiment of the present disclosure comprises the following steps: obtaining sample data, obtaining target aggregate data based on the sample data, and determining training data of the sample data according to the target aggregate data; marking black sample data in the training data to obtain black sample data and unlabeled sample data, and performing model training using data corresponding to the training features included in the black sample data to obtain a classification model; determining the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data; performing model training using the target black sample data and the target white sample data to obtain a target classification model, and classifying the sample data through the target classification model; on the one hand, obtaining target aggregate data , according to the target aggregate data, the training data participating in the model training in the sample data of the participants is determined, and the model is trained using the training data to obtain a classification model, thereby ensuring the security of the sample data of each participant and improving the classification accuracy of the classification model; on the other hand, the black sample data in the training data is marked, and the model is trained using the black sample data and the unlabeled sample data to obtain a classification model, and the black sample data included in the unlabeled sample data is determined through the classification model to obtain the target black sample and the target white sample, thereby solving the problem in the related technology that all sample data other than the black sample data is assumed to be white sample data, which brings a lot of noise to the model training, improves the accuracy of the sample data, reduces the noise in the model training, and improves the classification accuracy of the target classification model.

[0048] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present invention, and together with the specification are used to explain the principles of the present invention. Obviously, the accompanying drawings described below are only some embodiments of the present invention, and for those of ordinary skill in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0050] Figure 1 The following schematically shows a flow chart of a data classification method according to an exemplary embodiment of the present disclosure.

[0051] Figure 2 A block diagram of a data classification system according to an exemplary embodiment of the present disclosure is schematically shown.

[0052] Figure 3 A flowchart of a method for obtaining target aggregated data based on sample data according to an exemplary embodiment of the present disclosure is schematically shown.

[0053] Figure 4 A flowchart of a method for obtaining black sample data and unlabeled sample data according to an exemplary embodiment of the present disclosure is schematically shown.

[0054] Figure 5 A flowchart of a method for performing model training using the black sample data to obtain a classification model according to an exemplary embodiment of the present disclosure is schematically shown.

[0055] Figure 6 A flowchart of a method for obtaining target black sample data and target white sample data according to an exemplary embodiment of the present disclosure is schematically shown.

[0056] Figure 7 A flowchart of a method for obtaining target black sample data and target white sample data according to a distribution threshold and a classification result of unlabeled sample data according to an exemplary embodiment of the present disclosure is schematically shown.

[0057] Figure 8 A flowchart of a data classification method before model training using target black sample data and target white sample data according to an exemplary embodiment of the present disclosure is schematically shown.

[0058] Fig. 9 A flowchart of a method for performing model training using target black sample data and target white sample data to obtain a target classification model according to an exemplary embodiment of the present disclosure is schematically shown.

[0059] Fig.10 A block diagram schematically shows a data classification device according to an example embodiment of the present disclosure.

[0060] Fig.11 An electronic device for implementing the above data classification method according to an exemplary embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0061] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present invention will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present invention. However, those skilled in the art will appreciate that the technical solutions of the present invention may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present invention.

[0062] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0063] In the related technologies, sample data is classified through manual judgment or multi-channel risk aggregation. In the manual judgment method, experts in related fields mainly score sample data that are prone to risks based on their own experience, and finally classify the sample data owned by the participants based on the scores of each sample data; in the multi-channel risk aggregation method, machine learning methods are used to model the sample data of the participants, obtain the weights of each feature in the sample data, and then classify the sample data of the participants.

[0064] However, the above manual judgment method is extremely dependent on the experience of experts and is highly subjective, requiring a lot of adjustments and monitoring in the later stages of classification; and in the manual judgment method, when experts only analyze the black sample data in the sample data, they believe that the common problems of the black sample data are common problems of all samples, but in actual scenarios, such as the number of insured persons, in the overall sample data, the performance of this feature in the black sample data is no different from that in the white sample data, and cannot be used as one of the features for judging whether the sample data has risks. In the above multi-channel risk aggregation method, when each participant uses the sample to train the model, it assumes that the remaining sample data other than the black sample data are all white sample data, resulting in a large amount of noise in the model training process, which reduces the accuracy of the model in classifying the sample data.

[0065] Based on one or more of the above problems, this exemplary embodiment first provides a data classification method, which can be run on a device terminal, which can include a desktop computer, a portable computer, a smart phone, a tablet computer, etc. Of course, those skilled in the art can also run the method of the present invention on other platforms as required, and this exemplary embodiment does not specifically limit this. Figure 1 As shown, the data classification method may include the following steps:

[0066] Step S110. Acquire sample data, obtain target aggregate data based on the sample data, and determine training features of the sample data according to features included in the target aggregate data;

[0067] Step S120: marking the black sample data in the sample data to obtain the black sample data and the unmarked sample data, and performing model training using the data corresponding to the training features included in the black sample data to obtain a classification model;

[0068] Step S130. Determine the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data;

[0069] Step S140: Perform model training using the target black sample data and the target white sample data to obtain a target classification model, and classify the sample data using the target classification model.

[0070] The above-mentioned data classification method obtains sample data, obtains target aggregate data based on the sample data, and determines the training data of the sample data according to the target aggregate data; marks the black sample data in the training data to obtain black sample data and unlabeled sample data, and uses the data corresponding to the training features included in the black sample data to perform model training to obtain a classification model; determines the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data; uses the target black sample data and the target white sample data to perform model training to obtain a target classification model, and classifies the sample data through the target classification model; on the one hand, obtain the target aggregate data, and determine the training data according to the target aggregate data; on the other hand, ... The labeled aggregated data is used to determine the training data participating in the model training in the sample data of the participants, and the training data is used to train the model to obtain a classification model, thereby ensuring the security of the sample data of each participant and improving the classification accuracy of the classification model; on the other hand, the black sample data in the training data is labeled, and the model is trained using the black sample data and the unlabeled sample data to obtain a classification model, and the black sample data included in the unlabeled sample data is determined through the classification model to obtain the target black sample and the target white sample, thereby solving the problem in the related technology that all sample data other than the black sample data is assumed to be white sample data, which brings a lot of noise to the model training, thereby improving the accuracy of the sample data, reducing the noise in the model training, and improving the classification accuracy of the target classification model.

[0071] Hereinafter, each step involved in the data classification method of the exemplary embodiment of the present disclosure is explained and illustrated in detail.

[0072] First, the application scenarios and invention objectives of the exemplary embodiments of the present disclosure are explained and illustrated.

[0073] Specifically, the exemplary embodiments of the present disclosure may be used to implement risk control management, mainly to improve the accuracy of sample classification.

[0074] The exemplary embodiment of the present disclosure is based on the sample data of each participant, performs feature alignment on the encrypted sample data of each participant, obtains target aggregate data, and determines the training data of each participant for training the model according to the target aggregate data, thereby ensuring the data security of each participant; then, each participant marks the black sample data in the training data to obtain black sample data and unlabeled sample data, performs random marking in the black sample data to obtain pseudo-white sample data, performs model training through the black sample data, pseudo-white sample data and unlabeled data to obtain a classification model, judges the pseudo-white sample data through the classification model to obtain a classification result, and determines a distribution threshold according to the classification result of the pseudo-white sample; then, obtains the classification result of the unlabeled sample data through the classification model, determines the black sample data included in the unlabeled sample data according to the distribution threshold and the classification result of the labeled sample, and obtains the target black sample data and the target white sample data; finally, performs model training through the target black sample and the target white sample to obtain a target classification model, and classifies the sample data through the classification model, thereby improving the accuracy of classification.

[0075] Next, the data classification system involved in the exemplary embodiment of the present disclosure is explained and illustrated. Figure 2As shown, the data classification system may include: a participant 210 and a third party 220. The participant 210 may include a feature aggregation module 211, a sample marking module 212, a model training module 213 and a sample determination module 214; the participant 210 includes at least two participants, the feature aggregation module 211 is used for each participant to encrypt sample data, perform sample alignment on the encrypted sample data of each participant, obtain target aggregated data, each participant determines the training data participating in model training in the sample data according to the target aggregated data, uses the training data for training, obtains an intermediate result, and uses the intermediate result to obtain the gradient value of each participant; the sample marking module 212 is connected to the feature aggregation module 212 network, and is used to mark the black sample data in the training data to obtain black sample data and unmarked sample data. The sample data is recorded, and random marking is performed multiple times in the black sample data until each sample data in the black sample data is marked with pseudo-white sample data; the model training module 213 is connected to the sample marking module 212 network, and is used to use the black sample data, the pseudo-white sample data and the unlabeled data for model training to obtain a classification model, and the classification results of the pseudo-white sample data and the classification results of the unlabeled sample data are obtained through the classification model; the sample determination module 214 is connected to the model training module 213 network, and is used to determine the distribution threshold according to the classification results of the pseudo-white samples, and determine the black sample data included in the unlabeled samples according to the distribution threshold and the classification results of the unlabeled samples, and obtain the target black sample data and the target white sample data. After obtaining the target black sample data and the target white sample data, the model training module 213 uses the target black sample data and the target white sample data for model training to obtain the target classification model, and classifies the sample data through the target classification model. The third party 220 is used to send a key to each participant. After each participant trains the local model, a gradient value is obtained, the gradient value is encrypted, and the encrypted gradient value is sent to the third party 220. The third party 220 processes the encrypted gradient value to obtain a gradient value, and sends the obtained gradient value to each participant. Each participant updates the model parameters according to the gradient value.

[0076] The following will be combined Figure 2 Steps S110 to S140 are explained and described.

[0077] In step S110, sample data is acquired, target aggregate data is obtained based on the sample data, and training data included in the sample data is determined according to the target aggregate data.

[0078] Among them, the sample data is the sample data obtained by each participant. When obtaining the target aggregate data, the sample data of each participant can be encrypted, and the aggregate data is obtained according to the overlapping data samples in the sample data of each participant. Each participant decrypts the aggregate data to obtain the target aggregate data.

[0079] In this example embodiment, reference Figure 3 As shown, obtaining sample data and obtaining target aggregated data based on the sample data may include steps S310 and S320:

[0080] Step S310. Obtain sample data, and perform homomorphic encryption on the sample data to obtain homomorphically encrypted sample data;

[0081] Step S320. Perform sample alignment and aggregation on the homomorphically encrypted sample data of each participant, decrypt the target aggregate data, and obtain the target aggregate data.

[0082] In the following, step S310 and step S320 will be further explained and illustrated. Specifically, first, each participant obtains the corresponding sample data, performs homomorphic encryption on the obtained sample data, and obtains the homomorphic encrypted data of each participant. Since there is a large overlap in the data samples in the sample data of each participant, and the overlap of sample features is not high, the overlapping data samples in the sample data of each participant can be aggregated to obtain aggregated data. Each participant decrypts the aggregated data to obtain the target aggregated data. After obtaining the target aggregated data, each participant can obtain the training data for model training in its own sample data. Among them, in homomorphic encryption, the homomorphically encrypted data is processed to obtain an output, and this output is decrypted, and the result is the same as the output result obtained by processing the unencrypted original data in the same way. When homomorphically encrypting the sample data, additive homomorphic encryption can be used, multiplicative homomorphic encryption can be used, and subtractive homomorphic encryption can also be used. In this example embodiment, the homomorphic encryption method is not specifically limited.

[0083] For example, participant A can be a shopping mall, and participant B can be a bank. After obtaining the sample data of the participants, participants A and B perform homomorphic encryption on their own sample data. Since there is a large overlap in users between participants A and participants B, and the overlap of user features is not high, the sample data of the overlapping users of participants A and participants B can be aggregated to obtain aggregated data. After participants A and participants B obtain the aggregated data, they can decrypt the aggregated data to obtain target aggregated data. After each participant obtains the target aggregated data, the sample data included in the target aggregated data is the training data for subsequent model training of each participant.

[0084] By encrypting the sample data of each participant and aligning the encrypted data, the data security of each participant is guaranteed. By training the model with target aggregated data, a better training model is obtained, which improves the classification accuracy of the classification model.

[0085] In step S120, black sample data in the training data is marked to obtain black sample data and unmarked sample data, and the black sample data is used to perform model training to obtain a classification model.

[0086] In this example embodiment, reference Figure 4 As shown, marking the black sample data in the training data to obtain the black sample data and the unlabeled sample data may include step S410 and step S420:

[0087] Step S410: Obtain black sample data included in the training data, mark the black sample data, and obtain black sample data;

[0088] Step S420: Obtain the unlabeled sample data through the sample data other than the black sample data.

[0089] In the following, step S410 and step S420 will be further explained and illustrated. Specifically, after each participant determines the training data for model training included in the sample data according to the target aggregate data, the black sample data is marked in the training data, and the sample data other than the black sample data is the unmarked sample data.

[0090] After obtaining black sample data and unlabeled sample data, refer to Figure 5 As shown, using the black sample data to perform model training to obtain a classification model may include steps S510 to S530:

[0091] Step S510. Randomly mark the sample data included in the black sample data as pseudo-white sample data;

[0092] Step S520. Add the pseudo-white sample data to the unlabeled sample data to obtain first sample data corresponding to the black sample and second sample data corresponding to the unlabeled sample;

[0093] Step S530: Perform model training using the first sample data and the second sample data to obtain a classification model.

[0094] In the following, step S510-step S530 will be further explained and illustrated. Specifically, first, the sample data in the black sample data is randomly marked as pseudo-white sample data. In this example embodiment, the number of sample data randomly marked each time is not specifically limited; when the black sample data is P, the pseudo-white sample data is S, and the unlabeled sample data is U, the pseudo-white sample data S is added to the unlabeled sample data U, and the first sample data PS corresponding to the black sample data and the second sample data U+S corresponding to the unlabeled sample data can be obtained; then, the federated model is trained by the first sample data PS and the second sample data U+S to obtain the classification model.

[0095] In this example embodiment, pseudo-white sample data is obtained by randomly labeling black sample data, first sample data and second sample data are constructed by the black sample data, pseudo-white sample data and unlabeled sample data, and the model is trained by the first sample data and the second sample data, thereby reducing noise in model training and improving the classification accuracy of the classification model.

[0096] In step S130, the black sample data included in the unlabeled sample data is determined by the classification model to obtain target black sample data and target white sample data.

[0097] In this example embodiment, reference Figure 6 As shown, by using the classification model, determining the black sample data included in the unlabeled sample data to obtain target black sample data and target white sample data may include steps S610 to S630:

[0098] Step S610: Randomly mark the pseudo-white sample data in the black sample data, and input each pseudo-white sample data into the classification model in turn to obtain the classification result of the pseudo-white sample data;

[0099] Step S620. Determine a distribution threshold according to the classification result of the pseudo-white sample data;

[0100] Step S630: Obtain the classification result of the unlabeled sample data through the classification model, and obtain the target black sample data and the target white sample data according to the distribution threshold and the classification result of the unlabeled sample data.

[0101] In the following, step S610 to step S630 will be further explained and illustrated. Specifically, first, the pseudo-white sample data is randomly marked in the black sample data, and each pseudo-white sample data is input into the classification model in turn to obtain the classification result of each pseudo-white sample data, wherein the classification result of the pseudo-white sample is the probability of the pseudo-white sample, and after obtaining the classification result of each pseudo-white sample data in each round, the distribution threshold can be determined according to the classification results of all pseudo-white samples, and the black sample data included in the unlabeled sample data can be determined by the distribution threshold; after obtaining the distribution threshold, the unlabeled sample data is input into the classification model in turn to obtain the classification result of each unlabeled sample data, and the black sample data included in the unlabeled sample data is obtained according to the classification result of each unlabeled sample data and the distribution threshold.

[0102] For further reference, Figure 7 As shown, obtaining the target black sample data and the target white sample data according to the distribution threshold and the classification result of the unlabeled sample data may include steps S710 to S740:

[0103] Step S710. Obtain the classification result of the unlabeled sample data, and compare the classification result of the unlabeled sample data with the distribution threshold;

[0104] Step S720: When it is determined that the classification result of the unlabeled sample data is less than the distribution threshold, the unlabeled sample data is marked as white sample data;

[0105] Step S730: When it is determined that the classification result of the unlabeled sample data is not less than the distribution threshold, the unlabeled sample data is marked as black sample data;

[0106] Step S740: Obtain the target black sample data and the target white sample data through the black sample data in the sample data and the black sample data and the white sample data included in the unlabeled sample data.

[0107] In the following, step S710 to step S740 will be further explained and illustrated. Specifically, the classification result of the unlabeled sample data is compared with the distribution threshold, and when it is determined that the classification result of any sample data in the unlabeled sample is less than the distribution threshold, the unlabeled data can be marked as white sample data; when it is determined that the classification result of any sample data in the unlabeled sample data is greater than or equal to the distribution threshold, the unlabeled sample data can be marked as black sample data, and the target white sample data and the target black sample data are obtained through the marked black sample data and white sample data in the unlabeled sample data.

[0108] It should be noted that the black sample data can be subjected to multiple rounds of pseudo-white sample random marking until each sample data in the black sample data has been marked with pseudo-white sample data. Through multiple rounds of random marking, the unlabeled sample data is marked by the distribution threshold of each round, thereby improving the accuracy of the sample.

[0109] In this example embodiment, the distribution threshold is determined by a pseudo-white sample, and the sample data included in the unlabeled sample data is labeled by the distribution threshold, which solves the problem in the related art that all sample data other than black samples are assumed to be white samples, improves the accuracy of black and white samples, and further ensures the classification accuracy of the target classification model obtained by training the target black sample data and the target white sample data.

[0110] In step S140, the target black sample data and the target white sample data are used to perform model training to obtain a target classification model, and the sample data is classified by using the target classification model.

[0111] In this example embodiment, reference Figure 8 As shown, before using the target black sample data and the target white sample data for model training, the data classification method may further include step S810 and step S820:

[0112] Step S810. Obtain the number of times each sample data in the unlabeled sample data is marked as black sample data, and obtain sample data whose number of times being marked as black sample data is greater than a preset number;

[0113] Step S820. According to the application scenario, analyze the features included in the sample data that is marked as black sample data more than a preset number of times, obtain features with low correlation with the application scenario, delete the features with low correlation in the sample data, and obtain the target features.

[0114] In the following, step S810 and step S820 will be further explained and illustrated. First, since the black sample data is randomly marked for multiple rounds, the features in the sample data can be eliminated by the number of times each sample data in the unlabeled sample data is marked as black sample data in each round. Specifically, the number of times each sample data in the unlabeled sample is marked as black sample data is obtained. When the number of times any sample data is marked as a black sample is greater than a preset number, the sample data marked as black sample data greater than the preset number can be analyzed according to the actual application scenario, and the features with low correlation with the application scenario or no explanation in the sample data marked greater than the preset number are obtained, and the features are deleted from the sample data to obtain the target features. Among them, the preset number can be 5 times or 7 times. In this example embodiment, the preset number is not specifically limited.

[0115] For example, when sample data is obtained in which the number of times the sample data is marked as black sample data is greater than the preset number, the features in the sample data can be verified according to the actual application scenario. For example, when predicting the risk of corporate bankruptcy, it is found that the number of patents held by the company or the number of equipment owned by the company in the sample data is highly correlated with the company's bankruptcy. In this case, the features of the number of patents held by the company or the number of equipment owned by the company can be analyzed and judged, and the features of the number of patents held by the company or the number of equipment owned by the company can be deleted from the sample data.

[0116] After obtaining the target features, refer to Fig. 9 As shown, using the target black sample data and the target white sample data to perform model training to obtain a target classification model may include step S910 and step S920:

[0117] Step S910: Acquire the third sample data corresponding to the target feature included in the target black sample data and the fourth sample data corresponding to the target feature included in the target white sample data;

[0118] Step S920: Train the federation model using the third sample data and the fourth sample data to obtain the target classification model.

[0119] In the following, step S910 and step S920 will be further explained and illustrated. Specifically, first, the third sample data corresponding to the target feature in the target black sample data and the fourth sample data corresponding to the target feature included in the target white sample data are obtained; then, the third sample data and the fourth sample data are used to train the federation model to obtain the target classification model.

[0120] After obtaining the target classification model, each participant can classify the sample data according to the target classification model to obtain classification results, thereby improving the accuracy of sample data classification.

[0121] Furthermore, in actual usage scenarios, when a new participant is admitted, it is necessary to eliminate the features in the target features that do not exist or cannot be obtained by the admitted participant, obtain sample data from the admitted participant corresponding to the remaining features in the target features, perform model training again, and deploy the trained model online.

[0122] The data classification method and data classification system provided by the exemplary embodiments of the present disclosure have at least the following advantages: on the one hand, each participant performs homomorphic encryption on the sample data and performs feature alignment on the homomorphic encrypted sample data, thereby ensuring the data security of each participant, and obtaining a training model with better effect by training the model with the target aggregate data, thereby improving the classification accuracy of the classification model; on the other hand, the sample data in the black sample data is randomly labeled to obtain pseudo-white sample data, and the model is trained by the black sample data, the pseudo-white sample data and the unlabeled sample data to obtain a classification model, and the classification result of the pseudo-white sample data is obtained by the classification model, and a distribution threshold is determined according to the classification result, and the black sample data included in the unlabeled sample data is determined by the distribution threshold to obtain the target black sample data and the target white sample data, thereby improving the accuracy of the sample; on the other hand, the features in the sample data are eliminated to obtain the target sample data, and the model is trained by the third sample data and the third sample data corresponding to the feature data included in the target black sample data and the target white sample data to obtain the target training model, thereby improving the classification accuracy of the target classification model.

[0123] The exemplary embodiment of the present disclosure also provides a data classification device, referring to Fig.10 As shown, the data classification device may include: a training feature determination module 1010, a classification model training module 1020, a sample data determination module 1030 and a data classification module 1040. Among them:

[0124] A training feature determination module 1010 is used to obtain sample data, obtain target aggregate data based on the sample data, and determine training data included in the sample data according to the target aggregate data;

[0125] The classification model training module 1020 is used to mark the black sample data in the training data to obtain the black sample data and the unlabeled sample data, and perform model training using the data corresponding to the training features included in the black sample data to obtain the classification model;

[0126] A sample data determination module 1030 is used to determine the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data;

[0127] The data classification module 1040 is used to perform model training using the target black sample data and the target white sample data to obtain a target classification model, and classify the sample data using the target classification model.

[0128] In an exemplary embodiment of the present disclosure, acquiring sample data, and obtaining target aggregated data based on the sample data includes:

[0129] Acquire sample data, and perform homomorphic encryption on the sample data to obtain homomorphically encrypted sample data;

[0130] The target aggregate data obtained by performing sample alignment and aggregation on the homomorphically encrypted sample data of each participant is decrypted to obtain the target aggregate data.

[0131] In an exemplary embodiment of the present disclosure, marking the black sample data in the training data to obtain the black sample data and the unlabeled sample data includes:

[0132] Acquire black sample data included in the training data, mark the black sample data, and obtain black sample data;

[0133] The unlabeled sample data is obtained through the sample data other than the black sample data.

[0134] In an exemplary embodiment of the present disclosure, model training is performed using data corresponding to the training features included in the black sample data to obtain a classification model, including:

[0135] Randomly marking the sample data included in the black sample data as pseudo-white sample data;

[0136] Adding the pseudo-white sample data to the unlabeled sample data to obtain first sample data corresponding to the black sample and second sample data corresponding to the unlabeled sample;

[0137] Model training is performed using the first sample data and the second sample data to obtain a classification model.

[0138] In an exemplary embodiment of the present disclosure, determining the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data includes:

[0139] Performing multiple rounds of random marking of pseudo-white sample data in the black sample data, inputting each pseudo-white sample data into the classification model in turn, and obtaining a classification result of the pseudo-white sample data;

[0140] Determining a distribution threshold according to the classification result of the pseudo-white sample data;

[0141] The classification result of the unlabeled sample data is obtained through the classification model, and the target black sample data and the target white sample data are obtained according to the distribution threshold and the classification result of the unlabeled sample data.

[0142] In an exemplary embodiment of the present disclosure, obtaining the target black sample data and the target white sample data according to the distribution threshold and the classification result of the unlabeled sample data includes:

[0143] Obtaining a classification result of the unlabeled sample data, and comparing the classification result of the unlabeled sample data with the distribution threshold;

[0144] When it is determined that the classification result of the unlabeled sample data is less than the distribution threshold, marking the unlabeled sample data as white sample data;

[0145] When it is determined that the classification result of the unlabeled sample data is not less than the distribution threshold, marking the unlabeled sample data as black sample data;

[0146] The target black sample data and the target white sample data are obtained through the black sample data in the sample data and the black sample data and the white sample data included in the unlabeled sample data.

[0147] In an exemplary embodiment of the present disclosure, before using the target black sample data and the target white sample data for model training, the data classification method further includes:

[0148] Acquire the number of times each sample data in the unlabeled sample data is marked as black sample data, and acquire the sample data whose number of times being marked as black sample data is greater than a preset number;

[0149] According to the application scenario, the features included in the sample data that are marked as black sample data more than a preset number of times are analyzed to obtain features with low correlation with the application scenario, and the features with low correlation are deleted from the sample data to obtain the target features.

[0150] In an exemplary embodiment of the present disclosure, the target black sample data and the target white sample data are used to perform model training to obtain a target classification model, including:

[0151] Acquire third sample data corresponding to the target feature and included in the target black sample data, and fourth sample data corresponding to the target feature and included in the target white sample data;

[0152] The federal model is trained using the third sample data and the fourth sample data to obtain the target classification model.

[0153] The specific details of each module in the above-mentioned data classification device have been described in detail in the corresponding data classification method, so they will not be repeated here.

[0154] It should be noted that, although several modules or units of the equipment for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0155] In addition, although the steps of the method of the present invention are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.

[0156] In an exemplary embodiment of the present invention, an electronic device capable of implementing the above data conversion method is also provided.

[0157] It will be appreciated by those skilled in the art that various aspects of the present invention may be implemented as a system, method or program product. Therefore, various aspects of the present invention may be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as a "circuit", "module" or "system".

[0158] Reference below Fig.11 The electronic device 1100 according to this embodiment of the present invention is described. Fig.11 The electronic device 1100 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0159] like Fig.11As shown, the electronic device is in the form of a general computing device. The components of the electronic device may include, but are not limited to: the at least one processing unit 1110, the at least one storage unit 1120, a bus 1130 connecting different system components (including the storage unit 1120 and the processing unit 1110), and a display unit 1140.

[0160] The storage unit stores program codes, which can be executed by the processing unit 1110, so that the processing unit 1110 performs the steps according to various exemplary embodiments of the present invention described in the above “Exemplary Method” section of this specification. For example, the processing unit 1110 can perform the following steps: Figure 1 The step S110 shown in the figure is: acquiring sample data, obtaining target aggregate data based on the sample data, and determining the training data included in the sample data according to the target aggregate data; step S120: marking the black sample data in the training data to obtain black sample data and unlabeled sample data, and using the black sample data to perform model training to obtain a classification model; step S130: determining the black sample data included in the unlabeled sample data through the classification model to obtain target black sample data and target white sample data; step S140: performing model training with the target black sample data and the target white sample data to obtain a target classification model, and classifying the sample data through the target classification model.

[0161] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 11201 and / or a cache storage unit 11202 , and may further include a read-only storage unit (ROM) 11203 .

[0162] The storage unit 1120 may also include a program / utility 11204 having a set (at least one) of program modules 11205, such program modules 1105 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0163] Bus 1130 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0164] The electronic device 1100 may also communicate with one or more external devices 1200 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 1100, and / or communicate with any device that enables the electronic device 1100 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1150. Furthermore, the electronic device 1100 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 1160. As shown, the network adapter 1160 communicates with other modules of the electronic device 1100 via a bus 1130. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1100, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0165] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present invention.

[0166] In an exemplary embodiment of the present invention, a computer-readable storage medium is also provided, on which a program product capable of implementing the above method of this specification is stored. In some possible implementations, various aspects of the present invention can also be implemented in the form of a program product, which includes a program code, and when the program product is run on a terminal device, the program code is used to enable the terminal device to perform the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.

[0167] The program product for implementing the above method according to an embodiment of the present invention may adopt a portable compact disk read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto, and in this document, a readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, apparatus, or device.

[0168] The program product may use any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0169] Computer readable signal media may include data signals propagated in baseband or as part of a carrier wave, in which readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Readable signal media may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0170] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0171] Program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0172] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to an exemplary embodiment of the present invention, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.

[0173] Other embodiments of the invention will readily occur to those skilled in the art after considering the specification and practicing the invention invented herein. This application is intended to cover any variations, uses or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art that are not invented by the present invention. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the claims.

Claims

1. A data classification method, characterized in that: The method is used for risk control management, including: Acquire sample data, obtain target aggregate data based on the sample data, and determine training data included in the sample data according to the target aggregate data; the sample data includes users or enterprises; Marking the black sample data in the training data to obtain black sample data and unmarked sample data, randomly marking the sample data included in the black sample data as pseudo-white sample data, and using the black sample data for model training to obtain a classification model; Perform multiple rounds of random marking of pseudo-white sample data in the black sample data, input each pseudo-white sample data into the classification model in turn, and obtain the classification result of the pseudo-white sample data; determine the distribution threshold according to the classification result of the pseudo-white sample data; obtain the classification result of the unlabeled sample data through the classification model, and compare the classification result of the unlabeled sample data with the distribution threshold; when it is determined that the classification result of the unlabeled sample data is less than the distribution threshold, mark the unlabeled sample data as white sample data; when it is determined that the classification result of the unlabeled sample data is not less than the distribution threshold, mark the unlabeled sample data as black sample data; obtain target black sample data and target white sample data through the black sample data in the sample data and the black sample data and white sample data included in the unlabeled sample data; Obtain the number of times each sample data in the unlabeled sample data is marked as black sample data, and obtain sample data whose number of times being marked as black sample data is greater than a preset number; according to the application scenario, analyze the features included in the sample data whose number of times being marked as black sample data is greater than a preset number, obtain features with low correlation with the application scenario, delete the features with low correlation in the sample data, and obtain target features; obtain third sample data corresponding to the target features included in the target black sample data and fourth sample data corresponding to the target features included in the target white sample; train a federal model using the third sample data and the fourth sample data to obtain a target classification model.

2. The data classification method according to claim 1, characterized in that: Acquiring sample data, and obtaining target aggregated data based on the sample data, including: Acquire sample data, and perform homomorphic encryption on the sample data to obtain homomorphically encrypted sample data; The target aggregate data obtained by performing sample alignment and aggregation on the homomorphically encrypted sample data of each participant is decrypted to obtain the target aggregate data.

3. The data classification method according to claim 1, characterized in that: Marking the black sample data in the training data to obtain the black sample data and the unlabeled sample data includes: Acquire black sample data included in the training data, mark the black sample data, and obtain black sample data; The unlabeled sample data is obtained through the sample data other than the black sample data.

4. The data classification method according to claim 3, characterized in that: The black sample data is used to perform model training to obtain a classification model, including: Adding the pseudo-white sample data to the unlabeled sample data to obtain first sample data corresponding to the black sample and second sample data corresponding to the unlabeled sample; Model training is performed using the first sample data and the second sample data to obtain a classification model.

5. A data classification device, characterized in that: The device is used for risk control management, including: A training feature determination module, used to obtain sample data, obtain target aggregate data based on the sample data, and determine training data included in the sample data according to the target aggregate data; the sample data includes users or enterprises; a classification model training module, used to mark the black sample data in the training data to obtain black sample data and unlabeled sample data, randomly mark the sample data included in the black sample data as pseudo-white sample data, and perform model training using the data included in the black sample data and corresponding to the training features to obtain a classification model; The sample data determination module is used to perform multiple rounds of pseudo-white sample data random marking in the black sample data, input each pseudo-white sample data into the classification model in turn, and obtain the classification result of the pseudo-white sample data; determine the distribution threshold according to the classification result of the pseudo-white sample data; obtain the classification result of the unlabeled sample data through the classification model, and compare the classification result of the unlabeled sample data with the distribution threshold; when it is determined that the classification result of the unlabeled sample data is less than the distribution threshold, mark the unlabeled sample data as white sample data; when it is determined that the classification result of the unlabeled sample data is not less than the distribution threshold, mark the unlabeled sample data as black sample data; obtain the target black sample data and the target white sample data through the black sample data in the sample data and the black sample data and white sample data included in the unlabeled sample data. A data classification module is used to obtain the number of times each sample data in the unlabeled sample data is marked as black sample data, and obtain sample data whose number of times being marked as black sample data is greater than a preset number; according to an application scenario, analyze the features included in the sample data whose number of times being marked as black sample data is greater than a preset number, obtain features with low correlation with the application scenario, delete the features with low correlation in the sample data, and obtain target features; obtain third sample data corresponding to the target features included in the target black sample data and fourth sample data corresponding to the target features included in the target white sample; train a federal model using the third sample data and the fourth sample data to obtain a target classification model.

6. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data classification method according to any one of claims 1 to 4 is implemented.

7. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to perform the data classification method described in any one of claims 1-4 by executing the executable instructions.

Citation Information

Patent Citations

  • Longitudinal federation modeling method, device and equipment, and computer readable storage medium

    CN112052960A

  • Behavior discrimination model training method and device thereof, electronic equipment and storage medium

    CN112990294A