Data Classification Method and Device, Computer Storage Medium, Electronic Device

Through the twin neural network, feature extraction of the detection data and marked sample data is solved, and the problem of inaccurate data classification in the existing technology is achieved, efficient and accurate data classification is achieved, and user experience is improved.

CN112883990BActive Publication Date: 2025-06-24JINGDONG ALLIANZ GENERAL INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911205188.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-29
Publication Date
2025-06-24
Estimated Expiration
2039-11-29

AI Technical Summary

Technical Problem

When identifying insurance risk data, it is difficult for the prior art to accurately classify the data, resulting in the model having a preference for the categories of most samples and having weak distinction ability.

Method used

A twin neural network is used to extract the detection data and marked sample data for feature, and the category of the data to be detected is determined through the association information.

Benefits of technology

It improves the accuracy of data classification, shortens data classification time, reduces server computing pressure, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112883990B_ABST
    Figure CN112883990B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of computer technologies, and provides a data classification method, apparatus, computer-readable medium, and electronic device. The data includes user data and / or product data. The method includes: obtaining data to be detected and labeled sample data, and forming a data pair to be detected according to the data to be detected and the labeled sample data; respectively extracting features of the data pair to be detected through a siamese neural network to obtain correlation information between the data to be detected and the labeled sample data; and determining the category of the data to be detected according to the correlation information and the category information of the labeled sample data. The present disclosure classifies data through a siamese neural network, thereby improving the accuracy of data classification.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of Internet technology, more and more industries use Internet technology to identify data categories. Especially in the financial industry, Internet technology is used to identify insurance risk data to achieve dynamic pricing for users with different insurance risk data. Therefore, how to accurately identify data categories is the primary consideration.

[0003] In the prior art, based on supervised learning algorithms, the problem of sample imbalance is solved to a certain extent by oversampling or undersampling, and positive and negative samples are classified by training a classifier. However, the methods of oversampling or undersampling do not fundamentally solve the problem of data imbalance. Therefore, the trained model will have a preference for the categories of most samples to a certain extent and has a weak ability to distinguish between positive and negative samples.

[0004] In view of this, there is an urgent need in the art to develop a new data classification method and device.

[0005] It should be noted that the information disclosed in the above Background Art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present disclosure is to provide a data classification method, a data classification device, a computer-readable storage medium, and an electronic device, thereby at least achieving accurate classification of data to a certain extent.

[0007] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be learned in part through the practice of the present disclosure.

[0008] According to one aspect of the present disclosure, a data classification method is provided. The data includes user data and / or product data. The method includes: obtaining data to be detected and labeled sample data, and forming a data pair to be detected according to the data to be detected and the labeled sample data; respectively extracting features of the data pair to be detected through a siamese neural network to obtain the association information between the data to be detected and the labeled sample data; and determining the category of the data to be detected according to the association information and the category information of the labeled sample data.

[0009] In some exemplary embodiments of the present disclosure, the siamese neural network includes a first branch structure and a second branch structure. The first branch structure includes a first convolutional layer, a first fully-connected layer, a similarity calculation layer, and a correlation information determination layer. The second branch structure includes a second convolutional layer, a second fully-connected layer, the similarity calculation layer, and the correlation information determination layer. Among them, the first convolutional layer and the second convolutional layer, and the first fully-connected layer and the second fully-connected layer share network weights.

[0010] In some exemplary embodiments of the present disclosure, the siamese neural network is used to extract features from the pair of data to be detected respectively to obtain the correlation information between the data to be detected and the labeled sample data, including: inputting the data to be detected into the first branch structure, and sequentially extracting features of the data to be detected through the first convolutional layer and the first fully-connected layer to obtain the feature information of the data to be detected; inputting the labeled sample data into the second branch structure, and sequentially extracting features of the labeled sample data through the second convolutional layer and the second fully-connected layer to obtain the feature information of the labeled sample data; inputting the feature information of the data to be detected and the feature information of the labeled sample data into the similarity calculation layer, and determining the similarity between the data to be detected and the labeled sample data through the similarity calculation layer; and determining the correlation information according to the similarity through the correlation information determination layer.

[0011] In some exemplary embodiments of the present disclosure, determining the correlation information according to the similarity through the correlation information determination layer includes: comparing the similarity with a first threshold; if the similarity is greater than or equal to the first threshold, determining that the data to be detected and the labeled sample data belong to the same category, and obtaining first correlation information; if the similarity is less than the first threshold, determining that the data to be detected and the labeled sample data belong to different categories, and obtaining second correlation information.

[0012] In some exemplary embodiments of the present disclosure, the labeled sample data includes first-type labeled sample data and second-type labeled sample data; determining the category of the data to be detected according to the correlation information and the category information of the labeled sample data includes: obtaining the number of first sample pairs with the first correlation information and the number of second sample pairs with the second correlation information corresponding to the first-type labeled sample data; obtaining the number of third sample pairs with the first correlation information and the number of fourth sample pairs with the second correlation information corresponding to the second-type labeled sample data; and determining the category of the data to be detected according to the number of the first sample pairs, the number of the second sample pairs, the number of the third sample pairs, and the number of the fourth sample pairs.

[0013] In some exemplary embodiments of the present disclosure, determining the category of the data to be detected according to the number of the first sample pairs, the number of the second sample pairs, the number of the third sample pairs, and the number of the fourth sample pairs includes: performing a summation operation on the number of the first sample pairs and the number of the fourth sample pairs, and dividing the result of the summation operation by the total number of the labeled sample data to obtain the probability that the data to be detected belongs to the first type of data; performing a summation operation on the number of the second sample pairs and the number of the third sample pairs, and dividing the result of the summation operation by the total number of the labeled sample data to obtain the probability that the data to be detected belongs to the second type of data; and determining the category of the data to be detected according to the probability of the first type of data and the probability of the second type of data.

[0014] In some exemplary embodiments of the present disclosure, the method further includes: obtaining a plurality of positive samples and a plurality of negative samples; generating training sample pairs according to the positive samples and the negative samples, and determining associated information samples corresponding to the training sample pairs, where the training sample pairs include positive-positive sample pairs, negative-negative sample pairs, and positive-negative sample pairs, and the sum of the number of the positive-positive sample pairs and the number of the negative-negative sample pairs is equal to the number of the positive-negative sample pairs; inputting the training sample pairs into a twin neural network to be trained, and training the twin neural network to be trained according to the training sample pairs to obtain the twin neural network.

[0015] In some exemplary embodiments of the present disclosure, the associated information samples include first associated information samples and second associated information samples, the positive-positive sample pairs and the negative-negative sample pairs have the first associated information samples, and the positive-negative sample pairs have the second associated information samples.

[0016] In some exemplary embodiments of the present disclosure, generating training sample pairs according to the positive samples and the negative samples includes: establishing a similarity matrix according to the positive samples and the negative samples, where the values of the similarity matrix are the sample similarities between the positive samples and the negative samples; arranging the sample similarities from large to small to form a sequence, sequentially selecting a preset number of target similarities in the sequence, and obtaining target positive samples and target negative samples corresponding to the target similarities; and forming the positive-negative sample pairs according to the target positive samples and the target negative samples.

[0017] In some exemplary embodiments of the present disclosure, inputting the training sample pair into the twin neural network to be trained, and training the twin neural network to be trained according to the training sample pair to obtain the twin neural network, includes: inputting the training sample pair into the twin neural network to be trained, and performing feature extraction on the training sample pair through the twin neural network to be trained to obtain prediction correlation information corresponding to the training sample pair; determining a loss function according to the prediction correlation information and the correlation information sample, and adjusting the parameters of the twin neural network to be trained until the loss function reaches the minimum to obtain the twin neural network.

[0018] In some exemplary embodiments of the present disclosure, the twin neural network to be trained includes a first branch structure to be trained and a second branch structure to be trained. The first branch structure to be trained includes a first convolutional layer to be trained, a first fully connected layer to be trained, a similarity calculation layer to be trained, and a correlation information determination layer to be trained. The second branch structure to be trained includes a second convolutional layer to be trained, a second fully connected layer to be trained, the similarity calculation layer to be trained, and the correlation information determination layer to be trained. Among them, the first convolutional layer to be trained and the second convolutional layer to be trained, and the first fully connected layer to be trained and the second fully connected layer to be trained share network weights.

[0019] In some exemplary embodiments of the present disclosure, performing feature extraction on the training sample pair through the twin neural network to be trained to obtain prediction correlation information corresponding to the training sample pair, includes: performing feature extraction on the training sample pair through the twin neural network to be trained to obtain the sample similarity corresponding to the training sample pair, and determining the prediction correlation information according to the sample similarity.

[0020] In some exemplary embodiments of the present disclosure, performing feature extraction on the training sample pair through the twin neural network to be trained to obtain the sample similarity corresponding to the training sample pair, includes: performing feature extraction on the first training sample in the training sample pair through the first convolutional layer to be trained and the first fully connected layer to be trained to obtain the feature information of the first training sample; performing feature extraction on the second training sample in the training sample pair through the second convolutional layer to be trained and the second fully connected layer to be trained to obtain the feature information of the second training sample; calculating the sample similarity between the training sample pairs according to the feature information of the first training sample and the feature information of the second training sample through the similarity calculation layer to be trained.

[0021] In some exemplary embodiments of the present disclosure, determining the predicted association information according to the sample similarity includes: through the layer for determining the association information to be trained, judging whether the two training samples in the training sample pair belong to the same category according to the sample similarity, and determining the predicted association information corresponding to the training sample pair according to the judgment result.

[0022] In some exemplary embodiments of the present disclosure, determining the predicted association information corresponding to the training sample pair according to the judgment result includes: if the sample similarity is greater than or equal to a second threshold, determining that the two training samples in the training sample pair belong to the same category, and generating first predicted association information corresponding to the training sample pair; if the sample similarity is less than the second threshold, determining that the two training samples in the training sample pair belong to different categories, and generating second predicted association information corresponding to the training sample pair.

[0023] In some exemplary embodiments of the present disclosure, after inputting the training sample pair into the twin neural network to be trained and generating the sample similarity corresponding to the training sample pair through the twin neural network to be trained, the method further includes: updating the similarity matrix according to the sample similarity, and determining updated positive and negative sample pairs according to the updated similarity matrix; performing iterative training on the twin neural network to be trained according to the positive-positive sample pairs, the negative-negative sample pairs, and the updated positive and negative sample pairs.

[0024] According to one aspect of the present disclosure, there is provided a data classification device, where the data includes user data and / or product data, and the data classification device includes: an acquisition data module, configured to acquire data to be detected and labeled sample data, and form a data pair to be detected according to the data to be detected and the labeled sample data; a feature extraction module, configured to extract features of the data pair to be detected through a twin neural network to obtain the association information between the data to be detected and the labeled sample data; a category determination module, configured to determine the category of the data to be detected according to the association information and the category information of the labeled sample data.

[0025] According to one aspect of the present disclosure, there is provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, it implements the data classification method as described in the above embodiments.

[0026] According to one aspect of the present disclosure, there is provided an electronic device, including: one or more processors; a storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the data classification method as described in the above embodiments.

[0027] As can be seen from the above technical solutions, the data classification method, apparatus, computer-readable storage medium, and electronic device in the exemplary embodiments of the present disclosure at least have the following advantages and positive effects:

[0028] The data classification method in the exemplary embodiments of the present disclosure obtains the data to be detected and the labeled sample data, and inputs them into the siamese neural network. The siamese neural network is used to extract features from the data to be detected and the labeled sample data to obtain the association information between the data to be detected and the labeled sample data, and determines the category of the data to be detected according to the association information and the category information of the labeled sample data. On the one hand, the siamese neural network idea is applied to data classification in the present disclosure, shortening the time of data classification, improving the efficiency of data classification, and reducing the computing pressure on the server; on the other hand, the category of the data to be detected is determined through the association information, making full use of the category relationship between the data, greatly improving the accuracy of data classification and enhancing the user experience.

[0029] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0031] Figure 1 Schematically shows a flowchart of a data classification method according to an embodiment of the present disclosure;

[0032] Figure 2 Schematically shows a structural diagram of a siamese neural network according to an embodiment of the present disclosure;

[0033] Figure 3 Schematically shows a flowchart of obtaining association information according to an embodiment of the present disclosure;

[0034] Figure 4 Schematically shows a flowchart of determining the category of the data to be detected according to an embodiment of the present disclosure;

[0035] Figure 5 Schematically shows a structural diagram of a siamese neural network to be trained according to an embodiment of the present disclosure;

[0036] Figure 6Schematically shown is a flowchart of training a twin neural network to be trained according to an embodiment of the present disclosure;

[0037] Figure 7 Schematically shown is a block diagram of a data classification device according to an embodiment of the present disclosure;

[0038] Figure 8 Schematically shown is a module diagram of an electronic device according to an embodiment of the present disclosure;

[0039] Figure 9 Schematically shown is a schematic diagram of a program product according to an embodiment of the present disclosure. Detailed implementation manners

[0040] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0041] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.

[0042] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0043] The flowcharts shown in the drawings are only exemplary illustrations and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0044] Analyzing user data and / or product data can better formulate corresponding service strategies adaptively for different users. Taking the risk identification field in the insurance industry as an example, it mainly identifies the probability of claims and the probability of insurance fraud for the applicant and the insured. By conducting risk identification on the user's information data, different access principles can be implemented for users. For example, users with high risks are refused insurance, or dynamic pricing for users with different claim probabilities can be achieved. However, due to the large amount of underwriting data but small amount of claims data in the current insurance industry, and in some other scenarios, there will also be category biases in user data. Therefore, there is a serious data skew between positive and negative samples, that is, the problem of unbalanced sample data. Inputting such a sample ratio into the model will make the model more likely to be biased towards positive samples, and negative samples will also be misidentified as positive samples, thus losing good classification generalization performance.

[0045] According to the above technical problems, in the related technologies in this field, the following two processing algorithms are mainly used to solve:

[0046] First, the classification-based algorithm. This classification algorithm is mainly based on supervised learning and uses the labels of existing positive and negative samples in the insurance data for classification. Common classification algorithms include: SVM, XGboost, GBDT, and classification algorithms based on neural networks. This method regards the risk identification problem as a binary classification problem and classifies positive and negative samples by training the classifier. However, the classification algorithm directly classifies positive and negative samples. Although the samples can be increased or decreased through undersampling and oversampling, to a certain extent, the problem of unbalanced samples is solved. However, oversampling will cause the repeated appearance of data and does not generate more samples fundamentally. And undersampling is equivalent to throwing away some samples. From the perspective of samples, it is not beneficial to the training information of the model. Therefore, the traditional classification method does not fundamentally solve the problem of data imbalance, and the trained model will have a preference for the categories of most samples to a certain extent and has a weak ability to distinguish between positive and negative samples.

[0047] Second, the outlier detection-based algorithm. This algorithm is mainly based on unsupervised learning. Since the proportion of negative samples in the insurance risk data is relatively small, the negative samples are directly regarded as outliers for processing, and the positive samples are the samples for learning. Based on this method, the characteristics of positive samples can be obtained. Input an unknown data, and judge whether the data is an outlier to determine positive and negative samples. The outlier detection-based algorithm directly uses the category of most samples for training, and the category of minority samples does not participate in training, thus solving the problem of data imbalance. However, the essence of outlier detection is unsupervised learning, that is, there is no corresponding label. Therefore, a part of the accuracy will be lost, and there is data with labels that is not used, and the use of data is not sufficient.

[0048] In view of the problems existing in the related art, in an embodiment of the present disclosure, a data classification method is proposed. The data classification method can be applied to, but not limited to, the following scenarios: insurance risk identification field, financial risk identification field, user portrait construction, user behavior analysis, permission identification, access discrimination, content recommendation, service customization, etc. The present disclosure does not specifically limit the specific application scenarios of the data classification method, and changes in specific application scenarios should be understood to fall within the protection scope of the present disclosure. Figure 1 shows a schematic flowchart of the data classification method, as Figure 1 shown, the data classification method at least includes the following steps:

[0049] Step S110: Obtain the data to be detected and the labeled sample data, and form a data pair to be detected according to the data to be detected and the labeled sample data;

[0050] Step S120: Respectively extract features from the data pair to be detected through a siamese neural network to obtain the association information between the data to be detected and the labeled sample data;

[0051] Step S130: Determine the category of the data to be detected according to the association information and the category information of the labeled sample data.

[0052] On the one hand, the data classification method in the embodiment of the present disclosure applies the idea of a siamese neural network to data classification, shortening the time of data classification, improving the efficiency of data classification, and reducing the computing pressure on the server; on the other hand, it determines the category of the data to be detected through the association information, making full use of the category relationship between the data, greatly improving the accuracy of data classification and enhancing the user experience. User data includes, but is not limited to: user information, user behavior data, user portrait data, user historical behavior data, user current behavior data, etc.

[0053] It should be noted that the data classification method provided in the embodiment of the present disclosure is generally executed by a processor with computing functions. Among them, the processor may include a terminal device, a server, or a processor with computing functions composed of a combination of a terminal device and a server. The present disclosure does not make specific limitations on this.

[0054] To make the technical solution of the present disclosure clearer, next, the data classification method in this exemplary embodiment will be described in detail with an example.

[0055] In step S110, obtain the data to be detected and the labeled sample data, and form a data pair to be detected according to the data to be detected and the labeled sample data.

[0056] In an exemplary embodiment of the present disclosure, the data to be detected may be user data or product data. Correspondingly, the labeled sample data and the data to be detected belong to the same type. For example, if the data to be detected is user data, the labeled sample data is also user data. The difference between the data to be detected and the labeled sample data is that the data category of the data to be detected is unknown, while the data type of the labeled sample data is known. Specifically, in the field of insurance risk identification, the data to be detected may include user data for purchasing insurance, specifically, user data with insurance fraud behavior or user data without insurance fraud behavior. The present disclosure does not make specific limitations in this regard. The labeled sample data may be user data with known insurance fraud behavior or user data with known non-insurance fraud behavior. In addition, the labeled sample data is one or more. In actual applications, to ensure the accuracy of data classification, the more the number of labeled sample data, the better.

[0057] In an exemplary embodiment of the present disclosure, before obtaining the data to be detected and the labeled sample data, it is also necessary to preprocess the obtained user data to remove the noise in the data. The preprocessing of the user data may specifically include removing the timestamp, categorical variables, and null values in the user data, etc.

[0058] In an exemplary embodiment of the present disclosure, forming the data pairs to be detected according to the data to be detected and the labeled sample data is to combine multiple data to be detected and multiple labeled sample data respectively to form multiple data pairs to be detected. For example, combining 10 data to be detected and 15 labeled sample data to form 150 data pairs to be detected. Of course, the way of combining the data to be detected and the labeled sample data can be any way, and the present disclosure does not make specific limitations in this regard.

[0059] In step S120, the twin neural network is used to extract features from the data pairs to be detected respectively to obtain the association information between the data to be detected and the labeled sample data.

[0060] In an exemplary embodiment of the present disclosure, the twin neural network is a neural network composed of two sub-networks with the same structure, and the two sub-networks share network weights. The twin neural network may also be a neural network composed of three sub-networks with the same structure, and its three sub-networks share network weights. Of course, it may also be a neural network composed of multiple sub-networks. The present disclosure does not make specific limitations in this regard.

[0061] In an exemplary embodiment of the present disclosure, Figure 2 Schematically shows the structural diagram of the twin neural network 200, as Figure 2As shown, the Siamese neural network 200 may include a first branch structure and a second branch structure. The first branch structure may include a first convolutional layer 211, a first fully-connected layer 212, a similarity calculation layer 203, and a correlation information determination layer 204. The second branch structure may include a second convolutional layer 221, a second fully-connected layer 222, the similarity calculation layer 203, and the correlation information determination layer 204. The first branch structure and the second branch structure may share the similarity calculation layer 203 and the correlation information determination layer 204. Among them, the first convolutional layer 211 and the second convolutional layer 221, and the first fully-connected layer 212 and the second fully-connected layer 222 may share network weights.

[0062] In an exemplary embodiment of the present disclosure, Figure 3 A flowchart of obtaining correlation information is schematically shown, as Figure 3 shown, the process of obtaining the correlation information between the data to be detected and the labeled sample data by the Siamese neural network 200 through feature extraction respectively includes the following steps:

[0063] In step S310, the data to be detected is input into the first branch structure, and the first convolutional layer 211 and the first fully-connected layer 212 are used to perform feature extraction on the data to be detected in sequence to obtain the feature information of the data to be detected.

[0064] In step S320, the labeled sample data is input into the second branch structure, and the second convolutional layer 221 and the second fully-connected layer 222 are used to perform feature extraction on the labeled sample data in sequence to obtain the feature information of the labeled sample data.

[0065] In an exemplary embodiment of the present disclosure, step S310 and step S320 may be executed simultaneously or may not be executed simultaneously. For example, step S310 may be executed first and then step S320, or step S320 may be executed first and then step S310. The present disclosure does not make specific limitations on this.

[0066] In step S330, the feature information of the data to be detected and the feature information of the labeled sample data are input into the similarity calculation layer, and the similarity between the data to be detected and the labeled sample data is determined by the similarity calculation layer 203.

[0067] In an exemplary embodiment of the present disclosure, the feature information of the data to be detected and the feature information of the labeled sample data are matched, and the similarity is generated according to the matching result, which is the similarity between the data to be detected and the labeled sample data.

[0068] In step S340, the correlation information is determined according to the similarity through the correlation information determination layer 204.

[0069] In an exemplary embodiment of the present disclosure, the association information includes the association between two pieces of data. For example, the association information between the data to be detected and the labeled sample data may include whether the data to be detected and the labeled sample data belong to the same category or different categories. The association information determination layer 204 determines the association information according to the similarity. Specifically, it may include: comparing the similarity with a first threshold; if the similarity is greater than or equal to the first threshold, it is determined that the data to be detected and the labeled sample data belong to the same category, and the first association information is obtained; if the similarity is less than the first threshold, it is determined that the data to be detected and the labeled sample data belong to different categories, and the second association information is obtained. Among them, the first threshold can be set according to specific circumstances. For example, the first threshold can be a similarity of 0.5 or a similarity of 0.6. The present disclosure does not make specific limitations on this.

[0070] Continuing to refer to Figure 1 shown, in step S130, the category of the data to be detected is determined according to the association information and the category information of the labeled sample data.

[0071] In an exemplary embodiment of the present disclosure, the labeled sample data may include the first type of labeled sample data and the second type of labeled sample data. For example, in the field of insurance risk identification, when the first type of labeled sample data is the user data of those with insurance fraud behavior, the second type of labeled sample data is the user data of those without insurance fraud behavior; when the first type of labeled sample data is the user data of those without insurance fraud behavior, the second type of labeled sample data is the user data of those with insurance fraud behavior. In addition, to ensure the accuracy of data classification, the number of the first type of labeled sample data and the second type of labeled sample data is the same.

[0072] In an exemplary embodiment of the present disclosure, Figure 4 schematically shows a flowchart of determining the category of the data to be detected, as Figure 4As shown, in step S410, obtain the number of first sample pairs with the first association information corresponding to the first type of labeled sample data and the number of second sample pairs with the second association information; wherein, the first association information or the second association information is respectively that the data to be detected and the first type of labeled sample data belong to the same category or the data to be detected and the first type of labeled sample data belong to different categories. When the first association information is that the data to be detected and the first type of labeled sample data belong to the same category, then the second association information is that the data to be detected and the first type of labeled sample data belong to different categories. Correspondingly, the number of first sample pairs is the number of the data to be detected and multiple pieces of the first type of labeled sample data belonging to the same category, and the number of second sample pairs is the number of the data to be detected and multiple pieces of the first type of labeled sample data belonging to different categories. When the first association information is that the data to be detected and the first type of labeled sample data belong to different categories, then the second association information is that the data to be detected and the first type of labeled sample data belong to the same category. The present disclosure does not make specific limitations on this. In step S420, obtain the number of third sample pairs with the first association information corresponding to the second type of labeled sample data and the number of fourth sample pairs with the second association information; wherein, the number of third sample pairs and the number of fourth sample pairs are defined in the same way as the above-mentioned number of first sample pairs and the number of second sample pairs. The number of first sample pairs and the number of second sample pairs have been described in detail above and will not be elaborated here. In step S430, perform a summation operation on the number of first sample pairs and the number of fourth sample pairs, and divide the result of the summation operation by the total number of labeled sample data to obtain the probability that the data to be detected belongs to the first type of data; in step S440, perform a summation operation on the number of second sample pairs and the number of third sample pairs, and divide the result of the summation operation by the total number of labeled sample data to obtain the probability that the data to be detected belongs to the second type of data; in step S450, determine the category of the data to be detected according to the probability of the first type of data and the probability of the second type of data. Among them, the category of the first type of labeled sample data is the same as the category of the first type of data, and the category of the second type of labeled sample data is the same as the category of the second type of data.

[0073] For example, 200 labeled sample data are selected, among which 100 are user data with insurance fraud behavior and 100 are user data without insurance fraud behavior. The data A to be detected and the 200 labeled sample data are respectively input into the siamese neural network 200 to obtain the association information between the data A to be detected and the labeled sample data, and record the association information corresponding to the labeled sample data and the number of sample pairs corresponding to the association information. The results of feature extraction and calculation of the association information through the siamese neural network 200 are as follows: the number of the same category between the data A to be detected and 100 user data with insurance fraud behavior is 60, the number of different categories between the data A to be detected and 100 user data with insurance fraud behavior is 40, the number of the same category between the data A to be detected and 100 user data without insurance fraud behavior is 30, and the number of different categories between the data A to be detected and 100 user data without insurance fraud behavior is 70. Further probability statistics on this result shows that the probability that the data A to be detected is user data with insurance fraud behavior is (60 + 70) / 200 = 65%, and the probability that the data A to be detected is user data without insurance fraud behavior is (40 + 30) / 200 = 35%. By calculating the probability that the data A to be detected is user data with insurance fraud behavior in the embodiments of the present disclosure, different access principles can be implemented for users, and dynamic pricing can be achieved for users with different probabilities.

[0074] It should be noted that step S410 and step S420 can be executed simultaneously or not simultaneously, and step S430 and step S440 can be executed simultaneously or not simultaneously. For example, step S410 can be executed first and then step S420, or step S420 can be executed first and then step S410. The present disclosure does not make specific limitations on this.

[0075] In the exemplary embodiment of the present disclosure, before classifying data through the siamese neural network 200, the siamese neural network to be trained also needs to be trained to obtain the siamese neural network 200. Figure 5 Schematically shows the structural schematic diagram of the siamese neural network 500 to be trained, as Figure 5 shown, the siamese neural network 500 to be trained includes a first branch structure to be trained and a second branch structure to be trained. The first branch structure to be trained includes a first convolutional layer 511 to be trained, a first fully connected layer 512 to be trained, a similarity calculation layer 503 to be trained, and an association information determination layer 504 to be trained. The second branch structure 520 to be trained includes a second convolutional layer 521 to be trained, a second fully connected layer 522 to be trained, a similarity calculation layer 503 to be trained, and an association information determination layer 504 to be trained. Among them, the first convolutional layer 511 to be trained and the second convolutional layer 521 to be trained, and the first fully connected layer 512 to be trained and the second fully connected layer 522 to be trained share network weights.

[0076] In an exemplary embodiment of the present disclosure, Figure 6 A schematic diagram showing the training process of the twin neural network 500 to be trained is schematically illustrated, as Figure 6 shown. This process at least includes steps S610 to S630, which are described in detail as follows:

[0077] In step S610, a plurality of positive samples and a plurality of negative samples are obtained.

[0078] In an exemplary embodiment of the present disclosure, the positive sample or negative sample may be user data with insurance fraud behavior or user data without insurance fraud behavior. Usually, user data without insurance fraud behavior is used as the positive sample, and user data with insurance fraud behavior is used as the negative sample. Of course, the present disclosure is not limited thereto.

[0079] In step S620, training sample pairs are generated based on the positive samples and negative samples, and associated information samples corresponding to the training sample pairs are determined. The training sample pairs include positive-positive sample pairs, negative-negative sample pairs, and positive-negative sample pairs, and the sum of the numbers of positive-positive sample pairs and negative-negative sample pairs is equal to the number of positive-negative sample pairs.

[0080] In an exemplary embodiment of the present disclosure, generating training sample pairs based on the positive samples and negative samples may specifically include: establishing a similarity matrix based on the positive samples and negative samples, where the values of the similarity matrix are the sample similarities between the positive samples and negative samples; arranging the sample similarities in descending order to form a sequence, sequentially selecting a preset number of target similarities in the sequence, and obtaining the target positive samples and target negative samples corresponding to the target similarities; forming positive-negative sample pairs based on the target positive samples and target negative samples. Among them, the preset number is set according to the actual situation. The preset number refers to the number of target similarities, and further also refers to the number of target positive samples and target negative samples and the number of positive-negative sample pairs formed.

[0081] In an exemplary embodiment of the present disclosure, the associated information samples include first associated information samples and second associated information samples. The positive-positive sample pairs and negative-negative sample pairs have first associated information samples, and the positive-negative sample pairs have second associated information samples. The first associated information sample may be a same-class label, and the second associated information sample may be a different-class label.

[0082] In step S630, the training sample pairs are input into the twin neural network 500 to be trained.

[0083] In an exemplary embodiment of the present disclosure, the training sample pairs include positive-positive sample pairs, negative-negative sample pairs, and positive-negative sample pairs. These three types of sample pairs are randomly shuffled, and the twin neural network 500 to be trained is trained according to the sample pairs and the associated information samples corresponding to the sample pairs.

[0084] In step S640, the first training sample in the training sample pair is subjected to feature extraction through the to-be-trained first convolutional layer 511 and the to-be-trained first fully-connected layer 512 to obtain the feature information of the first training sample.

[0085] In an exemplary embodiment of the present disclosure, the first training sample may be a positive sample or a negative sample, and the second training sample may be a positive sample or a negative sample.

[0086] In step S650, the second training sample in the training sample pair is subjected to feature extraction through the to-be-trained second convolutional layer 521 and the to-be-trained second fully-connected layer 522 to obtain the feature information of the second training sample.

[0087] In an exemplary embodiment of the present disclosure, step S640 and step S650 may be executed simultaneously or may not be executed simultaneously. For example, step S640 may be executed first and then step S650, or step S650 may be executed first and then step S640. The present disclosure does not make specific limitations thereto.

[0088] In step S660, through the to-be-trained similarity calculation layer 503, the sample similarity between the training sample pairs is calculated according to the feature information of the first training sample and the feature information of the second training sample.

[0089] In an exemplary embodiment of the present disclosure, the feature information of the first training sample and the feature information of the second training sample are matched, and the matching result is generated as the similarity between the first training sample and the second training sample.

[0090] In step S670, the similarity matrix is updated according to the sample similarity.

[0091] In an exemplary embodiment of the present disclosure, after the training sample pair is input into the to-be-trained siamese neural network 500 and the sample similarity corresponding to the training sample pair is generated through the to-be-trained siamese neural network 500, the similarity matrix may be updated according to the sample similarity, and the positive and negative sample pairs may be determined according to the updated similarity matrix; then, the to-be-trained siamese neural network 500 is iteratively trained according to the positive-positive sample pairs, negative-negative sample pairs, and updated positive and negative sample pairs.

[0092] In step S680, through the to-be-trained association information determination layer 504, it is determined whether the two training samples in the training sample pair belong to the same category according to the sample similarity, and the predicted association information corresponding to the training sample pair is determined according to the determination result.

[0093] In an exemplary embodiment of the present disclosure, determining prediction association information corresponding to the training sample pair according to the judgment result may specifically include: if the sample similarity is greater than or equal to the second threshold, it is determined that the two training samples in the training sample pair belong to the same category, and first prediction association information corresponding to the training sample pair is generated; if the sample similarity is less than the second threshold, it is determined that the two training samples in the training sample pair belong to different categories, and second prediction association information corresponding to the training sample pair is generated. Wherein, the second threshold can be set according to specific circumstances. For example, the second threshold can be a similarity of 0.5 or a similarity of 0.6. The present disclosure does not make specific limitations on this.

[0094] In step S690, a loss function is determined according to the prediction association information and the association information sample, and the parameters of the twin neural network 500 to be trained are adjusted until the loss function reaches the minimum to obtain the twin neural network 200.

[0095] The data classification method in the exemplary embodiment of the present disclosure solves the problem of the imbalance in the number of positive and negative samples by forming positive-positive sample pairs, negative-negative sample pairs, and positive-negative sample pairs from positive and negative sample pairs, fully utilizes the relationship between each positive sample and negative sample, and through the training of the data format by the double-branch twin neural network, makes the intra-class gap very small and the inter-class gap large in a certain feature dimension extracted by the twin neural network, thereby realizing high-precision classification of data.

[0096] The following introduces the apparatus embodiments of the present disclosure, which can be used to execute the above data classification method of the present disclosure. For details not disclosed in the apparatus embodiments of the present disclosure, please refer to the embodiments of the above data classification method of the present disclosure.

[0097] Figure 7 The block diagram of a data classification apparatus according to an embodiment of the present disclosure is schematically shown.

[0098] Refer to Figure 7 As shown, a data classification apparatus 700 according to an embodiment of the present disclosure, the data classification apparatus 700 includes: an acquiring data module 701, an extracting information module 702, and a determining category module 703. Specifically:

[0099] The acquiring data module 701 is configured to acquire the data to be detected and the labeled sample data, and form a data pair to be detected according to the data to be detected and the labeled sample data;

[0100] The extracting information module 702 is configured to respectively perform feature extraction on the data pair to be detected through the twin neural network 200 to obtain the association information between the data to be detected and the labeled sample data;

[0101] A determining category module 703, configured to determine the category of the data to be detected according to the association information and the category information of the labeled sample data.

[0102] The specific details of each of the above data classification devices have been described in detail in the corresponding data classification method, and thus will not be elaborated here.

[0103] It should be noted that although several modules or units of the devices for execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0104] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0105] Those skilled in the art to which the present invention pertains can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to herein as "circuitry", "module", or "system".

[0106] Next, refer to Figure 8 to describe the electronic device 800 according to this embodiment of the present invention. Figure 8 The illustrated electronic device 800 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0107] As Figure 8 shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.

[0108] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit 810 can execute as Figure 1In step S110 shown in the figure, obtain the data to be detected and the labeled sample data, and form a data pair to be detected based on the data to be detected and the labeled sample data; in step S120, respectively perform feature extraction on the data pair to be detected through a siamese neural network to obtain the association information between the data to be detected and the labeled sample data; in step S130, determine the category of the data to be detected according to the association information and the category information of the labeled sample data.

[0109] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.

[0110] The storage unit 820 may further include a program / utility 8204 having a set (at least one) of program modules 8205. Such program modules 805 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0111] The bus 830 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.

[0112] The electronic device 800 may also communicate with one or more external devices 1000 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a viewer to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be performed through the input / output (I / O) interface 850. Moreover, the electronic device 800 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0113] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0114] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of this specification is stored. In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.

[0115] Referring to Figure 9 As shown, a program product 900 for implementing the above method according to an embodiment of the present invention is described, which can be a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0116] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0117] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0118] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0119] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0120] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present invention, and are not for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0121] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0122] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A data classification method, applied to insurance risk identification, characterized in that The data includes user data and / or product data, and the method includes: Obtaining data to be detected and labeled sample data, and forming a data pair to be detected according to the data to be detected and the labeled sample data; wherein, the data to be detected includes user data of purchasing insurance, and the labeled sample data includes first-type labeled sample data and second-type labeled sample data; the first-type labeled sample data is one of user data with insurance fraud behavior or user data without insurance fraud behavior, and the second-type labeled sample data is one of user data with insurance fraud behavior or user data without insurance fraud behavior, wherein the first-type labeled sample data and the second-type labeled sample data are of different types; Respectively extracting features from the data pair to be detected through a siamese neural network to obtain the association information between the data to be detected and the labeled sample data; the association information includes first association information and second association information; Determining the category of the data to be detected according to the association information and the category information of the labeled sample data, including: Obtaining the number of first sample pairs with the first association information and the number of second sample pairs with the second association information corresponding to the first-type labeled sample data; Obtaining the number of third sample pairs with the first association information and the number of fourth sample pairs with the second association information corresponding to the second-type labeled sample data; Determining the category of the data to be detected according to the number of first sample pairs, the number of second sample pairs, the number of third sample pairs, and the number of fourth sample pairs.

2. The data classification method according to claim 1, wherein The siamese neural network includes a first branch structure and a second branch structure. The first branch structure includes a first convolutional layer, a first fully connected layer, a similarity calculation layer, and an association information determination layer. The second branch structure includes a second convolutional layer, a second fully connected layer, the similarity calculation layer, and the association information determination layer. Among them, the first convolutional layer and the second convolutional layer, and the first fully connected layer and the second fully connected layer share network weights.

3. The data classification method according to claim 2, wherein Respectively extracting features from the data pair to be detected through a siamese neural network to obtain the association information between the data to be detected and the labeled sample data, including: Inputting the data to be detected into the first branch structure, and sequentially extracting features from the data to be detected through the first convolutional layer and the first fully connected layer to obtain the feature information of the data to be detected; Inputting the labeled sample data into the second branch structure, and sequentially extracting features from the labeled sample data through the second convolutional layer and the second fully connected layer to obtain the feature information of the labeled sample data; Inputting the feature information of the data to be detected and the feature information of the labeled sample data into the similarity calculation layer, and determining the similarity between the data to be detected and the labeled sample data through the similarity calculation layer; Determining the association information according to the similarity through the association information determination layer.

4. The data classification method according to claim 3, characterized in that Determine the layer based on the associated information, and determine the associated information according to the similarity, including: Compare the similarity with a first threshold; If the similarity is greater than or equal to the first threshold, determine that the data to be detected and the labeled sample data belong to the same category, and obtain first associated information; If the similarity is less than the first threshold, determine that the data to be detected and the labeled sample data belong to different categories, and obtain second associated information.

5. The data classification method according to claim 1, wherein Determine the category of the data to be detected according to the first sample pair quantity, the second sample pair quantity, the third sample pair quantity, and the fourth sample pair quantity, including: Perform a summation operation on the first sample pair quantity and the fourth sample pair quantity, and divide the result of the summation operation by the total number of the labeled sample data to obtain the probability that the data to be detected belongs to the first type of data; Perform a summation operation on the second sample pair quantity and the third sample pair quantity, and divide the result of the summation operation by the total number of the labeled sample data to obtain the probability that the data to be detected belongs to the second type of data; Determine the category of the data to be detected according to the probability of the first type of data and the probability of the second type of data.

6. The data classification method according to claim 1, wherein The method further includes: Obtain a plurality of positive samples and a plurality of negative samples; Generate training sample pairs according to the positive samples and the negative samples, and determine the associated information samples corresponding to the training sample pairs. The training sample pairs include positive-positive sample pairs, negative-negative sample pairs, and positive-negative sample pairs, and the sum of the number of the positive-positive sample pairs and the negative-negative sample pairs is equal to the number of the positive-negative sample pairs; Input the training sample pairs into the twin neural network to be trained, and train the twin neural network to be trained according to the training sample pairs to obtain the twin neural network.

7. The data classification method according to claim 6, wherein The associated information samples include first associated information samples and second associated information samples. The positive-positive sample pairs and the negative-negative sample pairs have the first associated information samples, and the positive-negative sample pairs have the second associated information samples.

8. The data classification method according to claim 6, wherein Generate training sample pairs according to the positive samples and the negative samples, including: Establish a similarity matrix according to the positive samples and the negative samples. The values of the similarity matrix are the sample similarities between the positive samples and the negative samples; Arrange the sample similarities from large to small to form a sequence, sequentially select a preset number of target similarities in the sequence, and obtain the target positive samples and target negative samples corresponding to the target similarities; Form the positive-negative sample pairs according to the target positive samples and the target negative samples.

9. The data classification method according to claim 6 or 8, characterized in that, Input the training sample pairs into the twin neural network to be trained, and train the twin neural network to be trained according to the training sample pairs to obtain the twin neural network, including: Input the training sample pairs into the twin neural network to be trained, and perform feature extraction on the training sample pairs through the twin neural network to be trained to obtain the predicted associated information corresponding to the training sample pairs; Determine a loss function based on the predicted association information and the association information sample, and adjust the parameters of the twin neural network to be trained until the loss function reaches the minimum to obtain the twin neural network.

10. A data classification device, applied to insurance risk identification, characterized in that, The data includes user data and / or product data, and the device includes: An obtaining data module, configured to obtain data to be detected and labeled sample data, and form a data pair to be detected according to the data to be detected and the labeled sample data; wherein, the data to be detected includes user data of purchasing insurance, and the labeled sample data includes first-type labeled sample data and second-type labeled sample data; the first-type labeled sample data is one of user data with insurance fraud behavior or user data without insurance fraud behavior, and the second-type labeled sample data is one of user data with insurance fraud behavior or user data without insurance fraud behavior, and the types of the first-type labeled sample data and the second-type labeled sample data are different; An information extraction module, configured to extract features from the data pair to be detected through a twin neural network to obtain the association information between the data to be detected and the labeled sample data; the association information includes first association information and second association information; A category determination module, configured to determine the category of the data to be detected according to the association information and the category information of the labeled sample data, including: obtaining the number of first sample pairs with the first association information corresponding to the first-type labeled sample data and the number of second sample pairs with the second association information; obtaining the number of third sample pairs with the first association information corresponding to the second-type labeled sample data and the number of fourth sample pairs with the second association information; determining the category of the data to be detected according to the number of the first sample pairs, the number of the second sample pairs, the number of the third sample pairs, and the number of the fourth sample pairs.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the data classification method according to any one of claims 1 to 9.

12. An electronic device, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the data classification method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Electronic device, insured domestic animal recognition method and computer readable storage medium

    CN107766807A

  • Health risk assessment method, device, and apparatus based on character traits

    CN107767959A