Abnormal sample data determination method and device, program product and storage medium

By classifying and converting the output data of the initial network model in machine learning, combining model parameters, efficiently finding abnormal sample data, the problem of inefficiency in the existing technology is solved and fast and accurate abnormal detection is achieved.

CN120067678APending Publication Date: 2025-05-30JIANXIN CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510103412.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the field of machine learning, it is difficult to efficiently determine abnormal sample data in the prior art, especially when modifying the prediction labels of samples one by one in the dataset, it is necessary to traverse the entire dataset, resulting in inefficiency.

Method used

By training the initial network model, obtaining the output data set, classifying and converting it, and generating multi-class data sets. Use these data sets and model parameters to find exception sample data from the sample data set.

Benefits of technology

Through batch processing and comprehensive analysis of model parameters, the speed and accuracy of abnormal detection are significantly improved, and the problem of inability to efficiently determine abnormal sample data is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067678A_ABST
    Figure CN120067678A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an abnormal sample data determination method and device, a program product and a storage medium, and the method comprises the steps: obtaining an input sample data set and an output first data set of an initial network model in a process of training the initial network model through employing the sample data set; performing classification operation on a plurality of first data included in the first data set to obtain N types of second data sets; performing data conversion on each type of data included in the N types of second data sets according to the data type in each type of second data set to obtain N types of third data sets; and searching abnormal sample data from the sample data set by using the N types of second data sets, the N types of third data sets and the model parameters of the initial network model. The problem that abnormal sample data cannot be efficiently determined in related technologies is solved, and the effect of efficiently determining the abnormal sample data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technologies, and more particularly, to a method and apparatus for determining abnormal sample data, a program product, and a storage medium. Background Art

[0002] In the field of machine learning, when the true labels of the dataset samples are unknown, by batch-modifying the predicted labels of the samples and observing the changes in the overall data evaluation metrics, the positions of the mislabeled samples in the dataset can be determined. However, modifying the predicted labels of the samples one by one is laborious and complex, and the performance of the observable evaluation metrics needs to be tested for each modification. In extreme cases, when the predicted label of the last sample in the predicted sample labels is incorrect, starting from the beginning of the dataset and modifying the predicted labels of each sample one by one, the entire dataset needs to be traversed to discover the abnormal sample, resulting in the problem of being unable to efficiently determine abnormal sample data. Summary of the Invention

[0003] The embodiments of the present application provide a method and apparatus for determining abnormal sample data, a program product, and a storage medium, so as to at least solve the problem of being unable to efficiently determine abnormal sample data in the related art.

[0004] According to an embodiment of the present application, a method for determining abnormal sample data is provided, including: during the process of training an initial network model using a sample data set, obtaining a first data set output by the initial network model based on the input sample data set, where one sample data in the sample data set corresponds to one first data in the first data set; performing a classification operation on multiple first data included in the first data set to obtain N second data sets, where N is a natural number greater than 1; respectively performing data conversion on each type of data included in the N second data sets according to the data categories of each type of second data set to obtain N third data sets; using the N second data sets, the N third data sets, and the model parameters of the initial network model to find abnormal sample data from the sample data set, where the data output by the initial network model corresponding to the abnormal sample data is abnormal data.

[0005] In an exemplary embodiment, performing a classification operation on multiple first data included in the first data set to obtain N second data sets includes: when the first data is a predicted label, respectively determining the data categories of multiple predicted labels according to the attribute values of the multiple predicted labels; performing the classification operation on the multiple predicted labels according to the data categories of the multiple predicted labels to obtain N second data sets.

[0006] In an exemplary embodiment, according to the data categories in each of the above-mentioned second data sets, data conversion is respectively performed on each type of data included in the N types of the above-mentioned second data sets to obtain N types of third data sets, including: performing a splitting operation on the N types of the above-mentioned second data sets according to the number of second data included in the above-mentioned second data sets, and splitting each type of the above-mentioned second data sets into M second sub-data sets, where M is a natural number greater than 1; respectively performing a modification operation on the M second sub-data sets to perform the above-mentioned data conversion on the attribute values of the second data in the M second sub-data sets to obtain N types of the above-mentioned third data sets; where the above-mentioned modification operation includes: modifying the attribute values of the data included in the M sub-data sets in the second target data set according to the attribute values of the data included in other data sets, to obtain third data, where the above-mentioned other data set and the above-mentioned second target data set are both data sets in the N types of the above-mentioned second data sets, and the attribute value of the above-mentioned third data is the same as the attribute value of the data in the above-mentioned other data set.

[0007] In an exemplary embodiment, performing a splitting operation on the N types of the above-mentioned second data sets according to the number of second data included in the above-mentioned second data sets, and splitting each type of the above-mentioned second data sets into M second sub-data sets, includes: determining the number of the second data in each type of the above-mentioned second data sets to obtain N second data numbers; based on the N second data numbers, determining a quantity ratio; based on the above-mentioned quantity ratio, splitting each type of the above-mentioned second data sets into M of the above-mentioned second sub-data sets.

[0008] In an exemplary embodiment, using the N types of the above-mentioned second data sets, the N types of the above-mentioned third data sets, and the model parameters of the above-mentioned initial network model to find abnormal sample data from the above-mentioned sample data set, includes: performing the following operations on each type of the above-mentioned second data sets to find abnormal sample data from the above-mentioned sample data set: calculating M initial evaluation indicators of the M second sub-data sets according to the attribute values of the second data in the M second sub-data sets; calculating M target evaluation indicators of the M third sub-data sets in the above-mentioned third data set according to the attribute values of the above-mentioned third data in the M third sub-data sets; based on the M initial evaluation indicators, the M target evaluation indicators, and the model parameters of the above-mentioned initial network model, finding the above-mentioned abnormal sample data in the above-mentioned sub-data sets.

[0009] In an exemplary embodiment, based on the M initial evaluation metrics, the M target evaluation metrics, and the model parameters of the initial network model, finding the abnormal sample data in the sub-data set includes: calculating the difference between the target evaluation metrics corresponding to the target sub-data set and the initial evaluation metrics, where the target sub-data set is any one of the M third sub-data sets; finding abnormal data in the target sub-data set when the difference is less than a preset threshold and the model parameters of the initial network model meet the preset conditions; determining the abnormal sample data based on the abnormal data, where the abnormal sample data is sample data that does not meet the training rules of the initial network model, and the data output by the initial network model corresponding to the abnormal sample data is abnormal data.

[0010] According to another embodiment of the present application, there is also provided a device for determining abnormal sample data, including: a first acquisition module, configured to acquire, during the process of training an initial network model using a sample data set, a first data set output by the initial network model based on the input sample data set, where one sample data in the sample data set corresponds to one first data in the first data set; a first execution module, configured to perform a classification operation on multiple first data included in the first data set to obtain N second data sets, where N is a natural number greater than 1; a first conversion module, configured to perform data conversion on each type of data included in the N second data sets respectively according to the data category of each type of second data set to obtain N third data sets; a first search module, configured to use the N second data sets, the N third data sets, and the model parameters of the initial network model to find abnormal sample data in the sample data set, where the data output by the initial network model corresponding to the abnormal sample data is abnormal data.

[0011] In an exemplary embodiment, the first execution module includes: a first determination sub-module, configured to, when the first data is a prediction label, determine the data categories of multiple prediction labels respectively according to the attribute values of the multiple prediction labels; a first execution sub-module, configured to perform the classification operation on the multiple prediction labels according to the data categories of the multiple prediction labels to obtain N second data sets.

[0012] In an exemplary embodiment, the first conversion module includes: a second execution sub-module, configured to perform a splitting operation on each of the N types of the second data sets according to the number of the second data included in the second data sets, and split each type of the second data sets into M second sub-data sets, where M is a natural number greater than 1; a third execution sub-module, configured to perform a modification operation on each of the M second sub-data sets respectively, so as to perform the data conversion on the attribute values of the second data in the M second sub-data sets, and obtain N types of the third data sets; wherein, the modification operation includes: modifying the attribute values of the data included in the M sub-data sets in the second target data set according to the attribute values of the data included in the other data set, to obtain third data, where the other data set and the second target data set are both data sets among the N types of the second data sets, and the attribute value of the third data is the same as the attribute value of the data in the other data set.

[0013] In an exemplary embodiment, the second sub-execution module includes: a first determination unit, configured to determine the number of the second data in each type of the second data sets, and obtain N second data numbers; a second determination unit, configured to determine a quantity ratio based on the N second data numbers; a first splitting unit, configured to split each type of the second data sets into M second sub-data sets based on the quantity ratio.

[0014] In an exemplary embodiment, the first search module includes: a fourth execution sub-module, configured to perform the following operations on each type of the second data sets to search for abnormal sample data from the sample data set: calculating M initial evaluation indexes of the M second sub-data sets according to the attribute values of the second data in the M second sub-data sets; calculating M target evaluation indexes of the M third sub-data sets according to the attribute values of the third data in the M third sub-data sets in the third data set; searching for the abnormal sample data in the sub-data sets based on the M initial evaluation indexes, the M target evaluation indexes, and the model parameters of the initial network model.

[0015] In an exemplary embodiment, the above-mentioned first search module includes: a first calculation sub-module, configured to calculate the difference between the above-mentioned target evaluation index corresponding to the target sub-data set and the above-mentioned initial evaluation index, wherein the target sub-data set is any one of the above-mentioned M third sub-data sets; a first search sub-module, configured to search for abnormal data in the target sub-data set when the difference is less than a preset threshold and the model parameters of the above-mentioned initial network model meet the preset conditions; a second determination sub-module, configured to determine the above-mentioned abnormal sample data based on the above-mentioned abnormal data, wherein the abnormal sample data is sample data that does not meet the training rules of the initial network model, and the data output by the above-mentioned initial network model corresponding to the abnormal sample data is abnormal data.

[0016] According to another embodiment of the present application, there is also provided a computer program product, including a computer program, and the above-mentioned computer program is configured to be executed by a processor to perform the steps in any one of the above-mentioned method embodiments.

[0017] According to another embodiment of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, and the above-mentioned computer program is configured to be executed by a processor to perform the steps in any one of the above-mentioned method embodiments.

[0018] According to another embodiment of the present application, there is also provided an electronic device, including a memory, a processor, and a computer program stored on the above-mentioned memory and executable on the above-mentioned processor, and the above-mentioned processor is configured to execute the above-mentioned computer program to perform the steps in any one of the above-mentioned method embodiments.

[0019] Through the present application, a classification operation is performed on a plurality of first data included in a first data set output by an initial network model based on an input sample data set to obtain N second data sets; according to the data categories in each second data set, data conversion is performed on each type of data included in the N second data sets to obtain N third data sets; the N second data sets, the N third data sets, and the model parameters of the initial network model are used to search for abnormal sample data from the sample data set. Since the present application performs a classification operation on a plurality of first data to obtain N second data sets and performs data conversion on the same type of first data, it avoids the inefficiency of analyzing each first data one by one. Instead, through batch processing and comprehensive analysis of model parameters, the speed and accuracy of anomaly detection are greatly improved. Therefore, the problem of being unable to efficiently determine abnormal sample data in the related art is solved, and the effect of efficiently determining abnormal sample data is achieved. Description of the Drawings

[0020] Figure 1It is a hardware structure block diagram of a mobile terminal for a method of determining abnormal sample data according to an embodiment of the present application;

[0021] Figure 2 It is a flowchart of a method of determining abnormal sample data according to an embodiment of the present application;

[0022] Figure 3 It is a flowchart of a method of determining abnormal sample data in a specific embodiment of the present application;

[0023] Figure 4 It is a schematic diagram of a method of determining abnormal sample data in a specific embodiment of the present application;

[0024] Figure 5 It is a structure block diagram of a device for determining abnormal sample data according to an embodiment of the present application. Specific Embodiments

[0025] In the following, embodiments of the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0027] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 It is a hardware structure block diagram of a mobile terminal for a method of determining abnormal sample data according to an embodiment of the present application. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 the structure shown is only schematic, and it does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.

[0028] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to a method for determining abnormal sample data in an embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.

[0029] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0030] In this embodiment, a method for determining abnormal sample data is provided. Figure 2 It is a flowchart of a method for determining abnormal sample data according to an embodiment of the present application, as Figure 2 shown, and the process includes the following steps:

[0031] Step S202, in the process of training the initial network model using the sample data set, obtain a first data set output by the initial network model based on the input sample data set, where one sample data in the sample data set corresponds to one first data in the first data set;

[0032] Optionally, the sample data in the sample data set may be sample data in fields such as financial risk control, medical diagnosis, and autonomous driving. This embodiment does not limit the field of the sample data in the sample data set.

[0033] Optionally, the initial network model includes but is not limited to a multi-layer perceptron model, a convolutional neural network model, and a recurrent neural network model. This embodiment does not limit the type of the initial network model.

[0034] Optionally, the first data set consists of output data generated by the initial network model based on the input sample data set. Each input sample data corresponds to an output data, and these output data reflect the model's preliminary prediction results for the input. During the training process, the first data set changes as the model is optimized until the model reaches the expected performance standard. The first data in the first data set can be of various types, depending on the model's task and the design of the output layer. For example, in a classification task, the first data can be class labels, while in a regression task, the first data can be continuous numerical predictions. The first data set can be a set of prediction labels.

[0035] Step S204: Perform a classification operation on the multiple first data included in the first data set above to obtain N sets of second data, where N is a natural number greater than 1.

[0036] Step S206: According to the data categories in each set of second data above, perform data conversion on each type of data included in the N sets of second data above to obtain N sets of third data.

[0037] Optionally, the data category is the basis for constructing the second data set, and each data category corresponds to a set of second data. For example, in a classification task, the model predicts which category each sample belongs to, and the data category is these predicted categories.

[0038] Step S208: Use the N sets of second data above, the N sets of third data above, and the model parameters of the initial network model above to find abnormal sample data from the sample data set above, where the data output by the initial network model corresponding to the abnormal sample data above is abnormal data.

[0039] Optionally, the model parameters include but are not limited to weights and biases.

[0040] In this embodiment, the execution subject of the above steps can be a terminal, a server, a specific processor set in the terminal or the server, or a processor or processing device set relatively independently of the terminal or the server, but is not limited thereto.

[0041] Through the above steps, classification operations are performed on multiple first data included in the first data set output based on the input sample data set of the initial network model, and N types of second data sets are obtained; according to the data categories in each type of second data set, data conversion is performed on each type of data included in the N types of second data sets to obtain N types of third data sets; the N types of second data sets, the N types of third data sets, and the model parameters of the initial network model are used to find abnormal sample data from the sample data set. Since in this embodiment, classification operations are performed on multiple first data to obtain N types of second data sets, and data conversion is performed on the first data of the same type, the inefficiency of analyzing each first data one by one is avoided. Instead, through batch processing and comprehensive analysis of model parameters, the speed and accuracy of anomaly detection are greatly improved. Therefore, the problem in the related art of being unable to efficiently determine abnormal sample data is solved, and the effect of efficiently determining abnormal sample data is achieved.

[0042] In an exemplary embodiment, performing classification operations on multiple first data included in the first data set to obtain N types of second data sets includes: when the first data is a prediction label, respectively determining the data categories of the multiple prediction labels according to the attribute values of the multiple prediction labels; performing the classification operations on the multiple prediction labels according to the data categories of the multiple prediction labels to obtain N types of the second data sets.

[0043] Optionally, the attribute value is used to indicate the value of the prediction label.

[0044] Optionally, for example, a credit scoring system is used to predict a customer's credit rating based on the customer's financial history and credit record to identify potential fraudulent applications. The output of the model is the credit rating prediction of the customer, which is a discrete category label. The prediction labels in the first data set are classified according to the risk level (corresponding to the above attribute value) to obtain three types of second data sets, namely, the second data sets of high-risk, medium-risk, and low-risk customers.

[0045] Optionally, for example, a deep learning model is used to analyze a patient's imaging data to assist a doctor in diagnosing whether the patient has a specific disease. The output of the model is a disease prediction based on the imaging. The prediction labels in the first data set are classified according to the predicted disease type (corresponding to the above attribute value) to obtain three types of second data sets: the second data sets of pneumonia, fracture, and normal imaging, respectively.

[0046] Optionally, for example, a sentiment analysis model is used to predict the sentiment tendency of the text, such as positive, negative, or neutral. The output of the model is the sentiment prediction label of the text. According to the prediction result of the model, the prediction labels in the first data set are classified according to the sentiment tendency (corresponding to the above attribute values), and three types of second data sets are obtained: namely, the second data sets of positive, negative, and neutral.

[0047] In this embodiment, through the classification operation, the first data is organized according to categories, and abnormal samples can be more accurately located in specific categories, achieving the purpose of improving the efficiency and accuracy of detecting abnormal data.

[0048] In an exemplary embodiment, according to the data categories in each of the above second data sets, data conversion is respectively performed on each type of data included in the N types of the above second data sets, and N types of third data sets are obtained, including: according to the number of the second data included in the above second data set, a splitting operation is performed on the N types of the above second data sets, and each type of the above second data set is split into M second sub-data sets, where M is a natural number greater than 1; a modification operation is respectively performed on the M second sub-data sets to perform the above data conversion on the attribute values of the second data in the M second sub-data sets, and N types of the above third data sets are obtained; where the above modification operation includes: according to the attribute values of the data included in other data sets, modifying the attribute values of the data included in the M sub-data sets in the second target data set to obtain third data, where the above other data set and the above second target data set are both data sets in the N types of the above second data sets, and the attribute value of the above third data is the same as the attribute value of the data in the above other data set.

[0049] Optionally, the second data may be a prediction label, and the second data is the data in the second data set after the first data is classified.

[0050] Optionally, the modification operation is used to convert the initial attribute value of the second data into the attribute value of the second data of different data categories.

[0051] Optionally, for example, the initial training network model deals with a text classification problem, and classifies the text into three themes: data category - technology (attribute value 1), data category - sports (attribute value 2), and data category - entertainment (attribute value 3). The second data sets include: second data set 1 (attribute value 1): [M1, M3, M4, M6, M8, M9], second data set 2 (attribute value 2): [M2, M5, M7], second data set 3 (attribute value 3): [M10, M11, M12, M13].

[0052] Optionally, based on the number of second data in each type of second data set, perform a splitting operation on the second data in the second data set to split the second data set into multiple second sub-data sets. For example, in an equal division manner, split the second data set 1 (attribute value 1): [M1, M3, M4, M6, M8, M9] into 3 second sub-data sets or split it into 2 second sub-data sets. For example, in a random splitting manner, split the second data set 2 (attribute value 2): [M2, M5, M7] into 2 second sub-data sets.

[0053] Optionally, for example, convert the attribute value 1 of the second data in the 3 second sub-data sets of the second data set 1: [(M1, M3), (M4, M6), (M8, M9)] to the attribute value 2 of the second data in the second data set 2 to obtain the third data set 1 (attribute value 2): [(M1, M3), (M4, M6), (M8, M9)].

[0054] Optionally, convert the attribute value 2 of the second data in the 2 second sub-data sets of the second data set 2: [M2, (M5, M7)] to the attribute value 3 of the second data in the second data set 3 to obtain the third data set 2 (attribute value 3) [M2, (M5, M7)].

[0055] In this embodiment, through the splitting operation, the modification operation task is decomposed into smaller sub-tasks, achieving the purpose of improving the data processing efficiency. In addition, the modification operation focuses on the second sub-data set, avoiding unnecessary processing of the entire data set, and achieving the purpose of reducing computing resources.

[0056] In an exemplary embodiment, according to the number of second data included in the above-mentioned second data set, perform a splitting operation on N types of the above-mentioned second data sets, and split each type of the above-mentioned second data set into M second sub-data sets, including: determining the number of the above-mentioned second data in each type of the above-mentioned second data set to obtain N second data quantities; based on the N second data quantities, determining the quantity ratio; based on the quantity ratio, splitting each type of the above-mentioned second data set into M of the above-mentioned second sub-data sets.

[0057] Optionally, the quantity ratio is the ratio of the number of second data in N types of second data sets.

[0058] Optionally, based on the quantity ratio, each type of second data set is divided into M second sub-data sets, including: when the value of the largest part in the quantity ratio accounts for less than the preset ratio of the total, the M values of each type of second data set are the same, where the preset ratio is used to determine whether the quantity of the second data is balanced; when the value of the largest part in the quantity ratio accounts for greater than or equal to the predicted ratio, the larger the quantity of the second data in the second data set, the larger the corresponding M value.

[0059] Optionally, for example, the initial training network model is used to perform sentiment analysis on online product reviews, and the reviews are classified into three sentiment tendencies: positive, neutral, and negative. For example, three types of second data sets are obtained, corresponding to the second data sets of the above three sentiment tendencies respectively.

[0060] Count the quantity of the second data in each type of second data set, denoted as N1, N2, and N3, which respectively represent the quantity of the second data predicted as positive, neutral, and negative sentiments. For example, N1 = 1000 (the second data representing positive reviews), N2 = 500 (the second data representing neutral reviews), and N3 = 200 (the second data representing negative reviews).

[0061] Based on the quantities of N1, N2, and N3, determine the quantity ratio 10:5:2. The value of the largest part in the quantity ratio accounting for the total = 10 / (10 + 5 + 2) = 10 / 17, and this value is greater than the preset ratio of 8 / 17, indicating that the quantity of the second data in the second data set is unbalanced. The second data set of positive reviews (N1 = 1000) can be divided into 10 second sub-data sets, with each second sub-data set containing 100 second data (1000 / 10). The second data set of neutral reviews (N2 = 500) is divided into 5 second sub-data sets, with each second sub-data set containing 100 second data (500 / 5). The second data set of negative reviews (N3 = 200) is divided into 2 second sub-data sets, with each second sub-data set containing 100 second data (200 / 2).

[0062] In this embodiment, by dividing the large data set into smaller subsets, the purpose of improving data processing efficiency is achieved. Furthermore, if the original data set categories are unbalanced, the data quantity in each subset can be balanced through division, achieving the purpose of reducing computing resources.

[0063] In an exemplary embodiment, using N classes of the above-mentioned second data sets, N classes of the above-mentioned third data sets, and the model parameters of the above-mentioned initial network model to find abnormal sample data from the above-mentioned sample data set includes: performing the following operations on each class of the above-mentioned second data sets to find abnormal sample data from the above-mentioned sample data set: calculating M initial evaluation metrics for M of the above-mentioned second sub-data sets according to the attribute values of the second data in the M above-mentioned second sub-data sets; calculating M target evaluation metrics for M of the above-mentioned third sub-data sets in the above-mentioned third data set according to the attribute values of the third data in the M above-mentioned third sub-data sets; and finding the above-mentioned abnormal sample data in the above-mentioned sub-data sets based on the M above-mentioned initial evaluation metrics, the M above-mentioned target evaluation metrics, and the model parameters of the above-mentioned initial network model.

[0064] Optionally, the initial evaluation metric is used to represent the training performance of the initial network model on the sample data set, including but not limited to confidence, F1 value, etc.

[0065] In this embodiment, by splitting the second data set into second sub-data sets and calculating the initial evaluation metrics for each sub-set, the purpose of efficiently locating abnormal sample data is achieved. Compared with evaluating the entire data set or each sample one by one, the amount of calculation and the required time can be significantly reduced. Especially when dealing with large-scale data sets, the rapid detection of abnormal samples becomes particularly important. Further, by comparing the evaluation metrics before and after modification, the purpose of reducing the dependence on the true labels is achieved.

[0066] In an exemplary embodiment, finding the above-mentioned abnormal sample data in the above-mentioned sub-data sets based on the M above-mentioned initial evaluation metrics, the M above-mentioned target evaluation metrics, and the model parameters of the above-mentioned initial network model includes: calculating the difference between the above-mentioned target evaluation metric corresponding to the target sub-data set and the above-mentioned initial evaluation metric, where the target sub-data set is any one of the M above-mentioned third sub-data sets; finding abnormal data in the above-mentioned target sub-data set when the above-mentioned difference is less than a preset threshold and the model parameters of the above-mentioned initial network model meet the preset conditions; and determining the above-mentioned abnormal sample data based on the above-mentioned abnormal data, where the above-mentioned abnormal sample data is sample data that does not meet the training rules of the initial network model, and the data output by the above-mentioned initial network model corresponding to the above-mentioned abnormal sample data is abnormal data.

[0067] Optionally, the preset conditions include but are not limited to the stability of the parameters and the range of the parameters. For example, the weights and bias terms are within a reasonable range.

[0068] Optionally, after determining that the abnormal data is located in the target sub-data set, a segmentation operation is performed on the target sub-data set to segment the target sub-data set into subsets with smaller data volumes, and the abnormal data is searched for in the subsets with smaller data volumes.

[0069] This embodiment avoids redundant analysis of all data by performing anomaly detection only on those sub-data sets whose evaluation index differences are less than a preset threshold, thereby achieving the purpose of reducing the computational cost. At the same time, the setting of preset conditions can avoid performing anomaly detection when the model parameters are in a poor state, thereby achieving the purpose of ensuring the efficiency and reliability of the detection process.

[0070] The present invention will be described below in conjunction with specific embodiments:

[0071] This embodiment takes a binary classification problem, where the first data is a predicted label, and only one sample data in the sample data set has an incorrect predicted label as an example for explanation. Figure 3 This is a flowchart of a method for determining abnormal sample data in a specific embodiment of the present application, comprising the following steps:

[0072] S302, performing a classification operation on the multiple prediction labels to obtain two sets of prediction labels: prediction labels with a prediction result of 0 and prediction labels with a prediction result of 1.

[0073] S304: Divide the above two types of prediction label sets into two equal parts (or n equal parts) to obtain prediction label subsets.

[0074] S306, convert the prediction results 0 of "the first copy of predicted label 0" and "the second copy of predicted label 0" to 1; convert the prediction results 1 of "the first copy of predicted label 1" and "the second copy of predicted label 1" to 0. Calculate the F1 value before conversion to obtain the initial evaluation index. Calculate the F1 value after conversion to obtain the target evaluation index.

[0075] S308, if the F1 values ​​of the prediction label subsets after the two conversion prediction results (for example, the F1 values ​​are both 0.7) are consistent with the evaluation index of the prediction label before the conversion (for example, the F1 values ​​are both 0.9), then the error sample is not in the above two prediction label subsets; if the evaluation index is inconsistent, then the error sample exists in the prediction label subset with a smaller decrease than the initial evaluation index. The prediction label subset with error samples can be further divided until the error sample is found. The time complexity is O(log 2 N). Figure 4 As shown, due to F1 1 Less than F1 2 , so the misclassified samples are in the first prediction label subset with a prediction result of 0; due to F1 3 With F14 They are the same, so the misclassified error samples are not in the set of predicted labels with a predicted result of 1.

[0076] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.

[0077] In this embodiment, a device for determining abnormal sample data is also provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0078] Figure 5 is a structural block diagram of a device for determining abnormal sample data according to an embodiment of the present application, as Figure 5 shown. The device includes:

[0079] A first acquisition module 502, configured to obtain a first data set output by the initial network model based on the input sample data set during the process of training the initial network model using the sample data set. One sample data in the sample data set corresponds to one first data in the first data set;

[0080] A first execution module 504, configured to perform a classification operation on multiple first data included in the first data set to obtain N second data sets, where N is a natural number greater than 1;

[0081] A first conversion module 506, configured to perform data conversion on each type of data included in the N second data sets respectively according to the data category in each type of second data set to obtain N third data sets;

[0082] A first search module 508, configured to search for abnormal sample data from the sample data set by using the N second data sets, the N third data sets, and the model parameters of the initial network model, where the data output by the initial network model corresponding to the abnormal sample data is abnormal data.

[0083] In an exemplary embodiment, the first execution module includes: a first determination sub-module, configured to, when the first data is a prediction label, determine data categories of the plurality of prediction labels respectively according to attribute values of the plurality of prediction labels; a first execution sub-module, configured to perform the classification operation on the plurality of prediction labels according to the data categories of the plurality of prediction labels to obtain N sets of the second data.

[0084] In an exemplary embodiment, the first conversion module includes: a second execution sub-module, configured to perform a splitting operation on each of the N sets of the second data according to the number of the second data included in the set of the second data, and split each set of the second data into M sets of second sub-data, where M is a natural number greater than 1; a third execution sub-module, configured to perform a modification operation on each of the M sets of the second sub-data respectively to perform the data conversion on the attribute values of the second data in the M sets of the second sub-data to obtain N sets of the third data; where the modification operation includes: modifying the attribute values of the data included in the M sets of sub-data in the second target data set according to the attribute values of the data included in another data set to obtain third data, where the another data set and the second target data set are both data sets in the N sets of the second data, and the attribute value of the third data is the same as the attribute value of the data in the another data set.

[0085] In an exemplary embodiment, the second sub-execution module includes: a first determination unit, configured to determine the number of the second data in each set of the second data to obtain N numbers of the second data; a second determination unit, configured to determine a quantity ratio based on the N numbers of the second data; a first splitting unit, configured to split each set of the second data into M sets of the second sub-data based on the quantity ratio.

[0086] In an exemplary embodiment, the first search module includes: a fourth execution sub-module, configured to perform the following operations on each set of the second data to search for abnormal sample data from the sample data set: calculate M initial evaluation indexes of the M sets of the second sub-data according to the attribute values of the second data in the M sets of the second sub-data; calculate M target evaluation indexes of the M sets of the third sub-data in the third data set according to the attribute values of the third data in the M sets of the third sub-data; search for the abnormal sample data in the sub-data set based on the M initial evaluation indexes, the M target evaluation indexes, and the model parameters of the initial network model.

[0087] In an exemplary embodiment, the above-mentioned first search module includes: a first calculation sub-module, configured to calculate the difference between the above-mentioned target evaluation index corresponding to the target sub-data set and the above-mentioned initial evaluation index, where the target sub-data set is any one of the above-mentioned M third sub-data sets; a first search sub-module, configured to search for abnormal data in the target sub-data set when the difference is less than a preset threshold and the model parameters of the above-mentioned initial network model meet the preset conditions; a second determination sub-module, configured to determine the above-mentioned abnormal sample data based on the above-mentioned abnormal data, where the abnormal sample data is sample data that does not meet the training rules of the initial network model, and the data output by the above-mentioned initial network model corresponding to the abnormal sample data is abnormal data.

[0088] An embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above-mentioned method embodiments.

[0089] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is set to execute the steps in any one of the above-mentioned method embodiments when running.

[0090] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.

[0091] An embodiment of the present application also provides an electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is set to run the computer program to execute the steps in any one of the above-mentioned method embodiments.

[0092] In an exemplary embodiment, the above-mentioned electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.

[0093] The specific examples in this embodiment may refer to the examples described in the above-mentioned embodiments and exemplary embodiments, and will not be repeated here.

[0094] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.

[0095] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for determining abnormal sample data, characterized in that: include: In the process of training the initial network model using the sample data set, obtaining a first data set output by the initial network model based on the sample data set input, wherein one sample data in the sample data set corresponds to one first data in the first data set; Performing a classification operation on the plurality of first data included in the first data set to obtain N types of second data sets, where N is a natural number greater than 1; According to the data category in each type of the second data set, respectively perform data conversion on each type of data included in the N types of the second data sets to obtain N types of third data sets; Abnormal sample data is searched from the sample data set using the N types of the second data sets, the N types of the third data sets and the model parameters of the initial network model, wherein the data output by the initial network model corresponding to the abnormal sample data is the abnormal data.

2. The method according to claim 1, characterized in that Performing a classification operation on the plurality of first data included in the first data set to obtain N types of second data sets, including: In the case where the first data is a predicted label, determining data categories of the plurality of predicted labels according to attribute values ​​of the plurality of predicted labels respectively; The classification operation is performed on the multiple predicted labels according to their data categories to obtain N types of the second data sets.

3. The method according to claim 1, characterized in that According to the data category in each type of the second data set, data conversion is performed on each type of data included in the N types of the second data sets to obtain N types of third data sets, including: According to the amount of second data included in the second data set, a segmentation operation is performed on N types of the second data set, and each type of the second data set is segmented into M second sub-data sets, where M is a natural number greater than 1; Performing modification operations on the M second subset data sets respectively, so as to perform the data conversion on the attribute values ​​of the second data in the M second subset data sets, to obtain N types of the third data sets; Wherein, the modification operation includes: modifying the attribute values ​​of the data included in the M sub-data sets in the second target data set according to the attribute values ​​of the data included in the other data sets, to obtain third data, wherein the other data sets and the second target data set are both data sets in the N types of the second data sets, and the attribute values ​​of the third data are the same as the attribute values ​​of the data in the other data sets.

4. The method according to claim 3, characterized in that According to the amount of second data included in the second data set, a segmentation operation is performed on N types of the second data sets, and each type of the second data set is segmented into M second sub-data sets, including: Determine the amount of the second data in each type of the second data set to obtain N amounts of second data; Determining a quantity ratio based on the N quantities of the second data; Based on the quantity ratio, each type of the second data set is divided into M second sub-data sets.

5. The method according to claim 3, characterized in that: Using the N types of the second data sets, the N types of the third data sets, and the model parameters of the initial network model, searching for abnormal sample data from the sample data set includes: The following operations are performed for each type of the second data set to find abnormal sample data from the sample data set: Calculating M initial evaluation indicators of the M second sub-data sets according to the attribute values ​​of the second data in the M second sub-data sets; Calculating M target evaluation indicators of the M third sub-data sets according to the attribute values ​​of the third data in the M third sub-data sets in the third data set; Based on the M initial evaluation indicators, the M target evaluation indicators and the model parameters of the initial network model, the abnormal sample data is searched in the sub-data set.

6. The method according to claim 5, characterized in that Based on the M initial evaluation indicators, the M target evaluation indicators and the model parameters of the initial network model, searching for the abnormal sample data in the sub-data set includes: Calculating the difference between the target evaluation index corresponding to the target sub-data set and the initial evaluation index, wherein the target sub-data set is any one of the M third sub-data sets; When the difference is less than a preset threshold and the model parameters of the initial network model meet preset conditions, searching for abnormal data in the target sub-data set; Based on the abnormal data, the abnormal sample data is determined, wherein the abnormal sample data is sample data that does not meet the initial network model training rules, and the data output by the initial network model corresponding to the abnormal sample data is abnormal data.

7. A device for determining abnormal sample data, characterized in that: include: A first acquisition module is used to acquire, during the process of training the initial network model using the sample data set, a first data set output by the initial network model based on the sample data set as input, wherein one sample data in the sample data set corresponds to one first data in the first data set; A first execution module, configured to perform a classification operation on a plurality of the first data included in the first data set to obtain N types of second data sets, wherein N is a natural number greater than 1; A first conversion module, configured to perform data conversion on each type of data included in the N types of the second data sets according to the data category in each type of the second data sets, to obtain N types of third data sets; The first search module is used to use the N types of the second data sets, the N types of the third data sets and the model parameters of the initial network model to search for abnormal sample data from the sample data set, wherein the data output by the initial network model corresponding to the abnormal sample data is the abnormal data.

8. A computer program product, characterized in that The method comprises a computer program, wherein when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method described in any one of claims 1 to 6 when executed by a processor.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 6 are implemented.