Data processing method, electronic device, storage medium, and program product

By dynamically adjusting the probability threshold of the machine learning model and combining statistical methods with model iterative updates, the problem of decreased accuracy of machine learning models due to changes in data distribution during sample labeling is solved, achieving an efficient and reliable automated labeling process.

WO2025222334A1PCT designated stage Publication Date: 2025-10-30BEIJING ZITIAO NETWORK TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/089110
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

In existing technologies, machine learning models experience a decline in prediction accuracy during sample labeling due to training data bias and changes in data distribution. Furthermore, simply increasing the probability threshold may reduce recall, affecting the consistency and efficiency of labeling results.

Method used

By dynamically adjusting the probability threshold of the machine learning model, and optimizing the probability threshold of each category to maximize the number of positive examples based on confidence and precision requirements, combined with statistical methods and iterative model updates, the model adapts to changes in data distribution, thereby improving the reliability and recall of annotations.

Benefits of technology

It achieves an automated annotation process that improves the annotation accuracy and robustness of machine learning models, reduces annotation costs, and adapts to dynamic data changes while meeting confidence and recall requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024089110_30102025_PF_FP_ABST
    Figure CN2024089110_30102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of data processing, and relates to a data processing method, an electronic device, a storage medium, and a program product. The data processing method comprises: using a first machine learning model to process each sample in a first sample set to obtain a predicted probability that each sample is classified into each of one or more categories; and for each category, determining a probability threshold corresponding to the category, so that when a confidence level corresponding to the category is not lower than a confidence level threshold, the number of positive examples of the category is maximized, wherein the probability threshold is used for determining, on the basis of the predicted probability of each sample, a category to which each sample belongs, and the confidence level corresponding to the category is a confidence level at which the true precision of classification based on the probability threshold satisfies a precision condition.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing methods, electronic devices, storage media and software products Technical Field

[0001] This disclosure relates to the field of data processing, and in particular to a data processing method, electronic device, storage medium, and program product. Background Technology

[0002] In business scenarios such as content moderation and data annotation, it is necessary to accurately classify or label samples. While manual review can be used to label samples, this method is time-consuming, labor-intensive, costly, and susceptible to subjective influences, leading to inconsistent results. To improve annotation efficiency and consistency, machine learning models can be used to automate the annotation process.

[0003] Summary of the Invention

[0004] This summary section is provided to briefly introduce the concepts, which will be described in detail in the subsequent detailed description section. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] According to some embodiments of this disclosure, a data processing method is provided, comprising: processing each sample in a first sample set using a first machine learning model to obtain a predicted probability that each sample is classified into each of one or more categories; for each category, determining a probability threshold corresponding to the category, such that the number of positive examples of the category is maximized when the confidence level corresponding to the category is not lower than the confidence level threshold, wherein the probability threshold is used to determine the category to which each sample belongs based on the predicted probability of each sample, and the confidence level corresponding to the category is the confidence level that the true accuracy of classification based on the probability threshold satisfies the accuracy condition.

[0006] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to perform a data processing method of any embodiment of the present disclosure based on instructions stored in the memory.

[0007] According to some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the data processing method of any embodiment described in the present disclosure.

[0008] According to some embodiments of the present disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to implement the data processing method of any embodiment of the present disclosure.

[0009] According to some embodiments of the present disclosure, a computer program is provided, comprising: instructions that, when executed by a processor, cause the processor to perform a data processing method according to any embodiment of the present disclosure.

[0010] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0011] Preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The accompanying drawings, which are included to provide a further understanding of the present disclosure, and which, together with the following detailed description, are incorporated in and form a part of this specification and are used to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and are not intended to limit the present disclosure. In the drawings:

[0012] Figure 1 shows a schematic flowchart of a data processing method according to some embodiments of the present disclosure.

[0013] Figure 2 shows a flowchart illustrating a method for determining a probability threshold according to some embodiments of the present disclosure.

[0014] Figure 3 illustrates a data processing flow diagram according to some embodiments of the present disclosure.

[0015] Figure 4 shows a schematic diagram of the structure of a data processing apparatus according to some embodiments of the present disclosure.

[0016] Figure 5 shows a schematic diagram of the structure of an electronic device according to some embodiments of the present disclosure.

[0017] Figure 6 shows a schematic diagram of the structure of a computer system according to some embodiments of the present disclosure.

[0018] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation

[0019] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. However, it is obvious that the described embodiments are only some embodiments of this disclosure, and not all embodiments. The following description of the embodiments is merely illustrative and is in no way intended to limit this disclosure or its application or use. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.

[0020] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.

[0021] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Furthermore, as used in this disclosure, the term "including" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Therefore, "comprising" and "including" are synonymous. The term "based on" means "at least partially based on".

[0022] Throughout this specification, the terms "one embodiment," "some embodiments," or "embodiment" mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the appearance of the phrases "in one embodiment," "in some embodiments," or "in an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment, but may refer to the same embodiment.

[0023] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.

[0024] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0025] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0026] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0027] Research revealed several issues with using machine learning models for automated labeling. Model predictions can be inaccurate due to factors such as bias in the training data and the model's generalization ability. Furthermore, labeling (classifying) samples requires combining the model's predicted probability for each category with a probability threshold. However, consistently using the same threshold may not adapt to dynamically changing data distributions, impacting prediction accuracy. Simply increasing the probability threshold to exclude potential false positives may reduce recall.

[0028] To improve the accuracy of model predictions, this disclosure proposes a method for dynamically adjusting thresholds to adapt to data changes and meet the requirements for precision and confidence, thereby improving the reliability of automatic labeling. In some embodiments of this disclosure, after the machine learning model used for sample classification has been trained, the probability threshold of the machine learning model is determined using the sample set, so as to maintain or improve the recall rate as much as possible while meeting the requirements for precision and confidence, in order to meet business needs.

[0029] Below, we will first define some of the technical terms used in this disclosure.

[0030] A test set is a collection of samples used to test a machine learning model. During the training and testing of a machine learning model, pre-labeled samples can be obtained and divided into training and test sets. For example, in a multi-class classification problem, the class labels are n. i ∈N(||N||=C), where N represents a set of one or more categories, and C represents the number of categories. The labeled samples are divided into training sets T(||T||=N). T) and test set S(||S||=N S ), N T N represents the number of samples in the training set. S This represents the number of samples in the test set. A machine learning model can be trained based on the training set T (||T||=N). T ) and test set S(||S||=N S ) merged into a large dataset D ALL =T∪S, which contains N ALL =N T +N S One sample. (Can be used) Represents category n i In the union D of the training set and the test set ALL The probability distribution in the matrix. Then, for each category n... i , in and Representing categories n i The number of samples in the training and test sets. When the sample size is large enough... Can be used as category n i An approximation of the true distribution.

[0031] Machine learning models are used to process input samples and determine output results. In some embodiments of this disclosure, machine learning models are used to classify samples. For example, determining the probability that a sample belongs to each category. The machine learning model can be a neural network model. Other types of models may also be used as needed, which will not be elaborated here.

[0032] Probability thresholds are used to classify samples based on probabilities predicted by a machine learning model. In some embodiments of this disclosure, probability thresholds can be set for each of one or more categories. A sample is determined to belong to a category if the probability of it belonging to that category is higher than the probability threshold corresponding to that category. For example, some machine learning models may include a softmax layer, or a softmax layer may be chained after the machine learning model. Alternatively, the machine learning model may process the samples first, and then the softmax layer may process the input vectors to obtain the probability of the sample belonging to each category. Based on the probability threshold (or softmax threshold), the predicted category to which the sample belongs can be further determined.

[0033] A positive example within a category refers to a sample that is assigned to that category based on the predictions of a machine learning model and a probability threshold. Positive examples include true positives and false positives. A true positive is a positive example whose pre-labeling matches the classification result; for example, a sample is determined to belong to category 1 based on a machine learning model and a probability threshold, and its label also indicates that it belongs to category 1. A false positive is a positive example whose labeling does not match the classification result; for example, a sample is determined to belong to category 1 based on a machine learning model and a probability threshold, but its label indicates that it does not belong to category 1.

[0034] Precision for a given category is the ratio of the number of true positives to the total number of positives in that category. For example, if 100 samples belong to category A, then the number of positives in category A is 100. Of these 100 samples, 80 are pre-labeled as category A (i.e., they actually belong to category A), and 20 are pre-labeled as one or more other categories (i.e., they do not actually belong to category A). Therefore, the precision for category A is 0.8.

[0035] Confidence level refers to a measure of the reliability of an inference result. For example, when making predictions, there may be a certain requirement for the prediction accuracy of a certain category, but this accuracy requirement should also be accompanied by a certain level of confidence.

[0036] An embodiment of the sample processing method of this disclosure is described below with reference to FIG1.

[0037] Figure 1 shows a schematic flowchart of a data processing method according to some embodiments of the present disclosure. As shown in Figure 1, the sample processing method of this embodiment includes steps S102 to S104.

[0038] In step S102, each sample in the first sample set is processed using the first machine learning model to obtain the predicted probability that each sample is classified into each of one or more categories.

[0039] The first machine learning model is used to process samples in the first input sample set and determine the output result. In some embodiments of this disclosure, the machine learning model is used to classify the samples. The first machine learning model can process the input samples to obtain vectors corresponding to the samples, and then pass them through activation layers such as softmax to obtain the probability that the sample belongs to each category. The first machine learning model can be a neural network model. Other types of models can also be used as needed, which will not be elaborated here.

[0040] The first machine learning model is used to classify or label samples. It can be a binary classification model, a multi-class classification model, a multi-label classification model, or a hierarchical classification model. Among them, multi-label classification and hierarchical classification can also be regarded as a type of multi-class classification.

[0041] In some embodiments, the machine learning model includes a softmax layer, the predicted probability of each sample being classified into one or more categories is output by the softmax layer, and the probability threshold is a softmax threshold.

[0042] The first sample set is a collection of multiple samples. It could be, for example, a test set. Thus, after the first machine learning model has been trained, it can be tested using the test set, and a probability threshold can be determined using the test set.

[0043] In step S104, for each category, a probability threshold corresponding to that category is determined such that the number of positive examples for that category is maximized while the confidence level corresponding to that category is not lower than the confidence threshold.

[0044] Probability thresholds are used to determine the category to which each sample belongs based on its predicted probability. For example, for each category, a sample belongs to that category if its predicted probability is not lower than the category's probability threshold; a sample does not belong to that category if its predicted probability is lower than the category's probability threshold. The number of positive examples for each category refers to the number of samples assigned to that category.

[0045] The confidence level corresponding to the category is the confidence level of the true accuracy (or ground truth accuracy) of the classification based on the probability threshold, which satisfies the accuracy condition. The accuracy condition, for example, is that the true accuracy is greater than a certain level, meaning the probability threshold allows the predicted accuracy to be greater than a certain level, and this accuracy requirement has a certain level of confidence.

[0046] True precision is the overall precision determined based on the first machine learning model and probability thresholds. It can be measured with a sufficiently large sample size, but it is difficult to obtain. In contrast, observational precision (or empirical precision or estimating precision) is based on observations from the first sample set. Observational precision is easier to obtain, but it may differ from true precision. However, observational precision affects the number of observed positive examples and, consequently, the overall observational precision. Therefore, if true precision can be determined, the corresponding number of positive examples and the observational precision can be determined as well.

[0047] Formula (1) exemplarily illustrates an objective function, and Formula (2) exemplarily illustrates the constraints of that objective function. Those skilled in the art can make appropriate modifications to these formulas as needed, which will not be described further in this disclosure. confidence (pr) actual ≥x)≥confidence threshold (2)

[0048] In the above formula, softmax_threshold represents the softmax threshold, and can also be other types of probability thresholds as needed; n pos This indicates the number of positive examples determined based on the softmax_threshold; pr actual This refers to the true accuracy rate, also known as ground truth accuracy rate; confidence (pr actual ≥x) represents the confidence level that the true precision is greater than x, where x is the value to be determined; threshold This represents the confidence threshold, which is a known value.

[0049] Since the determination of the probability threshold affects the number of positive examples in each category in the prediction results, i.e., there is a correlation between the probability threshold and the number of positive examples, maximizing the number of positive examples in each category can be taken as the solution objective, with the confidence level not being lower than the confidence threshold as a constraint. Since the confidence threshold is known, the precision condition that satisfies the confidence requirement can be obtained. Based on the precision condition, the number of positive examples and the observation precision can be determined, and then the probability threshold can be further determined.

[0050] Once the probability thresholds are determined, the samples to be processed can be classified based on the first machine learning model and the probability thresholds corresponding to each category.

[0051] During the labeling process of the samples to be processed, the first machine learning model is used to process the samples to obtain the probability of the samples belonging to each category. Then, based on the probability threshold determined in step S104, the category to which the samples belong is determined, thereby performing classification. Thus, the processing results of the samples to be processed can be obtained while meeting the requirements of confidence and recall.

[0052] The above embodiments can precisely and automatically determine the probability threshold for each category based on the data distribution in the first sample set. Therefore, the embodiments of this disclosure can scientifically maintain the confidence level within a business-acceptable range, significantly improving robustness, and maximizing recall while meeting confidence requirements to satisfy business needs.

[0053] In some embodiments, the confidence level corresponding to the category is determined based on a target probability, which is the probability of obtaining the number of positive examples and the observation precision determined based on the first sample set, provided that the probability threshold has the true precision.

[0054] The target probability is a priori conditional on the true accuracy. For any given class, it can be expressed as P(n pos ,pr eval |pr actual ) represents the target probability, where n pos pr represents the number of positive examples of that class in the prediction results for the first sample set. eval pr represents the observational precision for that category in the prediction results for the first sample set. actual This represents the true accuracy rate. The target probability can be determined, for example, by formula (3).

[0055] In formula (3), the number of true positives among the positives can be determined based on the number of positives predicted using the first sample set and the observation precision. That is, when the number of positives n is determined... pos and observation accuracy pr eval Then, the number of true positives can be calculated by multiplying the two. Furthermore, by considering the possibilities of all combinations of true positives, and by determining the true probability of each true positive and the true probability of each false positive based on the true accuracy rate, the target probability can be determined.

[0056] When processing formula (3), the computational cost of factorial operations is relatively high when the sample size is large and the number of positive examples is also relatively large. To improve computational efficiency, Stirling's approximation can be used to simplify the factorial operation. Of course, those skilled in the art can also use other methods to perform the calculation, and this disclosure does not limit this.

[0057] In some embodiments, the confidence level corresponding to the category is determined based on a first sum and a second sum. The first sum is the sum of the target probability values ​​that satisfy the precision condition, and the second sum is the sum of all values ​​of the target probability. The first and second sums can be determined by integrating the true precision rate. That is, the first sum is the integration of the target probability over the interval satisfying the precision condition, and the second sum is the integration of the target probability over the entire interval [0,1]. In other words, the confidence level can be determined by the area under the function curve of the target probability. For example, with the true precision rate on the horizontal axis and the target probability on the vertical axis, a value x is determined such that the ratio of the area under the curve corresponding to the horizontal axis greater than this value to the total area under the curve equals the confidence threshold. This allows us to solve for the precision condition: the true precision rate is greater than x. Therefore, the confidence level for a true precision rate greater than x can be greater than the confidence threshold.

[0058] In some embodiments, the probability density function of the target probability follows a conjugate prior distribution. The conjugate prior distribution can be, for example, a beta distribution. An exemplary formula for the confidence level is shown in Equation (4), where α and β are parameters in the beta distribution. Initially, α and β can be set to initial values.

[0059] Since the modeling process may be ongoing, multiple probability thresholds corresponding to each category can be determined repeatedly, and multiple observation precipitates determined by these probability thresholds can be established. For example, during K-fold cross-validation, parameters can be updated using the observed precipitates from multiple evaluations. Then, the parameters of the probability density function distribution of the true precision are determined using these multiple observation precipitates. Thus, the updated parameters allow the distribution to better reflect the actual true precision, thereby improving the accuracy and reliability of threshold determination through multiple iterations.

[0060] The following describes an embodiment of the method for determining the probability threshold with reference to Figure 2.

[0061] Figure 2 illustrates a flowchart of a method for determining a probability threshold according to some embodiments of the present disclosure. As shown in Figure 2, the method for determining the probability threshold in this embodiment includes steps S202 to S204. These steps can be applied to any category.

[0062] In step S202, for each category, a precision condition is determined based on the confidence level corresponding to that category and the confidence threshold. As explained in the foregoing embodiments, there is a certain correlation between confidence level and true precision. When the confidence threshold is known, the precision condition can be solved. For example, the relationship between confidence level and true precision is monotonic, so a bisection method can be used to determine the true precision corresponding to a specified confidence threshold, i.e., the precision threshold x in the precision condition.

[0063] In step S204, a probability threshold is determined that maximizes the number of positive examples while maintaining an observation precision greater than a precision threshold. This probability threshold is then used as the probability threshold corresponding to that category. For example, the optimal solution can be determined by iterating through all possible values ​​of the probability threshold. Of course, other methods can also be used, which will not be elaborated here.

[0064] The above embodiments describe methods for determining the probability threshold corresponding to each category based on a first sample set. These embodiments dynamically determine the threshold using an inference algorithm, enabling real-time processing of large amounts of data without inefficiency due to computational resource consumption, thus improving the performance of automated processes.

[0065] As needed, the method of the above embodiments can be executed multiple times to continuously update the probability threshold according to changes in the data. For example, in response to an update of the first sample set, the probability threshold corresponding to each category is redefined. Thus, the embodiments of this disclosure can adapt to changes in data distribution, so as to apply the first machine learning model to different application scenarios or at different times, and achieve better prediction results.

[0066] The embodiments disclosed herein can be applied to scenarios such as automatic data labeling and data review. These two scenarios are described below as examples.

[0067] In the scenario of automatic data labeling, the method described in the above embodiment is first used to determine the probability threshold corresponding to each category based on a first sample set. Then, based on the first machine learning model and the probability threshold corresponding to each category, the classification result of the samples in the second sample set to be processed is determined, where the second sample set is the training data of the second machine learning model; according to the classification result, the samples in the second sample set are labeled. Therefore, the labeling results of the samples in the labeled second sample set can have a higher confidence level, enabling the second machine learning model to achieve better training results.

[0068] In the data review scenario, the method described in the above embodiment is first used to determine the probability threshold corresponding to each category based on a first sample set. Then, based on the first machine learning model and the probability threshold corresponding to each category, the samples to be processed are classified, and the classification results include "approved" and "unapproved". In the case of "unapproved", the data can be output to the data stream for manual review.

[0069] The samples to be reviewed are, for example, online samples. If the online sample's labeling result indicates approval, the online sample is output online. If the online sample's labeling result indicates disapproval, the sample pool is updated using the online sample. This sample pool is used to train or test the first machine learning model. Thus, the first sample set can be updated using the labeled samples, and the first machine learning model can be retrained, tested, or the probability threshold re-determined using the updated first sample set.

[0070] Figure 3 illustrates a data processing flow diagram according to some embodiments of the present disclosure. In Figure 3, the data flow controller selects samples from the sample pool according to demand and sends them to the training set sampling stream and the test set sampling stream, respectively. The training set sampling stream is used to train the machine learning model, and the test set sampling stream is used to test the trained model. These samples are unlabeled, so manual data labeling is required so that the labeled samples can be used for model training and testing. In the model training stream, the model can be trained multiple times, and the multiple training model versions are labeled, for example, i, i+1, i+2, etc. After completing the training process of each version of the model, the probability threshold can be determined using the methods of embodiments of the present disclosure.

[0071] The updated model can replace the existing online model, and the new probability thresholds determined based on the updated model can also replace the original thresholds. Then, an online machine review stream can be provided through a data flow controller. The online model processes the samples in the online machine review stream to obtain predicted probabilities, and then uses these probabilities in conjunction with the new probability thresholds to control the threshold and determine whether a sample passes the review. If it passes the review, it is output online; if it fails, it undergoes online human review. Once the human review is passed, it is output online, and the sample is simultaneously returned to the sample pool.

[0072] The above embodiments employ a data-centric model iteration mechanism, which integrates real-time feedback into the training sample control flow, enabling on-demand labeling and effectively reducing labeling costs. Furthermore, by introducing statistical derivation, it is possible to perform refined and automated control over different target labels based on the distribution of test samples and other relevant inputs. This not only ensures the stability of accuracy but also scientifically maintains the confidence interval within an acceptable business range, significantly improving robustness.

[0073] The above describes relevant embodiments of the data processing method of this disclosure. The following describes some related apparatus and devices in some embodiments of this disclosure.

[0074] Figure 4 shows a schematic diagram of the structure of a data processing apparatus according to some embodiments of the present disclosure. As shown in Figure 4, the data processing apparatus 40 of this embodiment includes: a prediction module 402, configured to process each sample in a first sample set using a first machine learning model to obtain a predicted probability that each sample is classified into each of one or more categories; and a determination module 404, configured to determine a probability threshold corresponding to each category for each category, such that the number of positive examples of a category is maximized when the confidence level corresponding to the category is not lower than the confidence level threshold, wherein the probability threshold is used to determine the category to which each sample belongs based on the predicted probability of each sample, and the confidence level corresponding to the category is the confidence level that the true accuracy of classification based on the probability threshold satisfies the accuracy condition.

[0075] In some embodiments, the confidence level corresponding to the category is determined based on the target probability, which is the probability of obtaining the number of positive examples and the observation precision determined based on the first sample set, provided that the probability threshold has the true precision.

[0076] In some embodiments, the confidence level corresponding to a category is determined based on a first sum and a second sum, wherein the first sum is the sum of the target probability values ​​that satisfy the precision condition, and the second sum is the sum of all values ​​of the target probability.

[0077] In some embodiments, the probability density function of the target probability follows a conjugate prior distribution.

[0078] In some embodiments, the determining module 404 is further configured to: for each category, determine a precision condition based on the confidence level corresponding to the category and a confidence threshold; and determine a probability threshold that maximizes the observation precision and the number of positive examples, as the probability threshold corresponding to the category.

[0079] In some embodiments, the accuracy condition is that the true accuracy is greater than a specified value.

[0080] In some embodiments, the data processing apparatus 40 further includes an update module configured to redetermine the probability threshold corresponding to each category in response to an update of the first sample set.

[0081] In some embodiments, the data processing apparatus 40 further includes a parameter determination module configured to: for each category, acquire multiple determined probability thresholds corresponding to the category; determine multiple observation accuracies determined with the multiple probability thresholds; and use the multiple observation accuracies to determine the parameters of the distribution of the probability density function of the true accuracy.

[0082] In some embodiments, the machine learning model includes a softmax layer, the predicted probability is output by the softmax layer, and the probability threshold is the softmax threshold.

[0083] In some embodiments, the data processing apparatus 40 further includes a classification module configured to classify the samples to be processed based on a first machine learning model and a probability threshold corresponding to each category.

[0084] In some embodiments, the classification module is further configured to: determine the classification result of the samples in the second sample set to be processed based on the first machine learning model and the probability threshold corresponding to each category, wherein the second sample set is the training data of the second machine learning model; and label the samples in the second sample set according to the classification result.

[0085] In some embodiments, the sample to be processed is an online sample to be reviewed, and the data processing device 40 further includes a review module configured to: output the online sample online in response to the online sample's tagging result indicating that the review has passed.

[0086] In some embodiments, the review module is further configured to: update the sample pool using the online sample in response to the online sample's labeling result indicating that the review is not passed, the sample pool being used to train or test the first machine learning model.

[0087] It should be noted that the above-described units are merely logical modules divided according to their specific functions, and are not intended to limit the specific implementation method. For example, they can be implemented in software, hardware, or a combination of both. In actual implementation, the above-described units can be implemented as independent physical entities, or they can be implemented by a single entity (e.g., a processor (CPU or DSP, etc.), integrated circuit, etc.). Furthermore, the units shown in the accompanying drawings with dashed lines indicate that these units may not actually exist, and the operations / functions they perform can be implemented by the processing circuitry itself.

[0088] In addition, although not shown, the device may also include a memory that can store various information generated by the device and its constituent units during operation, programs and data used for operation, data to be transmitted by the communication unit, etc. The memory can be volatile memory and / or non-volatile memory. For example, the memory may include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Of course, the memory may also be located outside the device. Optionally, although not shown, the device may also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in a manner known in the art, such as including communication components such as antenna arrays and / or radio frequency links, various types of interfaces, communication units, etc. These will not be described in detail here. Furthermore, the device may also include other components not shown, such as radio frequency links, baseband processing units, network interfaces, processors, controllers, etc. These will not be described in detail here.

[0089] Some embodiments of this disclosure also provide an electronic device. Figure 5 shows a schematic diagram of the structure of an electronic device according to some embodiments of this disclosure. For example, in some embodiments, the electronic device 5 can be various types of devices, such as mobile terminals including but not limited to mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. For example, the electronic device 5 may include a display panel for displaying data and / or execution results utilized in the scheme according to this disclosure. For example, the display panel can be of various shapes, such as a rectangular panel, an elliptical panel, or a polygonal panel. In addition, the display panel can be not only a planar panel, but also a curved panel, or even a spherical panel.

[0090] As shown in FIG. 5, the electronic device 5 of this embodiment includes a memory 51 and a processor 52 coupled to the memory 51. It should be noted that the components of the electronic device 5 shown in FIG. 5 are merely exemplary and not limiting; the electronic device 5 may also have other components depending on the actual application requirements. The processor 52 can control other components in the electronic device 5 to perform desired functions.

[0091] In some embodiments, memory 51 is used to store one or more computer-readable instructions. When processor 52 executes the computer-readable instructions, the computer-readable instructions are executed by processor 52 to implement the method according to any of the above embodiments. For specific implementations and related explanations of the various steps of the method, please refer to the above embodiments; repeated details will not be elaborated here.

[0092] For example, processor 52 and memory 51 can communicate with each other directly or indirectly. For example, processor 52 and memory 51 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. Processor 52 and memory 51 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0093] For example, processor 52 can be embodied in various suitable processors, processing devices, such as central processing unit (CPU), graphics processing unit (GPU), network processor (NP), etc.; it can also be digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) can be an x86 or ARM architecture, etc. For example, memory 51 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Memory 51 can include, for example, system memory, which stores, for example, the operating system, application programs, boot loader, database, and other programs. Various application programs and various data can also be stored in the storage medium.

[0094] Furthermore, according to some embodiments of this disclosure, various operations / processes according to this disclosure, implemented via software and / or firmware, can install programs constituting the software from a storage medium or network onto a computer system with a dedicated hardware architecture, such as the computer system 60 shown in FIG. 6. When various programs are installed, the computer system is capable of performing various functions, including those described above. FIG. 6 shows a schematic diagram of the structure of a computer system according to some embodiments of this disclosure.

[0095] In Figure 6, the Central Processing Unit (CPU) 601 performs various processes according to a program stored in the Read-Only Memory (ROM) 602 or a program loaded from the storage portion 608 into the Random Access Memory (RAM) 603. The RAM 603 also stores data required as needed when the CPU 601 performs various processes, etc. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. The ROM 602, RAM 603, and storage portion 608 can be various forms of computer-readable storage media, as described below. It should be noted that although the ROM 602, RAM 603, and storage device 608 are shown separately in Figure 6, one or more of them may be combined or located in the same or different memories or storage modules.

[0096] CPU 601, ROM 602 and RAM 603 are interconnected via bus 604. Input / output interface 605 is also connected to bus 604.

[0097] The following components are connected to the input / output interface 605: input section 606, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 607, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 608, including hard disks, magnetic tapes, etc.; and communication section 609, including network interface cards such as LAN cards, modems, etc. The communication section 609 allows communication processing to be performed via a network such as the Internet. It is readily understood that although the various devices or modules in the computer system 60 shown in Figure 6 communicate via bus 604, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.

[0098] As needed, drive 610 is also connected to input / output interface 605. Removable media 611, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 610 as needed, so that computer programs read from them can be installed into storage section 608 as needed.

[0099] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or from a storage medium such as removable media 611.

[0100] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the CPU 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0101] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0102] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0103] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the method of any of the above embodiments. For example, the instructions may be embodied in computer program code.

[0104] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0105] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0106] The modules, components, or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules, components, or units do not necessarily constitute a limitation on the module, component, or unit itself.

[0107] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0108] The above description is merely an embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0109] Many specific details are set forth in the description provided herein. However, it is understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of the description.

[0110] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0111] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. A data processing method, comprising: The first machine learning model is used to process each sample in the first sample set to obtain the predicted probability that each sample is classified into each of one or more categories. For each category, a probability threshold is determined such that the number of positive examples for the category is maximized when the confidence level corresponding to the category is not lower than the confidence threshold. The probability threshold is used to determine the category to which each sample belongs based on the predicted probability of each sample, and the confidence level corresponding to the category is the confidence level that the true accuracy of classification based on the probability threshold satisfies the accuracy condition.

2. The data processing method according to claim 1, wherein, The confidence level corresponding to the category is determined based on the target probability, which is the probability of obtaining the number of positive examples and the observation precision based on the first sample set, provided that the probability threshold has the true precision.

3. The data processing method according to claim 2, wherein, The confidence level corresponding to the category is determined based on a first sum and a second sum, wherein the first sum is the sum of the values ​​of the target probability that satisfy the precision condition, and the second sum is the sum of all values ​​of the target probability.

4. The processing method according to claim 2, wherein, The probability density function of the target probability follows a conjugate prior distribution.

5. The data processing method according to claim 1, wherein, The step of determining a probability threshold corresponding to each category, such that the number of positive examples for a category is maximized while the confidence level corresponding to that category is not lower than the confidence threshold, includes: For each category, the precision condition is determined based on the confidence level corresponding to the category and the confidence threshold. A probability threshold is determined that maximizes the observation accuracy and the number of positive examples, and is used as the probability threshold corresponding to the category.

6. The data processing method according to claim 1, wherein, The accuracy condition is that the actual accuracy is greater than a specified value.

7. The data processing method according to claim 1 further includes: In response to an update to the first sample set, the probability threshold corresponding to each category is redefined.

8. The data processing method according to claim 7 further includes: For each category, obtain multiple determined probability thresholds corresponding to the category; Determine the accuracy of multiple observations determined by multiple probability thresholds; The parameters of the distribution of the probability density function of the true accuracy are determined using the multiple observation accuracies.

9. The data processing method according to claim 1, wherein, The machine learning model includes a softmax layer, the predicted probability is output by the softmax layer, and the probability threshold is the softmax threshold.

10. The data processing method according to any one of claims 1 to 9, further comprising: Based on the first machine learning model and the probability threshold corresponding to each category, the samples to be processed are classified.

11. The data processing method according to claim 10, wherein, The classification of samples to be processed based on the first machine learning model and the probability threshold corresponding to each category includes: Based on the first machine learning model and the probability threshold corresponding to each category, the classification result of the samples in the second sample set to be processed is determined, and the second sample set is the training data of the second machine learning model. Based on the classification results, the samples in the second sample set are labeled.

12. The data processing method according to claim 10, wherein, The samples to be processed are online samples awaiting review, and the processing method further includes: In response to the online sample's tagging result indicating approval, the online sample is output online.

13. The data processing method according to claim 12, further comprising: In response to the online sample's labeling result indicating that the review was not passed, the online sample is used to update the sample pool, which is used to train or test the first machine learning model.

14. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data processing method of any one of claims 1 to 13.

15. A computer program product, when run on a computer, causes the computer to implement the data processing method according to any one of claims 1 to 13.

16. A computer program comprising: Instructions, which, when executed by a processor, cause the processor to perform the data processing method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Classification model training method and device and storage medium

    CN114549897A

  • Machine Learning Classification with Confidence Thresholds

    US20190102683A1

  • Automatic thresholding for classification models

    US20230418909A1