Multi-class model calibration method and system based on unified threshold loss function
By using a unified threshold loss function in a multi-category model to optimize the decision boundary, the problem of inconsistent decision thresholds among categories in multi-category classification is solved, the reliability and accuracy of the model are improved, and the accurate probability calibration of the multi-category classification model is achieved.
Patent Information
- Application Number
- CN202411847678.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art ignores the problem of inconsistent decision thresholds between different categories in multi-category classification scenarios, resulting in deviations in the model when comparing the output probability of different categories, affecting the overall reliability of the model. Existing evaluation indicators such as expected calibration error (ECE) also have limitations in multiple categories of scenarios and cannot fully reflect calibration differences between categories.
A multi-category model calibration method based on a unified threshold loss function is adopted to minimize the threshold difference between categories by a unified threshold loss function and optimize the decision boundary. The method includes obtaining training data and category tags, preprocessing, building a deep learning model, and optimizing the decision threshold using a unified threshold loss function, and optimizing the model parameters through a gradient descent algorithm.
By reducing the difference in decision thresholds between categories, the reliability and accuracy of the model in multi-category classification can be improved, ensuring the predicted probability consistency and reliability of model outputs, while maintaining high classification performance.
Smart Images

Figure CN120030447A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model calibration, and in particular to a multi-category model calibration method and system based on a unified threshold loss function. Background Art
[0002] The probability calibration problem of deep neural networks, especially the calibration problem in multi-class classification scenarios, has always been an important research direction in the field of machine learning. Although traditional post-processing calibration methods, such as Platt scaling and temperature scaling, have achieved good results in binary classification problems, they still face significant challenges in multi-class classification tasks. These methods mainly adjust the output probability after model training, but since they cannot affect the feature learning during model training, their improvement effect is relatively limited.
[0003] In recent years, calibration-aware training methods have gradually become a research hotspot. In 2017, Guo et al. proposed a temperature scaling method at the top conference ICML, which provided a simple and effective solution for probability calibration by introducing a single parameter to adjust the output logits of the neural network. In 2019, the Mixup training method proposed by Thu l as idasan et al. improved the calibration performance of the model through data augmentation. In 2021, the trainable maximum mean calibration error (MMCE) loss function published by Kumar et al. directly integrated the calibration objective into the training process, demonstrating superior results over traditional cross entropy loss. These studies have significantly advanced the development of calibration technology for deep learning models.
[0004] For example, Chinese patent publication number: CN116452318A discloses a probability calibration model training method, device, equipment, medium and program product. The method includes: obtaining training data; dividing the training samples into windows according to the predicted probability of the training samples, and calculating the true probability of the training samples in each window; the window is determined according to a preset window division strategy; the determined length of each window and the window spacing are monotonically increasing or monotonically decreasing; the probability calibration model is trained according to the predicted probability and the true probability; the probability calibration model is an order-preserving regression calibration model, and the probability calibration model is used to calibrate the predicted probability output by the binary classification model. Since the determined length of each window and the window spacing are monotonically increasing or monotonically decreasing with the predicted probability of the training sample, the window length and spacing are more consistent with the sample distribution, which can improve the accuracy of the calculated true probability, and further improve the calibration effect of the trained probability calibration model on the predicted probability.
[0005] However, there are still the following problems in the prior art:
[0006] Existing studies mainly focus on the probability calibration of a single category, or adopt a unified calibration strategy to handle all categories, ignoring the problem of inconsistent decision thresholds between different categories in multi-category classification. Since the learning difficulty and data distribution of each category may be different, this inconsistency will cause the model to be biased when comparing the output probabilities of different categories, affecting the overall reliability of the model. In addition, existing evaluation metrics such as expected calibration error (ECE) also have limitations when dealing with multi-category scenarios and cannot fully reflect the calibration differences between categories. Summary of the invention
[0007] To this end, the present invention provides a multi-category model calibration method and system based on a unified threshold loss function to overcome the problem that the learning difficulty and data distribution of each category in the prior art may be different. This inconsistency will cause the model to produce deviations when comparing the output probabilities of different categories, affecting the overall reliability of the model. In addition, existing evaluation indicators such as expected calibration error (ECE) also have limitations when dealing with multi-category scenarios and cannot fully reflect the calibration differences between categories.
[0008] To achieve the above object, on the one hand, the present invention provides a multi-category model calibration method based on a unified threshold loss function, which comprises:
[0009] Obtaining training data and category labels, wherein the training data includes input feature data, category probability distribution data, and true label data;
[0010] Preprocessing the training data to obtain preprocessed data, including a training set and a test set;
[0011] Constructing a deep learning model based on a unified threshold loss function, and inputting the training set and the category label into the deep learning model to train and optimize model parameters;
[0012] A trained deep learning model is obtained, and the test set is input into the deep learning model to obtain an accurately calibrated multi-category prediction result.
[0013] Furthermore, the unified threshold loss function is expressed by formula (1):
[0014]
[0015] In formula (1), L uni represents the uniform threshold loss function, α 1 represents the first loss function, α 2 represents the second loss function, l represents the first threshold boundary hyperparameter, h represents the second threshold boundary hyperparameter, p represents the predicted probability, C represents the first penalty strength hyperparameter, γ represents the second penalty strength hyperparameter, and y represents the true label.
[0016] Furthermore, the unified threshold loss function is used to minimize the threshold difference between categories to optimize the decision boundary.
[0017] Furthermore, the first threshold boundary hyperparameter, the second threshold boundary hyperparameter, the first penalty intensity hyperparameter and the second penalty intensity hyperparameter are pre-set.
[0018] Furthermore, the process of training and optimizing model parameters includes:
[0019] Perform deep feature extraction on the input data to obtain feature representation, which is used to construct a unified threshold loss function;
[0020] Optimizing the decision thresholds of each category based on the unified threshold loss function to obtain a calibrated probability output;
[0021] Based on the calibrated probability output and the true label calculation loss, the model is optimized using a gradient descent algorithm until convergence, thereby obtaining a trained deep learning model.
[0022] Furthermore, the deep feature extraction includes:
[0023] Extract features from input data through a multi-layer neural network;
[0024] Map the original data into a high-dimensional feature space to obtain the feature representation of the data.
[0025] Furthermore, preprocessing the training data includes standardization and normalization.
[0026] Furthermore, the neural network model includes three hidden layers, and a ReLU activation function and a Dropout regularization layer are added after each hidden layer.
[0027] Furthermore, on the other hand, a system for a multi-category model calibration method based on a unified threshold loss function is provided, comprising:
[0028] A data acquisition module, which is used to obtain training data and category labels;
[0029] A preprocessing module, connected to the data acquisition module, for preprocessing a number of the training data to obtain preprocessed data;
[0030] A model training module, which is connected to the preprocessing module and is used to construct a deep learning model based on a unified threshold loss function, and input the training set and the category label into the deep learning model to train and optimize model parameters;
[0031] The result output module and the model training module are used to input the test set into the deep learning model to obtain accurately calibrated multi-category prediction results.
[0032] Furthermore, the model training module includes a feature extraction unit and a unified threshold calibration module.
[0033] The feature extraction unit is used to perform deep feature extraction on input data;
[0034] The unified threshold calibration module is used to receive the feature representation extracted by the feature extraction unit, and stores a unified threshold loss function internally, so as to optimize the decision thresholds of each category based on the unified threshold loss function to obtain a calibrated probability output.
[0035] Compared with the prior art, the present invention standardizes and normalizes the raw data through the data preprocessing link to provide standardized input data for subsequent model training. In the model optimization process, a multi-layer neural network is first used to extract features from the input data, and the raw data is mapped to a high-dimensional feature space, thereby obtaining an effective representation of the data. Based on these feature representations, a unified threshold loss function is calculated, and the decision threshold differences between different categories are evaluated to guide model training. Subsequently, a gradient descent optimization algorithm is used to continuously adjust the model parameters through back propagation, so that the decision boundaries of each category gradually converge. The entire training process is carried out in an end-to-end manner. Data flows from the data acquisition module to the calibration model training module in the system. Through a cyclic iterative optimization process, accurate probability calibration of the multi-category classification model is finally achieved, ensuring that the model can output reliable and consistent prediction probabilities while maintaining high classification performance.
[0036] In particular, the present invention constructs a unified threshold loss function, which guides model training by evaluating the difference in decision thresholds between different categories, and constructs a mechanism to "push the prediction probability to the extremes" - that is, to encourage the model to make more certain predictions, thereby forming a natural decision boundary and ensuring accurate probability calibration of multi-category classification models. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A schematic diagram of the steps of a multi-category model calibration method based on a unified threshold loss function according to an embodiment of the invention;
[0038] Figure 2 A schematic diagram of the probability preference of a unified threshold loss function according to an embodiment of the invention;
[0039] Figure 3 This is a confusion matrix diagram of an embodiment of the invention when a uniform threshold loss function is not used;
[0040] Figure 4This is a confusion matrix diagram after using a unified threshold loss function according to an embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0043] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0044] See also Figure 1 As shown, it is a schematic diagram of the steps of a multi-category model calibration method based on a unified threshold loss function according to an embodiment of the present invention. The multi-category model calibration method based on a unified threshold loss function according to the present invention includes:
[0045] S1, obtaining training data and category labels, wherein the training data includes input feature data, category probability distribution data and true label data;
[0046] S2, preprocessing the plurality of training data to obtain preprocessed data, including a training set and a test set;
[0047] S3, constructing a deep learning model based on a unified threshold loss function, and inputting the training set and the category label into the deep learning model to train and optimize model parameters;
[0048] S4, obtaining a trained deep learning model, inputting the test set into the deep learning model, and obtaining an accurately calibrated multi-category prediction result.
[0049] Specifically, the unified threshold loss function is expressed by formula (1):
[0050]
[0051] In formula (1), L uni represents the uniform threshold loss function, α 1 represents the first loss function, α2 Let \( \alpha \) denote the second loss function, \( l \) denote the first threshold boundary hyperparameter, \( h \) denote the second threshold boundary hyperparameter, \( p \) denote the predicted probability, \( C \) denote the first penalty intensity hyperparameter, \( \gamma \) denote the second penalty intensity hyperparameter, and \( y \) denote the true label. It can be understood that the true label is generally a one - hot vector, that is, a vector with the same length as the number of labels.
[0052] Specifically, please refer to Figure 2 As shown, it is a schematic diagram of the probability preference of the unified threshold loss function of the invention embodiment. The first threshold boundary hyperparameter \( l \) and the second threshold boundary hyperparameter \( p \) are introduced to define the threshold boundaries of high probability and low probability respectively. The two broken lines respectively represent the first loss function \( \alpha \) for the high - probability and low - probability regions 1 and the second loss function \( \alpha \) 2 . For the high - probability region (\( p>h \)), when the predicted probability \( p \) is lower than the threshold \( h \), the model will be subject to a larger penalty; as the value of \( p \) gradually exceeds \( h \), the penalty coefficient shows a decreasing trend. Similarly, for the low - probability region (\( p < l \)), when the predicted probability \( p \) exceeds the threshold \( l \), a larger penalty will be generated; when the value of \( p \) is lower than \( l \), the penalty coefficient gradually decreases.
[0053] It can be understood that the superposition effect of the two loss functions actually constructs a mechanism to "push the predicted probability \( p \) to the two poles" - that is, it encourages the model to make more definite predictions, either tending to high probability or tending to low probability, thus forming a natural decision - boundary interval between \( h \) and \( l \).
[0054] Specifically, the first penalty intensity hyperparameter \( C \) and the second penalty intensity hyperparameter \( \gamma \) control the intensity of the penalty mechanism: the larger the value of \( C \), the stricter the penalty imposed by the model on the samples whose predicted probabilities fall into the middle interval, thus more strongly driving the prediction results to polarize; similarly, the \( \gamma \) parameter adjusts the gradient change rate of the penalty function and affects the sensitivity of the model to different probability intervals. The reasonable setting of these hyperparameters is of great significance for balancing the discriminative ability and stability of the model.
[0055] Specifically, the unified threshold loss function is used to minimize the threshold difference between classes to optimize the decision boundary.
[0056] Specifically, the first threshold boundary hyperparameter, the second threshold boundary hyperparameter, the first penalty intensity hyperparameter, and the second penalty intensity hyperparameter are preset.
[0057] Specifically, in implementation, \( \gamma = 1.0 \), \( C = 1.5 \), \( h = 0.8 \), \( l = 0.2 \).
[0058] Specifically, the process of training and optimizing the model parameters includes
[0059] Perform deep feature extraction on the input data to obtain feature representation, which is used to construct a unified threshold loss function;
[0060] Optimizing the decision thresholds of each category based on the unified threshold loss function to obtain a calibrated probability output;
[0061] Based on the calibrated probability output and the true label calculation loss, the model is optimized using a gradient descent algorithm until convergence, thereby obtaining a trained deep learning model.
[0062] Specifically, the deep feature extraction includes:
[0063] Extract features from input data through a multi-layer neural network;
[0064] Map the original data into a high-dimensional feature space to obtain the feature representation of the data.
[0065] Specifically, preprocessing the training data includes standardization and normalization.
[0066] Specifically, the neural network model includes three hidden layers, and a ReLU activation function and a Dropout regularization layer are added after each hidden layer.
[0067] Specifically, in the implementation, a disease classification task based on human gene information was selected as the experimental object. This task aims to determine whether the subject is a healthy individual or a patient with a specific type of cancer by analyzing human gene expression profile data. In terms of model construction, a deep learning architecture was designed. The architecture uses a multi-layer fully connected neural network as the basic structure, which contains three hidden layers, each containing 512, 256 and 128 neurons respectively, and adds a ReLU activation function and a Dropout regularization layer after each hidden layer to enhance the nonlinear expression ability and generalization performance of the model. The input layer dimension of the model matches the number of gene features, and the output layer uses the Softmax function for multi-category probability prediction.
[0068] Specifically, in terms of the experimental design implemented, a strict data partitioning strategy was adopted, and the entire data set was randomly divided into training set, validation set and test set in a ratio of 5:1:4. Among them, the training set is used for model parameter learning, the validation set is used for hyperparameter optimization selection, and the test set is strictly reserved for the final performance evaluation. During the model training process, a small batch stochastic gradient descent algorithm was used for optimization, the batch size was set to 64, the initial learning rate was set to 0.001, and the Adam optimizer was used to adaptively adjust the learning rate. By monitoring the model performance on the validation set, we adopted an early stopping strategy to prevent overfitting and selected the model parameter configuration when the validation set performance was optimal. Finally, we evaluated the model performance on a completely independent test set.
[0069] See also Figure 3 as well as Figure 4 As shown, the confusion matrix diagram of the embodiment of the invention when the unified threshold loss function is not used and the confusion matrix diagram after the unified threshold loss function is used are respectively shown. When the unified threshold loss function is not used, the effect of the confusion matrix in the disease classification task is average. The classification probability on the diagonal is only around 0.4-0.7, and the result is poor.
[0070] However, after adopting the unified threshold loss function, the classification probability on each diagonal line increased significantly. In particular, in the identification of healthy people, the accuracy rate increased from 0.68 to 1.00, achieving perfect classification; the classification accuracy rate of liver cancer patients also increased to 1.00, and the classification accuracy rate of esophageal cancer patients remained at a high level of 0.87. The accurate recognition rate of gastric cancer patients increased significantly from 0.43 to 0.71, and the classification accuracy rate of colorectal cancer patients also increased significantly from 0.36 to 0.75. At the same time, it can be observed that after using the unified threshold loss function, the probability of misclassification on the non-diagonal line generally decreased, and most of the misclassification probabilities dropped to 0 or close to 0, indicating that the classification boundary of the model has become clearer and the degree of confusion between categories has been greatly reduced. It fully proves the significant effect of the unified threshold loss function in improving the classification performance of the model and probability calibration.
[0071] Specifically, a system for a multi-category model calibration method based on a unified threshold loss function is also provided, which includes:
[0072] A data acquisition module, which is used to obtain training data and category labels;
[0073] A preprocessing module, connected to the data acquisition module, for preprocessing a number of the training data to obtain preprocessed data;
[0074] A model training module, which is connected to the preprocessing module and is used to construct a deep learning model based on a unified threshold loss function, and input the training set and the category label into the deep learning model to train and optimize model parameters;
[0075] The result output module and the model training module are used to input the test set into the deep learning model to obtain accurately calibrated multi-category prediction results.
[0076] Specifically, the model training module includes a feature extraction unit and a unified threshold calibration module.
[0077] The feature extraction unit is used to perform deep feature extraction on input data;
[0078] The unified threshold calibration module is used to receive the feature representation extracted by the feature extraction unit, and stores a unified threshold loss function internally, so as to optimize the decision thresholds of each category based on the unified threshold loss function to obtain a calibrated probability output.
[0079] Specifically, there is no limitation on the specific structures of the data acquisition module, preprocessing module, model training module and result output module, and they themselves or each unit therein can be composed of logical components, and the logical components include field programmable processors, computers or microprocessors in computers.
[0080] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A multi-class model calibration method based on a unified threshold loss function, characterized in that: include: Obtaining training data and category labels, wherein the training data includes input feature data, category probability distribution data, and true label data; Preprocessing the training data to obtain preprocessed data, including a training set and a test set; Constructing a deep learning model based on a unified threshold loss function, and inputting the training set and the category label into the deep learning model to train and optimize model parameters; A trained deep learning model is obtained, and the test set is input into the deep learning model to obtain an accurately calibrated multi-category prediction result.
2. The multi-category model calibration method based on a unified threshold loss function according to claim 1, characterized in that: The unified threshold loss function is expressed by formula (1): In formula (1), L uni represents the unified threshold loss function, α1 represents the first loss function, α2 represents the second loss function, l represents the first threshold boundary hyperparameter, h represents the second threshold boundary hyperparameter, p represents the predicted probability, C represents the first penalty strength hyperparameter, γ represents the second penalty strength hyperparameter, and y represents the true label.
3. The multi-category model calibration method based on a unified threshold loss function according to claim 2, characterized in that: The unified threshold loss function is used to minimize the threshold difference between categories to optimize the decision boundary.
4. The multi-category model calibration method based on a unified threshold loss function according to claim 3, characterized in that: The first threshold boundary hyperparameter, the second threshold boundary hyperparameter, the first penalty intensity hyperparameter and the second penalty intensity hyperparameter are preset.
5. The multi-category model calibration method based on a unified threshold loss function according to claim 1, characterized in that: The process of training and optimizing model parameters includes: Perform deep feature extraction on the input data to obtain feature representation, which is used to construct a unified threshold loss function; Optimizing the decision thresholds of each category based on the unified threshold loss function to obtain a calibrated probability output; Based on the calibrated probability output and the true label calculation loss, the model is optimized using a gradient descent algorithm until convergence, thereby obtaining a trained deep learning model.
6. The multi-category model calibration method based on a unified threshold loss function according to claim 1, characterized in that: The deep feature extraction comprises: Extract features from input data through a multi-layer neural network; Map the original data into a high-dimensional feature space to obtain the feature representation of the data.
7. The multi-class model calibration method based on a unified threshold loss function according to claim 1, characterized in that: The preprocessing of the training data includes standardization and normalization.
8. The multi-category model calibration method based on a unified threshold loss function according to claim 1, characterized in that: The neural network model includes three hidden layers, and a ReLU activation function and a Dropout regularization layer are added after each hidden layer.
9. A system using the multi-class model calibration method based on a unified threshold loss function according to any one of claims 1 to 8, characterized in that: include: A data acquisition module, which is used to obtain training data and category labels; A preprocessing module, connected to the data acquisition module, for preprocessing a number of the training data to obtain preprocessed data; A model training module, which is connected to the preprocessing module and is used to construct a deep learning model based on a unified threshold loss function, and input the training set and the category label into the deep learning model to train and optimize model parameters; The result output module and the model training module are used to input the test set into the deep learning model to obtain accurately calibrated multi-category prediction results.
10. The system according to claim 9, characterized in that The model training module includes a feature extraction unit and a unified threshold calibration module. The feature extraction unit is used to perform deep feature extraction on input data; The unified threshold calibration module is used to receive the feature representation extracted by the feature extraction unit, and stores a unified threshold loss function internally, so as to optimize the decision thresholds of each category based on the unified threshold loss function to obtain a calibrated probability output.
Citation Information
Patent Citations
Probability calibration model training method and device, equipment, medium and program product
CN116452318A