Model training method and device, equipment, storage medium and program product
By calculating the complexity of portrait data to generate a training dataset and training the model, the problem of insufficient model generalization ability and robustness in existing technologies is solved, and stronger facial data recognition capabilities are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, model training relies on simply divided training and test sets, resulting in poor model generalization ability and robustness.
By acquiring multiple first-image data sets, calculating the complexity of each category, generating a training dataset, and training the initial model based on this dataset, a target model is generated. The target model is used to identify different categories of image data.
It improves the model's generalization ability and robustness, enabling it to better cover facial data in different scenarios and enhance the model's recognition capabilities.
Smart Images

Figure CN121808538A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and in particular to a model training method and device, equipment, a storage medium and a program product. BACKGROUND
[0002] Face data is identified to determine the actual access personnel, which is a key technology in the field of data security. In the related art, a model for identifying face data is usually trained, and the face data of a user is identified by the model to determine whether it is a preset personnel, and the personnel is allowed to access data if it is a preset personnel. However, in the related art, the training of the model depends on the pre-prepared training set data and test set data, and the training set data and test set data are usually obtained by simply dividing the original data. The training set data and test set data cannot generalize the face data in different situations, resulting in poor generalization ability and robustness of the trained model.
[0003] It can be seen that the related art has the problem of poor generalization ability and robustness of the model. SUMMARY
[0004] The embodiments of the present application provide a model training method, device, equipment, storage medium and program product to solve the problem of poor generalization ability and robustness of the model in the related art.
[0005] To solve the above problems, the present application is implemented as follows:
[0006] In a first aspect, the embodiments of the present application provide a model training method, comprising:
[0007] Obtaining a plurality of first portrait data, the plurality of first portrait data being portrait data of a plurality of categories;
[0008] Calculating a complexity corresponding to each category based on the plurality of first portrait data, the complexity being used to represent the distribution complexity, feature complexity and / or balance degree between the plurality of categories of the portrait data of the corresponding category;
[0009] Generating a training data set based on the plurality of first portrait data and the complexity of each category;
[0010] Training an initial model based on the training data set to obtain a target model, the initial model being a model for identifying portrait data, and the target model being used to identify portrait data of different categories.
[0011] In a second aspect, the embodiments of the present application further provide a model training device, comprising:
[0012] The first obtaining module is configured to obtain a plurality of first portrait data, wherein the plurality of first portrait data is portrait data of a plurality of categories;
[0013] The first calculating module is configured to calculate a complexity corresponding to each category based on the plurality of first portrait data, wherein the complexity is used to represent a distribution complexity, a feature complexity and / or a balance degree between the plurality of categories of the portrait data of the corresponding category;
[0014] The first generating module is configured to generate a training data set based on the plurality of first portrait data and the complexity of each category.
[0015] The training module is configured to train an initial model based on the training data set to obtain a target model, wherein the initial model is a model used for recognizing portrait data, and the target model is used for recognizing portrait data of different categories.
[0016] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a transceiver and a processor,
[0017] The processor is configured to obtain a plurality of first portrait data, wherein the plurality of first portrait data is portrait data of a plurality of categories.
[0018] The processor is further configured to calculate a complexity corresponding to each category based on the plurality of first portrait data, wherein the complexity is used to represent a distribution complexity, a feature complexity and / or a balance degree between the plurality of categories of the portrait data of the corresponding category.
[0019] The processor is further configured to generate a training data set based on the plurality of first portrait data and the complexity of each category.
[0020] The processor is further configured to train an initial model based on the training data set to obtain a target model, wherein the initial model is a model used for recognizing portrait data, and the target model is used for recognizing portrait data of different categories.
[0021] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory and a program stored in the memory and executable on the processor, wherein the program is executed by the processor to implement the steps of the model training method of the first aspect.
[0022] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the model training method of the first aspect.
[0023] In a sixth aspect, the present application also provides a computer program product comprising computer instructions which, when executed by a processor, implement the steps of the model training method of the first aspect.
[0024] In the embodiment of the present application, a plurality of first image data is obtained, the plurality of first image data being image data of a plurality of categories; a complexity corresponding to each category is calculated based on the plurality of first image data, the complexity being used to represent a distribution complexity, a feature complexity and / or a balance degree between the plurality of categories of the image data of the corresponding category; a training data set is generated based on the plurality of first image data and the complexity of each category; and an initial model is trained based on the training data set to obtain a target model, the initial model being a model used for identifying image data, and the target model being used for identifying image data of different categories. In this way, the training data set is generated based on the plurality of first image data and the complexity of each category, so that part of the first image data included in the training data set is first image data under different categories; and the initial model is trained based on the training data set to obtain the target model, so that the model generalization ability and robustness of the target model obtained by training can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0025] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0026] Figure 1 is a flowchart of a model training method provided by an embodiment of the present application;
[0027] Figure 2 is a schematic diagram of the overall flow of model training provided by an embodiment of the present application;
[0028] Figure 3 is a structural diagram of a model training device provided by an embodiment of the present application;
[0029] Figure 4 is a structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0030] The technical solutions of the embodiments of the present application will be described clearly and completely in the following description with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0031] Please refer to Figure 1 , Figure 1 is a flowchart of a model training method provided by an embodiment of the present application, as shown in Figure 1 , comprising the following steps:
[0032] Step 101, obtaining a plurality of first portrait data, the plurality of first portrait data being portrait data of a plurality of categories.
[0033] The plurality of first portrait data is portrait data of different users under different scene conditions, and the plurality of first portrait data can represent different facial features, so as to facilitate subsequent training of the model.
[0034] The plurality of categories is a category for dividing the plurality of first portrait data, and the plurality of categories can be configured according to the requirements of model training. For example, if the model to be trained needs to identify whether the portrait data is real data, the category can be set as a real face category and a non-real face category.
[0035] Step 102, calculating a complexity corresponding to each category based on the plurality of first portrait data, the complexity being used to represent a distribution complexity, a feature complexity and / or a balance degree between the plurality of categories of the portrait data of the corresponding category.
[0036] The complexity is used to represent the case of different categories of first portrait data, and can specifically represent the distribution complexity, the feature complexity and / or the balance degree between the plurality of categories of the portrait data of the corresponding category. The complexity can be used to determine whether the specific case of different category data is needed, and then the complexity can be used to determine the number of different categories of first portrait data in the generated training data set.
[0037] In some embodiments, a first complexity sub-parameter for representing the distribution complexity of the portrait data of the corresponding category, a second complexity sub-parameter for representing the feature complexity, and a third complexity sub-parameter for representing the balance degree between the plurality of categories can be calculated respectively, and then the first complexity sub-parameter, the second complexity sub-parameter and the third complexity sub-parameter are weighted to calculate the complexity corresponding to each category.
[0038] Specifically, the complexity corresponding to each category can be calculated by the following formula:
[0039] ;
[0040] In the formula, is the category entropy, which is used to measure the complexity of the category distribution; Variance(D) is the sum of feature variances, which is used to reflect the complexity of the features; For category balance, the greater the value, the more uniform the category distribution; w1, w2 and w3 are weight parameters for adjusting the three types of indexes.
[0041] Step 103, generating a training data set based on the plurality of first image data and the complexity of each category.
[0042] The training data set includes part of the first image data in the plurality of first image data. It can be understood that the training data set is generated based on the plurality of first image data and the complexity of each category, so that the part of the first image data included in the training data set is the first image data under different categories, the training data set can better cover different user facial conditions, and thus improve the prediction effect of the model trained by the training data set.
[0043] In some embodiments, the proportion of different categories can be calculated by complexity, so that the training data set can be generated by the proportion of different categories.
[0044] Specifically, the proportion can be calculated by the following formula:
[0045] ;
[0046] In the formula, OptimalSplit represents the optimal split function, Complexity D represents the comprehensive complexity of the first image data D, Size D represents the sample size of the data D, Task type represents the classification label of the task.
[0047] Further, the verification data set and the test data set can also be generated by the proportion, which are respectively used for verifying and testing the model.
[0048] Exemplarily, the specific implementation is as follows:
[0049] When , the proportion of the training set is increased, and the proportion of the test set is reduced; wherein represents a high complexity threshold preset by the system;
[0050] When , the proportion of the verification set is reduced to avoid over-segmentation; a predefined size threshold;
[0051] According to Task type (classification, regression, generation, etc.), the division strategy is adjusted, which is specifically represented as:
[0052] ;
[0053] In the formula, D trainFor a training data set, D val For a verification data set, D test For a test data set.
[0054] Step 104, training an initial model based on the training data set to obtain a target model, the initial model is a model for identifying portrait data, and the target model is used to identify portrait data of different categories.
[0055] The initial model is a model constructed before training, which is used to identify portrait data when constructed; the target model is a model obtained by training the initial model based on the training data set, and the target model can effectively identify portrait data of different categories. Since the part of the first portrait data included in the training data set is the first portrait data of different categories, the model generalization ability and robustness of the target model obtained by training can be improved.
[0056] In an embodiment of the present application, a plurality of first portrait data is obtained, the plurality of first portrait data is portrait data of a plurality of categories; the complexity corresponding to each category is calculated based on the plurality of first portrait data, the complexity is used to represent the distribution complexity, feature complexity and / or balance degree between the plurality of categories of the corresponding category of portrait data; a training data set is generated based on the plurality of first portrait data and the complexity of each category; an initial model is trained based on the training data set to obtain a target model, the initial model is a model for identifying portrait data, and the target model is used to identify portrait data of different categories. In this way, the training data set is generated based on the plurality of first portrait data and the complexity of each category, so that part of the first portrait data included in the training data set is the first portrait data of different categories; and the initial model is trained based on the training data set to obtain the target model, so that the model generalization ability and robustness of the target model obtained by training can be improved.
[0057] Further, the training process of the model in the present application is as shown in Figure 2 , including: data preparation, obtaining a plurality of first portrait data, and constructing a training data set; training a model, training a target model through the training data set; model evaluation, to determine the model generalization ability and robustness of the target model; cross-validation, to verify the generalization ability of the model; use of independent data set evaluation, specifically model robustness test, to determine the robustness of the model; model improvement and retraining, to optimize the model; final evaluation, to determine the final prediction effect of the model.
[0058] In one embodiment, the plurality of first portrait data is obtained, including:
[0059] Obtaining a plurality of initial portrait data;
[0060] Extract the distribution feature vectors of the multiple initial portrait data;
[0061] The enhancement strategy is determined based on the aforementioned distribution feature vector;
[0062] The multiple initial portrait data are adjusted based on the enhancement strategy to obtain the multiple first portrait data.
[0063] The aforementioned initial profile data are obtained by directly collecting users' facial data. In some implementations, initial profile data for different business scenarios, environments, and / or users can be collected using different devices. These different devices can be mobile phones, tablets, and / or computer cameras, etc.; different business scenarios can be financial business scenarios with relatively high security requirements and access control scenarios with relatively low security requirements, etc.; different environments can be lighting environments, such as indoor, outdoor, low light, and strong light environments; and different users can be different groups of people with different characteristics, such as age, gender, and ethnicity. It is understood that by acquiring multiple initial profile data sets, including those for different business scenarios, environments, and / or users, diverse data sources are ensured, thereby enabling the model to cover as many application scenarios as possible and improving the model's generalization ability.
[0064] It should be noted that the initial profile data collected in this invention are collected with the user's permission and authorization. If the user does not allow or authorize the collection, the user's initial profile data will not be collected.
[0065] Furthermore, distribution feature vectors are extracted from multiple initial portrait data. These distribution feature vectors can determine the distribution of multiple initial portrait data, such as the mean vector, variance, skewness, kurtosis, entropy, and other parameters of the multiple initial portrait data.
[0066] Among them, for the initial dataset D obtained by collection raw (Including multiple initial portrait data), calculate its distribution feature vector D. profile Specifically, it can be expressed by the following formula:
[0067] ;
[0068] In the formula This is a feature mean vector used to reflect the trend of multiple initial profile data centers, where n is the number of initial profile data points, and x... i Let i be the feature value of the i-th initial portrait data; This is the variance matrix, used to measure the degree of dispersion of the data; Skewness is used to describe the asymmetry of the data distribution, where X represents multiple random variables of initial profile data. is a mean value of the plurality of initial image data; is a kurtosis, used to describe the sharpness of the data distribution; is a data entropy, used to quantify the information content of the data, p i is a probability of the i-th initial image data occurring, log is a logarithm base. In this way, by the distribution feature vector D profile The features of the plurality of initial image data, such as data skewness, abnormal value distribution, etc., can be automatically identified, so as to facilitate subsequent enhancement of the initial image data.
[0069] The above determination of the enhancement strategy based on the distribution feature vector can achieve different initial image data being enhanced by different enhancement strategies. It should be noted that in the prior art, a fixed enhancement strategy (such as simple image flipping and rotation) is used to enhance data, but the use of a fixed enhancement strategy will destroy the distribution of the data, resulting in weak model generalization ability. In the present application, a suitable enhancement strategy is dynamically selected according to the distribution feature vector, such as preferentially using brightness adjustment and noise enhancement for the case of insufficient data under low light conditions; for insufficient different shooting angles, rotation, affine transformation, etc. are selected, and attack samples (such as screen replay attack and mask attack) for living body recognition are generated by a generative adversarial network, so as to significantly improve the diversity and coverage of the data while maintaining the consistency of the distribution.
[0070] In some embodiments, the enhancement strategy is determined based on the distribution feature vector, and can be represented by the following formula:
[0071] ;
[0072] In the formula, Aug i represents the i-th enhancement strategy (such as rotation, scaling, noise addition, color transformation, etc.), w j is the importance weight of the j-th distribution feature in the distribution feature vector, is the probability of selecting strategy i under a given distribution feature.
[0073] Further, after the enhancement strategy is determined, the plurality of initial image data is adjusted based on the enhancement strategy to obtain a plurality of first image data. The parameters corresponding to different enhancement strategies can be pre-configured, and the corresponding initial image data is adjusted by the parameters.
[0074] In some embodiments, the plurality of initial image data is adjusted to obtain an enhanced data set D aug (i.e. the plurality of first image data), which can be represented by the following formula:
[0075] ;
[0076] In the formula These are parameters used to enhance the initial portrait data. The Euclidean distance between the current initial image data and the target distribution. This is the distribution feature vector corresponding to the current dataset. The desired target distribution feature vector.
[0077] Among them, Div target The preset target diversity index is calculated using the following formula:
[0078] ;
[0079] In the formula, N is the total number of initial image data, (x i ,x j The first and second initial portrait data in the dataset are represented by the first and second initial portrait data.
[0080] In this embodiment of the invention, multiple initial portrait data are acquired; distribution feature vectors of the multiple initial portrait data are extracted; an enhancement strategy is determined based on the distribution feature vectors; and the multiple initial portrait data are adjusted according to the enhancement strategy to obtain the multiple first portrait data. Thus, by adjusting multiple initial portrait data based on the enhancement strategy, different enhancements can be applied to different initial portrait data, significantly improving data diversity and coverage while maintaining distribution consistency.
[0081] In one embodiment, the method further includes:
[0082] Obtain the first similarity, first distance, and first score corresponding to each first portrait data. The first similarity is used to characterize the similarity between the corresponding first portrait data and the portrait data in the training dataset. The first distance is the distance between the corresponding first portrait data and the feature vector corresponding to the training dataset. The first score is used to characterize the difficulty of the corresponding first portrait data.
[0083] The difficulty level of the corresponding first profile data is calculated based on the first similarity, the first distance, and the first score.
[0084] A test dataset is generated based on a preset ratio and the difficulty level of each first portrait data, wherein the preset ratio is the proportion of first portrait data at different difficulty levels;
[0085] The step of training the initial model based on the training dataset to obtain the target model includes:
[0086] The initial model is trained based on the training dataset and the test dataset to obtain the target model.
[0087] It should be noted that the data attacks faced by the model differ in different application scenarios, therefore the model needs to identify different attacks. For example, in the financial field, where security requirements are high, the model will face data from more challenging video playback attacks or 3D mask attacks, and the model needs to effectively identify the more challenging attack data. In contrast, for application scenarios with lower security requirements, such as access control, the model will only face less challenging photo attacks, and in this case, the model only needs to identify the less challenging attack data. Therefore, in this invention, different validation datasets are designed for different application scenarios to test the model's ability to identify attack data.
[0088] Specifically, by calculating the difficulty level of different first portrait data and generating test datasets based on preset ratios and the difficulty level of the first portrait data, test datasets of different difficulties can be generated.
[0089] For example, the test set can be divided into three levels of difficulty: easy, medium, and hard, as shown below:
[0090] D indep =D easy D medium D hard ;
[0091] D indep In the formula, D represents the test set. easy For a simple level test set, D medium For a medium-level test set, D hard This is the test set for the difficult level.
[0092] Furthermore, the difficulty level of the corresponding first profile data is calculated based on the first similarity, the first distance, and the first score, which can be specifically expressed by the following formula:
[0093] ;
[0094] In the formula, x represents the first portrait data. The training set D represents train The vector value, Indicates the first similarity. Indicates the first distance, Human rating (x) represents the first score of the first portrait data x.
[0095] Furthermore, the preset ratio can be expressed by the following formula:
[0096] ;
[0097] In the formula This ratio is used to ensure that the test set has a reasonable difficulty gradient distribution.
[0098] In some implementations, to enhance the challenge of the test set, additional adversarial profiling data can be generated and added. It can be expressed by the following formula:
[0099] ;
[0100] In the formula Indicates the intensity of the disturbance. This represents the gradient of the model's predicted loss f(x) with respect to the true label y, where x is the gradient of the loss with respect to the input x. adversarial This indicates new profile data.
[0101] The test set D is constructed in this way. indep It not only has distributional differences, but also different levels of difficulty, which can comprehensively evaluate the generalization ability of the model.
[0102] In this embodiment of the invention, a first similarity, a first distance, and a first score are obtained for each first portrait data. The first similarity characterizes the similarity between the corresponding first portrait data and the portrait data in the training dataset. The first distance is the distance between the corresponding first portrait data and the feature vector corresponding to the training dataset. The first score characterizes the difficulty level of the corresponding first portrait data. The difficulty level of the corresponding first portrait data is calculated based on the first similarity, the first distance, and the first score. A test dataset is generated based on a preset ratio and the difficulty level of each first portrait data, where the preset ratio is the proportion of first portrait data at different difficulty levels. Thus, by constructing a test dataset using a preset ratio, model testing under different application scenarios is achieved, enabling the trained target model to effectively cope with attack data.
[0103] In one embodiment, training the initial model based on the training dataset to obtain the target model includes:
[0104] The initial model is trained multiple times based on the training dataset to obtain the target model;
[0105] The multiple rounds include a first round and a second round. The first round is the round preceding the second round. The weight of the i-th sample in the second round is calculated by the weight of the i-th sample in the first round and the performance gradient of the i-th sample. The i-th sample is the first portrait data included in the training dataset. The regularization parameter in the second round is calculated by the regularization parameter of the first round and the validation loss.
[0106] In this embodiment of the invention, by dynamically updating the weights of the samples, overfitting or underfitting of the model during training can be avoided. For example, in difficult-to-identify 3D mask attack samples, the weights are continuously increased, enabling the model to more accurately identify the most likely attack methods to cause risk in different scenarios (such as financial scenarios).
[0107] The calculation process for adjusting the weights in different rounds is expressed by the following formula:
[0108] ;
[0109] In the formula, The weight for round t+1, The weight for round t, This is the learning rate decay coefficient. Let represent the performance gradient of the i-th sample (i.e., the first portrait data), reflecting the contribution of that sample to the improvement of model performance. In this way, when the loss of a sample remains high, the system will automatically increase its weight, prompting the model to learn more difficult samples; when a sample has been fully learned, its weight will be reduced to avoid overfitting.
[0110] It should be noted that existing technologies use fixed loss functions and static regularization parameters for model training, which cannot adapt to changes in data distribution and model state during training. In contrast, this invention uses dynamic regularization parameter adjustment to monitor model complexity and stability, and to strengthen or relax regularization as needed, thereby avoiding overfitting or underfitting in financial risk control scenarios.
[0111] Among them, the dynamic R-values of the model at different rounds. The following formula is used for calculation:
[0112] ;
[0113] In the formula, For model complexity, the L0 norm is used to measure parameter sparsity; Model stability is measured by batch-to-batch prediction variance; regularization parameter Adjust dynamically based on training progress and validation set performance.
[0114] Specifically, the calculation process for the regularization parameter is represented by the following formula:
[0115] ;
[0116] In the formula, This represents the weight value of the k-th parameter in the (t+1)-th round. This represents the weight value of the k-th parameter in round t. Hyperparameters used to control the magnitude of weight changes are typically , This represents the loss value in the t-th round. This represents the loss value at round t-1. When the validation loss increases, the regularization strength is increased; when the validation loss decreases, the regularization is appropriately relaxed, allowing the model to learn more complex patterns.
[0117] It should be noted that during the training process, the following metrics are continuously monitored: training loss convergence speed, validation set performance change trend, model parameter update magnitude, and gradient explosion / vanishing detection. Based on this monitoring information, hyperparameters such as learning rate and batch size are dynamically adjusted to ensure the stability and effectiveness of the training process.
[0118] In one embodiment, the method further includes:
[0119] A verification dataset is generated based on the multiple first profile data and the complexity of each category;
[0120] The target model is cross-validated based on the validation dataset to obtain the accuracy index, stability index, robustness index and / or consistency index;
[0121] A comprehensive index is calculated based on the accuracy index, the stability index, the robustness index, and / or the consistency index.
[0122] An evaluation report is generated based on the comprehensive index, and the evaluation report is used to evaluate the generalization ability of the target model.
[0123] The aforementioned comprehensive index is the Generalization Comprehensive Index (GCI), which is used to determine the generalization performance of a model.
[0124] Furthermore, the composite index is calculated using the following formula:
[0125] ;
[0126] In the formula For accuracy index, As a stability index, The robustness index is the consistency index, and w1, w2, w3, and w4 are coefficients.
[0127] The certainty index is calculated using the following formula:
[0128] ;
[0129] In the formula P train P represents the accuracy of the training set.val The accuracy index represents the accuracy of the validation dataset. It measures the consistency of the model's performance on the training and validation sets. The closer the value is to 1, the better the generalization ability.
[0130] The stability index is calculated using the following formula:
[0131] ;
[0132] In the formula P cv The stability index is used to measure the stability of the model performance by the coefficient of variation, representing the performance score of each fold in the cross-validation.
[0133] The robustness index is calculated using the following formula:
[0134] ;
[0135] In the formula P adversarial The prediction accuracy of adversarial examples, P clean Sensitivity represents the prediction accuracy for unperturbed samples. noise Threshold represents noise sensitivity. max This represents the preset upper limit of sensitivity. The robustness index is used to comprehensively consider the model's resistance to adversarial examples and its sensitivity to noise.
[0136] The consistency index is calculated using the following formula:
[0137] ;
[0138] In the formula, K represents the total number of samples, Prediction k Truth represents the model's output for the k-th sample. k Range represents the baseline truth value of the k-th sample. output The consistency index represents the theoretical range of the model's output values and is used to measure the consistency and reliability of the model's prediction results.
[0139] In some implementations, w1, w2, w3, and w4 can be adjusted using the following formula:
[0140] ;
[0141] w in the formula i For any one of the coefficients w1, w2, w3, and w4, Importance i It is dynamically calculated based on factors such as task type, data scale, and business needs.
[0142] After calculating the comprehensive index, an evaluation report is generated based on the comprehensive index, and the evaluation report is used to evaluate the generalization ability of the target model. For example, when GCI > 0.4, an evaluation report with excellent model generalization ability is generated; when 0.6 < GCI ≤ 0.4, the model generalization ability is good, and an evaluation report suggesting further optimization is generated; when GCI ≤ 0.6, there is a risk of overfitting in the model, and an evaluation report requiring redesign is generated.
[0143] In an embodiment of the present invention, a validation dataset is generated based on the multiple first portrait data and the complexity of each category; the target model is cross-validated based on the validation dataset to obtain an accuracy index, a stability index, a robustness index, and / or a consistency index; a comprehensive index is calculated based on the accuracy index, the stability index, the robustness index, and / or the consistency index; an evaluation report is generated based on the comprehensive index, and the evaluation report is used to evaluate the generalization ability of the target model. In this way, the model is verified through the validation dataset, and an evaluation report is generated to facilitate determining the generalization ability of the target model through the evaluation report.
[0144] In one embodiment, the cross-validating the target model based on the validation dataset to obtain an accuracy index, a stability index, a robustness index, and / or a consistency index includes:
[0145] Obtain the sample quantity corresponding to the validation dataset, and at least one of a sample category distribution parameter, a sample resource constraint, and a model stability parameter;
[0146] Generate an initial value of the cross-validation fold number based on the sample quantity;
[0147] Adjust the initial value based on at least one of the sample category distribution parameter, the sample resource constraint, and the model stability parameter to obtain an expected value of the cross-validation fold number;
[0148] Cross-validate the target model based on the validation dataset and the expected value of the cross-validation fold number to obtain the accuracy index, the stability index, the robustness index, and / or the consistency index.
[0149] The above cross-validation fold number is the K value. It should be noted that the K-value cross-validation in the prior art uses a fixed K value (usually 5 or 10), which cannot adapt to the characteristics of different datasets. In the present invention, the K value is automatically selected according to the data scale and complexity of the application scenario to ensure that each data fold contains enough minority class (attack class) samples and achieve the reliability of cross-validation. Meanwhile, monitor the validation stability and adjust dynamically to ensure that the model has a highly stable generalization performance when applied in the actual business scenario.
[0150] Specifically, an initial value for the cross-validation fold K is generated based on the number of samples in the dataset to avoid insufficient or excessive training data, ensuring that the model has sufficient training data and reliable generalization.
[0151] Furthermore, the sample class distribution parameter is used to characterize the complexity and imbalance of the data class distribution. The cross-validation fold number is dynamically adjusted by the sample class distribution parameter to ensure that the class distribution of each validation set and training set is uniform, especially that minority class samples are reasonably distributed between each fold, thereby reducing the risk of overfitting.
[0152] In addition, sample resource constraints are used to characterize the resource constraints used for model vectors. For example, within a set budget constraint for computational resources (such as memory and time), the cross-validation fold number K is automatically adjusted to ensure that cross-validation effectively controls computational costs while achieving accurate estimation of generalization ability.
[0153] The model stability parameter is used to characterize the stability of model performance (such as performance fluctuation) during cross-validation. If the stability is lower than the preset threshold, the cross-validation fold number K is dynamically increased to improve the reliability of performance evaluation.
[0154] In this way, the initial value is adjusted based on at least one of the sample category distribution parameters, the sample resource constraints, and the model stability parameters to obtain the expected value of the cross-validation fold, so that cross-validation using the expected value can more comprehensively reflect the generalization ability and stability of the model in practical applications.
[0155] In this embodiment of the invention, the number of samples corresponding to the validation dataset, and at least one of the following: sample class distribution parameters, sample resource constraints, and model stability parameters, are obtained. An initial value for the cross-validation fold is generated based on the number of samples. The initial value is adjusted based on at least one of the sample class distribution parameters, sample resource constraints, and model stability parameters to obtain an expected value for the cross-validation fold. Cross-validation is performed on the target model based on the validation dataset and the expected value of the cross-validation fold to obtain the accuracy index, the stability index, the robustness index, and / or the consistency index. Thus, by generating an initial value for the cross-validation fold using the number of samples, and then adjusting the initial value to obtain an expected value, cross-validation can be performed using the expected value, improving the accuracy of model generalization evaluation.
[0156] In some implementations, in addition to adjusting the K value and performing cross-validation to evaluate the model's generalization ability, the model's generalization ability can also be evaluated through multiple layers. It should be noted that existing technologies only calculate overall performance indicators and cannot deeply analyze the source and layers of generalization ability; however, this invention employs a four-layer generalization evaluation (data layer, feature layer, model layer, and task layer) for models in fields with high security requirements (such as the financial sector). For example, it analyzes the feature invariance of portrait data (robustness under changes in facial pose), the model architecture's defense capabilities against different attacks, and the model's transfer effect on other financial risk control tasks to determine the model's generalization ability at different levels.
[0157] For example, the evaluation of the model's data layer can be achieved using the following formula:
[0158] ;
[0159] In the formula G data This indicates the generalization of the data layer and is used to measure the model's adaptability to different data domains; This represents a transformation function that maps domain differences to adaptation scores; The hyperparameter used to control the rate of score decline is the decay coefficient. >0), MMD stands for Maximum Mean Discrepancy, which measures the degree of difference between two distributions.
[0160] MMD is calculated using the following formula:
[0161] ;
[0162] In the formula, D1={x i} represents the set of samples from distribution D1 (such as the training dataset), and D2={y j} represents the set of samples from distribution D2 (such as the test dataset).
[0163] It should be noted that the evaluation methods for other layers (including the feature layer, model layer, and task layer) can be the same as those for the data layer, and can be calculated in a similar way.
[0164] In this way, the generalization ability of the model is deeply evaluated through four layers: the data layer (domain variability), the feature layer (feature invariance and discriminability), the model layer (structural complexity and transfer adaptability), and the task layer (task transferability). Based on the evaluation results of each layer, a comprehensive generalization index can be calculated, thereby identifying generalization bottlenecks and providing specific optimization suggestions. Simultaneously, the generalization performance gap is quantified based on the ratio of model complexity to data size, and corresponding adjustment strategies are proposed, generating a multi-dimensional generalization ability analysis report.
[0165] For example, based on the multi-level evaluation results, a detailed report can be generated that includes at least one of the following: generalization ability scores and rankings at each level; identification of generalization ability bottlenecks and suggestions for improvement; comparative analysis with similar models; and description of applicable scenarios and limitations.
[0166] In one embodiment, the method further includes:
[0167] Acquire the plurality of second portrait data, wherein the plurality of second portrait data is a portion of the portrait data in the plurality of first portrait data;
[0168] Based on the multiple second portrait data, generate Fast Gradient Sign Method (FGSM) attack parameters, Projected Gradient Descent (PGD) attack parameters, and / or C&W attack parameters for each second portrait data.
[0169] Based on the FGSM attack parameters, PGD attack parameters and / or C&W attack parameters of each second profile data, attack sample data corresponding to the second profile data is generated.
[0170] The target model is tested based on the attack sample data corresponding to the multiple second profile data, and the test results are used to evaluate the robustness of the target model.
[0171] In this embodiment of the invention, a plurality of second portrait data are acquired, wherein the plurality of second portrait data are partial portrait data from a plurality of first portrait data; based on the plurality of second portrait data, Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), and / or Contrast & Warring States (C&W) attack parameters are generated for each second portrait data; based on the FGSM, PGD, and / or C&W attack parameters of each second portrait data, attack sample data corresponding to the respective second portrait data is generated; the target model is tested based on the attack sample data corresponding to the plurality of second portrait data to obtain test results, which are used to evaluate the robustness of the target model. Thus, by generating attack sample data using FGSM, PGD, and / or C&W attack parameters, and testing the target model using the attack sample data to obtain test results, the robustness of the target model can be evaluated through the test results.
[0172] The attack sample data is generated using FGSM attack parameters, PGD attack parameters, and / or C&W attack parameters, which can be specifically represented by the following formula:
[0173] ;
[0174] In the formula For attack sample data, For the second portrait data, For FGSM attack parameters, These are the parameters for PGD attack. For C&W attack parameters, To preset hyperparameters, , and is a coefficient.
[0175] The FGSM attack parameters are calculated using the following formula:
[0176] ;
[0177] in the formula For the second portrait data, Let y represent the trainable weights (fixed) in the model. The true label. The FGSM attack parameters utilize the sign direction of the gradient of the loss function to generate adversarial perturbations, which are fast but relatively weak in strength.
[0178] PGD attack parameters are calculated using the following formula:
[0179] ;
[0180] PGD is an iterative and enhanced version of FGSM, generating stronger adversarial examples through multiple iterations and projection operations.
[0181] C&W attack parameters are calculated using the following formula:
[0182] ;
[0183] C&W attack optimizes adversarial perturbations under L2 norm constraints, generating more perceptually natural adversarial examples.
[0184] In this way, composite adversarial attack samples are generated using methods such as FGSM, PGD, and C&W, and the attack intensity is dynamically adjusted to evaluate the robustness of the model against multimodal attacks, thereby ensuring that the model has a certain defensive capability under all attack types.
[0185] Furthermore, a robustness analysis report can be generated based on the test results. This report includes at least one of the following: statistical results of the success rate of various attacks, vulnerability analysis results of the model (which types of samples are most vulnerable to attack), defense strategy recommendations (data augmentation, adversarial training, model regularization, etc.), and robustness comparison results with the benchmark model.
[0186] In one embodiment, the method further includes:
[0187] Obtain the generalization loss, robustness loss, resource consumption parameters, and / or performance expectation deviation parameters of the target model;
[0188] Constraints are constructed based on the generalization loss, the robustness loss, the resource consumption parameter, and / or the performance expectation deviation parameter;
[0189] The optimized generalization loss, robustness loss, resource consumption parameter, and / or performance expectation deviation parameter are calculated based on the constraints.
[0190] The target model is adjusted based on the optimized generalization loss, robustness loss, resource consumption parameter, and / or performance expectation deviation parameter.
[0191] It should be noted that existing technologies typically rely on human experience for model improvement, which is inefficient and prone to overlooking key issues. In contrast, this invention automatically diagnoses model problems through evaluation results, such as insufficient defense against low light or video attacks, triggering data augmentation and architecture adjustments to optimize the model.
[0192] This allows for the construction of a problem diagnosis matrix based on the evaluation results, which in turn determines the adjustment strategy. For example, the problem diagnosis matrix can be constructed using the following formula:
[0193] ;
[0194] In the formula, w i This represents the importance weight of the i-th evaluation metric (such as generalization ability, robustness, accuracy, etc.) in the current scenario, and can be adaptively set through expert experience or data-driven approaches. i The deviation parameter represents the i-th indicator, used to characterize the degree of deviation from the target value and the degree of inadequacy in performance under the evaluation criteria. N represents the total number of evaluation indicators considered, such as generalization loss, robustness loss, resource consumption parameters, and / or performance expectation deviation.
[0195] The constraints constructed above based on the generalization loss, the robustness loss, the resource consumption parameter, and / or the expected performance deviation parameter can be specifically expressed by the following formula:
[0196] ;
[0197] ;
[0198] In the formula, L genL represents the generalization loss, used to measure performance on unseen samples; rob Represents robustness loss, used to assess performance gaps under adversarial attacks; C comp Indicates resource consumption parameters (such as FLOPs, latency, memory); D gap P(Ω) represents the performance expectation deviation parameter and its distance from the business KPI; P(Ω) represents the Pareto solution set that satisfies the constraints.
[0199] The NSGA-III algorithm can be used to find the Pareto optimal solution set.
[0200] ;
[0201] The dominance relationship in the formula is defined as follows: .
[0202] In this way, model problems are automatically diagnosed based on evaluation results, model improvement schemes are determined using multi-objective optimization methods, and continuous optimization is achieved through intelligent convergence criteria and meta-learning mechanisms.
[0203] The final output is a model containing the complete improved trajectory (Mimproved), specifically represented by the following formula:
[0204] ;
[0205] in the formula To optimize the model in the final stage, To improve historical records, For the application's strategy sequence, This is a performance improvement report.
[0206] In this embodiment of the invention, the generalization loss, robustness loss, resource consumption parameters, and / or expected performance deviation parameters of the target model are obtained; constraints are constructed based on the generalization loss, robustness loss, resource consumption parameters, and / or expected performance deviation parameters; optimized generalization loss, robustness loss, resource consumption parameters, and / or expected performance deviation parameters are calculated based on the constraints; and the target model is adjusted based on the optimized generalization loss, robustness loss, resource consumption parameters, and / or expected performance deviation parameters. Thus, by constructing constraints based on the generalization loss, robustness loss, resource consumption parameters, and / or expected performance deviation parameters, and calculating optimized generalization loss, robustness loss, resource consumption parameters, and / or expected performance deviation parameters based on the constraints, the target model can be adjusted according to the optimized generalization loss, robustness loss, resource consumption parameters, and / or expected performance deviation parameters, thereby achieving model optimization.
[0207] Furthermore, existing technologies only calculate a single metric on the test set, which cannot comprehensively verify the model's practical application capabilities. In this invention, for application scenarios with high security requirements (such as financial risk control), the model's deployment stability, cross-domain consistency, temporal stability, and resource consumption are measured to output a clear comprehensive financial scenario evaluation report for the final assessment. This report includes a panoramic view of model performance, risk assessment, deployment recommendations, monitoring metrics, and version update suggestions, thus obtaining a multi-dimensional, multi-scenario comprehensive evaluation report to ensure the model's reliability in real-world environments.
[0208] In some implementations, a multi-dimensional boundary test matrix can be constructed to achieve stress testing under different conditions. Specifically, the construction of a multi-dimensional boundary test matrix can be expressed by the following formula:
[0209] ;
[0210] in the formula The boundary test matrix represents the set of test inputs constructed under all test dimensions and boundary conditions; x i,j The input sample generated after applying the j-th boundary condition on the i-th dimension; To test the perturbation generation function, which characterizes the boundary input constructed in the i-th dimension (such as illumination changes, face occlusion, resolution anomalies, etc.); c j The j-th boundary test condition (e.g., brightness = 0, delay > 1s, etc.); m is the total number of test dimensions, such as the number of modalities and scene types; n is the number of boundary conditions under each dimension.
[0211] Furthermore, a multi-dimensional evaluation matrix is constructed to evaluate the model from different dimensions. Specifically, the multi-dimensional evaluation matrix can be represented by the following formula:
[0212] ;
[0213] In the formula, GCI is the generalization ability index, used to measure the model's performance on non-training samples; R atk为 Robustness score against attacks (higher is more robust); U real The practicality score is used to characterize user feedback or business metrics based on actual deployment scenarios; V time The model's performance score is given under the time stability test. The variation in the model's generalization ability across different data domains; The degree of variation in robustness under different attack intensities; User experience costs, such as the subjective burden caused by response time and false alarm rate; System performance variance reflects the stability of various model metrics on different test sets.
[0214] In this way, the stability of the model in extreme environments is evaluated through multi-dimensional boundary condition testing, and a multi-dimensional evaluation matrix is used to integrate multi-dimensional indicators such as generalization ability, robustness, and practicality to form a clear comprehensive evaluation report.
[0215] For example, the output evaluation report includes at least one of the following: a panoramic view of model performance, risk assessment and mitigation recommendations, recommended deployment configurations, definitions of monitoring metrics, update trigger conditions, and version compatibility analysis.
[0216] Please see Figure 3 , Figure 3 This is a structural diagram of a model training device provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the model training device 300 includes:
[0217] The first acquisition module 301 is used to acquire multiple first portrait data, wherein the multiple first portrait data are portrait data of multiple categories;
[0218] The first calculation module 302 is used to calculate the complexity corresponding to each category based on the plurality of first portrait data. The complexity is used to characterize the distribution complexity, feature complexity and / or the balance between the plurality of categories of portrait data of the corresponding category.
[0219] The first generation module 303 is used to generate a training dataset based on the plurality of first portrait data and the complexity of each category;
[0220] The training module 304 is used to train the initial model based on the training dataset to obtain the target model. The initial model is a model for recognizing portrait data, and the target model is used to recognize different categories of portrait data.
[0221] In one embodiment, the first acquisition module 301 includes:
[0222] The first acquisition unit is used to acquire multiple initial portrait data;
[0223] The extraction unit is used to extract the distribution feature vectors of the multiple initial portrait data;
[0224] A determining unit is configured to determine an enhancement strategy based on the distribution feature vector;
[0225] The first adjustment unit is used to adjust the plurality of initial portrait data according to the enhancement strategy to obtain the plurality of first portrait data.
[0226] In one embodiment, the model training device 300 further includes:
[0227] The second acquisition module is used to acquire the first similarity, first distance and first score corresponding to each first portrait data. The first similarity is used to characterize the similarity between the corresponding first portrait data and the portrait data in the training dataset. The first distance is the distance between the corresponding first portrait data and the feature vector corresponding to the training dataset. The first score is used to characterize the difficulty of the corresponding first portrait data.
[0228] The second calculation module is used to calculate the difficulty level of the corresponding first portrait data based on the first similarity, the first distance and the first score;
[0229] The second generation module is used to generate a test dataset based on a preset ratio and the difficulty level of each first portrait data, wherein the preset ratio is the proportion of first portrait data with different difficulty levels.
[0230] The training module 304 includes:
[0231] The first training unit is used to train the initial model based on the training dataset and the test dataset to obtain the target model.
[0232] In one embodiment, the training module 304 includes:
[0233] The second training unit is used to train the initial model for multiple rounds based on the training dataset to obtain the target model;
[0234] The multiple rounds include a first round and a second round. The first round is the round preceding the second round. The weight of the i-th sample in the second round is calculated by the weight of the i-th sample in the first round and the performance gradient of the i-th sample. The i-th sample is the first portrait data included in the training dataset. The regularization parameter in the second round is calculated by the regularization parameter of the first round and the validation loss.
[0235] In one embodiment, the model training device 300 further includes:
[0236] The third generation module is used to generate a verification dataset based on the multiple first portrait data and the complexity of each category;
[0237] The validation module is used to perform cross-validation on the target model based on the validation dataset to obtain accuracy index, stability index, robustness index and / or consistency index;
[0238] The third calculation module is used to calculate a comprehensive index based on the accuracy index, the stability index, the robustness index, and / or the consistency index.
[0239] The fourth generation module is used to generate an evaluation report based on the comprehensive index, and the evaluation report is used to evaluate the generalization ability of the target model.
[0240] In one embodiment, the verification module includes:
[0241] The second acquisition unit is used to acquire the number of samples corresponding to the verification dataset, as well as at least one of the sample category distribution parameters, sample resource constraints, and model stability parameters.
[0242] A generation unit is used to generate an initial value for the cross-validation fold number based on the number of samples.
[0243] The second adjustment unit is used to adjust the initial value based on at least one of the sample category distribution parameters, the sample resource constraints, and the model stability parameters to obtain the expected value of the cross-validation fold.
[0244] The validation unit is used to perform cross-validation on the target model based on the validation dataset and the expected value of the cross-validation folds to obtain the accuracy index, the stability index, the robustness index and / or the consistency index.
[0245] In one embodiment, the model training device 300 further includes:
[0246] The third acquisition module is used to acquire the plurality of second portrait data, wherein the plurality of second portrait data is a portion of the portrait data in the plurality of first portrait data;
[0247] The fifth generation module is used to generate Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), and / or C&W attack parameters for each second portrait data based on the plurality of second portrait data.
[0248] The sixth generation module is used to generate attack sample data corresponding to the second profile data based on the FGSM attack parameters, the PGD attack parameters and / or the C&W attack parameters of each second profile data.
[0249] The testing module is used to test the target model based on the attack sample data corresponding to the multiple second profile data, and obtain test results. The test results are used to evaluate the robustness of the target model.
[0250] In one embodiment, the model training device 300 further includes:
[0251] The fourth acquisition module is used to acquire the generalization loss, robustness loss, resource consumption parameters, and / or performance expectation deviation parameters of the target model;
[0252] The construction module is used to construct constraints based on the generalization loss, the robustness loss, the resource consumption parameter, and / or the performance expectation deviation parameter;
[0253] The fourth calculation module is used to calculate the optimized generalization loss, robustness loss, resource consumption parameter, and / or performance expectation deviation parameter based on the constraints.
[0254] The adjustment module is used to adjust the target model based on the optimized generalization loss, the robustness loss, the resource consumption parameter, and / or the performance expectation deviation parameter.
[0255] The model training apparatus provided in this embodiment of the invention can implement each process of each embodiment of the above-described model training method, with one-to-one correspondence of technical features and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0256] It should be noted that the model training device in the embodiments of the present invention can be a device, or it can be a component, integrated circuit, or chip in an electronic device.
[0257] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the above-described functionality. Figure 1 The various processes of the model training method embodiment shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0258] For details, see Figure 4 As shown, this embodiment of the invention also provides an electronic device, including a bus 401, a transceiver 402, an antenna 403, a bus interface 404, a processor 405, and a memory 406.
[0259] The processor 405 is used to acquire multiple first portrait data, wherein the multiple first portrait data are portrait data of multiple categories;
[0260] The processor 405 is further configured to calculate the complexity corresponding to each category based on the plurality of first portrait data, wherein the complexity is used to characterize the distribution complexity, feature complexity and / or balance between the plurality of categories of portrait data of the corresponding category;
[0261] The processor 405 is also configured to generate a training dataset based on the plurality of first portrait data and the complexity of each category;
[0262] The processor 405 is further configured to train an initial model based on the training dataset to obtain a target model, wherein the initial model is a model for recognizing portrait data, and the target model is used to recognize different categories of portrait data.
[0263] In one embodiment, acquiring multiple first profile data includes:
[0264] Obtain multiple initial profile data;
[0265] Extract the distribution feature vectors of the multiple initial portrait data;
[0266] The enhancement strategy is determined based on the aforementioned distribution feature vector;
[0267] The multiple initial portrait data are adjusted based on the enhancement strategy to obtain the multiple first portrait data.
[0268] In one embodiment, the transceiver 402 is used to acquire a first similarity, a first distance, and a first score corresponding to each first portrait data. The first similarity is used to characterize the similarity between the corresponding first portrait data and the portrait data in the training dataset. The first distance is the distance between the corresponding first portrait data and the feature vector corresponding to the training dataset. The first score is used to characterize the difficulty of the corresponding first portrait data.
[0269] The processor 405 is also used to calculate the difficulty level of the corresponding first portrait data based on the first similarity, the first distance and the first score;
[0270] The processor 405 is also used to generate a test dataset based on a preset ratio and the difficulty level of each first portrait data, wherein the preset ratio is the ratio of first portrait data with different difficulty levels;
[0271] The step of training the initial model based on the training dataset to obtain the target model includes:
[0272] The initial model is trained based on the training dataset and the test dataset to obtain the target model.
[0273] In one embodiment, training the initial model based on the training dataset to obtain the target model includes:
[0274] The initial model is trained multiple times based on the training dataset to obtain the target model;
[0275] The multiple rounds include a first round and a second round. The first round is the round preceding the second round. The weight of the i-th sample in the second round is calculated by the weight of the i-th sample in the first round and the performance gradient of the i-th sample. The i-th sample is the first portrait data included in the training dataset. The regularization parameter in the second round is calculated by the regularization parameter of the first round and the validation loss.
[0276] In one embodiment, the processor 405 is further configured to generate a verification dataset based on the plurality of first profile data and the complexity of each category;
[0277] The processor 405 is further configured to perform cross-validation on the target model based on the validation dataset to obtain an accuracy index, a stability index, a robustness index, and / or a consistency index.
[0278] The processor 405 is further configured to calculate a comprehensive index based on the accuracy index, the stability index, the robustness index, and / or the consistency index;
[0279] The processor 405 is also configured to generate an evaluation report based on the comprehensive index, the evaluation report being used to evaluate the generalization ability of the target model.
[0280] In one embodiment, the step of cross-validating the target model based on the validation dataset to obtain accuracy, stability, robustness, and / or consistency indices includes:
[0281] Obtain the number of samples corresponding to the validation dataset, as well as at least one of the following: sample category distribution parameters, sample resource constraints, and model stability parameters;
[0282] An initial value for the cross-validation fold number is generated based on the stated sample size;
[0283] The initial value is adjusted based on at least one of the sample category distribution parameters, the sample resource constraints, and the model stability parameters to obtain the expected value of the cross-validation fold.
[0284] The target model is cross-validated based on the validation dataset and the expected value of the cross-validation folds to obtain the accuracy index, the stability index, the robustness index, and / or the consistency index.
[0285] In one embodiment, the transceiver 402 is further configured to acquire the plurality of second portrait data, wherein the plurality of second portrait data are a portion of the portrait data in the plurality of first portrait data;
[0286] The processor 405 is also configured to generate, based on the plurality of second portrait data, fast gradient symbolic method (FGSM) attack parameters, projected gradient descent method (PGD) attack parameters and / or C&W attack parameters corresponding to each second portrait data.
[0287] The processor 405 is further configured to generate attack sample data corresponding to the second profile data based on the FGSM attack parameters, the PGD attack parameters and / or the C&W attack parameters of each second profile data;
[0288] The processor 405 is further configured to test the target model based on the attack sample data corresponding to the plurality of second profile data, and obtain test results, the test results being used to evaluate the robustness of the target model.
[0289] In one embodiment, the processor 405 is further configured to acquire the generalization loss, robustness loss, resource consumption parameters, and / or performance expectation deviation parameters of the target model;
[0290] The processor 405 is further configured to construct constraints based on the generalization loss, the robustness loss, the resource consumption parameter, and / or the performance expectation deviation parameter;
[0291] The processor 405 is further configured to calculate the optimized generalization loss, robustness loss, resource consumption parameter and / or performance expectation deviation parameter based on the constraints.
[0292] The processor 405 is further configured to adjust the target model based on the optimized generalization loss, the robustness loss, the resource consumption parameter, and / or the performance expectation deviation parameter.
[0293] exist Figure 4 In this context, a bus architecture (represented by bus 401) is used. Bus 401 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 405 and memory represented by memory 406. Bus 401 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 404 provides an interface between bus 401 and transceiver 402. Transceiver 402 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 405 is transmitted over a wireless medium via antenna 403, which further receives data and transmits data to processor 405.
[0294] Processor 405 is responsible for managing bus 401 and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 406 can be used to store data used by processor 405 during operation.
[0295] Optionally, the processor 405 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a graphics processing unit (GPU).
[0296] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the above-described functions. Figure 1 The various processes of the corresponding model training method embodiments, which achieve the same technical effect, will not be described again here to avoid repetition. The computer-readable storage medium mentioned includes, for example, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0297] The present invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the above-described... Figure 1 The various processes of the corresponding model training method implementation examples can achieve the same technical effect, and will not be described again here to avoid repetition.
[0298] In the embodiments of this invention, the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing four possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.
[0299] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0300] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or second terminal device, etc.) to execute the methods of the various embodiments of this application.
[0301] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A model training method, characterized in that, include: Acquire multiple first portrait data, wherein the multiple first portrait data are portrait data of multiple categories; The complexity of each category is calculated based on the multiple first portrait data. The complexity is used to characterize the distribution complexity, feature complexity and / or balance between the multiple categories of portrait data of the corresponding category. A training dataset is generated based on the multiple first portrait data and the complexity of each category; The initial model is trained based on the training dataset to obtain the target model. The initial model is used to identify portrait data, and the target model is used to identify different categories of portrait data.
2. The method as described in claim 1, characterized in that, The acquisition of multiple first profile data includes: Obtain multiple initial profile data; Extract the distribution feature vectors of the multiple initial portrait data; The enhancement strategy is determined based on the aforementioned distribution feature vector; The multiple initial portrait data are adjusted based on the enhancement strategy to obtain the multiple first portrait data.
3. The method as described in claim 1, characterized in that, The method further includes: Obtain the first similarity, first distance, and first score corresponding to each first portrait data. The first similarity is used to characterize the similarity between the corresponding first portrait data and the portrait data in the training dataset. The first distance is the distance between the corresponding first portrait data and the feature vector corresponding to the training dataset. The first score is used to characterize the difficulty of the corresponding first portrait data. The difficulty level of the corresponding first profile data is calculated based on the first similarity, the first distance, and the first score. A test dataset is generated based on a preset ratio and the difficulty level of each first portrait data, wherein the preset ratio is the proportion of first portrait data at different difficulty levels; The step of training the initial model based on the training dataset to obtain the target model includes: The initial model is trained based on the training dataset and the test dataset to obtain the target model.
4. The method as described in claim 1, characterized in that, The step of training the initial model based on the training dataset to obtain the target model includes: The initial model is trained multiple times based on the training dataset to obtain the target model; The multiple rounds include a first round and a second round. The first round is the round preceding the second round. The weight of the i-th sample in the second round is calculated by the weight of the i-th sample in the first round and the performance gradient of the i-th sample. The i-th sample is the first portrait data included in the training dataset. The regularization parameter in the second round is calculated by the regularization parameter of the first round and the validation loss.
5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: A verification dataset is generated based on the multiple first profile data and the complexity of each category; The target model is cross-validated based on the validation dataset to obtain the accuracy index, stability index, robustness index and / or consistency index; A comprehensive index is calculated based on the accuracy index, the stability index, the robustness index, and / or the consistency index. An evaluation report is generated based on the comprehensive index, and the evaluation report is used to evaluate the generalization ability of the target model.
6. The method as described in claim 5, characterized in that, The process of cross-validating the target model based on the validation dataset to obtain accuracy, stability, robustness, and / or consistency indices includes: Obtain the number of samples corresponding to the validation dataset, as well as at least one of the following: sample category distribution parameters, sample resource constraints, and model stability parameters; An initial value for the cross-validation fold number is generated based on the stated sample size; The initial value is adjusted based on at least one of the sample category distribution parameters, the sample resource constraints, and the model stability parameters to obtain the expected value of the cross-validation fold. The target model is cross-validated based on the validation dataset and the expected value of the cross-validation folds to obtain the accuracy index, the stability index, the robustness index, and / or the consistency index.
7. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Acquire the plurality of second portrait data, wherein the plurality of second portrait data is a portion of the portrait data in the plurality of first portrait data; Based on the multiple second portrait data, generate the Fast Gradient Sign Method (FGSM) attack parameters, Projected Gradient Descent (PGD) attack parameters, and / or C&W attack parameters corresponding to each second portrait data. Based on the FGSM attack parameters, PGD attack parameters and / or C&W attack parameters of each second profile data, attack sample data corresponding to the second profile data is generated. The target model is tested based on the attack sample data corresponding to the multiple second profile data, and the test results are used to evaluate the robustness of the target model.
8. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Obtain the generalization loss, robustness loss, resource consumption parameters, and / or performance expectation deviation parameters of the target model; Constraints are constructed based on the generalization loss, the robustness loss, the resource consumption parameter, and / or the performance expectation deviation parameter; The optimized generalization loss, robustness loss, resource consumption parameter, and / or performance expectation deviation parameter are calculated based on the constraints. The target model is adjusted based on the optimized generalization loss, robustness loss, resource consumption parameter, and / or performance expectation deviation parameter.
9. A model training device, characterized in that, include: The first acquisition module is used to acquire multiple first portrait data, wherein the multiple first portrait data are portrait data of multiple categories; The first calculation module is used to calculate the complexity corresponding to each category based on the plurality of first portrait data. The complexity is used to characterize the distribution complexity, feature complexity and / or the balance between the plurality of categories of portrait data of the corresponding category. The first generation module is used to generate a training dataset based on the plurality of first portrait data and the complexity of each category; The training module is used to train the initial model based on the training dataset to obtain the target model. The initial model is a model for recognizing portrait data, and the target model is used to recognize different categories of portrait data.
10. An electronic device, characterized in that, Including transceivers and processors, The processor is configured to acquire multiple first portrait data, wherein the multiple first portrait data are portrait data of multiple categories; The processor is further configured to calculate the complexity corresponding to each category based on the plurality of first portrait data, wherein the complexity is used to characterize the distribution complexity, feature complexity and / or balance between the plurality of categories of portrait data of the corresponding category; The processor is also configured to generate a training dataset based on the plurality of first portrait data and the complexity of each category; The processor is further configured to train an initial model based on the training dataset to obtain a target model, wherein the initial model is a model for recognizing portrait data, and the target model is used to recognize different categories of portrait data.
11. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the model training method as described in any one of claims 1 to 4.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the model training method as described in any one of claims 1 to 4.
13. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the model training method as described in any one of claims 1 to 4.