Gradient adaptive quantitative perception training method, device and equipment for field generalization

By introducing smooth gradients and evaluating the gradient chaos of task gradients in the quantized convolutional neural network model, the task gradient is selectively used for learning, which solves the problem of insufficient generalization ability of the quantization model in unknown fields, and achieves better generalization performance in resource-constrained environments.

CN119942304AActive Publication Date: 2025-05-06TSINGHUA UNIVERSITY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510074953.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art is difficult to improve the generalization capability of quantitative models in unknown fields under limited computing resources, especially in environments where edge device deployment is carried out.

Method used

The parameters of the quantized convolutional neural network model are updated by introducing a smooth gradient, and the gradient chaos of the task gradient is evaluated during the training process, and the task gradient is selectively used for learning to optimize the generalization ability and training stability of the model.

Benefits of technology

The generalization ability of quantitative models in unknown fields is improved, making the models deployed on resource-constrained devices more robust and can better face data distribution changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942304A_ABST
    Figure CN119942304A_ABST
Patent Text Reader

Abstract

The invention provides a field-generalization-oriented gradient adaptive quantitative perception training method, device and equipment, relates to the technical field of computers, and aims to improve the generalization ability of a quantitative model in an unknown field. The method comprises the following steps: acquiring a training data set and a quantitative convolutional neural network model; inputting each sample image in the training data set into the quantitative convolutional neural network model for processing to obtain a task gradient and a flatness of each scale factor in different training domains; evaluating the gradient confusion degree of the task gradients of each scale factor in different training domains, wherein the gradient confusion degree represents the consistency degree of the task gradient direction; and according to the gradient confusion degree, determining a learning gradient of each scale factor in different training domains, and updating parameters of the quantized convolutional neural network model based on the learning gradient to obtain a trained quantized convolutional neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a gradient adaptive quantization perception training method, device and equipment for domain generalization. Background Art

[0002] Deep learning models have performed well in various computer vision tasks, such as classification, object detection, and semantic segmentation. In practical applications, these models often suffer from significant performance degradation due to domain shift, which manifests as poor generalization ability to previously unseen data distributions. Domain generalization aims to address this problem by enabling models trained on known source domains to effectively generalize to unknown target domains.

[0003] However, the existing domain generalization scheme is only applicable to full-precision models. In many real-world scenarios, especially those involving edge device deployment and susceptible to domain shift, models often need to run under very limited computing resources. Therefore, the models generated based on the current domain generalization scheme are not very practical in actual deployment. The low-precision calculation method (i.e., quantization method) commonly used in edge devices only considers the same distribution assumption, and the application of the current domain generalization scheme to the quantized model is not effective. Since the distribution shift of data has not been seen in actual applications, the generalization performance of the quantized model cannot be guaranteed.

[0004] Therefore, how to improve the generalization ability of quantitative models in unknown fields is a technical problem that needs to be solved urgently. Summary of the invention

[0005] In view of the above problems, the embodiments of the present application provide a domain-generalized gradient adaptive quantization perception training method, apparatus and device to overcome the above problems or at least partially solve the above problems.

[0006] In a first aspect of an embodiment of the present application, a gradient adaptive quantization perception training method for domain generalization is disclosed, the method comprising: Obtaining a training data set and a quantized convolutional neural network model, wherein the training data set includes a plurality of sample images of equal amounts in different training domains; Input each sample image in the training data set into the quantized convolutional neural network model for processing, and obtain the task gradient and smooth gradient of each scale factor in different training domains, wherein the scale factor characterizes the characteristics of the weight and activation value distribution of the quantized convolutional neural network model, the task gradient is used to optimize the task performance of the quantized convolutional neural network model, and the smooth gradient is used to optimize the generalization of the quantized convolutional neural network model; Evaluate the gradient chaos of the task gradients of each scale factor in different training domains, where the gradient chaos represents the consistency of the task gradient direction; According to the gradient confusion, the learning gradients of each scale factor in different training domains are determined, and the parameters of the quantized convolutional neural network model are updated based on the learning gradients to obtain a trained quantized convolutional neural network model, wherein the learning gradients at least include the smooth gradients.

[0007] Optionally, evaluate the gradient confusion of the task gradients of each scale factor in different training domains, including: Every target training steps, evaluate the gradient confusion of the task gradients of each scale factor in different training domains; Determining the learning gradients of each scale factor in different training domains according to the gradient confusion degree, and updating the parameters of the quantized convolutional neural network model based on the learning gradients, including: According to the gradient confusion, the learning gradients of each scale factor in different training domains within the next target training steps are determined, and the parameters of the quantized convolutional neural network model are updated based on the learning gradients within the next target training steps.

[0008] Optionally, every target training steps, evaluate the gradient confusion of the task gradients of each scale factor in different training domains, including: Determine a task gradient sequence according to the task gradient of the scale factor in the training domain within the target training steps; wherein the task gradient of the scale factor in the training domain within each training step is a task gradient in the task gradient sequence; The gradient confusion degree is obtained according to the direction change of two adjacent task gradients in the task gradient sequence.

[0009] Optionally, the task gradient sequence includes a first task gradient sequence and a second task gradient sequence, and two task gradients with the same order in the first task gradient sequence and the second task gradient sequence represent two task gradients of adjacent training steps; The gradient confusion degree is obtained according to the direction change of two adjacent task gradients in the task gradient sequence, including: Comparing the directions of two task gradients with the same order in the first task gradient sequence and the second task gradient sequence in turn, and determining the number of two task gradients in adjacent training steps with different directions; The ratio of the number to the target number of training steps is used as the gradient confusion.

[0010] Optionally, determining the learning gradients of each scale factor in different training domains according to the gradient confusion degree includes: For the learning gradient of each scale factor in each training domain, when the gradient chaos of the task gradient of the scale factor in the training domain is less than or equal to the gradient chaos threshold, determine the smooth gradient as the learning gradient of the scale factor in the training domain; When the gradient chaos of the task gradient of the scale factor in the training domain is greater than the gradient chaos threshold, the smooth gradient and the task gradient are determined as the learning gradient of the scale factor in the training domain.

[0011] Optionally, each sample image in the training data set is input into the quantized convolutional neural network model for processing to obtain the task gradient and smoothing gradient of each scale factor in different training domains, including: Constructing a loss function according to the quantization result of the quantized convolutional neural network model and the sample image, wherein the loss function includes a task loss and a smoothing loss; The task gradients of each scale factor in different training domains are determined according to the task loss, and the smoothing gradients of each scale factor in different training domains are determined according to the smoothing loss.

[0012] A second aspect of the embodiments of the present application discloses an image processing method, the method comprising: Get the image to be processed; The image to be processed is input into a trained quantized convolutional neural network model to obtain a target visual task result. The trained quantized convolutional neural network model is trained according to the domain generalization-oriented gradient adaptive quantized perception training method described in the first aspect of the embodiment of the present application.

[0013] In a third aspect of the embodiments of the present application, a gradient adaptive quantization perception training device for domain generalization is disclosed, the device comprising: An acquisition module, used to acquire a training data set and a quantized convolutional neural network model, wherein the training data set includes a plurality of sample images of equal amounts in different training domains; A quantization module, used to input each sample image in the training data set into the quantized convolutional neural network model for processing, and obtain task gradients and smooth gradients of each scale factor in different training domains, wherein the scale factor characterizes the characteristics of the weight and activation value distribution of the quantized convolutional neural network model, the task gradient is used to optimize the task performance of the quantized convolutional neural network model, and the smooth gradient is used to optimize the generalization of the quantized convolutional neural network model; An evaluation module is used to evaluate the gradient chaos of the task gradients of each scale factor in different training domains, where the gradient chaos represents the consistency of the task gradient direction; An updating module is used to determine the learning gradients of each scale factor in different training domains according to the gradient confusion, and to update the parameters of the quantized convolutional neural network model based on the learning gradients to obtain a trained quantized convolutional neural network model, wherein the learning gradients at least include the smooth gradients.

[0014] According to a third aspect of an embodiment of the present application, an electronic device is disclosed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the gradient adaptive quantization perception training method for domain generalization described in the first aspect of the embodiment of the present application are implemented, or the steps of the image processing method described in the second aspect of the embodiment of the present application are implemented.

[0015] In a fourth aspect of an embodiment of the present application, a computer-readable storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the steps of the gradient adaptive quantization perception training method for domain generalization described in the first aspect of the embodiment of the present application are implemented, or the steps of the image processing method described in the second aspect of the embodiment of the present application are implemented.

[0016] In a fifth aspect of an embodiment of the present application, a computer program product is disclosed, including a computer program, which, when executed by a processor, implements the steps of the gradient adaptive quantization perceptual training method for domain generalization described in the first aspect of the embodiment of the present application, or the steps of the image processing method described in the second aspect of the embodiment of the present application.

[0017] The embodiments of the present application include the following advantages: In the embodiment of the present application, a smooth gradient is introduced to update the parameters of the quantized convolutional neural network model, so that the quantization process and the model smoothness can be jointly optimized, thereby improving the generalization ability of the model; and, during the training process, the gradient chaos of the task gradients of each of the scale factors in different training domains is evaluated, and the task gradient is selectively used for learning based on the gradient chaos, thereby enhancing the stability of the model training and ensuring the global convergence of the overall performance. In this way, the smooth gradient and task gradient conflict problem of the quantized model can be finely optimized at the granularity of the domain, thereby improving the generalization ability of the quantized model in unknown fields, and enhancing the generalization ability of the model deployed on resource-constrained devices, making the model more robust in the face of changes in data distribution. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0019] Figure 1 It is a flowchart of the steps of a gradient adaptive quantization perception training method for domain generalization provided in an embodiment of the present application; Figure 2 It is a flowchart of the steps of another gradient adaptive quantization perception training method for domain generalization provided in an embodiment of the present application; Figure 3 It is a schematic diagram of the architecture of a gradient adaptive quantization perception training method for domain generalization provided in an embodiment of the present application; Figure 4 is a flowchart of the steps of an image processing method provided by an embodiment of the present application; Figure 5 It is a structural schematic diagram of a gradient adaptive quantization perception training device for domain generalization provided in an embodiment of the present application; Figure 6 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0021] In real applications, when deploying machine learning models, the test data distribution may be different from the training distribution, which is a common phenomenon called distribution shift. Domain generalization (DG) aims to enhance the generalization of models to unseen domains. Common strategies include domain alignment, meta-learning, data augmentation, disentangled representation learning, and exploiting causal relationships. However, even with the development of these sophisticated techniques, basic empirical risk minimization can still achieve comparable out-of-distribution generalization performance when experimental conditions are carefully controlled.

[0022] At the same time, there is increasing interest in the geometry of loss function landscapes, especially in shared-aware minimization in pursuit of flatter minima during training. Flatter minima may lead to smaller domain-generalization gaps. Inspired by the study of flat minima, flatness-based generalization methods have begun to gain attention and show significant performance in domain generalization. For example, SAGM (a domain generalization scheme) improves generalization by optimizing the angle of weight gradients. Although flatness-based generalization methods have shown significant effects in improving out-of-distribution generalization performance, they are limited to full-precision training. This means that the models generated by current methods are not very practical in actual deployments and do not take into account quantization-specific factors. In other words, in many real-world scenarios, especially those involving edge device deployments and susceptible to domain shift, models often need to run on very limited computing resources.

[0023] Quantization-Aware Training (QAT) involves inserting simulated quantization nodes and retraining the model. QAT achieves a better balance between accuracy and compression ratio by simulating quantization in the forward-backward process, making the weights aware of numerical changes to improve runtime efficiency. A quantization-aware training method uses low-precision weights and activations in the forward pass and uses STE technology in backpropagation to estimate the gradient of the piecewise quantization function. Another quantization-aware training method adjusts the quantization function by introducing a learnable step-size scaling factor. There are also some quantization-aware training methods that offer the possibility of further improving quantization performance by freezing unstable weights.

[0024] However, achieving quantization-aware training for domain generalization remains challenging, specifically in the following two aspects: 1) Conflicting goals: low-precision computation aims to reduce model complexity, but this conflicts with maintaining generalization performance; 2) Training instability: how to ensure the correct convergence of low-precision weights remains a difficult problem, since both simulated quantization and sharpness-aware minimization (SAM) involve specific gradient approximations.

[0025] The embodiments of the present application found that when the domain generalization sharpness perception minimization method is directly applied to quantization perception training, the model generalization performance may unexpectedly decrease. For example, when 4-bit quantization is performed on the PACS dataset, the average out-of-distribution performance of the model decreases by 28.36%. Through in-depth analysis of the quantizer gradient behavior, the significant conflict between the task loss (empirical loss) and the smoothing loss caused by the gradient approximation leads to a decrease in the generalization ability of the training model, and even performs worse than the model that only optimizes a single objective.

[0026] In summary, the existing domain generalization solutions are only applicable to full-precision models, and the low-precision calculation methods (i.e., quantization methods) commonly used in edge devices only consider the same distribution assumption. Applying the current domain generalization solutions to quantized models has poor results. Since the distribution offset of data has not been seen in actual applications, the generalization performance of the quantized model cannot be guaranteed.

[0027] In order to overcome the limitations of the related art, the embodiment of the present application provides a gradient adaptive quantization perception training method for domain generalization, which defines the gradient chaos, which is used to describe the inconsistency of the gradient direction during the training process to quantize the degree of gradient conflict. Specifically, the smooth gradient introduces a quantized convolutional neural network model (quantizer) so that the quantization process and model smoothness can be jointly optimized to improve the generalization ability of the model. In addition, for the two different gradients (task gradient and smooth gradient) from the quantization target and the shared perception minimization target simultaneously received by the quantized convolutional neural network model, by evaluating the gradient chaos of the task gradients of each scale factor in different training domains, the task gradient is selectively used for learning based on the gradient chaos, which enhances the stability of the model training and ensures the global convergence of the overall performance. In this way, the smooth gradient and task gradient conflict problems of the quantized model are finely optimized at the granularity of the domain, which improves the generalization ability of the quantized model in unknown fields, so that the generalization ability of the deployed model is enhanced on resource-constrained devices, making the model more robust in the face of changes in data distribution.

[0028] In order to better understand the technical solution of the present application, the “quantization” and “flat minimum in domain generalization” involved in the embodiments of the present application are first explained.

[0029] Quantization: This embodiment of the application considers the uniform quantization function of the layer weights and activations: ,in Represents the rounding operator, is a learnable scaling factor in quantization-aware training (QAT), clip Function ensures that the value remains within bounds In b-bit quantization, for activation quantization, set and ; For weight quantization, set and . In addition, to overcome the non-differentiability of the rounding operation, a straight-through estimator (STE) is adopted to approximate the gradient:

[0030] Flat Minima in Domain Generalization: This embodiment of the application uses three optimization objectives to minimize the perceived sharpness: (a) Empirical Risk , (b) Perturbation loss , and (c) agency gap Among them, minimize and It is possible to find low loss areas while minimizing A flat minimum is ensured. This combination of optimization objectives improves both training performance and generalization. Therefore, the overall optimization is: ,in is a hyperparameter and can be further rewritten as: ,in, .

[0031] The following is a detailed description of the domain generalization-oriented gradient adaptive quantization perception training method of an embodiment of the present application in conjunction with the accompanying drawings.

[0032] Reference Figure 1 As shown, Figure 1 This is a flowchart of the steps of a gradient adaptive quantization perception training method for domain generalization provided by an embodiment of the present application. Figure 1 As shown, the domain generalization-oriented gradient adaptive quantization perception training method may include steps S110 to S140: Step S110: Acquire a training data set and a quantized convolutional neural network model, wherein the training data set includes a plurality of sample images of equal amounts in different training domains.

[0033] Among them, the quantized convolutional neural network model can be a quantized convolutional neural network model for processing various computer vision tasks.

[0034] The training data set refers to the data used to train the quantized convolutional neural network model. The training data set includes multiple sample images of equal amounts in different training domains, that is, the number of sample images corresponding to each training domain in the training data set is the same. For example, there are 3 training domains, each with 32 sample images, and the training data set has a total of 96 sample images. The training domain refers to the distribution of data. The data distribution of a training domain is the same, and the data distribution between different training domains is different.

[0035] Step S120: Input each sample image in the training data set into the quantized convolutional neural network model for processing to obtain the task gradient and smooth gradient of each scale factor in different training domains, wherein the scale factor characterizes the characteristics of the weight and activation value distribution of the quantized convolutional neural network model, the task gradient is used to optimize the task performance of the quantized convolutional neural network model, and the smooth gradient is used to optimize the generalization of the quantized convolutional neural network model.

[0036] Among them, the quantized convolutional neural network module includes multiple scale factors, which are used to describe (characterize) the characteristics of the weight and activation value distribution of the quantized convolutional neural network model. The scale factors are highly sensitive to perturbation losses. The obvious convergence of the scale factors to a suboptimal state does not necessarily mean satisfactory convergence, and may have a negative impact on performance outside the training domain. For each convolution layer of the quantized convolutional neural network model, there is a scale factor corresponding to the weight and a scale factor corresponding to the activation value. For example, if the quantized convolutional neural network model has 5 convolution layers, there are a total of 10 scale factors.

[0037] The task gradient refers to the task-related gradient of the quantized convolutional neural network model, and the smooth gradient refers to the flatness-related gradient. The sample image is input into the quantized convolutional neural network model for processing, and each individual training domain will cause characteristic task gradients and smooth gradients to the scale factor, thereby obtaining the task gradients and smooth gradients of each scale factor in different training domains. For example, if there are 5 scale factors and 3 training domains, the sample image is input into the quantized convolutional neural network model for processing, and the task gradients and smooth gradients of the 5 scale factors in the 3 training domains will be obtained respectively.

[0038] Step S130: Evaluate the gradient confusion of the task gradients of each scale factor in different training domains, where the gradient confusion represents the consistency of the task gradient direction.

[0039] In an embodiment of the present application, the inconsistency of the direction of the task gradient during the training process is quantified by the gradient confusion of the task gradients of each scale factor in different training domains. The smaller the gradient confusion, the more consistent the direction of the task gradient of the scale factor in the training domain, which means more stable training; the larger the gradient confusion, the more inconsistent the direction of the task gradient of the scale factor in the training domain, which means more unstable training. It should be noted that although a high gradient confusion does not necessarily mean an incorrect gradient, a low gradient confusion can provide some guarantee of gradient correctness.

[0040] Step S140: According to the gradient confusion degree, determine the learning gradient of each scale factor in different training domains, and update the parameters of the quantized convolutional neural network model based on the learning gradient to obtain a trained quantized convolutional neural network model, wherein the learning gradient at least includes the smooth gradient.

[0041] In the embodiment of the present application, the quantized convolutional neural network model processes the sample image, and the task gradient and smooth gradient of each scale factor in different training domains are obtained, that is, two different gradients of the scale factor in the training domain are obtained, and the learning of the task gradient may interfere with the learning of the smooth gradient. In order to avoid the conflict between the task gradient and the smooth gradient, the learning gradient of each scale factor in different training domains is determined according to the gradient confusion, and the learning gradient includes at least the smooth gradient, that is, according to the gradient confusion, the task gradient is discarded at certain scales, and only the smooth gradient is selected for learning to reduce the conflict between the task gradient and the smooth gradient, and ensure the global convergence of the overall performance.

[0042] Wherein, updating the parameters of the quantized convolutional neural network model based on the learning gradient means: updating the corresponding scale factor based on the learning gradient. For example, if the learning gradient of a scale factor in a training domain is a smooth gradient, the scale factor is updated based on the smooth gradient; if the learning gradient of a scale factor in a training domain is a smooth gradient and a task gradient, the scale factor is updated based on the smooth gradient and the task gradient. In this way, multiple training steps are performed according to the above steps, and after the training end conditions are met, a trained quantized convolutional neural network model is obtained.

[0043] By adopting the technical solution of the embodiment of the present application, a smooth gradient is introduced to update the parameters of the quantized convolutional neural network model, so that the quantization process and the model smoothness can be jointly optimized, thereby improving the generalization ability of the model; and, during the training process, the gradient chaos of the task gradients of each of the scale factors in different training domains is evaluated, and the task gradient is selectively used for learning based on the gradient chaos, thereby enhancing the stability of the model training and ensuring the global convergence of the overall performance. In this way, the smooth gradient and task gradient conflict problem of the quantized model can be finely optimized at the granularity of the domain, thereby improving the generalization ability of the quantized model in unknown fields, and enhancing the generalization ability of the model deployed on resource-constrained devices, making the model more robust in the face of changes in data distribution.

[0044] In combination with the above embodiments, in one implementation, the embodiment of the present application also provides a gradient adaptive quantization perception training method for domain generalization. In this method, step S120 of "inputting each sample image in the training data set into the quantized convolutional neural network model for processing to obtain the task gradient and smooth gradient of each scale factor in different training domains" specifically includes the following sub-steps S120-1 to S120-2: Step S120-1: construct a loss function according to the quantization result of the quantized convolutional neural network model and the sample image, wherein the loss function includes task loss and smoothing loss.

[0045] Step S120 - 2 : determining the task gradients of each scale factor in different training domains according to the task loss, and determining the smoothing gradients of each scale factor in different training domains according to the smoothing loss.

[0046] In an embodiment of the present application, a smoothing objective (smoothed loss) is introduced into the loss function of the quantized convolutional neural network model to perform generalization optimization in the potential weight space.

[0047] For example, the loss function can be expressed as:

[0048] in, Indicates mission loss, represents the smoothing loss, represents the parameters of the quantized convolutional neural network model, represents the scaling factor of the weight, Represents the quantization process of the quantized convolutional neural network model, represents the gradient of the task loss, represents the gradient adjustment parameter, Represents a sample image.

[0049] Compared to full precision training, there are several scaling factors in the loss function , each scale factor corresponds to two optimization objectives (i.e., task objective and smoothing objective), thus generating two sets of gradients, one of which is the task gradient from the task loss, i.e., task loss From , and the other group is the smoothed gradient from the smoothed loss, i.e., the smoothed gradient From .

[0050] By adopting the technical solution of the embodiment of the present application, a smoothing target is introduced into the loss function of the quantized convolutional neural network model, thereby simultaneously receiving two different gradients from the quantization target (task loss) and the perceptual minimization target (smoothing loss), so that the quantization process and model smoothness can be jointly optimized, thereby improving the generalization ability of the model.

[0051] In combination with the above embodiments, in one implementation, the present application embodiment further provides a gradient adaptive quantization perception training method for domain generalization. Specifically: In the method, "evaluating the gradient confusion of the task gradients of each scale factor in different training domains" in step S130 includes: evaluating the gradient confusion of the task gradients of each scale factor in different training domains every target training steps.

[0052] In this method, the step S140 of "determining the learning gradients of each scale factor in different training domains according to the gradient confusion, and updating the parameters of the quantized convolutional neural network model based on the learning gradients" includes: determining the learning gradients of each scale factor in different training domains within the next target training steps according to the gradient confusion, and updating the parameters of the quantized convolutional neural network model based on the learning gradients within the next target training steps.

[0053] In the embodiment of the present application, the training of each sample image corresponds to one training step, and there are a total of multiple training steps for training the quantized convolutional neural network model using the training data set. During the training process, the gradient confusion of the task gradient of each scale factor in different training domains is evaluated every target training steps, and then the learning gradient of each scale factor in different training domains in the next target training steps is determined based on the gradient confusion, so that the quantized convolutional neural network model performs parameter updates based on the learning gradient in the next target training steps.

[0054] For example, for the number of training steps T, the evaluation interval (target training steps) is K. During the training process, each K training steps evaluates the gradient confusion of the task gradient of each scale factor in different training domains, and then determines the learning gradient of each scale factor in different training domains in the next K training steps based on the gradient confusion, and in the next K training steps, updates the parameters of the quantized convolutional neural network model based on the learning gradient.

[0055] In this way, by periodically (every target training steps) re-evaluating the gradient chaos of the task gradients of each scale factor in different training domains and adjusting the learning gradient, this method allows smooth gradients to continue training while alleviating the adverse effects of gradient conflicts, improving overall convergence and enhancing the generalization performance of the model.

[0056] In an optional embodiment, every target training steps, the gradient confusion of the task gradients of each scale factor in different training domains is evaluated, including steps A1 to A2: Step A1: Determine a task gradient sequence according to the task gradient of the scale factor in the training domain within the target training steps; wherein the task gradient of the scale factor in the training domain within each training step is a task gradient in the task gradient sequence.

[0057] Step A2: Obtain the gradient confusion degree according to the direction change of two adjacent task gradients in the task gradient sequence.

[0058] In an embodiment of the present application, for each scale factor in each training domain, a task gradient sequence can be determined, and then based on the directional change of two adjacent task gradients in the task gradient sequence, the gradient confusion of the task gradient of the scale factor in the training domain can be determined.

[0059] Among them, the gradient confusion is obtained according to the direction change of two adjacent task gradients in the task gradient sequence. The number of direction changes of two adjacent task gradients in the task gradient sequence can be used as the gradient confusion; or the ratio of the number of direction changes of two adjacent task gradients in the task gradient sequence to the number of task gradients in the task gradient sequence can be used as the gradient confusion.

[0060] For example, for the task gradient sequence [-1, 1, 2], the directions of task gradient -1 and task gradient 1 are inconsistent, and one direction deflection (direction change) occurs; the directions of task gradient 1 and task gradient 2 are consistent, and there is no direction deflection. That is, in the task gradient sequence [-1, 1, 2], there is one direction deflection in the directions of two adjacent task gradients. The task gradient sequence has a total of 3 task gradients, and the gradient confusion degree can be 1 / 3.

[0061] Further, the task gradient sequence includes a first task gradient sequence and a second task gradient sequence, and two task gradients with the same order in the first task gradient sequence and the second task gradient sequence represent two task gradients of adjacent training steps; In the above step A1, "obtaining the gradient confusion degree according to the change in the directions of two adjacent task gradients in the task gradient sequence" specifically includes: comparing the directions of two task gradients with the same order in the first task gradient sequence and the second task gradient sequence in turn, and determining the number of two task gradients of adjacent training steps in different directions; and taking the ratio of the number to the number of the target training steps as the gradient confusion degree.

[0062] For example, for K training steps, and at each training step j, the corresponding task gradient is , the task gradient from the 1st training step to the K-1th training step can be used as the first task gradient sequence, and the task gradient from the 2nd training step to the Kth training step can be used as the second task gradient sequence. The first task gradient sequence is expressed as: , the second task gradient sequence is expressed as: .

[0063] The gradient chaos can be expressed as:

[0064] in, represents an element-wise sign function, is the indicator function, Indicates the ratio of the number of steps where the gradient direction is opposite to the previous step (i.e., gradient chaos).

[0065] By adopting the technical solution of the embodiment of the present application, the gradient chaos of the task gradient can be quantified, and then the task gradient can be discarded at certain scales according to the gradient chaos, and only the smooth gradient can be selected for learning to reduce the conflict between the task gradient and the smooth gradient and ensure the global convergence of the overall performance.

[0066] In combination with the above embodiments, in one implementation, the embodiment of the present application further provides a domain-generalized gradient adaptive quantization perception training method. In this method, "determining the learning gradients of each scale factor in different training domains according to the gradient confusion degree" in step S140 specifically includes the following sub-steps S140-1 to S140-2: Step S140-1: for the learning gradient of each scale factor in each training domain, when the gradient chaos of the task gradient of the scale factor in the training domain is less than or equal to the gradient chaos threshold, determine the smooth gradient as the learning gradient of the scale factor in the training domain.

[0067] Step S140-2: when the gradient chaos of the task gradient of the scale factor in the training domain is greater than the gradient chaos threshold, determine the smooth gradient and the task gradient as the learning gradient of the scale factor in the training domain.

[0068] There is a corresponding gradient chaos threshold in each training domain, and the gradient chaos thresholds corresponding to different training domains may be the same or different.

[0069] Specifically, for every target training steps (every K steps), the gradient chaos of the task gradients of each scale factor in different training domains is evaluated, and the gradient chaos thresholds corresponding to different training domains are ; If the scale factor The gradient confusion of the j-th training domain task at training step t Less than the gradient chaos threshold , then in the next K training steps, the smooth gradient is used as the learning gradient and the jth training domain data pair is frozen Gradient Otherwise, the gradient is learned in the training domain with the smooth gradient and the task gradient as the scale factor.

[0070] By adopting the technical solution of the embodiment of the present application, a dynamic selective freezing scheme is implemented, which selectively uses the task gradient for learning according to the gradient confusion, reduces the adverse effects of gradient conflicts, enhances the stability of model training, and ensures the global convergence of the overall performance.

[0071] The following is a specific example to illustrate the domain-generalized gradient adaptive quantization perception training method of the present application. Figure 2 As shown, Figure 2 : is a flowchart of another gradient adaptive quantization perception training method for domain generalization provided by an embodiment of the present application, the method comprising steps S210 to S240: Step S210: Obtain a training data set and a quantized convolutional neural network model, wherein the training data set includes a plurality of sample images of equal amounts in different training domains.

[0072] Step S220: Input each sample image in the training data set into the quantized convolutional neural network model for processing, and construct a loss function according to the quantization result of the quantized convolutional neural network model and the sample image, wherein the loss function includes a task loss and a smoothing loss; determine the task gradient of each scale factor in different training domains according to the task loss, and determine the smoothing gradient of each scale factor in different training domains according to the smoothing loss.

[0073] Step S230: every target training steps, evaluate the gradient confusion of the task gradients of each scale factor in different training domains, where the gradient confusion represents the consistency of the task gradient direction.

[0074] Specifically, a task gradient sequence is determined according to the task gradient of the scale factor in the training domain within the target training steps; wherein the task gradient of the scale factor in the training domain within each training step is a task gradient in the task gradient sequence; and the gradient confusion degree is obtained according to the direction change of two adjacent task gradients in the task gradient sequence.

[0075] Step S240: Determine the learning gradients of each scale factor in different training domains within the next target training steps according to the gradient confusion, and update the parameters of the quantized convolutional neural network model based on the learning gradients within the next target training steps.

[0076] Specifically, for the learning gradient of each scale factor in each training domain, when the gradient chaos of the task gradient of the scale factor in the training domain is less than or equal to the gradient chaos threshold, the smooth gradient is determined to be the learning gradient of the scale factor in the training domain; when the gradient chaos of the task gradient of the scale factor in the training domain is greater than the gradient chaos threshold, the smooth gradient and the task gradient are determined to be the learning gradient of the scale factor in the training domain.

[0077] In the embodiment of the present application, the parameters of the quantized convolutional neural network model are updated by introducing a smooth gradient, so that the quantization process and the model smoothness can be jointly optimized, thereby improving the generalization ability of the model; and, during the training process, the gradient chaos of the task gradients of each of the scale factors in different training domains is evaluated, and the task gradient is selectively used for learning based on the gradient chaos, thereby enhancing the stability of the model training and ensuring the global convergence of the overall performance. In this way, the smooth gradient and task gradient conflict problem of the quantized model is finely optimized at the granularity of the domain, improving the generalization ability of the quantized model in unknown fields, and enhancing the generalization ability of the model deployed on resource-constrained devices, making the model more robust in the face of changes in data distribution.

[0078] For example, the domain-generalized gradient-adaptive quantization-aware training method of the embodiment of the present application can be implemented by the dynamic selective freezing strategy of the scale factor in Table 1.

[0079] Table 1 Dynamic selective freezing strategy of scale factors:

[0080] For example, Figure 3 This is a schematic diagram of the architecture of a domain-generalized gradient adaptive quantization perception training method provided by an embodiment of the present application. Compared with the full-precision weight gradient, the domain-generalized gradient adaptive quantization perception training method of the embodiment of the present application has only two directions of tensor-level scale gradient, namely positive and negative, and selectively freezes the newly introduced task-related scale gradient (i.e., task gradient). By evaluating the task gradient of each scale for each training domain, The chaos degree of the task gradient is selected and the task gradients whose chaos degree is lower than the gradient chaos degree threshold are selectively frozen to improve the generalization ability of the model.

[0081] Specifically, each sample image in the training data set is input into the quantized convolutional neural network model for processing to obtain the task gradient of each scale factor in different training domains. and smooth gradient , every K training steps, the gradient disorder of the task gradients of each scale factor in different training domains is evaluated through the first task gradient sequence and the second task gradient sequence, and then the gradient disorder of the task gradient of the scale factor in the training domain is less than or equal to the gradient disorder threshold r (i.e. ), freeze the task gradient and use the smooth gradient as the learning gradient of the scale factor in the training domain; when the gradient chaos of the task gradient of the scale factor in the training domain is greater than the gradient chaos threshold (i.e. ), the learning gradient in the training domain is scaled by the smoothing gradient and the task gradient.

[0082] In this way, the smooth gradient and task gradient conflict problems of the quantization model are finely optimized at the domain granularity, which improves the generalization ability of the quantization model in unknown fields, enhances the generalization ability of the model deployed on resource-constrained devices, and makes the model more robust when facing changes in data distribution.

[0083] The present application also provides an image processing method, referring to Figure 4 As shown, Figure 4 4 is a flowchart of an image processing method provided in an embodiment of the present application, the image processing method comprising steps S410 to S420: Step S410: Acquire the image to be processed.

[0084] Step S420: Input the image to be processed into the trained quantized convolutional neural network model to obtain the target visual task result. The trained quantized convolutional neural network model is trained according to the domain generalization-oriented gradient adaptive quantized perception training method of an embodiment of the present application.

[0085] In the embodiment of the present application, the gradient adaptive quantization perception training method for domain generalization introduces a smooth gradient to update the parameters of the quantized convolutional neural network model, so that the quantization process and the model smoothness can be jointly optimized, thereby improving the generalization ability of the model; and, during the training process, the gradient chaos of the task gradients of each of the scale factors in different training domains is evaluated, and the task gradient is selectively used for learning based on the gradient chaos, thereby enhancing the stability of the model training and ensuring the global convergence of the overall performance. Therefore, according to the gradient adaptive quantization perception training method for domain generalization in the embodiment of the present application, the trained quantized convolutional neural network model has good generalization and task performance, and the image to be processed is processed based on the trained quantized convolutional neural network model to obtain accurate target visual task results.

[0086] The present application also provides a gradient adaptive quantization perception training device for domain generalization, referring to Figure 5 As shown, Figure 5 : is a schematic diagram of a gradient adaptive quantization perception training device for domain generalization provided in an embodiment of the present application, the device comprising: An acquisition module 510 is used to acquire a training data set and a quantized convolutional neural network model, wherein the training data set includes a plurality of sample images of equal amounts in different training domains; A quantization module 520 is used to input each sample image in the training data set into the quantized convolutional neural network model for processing, and obtain task gradients and smooth gradients of each scale factor in different training domains, wherein the scale factor characterizes the characteristics of the weight and activation value distribution of the quantized convolutional neural network model, the task gradient is used to optimize the task performance of the quantized convolutional neural network model, and the smooth gradient is used to optimize the generalization of the quantized convolutional neural network model; An evaluation module 530 is used to evaluate the gradient confusion of the task gradients of each scale factor in different training domains, where the gradient confusion represents the consistency of the task gradient direction; The updating module 540 is used to determine the learning gradients of each scale factor in different training domains according to the gradient confusion, and update the parameters of the quantized convolutional neural network model based on the learning gradients to obtain a trained quantized convolutional neural network model, wherein the learning gradients at least include the smooth gradients.

[0087] In an optional embodiment, the evaluation module is further used to evaluate the gradient confusion of the task gradients of each scale factor in different training domains every target training steps; The update module is also used to determine the learning gradients of each scale factor in different training domains within the next target training steps according to the gradient confusion, and update the parameters of the quantized convolutional neural network model based on the learning gradients within the next target training steps.

[0088] In an optional embodiment, the evaluation module includes: A sequence determination module is used to determine a task gradient sequence according to the task gradient of the scale factor in the training domain within the target training steps; wherein the task gradient of the scale factor in the training domain within each training step is a task gradient in the task gradient sequence; The chaos degree determination module is used to obtain the gradient chaos degree according to the direction change of two adjacent task gradients in the task gradient sequence.

[0089] In an optional embodiment, the task gradient sequence includes a first task gradient sequence and a second task gradient sequence, and two task gradients with the same order in the first task gradient sequence and the second task gradient sequence represent two task gradients of adjacent training steps; The chaos degree determination module is also used to compare the directions of two task gradients with the same order in the first task gradient sequence and the second task gradient sequence in turn, determine the number of two task gradients in adjacent training steps in different directions; and use the ratio of the number to the target number of training steps as the gradient chaos degree.

[0090] In an optional embodiment, the update module includes: A learning gradient determination module is used to determine, for each scale factor in each training domain, the learning gradient, when the gradient chaos of the task gradient of the scale factor in the training domain is less than or equal to a gradient chaos threshold, the smooth gradient as the learning gradient of the scale factor in the training domain; and when the gradient chaos of the task gradient of the scale factor in the training domain is greater than the gradient chaos threshold, determine the smooth gradient and the task gradient as the learning gradient of the scale factor in the training domain.

[0091] In an optional embodiment, the quantization module includes: A construction module, used to construct a loss function according to the quantization result of the quantized convolutional neural network model and the sample image, wherein the loss function includes a task loss and a smoothing loss; The gradient determination module is used to determine the task gradient of each scale factor in different training domains according to the task loss, and to determine the smoothing gradient of each scale factor in different training domains according to the smoothing loss.

[0092] The present application also provides an electronic device, referring to Figure 6 , Figure 6 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 6 As shown, the electronic device 600 includes: a memory 610 and a processor 620, the memory 610 and the processor 620 are connected via a bus communication, the memory 610 stores a computer program, and the computer program can be run on the processor 620, thereby implementing the steps of the domain generalization-oriented gradient adaptive quantization perception training method described in the embodiment of the present application.

[0093] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the domain generalization-oriented gradient adaptive quantization perception training method described in the embodiment of the present application are implemented.

[0094] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the domain generalization-oriented gradient adaptive quantization-aware training method described in the embodiment of the present application.

[0095] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0096] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods and devices according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0097] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0099] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.

[0100] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0101] The above is a detailed introduction to the domain-generalized gradient adaptive quantization perception training method, device and equipment provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A gradient adaptive quantization perception training method for domain generalization, characterized in that: include: Obtaining a training data set and a quantized convolutional neural network model, wherein the training data set includes a plurality of sample images of equal amounts in different training domains; Input each sample image in the training data set into the quantized convolutional neural network model for processing, and obtain the task gradient and smooth gradient of each scale factor in different training domains, wherein the scale factor characterizes the characteristics of the weight and activation value distribution of the quantized convolutional neural network model, the task gradient is used to optimize the task performance of the quantized convolutional neural network model, and the smooth gradient is used to optimize the generalization of the quantized convolutional neural network model; Evaluate the gradient chaos of the task gradients of each scale factor in different training domains, where the gradient chaos represents the consistency of the task gradient direction; According to the gradient confusion, the learning gradients of each scale factor in different training domains are determined, and the parameters of the quantized convolutional neural network model are updated based on the learning gradients to obtain a trained quantized convolutional neural network model, wherein the learning gradients at least include the smooth gradients.

2. The method according to claim 1, characterized in that Evaluate the gradient confusion of task gradients of each scale factor in different training domains, including: Every target training steps, evaluate the gradient confusion of the task gradients of each scale factor in different training domains; Determining the learning gradients of each scale factor in different training domains according to the gradient confusion degree, and updating the parameters of the quantized convolutional neural network model based on the learning gradients, including: According to the gradient confusion, the learning gradients of each scale factor in different training domains within the next target training steps are determined, and the parameters of the quantized convolutional neural network model are updated based on the learning gradients within the next target training steps.

3. The method according to claim 2, characterized in that Every target training steps, evaluate the gradient confusion of the task gradients of each scale factor in different training domains, including: Determine a task gradient sequence according to the task gradient of the scale factor in the training domain within the target training steps; wherein the task gradient of the scale factor in the training domain within each training step is a task gradient in the task gradient sequence; The gradient confusion degree is obtained according to the direction change of two adjacent task gradients in the task gradient sequence.

4. The method according to claim 3, characterized in that The task gradient sequence includes a first task gradient sequence and a second task gradient sequence, wherein two task gradients with the same order in the first task gradient sequence and the second task gradient sequence represent two task gradients of adjacent training steps; The gradient confusion degree is obtained according to the direction change of two adjacent task gradients in the task gradient sequence, including: Comparing the directions of two task gradients with the same order in the first task gradient sequence and the second task gradient sequence in turn, and determining the number of two task gradients in adjacent training steps with different directions; The ratio of the number to the target number of training steps is used as the gradient confusion.

5. The method according to any one of claims 1 to 4, characterized in that: According to the gradient confusion, the learning gradients of each scale factor in different training domains are determined, including: For the learning gradient of each scale factor in each training domain, when the gradient chaos of the task gradient of the scale factor in the training domain is less than or equal to the gradient chaos threshold, determine the smooth gradient as the learning gradient of the scale factor in the training domain; When the gradient chaos of the task gradient of the scale factor in the training domain is greater than the gradient chaos threshold, the smooth gradient and the task gradient are determined as the learning gradient of the scale factor in the training domain.

6. The method according to any one of claims 1 to 4, characterized in that: Input each sample image in the training data set into the quantized convolutional neural network model for processing to obtain the task gradient and smoothing gradient of each scale factor in different training domains, including: Constructing a loss function according to the quantization result of the quantized convolutional neural network model and the sample image, wherein the loss function includes a task loss and a smoothing loss; The task gradients of each scale factor in different training domains are determined according to the task loss, and the smoothing gradients of each scale factor in different training domains are determined according to the smoothing loss.

7. An image processing method, characterized in that: include: Get the image to be processed; The image to be processed is input into a trained quantized convolutional neural network model to obtain a target visual task result, wherein the trained quantized convolutional neural network model is trained according to the domain generalization-oriented gradient adaptive quantized perception training method described in any one of claims 1 to 6.

8. A gradient adaptive quantization perception training device for domain generalization, characterized in that: include: An acquisition module, used to acquire a training data set and a quantized convolutional neural network model, wherein the training data set includes a plurality of sample images of equal amounts in different training domains; A quantization module, used to input each sample image in the training data set into the quantized convolutional neural network model for processing, and obtain task gradients and smooth gradients of each scale factor in different training domains, wherein the scale factor characterizes the characteristics of the weight and activation value distribution of the quantized convolutional neural network model, the task gradient is used to optimize the task performance of the quantized convolutional neural network model, and the smooth gradient is used to optimize the generalization of the quantized convolutional neural network model; An evaluation module is used to evaluate the gradient chaos of the task gradients of each scale factor in different training domains, where the gradient chaos represents the consistency of the task gradient direction; An updating module is used to determine the learning gradients of each scale factor in different training domains according to the gradient confusion, and to update the parameters of the quantized convolutional neural network model based on the learning gradients to obtain a trained quantized convolutional neural network model, wherein the learning gradients at least include the smooth gradients.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the gradient adaptive quantization perception training method for domain generalization described in any one of claims 1 to 6 are implemented, or the steps of the image processing method described in claim 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the domain generalization-oriented gradient adaptive quantization perception training method described in any one of claims 1 to 6 are implemented, or the steps of the image processing method described in claim 7 are implemented.

Citation Information

Patent Citations

  • Domain generalization method and device of image processing model, and image processing method and device

    CN115880538A

  • Model abstract reasoning generalization ability evaluation method based on deep network

    CN116629363A

  • Neural network optimization method based on Taylor expansion momentum correction

    CN118036672A

  • Gradient granularity-based convolutional neural network field generalization classification method

    CN118365950A

  • Robust adaptation method and device based on antagonism explicit task distribution generation

    CN118940805A