Model optimization method, device and equipment for deployment stage and storage medium

CN118865011BActive Publication Date: 2026-10-09PEKING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410807160.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2026-10-09
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

[0004]有鉴于此,本公开提出了一种用于部署阶段的模型优化方法、装置、设备及存储介质,以解决相关技术中由于各种环境因素造成数据分布偏移,导致模型在实际部署环节的性能下降的问题

Benefits of technology

[0004] In view of this, this disclosure proposes a model optimization method, apparatus, device and storage medium for the deployment phase, in order to solve the problem in related technologies where data distribution shifts caused by various environmental factors lead to a decrease in model performance during actual deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118865011B_ABST
    Figure CN118865011B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model optimization method, device and equipment and storage medium for a deployment stage, which comprises: when an image recognition model is deployed to a target domain, obtaining a plurality of first sample data of the target domain; performing feature extraction on the plurality of first sample data respectively to obtain a plurality of feature vectors corresponding to the plurality of first sample data one by one; for any one of the plurality of feature vectors, calculating a value of an augmented entropy loss function corresponding to the feature vector; according to the value of the augmented entropy loss function corresponding to each feature vector, screening a plurality of second sample data from the plurality of first sample data; based on the plurality of second sample data and the corresponding augmented entropy loss function, performing model optimization on the image recognition model to obtain an optimized image recognition model. The present embodiment not only improves the adaptation capability of the model in the deployment stage to the target domain, but also significantly reduces the time and resource consumption of actual deployment while ensuring the adaptation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of deep learning technology, and specifically to a model optimization method, apparatus, device, and storage medium for the deployment phase. Background Technology

[0002] The application of deep neural networks in image recognition is often based on the assumption that the training set (also known as source domain data) and the test set (also known as target domain data) are sampled from the same distribution. In the real world, this assumption often does not hold true. Common problems such as weather changes, equipment damage, and signal noise can cause data distribution shifts, leading to a decrease in the performance of the model in actual deployment.

[0003] Therefore, there is an urgent need for a model optimization method for the deployment phase to solve the above-mentioned technical problems. Summary of the Invention

[0004] In view of this, this disclosure proposes a model optimization method, apparatus, device and storage medium for the deployment phase, in order to solve the problem in related technologies where data distribution shifts caused by various environmental factors lead to a decrease in model performance during actual deployment.

[0005] The first aspect of this disclosure proposes a model optimization method for the deployment phase, the method comprising:

[0006] When an image recognition model is deployed to a target domain, multiple first sample data of the target domain are acquired;

[0007] Feature extraction is performed on the plurality of first sample data respectively to obtain a plurality of feature vectors that correspond one-to-one with the plurality of first sample data;

[0008] For any one of the plurality of feature vectors, calculate the value of the augmented entropy loss function corresponding to the feature vector;

[0009] Based on the value of the augmented entropy loss function corresponding to each feature vector, multiple second sample data are selected from the multiple first sample data; each second sample data refers to the first sample data corresponding to the feature vector whose value is less than a preset threshold.

[0010] Based on the multiple second sample data and the corresponding augmented entropy loss function, the image recognition model is optimized to obtain the optimized image recognition model.

[0011] This embodiment of the disclosure uses the value of the augmented entropy loss function corresponding to each feature vector to select multiple second sample data from multiple first sample data, and optimizes the image recognition model based on the multiple second sample data and the corresponding augmented entropy loss function. This not only improves the model's adaptability to the target domain during the deployment stage, but also significantly reduces the actual deployment time and resource consumption while ensuring the adaptation effect.

[0012] In this embodiment of the disclosure, calculating the value of the augmented entropy loss function corresponding to the feature vector includes:

[0013] Based on the feature vector, the predefined neighborhood Gaussian distribution hyperparameters, the weight matrix of the model classifier, and the bias vector, the value of the augmented entropy loss function corresponding to the feature vector is calculated.

[0014] In this embodiment of the disclosure, based on the feature vector, predefined neighborhood Gaussian distribution hyperparameters, the weight matrix of the model classifier, and the bias vector, the value of the augmented entropy loss function corresponding to the feature vector is calculated, including:

[0015]

[0016] Among them, L AE θ refers to the numerical value of the augmented entropy loss function, C refers to the total number of categories, θ refers to the model parameters of the image recognition model, A refers to the weight matrix, used to linearly transform the feature vector z to a dimension corresponding to the total number of categories C, b refers to the bias vector, used to add an offset after the linear transformation, a i It is the row vector in matrix A corresponding to the i-th class, which is multiplied by the feature vector z to calculate the score of class i.

[0017] In this embodiment of the disclosure, second sample data is selected from the plurality of first sample data based on the value of the augmented entropy loss function corresponding to each feature vector, including:

[0018]

[0019] Where R(x) is the selection function used to select the second sample data from multiple first sample data, and L... AE (x) refers to the value of the augmented entropy loss function corresponding to each feature vector, L0 refers to the predefined preset threshold, 1 means to retain the corresponding feature vector, and 0 means to discard the corresponding feature vector.

[0020] In this embodiment of the disclosure, based on the plurality of second sample data and the corresponding augmented entropy loss function, the image recognition model is optimized to obtain an optimized image recognition model, including:

[0021] For any one of the plurality of second sample data, backpropagation is performed on the second sample data based on the augmented entropy loss function of the second sample data to update the model parameters of the image recognition model.

[0022] In this embodiment of the disclosure, backpropagation of the second sample data is performed based on the augmented entropy loss function of the second sample data, including:

[0023]

[0024] Where θ represents the model parameters, θ * This represents the updated model parameters, and η represents the model's learning rate. It is a partial differential operator, L AE (x i ) represents the augmented entropy loss function, n represents the total number of second sample data, and x represents the value of x. i This represents the i-th second sample data.

[0025] An embodiment of the second aspect of this disclosure provides a model optimization apparatus for the deployment phase, the apparatus comprising:

[0026] The data acquisition module is used to acquire multiple first sample data of the target domain when the image recognition model is deployed to the target domain;

[0027] The feature extraction module is used to extract features from the plurality of first sample data respectively, and obtain a plurality of feature vectors that correspond one-to-one with the plurality of first sample data;

[0028] The numerical calculation module is used to calculate the value of the augmented entropy loss function corresponding to any one of the plurality of feature vectors.

[0029] The data filtering module is used to filter out multiple second sample data from the multiple first sample data based on the value of the augmented entropy loss function corresponding to each feature vector; each second sample data refers to the first sample data corresponding to the feature vector whose value is less than a preset threshold.

[0030] The model optimization module is used to optimize the image recognition model based on the multiple second sample data and the corresponding augmented entropy loss function to obtain the optimized image recognition model.

[0031] An embodiment of the third aspect of this disclosure provides an electronic device including a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the model optimization method for the deployment phase described in the first aspect above.

[0032] An embodiment of the fourth aspect of this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to perform the model optimization method for the deployment phase described in the first aspect above.

[0033] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description or may be learned by practice of this disclosure. Attached Figure Description

[0034] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this disclosure. Furthermore, the same reference numerals denote the same parts throughout the drawings.

[0035] In the attached diagram:

[0036] Figure 1 A schematic flowchart of a model optimization method for the deployment phase provided in an embodiment of this disclosure is shown.

[0037] Figure 2 A flowchart illustrating another model optimization method for the deployment phase provided in an embodiment of this disclosure is shown.

[0038] Figure 3 A schematic diagram of the structure of a model optimization apparatus for the deployment phase provided in an embodiment of the present disclosure is shown.

[0039] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure is shown;

[0040] Figure 5 A schematic diagram of a storage medium provided according to an embodiment of the present disclosure is shown. Detailed Implementation

[0041] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0042] It should be noted that, unless otherwise stated, the technical or scientific terms used in this disclosure shall have the ordinary meaning as understood by one of ordinary skill in the art to which this disclosure pertains.

[0043] The following describes the technical scenarios involved in the embodiments of this disclosure.

[0044] To address the practical challenges of model deployment, existing methods for rapid model adaptation during the deployment phase primarily utilize edge computing power to quickly update the model at the deployment end to adapt to the new data distribution of the test input (which is the target domain). However, these existing methods have the following problems:

[0045] Insufficient utilization of reliable samples: During rapid model updates, input test samples can be categorized into reliable and harmful samples based on their beneficialness to the model. Current methods for rapid model adaptation in the deployment phase often adapt the model to the target domain by minimizing the entropy value of the predicted probability of reliable samples.

[0046] High computational cost: Because new data needs to be generated repeatedly and updated in reverse during the data augmentation process, the computational cost of these steps is high, which will increase the model adaptation time exponentially and fail to meet the actual deployment requirements.

[0047] To address the above shortcomings, this invention provides a highly efficient training method that improves model adaptation efficiency by utilizing single-step ensemble neighborhood augmentation. This method aims to enable rapid model adaptation to test data during the deployment phase, achieving the effect of multi-step augmented training through single-step training by modifying the training objective function. It aims to eliminate computational overhead by integrating multiple augmentations into a single optimization, fully releasing the potential of reliable samples and significantly enhancing the adaptation efficiency of single-step model optimization. This invention achieves the following technical effects:

[0048] 1. Improved sample utilization: The effect of approximating multi-step augmented training is achieved by using only a single backpropagation, which greatly improves the adaptation efficiency.

[0049] 2. Reduce computational cost: By optimizing the upper bound of the loss function, an original augmenting entropy is proposed as the loss function for optimization objectives. This can compress the augmentation process and accelerate model adaptation.

[0050] According to an embodiment of this disclosure, a model optimization method embodiment for the deployment phase is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0051] This embodiment provides a model optimization method for the deployment phase. Figure 1 This is a flowchart of a model optimization method for the deployment phase according to an embodiment of this disclosure, such as... Figure 1 As shown, the process includes the following steps:

[0052] Step S101: When the image recognition model is deployed to the target domain, multiple first sample data of the target domain are acquired.

[0053] In some specific embodiments, the target domain refers to the specific field or scenario in which the image recognition model will be applied, and the model needs to be evaluated and applied within the target domain. The data in the target domain may differ in distribution from the training data; this difference is called domain shift. If the model cannot adapt well to the target domain, performance degradation may occur.

[0054] In some specific embodiments, when an image recognition model is deployed to a target domain, it is necessary to acquire multiple first sample data in the target domain and optimize the image recognition model based on the multiple first sample data, so that the image recognition model can better adapt to the target domain.

[0055] Step S102: Perform feature extraction on the plurality of first sample data respectively to obtain a plurality of feature vectors that correspond one-to-one with the plurality of first sample data.

[0056] In some specific embodiments, the corresponding feature vector can be extracted from the first sample data by feature extraction. The feature extraction method is not specifically limited in this invention, and it can be any method that extracts the corresponding feature vector from the first sample data, such as a feature extractor.

[0057] Step S103: For any one of the plurality of feature vectors, calculate the value of the augmented entropy loss function corresponding to the feature vector.

[0058] In some specific embodiments, step S103 above includes step S1031:

[0059] Step S1031: Based on the feature vector, the predefined neighborhood Gaussian distribution hyperparameters, the weight matrix of the model classifier, and the bias vector, calculate the value of the augmented entropy loss function corresponding to the feature vector.

[0060] In this embodiment of the disclosure, the above step S1031 can be performed using the following formula to calculate the value of the augmented entropy loss function corresponding to the feature vector:

[0061]

[0062] Among them, L AE θ refers to the numerical value of the augmented entropy loss function, C refers to the total number of categories, θ refers to the model parameters of the image recognition model, A refers to the weight matrix, used to linearly transform the feature vector z to a dimension corresponding to the total number of categories C, b refers to the bias vector, used to add an offset after the linear transformation, a i It is the row vector in matrix A corresponding to the i-th class, which is multiplied by the feature vector z to calculate the score of class i.

[0063] Step S104: Based on the value of the augmented entropy loss function corresponding to each feature vector, select multiple second sample data from the multiple first sample data.

[0064] In some specific embodiments, each second sample data refers to the first sample data corresponding to the feature vector whose value is less than a preset threshold. The preset threshold can be set according to actual circumstances and is not specifically limited here.

[0065] In some specific embodiments, the second sample data can be selected from multiple first sample data using the following formula:

[0066]

[0067]

[0068] Where R(x) is the selection function used to select the second sample data from multiple first sample data, and L... AE (x) refers to the value of the augmented entropy loss function corresponding to each feature vector, L0 refers to the predefined preset threshold, 1 means to retain the corresponding feature vector, and 0 means to discard the corresponding feature vector.

[0069] In this embodiment of the disclosure, the above formula can be used to determine: when the L of the first sample data AE If (x) is greater than or equal to the preset threshold L0, then the first sample data can be discarded; similarly, if the first sample data L... AE If (x) is less than the preset threshold L0, then the first sample data can be retained.

[0070] Step S105: Based on the multiple second sample data and the corresponding augmented entropy loss function, the image recognition model is optimized to obtain the optimized image recognition model.

[0071] In this step, the image recognition model is optimized using the retained sample data, namely the second sample data, and the augmented entropy loss function corresponding to the second sample data, to obtain the optimized image recognition model; the optimized image recognition model can adapt well to the data distribution of the target domain.

[0072] In some specific embodiments, step S105 above includes step S1051:

[0073] Step S1051: For any one of the plurality of second sample data, backpropagation is performed on the second sample data based on the augmented entropy loss function of the second sample data to update the model parameters of the image recognition model.

[0074] In this step, the augmented entropy loss function of the second sample data is used to backpropagate the second sample data, which can simulate the effect of multiple data augmentations, thereby greatly improving the model's generalization and adaptability.

[0075] In some specific embodiments, the formula for calculating the updated model parameter θ using backpropagation is as follows:

[0076]

[0077] Where θ represents the model parameters, θ * This represents the updated model parameters, and η represents the model's learning rate. It is a partial differential operator, L AE (x i ) represents the augmented entropy loss function, n represents the total number of second sample data, and x represents the value of x. i This represents the i-th second sample data.

[0078] To address the issue of data distribution shifts caused by various environmental factors, leading to performance degradation in actual deployment of models in related technologies, this invention provides another specific embodiment, as follows:

[0079] This invention starts with the entropy loss function of multi-step augmentation, theoretically analyzes its upper bound, and provides an efficient training method to improve model adaptation efficiency by utilizing single-step ensemble neighborhood augmentation. This method aims to enable the model to quickly adapt to test data during the deployment phase, achieving the effect of multi-step augmentation training through single-step training by modifying the training objective function. This invention comprises two main parts: optimization of the augmented entropy loss of single-step ensemble and design of a selection mechanism.

[0080] Part 1: Optimization of Augmented Entropy Loss in Single-Step Ensemble

[0081] In optimizing the augmented entropy loss, this invention theoretically derives an upper bound on the expected entropy loss and indirectly supervises the training of augmented samples by optimizing this upper bound. This method can integrate the positive effects of multiple rounds of data augmentation into the single-step training of the model in a computationally efficient manner, thereby significantly improving the adaptation speed and reducing the consumption of computational resources.

[0082] Specifically, consider a sample and its corresponding features z∈R d The probability output by the classifier is given by the following formula:

[0083]

[0084] Where C is the total number of categories, p θ (z) i θ represents the probability that the classifier predicts a feature vector z as class i, where θ represents the model parameters. A refers to the weight matrix, used to linearly transform the input feature z to a dimension corresponding to the number of classes C. b refers to the bias vector, used to add an offset after the linear transformation. i It is the row vector in matrix A corresponding to the i-th class, which is multiplied by the input feature z to calculate the score of class i.

[0085] In the process of neighborhood augmentation, this invention treats the sample z [z is a feature vector] in the feature space as a Gaussian distribution within the neighborhood. New data is generated by Monte Carlo sampling from this distribution. Where ∑ is a predefined covariance matrix used to determine the range of the neighborhood. Since the feature z is augmented as... The resulting change in classifier prediction probability reflects the uncertainty of the model when encountering perturbations [“perturbation” refers to neighborhood augmentation of the model’s feature z]. This invention calculates the expected value of the output probability to include perturbation effects in all directions, thereby achieving more robust classification prediction. Specifically, the expected classification probability of the classifier is:

[0086]

[0087] The above passage describes a single neighborhood augmentation.

[0088] Considering a finite number of steps (i.e., N steps) of augmentation, the model will minimize the following loss function L. N Perform training. [Finite-step augmentation: Treat the neighborhood of z in the feature space as a Gaussian distribution] New features are obtained by Monte Carlo sampling based on this distribution. The sampling is performed N times in total, and the feature obtained from the i-th sampling is 】

[0089]

[0090] However, this method of directly augmenting N times leads to a linear increase in computational cost with the number of augmentations. When N is large, augmenting feature z to obtain robust classification probability predictions can enhance the adaptability of the model during the deployment phase, but it cannot meet the time requirements for model adaptation during the deployment phase. To solve this problem, this invention uses the above equation L... N The expression is generalized to the case where N approaches infinity, i.e., L ∞ Export L at the same time ∞ A closed-form upper bound is used as the new training loss function, and a single backpropagation is performed to achieve the effect of multiple rounds of data augmentation, thereby enabling rapid adaptation of the model during the deployment phase.

[0091] Specifically, L ∞ The upper bound can be represented in the following form:

[0092]

[0093] This invention will L ∞ The upper bound of the closed form is defined as the augmenting entropy L. AE And use it as the loss function for model training.

[0094] The statement "When the number of iterations N approaches infinity, the upper bound of this loss function can be obtained in the form of a closed-form solution" explains the upper bound of the loss function: As the number of samples N approaches infinity, the expression L can be derived through inequality scaling. ∞ It represents L N Behavior when N is infinite. Expression L ∞ This provides a theoretical upper bound for the loss function value, meaning that no matter how large the value of N is in a practical application, the value of the loss function will not exceed this upper bound.

[0095] Part 2: Screening Mechanism Design

[0096] In the design of the screening mechanism, to effectively filter out harmful samples that may lead to incorrect model adaptation, this invention introduces a screening strategy based on the augmented entropy loss function. This strategy evaluates the contribution of a sample to the augmented entropy loss and retains only those samples that have a positive impact on the model adaptation capability during the deployment phase, thereby ensuring the reliability and effectiveness of the adaptation process.

[0097] Specifically, this invention incorporates a selection function R(x) that retains only samples whose augmented entropy is less than a fixed threshold L0, allowing them to participate in model adaptation. This is achieved by combining the selection function R(x) and the augmented entropy L0. AE This invention proposes a loss function for deployment phase model adaptation with a filtering mechanism, the expression of which is as follows:

[0098] min θ R(x)L AE (x), where

[0099] Where I is the indicator function and L0 is a pre-set threshold.

[0100] In this embodiment of the invention, by applying the two key technologies described in Parts 1 and 2, the present invention not only improves the adaptability of the deployment phase model to the target domain, but also significantly reduces the actual deployment time and resource consumption while ensuring the adaptation effect. This enables the present invention to provide a more efficient and practical adaptive solution for the system when facing rapidly changing environments.

[0101] To address the performance degradation of models in actual deployment due to data distribution shifts caused by various environmental factors in related technologies, this invention also provides another specific embodiment, as follows: Figure 2 As shown:

[0102] Figure 2 (a) is a flowchart illustrating the finite-step augmentation process: the leftmost Data Stream puppy represents the input data stream, which passes through the Feature Extractor to obtain features z [blue circle]. Next, a finite number of vitinal augmentation steps [N times, or N times] are performed, each time treating the neighborhood of z [blue circle] in the feature space as a Gaussian distribution. Then, a Monte Carlo sampling is performed based on this distribution to obtain new features. [Green circle], sampling was performed N times in total, and the feature obtained from the i-th sampling is: Finally, the loss function L is used. N Backpropagation is performed on the model to update it, thereby improving the model's adaptability during the deployment phase.

[0103] Figure 2 (b) refers to the flowchart of the present invention: its technical solution focuses on [InfiniteAugmentation], indicating that the present invention will use the above formula L NThe expression is generalized to the case where N approaches infinity, that is, the expression is expanded around z [blue circle] in the feature space to obtain an infinite number of features. [The remaining circles], resulting in L ent That is, L ∞ Further derivation of L ∞ The upper bound is defined as the augmented entropy loss function L. AE Finally, the augmented entropy loss function is used to perform single-step backpropagation updates on the model, thereby improving the model's adaptability during the deployment phase.

[0104] Since this is a single-step process, rather than the N-step process in Figure (a), this invention not only improves the adaptability of the deployment phase model to the target domain, but also significantly reduces the actual deployment time and resource consumption while ensuring the adaptation effect.

[0105] To address the issue of data distribution shifts caused by various environmental factors, leading to performance degradation in actual deployment of models in related technologies, this invention also provides another embodiment, as follows:

[0106] 1. Data preprocessing: Obtain the original data sample x, and perform necessary cleaning and formatting on the data sample, such as standardization, to ensure that the data quality is suitable for subsequent processing.

[0107] 2. Using the model's feature extractor, extract the feature vector z from the preprocessed sample x.

[0108] 3. Using the weight matrix A, bias vector b, and predefined neighborhood Gaussian distribution hyperparameter Σ of the model classifier, calculate the augmented entropy loss function L. AE :

[0109]

[0110] 4. Introduce a screening mechanism for augmented entropy: Only retain samples whose augmented entropy is less than a fixed threshold L0, by combining the screening function R(x) and the augmented entropy L0. AE The loss function for adapting a deployment phase model with a filtering mechanism is calculated, and its expression is as follows:

[0111] min θ R(x)L AE (x), where Where I is the indicator function, i.e. L0 is a pre-set threshold.

[0112] 5. Model Update: Change R(x)L AE(x) serves as the loss function for optimization. A single-step backpropagation is applied to each sample (ensuring rapid model updates) to simulate multiple data augmentations, thereby improving the model's generalization ability and adaptation speed. The formula for updating the model parameters θ using backpropagation is as follows:

[0113]

[0114] Where η represents the learning rate of the model, It is a partial differential operator.

[0115] 6. Termination condition: The model can continue to be updated as long as test data is input.

[0116] Corresponding to the above implementation of the model optimization method for the deployment phase, this disclosure also provides a model optimization apparatus for the deployment phase, used to perform the above... Figures 1 to 2 The illustrated embodiment describes a model optimization method for the deployment phase. For example... Figure 3 As shown, the model optimization device for the deployment phase includes:

[0117] The data acquisition module is used to acquire multiple first sample data of the target domain when the image recognition model is deployed to the target domain;

[0118] The feature extraction module is used to extract features from the plurality of first sample data respectively, and obtain a plurality of feature vectors that correspond one-to-one with the plurality of first sample data;

[0119] The numerical calculation module is used to calculate the value of the augmented entropy loss function corresponding to any one of the plurality of feature vectors.

[0120] The data filtering module is used to filter out multiple second sample data from the multiple first sample data based on the value of the augmented entropy loss function corresponding to each feature vector; each second sample data refers to the first sample data corresponding to the feature vector whose value is less than a preset threshold.

[0121] The model optimization module is used to optimize the image recognition model based on the multiple second sample data and the corresponding augmented entropy loss function to obtain the optimized image recognition model.

[0122] Optionally, the numerical calculation module is also used to: calculate the value of the augmented entropy loss function corresponding to the feature vector based on the feature vector, the predefined neighborhood Gaussian distribution hyperparameters, the weight matrix of the model classifier, and the bias vector.

[0123] Optionally, the numerical calculation module is also used to calculate the value of the augmented entropy loss function corresponding to the feature vector using the following formula:

[0124]

[0125] Among them, L AE θ refers to the numerical value of the augmented entropy loss function, C refers to the total number of categories, θ refers to the model parameters of the image recognition model, A refers to the weight matrix, used to linearly transform the feature vector z to a dimension corresponding to the total number of categories C, b refers to the bias vector, used to add an offset after the linear transformation, a i It is the row vector in matrix A corresponding to the i-th class, which is multiplied by the feature vector z to calculate the score of class i.

[0126] Optionally, the data filtering module is also used to filter out second sample data from the plurality of first sample data using the following formula:

[0127]

[0128] Where R(x) is the selection function used to select the second sample data from multiple first sample data, and L... AE (x) refers to the value of the augmented entropy loss function corresponding to each feature vector, L0 refers to the predefined preset threshold, 1 means to retain the corresponding feature vector, and 0 means to discard the corresponding feature vector.

[0129] Optionally, the model optimization module is further configured to: for any one of the plurality of second sample data, perform backpropagation on the second sample data based on the augmented entropy loss function of the second sample data, so as to update the model parameters of the image recognition model.

[0130] Optionally, the model optimization module is also used to backpropagate the second sample data using the following formula:

[0131]

[0132] Where θ represents the model parameters, θ * This represents the updated model parameters, and η represents the model's learning rate. It is a partial differential operator, L AE (x i ) represents the augmented entropy loss function, n represents the total number of second sample data, and x represents the value of x. i This represents the i-th second sample data.

[0133] The model optimization apparatus for the deployment phase provided in the above embodiments of this disclosure and the model optimization method for the deployment phase provided in the embodiments of this disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0134] This disclosure also provides an electronic device for performing the model optimization method described above for the deployment phase. Please refer to... Figure 4 This illustrates a schematic diagram of an electronic device provided by some embodiments of the present disclosure. For example... Figure 4 As shown, the electronic device 4 includes: a processor 400, a memory 401, a bus 402, and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected via the bus 402. The memory 401 stores a computer program that can run on the processor 400. When the processor 400 runs the computer program, it executes the aforementioned provisions of this disclosure. Figures 1 to 2 The illustrated implementation provides a model optimization method for the deployment phase.

[0135] The memory 401 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 403 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network.

[0136] Bus 402 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Memory 401 is used to store programs, and the processor 400 executes the programs after receiving execution instructions. Figures 1 to 2 The model optimization method for the deployment phase disclosed in any of the illustrated embodiments can be applied to or implemented by the processor 400.

[0137] The processor 400 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 400 or by instructions in software form. The processor 400 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 401. The processor 400 reads the information in memory 401 and, in conjunction with its hardware, completes the steps of the above method.

[0138] The electronic device provided in this disclosure and the model optimization method for the deployment phase provided in this disclosure are based on the same inventive concept and have the same beneficial effects as the methods they employ, operate, or implement.

[0139] This disclosure also provides a computer-readable storage medium corresponding to the model optimization method for the deployment phase provided in the foregoing embodiments. Please refer to... Figure 5 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it executes the model optimization method for the deployment phase provided in any of the foregoing embodiments.

[0140] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.

[0141] The computer-readable storage medium provided in the above embodiments of this disclosure and the model optimization method for the deployment phase provided in the embodiments of this disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0142] It should be noted that:

[0143] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this disclosure may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0144] Similarly, it should be understood that, in order to simplify this disclosure and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of this disclosure, various features of this disclosure are sometimes grouped together in a single embodiment, figure, or description thereof. However, this approach to disclosure should not be construed as reflecting a schematic diagram in which the claimed disclosure requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, inventive aspects lie in fewer than all features of a single foregoing disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this disclosure.

[0145] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are intended to be within the scope of this disclosure and form different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0146] The above description is merely a preferred embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A model optimization method for the deployment phase, characterized in that, The method includes: When an image recognition model is deployed to a target domain, multiple first sample data of the target domain are acquired; Feature extraction is performed on the plurality of first sample data respectively to obtain a plurality of feature vectors that correspond one-to-one with the plurality of first sample data; For any one of the plurality of feature vectors, calculate the value of the augmented entropy loss function corresponding to the feature vector; Based on the value of the augmented entropy loss function corresponding to each feature vector, multiple second sample data are selected from the multiple first sample data; each second sample data refers to the first sample data corresponding to the feature vector whose value is less than a preset threshold. Based on the multiple second sample data and the corresponding augmented entropy loss function, the image recognition model is optimized to obtain the optimized image recognition model. Calculating the numerical value of the augmented entropy loss function corresponding to the feature vector includes: Based on the feature vector, predefined neighborhood Gaussian distribution hyperparameters, the weight matrix of the model classifier, and the bias vector, the numerical value of the augmented entropy loss function corresponding to the feature vector is calculated, including: in, This refers to the numerical value of the augmented entropy loss function. This refers to the total number of categories. This refers to the weight matrix, used to weight the eigenvectors. Linear transformation to the total number of categories Corresponding dimensions This refers to the bias vector, which is used to add an offset after a linear transformation. It is a matrix The corresponding number in the middle The row vectors of the class, and the feature vectors Multiplication is used to calculate categories. The score.

2. The method according to claim 1, characterized in that, Based on the value of the augmented entropy loss function corresponding to each feature vector, second sample data is selected from the plurality of first sample data, including: in, This refers to a filtering function used to select a second set of data from a set of first-sample data. This refers to the numerical value of the augmented entropy loss function corresponding to each feature vector. This refers to a predefined threshold; 1 indicates that the corresponding feature vector is retained, and 0 indicates that the corresponding feature vector is discarded.

3. The method according to claim 1, characterized in that, Based on the multiple second sample data and the corresponding augmented entropy loss function, the image recognition model is optimized to obtain an optimized image recognition model, including: For any one of the plurality of second sample data, backpropagation is performed on the second sample data based on the augmented entropy loss function of the second sample data to update the model parameters of the image recognition model.

4. The method according to claim 3, characterized in that, Based on the augmented entropy loss function of the second sample data, backpropagation is performed on the second sample data, including: in, Indicates model parameters, This represents the updated model parameters. This represents the model's learning rate. It is a partial differential operator. Represents the augmented entropy loss function. This indicates the total number of data points in the second sample. This represents the i-th second sample data.

5. A model optimization device for the deployment phase, characterized in that, The device includes: The data acquisition module is used to acquire multiple first sample data of the target domain when the image recognition model is deployed to the target domain; The feature extraction module is used to extract features from the plurality of first sample data respectively, and obtain a plurality of feature vectors that correspond one-to-one with the plurality of first sample data; The numerical calculation module is used to calculate the value of the augmented entropy loss function corresponding to any one of the plurality of feature vectors. The data filtering module is used to filter out multiple second sample data from the multiple first sample data based on the value of the augmented entropy loss function corresponding to each feature vector; each second sample data refers to the first sample data corresponding to the feature vector whose value is less than a preset threshold. The model optimization module is used to optimize the image recognition model based on the multiple second sample data and the corresponding augmented entropy loss function to obtain the optimized image recognition model. Calculating the numerical value of the augmented entropy loss function corresponding to the feature vector includes: Based on the feature vector, predefined neighborhood Gaussian distribution hyperparameters, the weight matrix of the model classifier, and the bias vector, the numerical value of the augmented entropy loss function corresponding to the feature vector is calculated, including: in, This refers to the numerical value of the augmented entropy loss function. This refers to the total number of categories. Weight matrix, used to weight feature vectors Linear transformation to the total number of categories Corresponding dimensions This refers to the bias vector, which is used to add an offset after a linear transformation. It is a matrix The corresponding number in the middle The row vectors of the class, and the feature vectors Multiplication is used to calculate categories. The score.

6. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the model optimization method for the deployment phase as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the model optimization method for the deployment phase as described in any one of claims 1 to 4.