Test-time adaptation method based on covariance matrix zero-order optimization

By inserting lightweight adapter and covariance matrix-driven perturbation strategy into the pretrained model, the problem of high memory and computing resources demand in edge devices during existing tests is solved, and efficient and stable model adaptation is achieved, which is suitable for the real-time optimization requirements of edge computing devices.

CN120258069APending Publication Date: 2025-07-04SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510290965.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing testing adaptive technology relies on backpropagation, resulting in high memory and computing resource requirements, difficult to apply in edge devices, and exhibit instability when dealing with high uncertainty or data contamination, especially when distribution changes in large-scale or complex architectures, and difficult to effectively constrain parameter updates.

Method used

The test-time adaptation method based on zero-order optimization of the covariance matrix is adopted. By inserting a lightweight adapter into the pretrained model, combining covariance-driven perturbation strategy and perturbation attenuation factor, the optimization direction is dynamically adjusted, and the model parameter update without backpropagation is achieved. The finite difference gradient estimation and dynamic update of the covariance matrix are used to update only the weight of the adapter layer.

Benefits of technology

It significantly reduces memory consumption, improves the stability and adaptability of the model, accelerates the convergence speed, and is suitable for resource-constrained edge computing devices, especially in fields such as intelligent manufacturing, medical image analysis and unmanned driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258069A_ABST
    Figure CN120258069A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptation method during testing based on covariance matrix zero-order optimization, and relates to the field of machine learning model optimizing.The method comprises the steps that a lightweight adapter is inserted into a backbone network of a pre-training model, and a sampling matrix is initialized; dynamically constructing a covariance matrix, dynamically adjusting the disturbance direction by using the covariance matrix, and sampling based on a step length control factor; finite difference gradient estimation is executed in forward propagation, and a gradient estimation value is updated in combination with a disturbance reduction factor; dynamically updating the covariance matrix and adapter parameters; and when feature distribution deviation is detected, model optimization is carried out through covariance matrix guided disturbance. The dynamic covariance matrix is constructed to guide the disturbance direction, a disturbance attenuation factor mechanism and an adapter framework are combined, back propagation model parameter optimization is not needed, and the method has the advantages of being high in precision, low in memory and high in stability, is suitable for edge calculation scenes such as industrial quality inspection and medical image analysis and has remarkable industrial application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of machine learning model optimization, and particularly to a test-time adaptation method based on covariance matrix zeroth-order optimization. Background Art

[0002] With the wide application of deep learning technology in practical scenarios, test-time adaptation (TTA) methods have received increasing attention, especially in edge devices where data distributions change dynamically. However, traditional TTA methods usually rely on backpropagation (BP), which adapts to new data distributions by explicitly calculating gradients and adjusting model parameters. Although this method can effectively improve the performance of the model during testing, its high computational and memory requirements make it difficult to implement in resource-constrained environments. For example, the BP method not only needs to store activation values and gradients but also retain optimizer states (such as momentum terms and learning rate estimates), resulting in a significant increase in memory requirements.

[0003] In addition, when traditional TTA methods handle situations with high uncertainty or contaminated data, their optimization processes often show instability. For example, unstructured parameter updates may lead to catastrophic forgetting (CF) of model features, thereby reducing the overall performance and robustness of the model. Especially when the model needs to adapt to distribution changes in large-scale or complex architectures (such as ViT-B / 16), existing optimization methods (such as stochastic zeroth-order optimization) are insufficient in terms of directional guidance ability and are difficult to effectively constrain parameter updates.

[0004] To solve the above problems, the present invention proposes a test-time adaptation method based on covariance matrix zeroth-order optimization (Covariance Matrix-based Zeroth-Order Optimization, COZO). By introducing a covariance-driven perturbation strategy and a perturbation decay factor (PRF), this method can dynamically adjust the optimization direction and achieve fine-grained guidance for parameter updates. In addition, the COZO method uses a lightweight adapter design to freeze the core parameters of the pre-trained model and only update the weights of the adapter layer, thereby ensuring the stability of the model while achieving efficient adaptation.

[0005] Therefore, those skilled in the art are committed to developing a test-time adaptation method based on covariance matrix zeroth-order optimization. Summary of the Invention

[0006] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is that existing test-time adaptation technologies rely on backpropagation, require high memory and computing resources, and are difficult to apply in edge devices.

[0007] To achieve the above object, the present invention provides a test-time adaptation method based on zero-order optimization of the covariance matrix. By combining the zero-order optimization method and the covariance-driven perturbation strategy, efficient model parameter updates can be achieved without backpropagation, significantly reducing memory consumption and enhancing stability and adaptation effects. The method includes the following steps:

[0008] S101: Insert a lightweight adapter into the backbone network of the pre-trained model and initialize the sampling matrix of the adapter;

[0009] S103: Dynamically construct a covariance matrix, use the covariance matrix to dynamically adjust the perturbation direction, and perform non-Gaussian random sampling based on a step-size control factor;

[0010] S105: Perform finite-difference gradient estimation in forward propagation and update the gradient estimation value in combination with a perturbation reduction factor;

[0011] S107: Dynamically update the covariance matrix and the adapter parameters;

[0012] S109: When a feature distribution shift is detected, optimize the model through covariance matrix-guided perturbation.

[0013] Further, in the step S101, the adapter is composed of two layers of fully connected networks. The adapter is inserted in parallel into the MLP module in the backbone network of the pre-trained model. The adapter adopts a single-layer or multi-layer configuration. The downsampling matrix of the adapter is initialized using the Kaiming initialization method, and the upsampling matrix of the adapter is initialized to zero.

[0014] Further, when the adapter adopts a single-layer configuration, the adapter adopts a compression ratio of 384:768 and is inserted into the 3rd layer MLP module of the backbone network of the pre-trained model.

[0015] Further, when the adapter adopts a multi-layer configuration, the adapter adopts a compression ratio of 256:768 and is inserted into different layer MLP modules of the backbone network of the pre-trained model.

[0016] Further, the adapter adopts a 3-layer configuration and is inserted into the 1st layer, 3rd layer, and 6th layer MLP modules of the backbone network of the pre-trained model.

[0017] Further, in the step S103, the sampling process combines the covariance matrix-guided sampling strategy to dynamically adjust the perturbation direction. The sampling process follows the following formula:

[0018] Zk ~m + τN(0, C)

[0019] Among them, Z k is the perturbation amount, m is the offset mean value to ensure that the sampling distribution does not deviate from the original feature center; τ is the scaling coefficient used to adjust the perturbation intensity; C is the covariance matrix, and N is the normal distribution sampling.

[0020] Furthermore, in the step S105, when performing the finite difference gradient estimation, a stochastic gradient estimation strategy is used, and the gradient direction of each iteration is calculated through the finite difference formula. The stochastic gradient estimation adopts the following method:

[0021]

[0022] The calculation method of the gradient estimation value:

[0023]

[0024] Among them, L is the loss function, W is the network parameter, Z i is the perturbation amount, q is the number of sampling times, σ is the perturbation reduction factor, and ∈ is the perturbation amplitude.

[0025] Furthermore, in the step S107, the covariance matrix is dynamically updated using the following update strategy:

[0026] Update in real time according to the input data: Calculate the outer product matrix of the perturbation vectors in real time according to the input data of each batch;

[0027] Fuse the historical covariance information using a forgetting factor;

[0028] Perform a positive semi - definite correction on the covariance matrix to control the stability of the covariance matrix data.

[0029] Furthermore, in the step S107, the adapter parameters are dynamically updated using the perturbation reduction factor mechanism, including:

[0030] In the initial stage, the perturbation reduction factor prevents excessive perturbation from causing model instability by adjusting the perturbation amplitude;

[0031] In the working stage, when the update amplitude meets the preset threshold of the perturbation reduction factor, the update magnification remains unchanged; when the amplitude deviates from the set range, the perturbation reduction factor dynamically adjusts the perturbation amplitude to enhance the stability of the entire adaptation process;

[0032] The perturbation reduction factor is defined as: σ = ∈ 2 , where σ is the perturbation reduction factor and ∈ is the perturbation amplitude.

[0033] Further, it is characterized in that the method realizes memory optimization by adopting the following mechanism:

[0034] Discard the computational graph cache related to backpropagation and only store the forward activation calculation results;

[0035] Store the perturbation vector in fixed-point format and reduce the storage requirement of the perturbation vector through floating-point optimization;

[0036] Dynamically release the invalid intermediate feature map cache.

[0037] In a preferred embodiment of the present invention, compared with the prior art, the present invention has the following beneficial effects:

[0038] 1. High optimization performance: The present invention accelerates the convergence speed through the perturbation strategy driven by the covariance matrix, achieving higher optimization accuracy than the traditional BP-Free method.

[0039] 2. Lower memory overhead: The present invention adopts a zero-order optimization design, avoiding the high memory occupation of backpropagation and significantly reducing the resource requirements during the adaptation process.

[0040] 3. Improved training stability: The present invention alleviates the feature collapse problem and improves the stability of the update process through the PRF mechanism and local update of the adapter layer.

[0041] 4. Strong integration and adaptability: The present invention takes the covariance matrix as the core and realizes efficient test-time adaptation in complex environments through an innovative zero-order optimization method. The present invention can be efficiently deployed in edge computing devices, can adapt to dynamic domain distribution offset scenarios, has strong scalability, and is particularly suitable for the real-time optimization requirements of edge computing devices, and has broad application prospects in the fields of intelligent manufacturing, medical image analysis, unmanned driving, etc.

[0042] The following will further illustrate the concept, specific structure and technical effects of the present invention with reference to the accompanying drawings, so as to fully understand the purpose, features and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a schematic flowchart of the test-time adaptation method based on covariance matrix zero-order optimization according to a preferred embodiment of the present invention;

[0044] Figure 2 is a schematic framework diagram of the test-time adaptation method based on covariance matrix zero-order optimization according to a preferred embodiment of the present invention;

[0045] Figure 3 is a schematic comparison diagram of the gradient distribution calculated by the COZO method and the gradient distribution calculated by the traditional zero-order optimization method according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0047] In the drawings, components with the same structure are denoted by the same numerical reference signs, and components with similar structures or functions everywhere are denoted by similar numerical reference signs. The size and thickness of each component shown in the drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. To make the illustration clearer, the thickness of some parts in the drawings is appropriately exaggerated.

[0048] To solve the problems of high memory overhead, low optimization efficiency, and poor stability in existing test-time adaptation technologies, the present invention proposes a test-time adaptation method based on zero-order optimization of the covariance matrix. By combining the zero-order optimization method (Zeroth-Order Optimization, ZOO) and the covariance-driven perturbation strategy, this method realizes efficient model parameter updates without backpropagation, significantly reduces memory consumption, and improves stability and adaptation effects. This method is particularly suitable for resource-constrained edge computing device scenarios, and can perform real-time and efficient adaptation to new data distributions during testing, with significant application value.

[0049] As Figure 1 - Figure 2 shown, a test-time adaptation method based on zero-order optimization of the covariance matrix provided by an embodiment of the present invention includes the following steps:

[0050] S101: Insert a lightweight adapter into the backbone network of the pre-trained model, and initialize and configure the sampling matrix of the adapter.

[0051] To further improve the model adaptation efficiency and reduce the interference to the backbone network parameters, the adapter layer of the present invention adopts the Kaiming-zero initialization strategy, and inserts a lightweight adaptation structure (parameter quantity <1.2M) in parallel in the MLP module of the ViT network to ensure that the core parameters of the model are frozen.

[0052] To further reduce the memory and parameter update overhead, the present invention only updates the parameters of the adapter layers (Adapter Layers) inserted into the backbone network, and freezes other parts of the pre-trained model. The adapter layer uses an architecture with a compression ratio of 384:768, etc., and adopts the design of Kaiming initialization for the downsampling matrix and zero initialization for the upsampling matrix, which can efficiently process large-scale models such as ViT.

[0053] In this embodiment, the lightweight adapter consists of two fully connected networks. The adapter is inserted in parallel into the MLP module in the backbone network of the pre-trained model. The adapter adopts a single-layer or multi-layer configuration. The downsampling matrix of the adapter is initialized using the Kaiming initialization method, and the upsampling matrix of the adapter is initialized to zero. MLP is a Multilayer Perceptron, which is a feedforward artificial neural network model that maps multiple input data sets to a single output data set.

[0054] When the adapter adopts a single-layer configuration, the adapter adopts a compression ratio of 384:768, and the adapter is inserted into the 3rd layer MLP module of the backbone network of the pre-trained model.

[0055] When the adapter adopts a multi-layer configuration, the adapter adopts a compression ratio of 256:768, and the adapter is inserted into the MLP modules of different layers of the backbone network of the pre-trained model.

[0056] In this embodiment, when the adapter adopts a multi-layer configuration, the adapter adopts a 3-layer configuration, and the adapter is inserted into the 1st layer, 3rd layer, and 6th layer MLP modules of the backbone network of the pre-trained model.

[0057] S103: Dynamically construct the covariance matrix, use the covariance matrix to dynamically adjust the perturbation direction, and perform non-Gaussian random sampling based on the step size control factor.

[0058] In this embodiment, a dynamic covariance matrix is introduced to accurately guide the perturbation sampling direction. The sampling process combines the sampling strategy guided by the covariance matrix to dynamically adjust the perturbation direction. The sampling process follows the following formula:

[0059] Z k ~m + τN(0, C)

[0060] where Z k is the perturbation amount, m is the offset mean to ensure that the sampling distribution does not deviate from the original feature center; τ is the scaling coefficient used to adjust the perturbation intensity; C is the covariance matrix, and N is the normal distribution sampling.

[0061] S105: Perform finite difference gradient estimation in the forward propagation and update the gradient estimation value in combination with the perturbation reduction factor.

[0062] In this embodiment, the Randomized Gradient Estimation (RGE) strategy is used to calculate the gradient through finite difference approximation. Using the forward pass of the model, the perturbation direction is dynamically calculated under the condition that the parameters remain unchanged, and this is used to guide the parameter update. This strategy eliminates the dependence on the gradient required for backpropagation and greatly reduces the memory usage.

[0063] When performing finite-difference gradient estimation, a stochastic gradient estimation strategy is utilized. The gradient direction for each iteration is calculated through the finite-difference formula. The stochastic gradient estimation adopts the following method:

[0064]

[0065] Method for calculating the gradient estimation value:

[0066]

[0067] where L is the loss function, W is the network parameter, Z i is the perturbation amount, q is the number of sampling times, σ is the perturbation reduction factor, and ∈ is the perturbation amplitude.

[0068] S107: Dynamically update the covariance matrix and adapter parameters.

[0069] 1. The covariance matrix is dynamically updated using the following update strategy:

[0070] 1) Update in real time according to the input data: Calculate the outer product matrix of the perturbation vectors in real time based on the input data of each batch;

[0071] 2) Incorporate historical covariance information using a forgetting factor;

[0072] 3) Perform positive semi-definite correction on the covariance matrix to control the stability of the covariance matrix data.

[0073] 2. To prevent over-updating of the parameters from causing feature collapse, the present invention proposes a perturbation reduction factor to dynamically adjust the update step size, ensure that the update amplitude is controlled, and improve the training stability.

[0074] The adapter parameters are dynamically updated using the perturbation reduction factor mechanism. The specific update process includes:

[0075] 1) In the initial stage, the perturbation reduction factor prevents excessive perturbation from affecting the model stability by adjusting the perturbation amplitude;

[0076] 2) During the working stage, when the update amplitude meets the preset threshold of the perturbation reduction factor, the update magnification remains unchanged; when the amplitude deviates from the set range, the perturbation reduction factor dynamically adjusts the perturbation amplitude to enhance the stability of the entire adaptation process;

[0077] 3) The perturbation reduction factor is defined as: σ = ∈ 2 , where σ is the perturbation reduction factor and ∈ is the perturbation amplitude.

[0078] S109: When a feature distribution shift is detected, optimize the model through covariance matrix-guided perturbation.

[0079] In addition, in this embodiment, by adopting zero-order optimization design, the high memory occupancy of backpropagation is avoided, and the resource requirements during the adaptation process are significantly reduced. The present invention realizes memory optimization through the following mechanism:

[0080] 1) Discard the computational graph cache related to backpropagation and only store the forward activation calculation results;

[0081] 2) Store the perturbation vector in fixed-point format and reduce the storage requirement of the perturbation vector through floating-point optimization;

[0082] 3) Dynamically release the cache of invalid intermediate feature maps.

[0083] The present invention constructs a dynamic covariance matrix to guide the perturbation direction, combines the perturbation attenuation factor mechanism and the adapter architecture to realize the optimization of model parameters without backpropagation. By analyzing the model feature distribution in real time through data, when a distribution shift is detected, the covariance matrix-driven finite difference gradient estimation is used for parameter update, and at the same time, the covariance matrix is updated by exponential moving average to maintain the optimization stability. The adapter layer of the present invention adopts the Kaiming-zero initialization strategy, and a lightweight adaptation structure is inserted in parallel in the MLP module of the ViT network to ensure the freezing of the core parameters of the model.

[0084] Compared with the prior art, the present invention has the following beneficial effects:

[0085] 1. The present invention accelerates the convergence speed through the covariance matrix-driven perturbation strategy and realizes higher optimization accuracy than traditional BP-Free methods.

[0086] 2. The present invention adopts zero-order optimization design, avoids the high memory occupancy of backpropagation, significantly reduces the resource requirements during the adaptation process, and has lower memory overhead.

[0087] 3. The present invention alleviates the feature collapse problem and improves the stability of the update process through the PRF mechanism and local update of the adapter layer, and the training stability is improved.

[0088] 4. The present invention takes the covariance matrix as the core and realizes efficient test-time adaptation in complex environments through an innovative zero-order optimization method, with strong integration and adaptability. The present invention can be efficiently deployed in edge computing devices, can adapt to dynamic domain distribution shift scenarios, has strong scalability, is particularly suitable for the real-time optimization requirements of edge computing devices, and has broad application prospects in the fields of intelligent manufacturing, medical image analysis, unmanned driving, etc.

[0089] The present invention will be described in detail below in conjunction with the preferred embodiments of the present invention.

[0090] Such as Figure 2As shown, a preferred embodiment of the present invention provides a test-time adaptation (TTA) method based on covariance matrix zero-order optimization (COZO), aiming to improve the test-time performance of the model under distribution shift and image corruption conditions, and at the same time making the application of zero-order optimization technology in this field a reality. This method combines random gradient estimation (RGE), covariance matrix-guided direction optimization, and adapter lightweight design, thereby realizing an efficient, stable, and resource-friendly model adaptation process, which is particularly suitable for resource-constrained environments such as edge computing.

[0091] The test-time adaptation method based on covariance matrix zero-order optimization provided by this embodiment includes the following steps:

[0092] Step 1: Random gradient estimation initialization

[0093] During testing, this embodiment first introduces a random gradient estimation (RGE) strategy, calculates the gradient direction of each iteration through a finite difference formula, and avoids the memory overhead required for backpropagation.

[0094] The random gradient estimation formula is:

[0095]

[0096] Where Z i is a perturbation amount randomly sampled from the standard normal distribution, representing the perturbation amount of each perturbation direction, L is the loss function, W is the network parameter, q is the number of samplings, σ is the perturbation reduction factor, ∈ is the perturbation amplitude, which is used to ensure the stability of gradient estimation and prevent numerical explosion. The incremental values of each sampling are accumulated and averaged to calculate the final gradient.

[0097] To reduce the variance of gradient estimation and improve its accuracy, this embodiment accumulates the incremental values through multiple samplings (q) and takes the average to form the final gradient estimation value.

[0098] The calculation method of the gradient estimation value:

[0099]

[0100] Where L is the loss function, W is the network parameter, Z i is the perturbation amount, q is the number of samplings, σ is the perturbation reduction factor, and ∈ is the perturbation amplitude.

[0101] Compared with the traditional backpropagation method, this step does not require explicit storage of gradients and activation values, significantly reducing memory usage.

[0102] Step 2: Covariance matrix-guided direction optimization

[0103] To overcome the problems of poor directionality and slow convergence speed in traditional random perturbation updates, this embodiment combines a sampling strategy guided by the covariance matrix to improve the efficiency of the optimization process by dynamically adjusting the perturbation direction.

[0104] The sampling process follows the following formula:

[0105] Z k ~m + τN(0, C)

[0106] where Z k is the perturbation amount, m is the offset mean, set to the zero vector to ensure that the sampling distribution does not deviate from the original feature center, τ is the scaling coefficient used to adjust the perturbation intensity, C is the covariance matrix, and N is the normal distribution sampling to ensure that the sampling distribution does not deviate from the original feature center.

[0107] The covariance matrix is continuously updated with the distribution difference of each test input, so the sampling direction is more in line with the current optimization needs. This design avoids the defect of excessive direction deviation in the traditional zero-order optimization random sampling process, thus achieving faster parameter update and convergence.

[0108] Step 3: Perturbation Decay Factor (PRF) mechanism

[0109] For the update of adapter parameters, a perturbation decay factor (PRF) is further introduced.

[0110] In the initial stage, PRF adjusts the perturbation amplitude to prevent excessive perturbation from causing model instability. In practical applications, when the update amplitude meets the preset threshold of PRF, the update magnification remains unchanged; when the amplitude deviates from the set range, PRF dynamically adjusts the perturbation amplitude to enhance the stability of the entire adaptation process.

[0111] The PRF formula is defined as follows:

[0112] σ = ∈ 2

[0113] PRF adaptively adjusts the amplitude limit of each perturbation to update the intensity, which can not only prevent model instability caused by excessive perturbation, but also maintain the continuity of the feature hierarchy and avoid the model falling into the state of overfitting or catastrophic forgetting.

[0114] Step 4: Network structure adjustment

[0115] To further improve the model adaptation efficiency and reduce the interference to the backbone network parameters, this embodiment inserts a lightweight adapter layer into the 3rd layer MLP module of the ViT - B / 16 (Vision Transformer) model.

[0116] 1. Adapter Design.

[0117] The adapter consists of two fully connected networks, which perform downsampling and upsampling operations with a compression ratio of 384:768. It can not only adjust features at a fine-grained level but also significantly reduce parameter overhead.

[0118] 1) Downsampling Matrix: Use Kaiming initialization to ensure stable feature distribution in the initial state;

[0119] 2) Upsampling Matrix: Use zero initialization to avoid introducing noisy features in the initial stage.

[0120] Through parallel design, the adapter layer is embedded in the MLP module in an insertable manner, rather than replacing the core functions of the main model.

[0121] Step Five: Adapter Parameter Freezing and Local Update

[0122] 1. Dynamic Update of Covariance Matrix

[0123] During the inflow of test data, regularly update the covariance matrix according to the distribution deviation of the adapter output features: Use the Exponential Moving Average (EMA) method to control the matrix update rate and avoid the adverse effects of short-term fluctuations on the domain adaptation effect.

[0124] The Exponential Moving Average method is a commonly used technical analysis tool. Different from the Simple Moving Average (SMA), EMA assigns higher weights to the most recent data points, so it can react faster to data changes. In this example, the EMA method is used to control the matrix update rate and avoid the adverse effects of short-term fluctuations on the domain adaptation effect.

[0125] 2. Distribution Difference Detection and Weight Adjustment

[0126] Use the KL (Kullback-Leibler, KL) divergence in the feature space to analyze the dynamic distribution difference between the training domain and the test domain. When the gap exceeds the set threshold, accelerate the adjustment frequency of the covariance matrix to avoid the adaptation process lagging behind the actual task requirements.

[0127] In addition, freeze all core parameters of the backbone network (such as weights and biases in the Transformer layer), and only update a small number of parameters (<1.2M) in the adapter layer. This not only reduces the memory requirements of large-scale models but also reduces the potential interference with the performance of the original network during training, thus achieving efficient and stable domain adaptation.

[0128] In another embodiment of the present invention, a multi-layer adapter design based on covariance dynamic adjustment is provided. By optimizing the adapter structure and covariance matrix update strategy, an enhanced design is provided for further improving the performance of test-time adaptation (TTA), enabling the gradient of the zero-order optimization estimate to be closer to the true gradient. Different from the single-layer adapter structure, this embodiment explores the insertion strategy of multi-layer adapters and combines the dynamically adjusted covariance matrix strategy, enabling the model to have stronger coping capabilities when dealing with complex data distribution shifts.

[0129] 1. Multi-layer adapter design

[0130] 1) Selection of adapter insertion position

[0131] The adapter module is inserted into the 1st layer, 3rd layer, and 6th layer of the ViT-B / 16 model, and each adapter module is used for cross-layer correction from middle and low-level features to high-level features.

[0132] Through extensive experiments, it is verified that these positions can effectively capture the different influence ranges of distribution differences on low-level features (such as texture) and high-level features (such as semantics).

[0133] 2) Optimization of adapter structure parameters

[0134] Setting of downsampling ratio: The adapter layer adopts a compression parameter ratio of 256:768 to adapt to a more flexible low-resource environment. Compared with the standard setting of 384:768, the 256 ratio is more lightweight, and the performance loss is only about 1.2%.

[0135] Initialization of the upsampling matrix: Zero initialization is used to ensure that the initial state of the adapter has no impact on the core feature parameters of the model, but it can flexibly adapt to the semantic alignment requirements of multi-layer features in the future.

[0136] 2. Covariance matrix dynamic sharing mechanism

[0137] 1) Joint update of cross-layer covariance distribution

[0138] A shared covariance matrix Cshared is used to be responsible for transmitting information between the update directions associated with different-level adapter modules, ensuring the consistency of the parameter optimization direction of the model throughout the entire depth range.

[0139] The update sampling of the covariance matrix still follows the formula:

[0140] Z k ~m + τN(0, C shared ),

[0141] where the result of each sampling can be dynamically fed back to all adapter modules, reducing the computational redundancy and the inconsistency of the distribution encoder caused by independent matrix updates.

[0142] 2) EMA-based local adjustment

[0143] Each layer of the adapter can still be multiplied by a dedicated scaling factor α local , expanding the adaptability of the covariance matrix to local feature differences:

[0144] C shared,new =(1 - β)C shared,old +β·C local

[0145] where β is a dynamic adjustment factor, which is particularly effective in identifying global and local distribution differences.

[0146] In another embodiment of the present invention, a dynamic intervention adaptation architecture based on scene detection is provided.

[0147] This embodiment proposes an improved solution that combines complex scene detection and a lightweight dynamic adaptation architecture, enhancing the applicability of test-time adaptation (TTA) in a multi-scenario deployment environment. By dynamically adjusting the optimization strategy through the scene detection module, the deployment requirements of multi-distribution data are met.

[0148] 1. Scene detection module design

[0149] 1) Detection target

[0150] Based on the KL divergence, calculate the difference between different input data distributions and the original training distribution to determine the type of the current scene (such as common scenes like noise, blur, weather, etc.).

[0151] 2) Design and implementation

[0152] Extract the depth features of the current input and match them with a preset scene template. If a significant deviation is matched, the scene detection module will output the optimal adjustment strategy.

[0153] 2. Dynamic optimization strategy switching

[0154] Different scenes control the update intensity of the perturbation direction. For example, in a noise scene, the weight update step size of the adapter layer is amplified to α×1.3 to strengthen the fuzzy compensation in feature reconstruction; in a weather scene, the PRF step size is adjusted to σ = ∈ 3 , compressing the model fluctuations that may be caused by excessive adjustment.

[0155] 3. Experimental comparison and scene tuning

[0156] 1) Scene tuning experiment

[0157] Noise scene: The dynamic strategy improves by 68.2% (compared to 67.1% of the static one).

[0158] Fuzzy scenario: The dynamic policy is improved to 67.9%.

[0159] Comprehensive scenario: The average improvement is 0.8 percentage points.

[0160] 2) Consumption comparison

[0161] Dynamic switching incurs an additional memory or inference time overhead of approximately 5%, but the overall performance is effectively enhanced.

[0162] Compared with the prior art, a test-time adaptation method based on zero-order optimization of the covariance matrix provided by the present invention aims to overcome the high memory requirements and computational overhead problems of traditional test-time adaptation methods in resource-constrained environments. By combining stochastic gradient estimation for zero-order optimization and using a covariance-driven perturbation strategy and a perturbation decay factor to guide parameter updates, the memory usage is significantly reduced, while efficient and stable model adaptation is achieved. The present invention further introduces an adapter layer that only adjusts a small number of parameters to support real-time domain adaptation, achieving excellent performance improvement while keeping the core weights of the pre-trained model unchanged.

[0163] Experimental verification shows that the COZO method demonstrates a leading advantage in various test scenarios. It not only achieves the state-of-the-art performance (67.1% accuracy) on the ImageNet-C dataset but also reduces the memory footprint to 1370MB, far lower than the over 5000MB of the traditional BP method, as Figure 3 shown. The present invention can be widely applied to various scenarios including edge computing, medical image analysis, and industrial quality inspection, providing an efficient and lightweight solution for the online adaptation of dynamically distributed data.

[0164] The adapter layer of the present invention adopts the Kaiming-zero initialization strategy and inserts a lightweight adaptation structure (with a parameter quantity <1.2M) in parallel in the MLP module of the ViT network to ensure the freezing of the core parameters of the model. Through experimental verification, this method achieves 67.1% accuracy in the test-time adaptation task of the ImageNet-C Level5 test set, with the memory consumption being only 26.5% of the traditional method, and the expected calibration error (ECE) being reduced to 4.1%. It has the characteristics of high accuracy, low memory, and strong stability. The present invention can be widely applied to edge computing scenarios such as industrial quality inspection and medical image analysis, effectively improving the real-time adaptation ability of intelligent devices in dynamic environments and having significant industrial application value.

[0165] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field according to the concept of the present invention through logical analysis, reasoning or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A test-time adaptation method based on zero-order optimization of the covariance matrix, characterized in that, The method includes the following steps: S101: Insert a lightweight adapter into the backbone network of the pre-trained model, and initialize the sampling matrix of the adapter; S103: Dynamically construct a covariance matrix, use the covariance matrix to dynamically adjust the perturbation direction, and perform non-Gaussian random sampling based on a step size control factor; S105: Perform finite difference gradient estimation in forward propagation, and update the gradient estimation value in combination with a perturbation reduction factor; S107: Dynamically update the covariance matrix and the adapter parameters; S109: When a feature distribution shift is detected, optimize the model through the covariance matrix-guided perturbation.

2. The method according to claim 1, characterized in that In step S101, the adapter consists of two fully connected networks, the adapter is inserted in parallel into the MLP module in the backbone network of the pre-trained model, the adapter adopts a single-layer or multi-layer configuration, and the downsampling matrix of the adapter is initialized using the Kaiming initialization method, and the upsampling matrix of the adapter is initialized to zero.

3. The method according to claim 2, wherein When the adapter adopts a single-layer configuration, the adapter adopts a compression ratio of 384:768, and the adapter is inserted into the 3rd layer MLP module of the backbone network of the pre-trained model.

4. The method according to claim 2, wherein When the adapter adopts a multi-layer configuration, the adapter adopts a compression ratio of 256:768, and the adapter is inserted into different layer MLP modules of the backbone network of the pre-trained model.

5. The method according to claim 4, characterized in that The adapter adopts a 3-layer configuration, and the adapter is inserted into the 1st layer, 3rd layer and 6th layer MLP modules of the backbone network of the pre-trained model.

6. The method according to claim 1, wherein In step S103, the sampling process combines the sampling strategy guided by the covariance matrix to dynamically adjust the perturbation direction, and the sampling process follows the following formula: Z k ~m + τN(0, C) Among them, Z k is the disturbance quantity, m is the offset mean to ensure that the sampling distribution does not deviate from the original feature center; τ is the scaling coefficient used to adjust the disturbance intensity; C is the covariance matrix, and N is the normal distribution sampling.

7. The method according to claim 1, wherein In step S105, when performing finite difference gradient estimation, use the stochastic gradient estimation strategy, calculate the gradient direction of each iteration through the finite difference formula, and the stochastic gradient estimation adopts the following method: The calculation method of the gradient estimation value: where L is the loss function, W is the network parameter, Z i is the perturbation, q is the number of sampling times, σ is the perturbation reduction factor, and ∈ is the perturbation amplitude.

8. The method according to claim 1, characterized in that, In step S107, the covariance matrix is dynamically updated using the following update strategy: Real-time update according to input data: Calculate the outer product matrix of the perturbation vector in real time according to the input data of each batch; Adopt a forgetting factor to fuse historical covariance information; Perform semi-positive definite correction on the covariance matrix to control the stability of the covariance matrix data.

9. The method according to claim 8, characterized in that, In step S107, the adapter parameters are dynamically updated using a perturbation reduction factor mechanism, including: In the initial stage, the perturbation reduction factor prevents the perturbation from being too large to cause model instability by adjusting the perturbation amplitude; In the working stage, when the update amplitude meets the preset threshold of the perturbation reduction factor, the update magnification remains unchanged; when the amplitude deviates from the set range, the perturbation reduction factor dynamically adjusts the perturbation amplitude to enhance the stability of the entire adaptation process; The disturbance reduction factor is defined as: σ = ∈ 2 , where σ is the disturbance reduction factor and ∈ is the disturbance amplitude.

10. The method according to any one of claims 1-9, characterized in that, The method adopts the following mechanism to achieve memory optimization: Discard the computational graph cache related to backpropagation and only store the forward activation calculation results; Store the perturbation vector in a fixed-point number format, and reduce the storage requirement of the perturbation vector through floating-point optimization; Dynamically release the cache of invalid intermediate feature maps.