Multi-level attention adaptive target tracking method based on motion estimation guidance
By introducing a multi-level attention adaptation mechanism based on motion estimation guidance in the target tracking method, combining the lightweight adaptive motion estimation module, significant hard attention sampling module and adaptive ViT attention head adjustment module, the problem of difficult to achieve high-precision tracking in complex scenarios in the prior art is solved, and efficient and accurate target tracking is achieved.
Patent Information
- Application Number
- CN202510098020.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing target tracking methods are difficult to achieve high-precision tracking in complex scenarios such as fast motion, occlusion, and deformation, and the waste of computing resources and unreasonable allocation of attention.
A multi-level attention adaptive target tracking method based on motion estimation guidance is adopted, combining a lightweight adaptive motion estimation module, a significant hard attention sampling module and an adaptive ViT attention head adjustment module to achieve accurate tracking of fast motion targets.
While maintaining high tracking accuracy, it significantly reduces computing overhead, improves the performance of target tracking algorithms, and achieves the unity of high precision and high efficiency.
Smart Images

Figure CN120047486A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology. Through innovative technologies such as lightweight adaptive motion estimation, saliency-based hard attention sampling, and adaptive ViT attention head adjustment, a multi-level attention adaptive object tracking method guided by motion estimation is specially designed to achieve accurate tracking of objects in complex scenarios such as fast motion, occlusion, and deformation. While maintaining high tracking accuracy, this method significantly reduces the computational overhead and has important theoretical research value and broad practical application prospects. Background Art
[0002] With the development of computer vision technology, object tracking has been widely applied in fields such as video surveillance, human-computer interaction, autonomous driving, and robot navigation. However, in actual scenarios, due to the influence of complex factors such as fast object motion speed, occlusion, deformation, illumination change, and background interference, existing object tracking methods still face the problem of tracking failure caused by fast object motion speed. When the object moves at a relatively fast speed, traditional tracking algorithms are difficult to accurately estimate the object position, which is mainly manifested in aspects such as the search area being unable to cover the actual position of the object resulting in the loss of the object, motion blur causing difficulties in feature extraction, fixed sampling step size leading to a decrease in positioning accuracy, and rapid changes in the object pose causing the appearance model to fail.
[0003] In addition, existing methods also have problems of waste of computing resources and unreasonable attention allocation. In terms of computing resources, existing methods often perform feature extraction and matching on the entire search area, ignoring the local characteristics of object motion, performing redundant calculations on non-object areas, having redundant feature extraction network structures resulting in low computing efficiency, and adopting a fixed computing resource allocation strategy without considering the scene complexity, causing waste of resources due to feature calculations of a large number of repetitive areas. In terms of the attention mechanism, traditional methods treat all areas equally and fail to highlight key areas, the number of attention heads is fixed and cannot be dynamically adjusted according to the scene complexity, lack of guidance from motion information leads to the disconnection between attention allocation and the object motion state, the multi-head attention mechanism has redundant parameter calculations and large computational overhead, and the attention weight calculation method is simple and fails to fully utilize context information. Summary of the Invention
[0004] To solve the above problems, the present invention proposes a multi-level attention adaptive object tracking method guided by motion estimation. This method realizes accurate tracking of fast-moving objects through the organic combination of a lightweight adaptive motion estimation module, a saliency-based hard attention sampling module, and an adaptive ViT attention head adjustment module.
[0005] To achieve the above object, the technical solution adopted by the present invention is: a multi-level attention adaptive object tracking method guided by motion estimation, including the following steps: S1: Construct a target tracking framework that includes a lightweight adaptive motion estimation module, a saliency-based hard attention sampling module, and an Adaptive ViT Attention Head Adjustment (AVAHA) module; S2: Pre-train the lightweight adaptive motion estimation module, and the pre-training method is as follows: S2.1: Construct a three-layer CNN network structure, including: the first layer is a 3×3 convolutional layer with a stride of 2, a padding of 1, an input channel number of 3, and an output channel number of 64; the second layer is a 1×1 convolutional layer with a stride of 1, an input channel number of 64, and an output channel number of 32; the third layer is a 3×3 convolutional layer with a stride of 1, a padding of 1, an input channel number of 32, and an output channel number of 1; each layer is followed by a BatchNorm layer and a ReLU activation function; S2.2: Adopt a multi-scale feature fusion strategy to perform channel attention weighted fusion after aligning the three-layer feature maps through upsampling; S2.3: Construct a loss function: , where is the motion estimation loss, is the regularization loss, is the smoothing loss; S3: Configure the saliency hard attention sampling module, including: S3.1: Design a saliency calculation network, with the input being the feature map , and the output being the saliency map ; S3.2: Adopt an adaptive threshold strategy for binarization: , where is the global mean of the saliency map , and the calculation formula is , is the global standard deviation of the saliency map , and the calculation formula is: , is an adjustable parameter used to control the strictness of the threshold. The larger the value, the fewer regions are retained after binarization, and it focuses more on high-saliency regions. The smaller the value, the more regions are retained after binarization, and more potential target regions can be captured. It is recommended that the value range is [0.5, 2.0]. The binarization operation converts the saliency map into a binary mask : .
[0006] S3.3: Design a region growing algorithm to merge or delete regions with a connected region area smaller than the threshold; S4: Configure the parameters of the adaptive ViT attention head adjustment module, including: S4.1: Set the value range of the number of attention heads ; S4.2: Define the attention head allocation strategy: , where: . is the number of basic attention heads, with a default value of 4, dynamically determined by the motion complexity , and the calculation formula is: , where is the scaling coefficient, with a default value of 8, is the complexity benchmark threshold, is the scaling factor, The function limits the value to the range [0, 8] to ensure that the total number of heads is in the range [4, 12]. The sigmoid function is used to smoothly map the complexity to the interval [0, 1].
[0007] S4.3: Construct the attention calculation formula: , where , are the query, key-value, and value matrices respectively, is the parameter matrix of the th attention head, is the output projection matrix, represents the concatenation operation on the last dimension, is the number of attention heads currently in use, dynamically determined by the strategy in S4.2.
[0008] S5: Configure the tracker parameters, including: S5.1: Set the size of the search area, defaulting to 4 times the target box; S5.2: Set the online update strategy: Update the template when the tracking confidence is greater than the threshold; The update weight α controls the fusion ratio of the new and old templates; Set the maximum number of templates, and delete the oldest template when it is exceeded.
[0009] S5.3: Set the failure recovery mechanism: Trigger a global search when the tracking confidence is lower than the threshold; Use the motion model to predict the possible target position; Combine the saliency map for candidate region screening.
[0010] S6: Perform online tracking, including: S6.1: Extract features according to the sampling steps determined by the lightweight adaptive motion estimation module; S6.2: Use the saliency mask generated by the saliency hard attention sampling module for regional attention; S6.3: Adaptively allocate the number of attention heads using the adaptive ViT attention head adjustment module; S6.4: Fuse multi - feature representations for target localization; S6.5: Update the target state and the online model.
[0011] Based on the above - mentioned technical solutions, the present invention further includes the following optimization strategies: Adopt a cross - layer feature fusion mechanism to adaptively fuse shallow - layer spatial detail features and deep - layer semantic features, introduce a channel attention mechanism to weight the importance of different - level features, and use residual connections to ensure information flow and avoid gradient vanishing; Combine Kalman filtering to smoothly predict the target state, introduce a motion consistency constraint to suppress the severe jitter of the tracking trajectory, and design an adaptive update strategy to dynamically adjust the model update rate according to the tracking confidence; Adopt the curriculum learning strategy to gradually increase the difficulty of training samples, introduce a contrastive learning loss to enhance the discriminability of feature representations, and use multi - scale training to enhance the scale invariance of the model.
[0012] Compared with the prior art, the advantages of the present invention are as follows: Firstly, the lightweight adaptive motion estimation module uses a lightweight CNN network to model the target motion pattern, which has the following advantages compared with traditional trackers: 1) Through a three - layer compact CNN structure, the number of parameters is only 1 / 10 of that of traditional motion estimation networks, significantly reducing the computational overhead; 2) Innovatively combine the joint constraints of motion amplitude and spatial gradient, making the motion prediction more robust, and the prediction accuracy is improved by more than 30% when the target moves rapidly; 3) Adopt an adaptive weight update strategy, which can dynamically adjust the model parameters according to the target motion state, avoiding the model drift problem caused by the fixed update rate of traditional methods.
[0013] Secondly, the saliency hard attention sampling module innovatively introduces a binary hard attention mechanism, which has significant advantages compared with traditional methods: 1) Through a saliency - driven sparse computing strategy, on average, only 20 - 30% of the search area needs to be used for feature extraction, and the computational efficiency is increased by 3 - 4 times; 2) Adopt a fast - normalized binary threshold calculation method, which only requires O(1) computational complexity, far lower than the O(n²) complexity of traditional soft attention; 3) Combine a hybrid mechanism of spatial and channel attention, and the feature expression ability is improved by 25% compared with single - attention, while keeping the computational overhead basically unchanged.
[0014] Third, the adaptive ViT attention head adjustment module realizes the dynamic adaptive adjustment of the number of attention heads, which has the following advantages compared with the traditional fixed-head scheme: 1) The dynamic allocation strategy based on motion complexity improves the utilization rate of attention resources by 40% and reduces the computational amount by 50% in simple scenarios; 2) The smooth mapping function based on sigmoid is adopted to avoid performance fluctuations caused by sudden changes in the number of heads, and the stability is improved by 15% compared with linear mapping; 3) The multi-scale feature fusion mechanism is innovatively introduced, and the feature expression ability is improved by 20% while only increasing the number of parameters by 5%.
[0015] The organic combination of these innovative technologies enables the present invention to achieve excellent results on mainstream benchmark datasets. Among them, the tracking success rate on the GOT-10k dataset reaches 78.5%, and the average precision reaches 81.1%. The performance improvement in fast-moving scenarios is particularly significant. At the same time, the method can reach a real-time processing speed of 32.1 FPS on RTX 3090, achieving the unity of high precision and high efficiency.
[0016] In summary, the present invention significantly improves the performance of the target tracking algorithm through the multi-level attention adaptive mechanism, realizes the real-time processing speed while maintaining high precision, and has important practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, wherein:
[0018] Figure 1 is the overall flowchart of the embodiment of the present invention; Figure 2 is the structural schematic diagram of the overall network model of the embodiment of the present invention; Figure 3 is the structural schematic diagram of the lightweight adaptive motion estimation module of the embodiment of the present invention; Figure 4 is the structural schematic diagram of the saliency hard attention sampling module of the embodiment of the present invention; Figure 5 is the structural schematic diagram of the adaptive ViT attention head adjustment module of the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.
[0020] The embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0021] As Figure 1 shown, the overall process of the multi-level attention adaptive target tracking method guided by motion estimation proposed by the present invention includes the following steps: First, input a continuous video frame sequence and an initial target box; Second, model and predict the target motion pattern through a lightweight adaptive motion estimation module; Then, the saliency hard attention sampling module performs saliency sampling on the search area; Finally, use the adaptive ViT attention head adjustment module for feature extraction and target localization.
[0022] As Figure 2 shown, the overall network model of the present invention mainly consists of three core modules: a lightweight adaptive motion estimation module, a saliency hard attention sampling module, and an adaptive ViT attention head adjustment module. These three modules are connected in a cascaded manner to form an end-to-end target tracking network structure.
[0023] As Figure 3 shown, the lightweight adaptive motion estimation module adopts a three-layer CNN network structure, including a 7×7 convolutional layer, a 1×1 convolutional layer, and a 3×3 convolutional layer. This module models the motion pattern of the target by analyzing the motion amplitude and spatial gradient information between consecutive frames. The module outputs three key parameters, namely motion amplitude M, motion speed V, and motion direction D, which will be used to guide the subsequent adjustment of the attention mechanism.
[0024] As Figure 4 shown, the core of the saliency hard attention sampling module is a binary hard attention mechanism. This module first calculates the saliency distribution of the feature map, and then performs binary processing based on a fast normalization threshold strategy to generate an attention mask. The regions with a value of 1 in the mask represent the regions that need to be focused on, and detailed feature extraction will be performed on these regions; while the regions with a value of 0 are skipped, thus achieving efficient utilization of computing resources.
[0025] As Figure 5As shown, the adaptive ViT attention head adjustment module adopts an improved Vision Transformer structure, and its greatest innovation lies in the dynamic adjustment mechanism of the number of attention heads. This module dynamically determines the required number of attention heads according to the motion complexity parameters output by the lightweight adaptive motion estimation module. When the target motion is simple, fewer attention heads are used to improve computational efficiency; when the target motion is complex, the number of attention heads is increased to enhance feature extraction ability.
[0026] In practical applications, the specific implementation process of the method of the present invention is as follows: 1) Input a continuous video frame sequence and the target box coordinates in the initial frame; 2) The lightweight adaptive motion estimation module analyzes the motion relationship between adjacent frames and outputs motion feature parameters; 3) The saliency hard attention sampling module generates a saliency map based on the current frame and performs binarization processing to obtain an attention mask; 4) The adaptive ViT attention head adjustment module dynamically adjusts the number of attention heads according to the motion complexity and extracts features from the region specified by the mask; 5) Based on the extracted features, target localization is performed to output the target position in the current frame.
[0027] In the implementation process of the method of the present invention, the parameters of each module can be appropriately adjusted according to the specific application scenario. For example, the CNN network structure in the lightweight adaptive motion estimation module can be appropriately simplified or expanded according to the computational resource limitations; the binarization threshold of the saliency hard attention sampling module can be dynamically adjusted according to the scene complexity; the range of the number of attention heads in the adaptive ViT attention head adjustment module can also be modified according to actual needs. This flexible design makes the method of the present invention have strong adaptability and practicality.
[0028] It should be noted that in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0029] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A multi-level attention adaptive target tracking method based on motion estimation guidance, characterized in that: The following steps are involved: S1: Build an object tracking framework consisting of a lightweight adaptive motion estimation module, a saliency-based hard attention sampling module, and an adaptive ViT attention head adjustment (AVAHA) module; S2: Pre-train the lightweight adaptive motion estimation module. The pre-training method is as follows: S2.1: Construct a three-layer CNN network structure, including: the first layer is a 3×3 convolution layer with a stride of 2, a padding of 1, 3 input channels, and 64 output channels; the second layer is a 1×1 convolution layer with a stride of 1, 64 input channels, and 32 output channels; the third layer is a 3×3 convolution layer with a stride of 1, a padding of 1, 32 input channels, and 1 output channel; each layer is followed by a BatchNorm layer and a ReLU activation function; S2.2: A multi-scale feature fusion strategy is adopted to align the three-layer feature maps through upsampling and then perform channel attention weighted fusion; S2.3: Construct loss function: ,in Estimated loss for motion, is the regularization loss, To smooth the loss; S3: Configure the saliency hard attention sampling module, including: S3.1: Design a saliency calculation network with feature maps as input , the output is a saliency map ; S3.2: Binarization using adaptive threshold strategy: ,in Saliency map The global mean of , Saliency map The global standard deviation is calculated as: , is an adjustable parameter used to control the strictness of the threshold. The larger the value, the fewer areas are retained after binarization, and the more focus is placed on high-saliency areas. The smaller the value, the more areas are retained after binarization, and more potential target areas can be captured. The value range of is [0.5, 2.0], and the binarization operation converts the saliency map Convert to binary mask : ; S3.3: Design a region growing algorithm to merge or delete connected regions whose area is smaller than a threshold; S4: Configure the adaptive ViT attention head adjustment module parameters, including: S4.1: Set the range of the number of attention heads ; S4.2: Define the attention head allocation strategy: ,in: . is the number of basic attention heads, the default value is 4, By motion complexity Dynamically determined, the calculation formula is: ,in is the scaling factor, the default value is 8, is the complexity benchmark threshold, is the scaling factor, The function limits the value to the range of [0,8] to ensure that the total number of heads is in the range of [4,12]. The sigmoid function is used to smoothly map the complexity to the interval of [0,1]. S4.3: Construct the attention calculation formula: ,in , are query, key and value matrices respectively, For the The parameter matrix of the attention head, is the output projection matrix, Indicates that the concatenation operation is performed on the last dimension. is the number of attention heads currently used, which is dynamically determined by the strategy in S4.2; S5: Configure tracker parameters, including: S5.1: Set the search area size, which is 4 times the target box by default; S5.2: Set the online update strategy: update the template when the tracking confidence is greater than the threshold; update the weight α to control the fusion ratio of the new and old templates; set the maximum number of templates, and delete the oldest template when it exceeds the limit; S5.3: Set up a failure recovery mechanism: trigger a global search when the tracking confidence is lower than the threshold; use the motion model to predict the possible target location; and use the saliency map to screen candidate regions; S6: Conduct online tracking, including: S6.1: performing feature extraction according to the number of sampling steps determined by the lightweight adaptive motion estimation module; S6.2: Use the saliency mask generated by the saliency hard attention sampling module to perform regional attention; S6.3: Adaptive allocation of the number of attention heads using the adaptive ViT attention head adjustment module; S6.4: Fusion of multiple feature representations for target localization; S6.5: Update target state and online model.
2. The multi-level attention adaptive target tracking method based on motion estimation guidance as claimed in claim 1, characterized in that: In the lightweight adaptive motion estimation module: Range of motion The calculation of is based on the following formula: ,in and Represent the feature maps of the current frame and the previous frame respectively, and denote the first-order and second-order spatial gradient operators, respectively, , , is the weight coefficient, and ; Sampling steps The formula is: ,in is the basic sampling step number, the default value is 1, is the scaling factor, the default value is 3, is the baseline threshold, is the scaling factor, The function clamps the values to the range [1,4].
3. The multi-level attention adaptive target tracking method based on motion estimation guidance according to claim 2 is characterized in that: In the saliency hard attention sampling module: Saliency Map The calculation of adopts multi-layer feature fusion, and the specific formula is as follows: ,in is the input feature map, and Respectively The weight matrix and bias vector of the layer, is the sigmoid activation function, which is used to map the output to the (0,1) interval. is the rectified linear unit activation function; Binarization uses an adaptive threshold mechanism: , ,in Saliency map The mean of is the standard deviation, is the adaptive coefficient, which is determined by the image contrast Dynamic Adjustment: , . is an indicative function; The regional screening criteria are defined as: ,in Indicates connected regions, is the area ratio threshold, usually 0.01-0.
05. are the height and width of the feature map, respectively. is the total number of connected areas.
4. The multi-level attention adaptive target tracking method based on motion estimation guidance according to claim 3 is characterized in that: In the adaptive ViT attention head adjustment module: Movement complexity The calculation formula is: ,in: is the motion amplitude, which indicates the displacement of the target between adjacent frames. is the target speed, indicating the target's movement rate, is the deformation degree, which indicates the deformation degree of the target appearance. is the occlusion degree, which indicates the degree to which the target is blocked by other objects. is the corresponding weight coefficient, and ; Number of attention heads Calculation: ,in: The default value is 6. is the scaling factor, which is used to adjust the influence of complexity on the number of heads. The function limits the value to the range [4,12] to ensure that the number of heads is within a reasonable range.
5. The multi-level attention adaptive target tracking method based on motion estimation guidance according to claim 4 is characterized in that: The feature fusion strategies of the three modules include: Feature map alignment: ,in For the The feature map of each module, is a bilinear interpolation function that scales the feature map to the target size. is the height and width of the target feature map; Channel attention weight calculation: ,in: It is a global average pooling operation, which compresses the feature map into dimensional vector, It is a two-layer perceptron network with the structure ,in is the dimensionality reduction ratio, Used to normalize weights, ensuring ; Final feature fusion: ,in: is the residual connection term, which is used to retain the original feature information. is the number of feature maps to be fused, For the Adaptive weights for feature maps.
6. The multi-level attention adaptive target tracking method based on motion estimation guidance according to claim 5, characterized in that: The training strategy of the lightweight adaptive motion estimation module includes: The loss function consists of: , , ,in: is the predicted motion amplitude, is the true range of motion, is the network weight parameter matrix, - is the regularization coefficient, which is used to balance various losses; Learning rate adjustment strategy: ,in: is the initial learning rate, set , is the decay exponent, set to 0.9, is the current training round number, is the total number of training rounds; Optimizer configuration: Use Optimizer, , , , .
7. The method for multi-level attention adaptive target tracking based on motion estimation guidance according to claim 6, characterized in that: The saliency calculation network structure of the saliency hard attention sampling module includes: Spatial Attention Module: ,in: for Convolutional layer, stride 1, padding 3, number of input channels , the number of output channels is 1, is the sigmoid activation function, which maps the output to interval, is the input feature map; Channel Attention Module: ,in: is the global average pooling operation, It is a two-layer perceptron network with the structure , the first layer uses the ReLU activation function, and the second layer uses the sigmoid activation function; Mixed Attention Features: ,in: represents element-wise multiplication (Hadamard product), For the final attention enhancement feature, is the spatial attention weight, is the channel attention weight.
8. The multi-level attention adaptive target tracking method based on motion estimation guidance according to claim 7, characterized in that: The attention calculation of the adaptive ViT attention head adjustment module includes: Query-key-value pair calculation: ,in: is the learnable transformation matrix, is the input feature sequence, is the sequence length, is the key-value pair dimension, is the model dimension; Attention score calculation: ,in: is the attention weight matrix, is the scaling factor used to prevent the gradient from disappearing; Multi-head attention output: ,in: For the The output of an attention head is is the output transformation matrix, is the number of attention heads, which is dynamically determined by the adaptive ViT attention head adjustment module.
9. The multi-level attention adaptive target tracking method based on motion estimation guidance according to claim 8, characterized in that: The target tracking method also includes the following components: Position Encoding Module: , ,in: is the position index, and its value range is , is the sequence length, is the dimension index, and its value range is , is the model dimension, usually an even number; Feedforward network module: ,in: is the first layer weight matrix, is the second layer weight matrix, is the first layer bias vector, is the second layer bias vector, is the hidden layer dimension of the feedforward network, usually , is the ReLU activation function; Residual connections and layer normalization: ,in: is a sub-layer network, which can be a multi-head attention layer or a feed-forward network layer. is the layer normalization operation, and the calculation formula is ,in and are the mean and standard deviation, and are learnable parameters.
10. The multi-level attention adaptive target tracking method based on motion estimation guidance according to claim 9, characterized in that: The training configuration of the target tracking framework includes: Data enhancement strategy: random horizontal flipping with probability , randomly flip vertically with probability , random rotation, angle range , random scaling, range , random brightness adjustment, range , the transformation formula is , random contrast adjustment, range , transformation formula: ,in is the image mean; Batch size setting: training batch , verification batch: , test batch ; Training schedule: total number of rounds , number of preheating rounds , learning rate warm-up factor , the learning rate decay strategy is cosine annealing, the formula is ,in is the current step number, is the total number of steps, and the early stopping strategy is the continuous performance of the validation set If the wheel is not lifted, it stops; Model saving strategy: Save the model with the best performance in the validation set ,Every Save a checkpoint once per round , save the last Checkpoints For model integration, the integration weight is ,in is the performance gap between each model and the optimal model, is the temperature coefficient.