An efficient parameter sharing method for transferring self-supervised pre-trained model to V-MoE

CN122655894BActive Publication Date: 2026-09-29QINGDAO UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611122840.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-29
Estimated Expiration
2046-07-28

AI Technical Summary

Technical Problem

[0003]然而,现有迁移方法仍存在不足

Benefits of technology

(1)本发明通过引入基于神经元激活选择性的参数保留与分组机制,能够根据预训练稠密视觉Transformer中前馈神经网络层各神经元的激活响应差异,将神经元划分为专有神经元和通用神经元,并进一步构建核心专家组和通用专家组。该机制在保留原始预训练权重有效表征能力的基础上,减少了无差别参数拆分造成的专家表征同质化和参数冗余问题,使核心专家能够更加集中地学习局部差异明显、对任务判断更关键的特征。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122655894B_ABST
    Figure CN122655894B_ABST
Patent Text Reader

Abstract

The application discloses an efficient parameter reuse method for transferring a self-supervised pre-training model to a V-MoE, and belongs to the technical field of deep learning and computer vision. An input feature block sequence is constructed based on a no-label image calibration set, a pool of to-be-transferred feedforward parameters in a dense model is extracted, and neuron activation selectivity is calculated; according to the activation selectivity, specific neurons and general neurons are divided, and a core expert group and a general expert group are respectively constructed; a labeled fine-tuning set is constructed, local feature complexity of the feature block is calculated, and low-complexity blocks and high-complexity blocks are shunted to different expert groups; a final feature representation is obtained based on expert output, and a downstream fine-tuning model is fine-tuned in combination with a task loss and a core expert load balancing loss. The application realizes differentiated reuse of pre-training parameters, reduces the risk of routing cold start, and realizes dynamic power allocation according to local feature complexity, thereby improving the inference efficiency and prediction performance of the model in downstream visual tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of deep learning and computer vision technology, and in particular to an efficient parameter reuse method for transferring self-supervised pre-trained models to V-MoE. Background Technology

[0002] In recent years, self-supervised learning-based visual Transformer models have been widely used in computer vision. However, as the model size continues to increase, dense visual Transformers typically face high computational and memory overhead when fine-tuning and deploying them for downstream tasks. To expand model capacity while controlling inference computation, transferring self-supervised pre-trained dense visual Transformers to sparse activation visual hybrid expert models, i.e., V-MoE, has become an important low-cost model reuse approach.

[0003] However, existing transfer methods still have shortcomings. First, the feedforward neural network layer of dense visual Transformers contains a large number of neurons with different functions. If the response characteristics of neurons are not distinguished during the transfer process, and the original parameters are directly copied or split, it is easy to cause similar representations among multiple experts, resulting in expert homogenization and parameter redundancy, making it difficult to fully utilize the specialized expressive power of the hybrid expert model. Second, the routers of existing V-MoE models usually lack effective pre-trained parameter priors and often need to start learning from random states. This can easily lead to unbalanced expert load, unstable routing, or even routing collapse in the early stages of training, thus affecting the stability of model fine-tuning. Finally, traditional hybrid expert routing mechanisms usually adopt a uniform routing method for all feature blocks, lacking differentiation of local feature complexity. This can easily generate unnecessary computational overhead for low-complexity feature blocks and expose core experts to background noise, making it difficult to achieve on-demand allocation of computing power. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention proposes a novel model transfer and parameter reuse method to achieve differentiated reuse of pre-trained parameters, reduce the risk of cold start in routing, and achieve dynamic computing power allocation based on local feature complexity, thereby improving the inference efficiency and prediction performance of the model in downstream vision tasks.

[0005] To achieve the above objectives, this invention provides an efficient parameter reuse method for transferring self-supervised pre-trained models to V-MoE, comprising: Step S1: Construct an input feature block sequence based on the unlabeled image calibration set, extract the feedforward parameter pool to be transferred in the dense model, and calculate the neuron activation selectivity; Obtaining unlabeled image calibration sets The unlabeled image calibration set The image is converted into a sequence of input feature blocks. ; The input feature block sequence Input a dense visual Transformer model that has been pre-trained under self-supervision Extract the dense visual Transformer model The original weight matrix of the feedforward neural network layer F is used to form the feedforward parameter pool to be transferred. ; The activation responses of each neuron in the feedforward neural network layer F to different feature block samples are statistically analyzed, and inactive neurons are removed based on the average activation. The activation selectivity index of each effective neuron is then calculated. ; Step S2: Based on the activation selectivity, divide the proprietary neurons and general neurons, construct the core expert group and the general expert group respectively, and initialize the router; Set activation selectivity threshold The activation selectivity index of each effective neuron With the activation selectivity threshold A comparison will satisfy The neuron is identified as a specialized neuron, and will satisfy the following... The neurons were identified as general neurons; Based on the weights of the first and second fully connected layers corresponding to the proprietary neurons, multiple core expert network sub-modules with consistent intermediate layer widths are constructed to form a core expert group. Initialize the router weight matrix ; Based on the preset requirements for the number and width of general experts, the weights of the first and second fully connected layers corresponding to the general neurons are balanced, grouped, and reassembled to construct multiple general expert network sub-modules with consistent intermediate layer widths, forming a general expert group. And by the core expert group General Expert Group and router Constructing the transferred V-MoE model The expert calculation part; Step S3: Construct a labeled fine-tuning set, calculate the local feature complexity of the feature blocks, and distribute low-complexity blocks and high-complexity blocks to different expert groups; From the labelless image calibration set Image samples are extracted and manually labeled to form a labeled fine-tuning set. The labeled fine-tuning set This includes training images and their corresponding ground truth labels P; The labeled fine-tuning set The training images are converted into a sequence of input feature blocks in the target domain. Calculate the target domain input feature block sequence Local feature complexity score of each feature block ; Set a complexity threshold The local feature complexity score of each feature block. With the complexity threshold A comparison will satisfy The feature blocks are determined to be low-complexity feature blocks, which will satisfy... The feature block is determined to be a high-complexity feature block; Low-complexity feature blocks are allocated to the general expert group according to a preset balanced allocation rule. A general expert network submodule is used to extract and maintain relatively stable basic visual information such as color, texture, and outline, while preserving their original location identifiers and grid order. The high-complexity feature block is input into the router R, and the router R performs Top-N routing selection based on the routing probability, assigning the high-complexity feature block to the top N core experts with the highest probability scores; The outputs of the first N core experts are weighted and fused according to their corresponding routing probabilities to obtain the output of a single core expert corresponding to a high-complexity feature block. The general expert group With core expert group The output feature blocks are recombined according to the original grid order and input into the subsequent Transformer encoding layer to obtain the final feature representation Y; Step S4: Obtain the final feature representation based on the expert output, and fine-tune the transfer model downstream by combining the task loss and the core expert load balancing loss. A prediction head H is constructed for the downstream visual task. The final feature representation Y is input into the prediction head H to obtain the prediction result. ; Based on the prediction results With the labeled fine-tuning set Calculate the task loss corresponding to the real label P. ; Computing Core Expert Group Calculate load balancing loss ; The task loss With core expert group Load balancing loss We perform a weighted summation to obtain the total loss L: ;

[0006] in, This is the preset loss balance coefficient; Based on the total loss L, the prediction head H, router R, and core expert group are analyzed. and V-MoE model Other trainable parameters involved in fine-tuning are backpropagated and updated, and the general expert group is also updated. The parameters used are lower than those of the core expert group. The learning rate is updated.

[0007] Optionally, in step S1, the unlabeled image calibration set is... The image is converted into a sequence of input feature blocks. ,include: Extract the unlabeled image calibration set The original image in For the original image Size normalization and tensor quantization are performed, and the images are processed according to a preset image block size. Perform non-overlapping mesh segmentation; The segmented local image patches are flattened and input into a linear projection layer, and positional encodings corresponding to the grid positions of the image patches are added to generate an input feature block sequence. .

[0008] Optionally, in step S1, the activation response of each neuron in the feedforward neural network layer F to different feature block samples is statistically analyzed, and inactive neurons are removed based on the average activation status, and the activation selectivity index of each effective neuron is calculated. ,include: For input feature block samples x Record the feedforward neural network layer F The Middle i Activation response amplitude of a single neuron And calculate the neuron in the unlabeled image calibration set. Activation mean and activation standard deviation ; Set active dead zone threshold , will satisfy The neurons were identified as inactive neurons and included in the feedforward parameter pool to be migrated. Simultaneously delete the corresponding first fully connected layer output weights and second fully connected layer input weights. For the remaining effective neurons, the first one is calculated according to the following formula. Activation selectivity index of neurons : ;

[0009] in, To prevent extremely small positive numbers with a denominator of zero.

[0010] Optionally, in step S2, multiple core expert network sub-modules with consistent intermediate layer widths are constructed based on the weights of the first fully connected layer and the weights of the second fully connected layer corresponding to the proprietary neurons, forming a core expert group. Initialize the router weight matrix ,include: Extract the weights of the first fully connected layer corresponding to the specialized neuron and construct a weight matrix. ,in, G This represents the total number of proprietary neurons. d Indicates the dimension of the hidden layer features; With the weight matrix Each weight vector in the dataset is used as a sample to be clustered. K-Means clustering is performed to obtain K feature clusters, and K cluster centers are calculated. ; Based on the clustering results, the weights of the first and second fully connected layers corresponding to the specialized neurons are sliced, capacity balanced, and reassembled to construct K core expert network sub-modules with consistent intermediate layer widths, forming a core expert group. ; The K cluster centers Arranged sequentially by expert number, forming a size of Router initial weight matrix The router's initial weight matrix Each column corresponds to an initial matching vector for a core expert.

[0011] Optionally, in step S3, the target domain input feature block sequence is calculated. Local feature complexity score of each feature block ,include: Input feature block sequence for the target domain The first in j For each feature block, before it enters the expert routing stage and before normalization is performed, the corresponding feature representation is extracted and processed along the feature dimension of the hidden layer. d Calculate the local feature complexity score : ;

[0012] in, Indicates the first jThe feature block in the th ... c Feature values ​​on each feature channel Indicates the first j Each feature block in the whole d The characteristic mean of each channel.

[0013] Optionally, in step S4, the core expert group is calculated. Calculate load balancing loss ,include: For the core expert group Based on the actual received frequencies of each core expert within a training batch and the average routing probability output by the router, the load balancing loss is calculated. : ;

[0014] in, Indicates the first k The normalized frequency at which core experts are actually selected This indicates that the router points to the first... k Average soft router probability of each core expert; When using the Top-N routing mechanism, the and Calculate according to the following formulas: ; ;

[0015] in, This represents the number of high-complexity feature blocks that participate in core expert routing within a training batch. Indicates the first z The first high-complexity feature block is routed to the second... k The probability of a core expert using a software router. Indicates an indicator function.

[0016] By adopting the above technical solution, the present invention has at least the following beneficial effects: (1) This invention introduces a parameter retention and grouping mechanism based on neuron activation selectivity. This mechanism can divide neurons into specialized neurons and general neurons according to the differences in activation responses of neurons in the feedforward neural network layer of a pre-trained dense visual Transformer, and further construct core expert groups and general expert groups. While preserving the effective representational capabilities of the original pre-trained weights, this mechanism reduces the homogenization of expert representations and parameter redundancy caused by indiscriminate parameter splitting, enabling core experts to more effectively learn features with significant local differences and more critical to task judgment.

[0017] (2) This invention uses the clustering results of the weights corresponding to the proprietary neurons to initialize the router weight matrix, replacing the completely random router initialization method. This allows the router to obtain prior information from the distribution of pre-trained parameters in the early stage of fine-tuning, thereby reducing the risk of cold start caused by random routing, alleviating the problem of unbalanced expert load and unstable routing, and improving the stability of the model fine-tuning process.

[0018] (3) This invention distributes feature blocks based on local feature complexity scores, assigning low-complexity feature blocks to a general expert group to preserve basic visual information such as color, texture, and contour, while assigning high-complexity feature blocks to a core expert group for focused modeling via a router. This reduces interference from low-complexity background features to core experts, lowers unnecessary expert routing and computational overhead, and enables on-demand allocation of computing resources.

[0019] (4) In the fine-tuning training process, the present invention introduces a load balancing loss that only applies to the core expert group, so that the high-complexity feature blocks are distributed relatively evenly among the core experts, while not imposing a forced balancing constraint on the general expert group. This balances the division of labor among experts, training stability and basic representation retention, which is beneficial to improving the inference efficiency and prediction performance of the V-MoE model after transfer in downstream visual tasks. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 A schematic diagram of the overall process of an efficient parameter reuse method for transferring a self-supervised pre-trained model to V-MoE provided by the present invention; Figure 2 A schematic diagram illustrating the process of extracting the feedforward parameter pool to be transferred from a dense visual Transformer model and calculating neuron activation selectivity; Figure 3 A flowchart illustrating the process of constructing a core expert group and a general expert group and initializing the router by selectively dividing proprietary neurons and general neurons based on activation; Figure 4 A flowchart illustrating the process of constructing labeled fine-tuning sets, calculating the local feature complexity of feature blocks, and routing low-complexity blocks and high-complexity blocks to different expert groups; Figure 5This is a flowchart illustrating the process of fine-tuning the transfer model downstream by combining task loss and core expert load balancing loss to obtain the final feature representation based on expert output. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] like Figure 1 As shown, this invention provides an efficient parameter reuse method for transferring self-supervised pre-trained models to V-MoE, comprising: Step S1: Construct an input feature block sequence based on the unlabeled image calibration set, extract the feedforward parameter pool to be transferred from the dense model, and calculate the neuron activation selectivity, such as... Figure 2 As shown.

[0024] Obtaining unlabeled image calibration sets Extract the unlabeled image calibration set The original image in For the original image Size normalization and tensor quantization are performed, and the images are processed according to a preset image block size. Non-overlapping mesh segmentation is performed. The resulting local image patches are flattened and input into a linear projection layer, where positional codes corresponding to the mesh positions of the image patches are added to generate an input feature block sequence. .

[0025] The input feature block sequence Input a dense visual Transformer model that has been pre-trained under self-supervision Extract the dense visual Transformer model The original weight matrix of the feedforward neural network layer F is used to form the feedforward parameter pool to be transferred. .

[0026] For input feature block samples x Record the feedforward neural network layer F The Middle i Activation response amplitude of a single neuron And calculate the neuron in the unlabeled image calibration set. Activation mean and activation standard deviation .

[0027] Set active dead zone threshold , will satisfy The neurons were identified as inactive neurons and included in the feedforward parameter pool to be migrated. Simultaneously delete the corresponding first fully connected layer output weights and second fully connected layer input weights.

[0028] For the remaining effective neurons, the first one is calculated according to the following formula. Activation selectivity index of neurons : ;

[0029] in, To prevent extremely small positive numbers with a denominator of zero.

[0030] Step S2: Based on activation selectivity, divide the neurons into specialized neurons and general neurons, construct the core expert group and the general expert group respectively, and initialize the router, as follows. Figure 3 As shown.

[0031] Set activation selectivity threshold The activation selectivity index of each effective neuron With the activation selectivity threshold A comparison will satisfy The neuron is identified as a specialized neuron, and will satisfy the following... The neurons were identified as general neurons.

[0032] Extract the weights of the first fully connected layer corresponding to the specialized neuron and construct a weight matrix. ,in, G This represents the total number of proprietary neurons. d This indicates the feature dimension of the hidden layer.

[0033] With the weight matrix Each weight vector in the dataset is used as a sample to be clustered. K-Means clustering is performed to obtain K feature clusters, and K cluster centers are calculated. .

[0034] Based on the clustering results, the weights of the first and second fully connected layers corresponding to the proprietary neurons are sliced, capacity balanced, and reassembled to construct K core expert network sub-modules with consistent intermediate layer widths, forming a core expert group. .

[0035] The K cluster centers Arranged sequentially by expert number, forming a size of Router initial weight matrix The router's initial weight matrix Each column corresponds to an initial matching vector for a core expert.

[0036] Based on the preset requirements for the number and width of general experts, the weights of the first and second fully connected layers corresponding to the general neurons are balanced, grouped, and reassembled to construct multiple general expert network sub-modules with consistent intermediate layer widths, forming a general expert group. And by the core expert group General Expert Group and router Constructing the transferred V-MoE model The expert calculation part.

[0037] Step S3: Construct a labeled fine-tuning set, calculate the local feature complexity of feature blocks, and distribute low-complexity blocks and high-complexity blocks to different expert groups, such as... Figure 4 As shown.

[0038] From the labelless image calibration set Image samples are extracted and manually labeled to form a labeled fine-tuning set. The labeled fine-tuning set This includes training images and their corresponding ground truth labels P.

[0039] The labeled fine-tuning set The training images are converted into a sequence of input feature blocks in the target domain. Input feature block sequence for the target domain The first in j For each feature block, before it enters the expert routing stage and before normalization is performed, the corresponding feature representation is extracted and processed along the feature dimension of the hidden layer. d Calculate the local feature complexity score : ;

[0040] in, Indicates the first j The feature block in the th ... c Feature values ​​on each feature channel Indicates the first j Each feature block in the whole d The characteristic mean of each channel.

[0041] Set a complexity threshold The local feature complexity score of each feature block. With the complexity threshold A comparison will satisfy The feature blocks are determined to be low-complexity feature blocks, which will satisfy... The feature block was determined to be a high-complexity feature block.

[0042] Low-complexity feature blocks are allocated to the general expert group according to a preset balanced allocation rule. A general expert network submodule is used to extract and maintain relatively stable basic visual information such as color, texture, and contour, while preserving their original location identifiers and grid order.

[0043] The high-complexity feature block is input into the router R, and the router R performs Top-N routing selection based on the routing probability, assigning the high-complexity feature block to the top N core experts with the highest probability scores.

[0044] The outputs of the first N core experts are weighted and fused according to their corresponding routing probabilities to obtain the output of a single core expert corresponding to a high-complexity feature block.

[0045] The general expert group With core expert group The output feature blocks are recombined according to the original grid order and input into the subsequent Transformer encoding layer to obtain the final feature representation Y.

[0046] Step S4: Obtain the final feature representation based on expert output, and fine-tune the transfer model downstream by combining task loss and core expert load balancing loss, such as... Figure 5 As shown.

[0047] A prediction head H is constructed for the downstream visual task. The final feature representation Y is input into the prediction head H to obtain the prediction result. .

[0048] Based on the prediction results With the labeled fine-tuning set Calculate the task loss corresponding to the real label P. In image classification tasks, the task loss is... This includes cross-entropy loss, a task loss mechanism in object detection or semantic segmentation tasks. Determined based on the loss function of the corresponding task header.

[0049] For the core expert group Based on the actual received frequencies of each core expert within a training batch and the average routing probability output by the router, the load balancing loss is calculated. : ;

[0050] in, Indicates the first k The normalized frequency at which core experts are actually selected This indicates that the router points to the first... k The average soft router probability of each core expert.

[0051] When using the Top-N routing mechanism, the and Calculate according to the following formulas: ; ;

[0052] in, This represents the number of high-complexity feature blocks that participate in core expert routing within a training batch. Indicates the first z The first high-complexity feature block is routed to the second... k The probability of a core expert using a software router. Indicates an indicator function.

[0053] The task loss With core expert group Load balancing loss We perform a weighted summation to obtain the total loss L: ;

[0054] in, This is the preset loss balance coefficient.

[0055] Based on the total loss L, the prediction head H, router R, and core expert group are analyzed. and V-MoE model Other trainable parameters involved in fine-tuning are backpropagated and updated, and the general expert group is also updated. The parameters used are lower than those of the core expert group. The learning rate is updated.

[0056] The load balancing loss Used only to constrain the core expert group Internal routing allocation is not applicable to the general expert group. Forced load balancing constraints are applied; when no high-complexity feature blocks participating in core expert routing exist within a training batch, the load balancing loss is not calculated. The load balancing loss of this batch Set to zero.

[0057] To further illustrate the technical solution of this invention, the following three specific application examples are provided: Case 1: A general natural image classification system based on the method of this invention (1) Construct a pool of feedforward parameters to be transferred and calculate neuron activation selectivity.

[0058] The system first obtains 30,000 natural images by uniformly and randomly sampling from the ImageNet-1K dataset, and then constructs an unlabeled image calibration set without using class labels. During the data preprocessing stage, the unlabeled image calibration set is... Each original image in the process is uniformly scaled and cropped to Image tensor of resolution; set image patch size to The system performs non-overlapping meshing on the image tensor, breaking each image down into 196 local image patches. These 196 patches are then flattened sequentially and input into a linear projection layer, mapping them to a hidden layer feature dimension of 768. Positional encoding corresponding to the mesh positions of the image patches is then added, resulting in a final image with dimensions of [missing information]. Input feature block sequence The system loads a dense visual Transformer model that has been pre-trained under self-supervised guidance. Locate the target feedforward neural network layer. Feedforward neural network layer The system employs a multilayer perceptron architecture with an expansion ratio of 4, containing 3072 intermediate layer neurons; the system extracts the feedforward neural network layer. The original pre-trained weight matrix forms the feedforward parameter pool to be transferred. Subsequently, the system will input the feature block sequence. Traversing the input dense visual Transformer model The total number is Feature block samples x For feedforward neural network layers For each neuron in the system, the system records the amplitude of its activation response to each feature block sample. And calculate the neuron in the unlabeled image calibration set. The activation mean and activation standard deviation are calculated based on a preset active dead zone threshold. The system detected and removed 72 inactive neurons, and added them to the feedforward parameter pool to be transferred. Simultaneously, the corresponding output weights of the first fully connected layer and input weights of the second fully connected layer are deleted to maintain the consistency of the dimension in subsequent matrix calculations; for the remaining 3000 effective neurons, the system calculates the... Activation selectivity index of neurons .

[0059] (2) Dividing neurons and constructing core expert groups, general expert groups, and routers includes: The system sets an activation selectivity threshold. The activation selectivity index of effective neurons With activation selectivity threshold Compare; the system will satisfy The 1536 neurons were divided into proprietary neurons, and their original pre-trained weights were used as the core expert group. The foundation for this construction is as follows: the remaining 1464 neurons are divided into general neurons, and their original pre-trained weights are used as the general expert group. The above-mentioned division method, while preserving the original pre-trained weights, allows high-selectivity neurons to undertake more feature modeling tasks such as target contours, local texture changes, and category discrimination-related regions, while allowing low-selectivity neurons to undertake more tasks of preserving basic visual information such as color, texture, and contour. For the 1536 dedicated neurons, the system extracts the corresponding first fully connected layer weight vectors, each with a dimension of 768, thus constructing a structure with a size of... The system sets the weight matrix; the number of core experts is 16, and K-Means clustering is performed on the weight vectors in the weight matrix to obtain 16 feature clusters, and the cluster center of each feature cluster is calculated; to ensure the consistency of the expert structure, the system performs capacity balancing on the number of neurons in each feature cluster, so that each core expert network submodule contains 96 intermediate layer neurons, thereby constructing 16 core expert network submodules with consistent intermediate layer widths, forming a core expert group. Simultaneously, the system arranges the 16 cluster centers sequentially according to the core expert numbers, forming a cluster with a size of [missing information]. Router initial weight matrix and assign it to the router. The linear weight layer allows the router to obtain prior information from the original pre-trained parameter distribution during the initial fine-tuning phase, reducing the cold start risk caused by random routing. For the 1464 general neurons, the system balances and reassembles the weights of the first and second fully connected layers corresponding to the general neurons according to the preset requirements for the number and width of general experts. The system constructs 12 general expert network sub-modules, each containing 122 intermediate layer neurons, thus forming a general expert group. At this point, the system has completed the core expert group. General Expert Group and router The construction of the V-MoE model and its composition as the transferred model. The expert calculation part.

[0060] (3) Construct a labeled fine-tuning set and perform feature block routing based on local feature complexity.

[0061] In general natural image classification tasks, the system is based on an unlabeled image calibration set. Images in the dataset can be retrieved or supplemented with category labels to form labeled fine-tuning sets. Labeled micro-tuning set Including training images and their corresponding ImageNet-1K class labels The system will fine-tune the tagged set. The training images are converted into a sequence of input feature blocks in the target domain in the same manner as in the calibration phase. Each image corresponds to 196 feature blocks, and the hidden layer feature dimension of each feature block is 768. During the fine-tuning stage, for the target domain input feature block sequence... The first in For each feature block, before the system enters the expert routing stage and performs normalization processing, it calculates the local feature complexity score along the hidden layer feature dimension. And the local feature complexity score With complexity threshold Compare; for those that satisfy Low-complexity feature blocks, such as large smooth backgrounds, weakly textured regions, or image patches with small edge variations, are assigned to the general expert group by the system. To preserve basic visual information such as color, texture, and outline, as well as their original mesh order; for those that meet the requirements Highly complex feature blocks, such as target contours, regions with significant local texture changes, or categories, are input into the router by the system. ;router Based on size weight matrix The system performs matching calculations on high-complexity feature blocks and outputs routing probabilities pointing to 16 core experts through a normalized exponential function. A Top-2 routing mechanism is used to assign each high-complexity feature block to the two core experts with the highest probability scores. For the outputs of the two core experts corresponding to the same high-complexity feature block, the system performs weighted fusion according to their routing probabilities to obtain the output of a single core expert corresponding to that feature block. Subsequently, the system groups a general expert group. Low-complexity feature block representation of the output and the core expert group The output high-complexity feature block representation is reassembled according to the grid order in the original image into a form of size. The feature sequences are then input into subsequent Transformer encoding layers. A global self-attention mechanism is used to re-establish the spatial and semantic relationships between feature blocks, resulting in the final feature representation. .

[0062] (4) Downstream visual task fine-tuning and core expert load balancing training.

[0063] The system will represent the final features. Classification prediction head with input and output dimensions of 1000 To obtain the prediction results The system is based on the prediction results. With real category labels Calculate the cross-entropy loss and use it as the task loss. Meanwhile, the system, based on the router Output of core expert routing probabilities and core expert groups Calculate the load balancing loss of each core expert based on their actual received frequency within the current training batch. Load balancing losses Used only to constrain the core expert group Internal routing allocation is not applicable to the general expert group. Forcibly imposing balance constraints; the system will suffer task losses. With load balancing losses We perform a weighted summation to obtain the total loss. Finally, the system is based on the total loss. For classification prediction head ,router Core Expert Group and V-MoE model Other trainable parameters involved in fine-tuning are updated via backpropagation; for the general expert group The system adopts a lower level than the core expert group. The learning rate is updated to reduce excessive drift of its underlying representation capabilities during training.

[0064] Case 2: Autonomous Driving Scene Object Detection and Semantic Segmentation System Based on the Method of the Invention (1) Construct a pool of feedforward parameters to be transferred and calculate neuron activation selectivity.

[0065] The system first randomly selects 10,000 images containing complex street scenes from the BDD100K autonomous driving public dataset, and constructs an unlabeled image calibration set without using bounding box annotations or semantic segmentation annotations. In the data preprocessing stage, to preserve local details of smaller targets such as pedestrians, traffic signs, and distant vehicles, the unlabeled image calibration set... Each original image in the process is uniformly scaled and cropped to Image tensor of resolution; set image patch size to The system performs non-overlapping meshing on the image tensor, breaking each image down into 1024 local image patches. These 1024 patches are then flattened sequentially and input into a linear projection layer, mapping them to a hidden layer feature dimension of 768. Positional encoding corresponding to the mesh positions of the image patches is then added, resulting in a final image with dimensions of [missing information]. Input feature block sequence The system loads a dense visual Transformer model that has been pre-trained under self-supervised guidance. Locate the target feedforward neural network layer. Feedforward neural network layer The system employs a multilayer perceptron architecture with an expansion ratio of 4, containing 3072 intermediate layer neurons; the system extracts the feedforward neural network layer. The original pre-trained weights form the feedforward parameter pool to be transferred. Subsequently, the system will input the feature block sequence. Traversing the input dense visual Transformer model The total number is Feature block samples; for feedforward neural network layers For each neuron in the system, the system records the amplitude of its activation response to each feature block sample. And calculate the neuron in the unlabeled image calibration set. The activation mean and activation standard deviation are calculated based on a preset active dead zone threshold. The system detects and removes 80 inactive neurons, and then transfers them to the feedforward parameter pool. Simultaneously, the corresponding output weights of the first fully connected layer and input weights of the second fully connected layer are deleted to maintain the consistency of the dimension in subsequent matrix calculations; for the remaining 3000 effective neurons, the system calculates the... Activation selectivity index of neurons .

[0066] (2) Divide neurons and construct core expert groups, general expert groups and routers.

[0067] The system sets an activation selectivity threshold. The activation selectivity index of effective neurons With activation selectivity threshold Compare; the system will satisfy The 1280 neurons were divided into proprietary neurons, and their original pre-trained weights were used as the core expert group. The foundation for this construction; the remaining 1712 neurons are divided into general neurons, and their original pre-trained weights are used as the general expert group. The above-mentioned division method, while preserving the original pre-trained weights, allows high-selectivity neurons to undertake more discriminative feature modeling tasks such as vehicle edges, pedestrian contours, traffic sign textures, and lane line intersections, while allowing low-selectivity neurons to undertake more tasks of preserving basic visual information such as sky, roads, and building facades. For the 1280 proprietary neurons, the system extracts their corresponding first fully connected layer weight vectors, each with a dimension of 768, thus constructing a structure with a size of... The system sets the number of core experts to 8, performs K-Means clustering on the weight vectors in the weight matrix to obtain 8 feature clusters, and calculates the cluster center of each feature cluster. To ensure the consistency of the expert structure, the system performs capacity balancing on the number of neurons in each feature cluster, so that each core expert network submodule contains 160 intermediate layer neurons, thereby constructing 8 core expert network submodules with consistent intermediate layer widths, forming a core expert group. Simultaneously, the system arranges the eight cluster centers sequentially according to the core expert numbers, forming a cluster with a size of [missing information]. Router initial weight matrix and assign it to the router. The linear weight layer allows the router to obtain prior information from the original pre-trained parameter distribution during the initial fine-tuning phase, reducing the cold start risk caused by random routing. For the 1712 general neurons, the system balances and reassembles the weights of the first and second fully connected layers corresponding to the general neurons according to the preset requirements for the number and width of general experts. The system constructs 8 general expert network sub-modules, each containing 214 intermediate layer neurons, thus forming a general expert group. At this point, the system has completed the core expert group. General Expert Group and router The construction of the V-MoE model and its composition as the transferred model. The expert calculation part.

[0068] (3) Construct a labeled fine-tuning set and perform feature block routing based on local feature complexity.

[0069] In autonomous driving target detection and semantic segmentation tasks, the system is based on an unlabeled image calibration set. The images in the dataset can be used to call upon or supplement their bounding box annotations and pixel-level semantic annotations to form labeled fine-tuning sets. Labeled micro-tuning set Including training images and their corresponding ground truth labels Authentic Labels This includes bounding box annotations or semantic segmentation annotations for categories such as vehicles, pedestrians, traffic signs, roads, lane areas, buildings, and sky; the system will fine-tune the labeled sets. The training images are converted into a sequence of input feature blocks in the target domain in the same manner as in the calibration phase. Each image corresponds to 1024 feature blocks, and the hidden layer feature dimension of each feature block is 768. During the fine-tuning stage, for the target domain input feature block sequence... The first in For each feature block, before the system enters the expert routing stage and performs normalization processing, it calculates the local feature complexity score along the hidden layer feature dimension. And the local feature complexity score With complexity threshold Compare; for those that satisfy Low-complexity feature blocks, such as large areas of sky, smooth asphalt roads, distant weakly textured backgrounds, or relatively uniform building wall areas, are assigned to the general expert group by the system. To preserve basic visual information such as color, texture, and outline, as well as their original mesh order; for those that meet the requirements The system inputs highly complex feature blocks, such as traffic signs, pedestrian limbs, vehicle edges, lane intersections, or areas containing distant small targets, into the router. ;router Based on size weight matrix The system performs matching calculations on high-complexity feature blocks and outputs routing probabilities pointing to 8 core experts through a normalized exponential function. A Top-2 routing mechanism is used to assign each high-complexity feature block to the two core experts with the highest probability scores. For the outputs of the two core experts corresponding to the same high-complexity feature block, the system performs weighted fusion according to their routing probabilities to obtain the output of a single core expert corresponding to that feature block. Subsequently, the system groups a general expert group. Low-complexity feature block representation of the output and the core expert group The output high-complexity feature block representation is reassembled according to the grid order in the original image into a form of size. The feature sequences are then input into subsequent Transformer encoding layers. A global self-attention mechanism is used to re-establish the spatial and semantic relationships between feature blocks, resulting in the final feature representation. .

[0070] (4) Perform downstream visual task fine-tuning and core expert load balancing training.

[0071] The system will represent the final features. Input dense prediction head built specifically for autonomous driving scenarios Dense prediction head It includes a bounding box regression branch and a pixel-level classification branch, and outputs the target bounding box coordinates, target category, and semantic segmentation mask as the prediction results. The system is based on the prediction results. With real labels Calculate the bounding box regression loss, class classification loss, and segmentation cross-entropy loss, and use them as the task loss. Meanwhile, the system, based on the router Output of core expert routing probabilities and core expert groups Calculate the load balancing loss of each core expert based on their actual received frequency within the current training batch. Load balancing losses Used only to constrain the core expert group Internal routing allocation is not applicable to the general expert group. Forcibly imposing balance constraints; the system will suffer task losses. With load balancing losses We perform a weighted summation to obtain the total loss. Finally, the system is based on the total loss. For dense prediction heads ,router Core Expert Group and V-MoE model Other trainable parameters involved in fine-tuning are updated via backpropagation; for the general expert group The system adopts a lower level than the core expert group. The learning rate is updated to reduce excessive drift of its underlying representation capabilities during training.

[0072] Case 3: High-resolution remote sensing image ground feature analysis and target extraction system based on the method of this invention

[0073] (1) Construct a pool of feedforward parameters to be transferred and calculate neuron activation selectivity.

[0074] The system first obtains 5000 remote sensing images containing complex terrain by uniformly and randomly sampling from the DOTA aerial remote sensing image dataset, and then constructs a labelless image calibration set of remote sensing scenes without using target bounding box annotations and ground feature region annotations. In the data preprocessing stage, to preserve the spatial structure and details of ground features such as buildings, vehicles, roads, and port facilities, the unlabeled image calibration set is... Each original image in the process is uniformly cropped to Image tensor of resolution; set image patch size to The system performs non-overlapping meshing on the image tensor, breaking down each remote sensing image into 4096 local image patches. These 4096 patches are then flattened sequentially and input into a linear projection layer, mapping them to a hidden layer feature dimension of 768. Positional encoding corresponding to the mesh location of each image patch is then added, resulting in a final image with dimensions of [missing information]. Input feature block sequence The system loads a dense visual Transformer model that has been pre-trained under self-supervised guidance. Locate the target feedforward neural network layer. Feedforward neural network layer The system employs a multilayer perceptron architecture with an expansion ratio of 4, containing 3072 intermediate layer neurons; the system extracts the feedforward neural network layer. The original pre-trained weights form the feedforward parameter pool to be transferred. Subsequently, the system will input the feature block sequence. Traversing the input dense visual Transformer model The total number is Feature block samples; for feedforward neural network layers For each neuron in the system, the system records the amplitude of its activation response to each feature block sample. And calculate the neuron in the unlabeled image calibration set. The activation mean and activation standard deviation are calculated based on a preset active dead zone threshold. The system detected and removed 56 inactive neurons, and then transferred them to the feedforward parameter pool. Simultaneously, the corresponding output weights of the first fully connected layer and input weights of the second fully connected layer are deleted to maintain the consistency of the dimension of subsequent matrix calculations; for the remaining 3016 effective neurons, the system calculates the... Activation selectivity index of neurons .

[0075] (2) Divide neurons and construct core expert groups, general expert groups and routers.

[0076] The system sets an activation selectivity threshold. Considering that core ground features such as buildings, vehicles, port facilities, and road nodes account for a relatively small proportion and their local structures are relatively concentrated in high-resolution remote sensing images, the system will effectively select the activation selectivity index of neurons. With activation selectivity threshold Compare; the system will satisfy The 768 neurons were divided into proprietary neurons, and their original pre-trained weights were used as the core expert group. The foundation for this construction is as follows: the remaining 2248 neurons are divided into general neurons, and their original pre-trained weights are used as the general expert group. The above-mentioned division method, while preserving the original pre-trained weights, allows high-selectivity neurons to undertake more tasks of modeling local features with significant differences, such as building edges, vehicle outlines, road intersections, and port facilities, while allowing low-selectivity neurons to undertake more tasks of preserving basic visual information, such as sea surfaces, deserts, farmland, clouds, and large-area background textures. For the 768 dedicated neurons, the system extracts the corresponding first fully connected layer weight vectors, each with a dimension of 768, thereby constructing a system with a size of... The system sets the number of core experts to 8, performs K-Means clustering on the weight vectors in the weight matrix to obtain 8 feature clusters, and calculates the cluster center of each feature cluster. To ensure the consistency of the expert structure, the system performs capacity balancing on the number of neurons in each feature cluster, so that each core expert network submodule contains 96 intermediate layer neurons, thereby constructing 8 core expert network submodules with consistent intermediate layer widths, forming a core expert group. Simultaneously, the system arranges the eight cluster centers sequentially according to the core expert numbers, forming a cluster with a size of [missing information]. Router initial weight matrix and assign it to the router. The linear weight layer allows the router to obtain prior information from the original pre-trained parameter distribution during the initial fine-tuning phase, reducing the cold start risk caused by random routing. For the 2248 general neurons, the system balances and reassembles the weights of the first and second fully connected layers corresponding to the general neurons according to the preset requirements for the number and width of general experts. The system constructs 8 general expert network sub-modules, each containing 281 intermediate layer neurons, thus forming a general expert group. At this point, the system has completed the core expert group. General Expert Group and router The construction of the V-MoE model and its composition as the transferred model. The expert calculation part.

[0077] (3) Construct a labeled fine-tuning set and perform feature block routing based on local feature complexity.

[0078] In the task of remote sensing image ground feature analysis and target extraction, the system is based on an unlabeled image calibration set. The image data is retrieved or supplemented with its target bounding box annotations, rotated bounding box annotations, or feature area annotations to form a labeled fine-tuning set. Labeled micro-tuning set Including training images and their corresponding real labels Authentic Labels Spatial labeling of land features including buildings, vehicles, ships, roads, bridges, port facilities, farmland, and water bodies; the system will fine-tune the labeled set. The training images are converted into a sequence of input feature blocks in the target domain in the same manner as in the calibration phase. Each image corresponds to 4096 feature blocks, and the hidden layer feature dimension of each feature block is 768. In the fine-tuning stage, for the target domain input feature block sequence... The first in For each feature block, before the system enters the expert routing stage and performs normalization processing, it calculates the local feature complexity score along the hidden layer feature dimension. And the local feature complexity score With complexity threshold Compare; for those that satisfy Low-complexity feature blocks, such as large areas of sea, desert, farmland, clouds, or background areas with weak texture variations, are assigned to the general expert group by the system. To preserve basic visual information such as color, texture, and outline, as well as their original mesh order; for those that meet the requirements Highly complex feature blocks, such as building edges, vehicle targets, road intersections, ship outlines, or areas where port facilities are located, are input into the router by the system. ;router Based on size weight matrix The system performs matching calculations on high-complexity feature blocks and outputs routing probabilities pointing to 8 core experts through a normalized exponential function. A Top-2 routing mechanism is used to assign each high-complexity feature block to the two core experts with the highest probability scores. For the outputs of the two core experts corresponding to the same high-complexity feature block, the system performs weighted fusion according to their routing probabilities to obtain the output of a single core expert corresponding to that feature block. Subsequently, the system groups a general expert group. Low-complexity feature block representation of the output and the core expert group The output high-complexity feature block representation is reassembled according to the grid order in the original image into a form of size [size missing]. The feature sequences are then input into subsequent Transformer encoding layers. A global self-attention mechanism is used to re-establish the spatial and semantic relationships between feature blocks, resulting in the final feature representation. .

[0079] Specifically, in Example 3: the high-resolution remote sensing image ground feature analysis and target extraction system based on the method of the present invention, the downstream visual task fine-tuning and core expert load balancing training include: The system will represent the final features. Input a dense prediction head specifically built for remote sensing imagery Dense prediction head It includes a rotated bounding box regression branch, an object category classification branch, and a ground feature region resolution branch, and outputs the oriented object bounding box coordinates, object category, and ground feature region resolution results as prediction results. The system is based on the prediction results. With real labels Calculate the rotation bounding box regression loss, category classification loss, and feature region resolution loss, and use them as the task loss. Meanwhile, the system, based on the router Output of core expert routing probabilities and core expert groups Calculate the load balancing loss of each core expert based on their actual received frequency within the current training batch. Load balancing losses Used only to constrain the core expert group Internal routing allocation is not applicable to the general expert group. Forcibly imposing balance constraints; the system will suffer task losses. With load balancing losses We perform a weighted summation to obtain the total loss. Finally, the system is based on the total loss. For dense prediction heads ,router Core Expert Group and V-MoE model Other trainable parameters involved in fine-tuning are updated via backpropagation; for the general expert group The system adopts a lower level than the core expert group. The learning rate is updated to reduce excessive drift of its underlying representation capabilities during training.

[0080] The present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. An efficient parameter reuse method for transferring self-supervised pre-trained models to V-MoE, characterized in that, include: Step S1: Construct an input feature block sequence based on the unlabeled image calibration set, extract the feedforward parameter pool to be transferred in the dense model, and calculate the neuron activation selectivity; Obtaining unlabeled image calibration sets The unlabeled image calibration set The image is converted into a sequence of input feature blocks. ; The input feature block sequence Input a dense visual Transformer model that has been pre-trained under self-supervision Extract the dense visual Transformer model The original weight matrix of the feedforward neural network layer F is used to form the feedforward parameter pool to be transferred. ; The activation responses of each neuron in the feedforward neural network layer F to different feature block samples are statistically analyzed, and inactive neurons are removed based on the average activation. The activation selectivity index of each effective neuron is then calculated. ; Step S2: Based on the activation selectivity, divide the proprietary neurons and general neurons, construct the core expert group and the general expert group respectively, and initialize the router; Set activation selectivity threshold The activation selectivity index of each effective neuron With the activation selectivity threshold A comparison will satisfy The neuron is identified as a specialized neuron, and will satisfy the following... The neurons were identified as general neurons; Based on the weights of the first and second fully connected layers corresponding to the proprietary neurons, multiple core expert network sub-modules with consistent intermediate layer widths are constructed to form a core expert group. Initialize the router weight matrix ; Based on the preset requirements for the number and width of general experts, the weights of the first and second fully connected layers corresponding to the general neurons are balanced, grouped, and reassembled to construct multiple general expert network sub-modules with consistent intermediate layer widths, forming a general expert group. And by the core expert group General Expert Group and router Constructing the transferred V-MoE model The expert calculation part; Step S3: Construct a labeled fine-tuning set, calculate the local feature complexity of the feature blocks, and distribute low-complexity blocks and high-complexity blocks to different expert groups; From the labelless image calibration set Image samples are extracted and manually labeled to form a labeled fine-tuning set. The labeled fine-tuning set This includes training images and their corresponding ground truth labels P; The labeled fine-tuning set The training images are converted into a sequence of input feature blocks in the target domain. Calculate the target domain input feature block sequence Local feature complexity score of each feature block ; Set a complexity threshold The local feature complexity score of each feature block. With the aforementioned complexity threshold A comparison will satisfy The feature blocks are determined to be low-complexity feature blocks, which will satisfy... The feature block is determined to be a high-complexity feature block; Low-complexity feature blocks are allocated to the general expert group according to a preset balanced allocation rule. A general expert network submodule is used to extract and maintain relatively stable basic visual information such as color, texture, and outline, while preserving their original location identifiers and grid order. The high-complexity feature block is input into the router R, and the router R performs Top-N routing selection based on the routing probability, assigning the high-complexity feature block to the top N core experts with the highest probability scores; The outputs of the first N core experts are weighted and fused according to their corresponding routing probabilities to obtain the output of a single core expert corresponding to a high-complexity feature block. The general expert group With core expert group The output feature blocks are recombined according to the original grid order and input into the subsequent Transformer encoding layer to obtain the final feature representation Y; Step S4: Obtain the final feature representation based on the expert output, and fine-tune the transfer model downstream by combining the task loss and the core expert load balancing loss. A prediction head H is constructed for the downstream visual task, and the final feature representation Y is input into the prediction head H to obtain the prediction result. ; Based on the prediction results With the labeled fine-tuning set Calculate the task loss corresponding to the real label P. ; Computing Core Expert Group Calculate load balancing loss ; The task loss With core expert group Load balancing loss We perform a weighted summation to obtain the total loss L: ; in, This is the preset loss balance coefficient; Based on the total loss L, the prediction head H, router R, and core expert group are analyzed. and V-MoE model Other trainable parameters involved in fine-tuning are backpropagated and updated, and the general expert group is also updated. The parameters used are lower than those of the core expert group. The learning rate is updated.

2. The efficient parameter reuse method for transferring self-supervised pre-trained models to V-MoE according to claim 1, characterized in that, In step S1, the unlabeled image calibration set The image is converted into a sequence of input feature blocks. ,include: Extract the unlabeled image calibration set The original image in For the original image Size normalization and tensor quantization are performed, and the images are processed according to a preset image block size. Perform non-overlapping mesh segmentation; The segmented local image patches are flattened and input into a linear projection layer, and positional encodings corresponding to the grid positions of the image patches are added to generate an input feature block sequence. .

3. The efficient parameter reuse method for transferring self-supervised pre-trained models to V-MoE according to claim 1, characterized in that, In step S1, the activation response of each neuron in the feedforward neural network layer F to different feature block samples is statistically analyzed, and inactive neurons are removed based on the average activation status. The activation selectivity index of each effective neuron is then calculated. ,include: For input feature block samples x Record the feedforward neural network layer F The Middle i Activation response amplitude of a single neuron And calculate the neuron in the unlabeled image calibration set. Activation mean and activation standard deviation ; Set active dead zone threshold , will satisfy The neurons were identified as inactive neurons and included in the feedforward parameter pool to be migrated. Simultaneously delete the corresponding first fully connected layer output weights and second fully connected layer input weights. For the remaining effective neurons, the first one is calculated according to the following formula. Activation selectivity index of neurons : ; in, To prevent extremely small positive numbers with a denominator of zero.

4. The efficient parameter reuse method for transferring a self-supervised pre-trained model to V-MoE according to claim 1, characterized in that, In step S2, multiple core expert network sub-modules with consistent intermediate layer widths are constructed based on the weights of the first and second fully connected layers corresponding to the proprietary neurons, forming a core expert group. Initialize the router weight matrix ,include: Extract the weights of the first fully connected layer corresponding to the specialized neuron and construct a weight matrix. ,in, G This represents the total number of proprietary neurons. d Indicates the dimension of the hidden layer features; With the weight matrix Each weight vector in the dataset is used as a sample to be clustered. K-Means clustering is performed to obtain K feature clusters, and K cluster centers are calculated. ; Based on the clustering results, the weights of the first and second fully connected layers corresponding to the specialized neurons are sliced, capacity balanced, and reassembled to construct K core expert network sub-modules with consistent intermediate layer widths, forming a core expert group. ; The K cluster centers Arranged sequentially by expert number, forming a size of Router initial weight matrix The router's initial weight matrix Each column corresponds to an initial matching vector for a core expert.

5. The efficient parameter reuse method for transferring a self-supervised pre-trained model to V-MoE according to claim 1, characterized in that, In step S3, the target domain input feature block sequence is calculated. Local feature complexity score of each feature block ,include: Input feature block sequence for the target domain The first in j For each feature block, before it enters the expert routing stage and before normalization is performed, the corresponding feature representation is extracted and processed along the feature dimension of the hidden layer. d Calculate the local feature complexity score : ; in, Indicates the first j The feature block in the th ... c Feature values ​​on each feature channel Indicates the first j Each feature block in the whole d The characteristic mean of each channel.

6. The efficient parameter reuse method for transferring a self-supervised pre-trained model to V-MoE according to claim 1, characterized in that, In step S4, the core expert group is calculated. Calculate load balancing loss ,include: For the core expert group Based on the actual received frequencies of each core expert within a training batch and the average routing probability output by the router, the load balancing loss is calculated. : ; in, Indicates the first k The normalized frequency at which core experts are actually selected This indicates that the router points to the first... k Average soft router probability of each core expert; When using the Top-N routing mechanism, the and Calculate according to the following formulas: ; ; in, This represents the number of high-complexity feature blocks that participate in core expert routing within a training batch. Indicates the first z The first high-complexity feature block is routed to the second... k The probability of a core expert using a software router. This indicates an indicator function.

Citation Information

Patent Citations

  • Scene and task dual-conditioned visual hybrid expert model construction and reasoning method

    CN121811180A

  • Incremental named entity recognition method based on hybrid expert network

    CN122334260A