Migration method of multi-modal pre-training model and related product
By assigning independent low-rank fine-tuning terms and gating parameters to the attention heads in the visual encoder of the multimodal pre-trained model, the problem of insufficient generalization ability when migrating in new fields is solved, and the accuracy and convergence speed of migration tasks are improved.
Patent Information
- Application Number
- CN202510243032.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-17
AI Technical Summary
Existing multimodal pretrained models are difficult to generalize effectively when migrating to new fields, resulting in poor performance in new fields.
The performance of the model in the migration task is optimized by assigning independent low-rank fine-tuning terms and gating parameters to each attention head of each Transformer encoder module in the visual encoder of the multimodal pre-trained model.
The accuracy and convergence speed of the model on downstream tasks are improved, and the robustness of the model to the distribution changes of input data in different fields is enhanced.
Smart Images

Figure CN120162741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for migrating a multi-modal pre-trained model and related products. Background Art
[0002] Data in the real environment can be divided into various different domains according to the pattern of distribution, and there is a more or less domain shift between different domains. This domain shift may stem from changes in lighting conditions, background settings, sensor types, or any other factors affecting data distribution.
[0003] In many practical applications, the data used to train a model may come from a specific distribution or domain. However, when the model is deployed, new domains that did not appear in the training data may be encountered. A model that is too specialized in the training domain may not generalize well to these new domains. The goal of domain generalization is to address this challenge. Specifically, domain generalization refers to the ability of a machine learning model to perform well in different domains after being trained with data from a specific domain. Its goal is to make the model robust to changes in the input data distribution that may occur in different real-world scenarios or domains. Domain generalization is of great significance in various fields such as computer vision, natural language processing, robotics, and healthcare, where models need to adapt to diverse and evolving environments.
[0004] In related technologies, Chinese Patent Application CN118097685A proposes a method for migrating a multi-modal pre-trained model based on self-supervised learning, focusing on how to optimize text prompts. Chinese Patent Application CN119295895A proposes a multi-modal model style embedding method and system for complex power vision scenarios. By establishing an instance-level style feature extraction model and aligning the style information of the instance-level style feature extraction model with the style information in the domain-level style information library, for a single image input during the inference process, it can generate efficient and accurate style prompt words, which is applicable to tasks such as defect recognition and object detection in real power scenarios, effectively enhancing the style generalization performance of the visual text pre-trained model in downstream tasks.
[0005] The present invention provides a new method for migrating a multi-modal pre-trained model. Summary of the Invention
[0006] The present invention provides a method for migrating a multi-modal pre-trained model and related products to improve the accuracy of model migration to downstream tasks and increase the convergence speed.
[0007] The present invention adopts the following technical solutions.
[0008] The present invention provides a method for migrating a multi-modal pre-trained model, including:
[0009] Provide a multi-modal pre-trained model and training samples. Among them, the multi-modal pre-trained model includes a text encoder, a visual encoder, and an output module. The visual encoder includes one or more Transformer encoder modules. A single Transformer encoder module includes a multi-head self-attention mechanism layer. The training samples are text-image pairs;
[0010] Input the text in the training sample into the text encoder to obtain text features;
[0011] Input the image in the training sample into the visual encoder to obtain image features. Among them, an independent low-rank fine-tuning term is assigned to the mapping matrix of each attention head of each Transformer encoder module in the visual encoder. Each low-rank fine-tuning term is the product of an independent A matrix and an independent B matrix, and an independent gating parameter is assigned to each attention head of each Transformer encoder module in the visual encoder. The output of the Transformer encoder module is determined according to the output features of its own attention heads and the gating parameters corresponding to the attention heads. The gating parameters are used to adjust the weights of the corresponding attention heads;
[0012] Input the text features and image features of the training sample into the output module to obtain a classification loss, and iterate on each A matrix, each B matrix, and each gating parameter according to the classification loss.
[0013] In some embodiments, the multi-modal pre-trained model is used for image classification, semantic segmentation, or object detection.
[0014] In some embodiments, the network structures of each Transformer encoder module are the same, or the network structures of at least two Transformer encoder modules are different.
[0015] In some embodiments, at least one Transformer encoder module includes a normalization layer, a multi-head self-attention mechanism layer, a residual layer, a normalization layer, a feed-forward neural network layer, and a residual layer arranged in sequence.
[0016] In some embodiments, the calculation formula of the classification loss is as follows:
[0017] Among them, is the classification loss, N is the total number of samples, i is the sample number, C is the total number of categories, c is the category number, S i is the image feature of the i-th sample, T y is the text feature corresponding to the true label, T c is the text feature corresponding to the category numbered c, and τ is the temperature coefficient.
[0018] In some embodiments, the output of the multi-head self-attention mechanism layer in a single Transformer encoder module in the visual encoder is denoted as f g , where f1, f2, ..., f H represent the features output by different attention heads, the subscripts of which represent the numbers of the attention heads, and H is the number of attention heads in the multi-head self-attention mechanism layer of this single Transformer encoder module. represent the normalized values of the gating parameters corresponding to different attention heads, the subscripts of which represent the numbers of the attention heads, and γ is a preset coefficient.
[0019] In some embodiments, the gating parameters of a single Transformer encoder module in the visual encoder are normalized according to the following formula: where g i is the gating parameter corresponding to the attention head encoded as i, and softmax(g i ) is the normalized value of g i , and g j is the gating parameter corresponding to the attention head encoded as j.
[0020] The present invention also provides a migration device for a multi-modal pre-training model, including:
[0021] An input module, configured to receive a multi-modal pre-training model and training samples, where the multi-modal pre-training model includes a text encoder and a visual encoder, the visual encoder includes a Transformer encoder module, the Transformer encoder module includes one or a plurality of multi-head self-attention mechanism layers arranged in sequence, and the training samples are text-image pairs;
[0022] A text processing module, configured to input the text in the training samples into the text encoder to obtain text features;
[0023] An image processing module, configured to input the images in the training samples into the visual encoder to obtain image features, where an independent low-rank fine-tuning term is assigned to the mapping matrix of each attention head, each low-rank fine-tuning term is the product of an independent A matrix and an independent B matrix, and an independent gating parameter is assigned to each attention head;
[0024] The image features are determined according to the output features of each attention head and the gating parameters corresponding to each attention head, and the gating parameters are used to adjust the weights of the corresponding attention heads;
[0025] An iterative update module is used to input the text features and image features of the training samples into an output module to obtain a classification loss, and iterate on each A matrix, each B matrix, and each gating parameter according to the classification loss.
[0026] The present invention also provides a computer device, including a memory and a processor, where the memory stores instructions, and the processor runs the instructions to execute the foregoing method for migrating a multi-modal pre-trained model.
[0027] The present invention also provides a program product, which executes the foregoing method for migrating a multi-modal pre-trained model when running on a processor.
[0028] The A matrix and B matrix of each attention head are independently optimized, which gives the optimization process greater optimization freedom in the parameter space, thereby increasing the accuracy of the model migrated to downstream tasks. Assigning independent gating parameters to each attention head helps to enhance the role of the attention heads associated with the tasks of model migration and suppress the role of the attention heads irrelevant to the tasks of model migration, thereby accelerating convergence and improving the performance of the model under specific tasks. The two optimization measures work together to improve the accuracy of the model migrated to downstream tasks and increase the convergence speed. Description of the Drawings
[0029] Figure 1 is a schematic flowchart of the method for migrating a multi-modal pre-trained model of the present invention.
[0030] Figure 2 is a schematic network structure diagram of a Transformer encoder module in the multi-modal pre-trained model of the present invention.
[0031] Figure 3 is a schematic structural diagram of the device for migrating a multi-modal pre-trained model of the present invention.
[0032] Figure 4 is a structural block diagram of the computer device of the present invention. Detailed Embodiments
[0033] The present invention will be further described below in conjunction with specific embodiments, but the protection scope of the present invention is not limited thereto.
[0034] Refer to Figure 1 , the present invention provides a method for migrating a multi-modal pre-trained model, including the following steps.
[0035] Step 101: Provide a multi-modal pre-trained model and training samples. The multi-modal pre-trained model includes a text encoder, a visual encoder, and an output module. The visual encoder includes one or more layers of Transformer encoder modules. Each individual Transformer encoder module includes a multi-head self-attention mechanism layer. The training samples are text-image pairs.
[0036] The multi-modal pre-trained model includes a text encoder and a visual encoder. It should also include an output module for comparative analysis of the output of the text encoder (also known as text embedding or text features) and the output of the visual encoder (also known as image embedding or image features). The visual encoder contains one layer of Transformer encoder module or multiple layers of Transformer encoder modules arranged in a stack. Each Transformer encoder module contains a multi-head self-attention mechanism layer. The network structure of the Transformer encoder module can be a typical structure or other known variant structures.
[0037] The training samples are text-image pairs. The text in the text-image pairs should be grammatically correct and have clear semantic expressions, highlighting specific category information.
[0038] The text in the text-image pairs is, for example: "a photo of + category name", "a + category name", "an image of + category name", "a photo of + category name in Chinese", "an image of + category name in Chinese".
[0039] In some embodiments, the multi-modal pre-trained model is a typical CLIP (Contrastive Language-Image Pretraining) network structure.
[0040] In some other embodiments, the multi-modal pre-trained model is a ClearCLIP network structure (reference: Lan M, Chen C, Ke Y, et al. Clearclip: Decomposing clip representations for dense vision-language inference[C] / / European Conference on Computer Vision. Springer, Cham, 2025: 143-160). This multi-modal pre-trained model can be used for semantic segmentation.
[0041] The following introduces the typical CLIP network structure, which includes a visual encoder, a text encoder, and a contrastive learning objective. The visual encoder converts the input image into an image embedding. The text encoder converts the input text into a text embedding. The text corresponding to each candidate class name is converted into a text embedding. The model is trained through contrastive learning to make the image embedding and text embedding corresponding to the image and text as close as possible in the feature space, while the image embedding and text embedding corresponding to unrelated images and texts are as far away as possible in the feature space. This network structure enables effective cross-modal retrieval and matching even for class names that have not been trained after the model training is completed.
[0042] The function of the multi-modal pre-training model can be to correctly describe the input image or to perform semantic segmentation on the input image. In the present invention, the multi-modal pre-training model can be used for image classification, semantic segmentation, or object detection, etc.
[0043] Step 102: Input the text in the training sample into the text encoder to obtain text features.
[0044] The role of the text encoder is to convert the input text (usually a natural language description) into a vector of a fixed dimension.
[0045] In the CLIP model, the text features are usually a 512-dimensional vector (for example, using ViT-B / 32 as the base model). This means that after being processed by the text encoder, each text input will be converted into a 512-dimensional embedding vector. The dimension of the text features is not limited to 512 dimensions and can be adjusted according to actual needs.
[0046] In some embodiments, the operation process of the text encoder includes a text preprocessing step and an encoding step. In the text preprocessing step, the input text is tokenized to convert the input text into word units. In the encoding step, the semantic information of the tokenized text is extracted through the self-attention mechanism to obtain a vector of a fixed length (typically a 512-dimensional vector), and this vector captures the global semantic information of the text.
[0047] Step 103: Input the images in the training samples into the vision encoder to obtain image features. Among them, an independent low-rank fine-tuning term is assigned to the mapping matrix of each attention head of each Transformer encoder module in the vision encoder. Each low-rank fine-tuning term is the product of an independent A matrix and an independent B matrix. And an independent gating parameter is assigned to each attention head of each Transformer encoder module in the vision encoder. The output of the Transformer encoder module is determined according to the output features of its own attention heads and the gating parameters corresponding to the attention heads. The gating parameters are used to adjust the weights of the corresponding attention heads.
[0048] A typical operation process of the vision encoder is as follows.
[0049] First step: Divide the input image that conforms to the preset image format (for example, the resolution meets the set requirements) into image patches of a fixed size. The size of the image patches is, for example, 16×16 or 32×32. Subsequently, each image patch is mapped to a high-dimensional vector. Typically, a 16×16 image patch is mapped to a 768-dimensional embedding vector. The embedding vectors corresponding to all the image patches form a feature sequence.
[0050] Second step: Perform positional encoding on the embedding vector corresponding to each image patch. To retain the spatial position information of each image patch, the positional encoding of each image patch is added to its corresponding embedding vector through vector addition to obtain the serialized vector containing position information corresponding to each image patch. Thus, the model can perceive the position information of the image patch in the original image.
[0051] Third step: Input the serialized vector obtained in the second step, along with a randomly initialized class vector (CLS token), into the Transformer encoder module to obtain the encoded information of each image patch and the class vector processed by the encoder.
[0052] Fourth step: Determine the image features output by the vision encoder according to the encoded information of each image patch output by the last layer of the Transformer encoder module and the class vector processed by the encoder. The fourth step is implemented through a fully connected layer or other task-related heads.
[0053] In another typical operation process of the vision encoder, the Transformer encoder module only outputs the encoded information of each image patch, and does not output the class vector.
[0054] In some embodiments, the network structures of the Transformer encoder modules are the same. In other embodiments, the network structures of at least two Transformer encoder modules are different.
[0055] In some embodiments, with reference to Figure 2 , at least one Transformer encoder module includes a normalization layer, a multi-head self-attention mechanism layer, a residual layer, a normalization layer, a feed-forward neural network layer, and a residual layer arranged in sequence.
[0056] Both normalization layers are layer normalization.
[0057] The first residual layer sums the input of the first normalization layer and the output of the self-attention mechanism layer.
[0058] The second residual layer sums the input of the second normalization layer and the output of the feed-forward neural network layer.
[0059] The feed-forward neural network layer is a first fully-connected layer, an activation layer, and a second fully-connected layer connected in sequence.
[0060] In the present invention, the self-attention mechanism layers in the Transformer encoder modules of the visual encoder are all multi-head self-attention mechanisms. Each attention head multiplies the input matrix passed from the previous layer (for example, the matrix formed by splicing the aforementioned serialized vectors) by a projection matrix respectively to obtain a Q matrix, a K matrix, and a V matrix. The three projection matrices corresponding to the attention head numbered 1 are spliced together and denoted as W1. The three projection matrices corresponding to the attention head numbered 2 are spliced together and denoted as W2, and so on. The projection matrices of all the self-attention mechanism layers in all the Transformer encoder modules of the visual encoder are spliced together and denoted as W.
[0061] Step 104: Input the text features and image features of the training sample into the output module to obtain a classification loss, and iterate on each A matrix, each B matrix, and each gating parameter according to the classification loss.
[0062] When the multi-modal pre-training model is used for image classification, the calculation formula of the classification loss is as follows:
[0063] Where is the classification loss, N is the total number of samples, i is the sample number, C is the total number of categories, c is the category number, S i is the image feature of the i-th sample, T y is the text feature corresponding to the true label, T c is the text feature corresponding to the category numbered c, and τ is the temperature coefficient.
[0064] Adjusting the temperature coefficient can change the "smoothness" or "confidence" of the model output. The introduction of the temperature coefficient is usually used to adjust the probability distribution, making the model output smoother or sharper. This is a pre-determined hyperparameter and is a hyperparameter in the loss function.
[0065] In other application scenarios, the calculation formula of the classification loss can be adaptively adjusted.
[0066] The present invention realizes model migration by performing low-rank fine-tuning on the projection matrix of the multi-head self-attention mechanism layer and adjusting the weights of each attention head of the multi-head self-attention mechanism layer.
[0067] The optimization objectives include: θ1 represents the low-rank mapping matrix parameters included in the low-rank fine-tuning of the attention head, and θ2 represents the gating parameters of the attention head.
[0068] Specifically, θ1 is {A1, A2, …, A G , B1, B2, …, B G}, which refers to the A matrix and B matrix introduced by the low-rank fine-tuning, and its subscript is the number of the corresponding attention head. G is the total number of attention heads of all Transformer encoder modules in the visual encoder. θ2 is {g1, g2, …, g G}, g is the gating parameter of the attention head, and the subscript is the encoding of the corresponding attention head.
[0069] The optimization algorithm uses common optimization algorithms such as stochastic gradient descent method or adaptive moment estimation.
[0070] After the projection matrix W is subjected to low-rank fine-tuning, the projection matrix is denoted as W′, then there is:
[0071]
[0072] In the present invention, each A matrix and each B matrix are used as parameters for model optimization.
[0073] In the first iteration, the A matrix is a zero matrix and the B matrix is randomly set.
[0074] For each attention head of the multi-head self-attention mechanism layer of all Transformer encoder modules, a gating parameter g1 to g G is assigned. The initial value of the gating parameter is 1.
[0075] Subsequently, the gating parameters of each attention head of the same multi-head self-attention mechanism layer are normalized. The normalization process uses the softmax function. The specific formula is as follows: Among them, g iis the gating parameter corresponding to the i-th attention head, j is the number of the attention head, and H is the number of attention heads in the same multi-head self-attention mechanism layer.
[0076] Then the output of a single multi-head self-attention mechanism layer is the concatenation of the output features of each attention head after adjustment:
[0077]
[0078] where f g is the output feature of a single multi-head self-attention mechanism layer, f1, f2,..., f H represent the features output by different attention heads, and their subscripts represent the numbers of the attention heads. represent the normalized values of the gating parameters corresponding to different attention heads, and their subscripts represent the numbers of the attention heads. H is the total number of attention heads in a single multi-head self-attention mechanism layer, and γ is a preset coefficient used to compensate for the scale change after the softmax operation.
[0079] Using the above optimization algorithm, the A matrix and B matrix of each attention head in each Transformer encoder module in the visual encoder are independently optimized, which gives the optimization process greater optimization freedom in the parameter space, thereby increasing the accuracy of the model transferred to downstream tasks. Assigning independent gating parameters to each attention head in each Transformer encoder module in the visual encoder helps to enhance the role of the attention heads related to the tasks of model transfer and suppress the role of the attention heads irrelevant to the tasks of model transfer, thereby accelerating convergence and improving the performance of the model under specific tasks. The two optimization measures work together to improve the accuracy of the model transferred to downstream tasks and increase the convergence speed.
[0080] The iteration termination condition can be the loss function threshold or the set number of iteration rounds. A typical number of iteration rounds is 30 times.
[0081] Based on the same inventive concept, referring to Figure 3 , an embodiment of the present invention further provides a migration device for a multi-modal pre-training model, including:
[0082] An input module for providing a multi-modal pre-training model and training samples, where the multi-modal pre-training model includes a text encoder, a visual encoder, and an output module, the visual encoder includes one or more Transformer encoder modules, a single Transformer encoder module includes a multi-head self-attention mechanism layer, and the training samples are text-image pairs;
[0083] A text processing module for inputting the text in the training sample into a text encoder to obtain text features;
[0084] An image processing module for inputting the image in the training sample into a vision encoder to obtain image features, wherein an independent low-rank fine-tuning term is assigned to the mapping matrix of each attention head of each Transformer encoder module in the vision encoder, each low-rank fine-tuning term is the product of an independent A matrix and an independent B matrix, and an independent gating parameter is assigned to each attention head of each Transformer encoder module in the vision encoder. The output of the Transformer encoder module is determined according to the output features of its own attention heads and the gating parameters corresponding to the attention heads, and the gating parameters are used to adjust the weights of the corresponding attention heads;
[0085] An iterative update module for inputting the text features and image features of the training sample into an output module to obtain a classification loss, and iterating on each A matrix, each B matrix, and each gating parameter according to the classification loss.
[0086] The above modules can be implemented by software code, by hardware circuits, or a combination of both. The running processes of the modules refer to the foregoing embodiments and will not be elaborated herein.
[0087] Based on the same inventive concept, refer to Figure 4 , an embodiment of the present invention further provides a computer device, including a memory and a processor. The memory stores instructions, and the processor runs the instructions to execute the foregoing migration method of the multi-modal pre-training model.
[0088] Based on the same inventive concept, an embodiment of the present invention further provides a program product that executes the foregoing migration method of the multi-modal pre-training model when running on a processor.
[0089] The processor is, for example, any known processor type such as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a combination thereof.
[0090] The memory and the storage medium are, for example, various media that can store program codes such as a USB flash drive, a solid-state drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0091] The various embodiments in the present invention are all described in a progressive manner. The same or similar parts among the various embodiments can be referred to each other, and the key points described in each embodiment are the differences from other embodiments.
[0092] The protection scope of the present invention is not limited to the above embodiments. Obviously, those skilled in the art can make various changes and deformations to the present invention without departing from the scope and spirit of the present invention. If these changes and deformations fall within the scope of the claims of the present invention and its equivalent technologies, the intention of the present invention also includes these changes and deformations.
Claims
1. A migration method for a multimodal pre-trained model, characterized in that: include: Providing a multimodal pre-trained model and training samples, wherein the multimodal pre-trained model includes a text encoder, a visual encoder and an output module, the visual encoder includes one or more layers of Transformer encoder modules, a single Transformer encoder module includes a multi-head self-attention mechanism layer, and the training samples are text-image pairs; Inputting the text in the training sample into a text encoder to obtain text features; Input the image in the training sample into a visual encoder to obtain image features, wherein an independent low-rank fine-tuning item is assigned to the mapping matrix of each attention head of each Transformer encoder module in the visual encoder, each low-rank fine-tuning item is the product of an independent A matrix and an independent B matrix, and an independent gating parameter is assigned to each attention head of each Transformer encoder module in the visual encoder, the output of the Transformer encoder module is determined according to the output features of each attention head of the Transformer encoder module and the gating parameters corresponding to each attention head, and the gating parameters are used to adjust the weight of the corresponding attention head; The text features and image features of the training samples are input into an output module to obtain a classification loss, and each A matrix, each B matrix, and each gating parameter are iterated according to the classification loss.
2. The migration method according to claim 1, characterized in that: The multimodal pre-trained model is used for image classification, semantic segmentation or object detection.
3. The migration method according to claim 1, characterized in that: The network structures of each Transformer encoder module are the same, or the network structures of at least two Transformer encoder modules are different.
4. The migration method according to claim 1, characterized in that: At least one Transformer encoder module includes a normalization layer, a multi-head self-attention mechanism layer, a residual layer, a normalization layer, a feedforward neural network layer and a residual layer arranged in sequence.
5. The migration method according to claim 1, characterized in that: The classification loss is calculated as follows: in, is the classification loss, N is the total number of samples, i is the sample number, C is the total number of categories, c is the category number, S i is the image feature of the i-th sample, T y is the text feature corresponding to the true label, T c is the text feature corresponding to the category numbered c, is the temperature coefficient.
6. The migration method according to claim 1, characterized in that: The output of the multi-head self-attention mechanism layer in a single Transformer encoder module in the visual encoder is denoted as f g , Among them, f1, f2, ..., f H represents the features of the output of different attention heads, and its subscript represents the number of the attention head. H is the number of attention heads in the multi-head self-attention mechanism layer in the single Transformer encoder module. It represents the normalized value of the gating parameter corresponding to different attention heads. Its subscript represents the number of the attention head, and γ is the preset coefficient.
7. The migration method according to claim 6, characterized in that: The gating parameters of a single Transformer encoder module in the visual encoder are normalized according to the following formula: Among them, g i is the gating parameter corresponding to the attention head encoded as i, softmax(g i ) is g i The normalized value of g j is the gating parameter corresponding to the attention head encoded as j.
8. A migration device for a multimodal pre-trained model, characterized in that: include: An input module, configured to receive a multimodal pre-trained model and training samples, wherein the multimodal pre-trained model includes a text encoder and a visual encoder, the visual encoder includes a Transformer encoder module, the Transformer encoder module includes one or multiple multi-head self-attention mechanism layers arranged in sequence, and the training samples are text-image pairs; A text processing module, used for inputting the text in the training sample into a text encoder to obtain text features; An image processing module, used for inputting the image in the training sample into a visual encoder to obtain image features, wherein an independent low-rank fine-tuning item is assigned to the mapping matrix of each attention head, each low-rank fine-tuning item is the product of an independent A matrix and an independent B matrix, and an independent gating parameter is assigned to each attention head; The image features are determined according to the output features of each attention head and the gating parameters corresponding to each attention head, and the gating parameters are used to adjust the weight of the corresponding attention head; The iterative update module is used to input the text features and image features of the training samples into the output module to obtain the classification loss, and iterate each A matrix, each B matrix and each gating parameter according to the classification loss.
9. A computer device, characterized in that: It includes a memory and a processor, the memory stores instructions, and the processor runs the instructions to execute the migration method of the multimodal pre-trained model according to any one of claims 1 to 7.
10. A program product, characterized in that When running on a processor, the method for migrating a multimodal pre-trained model according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Multi-modal pre-training model migration method based on self-supervised learning
CN118097685A
Multi-modal model style embedding method and system for complex electric power visual scene
CN119295895A