Visual language model adaptation method and device, computer equipment and storage medium
The visual language model adaptation method built through pre-training, lightweight adaptation and dynamic routing mechanism solves the problem of insufficient migration and adaptation of visual language models in specific scenarios, and achieves efficient and accurate multi-scenario adaptability and low-cost model migration.
Patent Information
- Application Number
- CN202510734393.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing visual language models lack the ability to transfer and adapt in specific scenarios, resulting in significant performance degradation. Full fine-tuning is expensive and prone to catastrophic forgetting.
The visual language model is pre-trained by obtaining task data from multiple scenes, and scene-related parameters are injected using a lightweight adapter. The scene weight is predicted in combination with a dynamic routing mechanism, and mapping and decoding are performed based on cross-modal semantics to build a target adaptation model.
It significantly enhances the migration capability of the visual language model in a variety of different scenarios, reduces computing resource consumption, improves the generalization and adaptability of the model, and achieves efficient and accurate performance.
Smart Images

Figure CN120671768A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a visual language model adaptation method, device, computer equipment and storage medium. Background Art
[0002] With the development of large multimodal models (such as CLIP, BLIP, and GPT-4V), visual language models (VLMs) have shown broad application potential in tasks such as decision support, image-text alignment, and object recognition. However, the training data for these models primarily comes from general internet scenarios. While they possess high abstractness and broad generalization capabilities, they lack fine-grained adaptability to specific scenarios. Real-world applications, such as indoor intelligent robots, assisted diagnosis and treatment systems, and smart logistics terminals, require models to understand language commands within specific environments and perform object recognition tasks. This is particularly true in settings where the scenarios vary significantly but the task types remain stable, such as medical institutions, home care, retirement communities, and bank service centers. When fine-tuning a visual language model (VLM) in one environment and transferring it to another, its performance can significantly degrade. For example, the instruction "take the red bottle" might refer to a potion in a hospital setting but a spice bottle in a home setting, indicating that both the semantics and visual morphology of the target differ. This domain gap cannot be addressed through simple fine-tuning or image augmentation. Furthermore, existing methods often rely on fine-tuning the VLM model on a full scale, which is costly and prone to catastrophic forgetting. Therefore, how to improve the migration and adaptation capabilities of visual language models is a problem that technical personnel in this field need to solve. Summary of the Invention
[0003] Embodiments of the present invention provide a visual language model adaptation method, apparatus, computer device, and storage medium, aiming to enhance the adaptation effect of the visual language model.
[0004] In a first aspect, an embodiment of the present invention provides a visual language model adaptation method, comprising:
[0005] Acquire task data for multiple scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the multiple scenes using the visual language model;
[0006] Using multiple lightweight adapters to inject scene-related parameters into each scene respectively, and combining with the initial joint semantic features to obtain scene features;
[0007] A dynamic routing mechanism is used to predict scene weights, and a fusion feature is obtained based on the fusion of the scene features and the scene weights;
[0008] Based on the scenario, the fused features are mapped using cross-modal semantics, and the mapped fused features are decoded and output by a task decoder to construct a target adaptation model;
[0009] The target adaptation model is used to perform migration adaptation on the visual language model of the specified scene.
[0010] In a second aspect, an embodiment of the present invention provides a visual language model adaptation device, comprising:
[0011] a pre-training unit, configured to obtain task data for a plurality of scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the plurality of scenes using the visual language model;
[0012] A scene adaptation unit, configured to inject scene-related parameters into each scene using a plurality of lightweight adapters, and fuse the parameters with the initial joint semantic features to obtain scene features;
[0013] A dynamic routing unit, configured to predict a scene weight using a dynamic routing mechanism, and obtain a fused feature based on the scene feature and the scene weight;
[0014] A mapping and decoding unit is used to map the fused features based on the scenario using cross-modal semantics, and decode and output the mapped fused features through a task decoder to construct a target adaptation model;
[0015] The migration and adaptation unit is used to use the target adaptation model to perform migration and adaptation on the visual language model of the specified scene.
[0016] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the visual language model adaptation method as described in the first aspect when executing the computer program.
[0017] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the visual language model adaptation method as described in the first aspect is implemented.
[0018] An embodiment of the present invention provides a visual language model adaptation method, apparatus, computer equipment and storage medium, the method comprising: obtaining task data of multiple scenes, pre-training a visual language model using the task data, and extracting initial joint semantic features from the task data of multiple scenes through the visual language model; injecting scene-related parameters into each scene respectively using multiple lightweight adapters, and fusing the initial joint semantic features to obtain scene features; predicting scene weights using a dynamic routing mechanism, and fusing the scene features and scene weights to obtain fused features; mapping the fused features using cross-modal semantics based on the scene, and decoding and outputting the mapped fused features through a task decoder to construct a target adaptation model; and migrating and adapting the visual language model of a specified scene using the target adaptation model.
[0019] The embodiments of the present invention pre-train a visual language model on multiple scene task data to extract initial joint semantic features. Scene-related parameters are injected through a lightweight adapter and fused to obtain scene features. A dynamic routing mechanism is used to predict scene weights, and the scene features and weights are fused to obtain fused features. Finally, cross-modal semantics are used for scene-based mapping, and the output is decoded by a task decoder. Based on this, the embodiments of the present invention implement a lightweight adaptation architecture based on the decoupling of tasks and scenes. The design goal of this architecture is to significantly enhance the transferability of the visual language model (VLM) when facing tasks in a variety of different scenarios. By adopting this lightweight adaptation structure, the model can more flexibly adapt to various task requirements, thereby achieving more efficient and accurate performance in different scenarios. By decoupling the dependencies between tasks and scenes, this structure enables the model to quickly adjust its parameters to adapt to new task requirements when handling tasks in new scenarios, without the need for large-scale retraining. This lightweight design not only improves the model's generalization ability but also reduces computing resource consumption, making the model more efficient and practical in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A flowchart of a visual language model adaptation method provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of a sub-process of a visual language model adaptation method provided by an embodiment of the present invention;
[0023] Figure 3 Another flowchart of a visual language model adaptation method provided by an embodiment of the present invention;
[0024] Figure 4 A schematic diagram of the principle architecture of a visual language model adaptation method provided by an embodiment of the present invention;
[0025] Figure 5 A schematic block diagram of a visual language model adaptation device provided by an embodiment of the present invention;
[0026] Figure 6 A schematic block diagram of a visual language model adaptation device provided by an embodiment of the present invention;
[0027] Figure 7 Another schematic block diagram of a visual language model adaptation device provided by an embodiment of the present invention;
[0028] Figure 8 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0031] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0032] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0033] The visual language model adaptation method provided in an embodiment of the present invention can be applied in an application environment where a client and a server interact, wherein the client communicates with the server via a network. The server obtains task data for multiple scenes, uses the task data to pre-train a visual language model, and uses the visual language model to extract initial joint semantic features from the task data for the multiple scenes. Multiple lightweight adapters are used to inject scene-related parameters into each scene, and the initial joint semantic features are combined to obtain scene features. A dynamic routing mechanism is used to predict scene weights, and fused features are obtained based on the scene features and scene weights. Based on the scene, the fused features are mapped using cross-modal semantics, and the mapped fused features are decoded and output by a task decoder to construct a target adaptation model. The client sends a specified scene, and the server uses the target adaptation model to transfer and adapt the visual language model for the specified scene. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented as a standalone server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0034] See below Figure 1 , an embodiment of the present invention provides a visual language model adaptation method, which specifically includes: steps S101 to S105.
[0035] Step S101: Acquire task data for multiple scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the multiple scenes using the visual language model;
[0036] Step S102: using multiple lightweight adapters to inject scene-related parameters into each scene respectively, and combining the initial joint semantic features to obtain scene features;
[0037] Step S103: using a dynamic routing mechanism to predict scene weights, and fusing the scene features and scene weights to obtain fused features;
[0038] Step S104: Based on the scenario, the fused features are mapped using cross-modal semantics, and the mapped fused features are decoded and output by a task decoder to construct a target adaptation model;
[0039] Step S105: Using the target adaptation model to perform migration adaptation on the visual language model of the specified scene.
[0040] In this embodiment, combined with Figure 4First, we extract initial joint semantic features by acquiring task data from multiple scenarios and pre-training them using a visual language model. Next, we inject specific parameters into each scenario and fuse these parameters with the initial features using a lightweight adapter to obtain scene features. Then, we use a dynamic router to predict scene weights and combine them with the scene features to obtain the final fused features. Finally, cross-modal semantics performs cross-modal semantic mapping on the fused features, and the task decoder outputs the final result.
[0041] This embodiment implements a lightweight adaptation architecture based on the decoupling of tasks and scenes. The design goal of this architecture is to significantly enhance the migration capability of the visual language model (VLM) when facing tasks in a variety of different scenarios. By adopting this lightweight adaptation structure, the model can adapt to various task requirements more flexibly, thereby achieving more efficient and accurate performance in different scenarios. By decoupling the dependencies between tasks and scenes, this structure enables the model to quickly adjust its own parameters to adapt to new task requirements when processing tasks in new scenarios without the need for large-scale retraining. This lightweight design not only improves the generalization ability of the model, but also reduces the consumption of computing resources, making the model more efficient and practical in practical applications.
[0042] In actual application scenarios, the visual language model adaptation method provided by this embodiment can be applied to the medical field. For example, in an intelligent auxiliary diagnosis system, the visual language model can understand the doctor's language instructions, such as "show the abnormal area in the patient's lung X-ray", and accurately mark the lesion site in the image. In addition, the method can also be applied to the financial field. For example, in an intelligent banking service system, the visual language model can recognize the customer's language instructions, such as "check my account balance", and quickly display relevant information on the interface. Similarly, the method can also be applied to the smart home field, such as an intelligent voice assistant can understand the instructions of family members, such as "turn on the lights in the living room", and automatically perform the corresponding operations. Through the visual language model adaptation method provided by the embodiment of the present invention, the model can be flexibly adjusted according to the task requirements in different scenarios to achieve efficient and accurate performance, thereby greatly improving the practicality and adaptability of the visual language model in actual applications.
[0043] In one embodiment, if Figure 2 As shown, the step S101 includes steps S201 to S203.
[0044] Step S201: Encode the image data using a visual encoder and output it as image features;
[0045] Step S202: Encode the text data into language features using a language encoder;
[0046] Step S203: Fusing the image features and language features into the initial joint semantic features.
[0047] Specifically, fusing the image features and language features into the initial joint semantic features includes:
[0048] The initial joint semantic features are obtained by fusion according to the following formula:
[0049] f=VLM frozen (x,t)
[0050] Among them, f represents the initial joint semantic features, x represents image data, t represents text data, VLM represents the visual language model containing visual encoder and language encoder, and frozen represents parameter freezing.
[0051] This embodiment uses frozen pre-training to train the visual language encoder VLM, which freezes the model parameters during the training phase. The visual language encoder VLM then extracts the joint semantic features of the image and text as the basic representation. Mainstream models can use architectures such as CLIP, BLIP, or Flamingo, and their main components include:
[0052] Visual Encoder (ViT / ResNet): Input image x → Output image features f x ;
[0053] Language encoder (Transformer): input text t → output language feature f t ;
[0054] Fusion module: For example, the BLIP architecture uses Cross-Attention to fuse images and text to obtain a cross-modal representation f, which is the initial joint semantic feature.
[0055] In this way, it can be ensured that the model can retain the key semantic connections between images and text when processing multimodal data, providing a solid foundation for subsequent scene adaptation. The advantage of frozen pre-training is that it can flexibly adjust the model in different scenarios by introducing new task data and adaptation strategies without changing the pre-trained model parameters. This method not only improves the stability and reliability of the model, but also greatly shortens the time and cost of migrating the model between different tasks. At the same time, using the visual language model (VLM) as the basis for feature extraction can fully utilize the model's powerful capabilities in visual and language understanding, further improving the model's performance in complex scenarios.
[0056] In one embodiment, step S102 includes:
[0057] The residual module is used to fuse the initial joint semantic features and scene-related parameters into scene features:
[0058] f i =f+Adapter i (f)
[0059] Among them, f i represents scene features, f represents initial joint semantic features, Adapter i (f) represents scene-related parameters, and i represents the i-th scene.
[0060] After extracting the initial joint semantic features, a lightweight Adapter module is inserted into the image-text fusion layer, and each scene has an independent Adapter to achieve the purpose of injecting a small amount of scene-related parameters for different task scenarios. This embodiment uses the classic residual BottleneckAdapter:
[0061] f i =f+Adapter i (f),i∈{family, community, institution...}
[0062] In practical applications, the initial joint semantic features can be down-projected, activated using GELU / ReLU activation functions, and then up-projected. The scene features are then obtained by summing the residuals. In this way, the features of each scene effectively combine the initial joint semantic features with the relevant parameters of the specific scene, enabling the model to better understand and handle task requirements in different scenarios. In this way, the lightweight adapter not only enhances the model's scene adaptability, but also maintains its lightweight nature, avoiding the increased computational burden caused by the introduction of too many parameters.
[0063] In one embodiment, step S103 includes:
[0064] The scene weights are predicted using a lightweight multi-layer perceptron and fused according to the following formula to obtain the fusion features:
[0065]
[0066] Among them, f * represents the fusion feature, α i represents the scene weight of the i-th scene.
[0067] This embodiment uses a dynamic routing mechanism to dynamically determine which weighted combination to use. The structure of the dynamic routing mechanism can use a lightweight MLP to predict the weights α = [α1, α2, ..., α i], and then fused the predicted weights to obtain the fused features. In practical applications, a lightweight MLP can consist of multiple fully connected layers, using ReLU or GELU activation functions between layers. Finally, softmax normalization is performed to obtain the weights for each scene. This mechanism allows the model to dynamically adjust its focus in different scenarios, thereby enhancing its adaptability to new scenarios.
[0068] Furthermore, the introduction of a dynamic routing mechanism improves model interpretability. Because the weight of each scenario is predicted using a lightweight multi-layer perceptron, it's possible to intuitively understand the importance the model places on each scenario when handling tasks in different scenarios. This not only helps us better understand the model's decision-making process but also facilitates subsequent model optimization and debugging.
[0069] In practical applications, data augmentation strategies can be used to improve the generalization and robustness of models. Specifically, when acquiring task data from multiple scenarios, random transformations and perturbations are performed on image and text data, such as image rotation, scaling, cropping, and text synonym substitution, to increase data diversity and complexity. This allows the model to be exposed to more data variations during training, thereby learning more robust feature representations and improving its adaptability to new data.
[0070] In one embodiment, step S104 includes:
[0071] Obtain the task type, and select the corresponding task decoder according to the task type to decode and output the mapped fusion features.
[0072] Different task types may correspond to different output formats and requirements. For example, in an image recognition task, the task decoder may need to map fused features to specific category labels; in a text generation task, the task decoder may need to convert fused features into natural language text. Therefore, choosing the appropriate task decoder based on the task type is crucial.
[0073] In practice, a set of task decoders can be predefined, each optimized for one or more specific task types. When a new task is needed, the most appropriate decoder can be selected from the predefined set of decoders for decoding output based on the task type. This approach not only improves the flexibility and scalability of the model but also ensures optimal performance across different tasks.
[0074] Furthermore, to further improve model performance, a multi-task learning strategy can be employed. Specifically, during the training phase, the model can be trained using data from multiple tasks simultaneously. This allows the model to learn how to handle a single task while also learning the relevance and commonalities between different tasks. This allows the model to better utilize the complementary information between different tasks, thereby improving its performance on each task.
[0075] Specifically, such as Figure 3 As shown, the step of obtaining the task type and selecting the corresponding task decoder according to the task type to decode and output the mapped fusion features includes: steps S301 to S303.
[0076] Step S301: When the task type is a classification task, a multi-layer perceptron is used as the task decoder;
[0077] Step S302: When the task type is a generation task, an LLM-style decoder is used as the task decoder;
[0078] Step S303: When the task type is an object detection task, a Transformer decoder is used as the task decoder.
[0079] When mapping cross-modal semantics to task outputs, this embodiment can select a corresponding task decoder depending on the task type. For example, in a classification task scenario (such as scene selection), MLP can be selected as the task decoder, and the decoding process can specifically include Transformer head output → Dropout → Linear → Softmax, that is, the fusion feature f* is input into the MLP for classification prediction to obtain the classification result. In a generation task scenario (such as text generation), an LLM-style decoder can be selected, and the decoding process can perform autoregressive text generation based on the fusion feature f*, the input embedding comes from the Adapter output, and the decoding target is the answer or path. In a target detection task scenario (such as target object detection), a Transformer decoder can be selected, and the decoding process can include position embedding → Transformer decoder → detection head output, thereby obtaining the position and label information of the target object. In this way, the visual language model adaptation method of the embodiment of the present invention can flexibly adapt to different types of task output requirements, thereby improving the versatility and practicality of the model.
[0080] In general, the visual language model adaptation method provided in this embodiment has the following advantages:
[0081] (1) Significantly improves multi-scenario adaptability. Unlike traditional static fine-tuning methods, this embodiment effectively improves the recognition robustness of the model in various environments by introducing a scenario adaptation mechanism. Experiments show that in three types of environments: home, hospital, and retirement community, the same model can automatically adapt to tasks without retraining.
[0082] (2) Low training cost and minimal parameter updates. All lightweight parameters are inserted as adapter modules, which do not affect the original VLM structure and can be flexibly loaded or unloaded. Compared with full fine-tuning, the parameter update amount is reduced by more than 90%.
[0083] (3) Strong structural compatibility, adaptable to a variety of large models. It is suitable for a variety of visual language model structures such as BLIP, CLIP, MiniGPT, and LLaVA, and can be embedded in mobile deployment solutions.
[0084] (4) Deployment-friendly and high inference efficiency. Scene judgment can rely solely on images, and only some adapters are activated during the inference phase, eliminating the need for redundant calculations of the entire model, making it suitable for edge scenario deployment.
[0085] Figure 5 A schematic block diagram of a visual language model adaptation device 500 provided in an embodiment of the present invention, the device 500 includes:
[0086] A pre-training unit 501 is configured to obtain task data for multiple scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the multiple scenes using the visual language model;
[0087] A scene adaptation unit 502 is configured to use a plurality of lightweight adapters to inject scene-related parameters into each scene respectively, and fuse the parameters with the initial joint semantic features to obtain scene features;
[0088] A dynamic routing unit 503 is configured to predict scene weights using a dynamic routing mechanism, and obtain fused features based on the scene features and the scene weights;
[0089] A mapping decoding unit 504 is configured to map the fused features based on the scenario using cross-modal semantics, and decode and output the mapped fused features through a task decoder to construct a target adaptation model;
[0090] The transfer adaptation unit 505 is configured to perform transfer adaptation on the visual language model of the specified scene using the target adaptation model.
[0091] In this embodiment, combined with Figure 4First, we extract initial joint semantic features by acquiring task data from multiple scenarios and pre-training them using a visual language model. Next, we inject specific parameters into each scenario and fuse these parameters with the initial features using a lightweight adapter to obtain scene features. Then, we use a dynamic router to predict scene weights and combine them with the scene features to obtain the final fused features. Finally, cross-modal semantics performs cross-modal semantic mapping on the fused features, and the task decoder outputs the final result.
[0092] This embodiment implements a lightweight adaptation architecture based on the decoupling of tasks and scenes. The design goal of this architecture is to significantly enhance the migration capability of the visual language model (VLM) when facing tasks in a variety of different scenarios. By adopting this lightweight adaptation structure, the model can adapt to various task requirements more flexibly, thereby achieving more efficient and accurate performance in different scenarios. By decoupling the dependencies between tasks and scenes, this structure enables the model to quickly adjust its own parameters to adapt to new task requirements when processing tasks in new scenarios without the need for large-scale retraining. This lightweight design not only improves the generalization ability of the model, but also reduces the consumption of computing resources, making the model more efficient and practical in practical applications.
[0093] In actual application scenarios, the visual language model adaptation method provided by this embodiment can be applied to the medical field. For example, in an intelligent auxiliary diagnosis system, the visual language model can understand the doctor's language instructions, such as "show the abnormal area in the patient's lung X-ray", and accurately mark the lesion site in the image. In addition, the method can also be applied to the financial field. For example, in an intelligent banking service system, the visual language model can recognize the customer's language instructions, such as "check my account balance", and quickly display relevant information on the interface. Similarly, the method can also be applied to the smart home field, such as an intelligent voice assistant can understand the instructions of family members, such as "turn on the lights in the living room", and automatically perform the corresponding operations. Through the visual language model adaptation method provided by the embodiment of the present invention, the model can be flexibly adjusted according to the task requirements in different scenarios to achieve efficient and accurate performance, thereby greatly improving the practicality and adaptability of the visual language model in actual applications.
[0094] In one embodiment, if Figure 6 As shown, the pre-training unit 501 includes:
[0095] A visual encoding unit 601 is configured to encode the image data into image features using a visual encoder;
[0096] A language encoding unit 602 is configured to encode the text data into language features using a language encoder;
[0097] The first fusion unit 603 is configured to fuse the image features and language features into the initial joint semantic features.
[0098] In one embodiment, the first fusion unit 603 includes:
[0099] The second fusion unit is used to obtain the initial joint semantic feature by fusing according to the following formula:
[0100] f=VLM frozen (x,t)
[0101] Among them, f represents the initial joint semantic features, x represents image data, t represents text data, VLM represents the visual language model containing visual encoder and language encoder, and frozen represents parameter freezing.
[0102] This embodiment uses frozen pre-training to train the visual language encoder VLM, which freezes the model parameters during the training phase. The visual language encoder VLM then extracts the joint semantic features of the image and text as the basic representation. Mainstream models can use architectures such as CLIP, BLIP, or Flamingo, and their main components include:
[0103] Visual encoder (ViT / ResNet): input image x → output image features;
[0104] Language encoder (Transformer): input text t → output language features;
[0105] Fusion module: For example, the BLIP architecture uses Cross-Attention to fuse images and text to obtain a cross-modal representation f, which is the initial joint semantic feature.
[0106] In this way, it can be ensured that the model can retain the key semantic connections between images and text when processing multimodal data, providing a solid foundation for subsequent scene adaptation. The advantage of frozen pre-training is that it can flexibly adjust the model in different scenarios by introducing new task data and adaptation strategies without changing the pre-trained model parameters. This method not only improves the stability and reliability of the model, but also greatly shortens the time and cost of migrating the model between different tasks. At the same time, using the visual language model (VLM) as the basis for feature extraction can fully utilize the model's powerful capabilities in visual and language understanding, further improving the model's performance in complex scenarios.
[0107] In one embodiment, the scene adaptation unit 502 includes:
[0108] A residual fusion unit is used to fuse the initial joint semantic features and scene-related parameters into scene features using a residual module:
[0109] f i =f+Adapter i (f)
[0110] Among them, f i represents scene features, f represents initial joint semantic features, Adapter i (f) represents scene-related parameters, and i represents the i-th scene.
[0111] After extracting the initial joint semantic features, a lightweight Adapter module is inserted into the image-text fusion layer, and each scene has an independent Adapter to achieve the purpose of injecting a small amount of scene-related parameters for different task scenarios. This embodiment uses the classic residual BottleneckAdapter:
[0112] f i =f+Adapter i (f),i∈{family, community, institution...}
[0113] In practical applications, the initial joint semantic features can be down-projected, activated using GELU / ReLU activation functions, and then up-projected. The scene features are then obtained by summing the residuals. In this way, the features of each scene effectively combine the initial joint semantic features with the relevant parameters of the specific scene, enabling the model to better understand and handle task requirements in different scenarios. In this way, the lightweight adapter not only enhances the model's scene adaptability, but also maintains its lightweight nature, avoiding the increased computational burden caused by the introduction of too many parameters.
[0114] In one embodiment, the dynamic routing unit 503 includes:
[0115] The weight prediction unit is used to predict the scene weight using a lightweight multi-layer perceptron and fuse it to obtain the fusion feature according to the following formula:
[0116]
[0117] Among them, f * represents the fusion feature, α i represents the scene weight of the i-th scene.
[0118] This embodiment uses a dynamic routing mechanism to dynamically determine which weighted combination to use. The structure of the dynamic routing mechanism can use a lightweight MLP to predict the weights α = [α1, α2, ..., αi ], and then fused the predicted weights to obtain the fused features. In practical applications, a lightweight MLP can consist of multiple fully connected layers, using ReLU or GELU activation functions between layers. Finally, softmax normalization is performed to obtain the weights for each scene. This mechanism allows the model to dynamically adjust its focus in different scenarios, thereby enhancing its adaptability to new scenarios.
[0119] Furthermore, the introduction of a dynamic routing mechanism improves model interpretability. Because the weight of each scenario is predicted using a lightweight multi-layer perceptron, it's possible to intuitively understand the importance the model places on each scenario when handling tasks in different scenarios. This not only helps us better understand the model's decision-making process but also facilitates subsequent model optimization and debugging.
[0120] In practical applications, data augmentation strategies can be used to improve the generalization and robustness of models. Specifically, when acquiring task data from multiple scenarios, random transformations and perturbations are performed on image and text data, such as image rotation, scaling, cropping, and text synonym substitution, to increase data diversity and complexity. This allows the model to be exposed to more data variations during training, thereby learning more robust feature representations and improving its adaptability to new data.
[0121] In one embodiment, the mapping decoding unit 504 includes:
[0122] The task decoding unit is used to obtain the task type and select the corresponding task decoder according to the task type to decode and output the mapped fusion features.
[0123] Different task types may correspond to different output formats and requirements. For example, in an image recognition task, the task decoder may need to map fused features to specific category labels; in a text generation task, the task decoder may need to convert fused features into natural language text. Therefore, choosing the appropriate task decoder based on the task type is crucial.
[0124] In practice, a set of task decoders can be predefined, each optimized for one or more specific task types. When a new task is needed, the most appropriate decoder can be selected from the predefined set of decoders for decoding output based on the task type. This approach not only improves the flexibility and scalability of the model but also ensures optimal performance across different tasks.
[0125] Furthermore, to further improve model performance, a multi-task learning strategy can be employed. Specifically, during the training phase, the model can be trained using data from multiple tasks simultaneously. This allows the model to learn how to handle a single task while also learning the relevance and commonalities between different tasks. This allows the model to better utilize the complementary information between different tasks, thereby improving its performance on each task.
[0126] In one embodiment, if Figure 7 As shown, the task decoding unit includes:
[0127] A first setting unit 701 is configured to use a multilayer perceptron as the task decoder when the task type is a classification task;
[0128] The second setting unit 702 is configured to use an LLM-style decoder as the task decoder when the task type is a generation task;
[0129] The third setting unit 703 is configured to use a Transformer decoder as the task decoder when the task type is a target detection task.
[0130] When mapping cross-modal semantics to task outputs, this embodiment can select a corresponding task decoder depending on the task type. For example, in a classification task scenario (such as scene selection), MLP can be selected as the task decoder, and the decoding process can specifically include Transformer head output → Dropout → Linear → Softmax, that is, the fusion feature f* is input into the MLP for classification prediction to obtain the classification result. In a generation task scenario (such as text generation), an LLM-style decoder can be selected, and the decoding process can perform autoregressive text generation based on the fusion feature f*, the input embedding comes from the Adapter output, and the decoding target is the answer or path. In a target detection task scenario (such as target object detection), a Transformer decoder can be selected, and the decoding process can include position embedding → Transformer decoder → detection head output, thereby obtaining the position and label information of the target object. In this way, the visual language model adaptation method of the embodiment of the present invention can flexibly adapt to different types of task output requirements, thereby improving the versatility and practicality of the model.
[0131] In general, the visual language model adaptation method provided in this embodiment has the following advantages:
[0132] (1) Significantly improves multi-scenario adaptability. Unlike traditional static fine-tuning methods, this embodiment effectively improves the recognition robustness of the model in various environments by introducing a scenario adaptation mechanism. Experiments show that in three types of environments: home, hospital, and retirement community, the same model can automatically adapt to tasks without retraining.
[0133] (2) Low training cost and minimal parameter updates. All lightweight parameters are inserted as adapter modules, which do not affect the original VLM structure and can be flexibly loaded or unloaded. Compared with full fine-tuning, the parameter update amount is reduced by more than 90%.
[0134] (3) Strong structural compatibility, adaptable to a variety of large models. It is suitable for a variety of visual language model structures such as BLIP, CLIP, MiniGPT, and LLaVA, and can be embedded in mobile deployment solutions.
[0135] (4) Deployment-friendly and high inference efficiency. Scene judgment can rely solely on images, and only some adapters are activated during the inference phase, eliminating the need for redundant calculations of the entire model, making it suitable for edge scenario deployment.
[0136] See also Figure 8 , Figure 8 This is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device is a device with wireless communication and wired communication.
[0137] The computer device includes a processor 802 , a memory, and a network interface 805 connected via a system bus 801 , wherein the memory may include a non-volatile storage medium 803 and an internal memory 804 .
[0138] The non-volatile storage medium 803 may store an operating system 8031 and a computer program 8032. When the computer program 8032 is executed, the processor 802 may execute a visual language model adaptation method.
[0139] The processor 802 is used to provide computing and control capabilities to support the operation of the entire computer device.
[0140] The internal memory 804 provides an environment for the operation of the computer program 8032 in the non-volatile storage medium 803. When the computer program 8032 is executed by the processor 802, the processor 802 can execute a visual language model adaptation method.
[0141] The network interface 805 is used to communicate with other devices over the network. Figure 8The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0142] The processor 802 is configured to run a computer program 8032 stored in a memory to implement any embodiment of the above-mentioned visual language model adaptation method.
[0143] It should be understood that in the embodiment of the present invention, the processor 802 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0144] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0145] Acquire task data for multiple scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the multiple scenes using the visual language model;
[0146] Using multiple lightweight adapters to inject scene-related parameters into each scene respectively, and combining with the initial joint semantic features to obtain scene features;
[0147] A dynamic routing mechanism is used to predict scene weights, and a fusion feature is obtained based on the fusion of the scene features and the scene weights;
[0148] Based on the scenario, the fused features are mapped using cross-modal semantics, and the mapped fused features are decoded and output by a task decoder to construct a target adaptation model;
[0149] The target adaptation model is used to perform migration adaptation on the visual language model of the specified scene.
[0150] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiments. The storage medium can include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.
[0151] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0152] Acquire task data for multiple scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the multiple scenes using the visual language model;
[0153] Using multiple lightweight adapters to inject scene-related parameters into each scene respectively, and combining with the initial joint semantic features to obtain scene features;
[0154] A dynamic routing mechanism is used to predict scene weights, and a fusion feature is obtained based on the fusion of the scene features and the scene weights;
[0155] Based on the scenario, the fused features are mapped using cross-modal semantics, and the mapped fused features are decoded and output by a task decoder to construct a target adaptation model;
[0156] The target adaptation model is used to perform migration adaptation on the visual language model of the specified scene.
[0157] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant description in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0158] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0159] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0160] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A visual language model adaptation method, characterized in that: include: Acquire task data for multiple scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the multiple scenes using the visual language model; Using multiple lightweight adapters to inject scene-related parameters into each scene respectively, and combining with the initial joint semantic features to obtain scene features; A dynamic routing mechanism is used to predict scene weights, and a fusion feature is obtained based on the fusion of the scene features and the scene weights; Based on the scenario, the fused features are mapped using cross-modal semantics, and the mapped fused features are decoded and output by a task decoder to construct a target adaptation model; The target adaptation model is used to perform migration adaptation on the visual language model of the specified scene.
2. The visual language model adaptation method according to claim 1, characterized in that: The task data includes image data and text data; pre-training a visual language model using the task data, and extracting initial joint semantic features from the task data of multiple scenes using the visual language model, includes: Encoding the image data using a visual encoder and outputting it as image features; Encoding the text data into language features using a language encoder; The image features and language features are fused into the initial joint semantic features.
3. The visual language model adaptation method according to claim 2, characterized in that: The fusing of the image features and the language features into the initial joint semantic features includes: The initial joint semantic features are obtained by fusion according to the following formula: f=VLM frozen (x,t) Among them, f represents the initial joint semantic features, x represents image data, t represents text data, VLM represents the visual language model containing visual encoder and language encoder, and frozen represents parameter freezing.
4. The visual language model adaptation method according to claim 1, characterized in that The method of injecting scene-related parameters into each scene using multiple lightweight adapters and fusing the initial joint semantic features to obtain scene features includes: The residual module is used to fuse the initial joint semantic features and scene-related parameters into scene features: f i =f+Adapter i (f) Among them, f i represents scene features, f represents initial joint semantic features, Adapter i (f) represents scene-related parameters, and i represents the i-th scene.
5. The visual language model adaptation method according to claim 1, characterized in that: The method of using a dynamic routing mechanism to predict scene weights and obtaining fused features based on the scene features and the scene weights includes: The scene weights are predicted using a lightweight multi-layer perceptron and fused according to the following formula to obtain the fusion features: Among them, f * represents the fusion feature, α i represents the scene weight of the i-th scene.
6. The visual language model adaptation method according to claim 1, characterized in that: The method of mapping the fused features based on the scenario by using cross-modal semantics and decoding and outputting the mapped fused features by a task decoder includes: Obtain the task type, and select the corresponding task decoder according to the task type to decode and output the mapped fusion features.
7. The visual language model adaptation method according to claim 6, characterized in that: The acquiring of the task type and selecting a corresponding task decoder according to the task type to decode and output the mapped fusion features include: When the task type is a classification task, a multilayer perceptron is used as the task decoder; When the task type is a generation task, an LLM-style decoder is used as the task decoder; When the task type is an object detection task, a Transformer decoder is used as the task decoder.
8. A visual language model adaptation device, characterized in that: include: a pre-training unit, configured to obtain task data for a plurality of scenes, pre-train a visual language model using the task data, and extract initial joint semantic features from the task data for the plurality of scenes using the visual language model; A scene adaptation unit, configured to inject scene-related parameters into each scene using a plurality of lightweight adapters, and fuse the parameters with the initial joint semantic features to obtain scene features; A dynamic routing unit, configured to predict a scene weight using a dynamic routing mechanism, and obtain a fused feature based on the scene feature and the scene weight; A mapping and decoding unit is used to map the fused features based on the scenario using cross-modal semantics, and decode and output the mapped fused features through a task decoder to construct a target adaptation model; The migration and adaptation unit is used to use the target adaptation model to perform migration and adaptation on the visual language model of the specified scene.
9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the visual language model adaptation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the visual language model adaptation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Representation understanding method based on multi-scale cross-modal feature fusion
CN115496991A
Image processing method and device, equipment and medium
CN117711001A
Pre-training language model fine tuning method based on lightweight feedforward network adapter
CN118885558A
Personnel abnormal behavior detection method and system based on visual language large model
CN119992641A