Vision-language model training method and device and related equipment

By constructing a multimodal training dataset and employing a freeze-thaw strategy, the parameters of the visual encoder and the heterogeneous language model are dynamically aligned. This addresses the issues of computational resource consumption and dimensionality conflicts in the upgrade of existing visual-language model architectures, enabling a fast and efficient training process and improving the robustness and generalization ability of the model.

CN121010853APending Publication Date: 2025-11-25AISINO CORPORATION

Patent Information

Application Number
CN202511129253.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing visual-language model architectures suffer from huge computational resource consumption and architectural dimension conflicts during upgrades and maintenance, making it difficult to embed heterogeneous LLMs efficiently and seamlessly, and lacking effective training methods to improve training efficiency and retain pre-training capabilities.

Method used

By constructing a training dataset that includes tasks such as image-text alignment, text-driven visual localization, and pure text inference, we employ a two-stage training strategy of dimensional dynamic alignment and 'freeze-thaw' to train the parameters of the visual encoder and heterogeneous language model. We use bilinear interpolation to expand the weight matrix and hierarchical initialization to ensure parameter matching of the cross-modal alignment module.

Benefits of technology

It enables rapid training of visual-language models, reduces computational resource consumption, improves model convergence speed, retains the pre-training capability of heterogeneous language models, avoids catastrophic forgetting, and enhances the robustness and generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010853A_ABST
    Figure CN121010853A_ABST
Patent Text Reader

Abstract

The invention provides a visual-language model training method and device and related equipment, and the method comprises the steps: constructing a training data set comprising an image-text alignment task, a text-driven visual positioning task and a plain text reasoning task through determining a visual encoder and a heterogeneous language model which form a multi-modal heterogeneous recognition model; carrying out dimension dynamic alignment adaptation on the heterogeneous language model, enabling a parameter architecture of the heterogeneous language model to be matched with a visual feature dimension output by a visual encoder, and carrying out supervision and fine adjustment on the visual-language model by adopting a freezing-unfreezing two-stage training strategy based on a training data set; and performing parameter training on a cross-modal alignment module connected with the visual encoder and the heterogeneous language model so as to obtain a trained visual-language model. According to the training method, time consumed by model training is saved, and the convergence speed is increased. Through dimension dynamic alignment and hierarchical weight mapping, the pre-training language ability is reserved to the maximum extent, the plain text task performance loss is reduced, and disastrous forgetting of the model is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence (AI), and in particular to a method and device for training a vision-language model and related equipment. BACKGROUND

[0002] With the rapid development of deep learning technology, especially the rise of the Transformer architecture, multi-modal heterogeneous recognition models have become a frontier hotspot in the field of artificial intelligence. These models integrate information from different modalities (such as vision, language, and speech) to achieve higher-level perception, understanding, and reasoning capabilities. As an important branch of multi-modal large models, vision-language models (VLMs) can handle complex applications such as image-text understanding, video analysis, and embodied intelligence by combining computer vision and natural language processing capabilities. The mainstream architecture of existing VLMs typically adopts a cascaded architecture of "vision encoder (such as ViT) + language model (LLM) + cross-modal alignment module (Aligner)". In this architecture, the vision encoder is responsible for extracting visual features from images or videos, the language model is responsible for processing text information and performing language reasoning, and the cross-modal alignment module acts as a bridge to map visual features to the semantic space of the language model to enable joint reasoning of the two modalities. Although existing VLM architectures have made significant progress, they still have some inherent defects in practical applications, especially in terms of model upgrading and maintenance. For example, breaking through the capability bottleneck of existing large language models requires huge computational resources and training periods, or due to the obvious architectural dimension conflicts between heterogeneous models, it is difficult to directly align the vision encoder and the language model. There are significant deficiencies in efficiently and seamlessly embedding heterogeneous LLMs into existing VLM architectures. There is a lack of a general solution that can effectively handle model architecture differences, suppress catastrophic forgetting, and significantly improve training efficiency. Therefore, how to provide a fast, efficient, and economical training method for vision-language models to meet the needs of the rapid development of multi-modal large models has become an important problem that needs to be solved in the industry. SUMMARY

[0003] In view of the above, the embodiments of the present application provide a training method, device, storage medium and electronic equipment for a vision-language model to at least or partially solve the above problems.

[0004] In a first aspect, the embodiments of the present application provide a training method for a vision-language model, comprising:

[0005] determining a vision encoder (ViT) and a heterogeneous language model (LLM) that constitute the vision-language model;

[0006] constructing a training data set containing an image-text alignment task, a text-driven visual positioning task, and a pure text reasoning task;

[0007] dimensionally dynamically aligning the heterogeneous language model to match a parameter architecture of the heterogeneous language model with a visual feature dimension output by the visual encoder and to retain pre-training language capabilities of the heterogeneous language model;

[0008] using the training data set, performing supervised fine-tuning of the visual- language model using a “freeze-thaw” two-stage training strategy to perform parameter training of a cross-modal alignment module (Aligner) connecting the visual encoder and the heterogeneous language model in the visual-language model, so that the visual feature dimension output by the visual encoder can be converted by the parameter-trained cross-modal alignment module to adapt to the input feature dimension of the heterogeneous language model, thereby obtaining a trained visual-language model.

[0009] Optionally, in an embodiment of the present application, the dimensionally dynamically aligning the heterogeneous language model comprises:

[0010] performing bilinear interpolation expansion of a weight matrix for at least one linear layer of the heterogeneous language model, and modifying a hidden layer dimension (hidden_size) of the heterogeneous language model to match the visual feature dimension output by the visual encoder;

[0011] performing layer number padding on the heterogeneous language model according to a difference in layer numbers between the heterogeneous language model and an original language model, wherein the newly added layers use cyclic initialization of weights of the original language model and add progressive noise;

[0012] performing initialization of an embedding layer of the heterogeneous language model according to a difference in a vocabulary size (vocab_size) between the heterogeneous language model and the original language model, and performing dropout processing and corresponding addition or adjustment of a head dimension (head_dim) parameter.

[0013] Optionally, in an embodiment of the present application, the supervised fine-tuning of the visual-language model using the training data set using the “freeze-thaw” two-stage training strategy comprises:

[0014] when using the training data set, at least the following two stages are divided:

[0015] first stage freeze training: freezing model parameters of the visual encoder and the heterogeneous language model to obtain predicted data corresponding to input training data; introducing cross-entropy loss (CE Loss) to calculate a first difference between the predicted data and corresponding real labels, adjusting parameters of the cross-modal alignment module based on the first difference, until the output dimension of the cross-modal alignment module matches the input dimension of the model of the heterogeneous language model;

[0016] Second stage joint optimization: the frozen state of the visual encoder, the heterogeneous language model and the cross-modal alignment module is released; using a hierarchical learning rate strategy, the visual encoder, the heterogeneous language model and the cross-modal alignment module are respectively set to different learning rates, and an alignment loss (Alignment Loss) is introduced as a regularization term on the basis of introducing the cross-entropy loss (CELoss), to calculate the second difference between the predicted probability distribution of the training data set input into the visual-language model and the real label distribution; according to the second difference and the differential learning rate, the parameters of the visual encoder, the heterogeneous language model and the cross-modal alignment module are adjusted, so as to keep the consistency of visual features and text features in the embedding space.

[0017] Optionally, in an embodiment of the present application, in the second stage joint optimization, the learning rate of the cross-modal alignment module is set to ten times or more than ten times the learning rate of the heterogeneous language model.

[0018] Optionally, in an embodiment of the present application, in the second stage joint optimization, the first 2 layers of the visual encoder are kept frozen, and the parameters of the remaining high layers are fine-tuned.

[0019] Optionally, in an embodiment of the present application, the pre-query template is used to identify the role of the user instruction provider, and the post-query template is used to identify the role of the user instruction executor.

[0020] Optionally, in an embodiment of the present application, the alignment loss (Lalign) is calculated as follows:

[0021] Lalign=1-cos(visual_embed,text_embed),

[0022] Wherein, visual_embed is the visual feature embedding output by the visual encoder, and text_embed is the text feature embedding processed by the heterogeneous language model.

[0023] In a second aspect, based on the training method of the visual-language model of the first aspect of the present application, an embodiment of the present application further provides a training device of a visual-language model, comprising:

[0024] An acquisition module is configured to determine a visual encoder (ViT) and a heterogeneous language model (LLM) constituting the visual-language model.

[0025] A construction module is configured to construct a training data set comprising a picture-text alignment task, a text-driven visual positioning task and a pure text reasoning task.

[0026] an alignment module configured to perform dynamic dimension alignment adaptation on the heterogeneous language model so that a parameter architecture of the heterogeneous language model matches a visual feature dimension output by the visual encoder, and the heterogeneous language model retains pre-training language capabilities thereof;

[0027] an adjustment module configured to perform supervised fine-tuning on the visual-language model using the training data set by adopting a "freeze-thaw" two-stage training strategy to perform parameter training on a cross-modal alignment module (Aligner) connecting the visual encoder and the heterogeneous language model in the visual-language model, so that the visual feature dimension output by the visual encoder can be converted by the parameter-trained cross-modal alignment module into an input feature dimension adapted to the heterogeneous language model, thereby obtaining a trained visual-language model.

[0028] In a third aspect, the embodiments of the present application further provide a computer storage medium, and the computer storage medium stores computer executable instructions. When the computer executable instructions are executed, any one of the training methods of the visual-language model according to the first aspect of the embodiments of the present application is performed.

[0029] In a fourth aspect, the embodiments of the present application further provide an electronic device, which comprises a processor, a memory, a communication interface and a communication bus. The processor, the memory and the communication interface complete communication with each other through the communication bus.

[0030] The memory is configured to store at least one executable instruction. The executable instruction causes the processor to perform any one of the training methods of the visual-language model according to the first aspect of the embodiments of the present application.

[0031] This application provides a training method, apparatus, and related equipment for a vision-language model. The method involves identifying the visual encoder (ViT) and heterogeneous language model (LLM) constituting the multimodal heterogeneous recognition model; constructing a training dataset including image-text alignment tasks, text-driven visual localization tasks, and pure text inference tasks; dynamically aligning and adapting the heterogeneous language model to match its parameter architecture with the visual feature dimensions output by the visual encoder, while preserving the pre-trained language capabilities of the heterogeneous language model; and using a two-stage "freeze-thaw" training strategy to supervise and fine-tune the vision-language model based on the training dataset. This involves training the parameters of the cross-modal alignment module (Aligner) connecting the visual encoder and the heterogeneous language model, enabling the visual feature dimensions output by the visual encoder to be converted by the trained cross-modal alignment module to adapt to the input feature dimensions of the heterogeneous language model, thereby obtaining a trained vision-language model. This training method saves model training time and improves convergence speed. By dynamically aligning dimensions and mapping hierarchical weights, the pre-trained language capabilities are preserved to the greatest extent, reducing performance loss in pure text tasks and avoiding catastrophic forgetting. This approach effectively overcomes the shortcomings of existing technologies in efficiently and seamlessly embedding heterogeneous LLMs into existing VLM architectures. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings.

[0033] Figure 1 A schematic diagram illustrating the workflow of a training method for a visual-language model provided in this application embodiment;

[0034] Figure 2 This is a schematic diagram of the structure of a training device for a visual-language model provided in an embodiment of this application.

[0035] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0036] In the specific embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art shall fall within the scope of protection of the embodiments of the present application.

[0037] It should be understood that each step described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0038] Embodiment one,

[0039] The embodiments of the present application provide a training method of a visual-language model, as shown in Figure 1 Figure 1 The workflow schematic diagram of the training method of the visual-language model provided by the embodiments of the present application includes:

[0040] Step S101, determining a visual encoder (ViT) and a heterogeneous language model (LLM) constituting the visual-language model. In the embodiments of the present application, the visual encoder is used to divide the input image into a series of fixed-size image patches, and each image patch is regarded as a “word”, and the relationship between the image patches is learned through a self-attention mechanism, so as to extract the global features of the image. It generally adopts a Transformer architecture to extract visual features through a self-attention mechanism, which can capture the relationship between elements at any position in the sequence, and solve the limitations of traditional recurrent neural networks (RNN) and convolutional neural networks (CNN) in long-distance dependence problems. The deep learning model trained based on large-scale text data can understand, generate and process human language. The heterogeneous language model also generally adopts a Transformer architecture, has a large number of parameters and strong language understanding and generation capabilities. For example, the Qwen series models (such as Qwen2.5-7B and Qwen3-8B) are currently popular LLMs, which perform well in text summarization, translation, question answering and other tasks. In this stage, the visual encoder and the heterogeneous language model constituting the visual-language model are determined, mainly to determine whether the visual feature dimension of the output of the visual encoder and the parameter architecture of the heterogeneous language model match, and only when the two do not match, the subsequent targeted processing process needs to be performed. In order to avoid the case of misoperation, which would otherwise affect the model processing capability of the visual-language model.

[0041] ​In the embodiments of the present application, due to the higher requirements of users on the performance of the model, developers are constantly seeking to integrate more powerful and updated LLMs into existing VLMs, which requires solving the architecture dimension conflict. For example, when integrating a new language model, the hidden_size of the model increases from the old 3584 to 4096, and the number of layers increases from the old 28 layers to 36 layers. This difference between the new and old architectures makes it impossible for a simple plug-and-play. However, the method described in the embodiments of the present application can better cope with this new-old work requirement.

[0042] Step S102, constructing a training data set containing a picture-text alignment task, a text-driven visual positioning task and a pure text reasoning task. In the embodiments of the present application, high-quality training data sets are required in the training process of the recognition model. Specifically, in an optional implementation manner of the embodiments of the present application, the training data set should at least contain about 100,000 data. The design of the data set aims to ensure that the visual dominant task and the language dominant task can be activated equally in joint training, so that the visual module and the language module can be fully optimized. This balanced activation is crucial for the robustness of the multi-modal model, which helps to avoid the model from excessively biasing to a certain modality during the training process, thereby leading to performance decline on another modality task.

[0043] Preferably, 50% of the data in the training data set can be set as graphic-text alignment task data (invoice recognition), 30% of the data as text-driven visual positioning task data, and 20% of the data as pure text reasoning data. Among them, the graphic-text alignment task data is the key to training the core application scenario recognition of the training process of the embodiments of the present application. It contains images (such as value-added tax invoices, supermarket receipts, bank return slip water, general receipts) and their corresponding text information (such as recognized fields, receipt categories, etc.). This ensures that the model can learn the accurate correspondence between visual information and text information in the image, such as identifying the amount, date, and purchaser on the invoice. This kind of data directly trains the model to establish a strong correlation between vision and language, which is the basis for accurate receipt recognition. The text-driven visual positioning data is used to train the model's ability to locate specific visual elements in the image according to the text instructions. For example, given the instruction "find all red objects in the image", the model needs to identify the red objects in the image and indicate their location. This helps to enhance the model's understanding of visual details and precise visual positioning capabilities, which is crucial for locating specific fields (such as invoice number, tax amount location) in receipt recognition. In this way, the model not only understands the text, but also associates the text instructions with specific areas in the image. Pure text reasoning data is used to train the pure language understanding and reasoning capabilities of the LLM, ensuring that the performance of the new LLM on language tasks will not be compromised after integration. For example, text summarization, logical question answering, mathematical calculation, etc. The introduction of this kind of pure text data can effectively prevent the model from "modality collapse" or "language ability degradation" during multi-modal training, ensuring that the new LLM can still maintain its strong pre-training capabilities when processing pure language tasks.

[0044] The above graphic-text alignment data used in the training process of the embodiments of the present application, for example, if the model is used for tax invoice and other image recognition and feedback, the data in the high-quality (alignment) training data set covers value-added tax invoices, supermarket receipts, bank return slip water, general receipts, etc., ensuring the diversity of the distribution of training data. This diversity enables the trained model to have stronger generalization ability, thereby widely adapting to different types and formats of receipt recognition needs, thereby exhibiting higher robustness and accuracy in actual application. This strategic training data set design is one of the key factors that enable the method described in the embodiments of the present application to have a significant training effect, which ensures that the model not only integrates new LLM, but also performs well on various multi-modal and single-modal tasks, which is crucial for practical deployment in complex scenarios such as financial document processing.

[0045] Step S103, performing dimension dynamic alignment adaptation on the heterogeneous language model to match the parameter architecture thereof with the visual feature dimension output by the visual encoder and retain the pre-training language ability of the heterogeneous language model. In this stage, the embodiments of the present application match the visual feature dimension output by the visual encoder with the new generation of language model through dynamic alignment adaptation, so as to solve the architecture conflict between different generations of language models (such as Qwen3-8B and Qwen2.5-VL-7B), especially the mismatch of key parameters such as hidden_size and num_hidden_layers. Without retraining the entire model, automatic matching of the visual encoder output and the input dimension of the new LLM is realized.

[0046] Specifically, in an optional implementation of the embodiments of the present application, the dimension dynamic alignment adaptation on the heterogeneous language model comprises: performing bilinear interpolation expansion weight matrix on at least one linear layer of the heterogeneous language model, modifying the hidden layer dimension (hidden_size) of the heterogeneous language model to match the visual feature dimension output by the visual encoder; performing layer number padding on the heterogeneous language model according to the layer number difference between the heterogeneous language model and the original language model, wherein the newly added layer adopts cyclic initialization weight of the original language model and adds progressive noise; performing initialization discard processing on the embedding layer of the heterogeneous language model according to the vocabulary size (vocab_size) difference between the heterogeneous language model and the original language model, and correspondingly adding or adjusting the attention head dimension (head_dim) parameter.

[0047] The operation process described in the above steps is exemplarily illustrated herein: for the key linear layers in the Transformer architecture language model, including q_proj (query projection), k_proj (key projection), v_proj (value projection), and o_proj (output projection), perform Bilinear Interpolation to expand the weight matrix, and when the hidden_size of the language model increases from 3584 of the original model to 4096 of the new model, the weight matrix of these linear layers needs to be expanded from the original input / output dimension (e.g., 3584) to the new dimension (e.g., 4096). Bilinear Interpolation is creatively applied to the expansion of the weight matrix here, rather than traditional image processing. It generates new weight values smoothly by weighted averaging adjacent weights in the original weight matrix, thereby retaining the knowledge and feature representation contained in the original pre-training weight to the greatest extent while expanding the dimension. When the number of layers (num_hidden_layers) of the new LLM increases from 28 layers of the original model to 36 layers of the new model, the hierarchical padding strategy is adopted at this stage. Among them, the weights of layers 0-27 of Qwen3-8B are directly copied to the corresponding positions of the new model. This maximizes the powerful language capabilities and feature representations that Qwen3-8B has learned in these layers. For the newly added 8 layers (i.e., layers 28 to 35), the cyclic initialization strategy is adopted, and the weights of layers 20-27 of Qwen3-8B are reused for initialization. On this basis, progressive noise is added. This strategy effectively adapts the depth of the new LLM to the VLM architecture, enabling it to utilize deeper language understanding capabilities. By directly copying most of the pre-training layer weights, the huge cost of training these layers from scratch is avoided, and the language foundation capabilities of the new LLM are fully inherited. Cyclic initialization provides a meaningful initial state instead of completely random initialization, which helps to accelerate the convergence of new layers. The introduction of progressive noise helps to break the repetition pattern, increase the robustness and generalization ability of the model, prevent the model from falling into local optimum, and provide a better starting point for subsequent fine-tuning. Compared with completely retraining all layers, this hierarchical copying and strategic initialization method significantly reduces the computational resources and time consumption of training. In addition, since the vocab_size (vocabulary size) of Qwen3-8B may differ from the original model (for example, the vocab_size of Qwen3-8B is 128 fewer than the original model), the last 128 tokens of the original model are discarded during initialization. At the same time, a new parameter head_dim is added and set to 128.Since the difference in vocabulary size is also an important reason affecting the integration of heterogeneous LLMs, in this stage, the application ensures that the new LLM can use its own vocabulary by strategically discarding unmatched word embeddings, avoiding input parsing errors or performance degradation caused by vocabulary mismatches. head_dim is the dimension of each attention head in the Transformer model. Adding or adjusting this parameter ensures that the attention mechanism can correctly process and project features, consistent with the new hidden_size (hidden_size=num_attention_heads*head_dim), thereby ensuring the correctness and efficiency of internal model calculations. The embedding layer is the entrance of the LLM, and its correct processing performance is crucial for the stable operation of the entire model. The precise adjustment of the embedding layer in the embodiments of the application ensures that the features output by the visual encoder can be correctly understood and processed by the LLM, thereby achieving effective fusion of visual and language modalities.

[0048] In neural networks, especially in Transformer architecture neural network models, hidden_size (or d_model) refers to the dimension of the model's internal hidden layer representation, i.e., the length of the feature vector corresponding to each token after the model processes it. It determines the richness of information that the model can capture and represent. For example, the hidden_size of Qwen2.5-7B is 3584, while the hidden_size of Qwen3-8B is 4096. q_proj, k_proj, v_proj, and o_proj are the core linear projection layers in the Self-Attention mechanism of the Transformer model. q_proj linearly transforms the input features into Query (Q); k_proj linearly transforms the input features into Key (K); v_proj linearly transforms the input features into Value (V). The Self-Attention mechanism determines the correlation between different tokens (attention scores) by calculating the dot product of Q and K, and then weights and sums V to get the attention output. o_proj concatenates the outputs of all attention heads and then linearly transforms them to get the final output in the Multi-Head Attention mechanism. Bilinear Interpolation is a method of function interpolation on a two-dimensional grid, commonly used in image processing for scaling operations. It calculates the value of a new pixel by weighted averaging the four neighboring pixels in a 2x2 region around the target point, generating a smooth image. In the embodiments of the present application, it is applied to the expansion of neural network weight matrices, i.e., by combining the existing elements in the original weight matrix, the weight values in the new dimension are generated to achieve smooth and meaningful weight expansion, rather than simple zero padding or random initialization. The method described in the embodiments of the present application can directly solve the mismatch between the output feature dimension of the visual encoder (e.g., 3584) and the input dimension of the new LLM (e.g., 4096). Thus, it effectively overcomes the information loss or noise introduced by traditional simple truncation or zero padding. Bilinear interpolation ensures that the expanded weight matrix can inherit the patterns and knowledge in the original pre-trained weights as much as possible due to its smooth interpolation characteristics. The method described in the embodiments of the present application effectively avoids the problem of unstable training or longer convergence time caused by random initialization of new part of the weight, thereby reducing the damage to the original ability of the model when adapting to the new dimension. Through this smooth weight expansion, the continuity of the model in the feature space is maintained, reducing the feature space deviation caused by dimension change, and to some extent, suppressing the possibility of catastrophic forgetting of the model.`num_hidden_layers` represents the number of hidden layers in the decoder of a Transformer architecture neural network model, signifying the model's depth. More layers typically allow the model to learn more complex and higher-level feature representations. For example, Qwen2.5-7B has 28 layers, while Qwen3-8B has 36. Cyclic initialization, a weight initialization strategy, is particularly suitable for model expansion (e.g., adding layers). It initializes new layers by reusing the weights of existing, well-performing layers in the model. For example, if the model needs to expand from 28 to 36 layers, the weights of several well-performing layers in the original model (e.g., layers 20-27) are copied and cyclically used to initialize the 8 new layers, providing a better starting point than random initialization. Progressive noise, on the other hand, is a random perturbation introduced gradually or at specific stages during model training or initialization. This noise helps prevent the model from overfitting the training data, increases the model's robustness, and encourages the model to explore a wider parameter space, potentially finding better local or global optima. For example, adding a small amount of random noise to new layers after cyclic initialization avoids all repeated layers having the exact same initial state, thus promoting the model's learning of richer features. Vocabulary size refers to the number of unique tokens the model can process. Each token has a unique ID in the vocabulary and corresponds to an embedding vector. For example, a model with a vocab size of 151936 can process 151936 different tokens. A token is the smallest semantic unit of text and can be a word, subword, or character. For example, the sentence "ticket recognition is an AI application" might be broken down into multiple tokens such as "ticket," "recognition," "is," "AI," and "application." The model understands its semantics by learning the embedding vectors of these tokens. The embedding layer is responsible for mapping discrete inputs (such as token IDs) to a continuous, low-dimensional vector space. These vectors are called word embeddings, and they capture the semantic and syntactic relationships between tokens. For example, in LLM, each input token is first converted into a vector representation by the embedding layer, and this vector will serve as the input to the subsequent Transformer layer. `head_dim` refers to the dimension of the attention head in the neural network model. In multi-head attention mechanisms, the model's `hidden_size` is divided into multiple subspaces of size `head_dim`, each corresponding to one attention head. The setting of `head_dim` affects the granularity and capacity of information processed by each attention head.

[0049] Step S104, using the training data set, adopting a "freeze-thaw" two-stage training strategy to supervise fine-tuning of the visual-language model, to perform parameter training on a cross-modal alignment module (Aligner) connecting the visual encoder and the heterogeneous language model in the visual-language model, so that the visual feature dimension output by the visual encoder can be converted by the parameter-trained cross-modal alignment module to adapt to the input feature dimension of the heterogeneous language model, thereby obtaining a trained visual-language model. The embodiment of the present application can effectively repair the misplacement of visual-language feature space caused by the dimension upgrade of the language model, so that the Aligner output is accurately matched with the input dimension of the new LLM. Specifically, the "freeze-thaw" two-stage training strategy can greatly reduce the amount of parameters that need to be optimized during the training of the Aligner, and accelerate the convergence speed. At the same time, the parameters of ViT and LLM are frozen to ensure that their respective pre-training capabilities (visual feature extraction and language understanding) are not damaged in the initial alignment stage. This way, the model can quickly and focusedly learn the mapping relationship of visual features to the semantic space of the new LLM, without the need to update the large and pre-trained ViT and LLM parameters during the implementation process, which can greatly compress the actual training time to 5% of the full parameter fine-tuning time, greatly improving the training efficiency and reducing the consumption of computing resources.

[0050] Specifically, in an optional implementation of the embodiment of the present application, the supervised fine-tuning of the visual-language model based on the training data set using the "freeze-thaw" two-stage training strategy comprises:

[0051] When training using the training data set, at least the following two stages are divided:

[0052] First stage freeze training: freeze the model parameters of the visual encoder and the heterogeneous language model to obtain predicted data corresponding to the training data input; introduce cross-entropy loss (CE Loss) to calculate the first difference between the predicted data and the corresponding true label, adjust the parameters of the cross-modal alignment module based on the first difference, until the output dimension of the cross-modal alignment module matches the input dimension of the model of the heterogeneous language model;

[0053] Second stage joint optimization: release the frozen state of the visual encoder, the heterogeneous language model and the cross-modal alignment module, set different learning rates for the visual encoder, the heterogeneous language model and the cross-modal alignment module respectively by using a hierarchical learning rate strategy, introduce an alignment loss (Alignment Loss) as a regular term on the basis of introducing the cross-entropy loss (CE Loss), calculate the second difference between the predicted probability distribution of the training data set input into the visual-language model and the real label distribution, and adjust the parameters of the visual encoder, the heterogeneous language model and the cross-modal alignment module according to the second difference and the different learning rates, so that the visual-language model can learn new tasks while keeping the consistency of visual features and text features in the embedding space.

[0054] In the above steps of the embodiments of the present application, the regular term is an additional term added to the loss function, which is used to punish the complexity of the model, thereby preventing the overfitting phenomenon and improving the generalization ability of the model. Overfitting refers to the phenomenon that the model performs well on the training data but poorly on new data. The alignment loss is used as a regular term to ensure that the model maintains the consistency of cross-modal features while optimizing the performance of the main task, thereby improving the robustness of the model. By introducing the alignment loss as a regular term, the model is forced to maintain the consistency of visual features and text features in the embedding space while learning new tasks (bill recognition). This effectively alleviates the catastrophic forgetting problem and ensures that the model can still maintain the understanding ability of the original task and modality after upgrading. The hierarchical learning rate strategy allows each part of the model to be optimized according to its characteristics and task requirements. For example, the fine-tuning of the ViT high layer and the LLM non-embedding layer enables the entire VLM to better adapt to the specific semantics and visual patterns of bill recognition, thereby improving the overall performance. The fast learning rate of the Aligner ensures that the bridge between the modalities is always in the optimal state. The alignment loss directly optimizes the cosine similarity between the visual embedding and the text embedding, so that the matched image-text pairs are closer in the shared embedding space and more consistent in direction. This enhances the cross-modal understanding ability of the model, enabling it to more accurately associate image content and text descriptions. The introduction of the regular term not only prevents overfitting, but also improves the generalization ability of the model on unseen data by encouraging the consistency of features

[0055] In the above steps of the embodiments of the present application, during the first stage freezing training process, freezing some modules or some parameters of the model means that the gradients of these parameters are not calculated during backpropagation, and their weights are not updated. This method is used to preserve the capabilities of the pre-trained model or to focus on optimizing specific modules in multi-stage training. Specifically, freezing ViT in VLM training can ensure that the visual features extracted by it remain stable and do not change dramatically due to the introduction of a new LLM. Only the cross-entropy loss (CE Loss) is calculated to measure the difference between the model's predicted probability distribution and the true label distribution to guide the Aligner to learn how to correctly map visual features to the input space of the LLM to support subsequent text generation or classification tasks. The commonly used loss function in machine learning is particularly suitable for classification tasks. It measures the difference between the probability distribution predicted by the model and the probability distribution of the true label. The smaller the cross-entropy value, the closer the model's prediction is to the true label. For example, in invoice recognition, if the model predicts that a picture is a "value-added tax invoice" with a probability of 0.9, and the true label is also "value-added tax invoice", then the cross-entropy loss will be small, indicating accurate prediction. The second stage joint optimization removes the frozen state of ViT, LLM, and Aligner, allowing all their parameters to participate in gradient update, while introducing different learning rates to make different parts of the model learn at different speeds. Thus, based on the alignment in the first stage, the visual and language modules are optimized in coordination, further improving the model's performance on multi-modal tasks, while further effectively suppressing catastrophic forgetting, i.e., by using alignment loss and hierarchical learning rate strategies, the model actively avoids forgetting old knowledge when adapting to new LLM and new tasks.

[0056] Optionally, in an embodiment of the present application, in the second stage joint optimization, the learning rate of the cross-modal alignment module is set to be ten times or more than the learning rate of the heterogeneous language model. This ensures that the parameters of the Aligner can be updated at a higher speed, allowing it to adapt to the feature changes between modalities more quickly. Further helping the model training process to converge faster and more stably.

[0057] Optionally, in an embodiment of the present application, in the second stage joint optimization, the first 2 layers of the visual encoder are kept frozen, and the parameters of the remaining high layers are fine-tuned. Continuing to freeze the first 2 layers can preserve the stability of their low-level visual feature extraction, which is generally task-independent. The remaining high layers of ViT (e.g., layer 3 and above) are fine-tuned to enable them to better adapt to the semantic space of the new LLM and learn visual features more relevant to the language modality.

[0058] Optionally, in an implementation form of the embodiment of the application, the calculation formula of the alignment loss (Lalign) can be set as:

[0059] Lalign = 1 - cos (visual_embed, text_embed) ;

[0060] and the alignment loss Lalign is taken as a regularization term, and is weighted and summed with the cross-entropy loss (L CE, Cross-Entropy Loss), and the total loss (Ltotal) is calculated as follows:

[0061] Ltotal = L CE + λ · Lalign

[0062] Wherein, visual_embed is the visual feature embedding output by the visual encoder (ViT), text_embed is the text feature embedding processed by the heterogeneous language model (LLM). λ is a hyperparameter, which can be determined as 0.05 through grid search. This value aims to balance the cross-entropy loss (task performance) and the alignment loss (cross-modal consistency) ;

[0063] The second difference is determined according to the total loss value Ltotal, which is used to adjust the parameters of the visual encoder, the heterogeneous language model and the cross-modal alignment module, so as to ensure that the cross-modal alignment effect is significantly improved without sacrificing the language task performance.

[0064] In the practical application scenario of the embodiment of the application, in the loss function containing the regularization term, lambda (Lambda) is a hyperparameter used to control the weight or intensity of the regularization term. It balances the importance of the main loss (such as the cross-entropy loss) and the regularization loss. By adjusting lambda, the trade-off between task performance and generalization ability (or cross-modal alignment effect here) of the model can be controlled. The application determines lambda = 0.05 through grid search, and this value is proved by experiments to be able to significantly improve the cross-modal alignment effect without sacrificing the language task performance.

[0065] The application provides a visual-language model training method, including: based on a preset instruction template of an alignment large language model, constructing a corresponding pre-query template and a post-query template; inputting the pre-query template into the preset alignment large language model to generate a user instruction based on the self-recurrence of the language model; inputting the user instruction into the preset alignment large language model again, combining the pre-query template and the post-query template to generate corresponding response data; combining the pre-query template, the post-query template, the user instruction and the corresponding response data to form a data set containing at least one "instruction-response" data pair as a target data set. The entire working process is an automatic process without human intervention, does not depend on predefined seed problems or prompt ranges, and only needs to input the pre-query template to automatically generate instructions and responses by using the self-recurrence characteristics of the large language model, without human annotation or prompt engineering throughout, thereby greatly reducing the labor cost. Therefore, the efficiency of obtaining a visual-language model meeting the training requirements of the model is significantly improved.

[0066] Embodiment two,

[0067] Based on the visual-language model training method provided in embodiment one of the application, a corresponding visual-language model training device is further provided herein, as shown in Figure 2 Figure 2 A structural schematic diagram of a visual-language model training device 20 provided in the application is shown in the figure, and the visual-language model training device 20 includes:

[0068] An acquisition module 201 is configured to determine a visual encoder (ViT) and a heterogeneous language model (LLM) constituting the visual-language model;

[0069] A construction module 202 is configured to construct a training data set containing a picture-text alignment task, a text-driven visual positioning task and a pure text reasoning task;

[0070] An alignment module 203 is configured to perform dynamic dimension alignment adaptation on the heterogeneous language model to match the parameter architecture of the heterogeneous language model with the visual feature dimension output by the visual encoder, and retain the pre-training language ability of the heterogeneous language model;

[0071] An adjustment module 204 is configured to use the training data set to perform supervised fine-tuning on the visual-language model by using a "freeze-thaw" two-stage training strategy to perform parameter training on a cross-modal alignment module (Aligner) connecting the visual encoder and the heterogeneous language model in the visual-language model, so that the visual feature dimension output by the visual encoder can be converted into an input feature dimension adapted to the heterogeneous language model by the cross-modal alignment module after parameter training, thereby obtaining a trained visual-language model.​

[0072] Optionally, in an implementation form of the embodiment of the application, the alignment module 203 is further configured to perform bilinear interpolation expansion weight matrix for at least one linear layer of the heterogeneous language model, modify the hidden layer dimension (hidden_size) of the heterogeneous language model to match the visual feature dimension output by the visual encoder, perform layer number padding on the heterogeneous language model according to the difference in the number of layers between the heterogeneous language model and the original language model, wherein the newly added layer adopts cyclic initialization weight of the original language model and adds progressive noise, and perform dropout processing on the initialization of the embedding layer of the heterogeneous language model according to the difference in the vocabulary size (vocab_size) between the heterogeneous language model and the original language model, and correspondingly add or adjust the head dimension (head_dim) parameter.

[0073] Optionally, in an implementation form of the embodiment of the application, the adjustment module 204 is further configured to perform supervised fine-tuning on the visual-language model based on the training data set by adopting a "freeze-thaw" two-stage training strategy, and at least perform the following two stages of implementation process:

[0074] First stage freeze training: freeze the model parameters of the visual encoder and the heterogeneous language model to obtain predicted data corresponding to the input training data; introduce cross-entropy loss (CE Loss) to calculate the first difference between the predicted data and the corresponding true label, adjust the parameters of the cross-modal alignment module based on the first difference, until the output dimension of the cross-modal alignment module matches the input dimension of the model of the heterogeneous language model;

[0075] Second stage joint optimization: the second stage joint optimization: remove the freeze state of the visual encoder, the heterogeneous language model and the cross-modal alignment module; use a hierarchical learning rate strategy to set different learning rates for the visual encoder, the heterogeneous language model and the cross-modal alignment module, and introduce an alignment loss (Alignment Loss) as a regularization term based on the cross-entropy loss (CE Loss), to calculate the second difference between the predicted probability distribution of the training data set input into the visual-language model and the true label distribution; adjust the parameters of the visual encoder, the heterogeneous language model and the cross-modal alignment module according to the second difference and the different learning rates, so as to keep the consistency of the visual features and the text features in the embedding space.

[0076] Optionally, in an implementation form of the embodiment of the application, in the second stage joint optimization, the adjustment module 204 is further configured to set the learning rate of the cross-modal alignment module to be ten times or more than the learning rate of the heterogeneous language model.

[0077] Optionally, in an implementation form of the embodiment of the application, in the second stage joint optimization, the adjusting module 204 further keeps the parameters of the first 2 layers of the visual encoder frozen and fine-tunes the parameters of the remaining high layers.

[0078] Optionally, in an implementation form of the embodiment of the application, the adjusting module 204 further sets the calculation formula of the alignment loss as:

[0079] Lalign = 1 - cos(visual_embed, text_embed);

[0080] and takes the alignment loss Lalign as a regularization term and performs weighted summation with the cross-entropy loss (L CE, Cross-Entropy Loss), and the total loss (L total ) is calculated as follows:

[0081] Ltotal = L CE + λ·Lalign

[0082] where visual_embed is the visual feature embedding output by the visual encoder (ViT), text_embed is the text feature embedding processed by the heterogeneous language model (LLM), and λ is a hyperparameter that can be determined as 0.05 through grid search, and this value aims to balance the cross-entropy loss (task performance) and the alignment loss (cross-modal consistency);

[0083] The second difference is determined according to the total loss value L total, so as to adjust the parameters of the visual encoder, the heterogeneous language model and the cross-modal alignment module, thereby ensuring that the cross-modal alignment effect is significantly improved without sacrificing the language task performance.

[0084] The application provides a device for training a visual-language model, a setting construction module is configured to construct a corresponding pre-query template and a post-query template based on a preset instruction template of an aligned large language model; a generation module is configured to input the pre-query template into the preset aligned large language model, so as to generate a user instruction based on the self-recurrence of the language model; a response module is configured to input the user instruction into the preset aligned large language model again, combine the pre-query template and the post-query template, and generate corresponding response data; and an obtaining module is configured to combine the pre-query template, the post-query template, the user instruction, and the corresponding response data to form a data set containing at least one "instruction-response" data pair as a target data set. The entire working process of the device is an automatic process without manual intervention, does not depend on predefined seed problems or prompt ranges, and only needs to input the pre-query template to automatically generate the instruction and the response by using the self-recurrence of the large language model, without manual labeling or prompting engineering throughout the process, thereby greatly reducing the labor cost and significantly improving the efficiency of obtaining the visual-language model meeting the training requirements.

[0085] Embodiment three,

[0086] The application also provides a storage medium having a computer program stored thereon, the program being executed by a processor to implement any one of the visual-language model training methods described in the foregoing embodiment one.

[0087] Embodiment four,

[0088] The application also provides an electronic device, such as Figure 3 as shown, Figure 3 The application provides a structural schematic diagram of an electronic device 30, which comprises:

[0089] one or more processors 401, a communication interface 402, a memory 403, and a communication bus 404, the processor 401, the memory 403, and the communication interface 402 complete communication with each other through the communication bus 404;

[0090] The memory 403 is configured to store one or more programs.

[0091] When the one or more programs are executed by the one or more processors 401, the one or more processors 401 implement any one of the visual-language model training methods described in the foregoing embodiment one.

[0092] To this end, particular embodiments of the present subject matter have been described. In some instances, the actions recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.

[0093] In the 1990s, it was relatively easy to distinguish whether an improvement in a technology was a hardware improvement (e.g., an improvement in the circuit structure of a diode, transistor, switch, etc.) or a software improvement (an improvement in a method flow). However, as technology has evolved, many improvements in method flows today can be considered as direct improvements in hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structures by programming the improved method flows into hardware circuits. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using hardware entity modules. For example, a programmable logic device (PLD) (e.g., a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by user programming of the device. A designer programs a digital system layer "integrated" on a PLD by himself / herself, without having to ask a chip manufacturer to design and manufacture a special integrated circuit chip. Moreover, instead of manually manufacturing an integrated circuit chip, this programming is now mostly implemented using "logic compiler" software, which is similar to a software compiler used when developing a program, and the original code before compilation must also be written in a specific programming language, which is called a hardware description language (HDL), and there are many types of HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc., and the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should be aware that, as long as the method flow is logically programmed in the above-mentioned hardware description languages and programmed into an integrated circuit, a hardware circuit that implements the logical method flow can be easily obtained.

[0094] The controller can be implemented in any suitable way, for example, the controller can take the form of a microprocessor or processor and a computer readable medium storing computer readable program code, such as software or firmware, executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320, the memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that, in addition to implementing the controller in pure computer readable program code, it is also possible to implement the controller in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers and embedded microcontrollers, etc. to achieve the same functionality by logically programming the method steps. Such a controller can therefore be considered as a hardware component, and the means included therein for implementing the various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing the various functions can even be considered as both a software module implementing the method and a structure within the hardware component.

[0095] The system layers, devices, modules or units illustrated by the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0096] For the sake of description, the above devices are described in various units by functions respectively. Of course, the functions of each unit can be implemented in one or more software and / or hardware in the implementation of the present application.

[0097] It should also be noted that the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, such that processes, methods, articles or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or other elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0098] Those skilled in the art will appreciate that embodiments of the present application can be readily used as a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.

[0099] The present application can be described in the general context of computer- executable instructions, such as program modules, being executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The present application can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote computer storage media including memory storage devices.

[0100] Each of the embodiments described in this specification has been described taking a progressive approach, and the same or similar parts between the embodiments can be mutually referred to, and each embodiment focuses on the difference from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0101] The above merely provides embodiments of the present application, but does not serve to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the scope of the claims of the present application.

Claims

1. A method for training a vision-language model, the method comprising: The method comprises the steps of: determining a visual encoder (ViT) and a heterogeneous language model (LLM) constituting the multi-modal heterogeneous recognition model; constructing a training data set containing a text-image alignment task, a text-driven visual positioning task and a pure text reasoning task; performing dimensional dynamic alignment adaptation on the heterogeneous language model to match the parameter architecture of the heterogeneous language model with the visual feature dimension output by the visual encoder, and retaining the pre-training language ability of the heterogeneous language model; using the training data set, performing supervised fine-tuning on the visual-linguistic model using a "freeze-thaw" two-stage training strategy to perform parameter training on a cross-modal alignment module (Aligner) connecting the visual encoder and the heterogeneous language model in the visual-linguistic model, so that the visual feature dimension output by the visual encoder can be converted by the parameter-trained cross-modal alignment module to adapt to the input feature dimension of the heterogeneous language model, thereby obtaining a trained visual-linguistic model. 2.The method of Claim 1, wherein, The dimensional dynamic alignment adaptation of the heterogeneous language model comprises: performing bilinear interpolation expansion weight matrix for at least one linear layer of the heterogeneous language model, modifying the hidden layer dimension (hidden_size) of the heterogeneous language model to match the visual feature dimension output by the visual encoder; performing layer number padding on the heterogeneous language model according to the difference in layer number between the heterogeneous language model and the original language model, wherein the newly added layer adopts cyclic initialization of the weight of the original language model and adds progressive noise; performing dropout processing on the initialization of the embedding layer of the heterogeneous language model according to the difference in vocabulary size (vocab_size) between the heterogeneous language model and the original language model, and correspondingly adding or adjusting the attention head dimension (head_dim) parameter. 3.The method of Claim 1, wherein, The supervised fine-tuning of the visual-linguistic model using the training data set and the "freeze-thaw" two-stage training strategy comprises: when using the training data set, at least the following two stages are divided: first stage freeze training: freeze the model parameters of the visual encoder and the heterogeneous language model to obtain predicted data corresponding to the training data input; introduce cross-entropy loss (CE Loss) to calculate the first difference between the predicted data and the corresponding true label, adjust the parameters of the cross-modal alignment module based on the first difference, until the output dimension of the cross-modal alignment module matches the input dimension of the model of the heterogeneous language model; Second-stage joint optimization: the frozen state of the visual encoder, the heterogeneous language model, and the cross-modal alignment module is released; a hierarchical learning rate strategy is used to set different learning rates for the visual encoder, the heterogeneous language model, and the cross-modal alignment module, and an alignment loss (Alignment Loss) is introduced as a regularization term on the basis of introducing the cross-entropy loss (CE Loss), to calculate a second difference between the predicted probability distribution of the training data set input into the visual-language model and the real label distribution; and the parameters of the visual encoder, the heterogeneous language model, and the cross-modal alignment module are adjusted according to the second difference and the differential learning rate, so as to keep the consistency of the visual features and the text features in the embedding space. 4.The method of Claim 3, wherein, In the second-stage joint optimization, the learning rate of the cross-modal alignment module is set to be ten times or more than ten times the learning rate of the heterogeneous language model. 5.The method of Claim 3, wherein, In the second-stage joint optimization, the first two layers of the visual encoder are kept frozen, and the parameters of the remaining high layers are fine-tuned. 6.The method of Claim 3, wherein, The pre-query template is used to identify the role of the user instruction provider, and the post-query template is used to identify the role of the user instruction executor. 7.The method of Claim 3, wherein, The alignment loss (Lalign) is calculated as follows: Lalign = 1 - cos(visual_embed, text_embed), where visual_embed is the visual feature embedding output by the visual encoder, and text_embed is the text feature embedding processed by the heterogeneous language model.

8. An apparatus for training a vision-language model, comprising: It comprises: An acquisition module for determining a visual encoder (ViT) and a heterogeneous language model (LLM) constituting the visual-language model; A construction module for constructing a training data set comprising a picture-text alignment task, a text-driven visual positioning task, and a pure text reasoning task; An alignment module for dynamically aligning the dimensions of the heterogeneous language model to match the parameter architecture of the visual feature dimension output by the visual encoder, while retaining the pre-training language ability of the heterogeneous language model; An adjustment module for supervising and fine-tuning the visual-language model using the training data set by adopting a "freeze-thaw" two-stage training strategy, to train the parameters of the cross-modal alignment module (Aligner) connecting the visual encoder and the heterogeneous language model in the visual-language model, so that the visual feature dimension output by the visual encoder can be converted into an input feature dimension adapted to the heterogeneous language model by the parameter-trained cross-modal alignment module, thereby obtaining a trained visual-language model.

9. A computer storage medium, characterized in that The computer storage medium stores computer executable instructions, which are executed to perform the training method of the visual-language model according to any one of claims 1-7.

10. An electronic device, comprising: It comprises: a processor, a memory, a communication interface, and a communication bus, the processor, the memory, and the communication interface being in communication with each other through the communication bus; The memory is configured to store at least one executable instruction, and the executable instruction is configured to enable the processor to perform operations corresponding to the method for training a visual-language model according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-modal model training method and device, equipment and storage medium

    CN117611938A

  • Visual language model training method, image tag prediction method and electronic equipment

    CN119229162A

Cited By

  • Large language model machine forgetting algorithm based on representation spatial offset

    CN121352046A

  • Visual enhancement method and device based on multi-modal language model

    CN121937578A