Model training method and device, equipment and storage medium
Through continuous visual encoders and pre-training methods, the problems of information loss and insufficient generalization caused by discrete visual encoders are solved, lossless encoding and cross-modal alignment of image information are achieved, and the processing accuracy of computer vision and the unified understanding ability of multimodal large language models are improved.
Patent Information
- Application Number
- CN202510713755.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-03
AI Technical Summary
In existing computer vision technology, discrete visual encoders lead to image information loss and insufficient feature generalization, lack cross-modal alignment capabilities, and make it difficult to achieve unified expression of generation and understanding.
A continuous visual encoder is used to convert visual data into a continuous latent space representation. Knowledge distillation and pre-training are performed through a teacher network and a diffusion model, and the visual encoder is optimized to improve cross-modal alignment capabilities and generate integrated expressions of understanding.
It achieves lossless encoding of image information, improves the fine-grained retention and cross-modal alignment capabilities of visual features, improves the accuracy and stability of computer vision processing, and supports efficient unified understanding and generation tasks of multimodal large language models.
Smart Images

Figure CN120747660A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to technical fields such as computer vision, deep learning, large models, image processing, and multimodal information processing. Background Art
[0002] Computer vision is a field of science that studies how to enable computers to understand and interpret visual information (such as images and videos). It combines knowledge from multiple disciplines such as computer science, mathematics, physics, and biology, aiming to enable computers to obtain meaningful information from visual data just like humans.
[0003] With the continuous development of artificial intelligence, the processing accuracy, efficiency and intelligence of computer vision technology are constantly improving. As the core component in the computer vision processing flow, the importance of visual encoders is becoming increasingly prominent. Summary of the Invention
[0004] The present disclosure provides a model training method, apparatus, device, and storage medium.
[0005] According to one aspect of the present disclosure, a model training method is provided, comprising:
[0006] Inputting the first sample image into a first visual encoder to map the first sample image to a continuous latent space to obtain an intermediate feature;
[0007] Input the intermediate features into the visual encoder to be trained to obtain the visual features to be optimized;
[0008] Determining a first training loss based on the visual features to be optimized;
[0009] Based on the first training loss, the visual encoder to be trained is optimized to obtain a continuous visual encoder including the first visual encoder and the second visual encoder; the second visual encoder is the visual encoder to be trained for which training convergence occurs.
[0010] According to another aspect of the present disclosure, there is provided a model training device, comprising:
[0011] a first feature extraction module, configured to input the first sample image into a first visual encoder to map the first sample image into a continuous latent space to obtain intermediate features;
[0012] The second feature extraction module is used to input the intermediate features into the visual encoder to be trained to obtain the visual features to be optimized;
[0013] a loss determination module, configured to determine a first training loss based on the visual features to be optimized;
[0014] The optimization module is used to optimize the visual encoder to be trained based on the first training loss to obtain a continuous visual encoder including the first visual encoder and the second visual encoder; the second visual encoder is the visual encoder to be trained whose training converges.
[0015] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0016] at least one processor; and
[0017] a memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0019] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0020] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.
[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0023] Figure 1 is a flowchart of a model training method according to an embodiment of the present disclosure;
[0024] Figure 2 is a schematic diagram of a process for determining a first training loss according to an embodiment of the present disclosure;
[0025] Figure 3 1 is a flow chart of optimizing a multimodal large language model for a multimodal unified understanding large model for image understanding tasks according to an embodiment of the present disclosure;
[0026] Figure 4 1 is a flow chart of optimizing a multimodal large language model and a diffusion model to be trained for a multimodal unified understanding large model for an image generation task according to an embodiment of the present disclosure;
[0027] Figure 5 This is another flowchart of optimizing a multimodal large language model and a diffusion model to be trained for a multimodal unified understanding large model for an image generation task according to an embodiment of the present disclosure;
[0028] Figure 6 is a schematic diagram of a data processing process in a pre-training phase of a continuous image encoder according to an embodiment of the present disclosure;
[0029] Figure 7 1 is a schematic diagram of the data processing process in the generation and understanding integrated training phase of a multimodal large model according to an embodiment of the present disclosure;
[0030] Figure 8A and Figure 8B is a structural diagram of a model training device according to an embodiment of the present disclosure;
[0031] Figure 9 It is a block diagram of an electronic device used to implement the model training method of an embodiment of the present disclosure. DETAILED DESCRIPTION
[0032] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0033] The terms "first," "second," and the like in this disclosure are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. Furthermore, the terms "including," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.
[0034] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations shown in the flowchart in the embodiments of the present disclosure, or there is a sequence of execution between different operations in technical implementation, otherwise, the execution order between multiple operations may not be prioritized, and multiple operations may also be executed simultaneously.
[0035] In computer vision tasks, image feature generation and understanding are often performed by separate models. While this separate approach can address each problem, it lacks a unified, integrated generation and understanding paradigm. Visual tokenization is crucial in achieving this integrated paradigm.
[0036] In related technologies, discrete visual encoders are often used for discrete visual encoding. Discrete visual encoding is used to discretize an image into a set of visual signals. The understanding of the image by this set of visual signals is discrete, and the features of different parts of the image are mapped to a discrete symbol space to obtain this set of visual signals. The discretized encoder of discrete visual encoding is used to map image features to a discrete symbol space, and the generated feature representation is usually a discrete independent image token (word). The image token obtained in this way is similar to the discrete expression of language modal words, so it is adopted by most mainstream research.
[0037] However, there are some problems with discretized encoders, including: on the one hand, the discretized expression of images can easily lead to image information loss, which is not conducive to image understanding; on the other hand, most discretized encoders use classic generative training methods, resulting in insufficient feature generalization and lack of cross-modal alignment capabilities.
[0038] In view of this, in order to solve at least one of the above problems, an embodiment of the present disclosure provides a model training method. The method trains a continuous visual encoder. Unlike traditional discretized encoders, the continuous visual encoder provided by the present disclosure converts the input visual data into a continuous latent space representation. Compared with discrete visual encoding, the output of the continuous visual encoding in the present disclosure is a continuous sequence of feature vectors, while the output of discrete visual encoding is a set of discrete visual symbol sequences.
[0039] In this disclosure, discretization operations such as vector quantization are not involved. The generated latent space is represented as a continuous numerical vector or tensor, which can avoid the information loss caused by discrete mapping during traditional discretization. Therefore, this disclosure retains the fine-grained information of the image through continuous visual encoding, while discrete visual encoding extracts the coarse-grained core visual concepts in the image.
[0040] The continuous visual encoder obtained through training disclosed in the present invention is suitable for large model architectures and can also be used as a component module of any visual coding network to achieve efficient extraction of continuous visual coding features.
[0041] For example, in a multimodal large model deployment scenario, the server side can deploy a multimodal unified understanding large model based on the continuous visual encoder and the multimodal large language model disclosed in this disclosure. The continuous visual encoder can efficiently convert the input visual data into a continuous latent space representation, thereby providing high-quality and high-precision visual features for the multimodal large language model, which can help the multimodal large language model to better understand and process multimodal data. These models (such as continuous visual encoders and large language models) can be deployed on the same server to achieve tight integration and efficient interaction; they can also be deployed on different servers for collaborative work according to actual needs.
[0042] In the present disclosure, the server can serve as the execution entity of the model training method provided by the present disclosure, and can flexibly call and manage these models to ensure the stable operation and efficient execution of the multimodal unified understanding large model, and meet the large-scale data processing needs in diverse scenarios.
[0043] like Figure 1 FIG. 1 is a flow chart of the model training method provided in an embodiment of the present disclosure, including the following contents:
[0044] S101: Input a first sample image into a first visual encoder to map the first sample image to a continuous latent space to obtain intermediate features.
[0045] During implementation, the first visual encoder can adopt a VAE (Variational Autoencoder) encoder and its extended deformations (such as β-VAE, InfoVAE, etc.). The VAE encoder is a model component that can convert original high-dimensional data (such as images, speech, etc.) into a low-dimensional representation in a latent space.
[0046] During implementation, the first sample image is input into the first visual encoder, which maps the first sample image to a continuous latent space through a variational inference mechanism to obtain intermediate features.
[0047] S102: Input the intermediate features into the visual encoder to be trained to obtain the visual features to be optimized.
[0048] The trained visual encoder is the part of the model that needs to be learned and optimized. It further processes the intermediate features of the input and generates the visual features to be optimized. In other words, the obtained low-dimensional intermediate features are further converted into high-dimensional visual features to be optimized through the visual encoder to be optimized.
[0049] S103: Determine a first training loss based on the visual feature to be optimized.
[0050] The first training loss is a metric that measures the gap between the current performance of the visual encoder being trained and the target performance. By determining the first training loss, we can quantify the deviation between the model's processing results for the first sample image under the existing parameters and the expected target, thus providing a data basis for subsequent optimization of the model parameters of the visual encoder being trained.
[0051] S104, optimizing the visual encoder to be trained based on the first training loss to obtain a continuous visual encoder including the first visual encoder and the second visual encoder; the second visual encoder is the visual encoder to be trained whose training converges.
[0052] Based on the calculated first training loss, the visual encoder to be trained is optimized to obtain a trained visual encoder that has converged. A trained visual encoder that has converged means that, after multiple iterations of optimization, its parameters have stabilized and its model performance has either stopped improving significantly or reached a satisfactory or stable level. At this point, the trained visual encoder is used as the second visual encoder, forming a complete continuous visual encoder together with the first visual encoder.
[0053] In the embodiment of the present disclosure, by inputting the first sample image into the first visual encoder, the first sample image can be mapped to a continuous latent space to obtain intermediate features, which can retain the basic information of the continuous features of the first sample image. The visual encoder to be trained further transforms the obtained intermediate features, which can more deeply mine the high-level semantic information of the first sample image and form more discriminative visual features to be optimized. The first training loss is determined based on the visual features to be optimized, and the visual encoder to be trained is optimized accordingly, and finally a second visual encoder with converged training is obtained, which then forms a continuous visual encoder with the first visual encoder. The continuous visual encoder disclosed in the present disclosure is composed of a first visual encoder and a second visual encoder, which can map the first sample image from the original pixel space to a continuous latent space, and further convert the intermediate features into high-dimensional semantic space understanding through the second visual encoder. The high-dimensional semantic space understanding can be directly used for subsequent image processing tasks, such as image content understanding tasks, multimodal tasks, etc. The continuous visual encoder obtained by the model training method provided in the present disclosure can losslessly encode the first sample image information by providing a continuous mapping relationship, which helps to maintain the semantic coherence and consistency of the first sample image features at different levels, and can provide more powerful visual features for subsequent image processing tasks, thereby improving the accuracy and stability of computer vision processing results.
[0054] In the embodiment of the present disclosure, based on freezing the model parameters of the first visual encoder, knowledge distillation can be performed on the visual encoder to be trained through the teacher network to determine the first training loss for optimizing the model parameters of the visual encoder to be trained.
[0055] Furthermore, in order to improve the ability of the visual encoder to be trained on the understanding side (i.e., understanding the image content), the embodiment of the present disclosure, based on the introduction of the teacher network for knowledge distillation, additionally adds a diffusion model to be trained to assist in training the visual encoder to be trained. Thus, the first training loss is determined based on the visual features to be optimized, such as Figure 2 As shown, it can also be implemented as optimizing the model parameters of the visual encoder to be trained by offline pre-training, including:
[0056] S201, determining the knowledge transfer loss based on the teacher network and the visual features to be optimized.
[0057] The teacher network is an image-text multimodal encoder pre-trained using contrastive learning. During implementation, a mature, stable, and reliable neural network model can be used as the teacher network. For example, a similar pre-trained model architecture such as CLIP (Contrastive Language-Image Pre-training) can be used as the teacher network. The core components of the teacher network include a teacher visual encoder and a teacher text encoder, which are responsible for encoding the first sample image and its corresponding natural language text into encoding features of the corresponding modality, respectively.
[0058] The knowledge transfer loss is determined based on the teacher network and the visual features to be optimized. This can be implemented by transferring the knowledge of the teacher model to the student model (the visual encoder to be trained) to improve the generalization performance of the student model on the understanding side. This allows the visual features output by the trained continuous visual encoder to have cross-modal alignment capabilities. Figure 2 As shown, the steps to determine the knowledge transfer loss are as follows:
[0059] S2011, input the first sample image into the teacher visual encoder of the teacher network to obtain reference visual features.
[0060] The teacher visual encoder is a module used to process image information, converting visual content associated with the first sample image into reference visual features. During implementation, the teacher visual encoder can extract spatial, texture, and color features from the first sample image through a series of neural network operations, including convolutional layers, pooling layers, and activation functions, converting raw pixel-level data into higher-level reference visual features. For example, the image encoder of the CLIP model can be used as the teacher visual encoder in this disclosure.
[0061] S2012: Input the first sample text corresponding to the first sample image into the teacher text encoder of the teacher network to obtain reference text features.
[0062] Unlike the teacher visual encoder, the teacher text encoder is a module used to process text information. It can convert the first sample text associated with the first sample image into text features. During implementation, the text encoder of the CLIP model can be used as the teacher text encoder in this disclosure.
[0063] S2013: Determine a first sub-loss based on the reference text features and the visual features to be optimized.
[0064] The first sub-loss measures the difference between the reference text features extracted by the teacher network and the visual features to be optimized obtained by the trained visual encoder, so as to align features across text and image modalities so that the trained continuous visual encoder can be suitable for cross-modal task understanding and processing.
[0065] During implementation, the first sub-loss can adopt some commonly used loss function forms, such as using contrastive loss to calculate the difference between the reference text features and the visual features to be optimized in the semantic space to obtain the first sub-loss.
[0066] S2014: Determine a second sub-loss based on the reference visual features and the visual features to be optimized.
[0067] The second sub-loss mainly measures the difference between the reference visual features extracted by the teacher network and the visual features to be optimized obtained by training the visual encoder, so as to guide the optimization process of the visual encoder to be trained.
[0068] During implementation, in order to accelerate convergence, the second sub-loss can be determined based on the following steps:
[0069] Step A1, performing dimensionality reduction processing on the reference visual feature to obtain a first dimensionality reduction feature;
[0070] Step A2, performing dimensionality reduction processing on the visual feature to be optimized to obtain a second dimensionality reduction feature;
[0071] Step A3: determine the loss between the first dimensionality reduction feature and the second dimensionality reduction feature to obtain a second sub-loss.
[0072] During implementation, the reference visual features output by the teacher visual encoder can be first subjected to dimensionality reduction processing using a pooling operation to obtain a first dimensionality-reduced feature. Simultaneously, the visual features to be optimized are processed using the same pooling operation to obtain a second dimensionality-reduced feature. Subsequently, a distillation loss (DistillLoss) can be calculated between the reference visual features and the visual features to be optimized to determine a second sub-loss. During implementation, this distillation loss can be represented by the L2 distance (Euclidean distance) between the first dimensionality-reduced feature and the second dimensionality-reduced feature.
[0073] In the disclosed embodiment, dimensionality reduction processing is performed on the reference visual features and the visual features to be optimized, which can retain the key information in the features while reducing the dimensionality, thereby reducing the time for calculating high-dimensional features, further reducing the model training time, and improving training efficiency.
[0074] S2015: Determine a knowledge transfer loss based on the first sub-loss and the second sub-loss.
[0075] During implementation, the first and second sub-losses can be summed according to preset weights to obtain the knowledge transfer loss. This loss comprehensively reflects the overall effectiveness of the trained visual encoder in learning the teacher network's knowledge of visual-textual associations and visual feature extraction. By minimizing the knowledge transfer loss, the trained visual encoder can be guided to align with the teacher network, thereby improving the model performance of the trained visual encoder.
[0076] During implementation, the model structure of the visual encoder to be trained can be the same as that of the teacher visual encoder. Initially, the model parameters of the visual encoder to be trained are assumed to be the same as those of the teacher visual encoder. This allows the visual encoder to be trained as a student model on a better basis for optimization training.
[0077] In the disclosed embodiment, based on the first sub-loss between the reference text features and the visual features to be optimized, the distance between the visual features to be optimized and the reference text features can be shortened, the expressive ability of the features learned by the teacher network by the visual encoder to be trained can be enhanced, its multimodal generalization performance can be improved, and the generated visual features to be optimized can be made to have more key information of the original image and potential correlation representations between key information. Based on the reference visual features and the visual features to be optimized, the second sub-loss is determined, and the knowledge of the teacher model on image encoding can be transferred to the visual encoder to be trained, thereby improving the understanding and representation ability of the visual encoder to be trained on the image content. The knowledge transfer loss combines the first sub-loss and the second sub-loss to provide more comprehensive supervision information for model training, avoids the training instability problem that may be caused by a single loss function, and further improves the stability and efficiency of model training.
[0078] S202: Processing intermediate features and noise information based on the diffusion model to be trained to obtain diffusion features.
[0079] Diffusion models are a type of generative model that simulates the gradual diffusion of data in noise while simultaneously learning the inverse denoising process to generate high-quality data that closely matches the original data distribution. Diffusion features are generated by processing intermediate features and noise information based on the diffusion model to be trained. This is done by adding noise to the intermediate features and then passing them through the diffusion model to be trained. The output of the diffusion model to be trained is the diffusion feature.
[0080] In some embodiments, a diffusion model with a complex structure may be selected as the diffusion model to be trained.
[0081] In other embodiments, the primary function of the diffusion model to be trained is to assist in training the visual encoder to be trained. Therefore, during implementation, the diffusion model to be trained may include multiple fully connected layers. This streamlined structure effectively shortens the processing time of the diffusion model to be trained during the diffusion feature extraction process, improves training efficiency, and facilitates debugging and optimization of the visual encoder to be trained within limited computing resources and time.
[0082] S203: Determine a first diffusion loss based on the diffusion characteristic and the intermediate characteristic.
[0083] The first diffusion loss reflects the degree of deviation between the diffusion feature and the intermediate feature. Through the first diffusion loss, the model parameters of the visual encoder to be trained can be effectively adjusted so that it can enhance the ability to understand the image based on the generation task of the diffusion model to be trained and improve the accuracy of continuous visual encoding.
[0084] S204: Determine a first training loss based on the knowledge transfer loss and the first diffusion loss.
[0085] During implementation, the recognition transfer loss and the first diffusion loss can be added together according to a preset weight distribution to determine the first training loss.
[0086] In an embodiment of the present disclosure, a teacher network provides reference visual features and reference text features, and a loss is calculated with the visual features to be optimized to obtain a knowledge transfer loss. The knowledge transfer loss can be used to transfer the teacher network knowledge to the visual encoder to be trained, prompting the visual features to be optimized to align with the reference features of the teacher network, thereby achieving semantic consistency of the visual features. The diffusion model to be trained processes intermediate features and noise information to generate diffusion features, which are used to determine a first diffusion loss with the intermediate features to improve the generation effect of the diffusion model to be trained. The knowledge transfer loss and the first diffusion loss are fused to determine a first training loss, which is used as a global optimization target to ensure that the model can both accurately understand the input and generate high-quality output. Overall, the present disclosure provides an image encoding-decoding pre-training scheme based on a diffusion model through offline pre-training. Multimodal contrastive learning and model distillation are introduced through the teacher network. This provides an offline pre-training paradigm that integrates multimodal contrastive learning, model distillation, and image reconstruction. This paradigm can enhance the feature expression of the visual encoder on the understanding side. This pre-training scheme can improve the generalization performance of the continuous visual encoder without compromising its generation capability, thereby improving its visual encoding capability in both understanding and generation tasks. At the same time, compared with online updates of the same multimodal large language model, the offline pre-training method has faster training efficiency and can significantly reduce training time and resource consumption.
[0087] Furthermore, after obtaining the first training loss, the model parameters of the diffusion model to be trained can be simultaneously optimized while optimizing the visual encoder to be trained based on the first training loss. This first training loss comprehensively considers the optimization requirements from both the knowledge transfer perspective and the diffusion model to be trained. When used to optimize the diffusion model to be trained, it balances feature extraction on the understanding side with the generation effect measured based on the diffusion model to be trained, further improving the optimization effect of the visual encoder to be trained and thus increasing training efficiency.
[0088] During offline pre-training, convergence conditions can be set to determine whether pre-training has achieved the desired results. These convergence conditions can be set as follows: the first training loss of the trained visual encoder on the validation set must not decrease significantly over multiple consecutive iterations, the fluctuation range of the first training loss must be less than a preset threshold, or the maximum number of training iterations must be reached. By monitoring these conditions in real time, ineffective overtraining can be avoided and the model can be stopped at the optimal state, effectively achieving the desired training results.
[0089] Verification can be implemented as follows: For each validation sample in the validation set, input it into the first visual encoder to obtain intermediate features. This intermediate feature is then input into the visual encoder to be trained to obtain the visual features to be optimized. This visual feature to be optimized is then input into the classifier to obtain the output class probability distribution. This output is the probability that the validation sample belongs to each class. The class corresponding to the maximum probability in the class probability distribution is obtained as the classification class, and the classification accuracy of multiple validation samples is then determined. If the classification accuracy is greater than or equal to the classification threshold, the visual encoder to be trained is considered to have passed verification.
[0090] During implementation, the Top-1 accuracy evaluation method of the ImageNet (ImageNet Large Scale Visual Recognition Challenge, a large and well-known visual recognition dataset) dataset can be used for testing. When its classification accuracy is greater than or equal to the classification threshold, the visual encoder to be trained is judged to have passed the verification. At this point, the pre-training is completed and the final required continuous visual encoder is obtained.
[0091] In the disclosed embodiment, not only is a continuous visual encoder trained through the aforementioned offline pre-training scheme, but a multimodal unified understanding large model can also be constructed based on it.
[0092] Specifically, based on the content explained above, in multimodal tasks, generation and comprehension tasks are often completed by different models, lacking a unified, integrated generation and comprehension paradigm. The output features of the continuous visual encoder obtained through the aforementioned pre-training will serve as key image semantic information input and can be transmitted to the multimodal large language model, enabling it to better extract visual features. For multimodal tasks, it can also fuse and process the semantic content of images and text based on high-quality visual features.
[0093] Therefore, in the disclosed embodiment, a multimodal unified understanding large model can be constructed based on the continuous visual encoder. The multimodal unified understanding large model also includes a multimodal large language model, a diffusion model to be trained, and a visual decoder.
[0094] Among them, the output features of the continuous visual encoder are used to input into the multimodal large language model;
[0095] Multimodal Large Language Models (MLLMs) are a type of AI model that integrates the processing capabilities of multiple data modalities, such as images, speech, video, and text, and is built on a large scale. Their core approach is to understand, generate, and reason about diverse information through cross-modal semantic alignment and joint modeling, overcoming the limitations of traditional single-modal large language models that only process text.
[0096] For image generation tasks, the output features of the multimodal large language model are used to input the diffusion model to be trained; the output features of the diffusion model to be trained are used to input the visual decoder.
[0097] For image understanding tasks, a multimodal large language model is used to output content understanding results, namely textual information. For example, if a user asks, "What kind of plant is in Image A, and how should it be planted?" the multimodal large language model can understand the content of Image A and, based on knowledge, determine the plant's variety and provide planting recommendations.
[0098] It is understandable that the multimodal unified large model can use different modules to process different tasks. For example, for image understanding tasks, a continuous visual encoder and a multimodal large language model can be used for processing, and the multimodal large language model outputs the understanding results. For image generation tasks, the second visual encoder in the continuous visual encoder, the multimodal large language, the diffusion model to be trained, and the image decoder can be used for processing to obtain the image generated by the image decoder.
[0099] In the disclosed embodiments, a continuous visual encoder, a multimodal large language model, a diffusion model to be trained, and a visual decoder are combined into a multimodal unified understanding large model, thereby achieving cross-modal information interaction between visual features, text features, and generation features. By inputting the output features of the continuous visual encoder into the multimodal large language model, the semantic information of the image can be fused with the semantic information of the text in the multimodal large language model. In this way, the multimodal large language model can simultaneously understand and process information from both image and text modalities, achieving multimodal unified understanding. For image generation tasks, the output features of the multimodal large language model are input into the diffusion model to be trained, and the diffusion model then inputs the processed features into the visual decoder to generate an image. As a result, the multimodal unified understanding large model can form a multi-task processing capability from image understanding to image generation. The multimodal unified understanding large model provided by the disclosed embodiments can simultaneously process generation tasks such as text-to-image and image-to-image and understanding tasks such as image question answering and image classification under the same framework, achieving coordinated optimization of generation tasks and understanding tasks, and improving the overall performance of the multimodal unified understanding large model in multimodal tasks.
[0100] It is understandable that the unified large model of multimodal understanding provided by the embodiments of the present disclosure can also handle separate text modality tasks, such as completing text understanding tasks, language translation tasks, etc.
[0101] During implementation, for the multimodal unified understanding large model, since the continuous visual encoder has been trained based on the aforementioned pre-training method, in order to improve the information processing capability of the multimodal unified understanding large model, the multimodal large language model and the diffusion model to be trained of the multimodal unified understanding large model can be optimized based on image understanding tasks and image generation tasks to achieve fine-tuning of the multimodal unified understanding large model.
[0102] The image understanding task primarily aims to enable a large multimodal language model to accurately parse and understand the semantics of input images. For example, a large multimodal language model can identify objects, scenes, and people in an image, understand the relationships between them, and answer questions related to the image content or generate descriptions of the image content, such as describing a picture.
[0103] Image generation tasks primarily include text-based image generation and image-based image generation. Text-based image generation tasks involve generating an image that matches a given text description and image content. Image-based image generation tasks involve modifying a given image based on a user's modification request (expressed in textual form) to generate a new image. Examples include modifying the style of a given image, adding or deleting content, or repairing a given image.
[0104] In the disclosed embodiment, the image understanding task requires the multimodal large language model to be able to accurately grasp the semantic information in the image based on the text description. By optimizing the multimodal large language model, it can better understand the complex semantic associations between images and texts. The image generation task requires the multimodal large language model to be able to generate images that meet the requirements based on input conditions such as text descriptions or partial images. By optimizing the diffusion model to be trained, it can better understand and utilize the features output by the multimodal large language model, so that the visual encoder can more accurately understand and generate images that meet the semantic and visual features. In summary, by optimizing the multimodal large language model and the diffusion model to be trained in image understanding and generation tasks, the multimodal large language model and the diffusion model to be trained can better adapt to various multimodal tasks.
[0105] In the embodiment of the present disclosure, for the image understanding task, the implementation method of optimizing or fine-tuning the multimodal large language model of the multimodal unified understanding large model can be as follows: Figure 3 Shown, including:
[0106] S301, using a continuous visual encoder to generate a first semantic vector of a second sample image.
[0107] That is, the pre-trained continuous visual encoder is used to process the second sample image. The continuous visual encoder, through its first visual encoder and second visual encoder, can losslessly encode the second sample image into a high-dimensional first semantic vector, which contains key semantic information of the second sample image.
[0108] S302: Input the first semantic vector and the second sample text for the second sample image into the multimodal large language model to obtain a text response generated by the multimodal large language model.
[0109] In the image understanding task, the second sample text for the second sample image refers to providing the multimodal large language model with semantic information, specific questions, or operation instructions related to the image. The multimodal large language model has the ability to process multimodal data. It can fuse the first semantic vector of the second sample image and the corresponding second text information. In this process, the multimodal large language model will combine the visual features of the second sample image and the textual semantic features of the second sample to generate a text response related to the content of the second sample image.
[0110] S303: Determine the text modality loss based on the text reply and the true value of the text reply.
[0111] The true value of the text response refers to the accurate and satisfactory text result that the multimodal large language model is expected to generate for the second sample image and second sample text input. By comparing the difference between the text response actually generated by the multimodal large language model and the true value of the text response, the text modality loss can be calculated. The text modality loss value reflects the gap between the performance of the multimodal large language model on the text generation task and the ideal result. It can be calculated using a loss function suitable for text generation tasks, such as cross-entropy loss.
[0112] S304: Optimize the multimodal large language model based on text modality loss.
[0113] That is, based on the obtained text modality loss, the parameters of the multimodal large language model are adjusted and optimized. This process is usually implemented through the backpropagation algorithm, that is, according to the gradient information of the loss function, the parameters of the multimodal large language model are updated. In the subsequent training process, the multimodal large language model can generate text responses that are closer to the ground truth, thereby improving the multimodal large language model's ability to grasp the semantics of image content and the accuracy of text generation in image understanding tasks.
[0114] In the disclosed embodiment, a continuous visual encoder is used to generate a first semantic vector for the second sample image. The continuous visual encoder losslessly encodes the features of the second sample image, providing an accurate image semantic basis for subsequent text response generation. The multimodal large language model is optimized based on text modality loss, so that the multimodal large language model adjusts the fusion method of the first semantic vector and the second sample text for the second sample image based on the true value feedback of the text response, making the information transfer between the two features smoother and further optimizing the performance of multimodal fusion.
[0115] In the embodiment of the present disclosure, for the image generation task, the implementation method of optimizing the multimodal large language model of the multimodal unified understanding large model and the diffusion model to be trained is as follows: Figure 4 Shown, including:
[0116] S401 , for an image-to-image task, determining a first initial vector based on a third sample image.
[0117] In the image generation task, the first initial vector can be obtained by using a mask token. For example, the third sample image used to generate the new image is input into the continuous visual encoder to obtain the first image token outputted by it. A portion of features can be randomly sampled from the first image token, and the sampled portion of features can be replaced with preset mask features, and the preset mask features are combined with the remaining features in the first image token as the first initial vector. The specific mask ratio can be controlled by a mask scheduling function (such as cosine scheduling). Which positions retain the initial features of the third sample image and which positions are initialized as mask features depends on the specific needs and goals of the image generation task, and the embodiments of the present disclosure do not limit this.
[0118] S402: Input the first initial vector into a second visual encoder of the continuous visual encoder to obtain a second semantic vector.
[0119] The second visual encoder is responsible for extracting higher-level semantic features from the first initial vector and converting it into a second semantic vector that can be processed by the multimodal large language model for subsequent multimodal fusion processing.
[0120] S403: Input the second semantic vector and the first text description of the image-to-image task into a multimodal large language model to obtain a first visual semantic vector output by the multimodal large language model.
[0121] The first textual description defines the requirements for the generated image, such as the image style. The multimodal large language model combines the second semantic vector of the third sample image with the first textual description to generate a first visual semantic vector that incorporates the semantics of both. This first visual semantic vector incorporates features of both the third sample image and the desired generated image as described in the first textual description, providing guidance for subsequent image generation.
[0122] S404: Input the first visual semantic vector into the diffusion model to be trained to obtain a second visual semantic vector.
[0123] The diffusion model to be trained can optimize and refine the first visual semantic vector during the diffusion process, and generate a second visual semantic vector with higher quality and more consistent with the target distribution by introducing noise and gradually denoising it.
[0124] S405 : Determine a second diffusion loss based on the first visual semantic vector and the second visual semantic vector.
[0125] The second diffusion loss measures the difference between the output and input of the diffusion model, specifically the degree of deviation between the second visual semantic vector and the first visual semantic vector. This reflects the model's optimization of the visual semantic vector during the diffusion process. By calculating the second diffusion loss, we can evaluate the ability of the multimodal large language model and the diffusion model to be trained to generate and restore image semantic features in image-to-image tasks.
[0126] S406: Optimize the multimodal large language model and the diffusion model to be trained based on the second diffusion loss.
[0127] Based on the second diffusion loss, the backpropagation algorithm is used to optimize the multimodal large language model and the diffusion model to be trained. The gradient of the multimodal large language model parameters is calculated based on the diffusion loss, and then the parameter values of the multimodal large language model are updated in the direction of gradient descent. After multiple iterative optimizations, the multimodal large language model is able to generate more accurate and instructive visual semantic vectors, while the diffusion model to be trained is able to better optimize and generate visual semantic vectors, thereby improving the quality and accuracy of generated images in the image-to-image task, making the generated images more consistent with the requirements of the third sample image and the first text description input.
[0128] In the disclosed embodiment, the first visual semantic vector is input into the diffusion model to be trained to obtain the second visual semantic vector, and the second diffusion loss is determined based on the two. The multimodal large language model and the diffusion model to be trained are optimized based on the second diffusion loss, which can adjust the parameters of the model more specifically, so that the fine-tuned multimodal unified understanding large model is more suitable for image-to-image tasks.
[0129] In the embodiment of the present disclosure, for the image generation task, the implementation method of optimizing the multimodal large language model of the multimodal unified understanding large model and the diffusion model to be trained is as follows: Figure 5 Shown, including:
[0130] S501, for the Wensheng graph task, determining a second initial vector.
[0131] Unlike image-to-image tasks, the second initial vector is usually a vector used to initialize the image generation process. It can be a random vector or a specifically designed vector to represent the initial state of the image to be generated. For example, the feature sequence of the image to be generated is completely replaced with a special mask feature as the second initial vector.
[0132] S502: Input the second initial vector into the second visual encoder of the continuous visual encoder to obtain a third semantic vector.
[0133] The second visual encoder is responsible for extracting higher-level semantic features from the second initial vector and converting it into a third semantic vector that can be processed by the multimodal large language model for subsequent multimodal fusion processing.
[0134] S503: Input the third semantic vector and the second text description of the text-to-image task into the multimodal large language model to obtain a third visual semantic vector output by the multimodal large language model.
[0135] The second text description is used to define the requirements of the new image to be generated.
[0136] The multimodal large language model combines the input third semantic vector and the second text description to generate a third visual semantic vector. The third visual semantic vector carries the image feature information described in the second text and guides the subsequent image generation process.
[0137] S504: Input the third visual semantic vector into the diffusion model to be trained to obtain a fourth visual semantic vector.
[0138] The third visual semantic vector is fed as input into the diffusion model to be trained. The diffusion model to be trained performs diffusion processing on the input third visual semantic vector to generate a fourth visual semantic vector.
[0139] S505 : Determine a third diffusion loss based on the fourth visual semantic vector and the third visual semantic vector.
[0140] The third diffusion loss reflects the degree of deviation between the fourth visual semantic vector generated after the diffusion model to be trained performs diffusion processing on the input third visual semantic information and the third visual semantic vector of the original input.
[0141] S506: Optimize the multimodal large language model and the diffusion model to be trained based on the third diffusion loss.
[0142] The third diffusion loss is used to adjust and optimize the parameters of the multimodal large language model and the diffusion model to be trained. The parameters of both models can be updated using the backpropagation algorithm based on the gradient information of the loss function.
[0143] In the disclosed embodiment, by optimizing the multimodal large language model and the diffusion model to be trained in the multimodal unified understanding large model based on the third diffusion loss, the accuracy and quality of image generation can be improved, the semantic understanding and generation capabilities can be enhanced, the generalization and adaptability of the model can be improved, the training process can be accelerated and the performance can be ensured to ensure stable, thereby realizing the integrated processing of multimodal tasks.
[0144] In summary, the overall model training process provided in the embodiments of the present disclosure may include a pre-training stage of a continuous image encoder and a generation-understanding integrated training and fine-tuning stage of a multimodal unified understanding large model.
[0145] Among them, the pre-training stage of the continuous image encoder is used to train the image encoding part of the understanding side of the multimodal unified understanding model (i.e., the visual encoder to be trained) and the image decoding part of the generation side (i.e., the diffusion model to be trained). In addition to these two modules to be trained, this stage introduces modules for freezing model parameters, including Figure 6 The first visual encoder 601, the teacher text encoder 604 and the teacher visual encoder 603 are shown in FIG. Specifically, the pre-training method of this stage is as follows: Figure 6 Shown, including:
[0146] The first sample image is input into the first visual encoder 601 to map the first sample image to a continuous latent space to obtain an intermediate feature.
[0147] The intermediate features are input into the visual encoder to be trained 602 to obtain the visual features to be optimized. The first sample image is input into the teacher visual encoder 603 of the teacher network to obtain the reference visual features. The first sample text corresponding to the first sample image is input into the teacher text encoder 604 of the teacher network to obtain the reference text features.
[0148] A first sub-loss is determined between the reference text feature and the visual feature to be optimized. A second sub-loss is determined based on the reference visual feature and the visual feature to be optimized. The intermediate features and noise information are processed based on the diffusion model to be trained 605 to obtain a diffusion feature. A first diffusion loss is determined based on the diffusion feature and the intermediate features.
[0149] The first sub-loss is combined with the second sub-loss and the first diffusion loss to determine the first training loss.
[0150] The visual encoder to be trained 602 and the diffusion model to be trained 605 are optimized based on the obtained first training loss.
[0151] The generation and understanding integrated training phase of the multimodal large model is used to train the multimodal large language model in the multimodal unified understanding large model and fine-tune the diffusion model to be trained on the image generation end. The continuous visual encoder comes from the pre-training phase of the continuous image encoder. The training process of this fine-tuning phase is as follows: Figure 7 Shown, including:
[0152] For the image understanding task, a continuous visual encoder 701 is used to generate a first semantic vector for the second sample image. The first semantic vector and the second sample text for the second sample image are input into a multimodal large language model 702 to obtain a text response generated by the multimodal large language model. Based on the text response and its true value, a text modality loss is determined. Based on the text modality loss, the multimodal large language model is optimized.
[0153] For the image-to-image task in the image generation task, a first initial vector can be determined based on the third sample image. The first initial vector is input into the second visual encoder 602 of the continuous visual encoder 701 to obtain a second semantic vector. The second semantic vector and the first text description of the image-to-image task are input into the multimodal large language model 702 to obtain a first visual semantic vector output by the multimodal large language model 702. The first visual semantic vector is input into the diffusion model to be trained 703 to obtain a second visual semantic vector. Based on the first visual semantic vector and the second visual semantic vector, a second diffusion loss is determined. Based on the second diffusion loss, the multimodal large language model 702 and the diffusion model to be trained 703 are optimized.
[0154] For the Vincent image task in the image generation task, a second initial vector is determined. The second initial vector is input into the second visual encoder 602 of the continuous visual encoder 701 to obtain a third semantic vector. The third semantic vector and the second text description of the Vincent image task are input into the multimodal large language model 702 to obtain a third visual semantic vector output by the multimodal large language model 702. The third visual semantic vector is input into the diffusion model to be trained 703 to obtain a fourth visual semantic vector. Based on the fourth visual semantic vector and the third visual semantic vector, a third diffusion loss is determined. Based on the third diffusion loss, the multimodal large language model 702 and the diffusion model to be trained 703 are optimized.
[0155] In summary, the model training method provided by the embodiments of the present disclosure can significantly improve the performance and efficiency of the multimodal unified understanding large model in multimodal tasks through the integrated architecture of continuous visual encoder and generation understanding: compared with the use of discrete visual encoding (such as VQ-VAE) resulting in blurred image details and insufficient cross-modal alignment, the continuous encoding encoder provided by the embodiments of the present disclosure can generate high-fidelity images based on continuous visual encoding technology (such as accurate restoration of texture and color gradient), and the multimodal unified understanding large model can simultaneously support generation and understanding tasks (such as semantic verification while generating); in addition, the offline pre-training scheme improves the robustness of the multimodal unified understanding large model in open domain scenarios (such as zero-sample generation), and supports rapid adaptation to vertical fields (such as medical care, autonomous driving, etc.).
[0156] Based on the same technical concept, the embodiment of the present disclosure also provides a model training device 800, such as Figure 8A and Figure 8B Shown, including:
[0157] A first feature extraction module 801 is configured to input a first sample image into a first visual encoder to map the first sample image into a continuous latent space to obtain intermediate features;
[0158] The second feature extraction module 802 is used to input the intermediate features into the visual encoder to be trained to obtain the visual features to be optimized;
[0159] A loss determination module 803 is configured to determine a first training loss based on the visual features to be optimized;
[0160] The optimization module 804 is configured to optimize the visual encoder to be trained based on the first training loss to obtain a continuous visual encoder including the first visual encoder and the second visual encoder; the second visual encoder is the visual encoder to be trained for which training convergence has occurred.
[0161] In some embodiments, the loss determination module 803 includes:
[0162] A first determining unit 8031 is configured to determine a knowledge transfer loss based on the teacher network and the visual features to be optimized;
[0163] A processing unit 8032 is configured to process intermediate features and noise information based on the diffusion model to be trained to obtain diffusion features;
[0164] A second determining unit 8033 is configured to determine a first diffusion loss based on the diffusion characteristic and the intermediate characteristic;
[0165] The third determining unit 8034 is configured to determine a first training loss based on the knowledge transfer loss and the first diffusion loss.
[0166] In some embodiments, the first training loss is further used to optimize the diffusion model to be trained.
[0167] In some embodiments, the first determining unit 8031 is specifically configured to:
[0168] Input the first sample image into the teacher visual encoder of the teacher network to obtain the reference visual features;
[0169] Inputting the first sample text corresponding to the first sample image into the teacher text encoder of the teacher network to obtain a reference text feature;
[0170] Determine the first sub-loss based on the reference text features and the visual features to be optimized;
[0171] Determining a second sub-loss based on the reference visual feature and the visual feature to be optimized;
[0172] Based on the first sub-loss and the second sub-loss, a knowledge transfer loss is determined.
[0173] In some embodiments, the first determining unit 8031 is specifically configured to:
[0174] Perform dimensionality reduction processing on the reference visual features to obtain the first dimensionality reduction feature;
[0175] Perform dimensionality reduction processing on the visual features to be optimized to obtain the second dimensionality reduction feature;
[0176] The loss between the first dimensionality reduced feature and the second dimensionality reduced feature is determined to obtain a second sub-loss.
[0177] In some embodiments, the diffusion model to be trained includes multiple layers of fully connected layers.
[0178] In some embodiments, further comprising:
[0179] Building a multimodal unified understanding model based on continuous visual encoder;
[0180] The multimodal unified understanding model also includes a multimodal large language model, a diffusion model to be trained, and a visual decoder.
[0181] The output features of the continuous visual encoder are used as input to the multimodal large language model;
[0182] The output features of the multimodal large language model are used as input to the diffusion model to be trained;
[0183] The output features of the diffusion model to be trained are used as input to the visual decoder.
[0184] In some embodiments, a fine-tuning module 805 is further included for:
[0185] Based on image understanding tasks and image generation tasks, optimize the multimodal large language model and the diffusion model to be trained of the multimodal unified understanding large model.
[0186] In some embodiments, wherein, for an image understanding task, optimizing a multimodal large language model of a multimodal unified understanding large model is performed, the fine-tuning module 805 includes:
[0187] A first generating unit 8051 is configured to generate a first semantic vector of a second sample image using a continuous visual encoder;
[0188] A second generating unit 8052 is configured to input the first semantic vector and the second sample text for the second sample image into the multimodal large language model to obtain a text response generated by the multimodal large language model;
[0189] A first loss determining unit 8053 is configured to determine a text modality loss based on the text reply and the ground truth of the text reply;
[0190] The first optimization unit 8054 is used to optimize the multimodal large language model based on text modality loss.
[0191] In some embodiments, for an image generation task, optimizing a multimodal large language model and a diffusion model to be trained of a multimodal unified understanding large model is performed, and the fine-tuning module 805 includes:
[0192] The third generating unit 8055 is configured to determine a first initial vector based on a third sample image for the image-to-image task;
[0193] A fourth generating unit 8056 is configured to input the first initial vector into a second visual encoder of the continuous visual encoder to obtain a second semantic vector;
[0194] a fifth generating unit 8057 , configured to input the second semantic vector and the first text description of the image-to-image task into the multimodal large language model to obtain a first visual semantic vector output by the multimodal large language model;
[0195] a sixth generating unit 8058, configured to input the first visual semantic vector into the diffusion model to be trained to obtain a second visual semantic vector;
[0196] A second loss determining unit 8059 is configured to determine a second diffusion loss based on the first visual semantic vector and the second visual semantic vector;
[0197] The second optimization unit 80510 is used to optimize the multimodal large language model and the diffusion model to be trained based on the second diffusion loss.
[0198] In some embodiments, for an image generation task, optimizing a multimodal large language model and a diffusion model to be trained of a multimodal unified understanding large model is performed, and the fine-tuning module includes:
[0199] The seventh generating unit 80511 is used to determine a second initial vector for the Wensheng graph task;
[0200] An eighth generating unit 80512 is configured to input the second initial vector into a second visual encoder of the continuous visual encoder to obtain a third semantic vector;
[0201] a ninth generating unit 80513, configured to input the third semantic vector and the second text description of the text-to-image task into the multimodal large language model to obtain a third visual semantic vector output by the multimodal large language model;
[0202] a tenth generating unit 80514, configured to input the third visual semantic vector into the diffusion model to be trained to obtain a fourth visual semantic vector;
[0203] A third loss determining unit 80515 is configured to determine a third diffusion loss based on the fourth visual semantic vector and the third visual semantic vector;
[0204] The third optimization unit 80516 is used to optimize the multimodal large language model and the diffusion model to be trained based on the third diffusion loss.
[0205] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0206] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0207] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0208] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0209] like Figure 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the device 900 can also be stored in the RAM 903. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0210] Various components in the device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0211] The computing unit 901 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the model training method. For example, in some embodiments, the model training method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the model training method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the model training method in any other appropriate manner (e.g., by means of firmware).
[0212] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0213] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0214] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0215] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0216] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0217] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0218] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0219] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A model training method, comprising: Inputting the first sample image into a first visual encoder to map the first sample image to a continuous latent space to obtain an intermediate feature; Inputting the intermediate features into the visual encoder to be trained to obtain the visual features to be optimized; Determining a first training loss based on the visual feature to be optimized; Optimizing the visual encoder to be trained based on the first training loss to obtain a continuous visual encoder including the first visual encoder and the second visual encoder; The second visual encoder is the visual encoder to be trained whose training converges.
2. The method according to claim 1, wherein The determining a first training loss based on the visual feature to be optimized includes: Determining a knowledge transfer loss based on the teacher network and the visual features to be optimized; Processing the intermediate features and noise information based on the diffusion model to be trained to obtain diffusion features; determining a first diffusion loss based on the diffusion characteristic and the intermediate characteristic; The first training loss is determined based on the knowledge transfer loss and the first diffusion loss.
3. The method according to claim 2, wherein: The first training loss is also used to optimize the diffusion model to be trained.
4. The method according to claim 2, wherein: The determining of the knowledge transfer loss based on the teacher network and the visual features to be optimized includes: Inputting the first sample image into the teacher visual encoder of the teacher network to obtain reference visual features; Inputting a first sample text corresponding to the first sample image into a teacher text encoder of the teacher network to obtain a reference text feature; Determining a first sub-loss based on the reference text feature and the visual feature to be optimized; determining a second sub-loss based on the reference visual feature and the visual feature to be optimized; The knowledge transfer loss is determined based on the first sub-loss and the second sub-loss.
5. The method according to claim 4, wherein The determining the second sub-loss based on the reference visual feature and the visual feature to be optimized includes: Performing dimensionality reduction processing on the reference visual feature to obtain a first dimensionality reduction feature; Performing dimensionality reduction processing on the visual feature to be optimized to obtain a second dimensionality reduction feature; Determine the loss between the first dimensionality reduction feature and the second dimensionality reduction feature to obtain the second sub-loss.
6. The method according to claim 2, wherein: The diffusion model to be trained includes multiple fully connected layers.
7. The method according to any one of claims 2 to 6, further comprising: Building a multimodal unified understanding model based on the continuous visual encoder; The multimodal unified understanding large model further includes a multimodal large language model, the diffusion model to be trained, and a visual decoder; The output features of the continuous visual encoder are used to input into the multimodal large language model; For the image generation task, the output features of the multimodal large language model are used to input into the diffusion model to be trained; The output features of the diffusion model to be trained are used to input into the visual decoder.
8. The method according to claim 7, further comprising: Based on the image understanding task and the image generation task, the multimodal large language model and the diffusion model to be trained of the multimodal unified understanding large model are optimized.
9. The method according to claim 8, wherein Optimizing the multimodal large language model of the multimodal unified understanding large model for the image understanding task includes: generating a first semantic vector of a second sample image using the continuous visual encoder; Inputting the first semantic vector and a second sample text corresponding to the second sample image into the multimodal large language model to obtain a text response generated by the multimodal large language model; determining a text modality loss based on the text response and the ground truth of the text response; The multimodal large language model is optimized based on the text modality loss.
10. The method according to claim 8, wherein For the image generation task, optimizing the multimodal large language model and the to-be-trained diffusion model of the multimodal unified understanding large model includes: For the image-to-image task, determining a first initial vector based on the third sample image; Inputting the first initial vector into the second visual encoder of the continuous visual encoder to obtain a second semantic vector; Inputting the second semantic vector and the first text description of the image-to-image task into the multimodal large language model to obtain a first visual semantic vector output by the multimodal large language model; Inputting the first visual semantic vector into the diffusion model to be trained to obtain a second visual semantic vector; determining a second diffusion loss based on the first visual semantic vector and the second visual semantic vector; Based on the second diffusion loss, the multimodal large language model and the diffusion model to be trained are optimized.
11. The method according to claim 8, wherein For the image generation task, optimizing the multimodal large language model and the to-be-trained diffusion model of the multimodal unified understanding large model includes: For the Wensheng graph task, determine the second initial vector; Inputting the second initial vector into the second visual encoder of the continuous visual encoder to obtain a third semantic vector; Inputting the third semantic vector and the second text description of the text-to-image task into the multimodal large language model to obtain a third visual semantic vector output by the multimodal large language model; Inputting the third visual semantic vector into the diffusion model to be trained to obtain a fourth visual semantic vector; determining a third diffusion loss based on the fourth visual semantic vector and the third visual semantic vector; Based on the third diffusion loss, the multimodal large language model and the diffusion model to be trained are optimized.
12. A model training device comprising: a first feature extraction module, configured to input a first sample image into a first visual encoder to map the first sample image into a continuous latent space to obtain intermediate features; A second feature extraction module is used to input the intermediate features into the visual encoder to be trained to obtain the visual features to be optimized; a loss determination module, configured to determine a first training loss based on the visual feature to be optimized; an optimization module, configured to optimize the visual encoder to be trained based on the first training loss to obtain a continuous visual encoder including the first visual encoder and a second visual encoder; The second visual encoder is the visual encoder to be trained whose training converges.
13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Cited By
Visual content generation method and training method and device of visual content generation model
CN121792813A