Method and device for training visual simulator, storage medium and electronic equipment
Through the three-stage alignment algorithm, the visual simulator is trained, and the problem of insufficient graphic and text alignment data is solved, and the visual simulator is plug-and-play on multimodal large language model is realized, which is suitable for low-resource scenarios such as multimodal content audit.
Patent Information
- Application Number
- CN202510413409.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, the lack of a large amount of high-quality text alignment data during training of multimodal large language models of graphics and text, making it difficult to effectively train the network in specific scenarios, especially in scenarios such as multimodal content review, data acquisition is difficult, the labeling process is slow, and the iteration cycle is long.
The three-stage alignment algorithm is used to train the vision simulator, which includes aligning the visual features output by the frozen vision encoder with the visual simulation features output by the vision simulator, and then aligning with the output of the frozen multimodal large language model, and final alignment with the text description data and problem instructions to build a vision simulator that can be plug-and-play on the multimodal large language model.
It realizes the good adaptation of the vision simulator and the multimodal large language model, and can be directly used on the multimodal large language model, solving the problem of feature utilization differences between visual features and text features in the multimodal large model, and is suitable for low-resource multimodal task scenarios.
Smart Images

Figure CN120451700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer technology, and in particular to a method, device, storage medium and electronic equipment for training a visual simulator. Background Art
[0002] With the development of the times, there are many solutions for multimodal large language models of images and texts. In order to train a model with the ability to understand images and texts through a large language model, data becomes particularly important. In particular, a large amount of high-quality image-text alignment data is needed for targeted training of the network. However, in many specific scenarios in actual applications, it is often impossible to obtain a large amount of image-text alignment data. Summary of the Invention
[0003] The purpose of the embodiments of this specification is to provide a method, apparatus, storage medium, and electronic device for training a visual simulator.
[0004] The embodiments of this specification provide a method for training a visual simulator. By using a three-stage alignment algorithm to train the visual simulator, the trained visual simulator can be better adapted to a multimodal large language model, and the trained visual simulator can be plug-and-play on the multimodal large language model. The method includes:
[0005] Inputting first image data into a frozen visual encoder to obtain visual features output by the visual encoder, inputting first text description data corresponding to the first image data into a visual simulator to obtain visual simulation features output by the visual simulator, and performing a first stage of training on the visual simulator by aligning the visual features with the visual simulation features;
[0006] Inputting the first text description data and the instruction text into the visual simulator trained in the first stage, inputting the output data of the visual simulator into a frozen large multimodal language model, and performing a second stage of training on the visual simulator by aligning the output of the large multimodal language model with the first text description data;
[0007] The first text description data and its corresponding first question instruction are input into the visual simulator after the second stage training, the output data of the visual simulator is input into the multimodal large language model, and the output of the multimodal large language model is aligned with the first answer information corresponding to the first question instruction, and the visual simulator is trained in the third stage to obtain a trained visual simulator, wherein the output of the trained visual simulator can be aligned with the input features of the connector in the multimodal large language model.
[0008] Furthermore, the method further comprises:
[0009] Obtaining a plurality of text description data corresponding to the first image data;
[0010] The first text description data is determined from the plurality of text description data.
[0011] Furthermore, determining the first text description data from the plurality of text description data includes:
[0012] The first text description data is determined from the plurality of text description data according to the minimum text unit quantity corresponding to each text description data.
[0013] Furthermore, the method further comprises:
[0014] The second image data in the downstream training sample set is input into the multimodal large language model to obtain the second text description data corresponding to the second image data output by the multimodal large language model, and the second text description data and the prompt template are input into the target large language model to obtain the task data output by the target large language model, wherein the task data includes the newly generated third text description information, the third question instruction and the third answer information corresponding to the third text description information, and the task data is used to fine-tune the visual simulator for the downstream task.
[0015] Furthermore, the prompt template includes synthesis instructions, task specifications, format descriptions and context demonstrations, wherein the context demonstrations are selected from the downstream training sample set.
[0016] Furthermore, the method further comprises:
[0017] The third text description information and the third question instruction are input into the multimodal large language model, and the output of the multimodal large language model is aligned with the third answer information, and the visual simulator is fine-tuned for downstream tasks to obtain a fine-tuned visual simulator.
[0018] Furthermore, the method further comprises:
[0019] The third text description information and the prompt template are input into the target large language model again to iteratively obtain the task data output by the target large language model.
[0020] The embodiments of this specification also provide a method for simulating visual features, including:
[0021] Obtain target text description information corresponding to the multimodal task;
[0022] The target text description information is input into a trained visual simulator to obtain the target visual simulation features corresponding to the multimodal task output by the visual simulator, and the output of the visual simulator can be aligned with the input features of the connector in the multimodal large language model, wherein the visual simulator is trained based on the method for training a visual simulator described in the embodiments of this specification.
[0023] The present invention also provides a device for training a visual simulator, comprising:
[0024] a first training module, configured to input first image data into a frozen visual encoder to obtain visual features output by the visual encoder, input first text description data corresponding to the first image data into a visual simulator to obtain visual simulation features output by the visual simulator, and perform a first phase of training on the visual simulator by aligning the visual features with the visual simulation features;
[0025] a second training module, configured to input the first text description data and instruction text into the visual simulator trained in the first stage, input output data of the visual simulator into a frozen large multimodal language model, and perform a second stage of training on the visual simulator by aligning the output of the large multimodal language model with the first text description data;
[0026] The third training module is used to input the first text description data and its corresponding first question instruction into the visual simulator after the second stage training, input the output data of the visual simulator into the multimodal large language model, and perform third stage training on the visual simulator by aligning the output of the multimodal large language model with the first answer information corresponding to the first question instruction to obtain a trained visual simulator, wherein the output of the trained visual simulator can be aligned with the input features of the connector in the multimodal large language model.
[0027] The embodiments of this specification also provide a device for simulating visual features, including:
[0028] The first acquisition module is used to obtain target text description information corresponding to the multimodal task;
[0029] The second acquisition module is used to input the target text description information into a trained visual simulator to obtain the target visual simulation features output by the visual simulator corresponding to the multimodal task, and the output of the visual simulator can be aligned with the input features of the connector in the multimodal large language model, wherein the visual simulator is trained based on the method for training a visual simulator described in the embodiments of this specification.
[0030] An embodiment of this specification further provides a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.
[0031] An embodiment of this specification further provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.
[0032] An embodiment of this specification also provides a computer program product having at least one instruction stored thereon, wherein the at least one instruction implements the steps of the above method when executed by a processor.
[0033] According to the solution of the embodiments of this specification, a three-stage alignment algorithm is used to train the visual simulator. Specifically, the visual simulator is trained in the first stage by aligning the visual features output by the frozen visual encoder with the visual simulation features output by the visual simulator. The visual simulator is trained in the second stage by aligning the output of the frozen multimodal large language model with the first text description data. The visual simulator is trained in the third stage by aligning the output of the multimodal large language model with the first answer information corresponding to the first text description data to obtain a trained visual simulator. This enables the visual simulator to be better adapted to the multimodal large language model, and the trained visual simulator can be plug-and-play on the multimodal large language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A flowchart of a method for training a visual simulator provided in an embodiment of this specification.
[0035] Figure 2 This is a schematic diagram of a framework for training a visual simulator, which is an example provided in an embodiment of this specification.
[0036] Figure 3 A flowchart of a method for simulating visual features provided in an embodiment of this specification.
[0037] Figure 4 This is a schematic diagram of the structure of a device for training a visual simulator provided in an embodiment of this specification.
[0038] Figure 5 This is a schematic diagram of the structure of a device for simulating visual features provided in an embodiment of this specification.
[0039] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0040] To make the objectives, technical solutions, and advantages of this specification more clear, the following will clearly and completely describe the technical solutions of this specification in conjunction with the specific embodiments of this specification and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.
[0041] See Figure 1 , is a flow chart of a method for training a visual simulator provided in an embodiment of this specification. In an embodiment of this specification, the method for training a visual simulator is applied to a device for training a visual simulator (hereinafter referred to as a "training device") or an electronic device equipped with a training device. Figure 1 The process shown in FIG. 1 is described in detail. The method for training a visual simulator may specifically include the following steps:
[0042] S102, input the first image data into the frozen visual encoder to obtain the visual features output by the visual encoder, input the first text description data corresponding to the first image data into the visual simulator to obtain the visual simulation features output by the visual simulator, and perform the first stage training on the visual simulator by aligning the visual features with the visual simulation features.
[0043] In some embodiments, the visual simulator includes a text encoder and a text projector, the output of the text encoder is the input of the text projector, and the target features aligned by the visual simulator are the input features of the connector in the multimodal large language model (MLLMs). As an example, the visual simulator consists of a text encoder and a text projector, the text encoder uses the T5 (Text-to-Text Transfer Transformer) model architecture, and the text projector uses an MLP (Multilayer Perceptron).
[0044] In some embodiments, the frozen visual encoder indicates that the weights in the visual encoder have been frozen. In some embodiments, during the first stage of training, the features of all patches (blocks) of the visual simulator and the frozen visual encoder are directly aligned coarsely through the MSE (Mean-Square Error) loss, so that the visual simulator can preliminarily simulate the image input through text description data (the text description data is also referred to as "caption" in this context), thereby achieving patch-level alignment of the visual simulator and the visual encoder after the first stage of training. In some embodiments, the loss function L1 in the first stage of training is shown as follows:
[0045] L1=L MSE (H v ,H c ,θ)
[0046] Among them H v represents the visual features output from the visual encoder (i.e., the visual features from the image data), H c represents the visual simulation features output by the visual simulator (i.e., the visual simulation features from the text description data), θ represents the parameters of the trainable text encoder and text projector, and L MSE Represents the MSE loss function.
[0047] In some embodiments, for different application scenarios, first image data corresponding to the application scenario is collected and first text description data corresponding to the first image data is obtained as training data used in the first stage of training. In some embodiments, after obtaining the first image data, various implementation methods can be used to obtain the first text description data corresponding to the first image data. For example, the first text description data corresponding to the first image data can be obtained based on image processing technology, or the first text description data of the first image data can be obtained based on manual annotation, or the first image data can be input into a large language model to obtain the first text description data corresponding to the first image data. This application does not limit this.
[0048] S104: Input the first text description data and instruction text into the visual simulator after the first stage training, input the output data of the visual simulator into the frozen multimodal large language model, and perform the second stage training on the visual simulator by aligning the output of the multimodal large language model with the first text description data.
[0049] In some embodiments, the frozen multimodal large language model includes a frozen connector and a frozen large language model. The connector is connected in series with the large language model, and the output of the connector serves as the input to the large language model. In some embodiments, a caption reconstruction task is introduced during the second phase of training. After freezing the multimodal large language model, the text encoder and text projector in the visual simulator can be fine-tuned with multimodal instructions. Through the caption reconstruction task, the pseudo-visual features of the visual simulator can be fine-grained, allowing them to be effectively processed and utilized by the large language model. In some embodiments, the instruction text is used to instruct a detailed text description of the corresponding image. In some embodiments, the caption and instruction text are input into the visual simulator, and the image features are replaced with the visual simulation features simulated by the caption. These are then input into the frozen multimodal large language model along with the instruction text, for example, the instruction text is "Describe this image in detail." The response output by the large language model is then aligned with the caption input to the visual simulator to complete the caption reconstruction task. In some embodiments, the second phase of alignment uses the same data as the first phase. In some embodiments, the second phase of training utilizes a cross-entropy loss function. After the second phase of training, the visual simulator and the large language model are pre-aligned.
[0050] S106: Input the first text description data and its corresponding first question instruction into the visual simulator after the second stage training, input the output data of the visual simulator into the multimodal large language model, and perform third stage training on the visual simulator by aligning the output of the multimodal large language model with the first answer information corresponding to the first question instruction to obtain a trained visual simulator, wherein the output of the trained visual simulator can be aligned with the input features of the connector in the multimodal large language model.
[0051] In some embodiments, first, task data based on text descriptions for the third stage of training is obtained, and the task data includes first text description data and its corresponding first question instruction and first answer information. The first text description data and its corresponding first question instruction in the task data are input into the visual simulator after the second stage of training, and the output data of the visual simulator is input into the frozen multimodal large language model to obtain the output of the large language model in the multimodal large language model. The loss function is constructed by comparing the output with the first answer information corresponding to the first question instruction to perform the third stage of training on the visual simulator. In some embodiments, the third stage training process adopts a cross-entropy loss function. In some embodiments, the task data used for the third stage of training (the first question instruction and the first answer information corresponding to the first text description data) are obtained based on the caption data generation link described below.
[0052] In the third stage, the text encoder and text projection layer in the visual simulator are fine-tuned with multimodal instructions on multiple different VQA (Visual Question Answering) tasks, and the visual features of the caption simulation and the first question instruction are input into the multimodal large language model to obtain the corresponding answer. The purpose of the third stage of training is to make the visual features of the caption simulation applicable to a wide range of multimodal scenarios, and to complete the understanding and reasoning of multimodal tasks through the frozen multimodal large language model, thereby achieving reasoning level alignment with the large language model. It should be noted that since the output of the trained visual simulator obtained through three stages of training can be aligned with the input features of the connector in the multimodal large language model, the trained visual simulator can be plug-and-play on the multimodal large language model.
[0053] According to the solution of the embodiments of this specification, a three-stage alignment algorithm is used to train the visual simulator. Specifically, the visual simulator is trained in the first stage by aligning the visual features output by the frozen visual encoder with the visual simulation features output by the visual simulator. The visual simulator is trained in the second stage by aligning the output of the frozen multimodal large language model with the first text description data. The visual simulator is trained in the third stage by aligning the output of the multimodal large language model with the first answer information corresponding to the first text description data to obtain a trained visual simulator. This enables the visual simulator to be better adapted to the multimodal large language model, and the trained visual simulator can be plug-and-play on the multimodal large language model.
[0054] This application found that in many scenarios, it is often impossible to obtain a large amount of image-text alignment data. For example, in the multimodal content review scenario, there is a long-tail distribution problem. It is difficult to obtain a large amount of multimodal data for various risk factors, such as special events, implicit expressions, and implied risk identification. Emerging multimodal review needs and scenarios may not have historical data to learn from and refer to, and rely on manual collection of multimodal data (images) and manual labeling and matching. The labeling process is slow, the iteration cycle is long, and there is also a long-tail distribution problem. According to the solution of the embodiments of this specification, a visual simulator that can be plug-and-play in a multimodal large language model can be constructed. The visual simulator can use the text description data corresponding to the limited multimodal images to simulate visual features, thereby using the text description data to replace the multimodal image data in the multimodal large language model to complete the training. This solution is particularly suitable for application scenarios where multimodal data is missing or the multimodal data is limited, that is, it is suitable for multimodal tasks under low-resource settings. For example, for the above-mentioned multimodal content review scenario, the visual features are simulated based on the trained visual simulator to complete the training of the multimodal large language model, which can achieve multimodal content review in low-resource scenarios without relying on the risk factor injection method of a large amount of multimodal data; and based on the solution of the embodiments of this specification, by adopting a three-stage progressive alignment strategy to align a caption-based visual simulator, it can better adapt to the multimodal large language model, solving the problem of feature utilization differences between visual features and text features in multimodal large models in the prior art.
[0055] In some embodiments, the method further includes: obtaining a plurality of text description data corresponding to the first image data; and determining the first text description data from the plurality of text description data. In some embodiments, the plurality of text description data corresponding to the first image data can be obtained by the same implementation method, or by different implementation methods, and this specification does not limit this. In some embodiments, the first text description data for training can be randomly determined from a plurality of text description data. In some embodiments, the first text description data can be selected from a plurality of text description data based on predetermined rules, and the selection factors associated with the predetermined rules include but are not limited to time factors, text unit quantity factors, visual feature quantity factors, etc. For example, the text description data containing the largest number of visual features can be selected from a plurality of text description data based on predetermined rules as the first text description data.
[0056] In some embodiments, the determining of the first text description data from the multiple text description data includes: determining the first text description data from the multiple text description data based on the minimum number of text units corresponding to each text description data. In some embodiments, based on the minimum number of text units corresponding to each text description data, selecting the one with the largest number of corresponding minimum text units from the multiple text description data as the first text description data. In some embodiments, based on the minimum number of text units corresponding to each text description data, at least one candidate text description data whose minimum number of text units exceeds a preset threshold is determined, and then the first text description data is determined from the at least one candidate text description data (which can be determined randomly or in combination with other factors). By filtering based on the minimum number of text units corresponding to the caption (that is, the number of tokens), more detailed captions can be retained for training the visual simulator. In some embodiments, the text encoder in the visual simulator adopts a T5 model structure with relative position encoding, combined with a filtering scheme based on the minimum number of text units corresponding to the caption, which can obtain text description data of sufficient length and align it with the visual features in the multimodal large language model, thereby solving the length utilization difference problem that may exist in the prior art (for example, in the CLIP (Contrastive Language-Image Pre-Training) scheme based on image-text alignment in the prior art, the text position embedding of the text-transformer is only 77, and relevant research shows that the actual effective text length is even less than 20).
[0057] In some embodiments, the method further includes: inputting the second image data in the downstream training sample set into the multimodal large language model, obtaining second text description data corresponding to the second image data output by the multimodal large language model, inputting the second text description data and a prompt template into a target large language model, and obtaining task data output by the target large language model, wherein the task data includes newly generated third text description information, a third question instruction corresponding to the third text description information, and third answer information, and the task data is used to fine-tune the visual simulator for the downstream task. In some embodiments, the downstream training sample set includes the second image data collected for the downstream task. In some embodiments, for a low-resource multimodal task that can only obtain a small amount of multimodal data containing images, the existing multimodal large language model is used to generate detailed second text description data for the available image (i.e., the second image data). Then, based on this second text description data, the target large language model can be further used to construct new third text description data and corresponding task data. In this way, a generation chain of task data for caption enhancement (which may also be referred to as a "caption data generation chain" in this context) can be constructed to obtain more task data for fine-tuning the low-resource multimodal downstream task. As an example, for a multimodal content review scenario, limited multimodal data (i.e., second image data) in this scenario is obtained, the second image data and its corresponding visual question-answering task data are input into an existing multimodal large language model, and the second text description data and its corresponding visual question-answering task data are obtained. Then, the target large language model is used to construct new third text description data and corresponding task data. In some embodiments, the prompt template is used to drive the target large language model to synthesize the new third text description data.
[0058] In some embodiments, the prompt template includes synthesis instructions, task specifications, format instructions and context demonstrations, wherein the context demonstration is selected from the downstream training sample set. In some embodiments, the task specification is used to specify a specific multimodal task scenario or theme to guide the target large language model to generate data that fits the task background, and the format instructions guide the target large language model to generate data in the same format as the task data. In some embodiments, multiple data samples are randomly selected from the downstream training sample set as context demonstrations, and the target large language model is explicitly guided to learn from real data to improve the fidelity of the synthesized data; in some embodiments, the synthesized data is continuously added to the multimodal task data as a feasible supplement, and serves as a new optional context demonstration.
[0059] In some embodiments, the method further includes: inputting the third text description information and the third question instruction into the multimodal large language model, and fine-tuning the visual simulator for the downstream task by aligning the output of the multimodal large language model with the third answer information to obtain a fine-tuned visual simulator. In some embodiments, the new task data generated based on the caption data generation link is used in the third stage of training of the visual simulator to fine-tune the visual simulator for a specific downstream task.
[0060] In some embodiments, the method further includes: re-inputting the third text description information and the prompt template into the target large language model to iteratively obtain task data output by the target large language model. In some embodiments, the prompt template can be adjusted based on the third text description information currently output by the target large model, thereby continuously iterating to obtain task data that meets the requirements.
[0061] It should be noted that by combining limited multimodal data and new caption-enhanced data continuously generated through the caption data generation link, it is possible to use an aligned caption-based visual simulator to more fully fine-tune the multimodal large language model for low-resource downstream tasks, allowing the multimodal large language model to benefit from the caption task data.
[0062] Figure 2 This is a schematic diagram of a framework for training a visual simulator, which is an example provided in an embodiment of this specification. Figure 2The three-stage alignment algorithm framework of the visual simulator is exemplified, specifically: the first stage is patch-level alignment with vision, wherein the image data is input into a frozen visual encoder, and the text description data corresponding to the image data is input into the visual simulator, which includes a text encoder and a text projector connected in series, and then the features of all patches of the visual simulator and the frozen visual encoder are coarse-grainedly aligned through the MSE loss; the second stage is pre-alignment with the large language model, freezing the connector and the large language model (that is, freezing the multimodal large language model), replacing the image features with the visual features simulated by the text description data, and inputting them into the frozen multimodal large language model together with the instruction text (not shown), wherein the instruction text The original text is "Describe the picture in detail", and then the response of the large language model is aligned with the input text description data based on the text reconstruction loss (that is, the cross entropy loss function used in the second stage); the third stage is to align with the inference level of the large language model, and input the text-based task data into the visual encoder (in the actual training process, the first description text data and the corresponding first question instruction are input into the visual encoder, and the first answer information in the task data is used to compare with the answer output by the large language model) to input the simulated visual features and question instructions into the large language model to obtain the corresponding answer, and then align the output of the large language model with the answer information of the task data based on the multimodal task loss (that is, the cross entropy loss function used in the third stage).
[0063] Figure 3 A flowchart of a method for simulating visual features provided in an embodiment of this specification is provided. In an embodiment of this specification, the method for simulating visual features is applied to a device for simulating visual features (hereinafter referred to as a "simulation device") or an electronic device equipped with a simulation device. The method for simulating visual features may specifically include the following steps:
[0064] S202: Obtain target text description information corresponding to the multimodal task.
[0065] In some embodiments, for a multimodal task in a specific application scenario, image data corresponding to the multimodal task is collected, and target text description information corresponding to the image data is obtained. In some embodiments, the implementation method for obtaining the target text description information corresponding to the multimodal task is the same or similar to the implementation method for obtaining the first text description information described above, and will not be repeated here.
[0066] S204, inputting the target text description information into a trained visual simulator, obtaining the target visual simulation features output by the visual simulator corresponding to the multimodal task, wherein the output of the visual simulator can be aligned with the input features of the connector in the multimodal large language model, wherein the visual simulator is trained based on the method for training a visual simulator described in the embodiments of this specification.
[0067] In some embodiments, based on the visual simulator trained in the embodiments of this specification, simulated target visual simulation features can be obtained based on target text description information, so that text data can be used instead of multimodal image data to complete training in MLLMs.
[0068] Figure 4 This is a schematic diagram of a device for training a visual simulator, provided in an embodiment of this specification. This device for training a visual simulator (hereinafter referred to as "training device 1") can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, training device 1 includes a first training module 11, a second training module 12, and a third training module 13.
[0069] The first training module 11 is used to input the first image data into a frozen visual encoder to obtain the visual features output by the visual encoder, input the first text description data corresponding to the first image data into a visual simulator to obtain the visual simulation features output by the visual simulator, and perform the first stage training on the visual simulator by aligning the visual features with the visual simulation features.
[0070] The second training module 12 is used to input the first text description data and instruction text into the visual simulator after the first stage training, input the output data of the visual simulator into the frozen multimodal large language model, and perform the second stage training on the visual simulator by aligning the output of the multimodal large language model with the first text description data.
[0071] The third training module 13 is used to input the first text description data and its corresponding first question instruction into the visual simulator after the second stage training, input the output data of the visual simulator into the multimodal large language model, and perform third stage training on the visual simulator by aligning the output of the multimodal large language model with the first answer information corresponding to the first question instruction to obtain a trained visual simulator, wherein the output of the trained visual simulator can be aligned with the input features of the connector in the multimodal large language model.
[0072] In some embodiments, the training device 1 is further used to: obtain a plurality of text description data corresponding to the first image data; and determine the first text description data from the plurality of text description data.
[0073] In some embodiments, determining the first text description data from the plurality of text description data includes: determining the first text description data from the plurality of text description data according to the minimum number of text units corresponding to each text description data.
[0074] In some embodiments, the training device 1 is also used to: input the second image data in the downstream training sample set into the multimodal large language model, obtain the second text description data corresponding to the second image data output by the multimodal large language model, input the second text description data and the prompt template into the target large language model, and obtain the task data output by the target large language model, wherein the task data includes newly generated third text description information, a third question instruction corresponding to the third text description information, and third answer information, and the task data is used to fine-tune the visual simulator for downstream tasks.
[0075] In some embodiments, the prompt template includes synthesis instructions, task specifications, format instructions, and context demonstration, wherein the context demonstration is selected from the downstream training sample set.
[0076] In some embodiments, the training device 1 is also used to: input the third text description information and the third question instruction into the multimodal large language model, and fine-tune the visual simulator for downstream tasks by aligning the output of the multimodal large language model with the third answer information to obtain a fine-tuned visual simulator.
[0077] In some embodiments, the training device 1 is further used to: input the third text description information and the prompt template into the target large language model again to iteratively obtain task data output by the target large language model.
[0078] Figure 5 This is a schematic diagram of the structure of a device for simulating visual features provided in an embodiment of this specification. This device for simulating visual features (hereinafter referred to as "simulation device 2") can be implemented as all or part of an electronic device through software, hardware, or a combination of both. According to some embodiments, simulation device 2 includes a first obtaining module 21 and a second obtaining module 22.
[0079] The first obtaining module 21 is used to obtain target text description information corresponding to the multimodal task.
[0080] The second acquisition module 22 is used to input the target text description information into a trained visual simulator to obtain the target visual simulation features output by the visual simulator corresponding to the multimodal task, and the output of the visual simulator can be aligned with the input features of the connector in the multimodal large language model, wherein the visual simulator is trained based on the method for training a visual simulator described in the embodiments of this specification.
[0081] The above-mentioned device embodiments correspond to the method embodiments. For detailed descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For detailed descriptions, please refer to the corresponding method embodiments.
[0082] The embodiments of this specification also provide a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor and executing the method of the embodiments of this specification.
[0083] An embodiment of the present specification further provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded by the processor to execute the method of the embodiment of the present specification.
[0084] The embodiments of this specification also provide Figure 6 The structural diagram of the electronic device shown in FIG. Figure 6 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its operations. The processor reads the corresponding computer program from the non-volatile storage into the memory and then runs it to implement the above method.
[0085] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0086] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0087] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0088] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0089] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0090] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0091] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.
[0092] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0093] The foregoing is merely an example of the present invention and is not intended to limit the present invention. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A method for training a visual simulator, comprising: Inputting first image data into a frozen visual encoder to obtain visual features output by the visual encoder, inputting first text description data corresponding to the first image data into a visual simulator to obtain visual simulation features output by the visual simulator, and performing a first stage of training on the visual simulator by aligning the visual features with the visual simulation features; Inputting the first text description data and the instruction text into the visual simulator trained in the first stage, inputting the output data of the visual simulator into a frozen large multimodal language model, and performing a second stage of training on the visual simulator by aligning the output of the large multimodal language model with the first text description data; The first text description data and its corresponding first question instruction are input into the visual simulator after the second stage training, the output data of the visual simulator is input into the multimodal large language model, and the output of the multimodal large language model is aligned with the first answer information corresponding to the first question instruction, and the visual simulator is trained in the third stage to obtain a trained visual simulator, wherein the output of the trained visual simulator can be aligned with the input features of the connector in the multimodal large language model.
2. The method according to claim 1, further comprising: Obtaining a plurality of text description data corresponding to the first image data; The first text description data is determined from the plurality of text description data.
3. The method according to claim 2, wherein determining the first text description data from the plurality of text description data comprises: The first text description data is determined from the plurality of text description data according to the minimum text unit quantity corresponding to each text description data.
4. The method according to claim 1, further comprising: The second image data in the downstream training sample set is input into the multimodal large language model to obtain the second text description data corresponding to the second image data output by the multimodal large language model, and the second text description data and the prompt template are input into the target large language model to obtain the task data output by the target large language model, wherein the task data includes the newly generated third text description information, the third question instruction and the third answer information corresponding to the third text description information, and the task data is used to fine-tune the visual simulator for the downstream task.
5. The method according to claim 4, wherein the prompt template includes synthesis instructions, task specifications, format instructions and context demonstration, wherein: The context demonstration is selected from the downstream training sample set.
6. The method according to claim 4, further comprising: The third text description information and the third question instruction are input into the multimodal large language model, and the output of the multimodal large language model is aligned with the third answer information, and the visual simulator is fine-tuned for downstream tasks to obtain a fine-tuned visual simulator.
7. The method according to claim 4, further comprising: The third text description information and the prompt template are input into the target large language model again to iteratively obtain the task data output by the target large language model.
8. A method for simulating visual features, comprising: Obtain target text description information corresponding to the multimodal task; The target text description information is input into a trained visual simulator to obtain target visual simulation features corresponding to the multimodal task output by the visual simulator, wherein the output of the visual simulator can be aligned with the input features of the connector in the multimodal large language model, wherein the visual simulator is trained based on the method described in any one of claims 1 to 7.
9. A device for training a visual simulator, comprising: a first training module, configured to input first image data into a frozen visual encoder to obtain visual features output by the visual encoder, input first text description data corresponding to the first image data into a visual simulator to obtain visual simulation features output by the visual simulator, and perform a first phase of training on the visual simulator by aligning the visual features with the visual simulation features; a second training module, configured to input the first text description data and instruction text into the visual simulator trained in the first stage, input output data of the visual simulator into a frozen large multimodal language model, and perform a second stage of training on the visual simulator by aligning the output of the large multimodal language model with the first text description data; The third training module is used to input the first text description data and its corresponding first question instruction into the visual simulator after the second stage training, input the output data of the visual simulator into the multimodal large language model, and perform third stage training on the visual simulator by aligning the output of the multimodal large language model with the first answer information corresponding to the first question instruction to obtain a trained visual simulator, wherein the output of the trained visual simulator can be aligned with the input features of the connector in the multimodal large language model.
10. A device for simulating visual features, comprising: The first acquisition module is used to obtain target text description information corresponding to the multimodal task; The second acquisition module is used to input the target text description information into a trained visual simulator to obtain the target visual simulation features corresponding to the multimodal task output by the visual simulator, and the output of the visual simulator can be aligned with the input features of the connector in the multimodal large language model, wherein the visual simulator is trained based on the method described in any one of claims 1 to 7.
11. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 8.
13. A computer program product having at least one instruction stored thereon, characterized in that: When the at least one instruction is executed by the processor, the steps of the method according to any one of claims 1 to 8 are implemented.