Task understanding method and device for images and model training method and device
By introducing a back-viewing and attention mechanism into the large model, the fine-grained features in multiple maps are dynamically retrieved, and the problem of insufficient information focus in the multi-graph understanding task in the existing technology is solved, achieving more efficient and accurate image understanding.
Patent Information
- Application Number
- CN202510290472.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art is difficult to focus key information based on specific tasks in multi-picture understanding tasks, resulting in inaccuracy and inefficiency of image or video understanding, especially in application scenarios such as property claims.
The back-looking and review mechanism is adopted, and after a global understanding is carried out through a large model, fine-grained features are dynamically retrieved and extracted based on user task instructions. Combined with attention mechanism and visual language adapter, humans are simulated to re-examine images with tasks and dynamically retrieve key information related to the task.
It improves the accuracy and efficiency of multi-graph understanding tasks, can better focus on key information, reduce redundant content, and improves the accuracy and efficiency of user task processing.
Smart Images

Figure CN120279387A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of image processing technology, and in particular, to a method for task understanding of images, a model training method, and a device therefor. Background Art
[0002] Computer vision is a technology that enables a computer to understand and analyze visual information such as pictures, similar to the ability of the human "eyes" and "brain" to process vision. Machine learning, especially deep learning, provides powerful algorithms and model support for computer vision, and can automatically extract features from images and learn. For example, a convolutional neural network (CNN) can automatically extract hierarchical features of an image and effectively process complex visual tasks, such as identifying damaged items or detecting anomalies. Moreover, large models have also been applied to the field of image processing and combined with models such as convolutional neural networks, promoting the rapid development of artificial intelligence in the field of vision and bringing transformative technological progress to many fields such as property claims settlement, criminal investigation, security, and medical imaging. During the process of data processing, the processing process of the model also protects the privacy of the involved private data.
[0003] Currently, there is a need for an improved solution that can enhance the understanding ability of multiple images or videos and the accuracy of executing user task instructions during the process of user task processing. Summary of the Invention
[0004] One or more embodiments of this specification describe a method for task understanding of images, a model training method, and a device therefor, to enhance the understanding ability of multiple images or videos and the accuracy of executing user task instructions during the process of user task processing. The specific technical solutions are as follows.
[0005] In a first aspect, an embodiment provides a method for task understanding of images, the method comprising:
[0006] extracting multiple fine-grained features of a set of images to be processed;
[0007] performing global feature extraction on the set of images to obtain an initial global understanding;
[0008] inputting the initial global understanding and a user task instruction into a large model, and determining the content to be reviewed through the large model;
[0009] determining a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features;
[0010] inputting the first fine-grained feature into the large model, and outputting the understanding content of the several images under the user task instruction through the large model.
[0011] In one embodiment, the step of extracting multiple fine-grained features of a set of images to be processed includes: for each image in the set of images, dividing the image into a plurality of sub-blocks, visually encoding the plurality of sub-blocks respectively to obtain corresponding visual features, and using the visual features of the plurality of sub-blocks as the multiple fine-grained features of the image;
[0012] The step of determining a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features includes: through a pre-trained review module, based on the attention parameters included in the review module, using an attention mechanism to determine relevant visual features for the content to be reviewed from the multiple fine-grained features, and converting the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0013] In one embodiment, the attention parameters include a first parameter, a second parameter, and a third parameter; the step of obtaining the first fine-grained feature includes:
[0014] Obtaining a query vector based on the product of the first parameter and the content to be reviewed;
[0015] Obtaining a key-value pair corresponding to each fine-grained feature based on the products of the second parameter and the third parameter respectively and the multiple fine-grained features;
[0016] Determining corresponding attention coefficients based on the products of the query vector and the keys in the multiple key-value pairs;
[0017] Obtaining the first fine-grained feature based on the products of multiple attention coefficients and the values in the multiple key-value pairs.
[0018] In one embodiment, the initial global understanding is a feature in the feature space of the large model; the step of performing global feature extraction on the set of images includes:
[0019] Performing overall visual encoding on the set of images respectively to obtain visual features of each image, and splicing the visual features of the set of images to obtain a global multi-image representation of the set of images;
[0020] Mapping the global multi-image representation to the feature space through a pre-trained vision-language adapter to obtain the initial global understanding.
[0021] In one embodiment, the step of determining the content to be reviewed by the large model includes:
[0022] Based on the understanding of the initial global understanding and the user task instructions, the large model generates a review identifier during the process of outputting prediction content, and the review identifier is used to indicate the content to be reviewed retrieved from the large model;
[0023] Based on the review identifier, retrieve the content to be reviewed from the large model.
[0024] In one implementation, the step of retrieving the content to be reviewed from the large model based on the review identifier includes:
[0025] When it is detected that the large model generates the review identifier during the process of outputting prediction content, input the first number of preset learnable embedding vectors into the large model, and obtain the first number of learned embedding vectors determined by the large model based on the first number of learnable embedding vectors, and determine the first number of learned embedding vectors as the content to be reviewed; the learnable embedding vectors are used to indicate the large model to extract key information that needs to be reviewed.
[0026] In one implementation, the method further includes: adding the review identifier to the vocabulary of the large model in advance.
[0027] In one implementation, the large model is an autoregressive large model; the content to be reviewed is not used as the output content of the large model.
[0028] In a second aspect, an embodiment provides a model training method, including:
[0029] Obtain a set of image samples and the standard understanding content for this set of image samples under the user task instructions;
[0030] Extract multiple fine-grained features of the set of image samples;
[0031] Perform global feature extraction on the set of image samples to obtain an initial global understanding;
[0032] Input the initial global understanding and the user task instructions into the large model, and determine the content to be reviewed through the large model;
[0033] Determine the first fine-grained feature related to the content to be reviewed from the multiple fine-grained features;
[0034] Input the first fine-grained feature into the large model, and output the understanding content for the set of image samples under the user task instructions through the large model;
[0035] Based on the difference between the understanding content and the standard understanding content, determine the prediction loss;
[0036] Fine-tune the large model at least based on the predicted loss.
[0037] In one embodiment, the standard understanding content includes a look-back identifier, and the look-back identifier is added before the key information in the standard understanding content.
[0038] In one embodiment, the multiple fine-grained features are visual features;
[0039] The step of determining, from the multiple fine-grained features, a first fine-grained feature related to the content to be looked back includes:
[0040] Through a look-back module, based on the attention parameters included in the look-back module, use an attention mechanism to determine, from the multiple fine-grained features, relevant visual features for the content to be looked back, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0041] In one embodiment, the initial global understanding is a feature in the feature space of the large model; the step of performing global feature extraction on the set of image samples includes:
[0042] Perform overall visual encoding on each of the set of image samples to obtain visual features of each image, and splice the visual features of this set of images to obtain a global multi-image representation of this set of images;
[0043] Map the global multi-image representation to the feature space through a vision-language adapter to obtain the initial global understanding.
[0044] In one embodiment, the model training method includes two stages; the step of fine-tuning the large model at least based on the predicted loss includes:
[0045] In the first stage, update the parameters of the look-back module and the vision-language adapter based on the predicted loss, without updating the parameters of the large model;
[0046] In the second stage, perform parameter fine-tuning on the look-back module, the vision-language adapter, and the large model based on the predicted loss.
[0047] In a third aspect, an embodiment provides a task understanding device for images, including:
[0048] A fine-grained module configured to extract multiple fine-grained features of a set of images to be processed;
[0049] A global module configured to perform global feature extraction on the set of images to obtain an initial global understanding;
[0050] A determination module, configured to input the initial global understanding and the user task instruction into a large model, and determine the content to be reviewed through the large model;
[0051] A review module, configured to determine a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features;
[0052] An understanding module, configured to input the first fine-grained feature into the large model, and output the understanding content of the several images under the user task instruction through the large model.
[0053] In a fourth aspect, an embodiment provides a model training device, including:
[0054] An acquisition module, configured to acquire a set of image samples and the standard understanding content of this set of image samples under a user task instruction;
[0055] An extraction module, configured to extract multiple fine-grained features of the set of image samples;
[0056] An initial module, configured to perform global feature extraction on the set of image samples to obtain an initial global understanding;
[0057] An input module, configured to input the initial global understanding and the user task instruction into a large model, and determine the content to be reviewed through the large model;
[0058] A review module, configured to determine a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features;
[0059] A prediction module, configured to input the first fine-grained feature into the large model, and output the understanding content of the set of image samples under the user task instruction through the large model;
[0060] A loss module, configured to determine a prediction loss based on the difference between the understanding content and the standard understanding content;
[0061] A training module, configured to at least fine-tune the large model based on the prediction loss.
[0062] In a fifth aspect, an embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method according to any one of the first aspect to the second aspect.
[0063] In a sixth aspect, an embodiment provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method according to any one of the first aspect to the second aspect is implemented.
[0064] In the method and device provided in the embodiments of this specification, image features are divided into fine-grained features and global features. Through the understanding of the user task instructions by the large model and the understanding of the global features of the image, the content to be reviewed is determined therefrom, that is, the large model is used to determine, like a human, the places that need to be carefully viewed and understood from multiple images, and the first fine-grained features related to the content to be reviewed are determined from the fine-grained features. Furthermore, the large model continues to implement a detailed understanding of the multiple images based on the first fine-grained features. Through this method, the embodiments can improve the understanding ability of multiple images or videos during the user task processing, as well as the accuracy when executing user task instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0066] Figure 1 It is a schematic diagram of the implementation scenario of an embodiment disclosed in this specification;
[0067] Figure 2 It is a schematic flowchart of a model training method provided by the embodiment;
[0068] Figure 3 It is a schematic flowchart of a process when extracting image features provided by the embodiment;
[0069] Figure 4 It is a flowchart of an execution method of the review module provided by the embodiment;
[0070] Figure 5 It is a schematic flowchart of a task understanding method for an image provided by the embodiment;
[0071] Figure 6 It is a schematic block diagram of a task understanding device for an image provided by the embodiment;
[0072] Figure 7 It is a schematic block diagram of a model training device provided by the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The following describes the solutions provided in this specification in conjunction with the drawings.
[0074] Figure 1Schematic diagram of an implementation scenario of an embodiment disclosed in this specification. It includes a computing device, a visual memory bank, and a large model. The visual memory bank and the large model can be implemented by other devices respectively, or can be implemented in the computing device. When it is necessary to understand multiple images under a user task instruction, the computing device can extract the fine-grained features of the multiple images and store them in the visual memory bank. At the same time, the computing device also performs an initial global understanding of the multiple images, and inputs the initial global understanding and the user task instruction into the large model to determine the content to be reviewed therein through the large model. The computing device obtains the content to be reviewed from the large model, retrieves the fine-grained features related to the content to be reviewed from the visual memory bank, and inputs the fine-grained features into the large model. The large model determines the understanding content of the multiple images based on the initial global understanding, the user task instruction, and the retrieved fine-grained features.
[0075] Figure 1 This is only an example of the application of this application in an implementation scenario. The implementation scenarios in actual applications may be different from or similar to this scenario. The following will detail the related concepts, problems to be solved, and methods provided by this application.
[0076] The large model is short for the Large Language Model (LLM). The large model is a type of natural language processing model based on deep learning, usually consisting of billions to hundreds of billions of parameters. Through training on large-scale text data, the model can understand and generate natural language. The large language model has a corresponding feature space and can process the feature vectors in this feature space. The feature space of the large model is also called the text space or text vector space. Text can be mapped to this feature space to obtain the embedding vector corresponding to this text. The embedding vectors in the feature space and the text do not correspond one-to-one. The feature space is a high-dimensional space that contains more embedding vectors than all the embedding vectors of the text.
[0077] The embedding vectors mentioned in this application, which can also be called tokens or feature vectors, are the basic units for the language model (including the large model) to process text. They facilitate the training and inference of the model by means of text segmentation, reducing model complexity, and capturing language structures. At the same time, the flexible use of tokens also enables the language model to adapt to different language and task requirements.
[0078] The large model applies multi-modal artificial intelligence technology and can well understand and generate information containing multiple modalities such as text, images, and speech through deep learning algorithms, and can be used to process multi-image understanding tasks.
[0079] Multi-modal artificial intelligence refers to the artificial intelligence (AI) technology that can process and understand multiple different types of data (such as text, images). For example, it can understand video content through visual and language information simultaneously.
[0080] Deep learning is a machine learning method based on artificial neural networks. By constructing a multi-layer network structure to simulate the learning process of the human brain, it automatically learns features and patterns from a large amount of data and is widely used in fields such as computer vision and natural language processing.
[0081] The multi-image understanding task refers to analyzing and processing content containing a large number of pictures, including scene classification, event recognition, key information localization, etc., to extract semantic information in the images or achieve automated high-level understanding.
[0082] When this application performs feature processing on multi-images, it uses a visual encoder and a vision-language adapter for image encoding and image feature alignment.
[0083] A visual encoder is a neural network module used to extract image features. By encoding visual input data, it converts it into a feature representation in the latent space.
[0084] A vision-language adapter is a model component that connects visual (image) and language (text) data. It can convert visual features into language data, that is, it can convert visual features into feature vectors in the feature space of the large model, align the feature space of the large model, and promote information interaction and alignment in multi-modal tasks. The vision-language adapter can be implemented using models such as a multi-layer perceptron (MLP).
[0085] When the large model performs the multi-image understanding task, it uses an improved new RAG technology. Retrieval-Augmented Generation (RAG) is an AI technology that combines retrieval and generation. When the traditional RAG technology uses AI to generate content, it first looks up relevant information from a specific database and then combines this information to generate more accurate and well-founded answers. For example, when asking the AI a specific question, it will first retrieve and review relevant materials from the database and then generate an answer instead of guessing out of thin air. This application provides a new retrieval-augmented generation technology. Before generating content, the large model first understands the image as a whole, then dynamically determines the details that need to be viewed during the content generation process according to specific user task instructions, then retrieves and extracts relevant information from the visual memory bank, and then the large model combines the retrieved information to generate a more detailed, more realistic, and more accurate understanding.
[0086] Retrieving and extracting relevant information from the visual memory bank can be performed based on the attention mechanism. The attention mechanism is a technique widely used in deep learning models, mainly used to solve the problem of focus selection in information processing. Its core idea is that among a large amount of input information, it is possible to selectively focus on more important parts instead of processing all information equally.
[0087] The multiple images to be processed can be understood as a set of images, and this set of images contains at least one image. There is a correlation among the multiple images in this set of images. For example, it can be a video, or pictures of different scenes under the same theme.
[0088] User task instructions have different meanings and types in different application scenarios. In the property claim settlement scenario, user task instructions may be "locate damaged items" or "identify property loss details", "locate the damaged area", "where is the most severely damaged part", etc. In the security scenario, user task instructions may be "locate the situation of people falling" "locate the fire", etc. In short, user task instructions are the target task descriptions indicating extraction from this set of images.
[0089] Taking the property claim settlement scenario as an example to illustrate the problems faced when the present application understands multiple images. In the property insurance claim settlement business, claims adjusters need to analyze a large number of on-site pictures and multi-picture reports submitted by users to confirm the loss situation, liability attribution, and claim amount. This business often faces challenges such as multi-picture information redundancy and cognitive burden. Applying large language models to some business areas of property insurance claim settlement can solve tasks such as claim material information induction, automatic generation and optimization of claim reports, multi-modal data analysis and understanding, intelligent review of materials submitted by users, knowledge Q&A, and auxiliary decision-making, reducing the manual workload while improving the efficiency of claims adjusters.
[0090] However, when a multi-picture understanding model based on deep learning (such as a large model) processes multi-picture scenario tasks, due to the lack of task-oriented attention guidance, it is difficult to focus on key information according to specific tasks, such as "identifying damaged items" or "extracting the accident time point", etc., and it is easy to miss important details in the images or text reports, thus affecting the accuracy of claim judgment. The same problem also exists in other application scenarios, that is, the large model cannot give a more detailed and more specific analysis and understanding.
[0091] Inspired by the human cognitive mechanism, this application introduces a review mechanism to simulate the cognitive process of claims adjusters "re-examining images and reports with a task". When initially viewing the image set, the large model will automatically form a global understanding of the overall information, but may overlook details. Subsequently, according to the specific task requirements in the user task instruction (such as "locate damaged items" or "extract liability description"), the review retrieval is initiated to centrally analyze the key image or report content. Specifically, the large model uses the fine-grained features of multi-image content as an "external knowledge base", iteratively retrieves key information related to the user task during the process of understanding the images, achieves precise focus, ignores redundant content, and improves processing efficiency.
[0092] The large model can also be replaced by other language models or non-language models with similar functions. The non-language models include image processing models, etc.
[0093] The present application will be described below separately from the model training stage and the model inference stage. The model inference stage is the application stage after the model training is completed. First, in combination with Figure 2 the embodiments of the model training stage will be described.
[0094] Figure 2 FIG. 12 is a schematic flowchart of a model training method provided for the embodiment. This method is executed by a computing device and includes the following steps.
[0095] Step S260, obtain a set of image samples S1 and the standard understanding content T1 for this set of image samples S1 under the user task instruction U1.
[0096] This set of image samples S1 contains one or more images. A set of image samples belongs to one sample, and when processing, the set of image samples S1 is processed as one sample. The standard understanding content can be text data. This set of image samples S1 and the corresponding standard understanding content T1 come from the training set. The training set contains multiple sets of image samples and the corresponding standard understanding content, and specifically may include a first type of training set and a second type of training set, and can both be in the image-text format or the video-text format.
[0097] The user task in the first type of training set is to perform a text description (caption) on the image, that is, a description-type user task. In this type of training set, the user task corresponding to each set of image samples is the same, which is "perform a text description on the image". This type of training set can use existing datasets such as LAION-CCSB or Valley.
[0098] The user tasks in the second type of training set are conversations based on images. In this type of training set, the user tasks corresponding to each group of image samples are different. For example, the user task can be to locate damaged items in the image or extract liability descriptions, etc. This type of training set can use existing datasets such as LLaVA-v1.5, VideoInstruct, or VCGPlus12k.
[0099] In this application, the user task and the user task instruction are essentially the same concept. The user task instruction is the implementation form of the user task, and the user task instruction can be in the form of text data, etc. There is a corresponding relationship between a group of image samples S1 and the user task instruction U1.
[0100] The standard understanding content T1 is the label of this group of image samples S1, which contains a review mark added before the key information in the standard understanding content. The review mark can be denoted as <rewind>, abbreviated as <rw>。At least one look-back identifier is included in the standard understanding content T1 corresponding to a set of image samples S1. The look-back identifier can appear at the beginning or in the middle of the sentence text corresponding to the standard understanding content. Look-back identifier <rewind>It is also a token in the vocabulary of the large model.
[0101] For example, the standard understanding content corresponding to a certain video is: A car's <rw>Headlight <rw>Damaged.
[0102] When constructing the sample, the standard understanding content of the image sample can be pre-generated by the large model, and the review identifier in the standard understanding content <rw>It can also be added in advance through a large model. For example, an image sample and an instruction text can be constructed into prompt data, the prompt is input into the large model, and the large model outputs a content containing a look-back identifier <rw>Standard understanding content. The indication text can be used to indicate that the large model adds a look-back identifier in front of the text specified by the rules <rw>Text specifying rules, which may include entity words, verb phrases, etc. Image samples and corresponding look-back identifiers can also be provided in the prompt <rw>The text is used as an example.
[0103] Step S210: Extract multiple fine-grained features of the set of image samples S1.
[0104] The set of image samples S1 contains multiple fine-grained features. Each fine-grained feature can be represented by an embedding vector. Each fine-grained feature is a local detailed feature in a certain part of the set of image samples S1, rather than the overall image feature or global feature of a certain image in the set of image samples S1. Fine-grained features are used to capture subtle visual differences in images.
[0105] When extracting multiple fine-grained features of the set of image samples S1, there are various implementation methods. For example, for each image in the set of image samples S1, the image can be segmented into several sub-blocks, and visual encoding is performed on the several sub-blocks respectively to obtain visual features corresponding to the several sub-blocks respectively. The visual features of the several sub-blocks are used as the multiple fine-grained features of the image. The fine-grained features of each image in the set of image samples S1 constitute the multiple fine-grained features of the set of image samples S1, and the multiple fine-grained features of the set of image samples S1 can be added to the visual memory bank. When needed, the fine-grained features are read from the visual memory bank.
[0106] Among them, the visual memory bank can be stored in a database or stored in the memory of a computing device in the form of a tensor. A sub-block is a local area of an image and is also a part of the image. When performing visual encoding on several sub-blocks respectively, it can be implemented through a visual encoder. The visual encoder can adopt OpenAI-CLIP ViT-L / 14, etc.
[0107] Figure 3 It is a schematic flow diagram for extracting image features provided for the embodiment. Among them, in the first row, the image is segmented into 4 sub-blocks, visual encoding is performed on each sub-block to obtain 4 visual features corresponding to each sub-block (the visual features are represented by small squares), one image corresponds to 16 visual features, and these 16 visual features are the fine-grained features of the image. Among them, the number of visual features obtained after encoding each sub-block is related to the parameters of the visual encoder and can be changed. The 16 visual features are stored in the visual memory bank for later use.
[0108] The multiple fine-grained features obtained in the above manner are visual features, not text data, and are data that cannot be directly processed by the large model.
[0109] In one implementation, multiple visual features can also be converted to the feature space of the large model through a vision-language adapter, and the multiple fine-grained features obtained are not visual features but feature vectors that the large model can directly process, that is, feature vectors in the feature space corresponding to the large model.
[0110] Step S220: Extract global features from the set of image samples S1 to obtain an initial global understanding.
[0111] Among them, the initial global understanding can be features in the feature space of the large model, that is, features that can be directly input into the large model and directly processed by the large model. The initial global understanding can include several embedding vectors, that is, several tokens. Several in this application means one or more. Global feature extraction refers to extracting features from the image as a whole. The extracted features are the overall image features and global features of the image, which are also the coarse-grained features of the image, rather than the local detail features or fine-grained features of the image. The initial global understanding includes the global features of all images in the set of image samples S1.
[0112] There are obvious differences between the fine-grained features in step S210 and the initial global understanding in step S220. For an image, there are obvious differences in the amount of feature information between segmenting it into multiple small blocks and encoding each small block separately using a visual encoder, and directly encoding the entire image (without segmentation). After the image is segmented into multiple small blocks, each small block is independently encoded, which can capture local features more meticulously. This processing method can better retain the local details of the image, and its information is more abundant. When directly encoding the entire image, the model extracts features from a global perspective, which may lose some local details, but can more efficiently capture the overall structure and semantic information of the image.
[0113] This step can be implemented in various ways. For example, the following steps 1 and 2 can be used to determine the initial global understanding.
[0114] Step 1: Perform overall visual encoding on the set of image samples S1 respectively to obtain the visual features of each image, and splice the visual features of the set of image samples S1 to obtain the global multi-image representation of the set of image samples S1.
[0115] Before performing visual encoding, each image in the set of image samples S1 can be downsampled, and the downsampled image can be visually encoded, which can reduce the amount of data.
[0116] Among them, when performing overall visual encoding on each image, the visual features of the image can be extracted through a visual encoder. The visual features are the overall features of the image. The number of visual features of each image can be one or more, and its number is related to the parameters of the visual encoder. Splice the visual features of all images in the set of image samples S1 into one feature, and this feature can be used as the global multi-image representation.
[0117] Step 2: Map the global multi-image representation to the feature space of the large model through a vision-language adapter to obtain an initial global understanding. Here, the vision-language adapter can be a model that needs to be trained.
[0118] Specifically, the global multi-image representation can be input into the vision-language adapter, and the vision-language adapter outputs the initial global understanding obtained after mapping the global multi-image representation. The initial global understanding is a feature in the feature space of the large model.
[0119] In Figure 3 In the example shown, in the second row, the original image is downsampled, that is, the width and height values are reduced while maintaining the aspect ratio, and then overall visual encoding is performed to obtain 4 embedding vectors. Suppose this set of image samples S1 contains 20 images, and each image corresponds to 4 embedding vectors representing visual features, for a total of 20 * 4, that is, 80 embedding vectors. These 80 embedding vectors are concatenated into a matrix, and the obtained matrix can be used as the global multi-image representation of this set of image samples S1. After concatenating the visual features of multiple images, a global multi-image representation can be obtained, and this global multi-image representation is used to be input into the large model after processing.
[0120] In one implementation, when the number of visual features is relatively large, spatial pooling sampling can also be performed on the visual features, and the sampled visual features are concatenated to obtain the global feature representation of this set of image samples S1.
[0121] The execution order of the above steps S210 and S220 is not sequential.
[0122] Step S230: Input the initial global understanding and the user task instruction U1 into the large model, and determine the content to be reviewed through the large model. The large model is used to generate text based on the input data and output it.
[0123] When determining the content to be reviewed in this step, the initial global understanding and the user task instruction U1 are input into the large model for determination, rather than directly inputting multiple fine-grained features and the user task instruction U1 into the large model. This can first enable the large model to understand the multi-images from a global perspective and then continue to understand the multi-images from a detailed perspective. Moreover, the number of tokens of the multiple fine-grained features of this set of image samples S1 is quite large, much larger than the number of tokens in the initial global understanding. Not directly inputting multiple fine-grained features into the large model can take into account the capacity bottleneck of the large model and avoid the problem that the large model cannot handle a large number of tokens.
[0124] The content to be reviewed is the key content that requires careful inspection of details when processing the set of image samples under the user task instruction, belonging to the review intention. The content to be reviewed determined from the large model can be in the form of an embedded vector in the feature space. The content to be reviewed can contain semantic information or can have no specific semantic information.
[0125] For example, under the user task of locating the damaged part of a vehicle, the embodiment hopes that the large model can learn to determine the damaged part of the vehicle as the content to be reviewed from the set of image samples S1. For example, the damaged part is a vehicle lamp, etc.
[0126] In one implementation, when determining the content to be reviewed through the large model, steps 3 and 4 can be adopted to implement.
[0127] Step 3, through the understanding of the initial global understanding and the user task instruction U1, the large model generates a review identifier during the process of outputting the prediction content <rw>。
[0128] Step 4, based on the review flag <rw>, obtain the content to be reviewed from the large model.
[0129] Among them, the review identifier <rw>Used to indicate retrieving content to be reviewed from a large model. Some large models generate multiple tokens step by step when generating predicted content, and review identifiers can be generated interspersed during the process of generating predicted content <rw>, thereby prompting the large model to output the content to be reviewed next. During the process of the large model outputting the predicted content, it is possible to detect whether there is a review identifier in the predicted content <rw>。When a look-back identifier appears in the predicted content <rw>When it is, step 4 is executed.
[0130] In specific implementation, the look-back identifier can be pre-set <rewind>Add it to the vocabulary of the large model. Initially, the embedding vector of the word "rewind" in the vocabulary can be used as the rewind identifier <rewind>The initialization token. This look-back identifier can be updated during subsequent training <rw>Embedding vector.
[0131] According to step S210, a look-back identifier has been added to the standard understanding content. <rw>, and the addition rule is preset. For example, the addition rule is to insert a recall identifier in front of the key information <rw>Therefore, in step 4, it is possible to be based on the look-back identifier <rw>and the preset addition rules, obtain the content for review from the large model. For example, the review identifier in the prediction content of the large model can be used <rw>Subsequently, a set number of embedding vectors output by the large model are used as the content to be reviewed. The large model includes multiple computational layers, including an input layer, intermediate hidden layers, and an output layer. In the output layer, the embedding vectors can be converted into corresponding texts based on the vocabulary and output. Specifically, when obtaining the content to be reviewed from the large model, a set number of embedding vectors can be obtained from the output layer.
[0132] In one implementation, to more accurately prompt the large model to generate the content to be reviewed, step 4 above can be implemented in the following manner:
[0133] When it is detected that the large model generates a review flag during the process of outputting the predicted content <rw>At this time, input Q preset learnable embedding vectors into the large model, and obtain Q learned embedding vectors determined by the large model based on the Q learnable embedding vectors. Determine the Q learned embedding vectors as the content to be reviewed.
[0134] Among them, the learnable embedding vectors are used to instruct the large model to extract the key information that needs to be reviewed, so as to determine the content to be reviewed based on the key information that needs to be reviewed, and output Q tokens as the content to be reviewed.
[0135] In specific implementation, the review flag can be <rw>Input into the large model together with Q learnable embedding vectors. After learning, the large model can determine Q + 1 tokens, and Q of these tokens are the learned embedding vectors.
[0136] Among them, the first quantity Q is a hyperparameter that can be determined in advance. For example, it can be, but is not limited to, values such as 5, 6, 7, 8, 9, etc.
[0137] In one implementation, the large model is an autoregressive large model. An autoregressive large model is a prediction model based on time series data. Its core idea is to use historical data to predict future data points. It assumes that the output at the current moment depends on a certain functional relationship of the output (or input) at the previous moment. Autoregressive large models such as the GPT series models generate coherent text by predicting the next word one by one. In the prediction process of the autoregressive large model, the following rule will be followed, that is: the token output at the previous time step is re - input into the large model, and the output token at the current time step is obtained. That is, the output token at the current time step is determined based on the output token at the historical time step.
[0138] In the autoregressive large model, when inputting Q preset learnable embedding vectors into the large model, the Q learnable embedding vectors can be input into the large model through the method of Teacher Forcing. In this method, instead of inputting the token generated by the large model at the previous time step into the large model, the Q learnable embedding vectors are input into the large model, so that the large model encodes according to the context to obtain a specific look - back intention, that is, the content to be looked back.
[0139] After generating the content to be looked back, the content to be looked back is not decoded as the text output content to be given to the user, nor is the content to be looked back input into the large model at the next time step.
[0140] The purpose of inserting Q learnable embedding vectors into the large model also includes clearly instructing the large model to encode the look - back intention in these Q embedding vectors, so that it can be clearly determined that after inputting Q preset learnable embedding vectors into the large model, the Q embedding vectors output by the large model are the content to be looked back.
[0141] In the process of the autoregressive large model outputting the understanding content (i.e., the prediction content) based on the initial global understanding and the user task instruction, when the autoregressive large model needs to look back at the fine - grained features of the image, the content to be looked back can be determined multiple times and dynamically, that is, look - back identifiers can be generated multiple times in the prediction content. <rw>whenever a replay identifier is detected <rw>When generating, Q preset learnable embedding vectors of the first quantity can be input into the autoregressive large model, and Q learned embedding vectors determined by the autoregressive large model based on the Q learnable embedding vectors can be obtained, and the Q learned embedding vectors are determined as the content to be reviewed.
[0142] Step S240, determining a first fine-grained feature related to the content to be reviewed from multiple fine-grained features.
[0143] Among them, multiple fine-grained features can be stored in the visual memory bank. The content to be reviewed is the embedding vector encoded by the large model, which contains rich context information and helps to retrieve the required fine-grained supplementary information from the visual memory bank. Multiple fine-grained features can be directly stored in the memory of the computing device. When retrieval is required, the computing device retrieves the first fine-grained feature related to the content to be reviewed from the multiple fine-grained features stored in the memory and extracts the first fine-grained feature from the memory.
[0144] This step can be executed by a rewinder module outside the large model. The rewinder module can be a program module in the computing device.
[0145] In one implementation, multiple fine-grained features and the first fine-grained feature are both feature vectors in the feature space of the large model and are vectors that the large model can directly process. In this scenario, when determining the first fine-grained feature related to the content to be reviewed from multiple fine-grained features, feature matching can be used, or the attention mechanism can be used to determine.
[0146] In one implementation, multiple fine-grained features are visual features, not feature vectors in the feature space of the large model and are vectors that the large model cannot directly process.
[0147] In this scenario, through the rewinder module, fine-grained features with high relevance can be retrieved from multiple fine-grained features, and the fine-grained feature can be transformed (that is, aligned) into the feature space of the large model to obtain the first fine-grained feature. Among them, the rewinder module contains learnable attention parameters. Initially, the attention parameters can be preset.
[0148] Specifically, the rewinder module can adopt the attention mechanism to determine the relevant visual features for the content to be reviewed from multiple fine-grained features based on the attention parameters it contains, and transform the relevant visual features into the feature space of the large model to obtain the first fine-grained feature.
[0149] Among them, the rewinder module adopts a minimalist structure, and its attention parameters include a first parameter W Q and a second parameter W K and the third parameter W V , that is, it only contains three linear projection layers. The rewinder module can determine the first fine-grained feature based on the attention parameter and the content to be rewound in various ways. For example, the first fine-grained feature can be obtained through the product of the attention parameter, the content to be rewound, and multiple fine-grained features. The first fine-grained feature can also be determined by the following steps 5 to 8, see Figure 4 as shown. Figure 4 It is a flowchart of an execution method of the rewinder module provided by the embodiment.
[0150] Step 5, based on the first parameter W Q and the product of the content to be rewound (such as token1), to obtain the query vector q1. Among them, the product is a linear multiplication. For example, multiply the first parameter W Q by the content to be rewound token1 to obtain the query vector q1. It can also be that the query vector q1 is obtained after performing a preset process on the product.
[0151] Here, an example is given where the content to be rewound contains one token. When the content to be rewound contains Q tokens or embedding vectors, the Q tokens can be constructed into a matrix, and based on the first parameter W Q and the product of this matrix, the corresponding query matrix is obtained.
[0152] Step 6, based on the second parameter W K and the third parameter W V respectively multiply with multiple fine-grained features to obtain the key-value pairs corresponding to each fine-grained feature. Among them, the product is a linear multiplication.
[0153] Multiple fine-grained features can be obtained from the visual memory bank. Taking a fine-grained feature F1 as an example, multiply the fine-grained feature F1 by the second parameter W K to obtain the key K, and multiply the fine-grained feature F1 by the third parameter W V to obtain the value V. The key and the value form the key-value pair corresponding to the fine-grained feature F1. In this way, the key-value pairs corresponding to all fine-grained features can be obtained.
[0154] Step 7, based on the product of the query vector q1 and the key K in multiple key-value pairs respectively, determine the corresponding attention coefficients.
[0155] For example, the query vector q1 can be multiplied by the key K in the key-value pair corresponding to the fine-grained feature F1 to obtain the attention coefficient corresponding to the fine-grained feature F1 or the key-value pair. It can also be that the result after performing a preset process on the product is used as the attention coefficient. The query vector q1 and the matrix formed by the key K in multiple key-value pairs are dot-producted to obtain the attention coefficient matrix.
[0156] Step 8: Based on the product of multiple attention coefficients and the values V in multiple key-value pairs, obtain the first fine-grained feature. Specifically, the attention coefficient matrix composed of multiple attention coefficients can be dot-multiplied with the matrix composed of the values V in multiple key-value pairs to obtain the first fine-grained feature. Alternatively, perform a preset process on the result of the dot product to obtain the first fine-grained feature.
[0157] It can be understood that when the content to be reviewed contains Q tokens (embedding vectors or feature vectors), after being processed through the above multiple steps, the obtained first fine-grained feature also contains Q tokens (embedding vectors or feature vectors).
[0158] To reduce the computational complexity and improve the stability and generalization ability of the model, the product of multiple attention coefficients and the values V in multiple key-value pairs can be normalized. The normalization process can be RMS Norm or other processing methods.
[0159] In this implementation, the rewinder module encodes the replay intention (i.e., the content to be reviewed) and the fine-grained features (visual features) in the visual memory bank through three linear projection layers respectively. This design ensures the modal alignment between the replay intention in the large model feature space and the features in the visual space. Subsequently, through dot product attention calculation, the rewinder module determines the correlation between the replay intention and the visual features and extracts the most important information.
[0160] Step S250: Input the first fine-grained feature into the large model, and through the large model, output the understanding content of this group of image samples S1 under the user task instruction. The first fine-grained feature contains Q tokens.
[0161] The large model can determine the understanding content of this group of image samples S1 based on the initial global understanding, the first fine-grained feature, and the user task instruction.
[0162] When the large model is an autoregressive large model, the teacher forcing method can be used to input Q tokens into the large model. In this method, instead of inputting the tokens generated by the large model in the previous time step into the large model, the first fine-grained feature is input into the large model, enabling the large model to generate the understanding content of this group of image samples S1 based on the input data and the previously learned information. Here, the tokens generated by the large model in the previous time step are Q learned embedding vectors. During execution, the first fine-grained feature is input into the large model, rather than the content to be reviewed generated in the previous time step.
[0163] The understanding content determined in this embodiment takes into account both fine-grained features and global features, and adopts a look-back review mechanism to obtain detailed feature information, so as to help the large model better and more pertinently understand the image and obtain more accurate understanding content.
[0164] Step S270: Determine the prediction loss based on the difference between this understanding content and the standard understanding content.
[0165] When determining the prediction loss, it can be determined based on the cross-entropy loss function. The understanding content generated by the large model is text data, and the standard understanding content is also text data. Specifically, the prediction loss can be determined based on the difference between the corresponding contents of this understanding content and the standard understanding content. The execution of this step can refer to the existing technology and will not be elaborated here.
[0166] Step S280: Fine-tune the large model based on the prediction loss at least.
[0167] Among them, the training method of this model can include two stages. In the first stage, the parameters of the look-back module rewinder and the vision-language adapter are updated based on the prediction loss, and the parameters of the large model are not updated. In the second stage, the parameters of the look-back module rewinder, the vision-language adapter, and the large model are fine-tuned based on the prediction loss.
[0168] In the training of the first stage, although the parameters of the large model are not updated, the parameters of the last layer of the large model can be modified to enable the large model to learn to add a look-back identifier to its output prediction content. <rw>。
[0169] In the training of the first stage and the second stage, the image samples in the first type of training set and the second type of training set can be used alternately. The above steps S210 - S270 are the forward process in model training, and step S280 is the backward process in model training and also the parameter update process. The specific parameter update and fine-tuning process can refer to the existing technology and will not be elaborated here.
[0170] The above steps S210 - S280 illustrate the process of training the large model, the rewinder, and the vision-language adapter using a set of image samples. In practical applications, steps S210 - S260 can be executed using multiple sets of image samples, and the total prediction loss can be determined based on the above differences corresponding to multiple sets of image samples, and the model can be updated once. The embodiments can train the model through multiple iterations. When the above prediction loss is less than the preset value, or the number of iterations exceeds the threshold, it can be considered that the model has reached the convergence state and the training process is completed.
[0171] The above is the description of the embodiments in the model training stage. After the large model, the rewinder, and the vision-language adapter are trained, they can be put into the application stage to perform model inference. The following combines Figure 5 to illustrate the embodiments in the model inference stage.
[0172] Figure 5 It is a schematic flowchart of a method for task understanding of images provided by the embodiment. This method can be executed by a computing device and specifically includes the following steps.
[0173] Step S510, extract multiple fine-grained features of a set of images M1 to be processed.
[0174] Among them, "to be processed" means that this set of images M1 needs to be processed under the user task instruction U2. This set of images M1 contains one or more images, and a set of images M1 is treated as a piece of data to be processed during processing. The user task instruction U2 can be in the form of text data, etc. There is a corresponding relationship between a set of images M1 and the user task instruction U2.
[0175] This set of images M1 contains multiple fine-grained features. Each fine-grained feature can be represented by an embedding vector. Each fine-grained feature is a detailed feature of a certain part in this set of images M1, rather than the overall image feature or global feature of a certain image in this set of images M1. Fine-grained features are used to capture subtle visual differences in the images.
[0176] There are multiple implementation methods when extracting multiple fine-grained features of the set of images M1. For example, for each image in the set of images M1, the image can be segmented into several sub-blocks, and visual encoding is performed on the several sub-blocks respectively to obtain visual features corresponding to the several sub-blocks respectively. The visual features of the several sub-blocks are used as the multiple fine-grained features of the image. The fine-grained features of each image in the set of images M1 constitute the multiple fine-grained features of the set of images M1, and the multiple fine-grained features of the set of images M1 can be added to the visual memory bank. When needed, the fine-grained features are read from the visual memory bank.
[0177] Among them, the visual memory bank can be implemented through a database. A sub-block is a local area of an image and is also a part of the image. When performing visual encoding on several sub-blocks respectively, it can be implemented through a visual encoder. The visual encoder can adopt OpenAI-CLIP ViT-L / 14, etc.
[0178] The process of extracting the fine-grained features of an image can refer to Figure 3 the first row in. Among them, the image is segmented into 4 sub-blocks, visual encoding is performed on each sub-block, and 4 visual features corresponding to each sub-block are obtained (the visual features are represented by small squares). One image corresponds to 16 visual features, and these 16 visual features are the fine-grained features of the image. Among them, the number of visual features obtained after each sub-block is encoded is related to the parameters of the visual encoder and can be changed. The 16 visual features are stored in the visual memory bank for later use.
[0179] The multiple fine-grained features obtained by the above method are visual features, not text data, and are data that the large model cannot directly process.
[0180] In one implementation method, multiple visual features can also be converted to the feature space of the large model through a visual language adapter. The multiple fine-grained features obtained are not visual features, but feature vectors that the large model can directly process, that is, feature vectors in the feature space corresponding to the large model.
[0181] Step S520, perform global feature extraction on the set of images M1 to obtain an initial global understanding.
[0182] Among them, the initial global understanding can be a feature in the feature space of the large model, that is, a feature that can be directly input into the large model and directly processed by the large model. The initial global understanding can include several embedding vectors, that is, several tokens. Several in this application means one or more. Global feature extraction refers to extracting features from the image as a whole, and the extracted features are the overall image features and global features of the image, and are also the coarse-grained features of the image, rather than the local detail features or fine-grained features of the image. The initial global understanding includes the global features of all images in the set of images M1.
[0183] For the fine-grained features in step S510 and the initial global understanding in S520, there are obvious differences between the two. For an image, segmenting it into multiple small blocks and encoding each small block separately using a visual encoder results in significantly different feature information compared to directly encoding the entire image (without segmentation). After the image is segmented into multiple small blocks, each small block is independently encoded, enabling more detailed capture of local features. This processing method can better retain the local details of the image, and its information content is more abundant. When directly encoding the entire image, the model extracts features from a global perspective, which may lose some local details but can more efficiently capture the overall structure and semantic information of the image.
[0184] Step S520 can be implemented in multiple ways. For example, the following steps 9 and 10 can be used to determine the initial global understanding.
[0185] Step 9: Perform overall visual encoding on each image in the set of images M1 to obtain the visual features of each image, and splice the visual features of the set of images M1 to obtain the global multi-image representation of the set of images M1.
[0186] Before performing visual encoding, each image in the set of images M1 can be downsampled, and the downsampled images can be visually encoded, which can reduce the amount of data.
[0187] Among them, when performing overall visual encoding on each image, the visual features of the image can be extracted through a visual encoder, and the visual features are the overall features of the image. The number of visual features of each image can be one or more, and its number is related to the parameters of the visual encoder. Splice the visual features of all images in the set of images M1 into one feature, and this feature can be used as the global multi-image representation.
[0188] Step 10: Map the global multi-image representation to the feature space of the large model through a visual language adapter to obtain the initial global understanding. Among them, the visual language adapter is a trained model.
[0189] Specifically, the global multi-image representation can be input into the visual language adapter, and the visual language adapter outputs the initial global understanding obtained after mapping the global multi-image representation. The initial global understanding is a feature in the feature space of the large model.
[0190] In Figure 3 In the example shown, in the second row, the original image is downsampled and then visually encoded globally to obtain 4 embedding vectors. Assume that the set of images M1 contains 20 images, and each image corresponds to 4 embedding vectors representing visual features, so there are a total of 20 * 4 = 80 embedding vectors. These 80 embedding vectors are concatenated into a matrix, and the resulting matrix can be used as the global multi-image representation of the set of images M1. After concatenating the visual features of multiple images, a global multi-image representation can be obtained, and this global multi-image representation is used to be input into a large model after being processed.
[0191] In one implementation, when the number of visual features is relatively large, spatial pooling sampling can also be performed on the visual features, and the sampled visual features are concatenated to obtain the global feature representation of the set of images M1.
[0192] The execution order of the above steps S510 and S520 is not sequential.
[0193] Step S530: Input the initial global understanding and the user task instruction U2 into the large model, and determine the content to be reviewed back through the large model. The large model is used to generate text based on the input data and output it.
[0194] In this step of determining the content to be reviewed back, the initial global understanding and the user task instruction U2 are input into the large model for determination, rather than directly inputting multiple fine-grained features and the user task instruction U2 into the large model. This can enable the large model to first understand the multi-images from a global perspective and then continue to understand the multi-images from a detailed perspective. Moreover, the number of tokens of the multiple fine-grained features of the set of images M1 is quite large, much larger than the number of tokens in the initial global understanding. Not directly inputting multiple fine-grained features into the large model can take into account the capacity bottleneck of the large model and avoid the problem that the large model cannot handle a large number of tokens.
[0195] The content to be reviewed back is the key content that needs to be carefully viewed for details when processing the set of images M1 under the user task instruction U2, and it belongs to the review intention. The content to be reviewed back determined from the large model can be in the form of embedding vectors in the feature space. The content to be reviewed back can contain semantic information or may not have specific semantic information.
[0196] For example, under the user task of locating the damaged part of a vehicle, the embodiment hopes that the large model can learn to determine the damaged part of the vehicle as the content to be reviewed back from the set of image samples S1, such as the damaged part being the headlight, etc.
[0197] In one implementation, when determining the content to be reviewed back through the large model, steps 11 and 12 can be used to achieve it.
[0198] Step 11, through the understanding of the initial global understanding and the user task instruction U2, the large model generates a review mark during the process of outputting the predicted content <rw>。
[0199] Step 12, based on the flashback identifier <rw>, obtain the content to be reviewed from the large model.
[0200] Among them, the review identifier <rw>Used to indicate obtaining content to be reviewed from a large model. Some large models generate multiple tokens step by step when generating prediction content, and review identifiers can be generated interspersed during the process of generating prediction content <rw>, thereby prompting the large model to output the content to be reviewed next. During the process of the large model outputting the predicted content, it is possible to detect whether there is a review identifier in the predicted content <rw>。When a recall identifier appears in the predicted content <rw>When, step 12 is executed.
[0201] In specific implementation, the replay identifier can be pre-set <rewind>Add it to the vocabulary of the large model. Initially, the embedding vector of the word "rewind" in the vocabulary can be used as the rewind identifier <rewind>Initialization token. The look-back identifier can be updated during subsequent training <rw>The embedded vector.
[0202] According to Figure 2 As can be seen from the embodiments, during the training phase, a review mark has been added to the standard understanding content <rw>, and the addition rule is preset. For example, the addition rule is to insert a recall identifier in front of the key information <rw>。After being trained, the large model has learned to add a review mark at the specified position <rw>Therefore, in step 12, it is possible to base on the look-back identifier <rw>and the preset addition rules, obtain the lookback content from the large model. For example, the lookback identifier in the prediction content of the large model can be used <rw>Subsequently, a set number of embedding vectors output by the large model are used as the content to be reviewed. The large model includes multiple computing layers, including an input layer, intermediate hidden layers, and an output layer. In the output layer, the embedding vectors can be converted into corresponding texts based on a vocabulary and output. Specifically, when obtaining the content to be reviewed from the large model, a set number of embedding vectors can be obtained from the output layer.
[0203] In one implementation, to more accurately prompt the large model to generate the content to be reviewed, step 12 above can be implemented in the following manner:
[0204] When it is detected that the large model generates a review identifier during the process of outputting prediction content <rw>When, input Q preset learnable embedding vectors of the first quantity into the large model, and obtain Q learned embedding vectors determined by the large model based on the Q learnable embedding vectors, and determine the Q learned embedding vectors as the content to be reviewed.
[0205] Among them, the learnable embedding vectors are used to instruct the large model to extract the key information that needs to be reviewed, so as to determine the content to be reviewed based on the key information that needs to be reviewed, and output Q tokens as the content to be reviewed.
[0206] In specific implementation, the review flag can be <rw>Input into the large model together with Q learnable embedding vectors. After learning, the large model can determine Q + 1 tokens, and Q of these tokens are the learned embedding vectors.
[0207] Among them, the first quantity Q is a hyperparameter that can be determined in advance. For example, it can be, but is not limited to, values such as 5, 6, 7, 8, 9, etc.
[0208] In one implementation, the large model is an autoregressive large model. An autoregressive large model is a prediction model based on time series data. Its core idea is to use historical data to predict future data points. It assumes that the output at the current time depends on a certain functional relationship of the output (or input) at the previous time. Autoregressive large models such as the GPT series of models generate coherent text by predicting the next word one by one. In the prediction process of the autoregressive large model, the following rule is followed, that is: the token output at the previous time step is re - input into the large model, and the output token at the current time step is obtained. That is, the output token at the current time step is determined based on the output tokens at historical time steps.
[0209] In the autoregressive large model, when inputting Q preset learnable embedding vectors into the large model, the Q learnable embedding vectors can be input into the large model by the teacher forcing method. In this method, the token generated by the large model in the previous time step is not input into the large model, but the Q learnable embedding vectors are input into the large model, so that the large model encodes a specific look - back intention according to the context, that is, the content to be looked back.
[0210] After generating the content to be looked back, the content to be looked back is not used as the output content of the large model, that is, the content to be looked back is not input into the large model in the next time step.
[0211] The purpose of inserting Q learnable embedding vectors into the large model also includes clearly instructing the large model to encode the look - back intention in these Q embedding vectors, so that it can be clearly determined that after inputting Q preset learnable embedding vectors into the large model, the Q embedding vectors output by the large model are the content to be looked back.
[0212] For example, in Figure 5 S530 of <images>And the user task instruction U2 is input into the autoregressive large model in the form of multimodal data. Initial global understanding <images>Contains a number of tokens, and the user task instruction U2 contains a number of tokens. Small boxes are used to represent tokens in the figure. The large model outputs the token "The", then inputs the token "The" into the large model, and then outputs "video", "shows", and " <rw>". At this time, it is detected that the large model outputs a look-back identifier <rw>Then, <rw>Input into the large model together with Q learnable embedding vectors, and the large model outputs Q + 1 learned tokens.
[0213] Step S540, determine a first fine-grained feature related to the content to be reviewed from multiple fine-grained features.
[0214] Among them, multiple fine-grained features can be stored in the visual memory bank. The content to be reviewed is the embedding vector encoded by the large model, which contains rich context information and helps to retrieve the required fine-grained supplementary information from the visual memory bank.
[0215] This step can be executed by a rewinder module outside the large model. The rewinder module can be a program module in the computing device.
[0216] In one implementation, multiple fine-grained features and the first fine-grained feature are both feature vectors in the feature space of the large model and are vectors that the large model can directly process. In this scenario, when determining the first fine-grained feature related to the content to be reviewed from multiple fine-grained features, feature matching or the attention mechanism can be used.
[0217] In one implementation, multiple fine-grained features are visual features and are not feature vectors in the feature space of the large model, and are vectors that the large model cannot directly process.
[0218] In this scenario, the rewinder module can be used to retrieve highly relevant fine-grained features from multiple fine-grained features and convert the fine-grained features to the feature space of the large model to obtain the first fine-grained feature. Among them, the rewinder module contains attention parameters to be learned. Initially, the attention parameters can be preset.
[0219] Specifically, the rewinder module can use the attention mechanism to determine the relevant visual features for the content to be reviewed from multiple fine-grained features based on the attention parameters it contains, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0220] Among them, the rewinder module adopts a minimalist structure, and its attention parameters include a first parameter W Q a second parameter W K and a third parameter W V , that is, it only includes three linear projection layers. The rewinder module can determine the first fine-grained feature based on the attention parameters and the content to be rewound in various ways. For example, the first fine-grained feature can be obtained by the product of the attention parameters, the content to be rewound, and multiple fine-grained features. The first fine-grained feature can also be determined by the following steps 13 to 16, see Figure 4 as shown.
[0221] Step 13, based on the product of the first parameter W Q and the content to be rewound (such as token1), obtain the query vector q1. Among them, this product is a linear multiplication. For example, multiply the first parameter W Q by the content to be rewound token1 to obtain the query vector q1. It can also be that the query vector q1 is obtained after performing a preset process on the product.
[0222] Here, an example is given with the content to be rewound containing one token. When the content to be rewound contains Q tokens or embedding vectors, the Q tokens can be constructed into a matrix, and based on the product of the first parameter W Q and this matrix, the corresponding query matrix is obtained.
[0223] Step 14, based on the product of the second parameter W K and the third parameter W V with multiple fine-grained features respectively, obtain the key-value pairs corresponding to each fine-grained feature. Among them, this product is a linear multiplication.
[0224] Multiple fine-grained features can be obtained from the visual memory bank. Taking a fine-grained feature F1 as an example, multiply the fine-grained feature F1 by the second parameter W K to obtain the key K, and multiply the fine-grained feature F1 by the third parameter W V to obtain the value V. The key and the value form the key-value pair corresponding to the fine-grained feature F1. In this way, the key-value pairs corresponding to all fine-grained features can be obtained.
[0225] Step 15, based on the product of the query vector q1 and the key K in multiple key-value pairs respectively, determine the corresponding attention coefficients.
[0226] For example, the query vector q1 can be multiplied by the key K in the key-value pair corresponding to the fine-grained feature F1 to obtain the attention coefficient corresponding to this fine-grained feature F1, that is, the key-value pair. It can also be that the result after performing a preset process on the product is used as the attention coefficient. The query vector q1 and the matrix formed by the key K in multiple key-value pairs are dot-producted to obtain the attention coefficient matrix.
[0227] Step 16: Obtain the first fine-grained feature based on the product of multiple attention coefficients and the values V in multiple key-value pairs. Specifically, the attention coefficient matrix composed of multiple attention coefficients can be dot-producted with the matrix composed of the values V in multiple key-value pairs to obtain the first fine-grained feature. Alternatively, perform a preset process on the result of the dot product to obtain the first fine-grained feature.
[0228] It can be understood that when the content to be reviewed contains Q tokens (embedding vectors or feature vectors), after the processing of the above multiple steps, the obtained first fine-grained feature also contains Q tokens (embedding vectors or feature vectors).
[0229] To reduce the computational complexity and improve the stability and generalization ability of the model, the product of multiple attention coefficients and the values V in multiple key-value pairs can be normalized. The normalization process can be RMS Norm or other processing methods.
[0230] In this implementation, the rewinder module encodes the playback intention (i.e., the content to be reviewed) and the fine-grained features (visual features) in the visual memory bank through three linear projection layers. This design ensures the modal alignment between the playback intention in the large model feature space and the features in the visual space. Subsequently, through dot-product attention calculation, the rewinder module determines the correlation between the playback intention and the visual features and extracts the most important information.
[0231] See Figure 5 , in step S540, input the Q tokens determined by the large model (i.e., Figure 5 the 2 small squares under the rewinder module in into the rewinder module. The rewinder module obtains all the fine-grained features corresponding to the set of images M1 from the visual memory bank and calculates the first fine-grained feature through steps 13-16.
[0232] Step S550: Input the first fine-grained feature into the large model, and output the understanding content of the set of images M1 under the user task instruction U2 through the large model. The first fine-grained feature contains Q tokens.
[0233] The large model can determine the understanding content of the set of images M1 based on the initial global understanding, the first fine-grained feature, and the user task instruction U2.
[0234] When the large model is an autoregressive large model, the teacher forcing method can be used to input Q tokens into the large model. In this method, instead of inputting the tokens generated by the large model at the previous time step into the large model, the first fine-grained feature is input into the large model, enabling the large model to generate the understanding content for the set of images M1 based on the input data and the information learned previously. Here, the Q learned embedding vectors are generated by the large model at the previous time step. During execution, the first fine-grained feature is input into the large model, rather than the content to be reviewed generated at the previous time step.
[0235] In Figure 5 the two small boxes enclosed by the dashed line are the first fine-grained feature. The first fine-grained feature is input into the large model through the teacher forcing method. Based on the first fine-grained feature and the tokens determined at the historical time steps, the large model continues to sequentially determine the following content: a person placing his clothes into a washing machine. It can be seen that the understanding content output by the large model after inputting the first fine-grained feature references the first fine-grained feature, that is, it combines the local detail features of the set of images M1. It should be noted that the prediction content "The video shows” output by the large model before determining the content to be reviewed and the understanding content output after inputting the first fine-grained feature together constitute the understanding content for the set of images M1.
[0236] The understanding content determined in this embodiment takes into account both fine-grained features and global features, and uses the review and scrutiny mechanism to obtain detailed feature information, thereby helping the large model to better and more specifically understand the images and obtain more accurate understanding content.
[0237] This embodiment simulates the cognitive process of "reviewing and scrutinizing with a task" when humans view multi-image content, and applies the task-oriented dynamic retrieval based on the attention mechanism to the understanding task of multi-images. The embodiment designs a lightweight review module rewinder to retrieve key information related to the task from the fine-grained features of multi-images and integrate this key information into the generation process, improving the context understanding ability and avoiding missing task-related content. Combining with the generative AI technology, the embodiment supports the automatic generation of image understanding reports, task result summaries, and multi-round interaction support. Different from the traditional RAG offline retrieval, this embodiment can perform online iterative retrieval during the generation process, dynamically activate the review operation, extract key information driven by task requirements, reduce the resource overhead of the task while ensuring performance improvement, and is suitable for deployment in actual business scenarios.
[0238] In this specification, words such as "first" in the first fine-grained feature, first parameter, and first quantity, and the corresponding "second" (if any) in the text are only for the convenience of distinction and description, and do not have any limiting meaning.
[0239] In this specification, a computing device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities.
[0240] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments, and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0241] Figure 6 A schematic block diagram of a task understanding device for images provided for an embodiment. This device embodiment corresponds to Figure 5 the method embodiment shown. The device 600 is deployed in a computing device and includes: a fine-grained module 610 configured to extract multiple fine-grained features of a set of images to be processed; a global module 620 configured to perform global feature extraction on the set of images to obtain an initial global understanding; a determination module 630 configured to input the initial global understanding and a user task instruction into a large model, and determine the content to be reviewed back through the large model; a review module 640 configured to determine a first fine-grained feature related to the content to be reviewed back from the multiple fine-grained features; an understanding module 650 configured to input the first fine-grained feature into the large model and output the understanding content of the several images under the user task instruction through the large model.
[0242] In one embodiment, the fine-grained module 610 is specifically configured to: for each image in the set of images, divide the image into several sub-blocks, perform visual encoding on the several sub-blocks respectively to obtain corresponding visual features, and use the visual features of the several sub-blocks as the multiple fine-grained features of the image;
[0243] The review module 640 is specifically configured to: based on the trained attention parameters included in the review module 640, use the attention mechanism to determine relevant visual features for the content to be reviewed back from the multiple fine-grained features, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0244] In one implementation, the attention parameters include a first parameter, a second parameter, and a third parameter. The look-back module 640 includes: a first sub-module to a fourth sub-module (not shown in the figure). Among them, the first sub-module is configured to obtain a query vector based on the product of the first parameter and the content to be looked back. The second sub-module is configured to obtain a key-value pair corresponding to each fine-grained feature based on the products of the second parameter and the third parameter respectively and the multiple fine-grained features. The third sub-module is configured to determine corresponding attention coefficients based on the products of the query vector and the keys in the multiple key-value pairs respectively. The fourth sub-module is configured to obtain the first fine-grained feature based on the products of the multiple attention coefficients and the values in the multiple key-value pairs.
[0245] In one implementation, the initial global understanding is a feature in the feature space of the large model. The global module 620 includes an encoding sub-module 621 and a mapping sub-module 622. Among them, the encoding sub-module 621 is configured to perform overall visual encoding on the set of images respectively to obtain visual features of each image, and splice the visual features of the set of images to obtain a global multi-image representation of the set of images. The mapping sub-module 622 is configured to map the global multi-image representation to the feature space through a pre-trained vision-language adapter to obtain the initial global understanding.
[0246] In one implementation, the determination module 630 includes an identification sub-module 631 and an acquisition sub-module 632. The identification sub-module 631 is configured to generate a look-back identifier during the process of the large model outputting prediction content through understanding the initial global understanding and the user task instruction, and the look-back identifier is used to indicate obtaining the content to be looked back from the large model. The acquisition sub-module 632 is configured to obtain the content to be looked back from the large model based on the look-back identifier.
[0247] In one implementation, the acquisition sub-module 632 is specifically configured to:
[0248] When it is detected that the large model generates the look-back identifier during the process of outputting prediction content, input a first number of preset learnable embedding vectors into the large model, and obtain a first number of learned embedding vectors determined by the large model based on the first number of learnable embedding vectors, and determine the first number of learned embedding vectors as the content to be looked back; the learnable embedding vectors are used to indicate the large model to extract key information that needs to be looked back.
[0249] In one implementation, the device 600 further includes: an adding module (not shown in the figure), configured to pre-add the look-back identifier to the vocabulary of the large model.
[0250] In one embodiment, the large model is an autoregressive large model; the content to be reviewed is not used as the output content of the large model.
[0251] Figure 7 FIG. is a schematic block diagram of a model training device provided for an embodiment. The device 700 corresponds to Figure 2 the method embodiment shown. The device 700 is deployed in a computing device and includes: an acquisition module 710 configured to acquire a set of image samples and standard understanding content for the set of image samples under a user task instruction; an extraction module 720 configured to extract multiple fine-grained features of the set of image samples; an initial module 730 configured to perform global feature extraction on the set of image samples to obtain an initial global understanding; an input module 740 configured to input the initial global understanding and the user task instruction into a large model to determine the content to be reviewed through the large model; a review module 750 configured to determine a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features; a prediction module 760 configured to input the first fine-grained feature into the large model and output, through the large model, the understanding content for the set of image samples under the user task instruction; a loss module 770 configured to determine a prediction loss based on the difference between the understanding content and the standard understanding content; and a training module 780 configured to fine-tune at least the large model based on the prediction loss.
[0252] In one embodiment, the standard understanding content includes a review identifier, and the review identifier is added before the key information in the standard understanding content.
[0253] In one embodiment, the multiple fine-grained features are visual features. The review module 750 is specifically configured to: based on the attention parameters included in the review module, use an attention mechanism to determine relevant visual features for the content to be reviewed from the multiple fine-grained features, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0254] In one embodiment, the initial global understanding is a feature in the feature space of the large model. The initial module 730 includes: a visual sub-module and a conversion sub-module (not shown in the figure). The visual sub-module is configured to perform overall visual encoding on each of the set of image samples to obtain visual features of each image, and splice the visual features of the set of images to obtain a global multi-image representation of the set of images. The conversion sub-module is configured to map the global multi-image representation to the feature space through a vision-language adapter to obtain the initial global understanding.
[0255] In one embodiment, the model training method includes two phases. The training module 780 includes: a training sub-module 781 and a fine-tuning sub-module 782. The training sub-module 781 is configured to update the parameters of the look-back module and the vision-language adapter based on the prediction loss in the first phase without updating the parameters of the large model. The fine-tuning sub-module 782 is configured to fine-tune the parameters of the look-back module, the vision-language adapter, and the large model based on the prediction loss in the second phase.
[0256] The above device embodiments correspond to the method embodiments. For specific descriptions, reference may be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference may be made to the corresponding method embodiments.
[0257] This specification embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute Figures 1 to 5 any of the methods described above.
[0258] This specification embodiment also provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, it implements Figures 1 to 5 any of the methods described above.
[0259] The various embodiments in this specification are all described in a progressive manner. The same or similar parts among the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the storage medium and computing device embodiments, since they are basically similar to the method embodiments, the descriptions are relatively simple. For the relevant parts, reference may be made to the partial descriptions of the method embodiments.
[0260] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0261] The above specific implementation manners further elaborate on the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above is only the specific implementation manners of the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.< / rw> can be <rw>Input into the large model together with Q learnable embedding vectors, and the large model outputs Q + 1 learned tokens.
[0213] Step S540, determine a first fine-grained feature related to the content to be reviewed from multiple fine-grained features.
[0214] Among them, multiple fine-grained features can be stored in the visual memory bank. The content to be reviewed is the embedding vector encoded by the large model, which contains rich context information and helps to retrieve the required fine-grained supplementary information from the visual memory bank.
[0215] This step can be executed by a rewinder module outside the large model. The rewinder module can be a program module in the computing device.
[0216] In one implementation, multiple fine-grained features and the first fine-grained feature are both feature vectors in the feature space of the large model and are vectors that the large model can directly process. In this scenario, when determining the first fine-grained feature related to the content to be reviewed from multiple fine-grained features, feature matching or the attention mechanism can be used.
[0217] In one implementation, multiple fine-grained features are visual features and are not feature vectors in the feature space of the large model, and are vectors that the large model cannot directly process.
[0218] In this scenario, the rewinder module can be used to retrieve highly relevant fine-grained features from multiple fine-grained features and convert the fine-grained features to the feature space of the large model to obtain the first fine-grained feature. Among them, the rewinder module contains attention parameters to be learned. Initially, the attention parameters can be preset.
[0219] Specifically, the rewinder module can use the attention mechanism to determine the relevant visual features for the content to be reviewed from multiple fine-grained features based on the attention parameters it contains, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0220] Among them, the rewinder module adopts a minimalist structure, and its attention parameters include a first parameter W Q a second parameter W K and a third parameter W V , that is, it only includes three linear projection layers. The rewinder module can determine the first fine-grained feature based on the attention parameters and the content to be rewound in various ways. For example, the first fine-grained feature can be obtained by the product of the attention parameters, the content to be rewound, and multiple fine-grained features. The first fine-grained feature can also be determined by the following steps 13 to 16, see Figure 4 as shown.
[0221] Step 13, based on the product of the first parameter W Q and the content to be rewound (such as token1), obtain the query vector q1. Among them, this product is a linear multiplication. For example, multiply the first parameter W Q by the content to be rewound token1 to obtain the query vector q1. It can also be that the query vector q1 is obtained after performing a preset process on the product.
[0222] Here, an example is given with the content to be rewound containing one token. When the content to be rewound contains Q tokens or embedding vectors, the Q tokens can be constructed into a matrix, and based on the product of the first parameter W Q and this matrix, the corresponding query matrix is obtained.
[0223] Step 14, based on the product of the second parameter W K and the third parameter W V with multiple fine-grained features respectively, obtain the key-value pairs corresponding to each fine-grained feature. Among them, this product is a linear multiplication.
[0224] Multiple fine-grained features can be obtained from the visual memory bank. Taking a fine-grained feature F1 as an example, multiply the fine-grained feature F1 by the second parameter W K to obtain the key K, and multiply the fine-grained feature F1 by the third parameter W V to obtain the value V. The key and the value form the key-value pair corresponding to the fine-grained feature F1. In this way, the key-value pairs corresponding to all fine-grained features can be obtained.
[0225] Step 15, based on the product of the query vector q1 and the key K in multiple key-value pairs respectively, determine the corresponding attention coefficients.
[0226] For example, the query vector q1 can be multiplied by the key K in the key-value pair corresponding to the fine-grained feature F1 to obtain the attention coefficient corresponding to this fine-grained feature F1, that is, the key-value pair. It can also be that the result after performing a preset process on the product is used as the attention coefficient. The query vector q1 and the matrix formed by the key K in multiple key-value pairs are dot-producted to obtain the attention coefficient matrix.
[0227] Step 16: Obtain the first fine-grained feature based on the product of multiple attention coefficients and the values V in multiple key-value pairs. Specifically, the attention coefficient matrix composed of multiple attention coefficients can be dot-producted with the matrix composed of the values V in multiple key-value pairs to obtain the first fine-grained feature. Alternatively, perform a preset process on the result of the dot product to obtain the first fine-grained feature.
[0228] It can be understood that when the content to be reviewed contains Q tokens (embedding vectors or feature vectors), after the processing of the above multiple steps, the obtained first fine-grained feature also contains Q tokens (embedding vectors or feature vectors).
[0229] To reduce the computational complexity and improve the stability and generalization ability of the model, the product of multiple attention coefficients and the values V in multiple key-value pairs can be normalized. The normalization process can be RMS Norm or other processing methods.
[0230] In this implementation, the rewinder module encodes the playback intention (i.e., the content to be reviewed) and the fine-grained features (visual features) in the visual memory bank through three linear projection layers. This design ensures the modal alignment between the playback intention in the large model feature space and the features in the visual space. Subsequently, through dot-product attention calculation, the rewinder module determines the correlation between the playback intention and the visual features and extracts the most important information.
[0231] See Figure 5 , in step S540, input the Q tokens determined by the large model (i.e., Figure 5 the 2 small squares under the rewinder module in into the rewinder module. The rewinder module obtains all the fine-grained features corresponding to the set of images M1 from the visual memory bank and calculates the first fine-grained feature through steps 13-16.
[0232] Step S550: Input the first fine-grained feature into the large model, and output the understanding content of the set of images M1 under the user task instruction U2 through the large model. The first fine-grained feature contains Q tokens.
[0233] The large model can determine the understanding content of the set of images M1 based on the initial global understanding, the first fine-grained feature, and the user task instruction U2.
[0234] When the large model is an autoregressive large model, the teacher forcing method can be used to input Q tokens into the large model. In this method, instead of inputting the tokens generated by the large model at the previous time step into the large model, the first fine-grained feature is input into the large model, enabling the large model to generate the understanding content for the set of images M1 based on the input data and the information learned previously. Here, the Q learned embedding vectors are generated by the large model at the previous time step. During execution, the first fine-grained feature is input into the large model, rather than the content to be reviewed generated at the previous time step.
[0235] In Figure 5 the two small boxes enclosed by the dashed line are the first fine-grained feature. The first fine-grained feature is input into the large model through the teacher forcing method. Based on the first fine-grained feature and the tokens determined at the historical time steps, the large model continues to sequentially determine the following content: a person placing his clothes into a washing machine. It can be seen that the understanding content output by the large model after inputting the first fine-grained feature references the first fine-grained feature, that is, it combines the local detail features of the set of images M1. It should be noted that the prediction content "The video shows” output by the large model before determining the content to be reviewed and the understanding content output after inputting the first fine-grained feature together constitute the understanding content for the set of images M1.
[0236] The understanding content determined in this embodiment takes into account both fine-grained features and global features, and uses the review and scrutiny mechanism to obtain detailed feature information, thereby helping the large model to better and more specifically understand the images and obtain more accurate understanding content.
[0237] This embodiment simulates the cognitive process of "reviewing and scrutinizing with a task" when humans view multi-image content, and applies the task-oriented dynamic retrieval based on the attention mechanism to the understanding task of multi-images. The embodiment designs a lightweight review module rewinder to retrieve key information related to the task from the fine-grained features of multi-images and integrate this key information into the generation process, improving the context understanding ability and avoiding missing task-related content. Combining with the generative AI technology, the embodiment supports the automatic generation of image understanding reports, task result summaries, and multi-round interaction support. Different from the traditional RAG offline retrieval, this embodiment can perform online iterative retrieval during the generation process, dynamically activate the review operation, extract key information driven by task requirements, reduce the resource overhead of the task while ensuring performance improvement, and is suitable for deployment in actual business scenarios.
[0238] In this specification, words such as "first" in the first fine-grained feature, first parameter, and first quantity, and the corresponding "second" (if any) in the text are only for the convenience of distinction and description, and do not have any limiting meaning.
[0239] In this specification, a computing device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities.
[0240] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments, and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0241] Figure 6 A schematic block diagram of a task understanding device for images provided for an embodiment. This device embodiment corresponds to Figure 5 the method embodiment shown. The device 600 is deployed in a computing device and includes: a fine-grained module 610 configured to extract multiple fine-grained features of a set of images to be processed; a global module 620 configured to perform global feature extraction on the set of images to obtain an initial global understanding; a determination module 630 configured to input the initial global understanding and a user task instruction into a large model, and determine the content to be reviewed back through the large model; a review module 640 configured to determine a first fine-grained feature related to the content to be reviewed back from the multiple fine-grained features; an understanding module 650 configured to input the first fine-grained feature into the large model and output the understanding content of the several images under the user task instruction through the large model.
[0242] In one embodiment, the fine-grained module 610 is specifically configured to: for each image in the set of images, divide the image into several sub-blocks, perform visual encoding on the several sub-blocks respectively to obtain corresponding visual features, and use the visual features of the several sub-blocks as the multiple fine-grained features of the image;
[0243] The review module 640 is specifically configured to: based on the trained attention parameters included in the review module 640, use the attention mechanism to determine relevant visual features for the content to be reviewed back from the multiple fine-grained features, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0244] In one implementation, the attention parameters include a first parameter, a second parameter, and a third parameter. The look-back module 640 includes: a first sub-module to a fourth sub-module (not shown in the figure). Among them, the first sub-module is configured to obtain a query vector based on the product of the first parameter and the content to be looked back. The second sub-module is configured to obtain a key-value pair corresponding to each fine-grained feature based on the products of the second parameter and the third parameter respectively and the multiple fine-grained features. The third sub-module is configured to determine corresponding attention coefficients based on the products of the query vector and the keys in the multiple key-value pairs respectively. The fourth sub-module is configured to obtain the first fine-grained feature based on the products of the multiple attention coefficients and the values in the multiple key-value pairs.
[0245] In one implementation, the initial global understanding is a feature in the feature space of the large model. The global module 620 includes an encoding sub-module 621 and a mapping sub-module 622. Among them, the encoding sub-module 621 is configured to perform overall visual encoding on the set of images respectively to obtain visual features of each image, and splice the visual features of the set of images to obtain a global multi-image representation of the set of images. The mapping sub-module 622 is configured to map the global multi-image representation to the feature space through a pre-trained vision-language adapter to obtain the initial global understanding.
[0246] In one implementation, the determination module 630 includes an identification sub-module 631 and an acquisition sub-module 632. The identification sub-module 631 is configured to generate a look-back identifier during the process of the large model outputting prediction content through understanding the initial global understanding and the user task instruction, and the look-back identifier is used to indicate obtaining the content to be looked back from the large model. The acquisition sub-module 632 is configured to obtain the content to be looked back from the large model based on the look-back identifier.
[0247] In one implementation, the acquisition sub-module 632 is specifically configured to:
[0248] When it is detected that the large model generates the look-back identifier during the process of outputting prediction content, input a first number of preset learnable embedding vectors into the large model, and obtain a first number of learned embedding vectors determined by the large model based on the first number of learnable embedding vectors, and determine the first number of learned embedding vectors as the content to be looked back; the learnable embedding vectors are used to indicate the large model to extract key information that needs to be looked back.
[0249] In one implementation, the device 600 further includes: an adding module (not shown in the figure), configured to pre-add the look-back identifier to the vocabulary of the large model.
[0250] In one embodiment, the large model is an autoregressive large model; the content to be reviewed is not used as the output content of the large model.
[0251] Figure 7 FIG. is a schematic block diagram of a model training device provided for an embodiment. The device 700 corresponds to Figure 2 the method embodiment shown. The device 700 is deployed in a computing device and includes: an acquisition module 710 configured to acquire a set of image samples and standard understanding content for the set of image samples under a user task instruction; an extraction module 720 configured to extract multiple fine-grained features of the set of image samples; an initial module 730 configured to perform global feature extraction on the set of image samples to obtain an initial global understanding; an input module 740 configured to input the initial global understanding and the user task instruction into a large model to determine the content to be reviewed through the large model; a review module 750 configured to determine a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features; a prediction module 760 configured to input the first fine-grained feature into the large model and output, through the large model, the understanding content for the set of image samples under the user task instruction; a loss module 770 configured to determine a prediction loss based on the difference between the understanding content and the standard understanding content; and a training module 780 configured to fine-tune at least the large model based on the prediction loss.
[0252] In one embodiment, the standard understanding content includes a review identifier, and the review identifier is added before the key information in the standard understanding content.
[0253] In one embodiment, the multiple fine-grained features are visual features. The review module 750 is specifically configured to: based on the attention parameters included in the review module, use an attention mechanism to determine relevant visual features for the content to be reviewed from the multiple fine-grained features, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
[0254] In one embodiment, the initial global understanding is a feature in the feature space of the large model. The initial module 730 includes: a visual sub-module and a conversion sub-module (not shown in the figure). The visual sub-module is configured to perform overall visual encoding on each of the set of image samples to obtain visual features of each image, and splice the visual features of the set of images to obtain a global multi-image representation of the set of images. The conversion sub-module is configured to map the global multi-image representation to the feature space through a vision-language adapter to obtain the initial global understanding.
[0255] In one embodiment, the model training method includes two phases. The training module 780 includes: a training sub-module 781 and a fine-tuning sub-module 782. The training sub-module 781 is configured to update the parameters of the look-back module and the vision-language adapter based on the prediction loss in the first phase without updating the parameters of the large model. The fine-tuning sub-module 782 is configured to fine-tune the parameters of the look-back module, the vision-language adapter, and the large model based on the prediction loss in the second phase.
[0256] The above device embodiments correspond to the method embodiments. For specific descriptions, reference may be made to the descriptions in the method embodiment section, which will not be elaborated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, reference may be made to the corresponding method embodiments.
[0257] This specification embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute Figures 1 to 5 any of the methods described above.
[0258] This specification embodiment also provides a computing device, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, it implements Figures 1 to 5 any of the methods described above.
[0259] The various embodiments in this specification are all described in a progressive manner. The same or similar parts among the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the storage medium and computing device embodiments, since they are basically similar to the method embodiments, the descriptions are relatively simple. For the relevant parts, reference may be made to the partial descriptions of the method embodiments.
[0260] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0261] The above specific implementation manners further elaborate on the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above is only the specific implementation manners of the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.< / rw> < / rw> < / rw> < / images> < / images> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rewind> < / rewind> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rewind> < / rewind> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rw> < / rewind> < / rw> < / rewind>
Claims
1. A method for task understanding of images, the method comprising: Extracting multiple fine-grained features of a set of images to be processed; Performing global feature extraction on the set of images to obtain an initial global understanding; Inputting the initial global understanding and user task instructions into a large model, and determining the content to be reviewed through the large model; Determining a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features; Inputting the first fine-grained feature into the large model, and outputting the understanding content of the several images under the user task instructions through the large model.
2. The method according to claim 1, wherein the step of extracting multiple fine-grained features of a set of images to be processed comprises: For each image in the set of images, segmenting the image into several sub-blocks, respectively performing visual encoding on the several sub-blocks to obtain corresponding visual features, and taking the visual features of the several sub-blocks as the multiple fine-grained features of the image; The step of determining a first fine-grained feature related to the content to be reviewed from the multiple fine-grained features comprises: Through a pre-trained review module, based on the attention parameters included in the review module, using an attention mechanism to determine relevant visual features for the content to be reviewed from the multiple fine-grained features, and converting the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
3. The method according to claim 2, wherein the attention parameters include a first parameter, a second parameter, and a third parameter; the step of obtaining the first fine-grained feature comprises: Based on the product of the first parameter and the content to be reviewed, obtaining a query vector; Based on the products of the second parameter and the third parameter and the multiple fine-grained features respectively, obtaining key-value pairs corresponding to each fine-grained feature; Based on the products of the query vector and the keys in the multiple key-value pairs respectively, determining corresponding attention coefficients; Based on the products of the multiple attention coefficients and the values in the multiple key-value pairs, obtaining the first fine-grained feature.
4. The method according to claim 1, wherein the initial global understanding is a feature in the feature space of the large model; the step of performing global feature extraction on the set of images comprises: Performing overall visual encoding on the set of images respectively to obtain visual features of each image, and splicing the visual features of the set of images to obtain a global multi-image representation of the set of images; Mapping the global multi-image representation to the feature space through a pre-trained vision-language adapter to obtain the initial global understanding.
5. The method according to claim 1, wherein the step of determining the content to be reviewed through the large model comprises: Through the understanding of the initial global understanding and user task instructions, the large model generates a review identifier during the process of outputting prediction content, and the review identifier is used to indicate obtaining the content to be reviewed from the large model; Based on the review identifier, obtaining the content to be reviewed from the large model.
6. The method according to claim 5, wherein the step of reading the content to be reviewed from the large model based on the review identifier comprises: When it is detected that the large model generates the review identifier during the process of outputting predicted content, input a first number of preset learnable embedding vectors into the large model, and obtain a first number of learned embedding vectors determined by the large model based on the first number of learnable embedding vectors, and determine the first number of learned embedding vectors as the content to be reviewed; The learnable embedding vectors are used to instruct the large model to extract key information that needs to be reviewed.
7. The method according to claim 5, further comprising: The review identifier is pre-added to the vocabulary of the large model.
8. The method according to claim 7, wherein the large model is an autoregressive large model; the content to be reviewed is not used as the output content of the large model.
9. A model training method, comprising: Obtain a set of image samples and standard understanding content for this set of image samples under a user task instruction; Extract a plurality of fine-grained features of the set of image samples; Perform global feature extraction on the set of image samples to obtain an initial global understanding; Input the initial global understanding and the user task instruction into a large model, and determine the content to be reviewed through the large model; Determine a first fine-grained feature related to the content to be reviewed from the plurality of fine-grained features; Input the first fine-grained feature into the large model, and output, through the large model, the understanding content for the set of image samples under the user task instruction; Determine a prediction loss based on the difference between the understanding content and the standard understanding content; Fine-tune at least the large model based on the prediction loss.
10. The method according to claim 9, wherein the standard understanding content contains a review identifier, and the review identifier is added before the key information in the standard understanding content.
11. The method according to claim 9, wherein the plurality of fine-grained features are visual features; The step of determining a first fine-grained feature related to the content to be reviewed from the plurality of fine-grained features comprises: Through a review module, based on the attention parameters included in the review module, use an attention mechanism to determine relevant visual features for the content to be reviewed from the plurality of fine-grained features, and convert the relevant visual features to the feature space of the large model to obtain the first fine-grained feature.
12. The method according to claim 11, wherein the initial global understanding is a feature in the feature space of the large model; the step of performing global feature extraction on the set of image samples comprises: Perform overall visual encoding on each image in the set of image samples to obtain visual features of each image, and splice the visual features of this set of images to obtain a global multi-image representation of this set of images; Map the global multi-image representation to the feature space through a vision-language adapter to obtain the initial global understanding.
13. The method according to claim 12, wherein the model training method comprises two stages; the step of fine-tuning at least the large model based on the prediction loss comprises: In the first stage, the parameters of the look-back module and the vision-language adapter are updated based on the prediction loss, and the parameters of the large model are not updated; In the second stage, the parameters of the look-back module, the vision-language adapter, and the large model are fine-tuned based on the prediction loss.
14. An apparatus for task understanding of images, comprising: A fine-grained module configured to extract multiple fine-grained features of a set of images to be processed; A global module configured to perform global feature extraction on the set of images to obtain an initial global understanding; A determination module configured to input the initial global understanding and a user task instruction into a large model, and determine the content to be looked back through the large model; A look-back module configured to determine a first fine-grained feature related to the content to be looked back from the multiple fine-grained features; An understanding module configured to input the first fine-grained feature into the large model, and output the understanding content of the several images under the user task instruction through the large model.
15. A model training apparatus, comprising: An acquisition module configured to acquire a set of image samples and standard understanding content of the set of image samples under a user task instruction; An extraction module configured to extract multiple fine-grained features of the set of image samples; An initial module configured to perform global feature extraction on the set of image samples to obtain an initial global understanding; An input module configured to input the initial global understanding and the user task instruction into a large model, and determine the content to be looked back through the large model; A look-back module configured to determine a first fine-grained feature related to the content to be looked back from the multiple fine-grained features; A prediction module configured to input the first fine-grained feature into the large model, and output the understanding content of the set of image samples under the user task instruction through the large model; A loss module configured to determine a prediction loss based on the difference between the understanding content and the standard understanding content; A training module configured to fine-tune at least the large model based on the prediction loss.
16. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed in a computer, the computer is made to execute the method according to any one of claims 1-13.
17. A computing device, comprising a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method according to any one of claims 1-13 is implemented.