Training method and retrieval method of generative multi-modal retrieval model and related device
By evolving and scoring sample instructions, an instruction scoring and result reward model is constructed. Combined with a proximal policy optimization algorithm, a multimodal large language model is reinforced to generate a generative multimodal retrieval model. This solves the problems of high computational complexity and retrieval accuracy dependence on instruction quality in existing technologies, and achieves efficient and accurate text-image cross-modal retrieval.
Patent Information
- Application Number
- CN202510835531.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-28
AI Technical Summary
In existing technologies, text-image cross-modal retrieval methods rely on global similarity calculation of high-dimensional vectors, resulting in high computational complexity. Furthermore, retrieval accuracy depends on the quality of the query command, and ambiguous commands are prone to causing retrieval bias, making it difficult to improve retrieval efficiency and accuracy.
By evolving and scoring sample instructions, an instruction scoring model and a result reward model are constructed. Combined with a proximal policy optimization algorithm, a multimodal large language model is trained through reinforcement learning to generate a generative multimodal retrieval model and optimize the text-image cross-modal retrieval process.
It improves the efficiency and accuracy of cross-modal text-image retrieval, reduces reliance on query command quality, and enhances retrieval stability and precision.
Smart Images

Figure CN120849633A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a training method, retrieval method, and related apparatus for a generative multimodal retrieval model. Background Technology
[0002] With the explosive growth of multimedia data in both quantity and form, text-image cross-modal retrieval technology has become an important research direction in the field of artificial intelligence. This technology accurately retrieves related image modal data through text modal queries.
[0003] Related technologies rely on cross-modal embedding learning or shared representation alignment to achieve retrieval by directly learning the similarity between modalities. However, the semantic representations between heterogeneous modalities such as text and images are significantly different, making it difficult to improve retrieval performance. Moreover, as the size of the retrieval database increases, the computational complexity of these methods increases dramatically due to their reliance on global similarity calculation of high-dimensional vectors (such as cosine similarity), resulting in a significant decrease in retrieval efficiency.
[0004] In addition, existing technologies also compensate for the gap between modalities by generating high-quality identifiers. For example, identifiers are generated by multi-large language models (MLLM) to help achieve semantic transfer between modalities. However, large model-based methods are highly sensitive to the quality of query instructions. Ambiguous instructions are prone to causing retrieval bias, resulting in retrieval accuracy being highly dependent on the quality of instruction input. Therefore, how to eliminate instruction dependence and further improve retrieval accuracy has become the main challenge of generative text-image cross-modal retrieval. Summary of the Invention
[0005] This invention provides a training method, retrieval method, and related apparatus for a generative multimodal retrieval model, which addresses the shortcomings of existing large-model-based methods that are highly sensitive to the quality of query commands, and where fuzzy commands can easily lead to retrieval biases, resulting in retrieval accuracy being highly dependent on the quality of command input. This invention improves retrieval efficiency and accuracy.
[0006] This invention provides a training method for a generative multimodal retrieval model, comprising: The sample instructions are evolved and scored to obtain an instruction score dataset; The instruction scoring model is trained by fine-tuning the large language model based on the instruction scoring dataset to obtain the instruction scoring model; the result reward model is trained by fine-tuning the large language model based on discrete image tokens, wherein the discrete image tokens are obtained based on the multimodal large language model MLLM. The MLLM is trained by a proximal policy optimization algorithm based on the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model.
[0007] According to the training method of the generative multimodal retrieval model provided by the present invention, the result reward model dataset is obtained through the following steps: Based on the MLLM, the candidate images are visually encoded to generate discrete image tokens, which are then stored in a trie structure to obtain tree structure data. The tree structure data is labeled to obtain the result reward model dataset.
[0008] According to the training method of the generative multimodal retrieval model provided by the present invention, the step of using a proximal policy optimization algorithm to perform reinforcement learning training on the MLLM based on the instruction scoring model and the result reward model to obtain the generative multimodal retrieval model includes: The sample query retrieval command is input into the command scoring model to obtain the command reward, and the retrieval candidate token corresponding to the sample query retrieval command is input into the result reward model to obtain the result supervision reward. The reward function corresponding to the near-end policy optimization algorithm is determined based on the instruction reward and the result supervision reward. The joint loss is determined based on the reward function, policy objective function, and entropy regularization term, and the MLLM is iteratively trained using the joint loss as the loss function to obtain the generative multimodal retrieval model.
[0009] The present invention also provides a retrieval method, comprising: The generative multimodal retrieval model processes the image command to be retrieved, resulting in multiple discrete token sequences; wherein the generative multimodal retrieval model is trained using the same training method as the generative multimodal retrieval model. The multiple sets of discrete token sequences are arranged in descending order, and the candidate image matched by the first set of discrete token sequences is determined as the retrieval result.
[0010] According to a retrieval method provided by the present invention, the step of processing the image command to be retrieved based on a generative multimodal retrieval model to obtain multiple sets of discrete token sequences includes: Based on the generative multimodal retrieval model, multiple scoring instructions are obtained according to the instructions for the image to be retrieved; A constrained bundle search is used to search for discrete tokens corresponding to the target scoring instruction to obtain the multiple sets of discrete token sequences; wherein, the target scoring instruction is the scoring instruction corresponding to the maximum score among the multiple scoring instructions.
[0011] According to a retrieval method provided by the present invention, the process of using a constrained bundle search to process the target scoring instruction and obtain the multiple sets of discrete token sequences includes: A multi-level decay optimization strategy combined with the constrained bundle search is used to search for the discrete tokens corresponding to the target scoring instruction, thereby obtaining the multiple sets of discrete token sequences.
[0012] The present invention also provides a training apparatus for a generative multimodal retrieval model, comprising: The instruction evolution and scoring module is used to evolve and score sample instructions to obtain an instruction scoring dataset. The first training module is used to train the large language model by fine-tuning based on the instruction scoring dataset to obtain the instruction scoring model; and to train the large language model by fine-tuning based on the result reward model dataset to obtain the result reward model, wherein the result reward model dataset is obtained based on the multimodal large language model MLLM according to discrete image tokens. The second training module is used to perform reinforcement learning training on MLLM using the proximal policy optimization algorithm based on the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model.
[0013] The present invention also provides a retrieval device, comprising: The retrieval module is used to process the image command to be retrieved based on a generative multimodal retrieval model to obtain multiple sets of discrete token sequences; wherein, the generative multimodal retrieval model is trained by the training method of the generative multimodal retrieval model; The matching module is used to sort the multiple sets of discrete token sequences in descending order and determine the candidate image that matches the first set of discrete token sequences as the retrieval result.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a training method or a retrieval method for a generative multimodal retrieval model as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a training method or retrieval method for a generative multimodal retrieval model as described above.
[0016] The generative multimodal retrieval model training method, retrieval method, and related apparatus provided by this invention evolve and score sample instructions to obtain an instruction scoring dataset. Based on the instruction scoring dataset, a large language model is trained through fine-tuning to obtain an instruction scoring model. Then, based on the result reward model dataset, the large language model is trained through fine-tuning to obtain a result reward model. Finally, a near-end policy optimization algorithm is used to perform reinforcement learning training on the MLLM based on the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model. By combining instruction evolution and reinforcement learning, the performance of the text-image cross-modal retrieval model is optimized and improved, thereby increasing the efficiency and accuracy of text-image cross-modal retrieval. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the training method for the generative multimodal retrieval model provided by the present invention.
[0019] Figure 2 This is a diagram illustrating the instruction derivation path provided by the present invention.
[0020] Figure 3 This is a flowchart illustrating the trie construction process provided by the present invention.
[0021] Figure 4 This is a flowchart illustrating the method for annotating discrete image tokens provided by the present invention.
[0022] Figure 5 This is one of the flowcharts illustrating the retrieval method provided by the present invention.
[0023] Figure 6 This is the second flowchart illustrating the retrieval method provided by the present invention.
[0024] Figure 7 This is a schematic diagram of the structure of the training device for the generative multimodal retrieval model provided by the present invention.
[0025] Figure 8 This is a schematic diagram of the generative retrieval device provided by the present invention. Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0027] The following combination Figures 1-8 The present invention describes the training method, retrieval method, and related apparatus of the generative multimodal retrieval model.
[0028] Figure 1 This is a flowchart illustrating the training method for the generative multimodal retrieval model provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps: Step 110: Evolve and score the sample instructions to obtain the instruction score dataset.
[0029] In this step, the sample instructions include text instructions or image instructions; text instructions include text data, such as, "Could you give me a picture of a rose cake?" In this embodiment, artificial intelligence-assisted tools (such as deeply customized large language models) can be used to perform deep or broad evolution of sample instructions to obtain multiple evolved sample instructions.
[0030] For example, the complexity of instructions can be increased by adding constraints, specifying inputs, adding reasoning steps, deepening logic, and complicating input formats, etc. For example, the complexity of instructions can be increased by performing the following operations to achieve the breadth evolution of instructions: topic expansion and task type expansion, etc.
[0031] In this embodiment, after obtaining multiple evolved sample instructions, each sample instruction is scored to obtain an instruction score dataset.
[0032] Specifically, different weights can be set and scored according to dimensions such as the quality, diversity, and complexity of each sample instruction.
[0033] For example, the weight of quality is set to 40%, the weight of diversity is set to 30%, and the weight of complexity is set to 30%.
[0034] Figure 2 This is a diagram illustrating the instruction derivation path provided by the present invention. Figure 2In the illustrated embodiment, the evolution process of the input original sample instruction includes upward evolution and downward evolution. If the sample instruction is "Can I have a picture of a rose cake?", the downward evolution yields the following instructions: "Can you show me a cake?", "Can I have a picture of a rose?", and "Please draw an exquisite cake with flower decorations"; the upward evolution yields the following instructions: "Please generate a double-layer chocolate frosting rose cake, 20 cm in diameter", "It needs to contain three layers of vanilla sponge cake, decorated with pink buttercream roses", and "Can you show me a picture of a yellow frosting rose cake placed on a wooden tray with a napkin at the bottom?"
[0035] Step 120: Train the large language model using fine-tuning based on the instruction scoring dataset to obtain the instruction scoring model; train the large language model using fine-tuning based on the result reward model dataset to obtain the result reward model, which is obtained based on the multimodal large language model MLLM using discrete image tokens.
[0036] In this step, the large language model can be a large language model based on the Transformer structure.
[0037] For example, the large language model is the LLaMA-7B model, which adopts a Transformer structure and is composed of multiple stacked encoders. Each layer contains a multi-head self-attention mechanism and a feed-forward network. The self-attention mechanism captures long-distance dependencies by calculating the association weights of words at different positions, while the feed-forward network enhances the model's expressive power through nonlinear transformations.
[0038] In this step, the fine-tuning methods during the construction of the instruction scoring model include: using LoRA or QLoRA techniques to freeze most of the parameters of the pre-trained model (such as LLaMA-7B), updating only the low-rank adapter parameters, reducing memory usage and adapting to the instruction scoring task; minimizing the cross-entropy loss function to optimize the model's ability to predict instruction quality, and using Alpaca-style (instruction-input-output triplet) or ShareGPT multi-turn dialogue format for the training data, ultimately obtaining the corresponding instruction scoring model.
[0039] In this step, discrete image tokens can be obtained by encoding candidate images using a visual encoder. Discrete image tokens include multiple tokens corresponding to sample instructions (images) and text descriptions corresponding to those sample instructions.
[0040] In this embodiment, the fine-tuning methods during the construction of the outcome reward model include: (1) The StableReinforce algorithm is used to improve the traditional PPO. By dynamically adjusting the loss function and advantage estimation strategy, the training stability is improved.
[0041] Among them, the reward design is divided into two categories: (1) result-supervised rewards: directly evaluate the matching degree between the output result and the labeled token (such as BLEU, ROUGE); (2) process-supervised rewards: introduce reward signals for intermediate reasoning steps (such as logical consistency scores). (2) Combine the instruction scoring model with the image token encoder to construct an end-to-end reward model, namely the result reward model.
[0042] For example, in GRPO (Group Relative Policy Optimization), the relative reward optimization model is compared between different response groups.
[0043] Step 130: Use the proximal policy optimization algorithm to perform reinforcement learning training on MLLM based on the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model.
[0044] In this step, the near-end policy optimization algorithm is used to optimize the reinforcement learning training process of the multimodal large language model (MLLM). For example, the instruction reward function corresponding to the sample instruction is first obtained through the instruction scoring model, and the result supervision reward function corresponding to the sample instruction is obtained through the result reward model. The total value loss is designed accordingly. Then, the joint loss is designed by combining the policy loss and entropy regularization term corresponding to the near-end policy optimization algorithm. Based on the joint loss, the model parameters are updated through backpropagation and optimization, and finally the trained generative multimodal retrieval model is obtained.
[0045] In this embodiment, the near-end strategy optimization algorithm loss function is used in result-supervised reinforcement learning to improve the efficiency of generative text-image cross-modal retrieval by optimizing the alignment and coordination between instruction quality and retrieval results, and dynamically balancing exploration and utilization strategies.
[0046] The generative multimodal retrieval model training method provided in this invention involves evolving and scoring sample instructions to obtain an instruction scoring dataset. Based on this dataset, a large language model is trained through fine-tuning to obtain an instruction scoring model. Then, the instruction scoring model is trained again using discrete image tokens through fine-tuning to obtain a result reward model. Finally, a near-end policy optimization algorithm is used to train the MLLM (Multi-Level Model) using reinforcement learning based on the instruction scoring model and the result reward model to obtain the generative multimodal retrieval model. This combined approach of instruction evolution and reinforcement learning optimizes and improves the performance of the text-image cross-modal retrieval model, thereby enhancing the efficiency and accuracy of text-image cross-modal retrieval.
[0047] In some embodiments, the result reward model dataset is obtained through the following steps: visually encoding candidate images based on MLLM to generate discrete image tokens, and storing them in a trie structure to obtain tree structure data; labeling the tree structure data to obtain the result reward model dataset.
[0048] In this embodiment, the images in the candidate set (Gallery) are converted into discrete image tokens by the visual encoder of MLLM and stored in a dictionary tree structure to obtain tree structure data; then, each node or path in the tree structure data is evaluated and labeled to construct a result-supervised reinforcement learning training set.
[0049] The annotation attributes include at least two of the following for each image (instruction): the discrete image token, the image description, and the image name.
[0050] For example, candidate images are first visually encoded to obtain multiple discrete tokens, which are then stored in a trie. Each trie node contains the following fields: token: The discrete encoded value of an image patch (such as a hash value or integer index); image_name: The name of the original image associated with it (used for reverse retrieval); Description: Image description text for the MS COCO dataset; children: An array of pointers to child nodes (indexed by token value).
[0051] The desired discrete image tokens, including the following attributes: token and description, are retrieved from the tree results using a restricted bundle search method.
[0052] Restricted-beam search strategies include: (1) Search space constraints: (1.1) Bundle width limit: Dynamically adjust the number of candidate paths (e.g., k=5), and prioritize retaining branches with high semantic similarity to the query description.
[0053] (1.2) Depth limit: Set the maximum search depth (e.g., d=32), and end the search after it is completed to prevent the search from exploding due to the complexity of the image.
[0054] (1.3) Trie tree constraint: MLLM uses an autoregressive method to generate discrete image tokens, selects all next nodes in the Trie tree for the current sequence, and ensures that the final generated discrete image token sequence corresponds to the candidate image.
[0055] (2) Search process: (2.1) Initialization: Starting from the root node, load the initial token substring (e.g., generate from 0) as candidate paths.
[0056] (2.2) Path expansion: (2.2.1) Breadth evolution: For the end node of the current path, traverse all possible subsequent token branches.
[0057] In the above evolution process, each selection process is as follows: the beam search method is used to select the next candidate token of the current generated sequence in the generated vocabulary, and the beam_size token with the highest probability is selected from the tokens selected in the previous step, which is the beam width.
[0058] (2.3) Result generation: The token sequences of the top-k paths are merged and reconstructed by the modified stream decoder to obtain the corresponding discrete image tokens.
[0059] Figure 3 This is a flowchart illustrating the trie construction process provided by the present invention. Figure 3 In the illustrated embodiment, the process of constructing the tree structure data (tree building) includes: encoding the images in the image library using a visual encoder to obtain 32 image tokens, vectorizing them to obtain 32 discrete tokens, and then decoding them to obtain 32 ID strings (visual tokens, for example, a string represented as...).<img_0667> Finally, these strings are stored in a trie, with each branch of the trie having 32 nodes, and each node corresponding to an image.
[0060] The training method for the generative multimodal retrieval model provided in this invention uses MLLM to visually encode candidate images, generate discrete image tokens, and store them in a dictionary tree structure to generate tree structure data. The tree structure data is labeled to obtain the result reward model dataset. By constructing the result reward model dataset through hierarchical encoding and labeling, the quality of the dataset is improved, thereby enhancing image processing efficiency and cross-modal alignment capability.
[0061] In some embodiments, a proximal policy optimization algorithm is used to train the MLLM using reinforcement learning based on an instruction scoring model and an outcome reward model to obtain a generative multimodal retrieval model, including: (1) Input the sample query retrieval command into the command scoring model to obtain the command reward, and input the retrieval candidate token corresponding to the sample query retrieval command into the result reward model to obtain the result supervision reward.
[0062] In this embodiment, the loss function of the instruction scoring model is expressed by the following formula: ; in, Indicates the actual score of the instruction. This represents the predicted score from the instruction scoring model, where N is the number of samples.
[0063] In this embodiment, the candidate set can be the public dataset MS COCO.
[0064] In this embodiment, discrete image tokens are generated using MLLM and labeled using a discriminant function. Specifically, the discrete image tokens are calculated using the following formula: ; in, This represents a sequence of discrete image tokens generated by the model; Indicates the target token sequence. For an indicator function, if and only if The output is 1 if the condition is met, otherwise the output is 0.
[0065] In this embodiment, the loss function of the outcome reward model is expressed as: ; Where N is the number of training samples. This represents the label of the i-th sample (1 for a perfect match, 0 otherwise). This represents the token sequence generated by the model. This represents the matching probability predicted by the reward model.
[0066] In this embodiment, the query retrieval command is input into the Instruction Rating Model (IRM) to obtain the instruction reward. The candidate set is input into the result-supervised reward model (ORM) to obtain the result-supervised reward. .
[0067] (2) Determine the reward function corresponding to the near-end policy optimization algorithm based on instruction reward and result supervision reward.
[0068] In this embodiment, the reward function Specifically, it is expressed by the following formula: ; in, The score output by the instruction scoring model. The score output by the model is used as a reward for the outcome.
[0069] (3) Determine the joint loss based on the reward function, policy objective function and entropy regularization term, and use the joint loss as the loss function to iteratively train MLLM to obtain a generative multimodal retrieval model.
[0070] In this embodiment, the core formula of the near-short strategy optimization algorithm is as follows: (3.1) The policy objective function is expressed as: ; (3.2) The corresponding value function loss is calculated based on the reward function using the following formula, specifically expressed as: ; (3.3) The entropy regularization term is expressed as: ; Correspondingly, the joint loss is expressed by the following formula: ; in, This is the current strategy (model to be optimized). This is the old strategy (the model before the update). This represents the advantage function calculated using generalized advantage estimation (GAE). The result of the total reward function calculation. , This represents the loss weighting coefficient. The pruning threshold is set; the model parameters W are updated through backpropagation and optimization; after 200 iterations, the model parameters with the best performance are saved as the optimal model, i.e., the generative multimodal retrieval model.
[0071] The training method for the generative multimodal retrieval model provided in this embodiment of the invention determines the reward function by combining the outputs of the instruction scoring model and the result reward model, and iteratively trains the MLLM with the joint loss determined by the reward function, the policy objective function and the entropy regularization term to obtain the generative multimodal retrieval model. The accuracy of the model in retrieval is further improved by the result-supervised reinforcement learning method based on the instruction scoring model.
[0072] exist Figure 4 In the illustrated embodiment, during the construction of the result reward model dataset, the result reward corresponding to the sample instruction "Can I have a picture of a rose cake?" is "012.jpg" (the target answer is in image format); this sample instruction is input into MLLM for the first restricted-bundle search, generating the discrete token sequence "T= t 1, t 2, t 3,…, t 32 The corresponding answer is "098. Jpg". × This indicates that the answer does not match the reward given above, and the corresponding result is marked as: y r =0; First constrained search generates discrete token sequence "T= t 1, t 2, t 3,…, t 32The corresponding answer is "012. Jpg√", indicating that the answer matches the reward given above. The corresponding result is marked as follows: y r =1; After performing multiple searches in this manner, the corresponding result reward model dataset is obtained.
[0073] The retrieval method provided by this invention will be described below. The retrieval method described below can be referred to in correspondence with the training method of the generative multimodal retrieval model described above.
[0074] Figure 5 This is one of the flowcharts illustrating the retrieval method provided by the present invention, such as... Figure 5 As shown, the retrieval method includes the following steps: Step 510: Process the image command to be retrieved based on the generative multimodal retrieval model to obtain multiple sets of discrete token sequences; wherein, the generative multimodal retrieval model is trained by the generative multimodal retrieval model training method.
[0075] In this step, the generative multimodal retrieval model is implemented through the following steps: The sample instructions are evolved and scored to obtain an instruction score dataset; The instruction scoring model is obtained by training a large language model through fine-tuning based on the instruction scoring dataset; the result reward model is obtained by training the instruction scoring model through fine-tuning based on discrete image tokens, and the discrete image tokens are obtained based on the multimodal large language model MLLM. A generative multimodal retrieval model is obtained by using a proximal policy optimization algorithm to train MLLM based on the instruction scoring model and the result reward model.
[0076] The implementation methods of the above steps correspond one-to-one with the implementation steps of steps 110-130 above, and will not be repeated in this embodiment.
[0077] In this embodiment, an image retrieval command is input, and the generative multimodal retrieval model scores the evolved command through a command scoring model, and selects the command with the highest score as input to obtain multiple sets of discrete image token sequences.
[0078] Step 520: Sort the multiple sets of discrete token sequences in descending order, and determine the candidate image that matches the first set of discrete token sequences as the retrieval result.
[0079] In this embodiment, a specified optimization strategy is used to optimize the token candidates, and the results are sorted in descending order. The resulting Rank-1 (the first set of discrete token sequences) is taken as the optimal token sequence result.
[0080] Specifically, the search process includes: (1) Input user retrieval instructions to the MLLM of instruction evolution and instruction reward model IRM, generate the evolved instructions and their corresponding reward scores; (2) Select the instruction with the highest score after evolution and input it into the optimal model. M This generates multiple candidate sequences; (3) Calculate the score of each generated sequence and sort them in descending order, where Rank-1 is the optimal retrieval result.
[0081] Finally, the images in the candidate set that match the optimal token sequence result are the final retrieval results.
[0082] The retrieval method provided in this embodiment of the invention processes the image command to be retrieved by a generative multimodal retrieval model trained by the generative multimodal retrieval model training method, and obtains multiple sets of discrete token sequences. The multiple sets of discrete token sequences are arranged in descending order, and the candidate image matched by the first set of discrete token sequences is determined as the retrieval result, thereby improving the efficiency and accuracy of generative text-image cross-modal retrieval.
[0083] In some embodiments, processing the image command to be retrieved based on the generative multimodal retrieval model to obtain multiple discrete token sequences includes: obtaining multiple rating commands based on the image command to be retrieved using the generative multimodal retrieval model; and using constrained bundle search to search for the discrete tokens corresponding to the target rating command to obtain multiple discrete token sequences; wherein, the target rating command is the rating command corresponding to the maximum rating among the multiple rating commands.
[0084] In this embodiment, the multimodal scoring instruction is obtained through the following steps: An open instruction generation strategy similar to the BGE-VL model is adopted, and MLLM (such as InternVL2-26B) is used to score the clarity, semantic coverage and other dimensions of the instructions.
[0085] For example, (1) scoring dimensions: including instruction completeness, clarity and task complexity (such as whether multi-step reasoning is involved); (2) screening strategy: retain the instruction with the highest score as the target scoring instruction and eliminate low-quality candidates (such as instructions with a score lower than the threshold of 3.0).
[0086] In this embodiment, the generated instructions need to be semantically aligned with the image token. For example, an EVA-CLIP visual encoder is used to extract image features, and residual quantization is used to map the text instructions and image features to a shared semantic space, ensuring that the instructions accurately describe the semantic hierarchy of the discrete token. In this embodiment, obtaining the discrete image token sequence using the Restricted Bundle Search (RBS) method includes the following steps: (1) Search space construction: Discrete tokens are generated by a visual encoder (such as ViT or ConvNext) and stored in a trie structure.
[0087] Each tree node contains attributes such as token value, image name, and description text.
[0088] (2) Dynamic constraint strategy: Bundle width control: Limit the number of candidate paths per layer expansion (e.g., k=5) to avoid computational explosion.
[0089] Semantic pruning: Calculate the cosine similarity between candidate paths and target instructions using CLIP embedding (threshold > 0.7) to filter irrelevant branches.
[0090] Path scoring: Joint probability = token matching degree × structural integrity weight (e.g., score when leaf node not reached × 0.8).
[0091] (3) Token sequence generation Starting from the tree node corresponding to the target scoring instruction, generate multiple token sequences through the following steps: (3.1) Path expansion: For the terminal node of the current path, traverse all possible subsequent token branches; (3.2) Path merging: Merge branches with high semantic similarity (such as the blue tokens corresponding to the descriptions of "sky" and "clouds"); (3.3) Result optimization: Modify the stream decoding of the token sequence of the Top-k path to generate the final image or cross-modal retrieval result.
[0092] The retrieval method provided in this invention obtains multiple scoring instructions based on the image instructions to be retrieved through a generative multimodal retrieval model; it uses constrained bundle search to search for discrete tokens corresponding to the target scoring instructions, obtaining multiple sets of discrete token sequences; it filters high-quality instructions through an instruction scoring model, and combines dynamic path optimization of constrained bundle search to achieve high efficiency and accuracy in multimodal retrieval tasks.
[0093] In some embodiments, processing the target scoring instruction using constrained bundle search to obtain multiple sets of discrete token sequences includes: using a multi-level decay optimization strategy combined with constrained bundle search to search for the discrete tokens corresponding to the target scoring instruction to obtain multiple sets of discrete token sequences.
[0094] In this embodiment, the multi-level attenuation optimization strategy is expressed by the following formula: ; in, Indicates the generation of the first i The conditional probability when there are 1 token. The decay coefficient (ranging from 0.8 to 1) indicates that a smaller value means that the preceding tokens have a significant impact on the score.
[0095] In this embodiment, the multi-level attenuation optimization strategy includes: (1) Dynamic adjustment of learning rate; combining cosine annealing and exponential decay strategies for adjustment at different stages: In the initial stage: Use linear warm-up (e.g., 5 epochs) to gradually increase the learning rate to avoid gradient explosion; Mid-term: Use exponential decay (γ=0.9) to accelerate convergence; Later stage: Switch to cosine annealing (T_max=100) and fine-tune the parameters.
[0096] (2) Search weight decay; introduce a decay factor into the scoring function of candidate paths: (2.1) Path length penalty: Apply an exponentially decaying weight (e.g., score × 0.8^d, where d is the current depth) to paths that do not reach the leaf nodes. (2.2) Duplicate token suppression: The priority of redundant paths is dynamically reduced by n-gram duplicate detection, and finally multiple sets of discrete token sequences are obtained.
[0097] The retrieval method provided in this embodiment of the invention uses a multi-level attenuation optimization strategy combined with constrained bundle search to search for discrete tokens corresponding to the target scoring instruction, thereby obtaining multiple sets of discrete token sequences. Through the synergistic effect of constraints and attenuation, the quality and efficiency of discrete image token generation are improved.
[0098] Figure 6 This is the second flowchart illustrating the retrieval method provided by the present invention. Figure 6 In the illustrated embodiment, the retrieval method includes three stages: Step 1-3. Step 1 involves training the instruction scoring model and the result-supervised reward model (same as the result-reward model described above). Specifically, sample instructions are input into a large language model for instruction evolution, resulting in 2-4 instructions. These instructions are then scored using the large language model to obtain an instruction scoring dataset, which is then used to train the LLaMa7B dataset to obtain the instruction scoring model. Additionally, the sample instruction "Can I have a picture of a rose?" is input into a generative retrieval system, and the generated discrete image tokens are labeled using a discriminant function (e.g., image 1 is labeled as...). × Mark image 1 with a √, and mark image 1 with a √. × The result reward model dataset is obtained and LLaMa7B is trained to obtain the corresponding result reward model.
[0099] Step 2: Proximal Policy Optimization (PPO) Reinforcement Learning Training; Specifically, the PPO algorithm is used to perform reinforcement learning on the MLLM, and the instruction scoring model and result reward model obtained in Step 1 are used to guide model training (the instruction reward r is generated through the instruction scoring model). i The outcome reward r is generated through the outcome reward model. o Finally, two types of rewards are combined: r t =r i ·r o (and thereby optimize the PPO algorithm parameters) to obtain the final generative multimodal retrieval model.
[0100] Step 3: Activate instruction evolution and perform reasoning; specifically, first, evolve the input instruction "Can I have a picture of a rose cake?" and obtain the corresponding instruction rating dataset through the instruction rating model. Then, input the instruction rating dataset into the trained multimodal large model, and use a multi-level decay optimization strategy combined with constrained search to search for discrete tokens (Token1, Token2, ... Token31 and Token32) corresponding to different rating instructions, thereby determining the optimal retrieval. The training device of the generative multimodal retrieval model provided by this invention is described below. The training device of the generative multimodal retrieval model described below can be referred to in correspondence with the training method of the generative multimodal retrieval model described above.
[0101] Figure 7 This is a schematic diagram of the structure of the training device for the generative multimodal retrieval model provided by the present invention, as shown below. Figure 7 As shown, the training device for the generative multimodal retrieval model includes: an instruction evolution and scoring module 710, a first training module 720, and a second training module 730.
[0102] The instruction evolution and scoring module 710 is used to evolve and score sample instructions to obtain an instruction scoring dataset; The first training module is used to train the large language model by fine-tuning based on the instruction scoring dataset to obtain the instruction scoring model; and to train the large language model by fine-tuning based on the result reward model dataset to obtain the result reward model. The result reward model dataset is obtained based on the multimodal large language model MLLM using discrete image tokens. The second training module is used to perform reinforcement learning training on MLLM based on the instruction scoring model and the result reward model using the proximal policy optimization algorithm to obtain a generative multimodal retrieval model.
[0103] The training device for the generative multimodal retrieval model provided in this embodiment of the invention evolves and scores sample instructions to obtain an instruction scoring dataset. Based on the instruction scoring dataset, a large language model is trained through fine-tuning to obtain an instruction scoring model. Then, based on discrete image tokens, the instruction scoring model is trained through fine-tuning to obtain a result reward model. Finally, a near-end policy optimization algorithm is used to perform reinforcement learning training on the MLLM based on the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model. By combining instruction evolution and reinforcement learning, the performance of the text-image cross-modal retrieval model is optimized and improved, thereby improving the efficiency and accuracy of text-image cross-modal retrieval.
[0104] Figure 8 This is a schematic diagram of the retrieval device provided by the present invention, as shown below. Figure 8 As shown, the retrieval device includes a retrieval module 810 and a matching module 820.
[0105] The retrieval module 810 is used to process the image command to be retrieved based on the generative multimodal retrieval model to obtain multiple sets of discrete token sequences; wherein, the generative multimodal retrieval model is trained by the generative multimodal retrieval model training method; The matching module 820 is used to sort multiple sets of discrete token sequences in descending order and determine the candidate image matched by the first set of discrete token sequences as the retrieval result.
[0106] The retrieval device provided in this embodiment of the invention processes the image command to be retrieved by a generative multimodal retrieval model trained by a generative multimodal retrieval model training method, and obtains multiple sets of discrete token sequences; the multiple sets of discrete token sequences are arranged in descending order, and the candidate image matched by the first set of discrete token sequences is determined as the retrieval result, thereby improving the efficiency and accuracy of generative text-image cross-modal retrieval.
[0107] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute a training method for a generative multimodal retrieval model. This method includes: evolving and scoring sample instructions to obtain an instruction scoring dataset; training a large language model (MLLM) based on the instruction scoring dataset through fine-tuning to obtain an instruction scoring model; training the MLLM based on a result reward model dataset through fine-tuning to obtain a result reward model, wherein the result reward model dataset is obtained based on the multimodal large language model (MLLM) using discrete image tokens; and using a proximal policy optimization algorithm to perform reinforcement learning training on the MLLM based on the instruction scoring model and the result reward model to obtain the generative multimodal retrieval model.
[0108] Alternatively, a retrieval method may be executed, comprising: processing the image command to be retrieved based on a generative multimodal retrieval model to obtain multiple sets of discrete token sequences; wherein the generative multimodal retrieval model is trained by a generative multimodal retrieval model training method; arranging the multiple sets of discrete token sequences in descending order, and determining the candidate image matched by the first set of discrete token sequences as the retrieval result.
[0109] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the training method of the generative multimodal retrieval model provided by the above methods. The method includes: evolving and scoring sample instructions to obtain an instruction scoring dataset; training a large language model based on the result reward model dataset through fine-tuning to obtain a result reward model, wherein the result reward model dataset is obtained based on the multimodal large language model MLLM according to discrete image tokens; and using a proximal policy optimization algorithm to perform reinforcement learning training on the MLLM according to the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model.
[0111] Alternatively, a retrieval method may be executed, comprising: processing the image command to be retrieved based on a generative multimodal retrieval model to obtain multiple sets of discrete token sequences; wherein the generative multimodal retrieval model is trained by a generative multimodal retrieval model training method; arranging the multiple sets of discrete token sequences in descending order, and determining the candidate image matched by the first set of discrete token sequences as the retrieval result.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a generative multimodal retrieval model, characterized in that, include: The sample instructions are evolved and scored to obtain an instruction score dataset; The instruction scoring model is trained by fine-tuning the large language model based on the instruction scoring dataset to obtain the instruction scoring model; the result reward model is trained by fine-tuning the large language model based on the result reward model dataset, wherein the result reward model dataset is obtained based on the multimodal large language model MLLM according to discrete image tokens. The MLLM is trained by a proximal policy optimization algorithm based on the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model.
2. The training method for the generative multimodal retrieval model according to claim 1, characterized in that, The resulting reward model dataset was obtained through the following steps: Based on the MLLM, the candidate images are visually encoded to generate discrete image tokens, which are then stored in a trie structure to obtain tree structure data. The tree structure data is labeled to obtain the result reward model dataset.
3. The training method for the generative multimodal retrieval model according to claim 1, characterized in that, The method employs a proximal strategy optimization algorithm to train the MLLM using reinforcement learning based on the instruction scoring model and the result reward model, resulting in a generative multimodal retrieval model, including: The sample query retrieval command is input into the command scoring model to obtain the command reward, and the retrieval candidate token corresponding to the sample query retrieval command is input into the result reward model to obtain the result supervision reward. The reward function corresponding to the near-end policy optimization algorithm is determined based on the instruction reward and the result supervision reward. The joint loss is determined based on the reward function, policy objective function, and entropy regularization term, and the MLLM is iteratively trained using the joint loss as the loss function to obtain the generative multimodal retrieval model.
4. A retrieval method, characterized in that, include: The generative multimodal retrieval model is used to process the image command to be retrieved, resulting in multiple discrete token sequences; wherein the generative multimodal retrieval model is trained by the training method of the generative multimodal retrieval model as described in any one of claims 1-3. The multiple sets of discrete token sequences are arranged in descending order, and the candidate image matched by the first set of discrete token sequences is determined as the retrieval result.
5. The retrieval method according to claim 4, characterized in that, The generative multimodal retrieval model processes the image command to be retrieved, resulting in multiple discrete token sequences, including: Based on the generative multimodal retrieval model, multiple scoring instructions are obtained according to the instructions for the image to be retrieved; A constrained-bind search is used to search for discrete tokens corresponding to the target scoring instruction to obtain the multiple sets of discrete token sequences; wherein, the target scoring instruction is the scoring instruction corresponding to the maximum score among the multiple scoring instructions.
6. The retrieval method according to claim 5, characterized in that, The process of using constrained bundle search to process the target scoring instruction to obtain the multiple sets of discrete token sequences includes: A multi-level decay optimization strategy combined with the constrained bundle search is used to search for the discrete tokens corresponding to the target scoring instruction, thereby obtaining the multiple sets of discrete token sequences.
7. A training device for a generative multimodal retrieval model, characterized in that, include: The instruction evolution and scoring module is used to evolve and score sample instructions to obtain an instruction scoring dataset. The first training module is used to train the large language model by fine-tuning based on the instruction scoring dataset to obtain the instruction scoring model; and to train the large language model by fine-tuning based on the result reward model dataset to obtain the result reward model, wherein the result reward model dataset is obtained based on the multimodal large language model MLLM according to discrete image tokens. The second training module is used to perform reinforcement learning training on MLLM using the proximal policy optimization algorithm based on the instruction scoring model and the result reward model to obtain a generative multimodal retrieval model.
8. A retrieval device, characterized in that, include: The retrieval module is used to process the image command to be retrieved based on a generative multimodal retrieval model to obtain multiple sets of discrete token sequences; wherein the generative multimodal retrieval model is trained by the training method of the generative multimodal retrieval model as described in any one of claims 1-3. The matching module is used to sort the multiple sets of discrete token sequences in descending order and determine the candidate image that matches the first set of discrete token sequences as the retrieval result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the training method of the generative multimodal retrieval model as described in any one of claims 1 to 3 or the retrieval method as described in any one of claims 4 to 6.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the training method of the generative multimodal retrieval model as described in any one of claims 1 to 3 or the retrieval method as described in any one of claims 4 to 6.
Citation Information
Cited By
Intelligent application interaction method and system driven by multi-mode end side model
CN121188100A