Model Inference Method, Apparatus, Device, and Medium for Vision-Language Models
Through semantic guidance network and hybrid mask strategy, the visual language model is optimized, the target problem set is generated and iteratively trained, which solves the accuracy and completeness of the visual language model in complex tasks, and improves the model reasoning ability and user experience.
Patent Information
- Application Number
- CN202510535090.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-27
AI Technical Summary
When existing visual language models deal with complex visual and text information fusion tasks, they cannot accurately and in-depth understanding and processing, resulting in poor accuracy and completeness of output results. Automated data construction faces the problem of modal inconsistency, making it difficult to deal with fine-grained visual language understanding tasks.
By determining the target entity set based on semantic guidance network, initial visual language model and visual samples, an initial single-step problem set is generated, and the model is fine-tuned in combination with preset semantic expansion strategies and mixed mask strategies. Finally, iterative training is used for distillation models to improve the model inference ability.
It improves the model reasoning ability of the visual language model, improves the user experience, and enhances the model's performance in complex tasks and the ability to relate visual and text information in detail.
Smart Images

Figure CN120046743B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of model reasoning, and in particular to a model reasoning method, device, equipment and medium for a visual language model. Background Art
[0002] In the field of visual language models (VLMs), with the continuous development of science and technology, their applications in joint visual and text information processing tasks have become more and more extensive, such as image caption generation and visual question answering, showing great potential.
[0003] However, current VLMs still face many severe challenges in practical applications. The existing way of constructing instruction datasets is difficult to meet large-scale needs. Methods based on self-supervised learning still have obvious limitations in multimodal semantic understanding and cross-modal collaboration. This makes it impossible for the model to accurately and deeply understand and process complex visual and text information fusion tasks, resulting in poor accuracy and completeness of the output results. In addition, attempts to automate data construction are plagued by modal inconsistency. This inconsistency makes it difficult for the model to handle fine-grained visual language understanding tasks, and it is impossible to make detailed and accurate associations and interpretations of visual and text information, which in turn affects the performance of the model in complex tasks.
[0004] Therefore, how to improve the model reasoning ability of visual language models is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a model reasoning method, device, equipment and medium for a visual language model, which can improve the model reasoning ability of the visual language model. The specific scheme is as follows:
[0006] In a first aspect, the present application provides a model reasoning method for a visual language model, comprising:
[0007] A target entity set is determined based on a semantic guidance network, an initial visual language model, a visual sample, and a first text description corresponding to the visual sample, and entities whose uncertainties meet a preset threshold condition are determined from the target entity set using a result of classification prediction of the target entity set, and an initial single-step question set is generated based on the entities whose uncertainties meet the preset threshold condition; a single-step question is a single logical relationship question;
[0008] Optimizing the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample to obtain a target single-step question set;
[0009] Determine a target multi-step problem set by using the target single-step problem set, a preset semantic expansion strategy, and a preset problem reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step problem set, the target multi-step problem set, a preset model reasoning strategy, and the initial vision-language model; a multi-step problem is multiple problems with a logical progressive relationship based on a single-step problem.
[0010] Determine a hybrid masking strategy based on the visual samples, the first text description, a preset visual masking probability, and a preset text masking probability, and use the hybrid masking strategy to fine-tune the first fine-tuned model to obtain a second fine-tuned model.
[0011] Determine the second fine-tuned model as the student model, and use a preset teacher model to distill the student model to obtain a distilled model, determine the sample difficulty of the training sample set, and iteratively train the distilled model based on the training sample set and the sample difficulty to trigger model reasoning using the obtained target vision-language model.
[0012] Optionally, the determining the target entity set based on a semantic guidance network, an initial vision-language model, visual samples, and a first text description corresponding to the visual samples includes:
[0013] Use the initial vision-language model to parse the visual samples and the first text description corresponding to the visual samples to obtain corresponding cross-modal representations.
[0014] Extract an initial entity set in the visual samples and the first text description by using the cross-modal representations through the semantic guidance network to complete the entity extraction operation.
[0015] Determine a first importance degree and a second importance degree of the initial entity set; the first importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the visual information, and the second importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the text description.
[0016] Determine the importance score of each initial entity in the initial entity set based on the first importance degree, the second importance degree, a preset visual weight control parameter, and a preset text weight control parameter.
[0017] In the initial entity set, determine the initial entities with importance scores greater than a preset importance threshold as target entities, and construct a target entity set based on each of the target entities.
[0018] Optionally, in the process of determining entities with uncertainty meeting a preset threshold condition from the target entity set by using the classification prediction result of the target entity set and generating an initial single-step question set based on the entities with uncertainty meeting the preset threshold condition, the following steps are included:
[0019] Perform classification prediction on the target entity set and obtain a classification prediction result;
[0020] Determine corresponding category information from the classification prediction result, and determine a first target confidence level with a category probability value greater than a preset probability threshold in the category information;
[0021] Determine the uncertainty results of each target entity in the target entity set based on a preset constant value and the first target confidence level, and use the uncertainty results to determine entities with uncertainty meeting the preset threshold condition from the target entity set;
[0022] Determine the target importance score and target uncertainty of the entities with uncertainty meeting the preset threshold condition;
[0023] Determine an entity weight result based on the target importance score, the target uncertainty, a preset importance control coefficient, and a preset uncertainty control coefficient;
[0024] Use the entity weight result to correspondingly adjust the question type of the generated single-step question, so as to determine an initial single-step question set based on the adjusted question type.
[0025] Optionally, the optimization of the initial single-step question set based on the initial single-step question set, the first text description, and a second text description having no preset corresponding relationship with the visual sample to obtain a target single-step question set includes:
[0026] Construct a positive sample by using the initial single-step question set and the first text description corresponding to the visual sample, and construct a negative sample by using the initial single-step question set and the second text description having no preset corresponding relationship with the visual sample;
[0027] In the positive sample, determine a first semantic similarity between the initial single-step question set and the first text description based on the embedding vectors of the initial single-step question set and the first text description;
[0028] In the negative sample, determine a second semantic similarity between the initial single-step question set and the second text description based on the embedding vectors of the initial single-step question set and the second text description;
[0029] Determine a target loss value based on the first semantic similarity, the second semantic similarity, and a preset contrastive learning loss function;
[0030] Optimize the initial single-step problem set using the target loss value to obtain a target single-step problem set.
[0031] Optionally, the determining of the training sample set and the first fine-tuned model based on the target single-step problem set, the target multi-step problem set, a preset model inference strategy, and the initial vision-language model includes:
[0032] Fine-tune the initial vision-language model using the target single-step problem set, the target multi-step problem set, and a preset model inference strategy to obtain a first fine-tuned model;
[0033] During the process of fine-tuning the initial vision-language model, determine the training sample set based on the target single-step problem set and the target multi-step problem set.
[0034] Optionally, the determining of the hybrid masking strategy based on the visual samples, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tuning the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model includes:
[0035] Divide the visual samples into several regions, and divide the first text description into several words;
[0036] Determine the importance degree of visual information based on the several regions and a first adaptive weight coefficient, and determine the importance degree of text information based on the several words and a second adaptive weight coefficient;
[0037] Determine the hybrid masking strategy based on the importance degree of visual information, the preset visual masking probability, the importance degree of text information, and the preset text masking probability;
[0038] Fine-tune the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model.
[0039] Optionally, the determining of the sample difficulty of the training sample set, and iteratively training the distillation model based on the training sample set and the sample difficulty to trigger model inference using the obtained target vision-language model includes:
[0040] Determine the second target confidence level of the distillation model for predicting each current problem in the training sample set, and determine the target loss value of the distillation model on each current problem;
[0041] Determine the sample difficulty of the training sample set based on the second target confidence level and the target loss value, and iteratively train the distillation model based on the training sample set and the sample difficulty to obtain target model parameters;
[0042] Adjust the current model parameters of the distillation model using the target model parameters, and obtain a target vision-language model, so as to trigger a preset model inference operation using the target vision-language model.
[0043] In a second aspect, the present application provides a model inference device for a vision-language model, including:
[0044] A single-step question generation module, configured to determine a target entity set based on a semantic guidance network, an initial vision-language model, a visual sample, and a first text description corresponding to the visual sample, and determine entities that meet a preset threshold condition of uncertainty from the target entity set using the result of classifying and predicting the target entity set, and generate an initial single-step question set based on the entities that meet the preset threshold condition of uncertainty; the single-step question is a single logical relationship question;
[0045] A single-step question optimization module, configured to optimize the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample, so as to obtain a target single-step question set;
[0046] A first model fine-tuning module, configured to determine a target multi-step question set using the target single-step question set, a preset semantic expansion strategy, and a preset question reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step question set, the target multi-step question set, a preset model inference strategy, and the initial vision-language model; the multi-step question is multiple questions with a logical progressive relationship based on the single-step question;
[0047] A second model fine-tuning module, configured to determine a mixed masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tune the first fine-tuned model using the mixed masking strategy to obtain a second fine-tuned model;
[0048] A model inference module, configured to determine the second fine-tuned model as a student model, and perform distillation on the student model using a preset teacher model to obtain a distillation model, determine the sample difficulty of the training sample set, and perform iterative training on the distillation model based on the training sample set and the sample difficulty, so as to trigger model inference using the obtained target vision-language model.
[0049] In a third aspect, the present application provides an electronic device, including:
[0050] A memory, configured to store a computer program;
[0051] A processor for executing the computer program to implement the foregoing model inference method for a vision-language model.
[0052] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing model inference method for a vision-language model is implemented.
[0053] In this application, a target entity set is determined based on a semantic guidance network, an initial vision-language model, vision samples, and a first text description corresponding to the vision samples. Entities that meet the preset threshold condition for uncertainty are determined from the target entity set using the result of classifying and predicting the target entity set. An initial single-step question set is generated based on the entities that meet the preset threshold condition for uncertainty. A single-step question is a question with a single logical relationship. The initial single-step question set is optimized based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the vision samples to obtain a target single-step question set. A target multi-step question set is determined using the target single-step question set, a preset semantic extension strategy, and a preset question reasoning strategy. A multi-step question is a series of questions with a logical progression based on a single-step question. A training sample set and a first fine-tuned model are determined based on the target single-step question set, the target multi-step question set, a preset model reasoning strategy, and the initial vision-language model. A hybrid masking strategy is determined based on the vision samples, the first text description, a preset vision masking probability, and a preset text masking probability. The first fine-tuned model is fine-tuned using the hybrid masking strategy to obtain a second fine-tuned model. The second fine-tuned model is determined as the student model, and the student model is distilled using a preset teacher model to obtain a distilled model. The sample difficulty of the training sample set is determined, and the distilled model is iteratively trained based on the training sample set and the sample difficulty to trigger model reasoning using the obtained target vision-language model. As can be seen from the above, in this application, first, a target entity set is determined by combining a semantic guidance network, an initial vision-language model, vision samples, and a corresponding first text description. The target entities are classified and predicted to find entities with uncertainty reaching the preset threshold, and an initial single-step question set with a single logical relationship is generated based on these entities. Next, the initial single-step questions are optimized using the initial single-step question set, the first text description, and a second text description that has no relation to the vision samples to obtain a target single-step question set. Subsequently, based on the target single-step question set, a target multi-step question set with a logical progression is obtained by applying a preset semantic extension and question reasoning strategy. In combination with a preset model reasoning strategy, the initial vision-language model is fine-tuned to determine a training sample set and obtain a first fine-tuned model. Then, based on the vision samples, the first text description, and preset vision and text masking probabilities, a hybrid masking strategy is determined. The first fine-tuned model is fine-tuned again using this strategy to obtain a second fine-tuned model. Finally, the second fine-tuned model is used as the student model, and a distilled model is obtained by distilling it with a preset teacher model. The sample difficulty of the training sample set is determined, and the distilled model is iteratively trained based on this to trigger model reasoning using the obtained target vision-language model. In this way, this application can improve the model reasoning ability of the vision-language model and enhance the user experience to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0055] Figure 1 Flowchart of a model inference method for a vision-language model disclosed in this application;
[0056] Figure 2 Flowchart of a specific model inference method for a vision-language model disclosed in this application;
[0057] Figure 3 Structure diagram of a model inference device for a vision-language model disclosed in this application;
[0058] Figure 4 Structure diagram of an electronic device disclosed in this application. Detailed implementation manners
[0059] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0060] Currently, VLMs face many severe challenges in practical applications. The existing construction methods of instruction datasets are difficult to meet large-scale requirements. The methods based on self-supervised learning still have obvious limitations in multi-modal semantic understanding and cross-modal collaboration. This makes it impossible for the model to accurately and deeply understand and process complex visual and text information fusion tasks, resulting in poor accuracy and integrity of the output results. Moreover, attempts at automated data construction are troubled by modal inconsistencies. Such inconsistencies make it difficult for the model to handle fine-grained vision-language understanding tasks, unable to make detailed and accurate associations and interpretations of visual and text information, thereby affecting the performance of the model in complex tasks. Therefore, this application provides a model inference method, device, device, and medium for vision-language models, which can improve the model inference ability of vision-language models.
[0061] See Figure 1 As shown, the embodiments of the present invention disclose a model inference method for a vision-language model, including:
[0062] Step S11: Determine a target entity set based on a semantic guidance network, an initial vision - language model, visual samples, and first - text descriptions corresponding to the visual samples, and determine entities in the target entity set whose uncertainty meets a preset threshold condition by using the result of classification prediction for the target entity set. Generate an initial single - step question set based on the entities whose uncertainty meets the preset threshold condition; a single - step question is a single - logical - relationship question.
[0063] In this embodiment, first, use the initial vision - language model to parse the visual samples and the first - text descriptions corresponding to the visual samples. Among them, the initial vision - language model has the ability to process visual and text information. Through its internal algorithms and structures, it can identify and understand visual elements such as objects, attributes, and spatial relationships in visual samples, and at the same time, it can also extract and analyze semantic information in the first - text descriptions. During the parsing process, the initial vision - language model fuses the visual information and the text information to obtain corresponding cross - modal representations. This cross - modal representation integrates information from both the visual and text dimensions, providing certain basic data for subsequent entity extraction.
[0064] Next, use the semantic guidance network (i.e., SGN, Semantics - Guided Neural Network) to extract the initial entity set in the visual samples and text descriptions by using the above - obtained cross - modal representation to complete the entity extraction operation. The semantic guidance network can extract entities from cross - modal information. It can identify entities in visual samples and text descriptions according to the feature information in the cross - modal representation. These entities may be objects, people, scenes, etc., which are the core elements constituting the visual scene and text description. During the extraction process, the semantic guidance network will deeply analyze the cross - modal representation, and use specific algorithms and rules to screen out entities that meet certain criteria, thus forming the initial entity set.
[0065] Then, determine the first importance degree and the second importance degree of the initial entity set. It should be noted that the first importance degree is the importance degree of the initial entity predicted by the initial vision - language model in visual information. This is obtained by analyzing the visual samples through the initial vision - language model and evaluating factors such as the prominence of each initial entity in the visual scene and the degree of association with other visual elements. The second importance degree is the importance degree of the initial entity predicted by the initial vision - language model in the text description. The initial vision - language model will analyze factors such as the mention frequency of each entity in the text and the detailed degree of description to determine its importance in the text description.
[0066] Furthermore, the importance score of each initial entity in the initial entity set is determined based on the first importance, the second importance, the preset visual weight control parameter and the preset text weight control parameter. Specifically, in order to comprehensively consider the impact of both visual and textual aspects on the importance of entities, preset visual weight control parameters and preset text weight control parameters are introduced. These two parameters are used to adjust the relative weights of visual information and textual information when determining the importance score of an entity. Through a specific calculation formula, the first importance is multiplied by the preset visual weight control parameter, the second importance is multiplied by the preset text weight control parameter, and then the two results are added to obtain the importance score of each initial entity. This method can flexibly balance the contribution of visual and textual information to the importance of entities according to actual needs.
[0067] In the initial entity set, the initial entities with importance scores greater than the preset importance threshold are determined as target entities, so as to construct the target entity set based on each target entity. In other words, only those initial entities with importance scores exceeding the preset importance threshold are identified as target entities, and these target entities constitute the target entity set. The target entity set contains the most critical information elements in the visual sample and text description, and is the core basis for subsequent reasoning and question generation.
[0068] Next, the target entity set is classified and predicted, and the classification prediction results are obtained. The initial visual language model will classify each target entity in the target entity set and determine the category to which it belongs. For example, a target entity may be judged to be a "fruit" category, or an "animal" category, etc. In addition, the classification prediction results contain the category information of each target entity and the corresponding category probability value.
[0069] In a specific implementation, the corresponding category information is determined from the classification prediction result, and the first target confidence level of the category probability value greater than the preset probability threshold is determined in the category information. That is, only when the probability value of a category of a target entity is greater than the preset probability threshold, the category is considered to be a reliable classification of the target entity, and the corresponding confidence level is the first target confidence level.
[0070] Then, in this embodiment, the uncertainty results of each target entity in the target entity set are determined based on the preset constant value and the first target confidence, and the entities whose uncertainties meet the preset threshold conditions are determined from the target entity set using the uncertainty results. Generally, the uncertainty results can be obtained by subtracting the first target confidence from the preset constant value. Then, the uncertainty results of each target entity are compared with the preset threshold conditions to screen out entities whose uncertainties meet the preset threshold conditions. In addition, these entities have high uncertainty in classification prediction, which may be caused by unclear visual features, vague text descriptions, and other reasons.
[0071] Further, in this embodiment, the target importance score and target uncertainty of the entities whose uncertainty meets the preset threshold condition are determined. For these selected entities, their target importance score and target uncertainty are determined again. The determination method of the target importance score is similar to that of determining the importance score of the initial entities before, except that it is calculated for these specific entities, and the target uncertainty is the uncertainty result obtained from the above calculation.
[0072] After obtaining the target importance score and target uncertainty, an entity weight result is determined based on the target importance score, target uncertainty, preset importance control coefficient, and preset uncertainty control coefficient. The preset importance control coefficient and preset uncertainty control coefficient are used to adjust the relative importance of the target importance score and target uncertainty when determining the entity weight result. Through a specific calculation formula, the target importance score is multiplied by the preset importance control coefficient, the target uncertainty is multiplied by the preset uncertainty control coefficient, and then these two results are added together to obtain the entity weight result of each entity whose uncertainty meets the preset threshold condition.
[0073] Finally, the problem type of the generated single-step question is adjusted accordingly using the entity weight result to determine the initial single-step question set based on the adjusted problem type. That is, according to the entity weight result, a probability adjustment is made to different types of single-step questions that may be generated. In this way, the generated initial single-step question set can be more targeted around those entities with higher uncertainty and importance, thereby better improving the model's reasoning ability and understanding ability of key information.
[0074] Step S12: Optimize the initial single-step question set based on the initial single-step question set, the first text description, and the second text description that has no preset corresponding relationship with the visual sample to obtain a target single-step question set.
[0075] In this embodiment, first, positive samples are constructed using the initial single-step question set and the first text description corresponding to the visual sample. It should be noted that the initial single-step question set is generated around the visual sample and the first text description, and there is an inherent logical connection and semantic association between them. Therefore, combining the initial single-step question with the first text description forms positive samples with correct semantic correspondence. At the same time, negative samples are constructed using the initial single-step question set and the second text description that has no preset corresponding relationship with the visual sample. The second text description has no direct corresponding relationship with the visual sample. Combining the initial single-step question with the second text description forms negative samples with semantic mismatches. Suppose the second text description is "There is a book on the table", then "How many kinds of fruits are there in the basket" and "There is a book on the table" form a negative sample.
[0076] Specifically, in the positive samples, the first semantic similarity between the initial single-step problem set and the first text description is determined based on the embedding vectors of the initial single-step problem set and the first text description. The embedding vector is to transform text information into a representation in the vector space. By calculating the similarity between the embedding vectors of the initial single-step problem and the first text description, the degree of their semantic proximity can be measured. The commonly used similarity calculation method is to calculate the cosine similarity, which represents the similarity by calculating the cosine value of the angle between two vectors. In the negative samples, the second semantic similarity between the initial single-step problem set and the second text description is also determined based on the embedding vectors of the initial single-step problem set and the second text description.
[0077] Furthermore, based on the first semantic similarity, the second semantic similarity, and a preset contrastive learning loss function, the target loss value is determined. The purpose of the contrastive learning loss function is to make the semantic similarity between positive samples as high as possible and the semantic similarity between negative samples as low as possible. Through this loss function, the initial vision-language model will learn to closely associate the initial single-step problem with the correct text description and distinguish it from irrelevant text descriptions at the same time. In addition, the target loss value reflects the quality of the current initial single-step problem set in semantic matching. The smaller the loss value, the higher the matching degree between the initial single-step problem set and the correct text description, and the better the discrimination from irrelevant text descriptions.
[0078] Next, in this embodiment, the initial single-step problem set is optimized using the target loss value to obtain the target single-step problem set. The optimization process usually adjusts the parameters of the model through the backpropagation algorithm to gradually reduce the target loss value. In each iteration, the model calculates the gradient based on the target loss value and updates the relevant parameters. As the iteration progresses, the semantic similarity between the initial single-step problem set and the first text description will continuously increase, and the semantic similarity between the initial single-step problem set and the second text description will continuously decrease. Finally, when the target loss value converges to a small value, the optimized single-step problem set, that is, the target single-step problem set, is obtained.
[0079] Step S13: Determine the target multi-step problem set using the target single-step problem set, a preset semantic expansion strategy, and a preset problem reasoning strategy, and determine the training sample set and the first fine-tuned model based on the target single-step problem set, the target multi-step problem set, a preset model reasoning strategy, and the initial vision-language model; The multi-step problem is multiple problems with a logical progressive relationship based on the single-step problem.
[0080] In this embodiment, first, the target single-step question set is processed using a preset semantic extension strategy to generate a more complex intermediate question structure. The preset semantic extension strategy aims to explore the potential semantic associations in the single-step questions and expand them into a more in-depth and broad question statement. This extension can be derived based on the information contained in the visual sample and the logic of the language expression.
[0081] Next, combined with the preset question reasoning strategy, the semantically expanded questions are further transformed into a target multi-step question set. The preset question reasoning strategy will construct a sequence of questions with progressive levels based on the nature and logical relationship of the questions. For example, it gradually transitions from simple factual questions to complex questions that require causal analysis and hypothesis reasoning. Through such a strategy, each multi-step question is based on a single-step question and is logically progressive, which increases the complexity of the problem and the challenge to the model's reasoning ability.
[0082] After determining the target multi-step question set, the initial visual language model is fine-tuned using the target single-step question set, the target multi-step question set, and the preset model reasoning strategy to obtain the first fine-tuned model. During the fine-tuning process, the model will learn and adjust parameters for the target single-step question set and the target multi-step question set based on the preset model reasoning strategy. For example, the preset model reasoning strategy may specify the reasoning path and calculation method that the model should use when dealing with different types of questions. For single-step questions, the model learns to directly extract key information from visual samples and text descriptions to answer; for multi-step questions, the model learns to deduce and draw conclusions step by step according to the logical progression of the questions. In this process, the parameters of the model will be continuously updated to better meet the needs of answering these questions.
[0083] At the same time, in the process of fine-tuning the initial visual language model, the training sample set is determined based on the target single-step problem set and the target multi-step problem set. These sample sets can provide learning materials for the model, so that the model can master the reasoning skills from simple to complex problems in continuous training, thereby improving its performance in visual language reasoning tasks, and finally obtaining the first fine-tuned model with improved reasoning ability.
[0084] Step S14: determining a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and using the hybrid masking strategy to fine-tune the first fine-tuned model to obtain a second fine-tuned model.
[0085] In this embodiment, first, the visual samples are divided into several regions, and the division method can be carried out according to specific visual features and task requirements. For example, regular grid division can be adopted to evenly divide the image into small blocks of the same size, or based on the semantic information of the image, an image segmentation algorithm can be used to divide the visual samples into regions with different semantics, such as dividing a landscape picture into regions such as sky, grassland, and trees. At the same time, the first text description is divided into several words. This step is relatively intuitive and usually performs word segmentation according to the grammar rules of natural language.
[0086] Next, the visual information importance is determined based on several regions and the first adaptive weight coefficient. Among them, the first adaptive weight coefficient is dynamically adjusted according to the characteristics of the visual samples and the situation of model training. For each divided visual region, its importance in the entire visual sample can be calculated. This can be achieved through various methods. For example, using the attention mechanism, the model will automatically learn the importance weights of each region, or it can be calculated based on the visual features of the region, such as color contrast and texture complexity.
[0087] Meanwhile, the text information importance is determined based on several words and the second adaptive weight coefficient. The second adaptive weight coefficient is also adaptively adjusted. For each word, its importance can be determined according to factors such as its semantic importance in the text description and its degree of association with the visual sample. For example, in the text describing a picture, the core object names and key action words often have a higher text information importance, while the importance of some function words and auxiliary words is relatively low.
[0088] Furthermore, a mixed masking strategy is determined based on the visual information importance, the preset visual masking probability, the text information importance, and the preset text masking probability. The preset visual masking probability and the preset text masking probability are used to control the proportion of masking. Specifically, for visual regions, according to their visual information importance and the preset visual masking probability, it is determined which regions need to be masked. Regions with lower importance may be more likely to be masked to prompt the model to learn to reason from a wider range of visual information. For text words, similarly, according to their text information importance and the preset text masking probability, it is determined which words need to be masked. Such a mixed masking strategy can consider the information of both the visual and text modalities at the same time, enabling the model to pay more attention to important information during the learning process and improving the generalization ability and reasoning ability of the model.
[0089] Finally, the first fine-tuned model is fine-tuned using a hybrid masking strategy to obtain a second fine-tuned model. During the fine-tuning process, the model masks the input visual samples and the first text description according to the hybrid masking strategy, and then tries to recover the masked part from the remaining information. By continuously adjusting the parameters of the model, the model can better understand and process visual and text information, thereby improving its performance in visual language reasoning tasks. After this round of fine-tuning, the model can better adapt to the complex situations in actual applications, further enhancing its reasoning ability and generalization ability.
[0090] Step S15: Determine the second fine-tuned model as the student model, and use a preset teacher model to distill the student model to obtain a distilled model. Determine the sample difficulty of the training sample set, and iteratively train the distilled model based on the training sample set and the sample difficulty to trigger model inference using the obtained target visual language model.
[0091] In this embodiment, the second fine-tuned model is determined as the student model, and a preset teacher model is used to perform a distillation operation on it to obtain a distilled model. It should be noted that the core idea of model distillation is to transfer the knowledge learned by the teacher model to the student model, so that the student model can achieve a performance similar to that of the teacher model while having a smaller scale and lower computational cost. The preset teacher model is usually a model that has been fully trained on a large amount of data and has high performance.
[0092] During the distillation process, the teacher model generates soft labels for each question in the training sample set. These soft labels contain the probability distribution information of the teacher model for each possible answer. The student model then tries to learn these soft labels and optimizes its own parameters by minimizing the difference from the output of the teacher model. Specifically, a loss function such as KL divergence can be used to measure the difference between the outputs of the student model and the teacher model, and the parameters of the student model are continuously adjusted to make the student model gradually approach the performance of the teacher model, thereby obtaining a distilled model.
[0093] After obtaining the distilled model, determine the second target confidence level of the distilled model for predicting each current question in the training sample set. The second target confidence level reflects the certainty of the distilled model about its own prediction results and can be calculated through the probability distribution output by the distilled model. For example, if the prediction result of the distilled model for a certain question is concentrated on a certain answer and the probability of this answer is very high, then it can be considered that the distilled model has a high confidence level in this prediction; conversely, if the probability distribution is relatively dispersed, then the confidence level is low.
[0094] Meanwhile, determine the target loss value of the distillation model for each current problem. The target loss value measures the difference between the prediction result of the distillation model and the true label. Common loss functions, such as the cross-entropy loss function, can be used for calculation. Then, based on the second target confidence and the target loss value, determine the sample difficulty of the training sample set. Generally speaking, when the prediction confidence of the distillation model for a certain problem is low and the target loss value is large, it indicates that this problem is relatively difficult for the distillation model, that is, the sample difficulty is high. On the contrary, when the prediction confidence is high and the target loss value is small, the sample difficulty is low.
[0095] Then, perform iterative training on the distillation model based on the training sample set and the sample difficulty. During the iterative training process, the training samples can be weighted according to the sample difficulty. For problems with higher sample difficulty, a higher weight can be given, enabling the model to pay more attention to these difficult samples during training, thereby improving the model's performance on complex problems. Moreover, in each round of iteration, the model will perform forward propagation based on the training sample set, calculate the prediction result and the target loss value, and then update the model's parameters through the backpropagation algorithm. As the iteration progresses, the performance of the model will continuously improve, and the target loss value will gradually decrease.
[0096] Finally, in this embodiment, obtain the target model parameters based on the iterative training. The target model parameters are the parameter combination that enables the model to achieve better performance on the training sample set after multiple rounds of iterative training. Use the target model parameters to adjust the current model parameters of the distillation model to obtain the target vision-language model. Further, trigger the preset model inference operation using the target vision-language model. In practical applications, the target vision-language model can receive new visual samples and text descriptions, and perform vision-language inference tasks, such as answering questions about the content of pictures, generating text descriptions related to pictures, etc., to provide accurate and efficient inference results for users.
[0097] As can be seen from the above, in this application, first, a target entity set is determined by combining a semantic guidance network, an initial vision-language model, visual samples, and corresponding first text descriptions. The target entities are classified and predicted to find entities with uncertainty reaching a preset threshold, thereby generating an initial single-step question set with a single logical relationship. Next, the initial single-step questions are optimized using the initial single-step question set, the first text description, and a second text description unrelated to the visual samples to obtain a target single-step question set. Subsequently, based on the target single-step question set, a preset semantic expansion and question reasoning strategy are used to obtain a target multi-step question set with a logical progression relationship. In combination with a preset model reasoning strategy, the initial vision-language model is fine-tuned to determine a training sample set and obtain a first fine-tuned model. Then, according to the visual samples, the first text description, and preset visual and text mask probabilities, a hybrid masking strategy is determined. This strategy is used to fine-tune the first fine-tuned model again to obtain a second fine-tuned model. Finally, the second fine-tuned model is used as the student model, and a distilled model is obtained by distilling with a preset teacher model. The sample difficulty of the training sample set is determined, and based on this, the distilled model is iteratively trained to trigger model reasoning using the obtained target vision-language model. In this way, this application can improve the model reasoning ability of the vision-language model and enhance the user experience to a certain extent.
[0098] Next, in combination with Figure 2 the schematic diagram shown below, the technical solution of the embodiment of the present application will be specifically described.
[0099] Specifically, first, the visual samples (denoted by V) and text descriptions (i.e., the first text descriptions, denoted by T) are parsed. The VLM (i.e., the vision-language model) is used to parse the objects, attributes, and relationships in the image, and combined with the text information, a cross-modal information representation is formed. Then, through the semantic guidance network, the initial entity set in the visual samples and text descriptions is extracted using the cross-modal representation. Let the initial entity set be , then the importance score of each entity can be expressed by the formula: . In the formula for the importance score, represents the importance degree of the entity predicted by the VLM in the visual information, while represents the importance degree of the entity in the text description. and are hyperparameters used to control the weights of visual and text information. After obtaining the importance score of each entity, entities in that are greater than this threshold are selected using the threshold as the core targets for question generation.
[0100] Next, the vision-language model classifies and predicts the selected entities. Let the confidence of the highest category be , then the uncertainty measure Among them, reflects the confidence of the model in the most likely class of the entity, representing the degree of uncertainty of the model.
[0101] Then, entities with higher uncertainty ( ) are selected , and questions are generated for these entities. Among them, is a preset threshold condition. The question design methods can include: object recognition questions, such as "What is this object?"; property recognition questions, such as "What color is this item?"; quantity reasoning questions, such as "How many similar items are here?"; relationship understanding questions, such as "What is this person doing?"
[0102] According to the combined weight of entity importance and uncertainty , the generation probability of different types of questions is balanced. The calculation formula for the combined weight is . Among them, controls the influence based on entity importance; controls the influence based on uncertainty.
[0103] Furthermore, to optimize the quality of question generation, through cross-modal contrastive learning, the model is made to learn to understand the same semantic concept in different contexts. Specifically, positive and negative samples are first constructed. Let be the question generated for , and the original text description is T. Then the positive sample is , and the negative sample is , where is irrelevant text. Then, the semantic similarity between the question and the text T is calculated. The calculation formula for the similarity is , where represents the extracted embedding vector. Then, a contrastive learning loss function (i.e., ContrastiveLoss) is used for training to optimize the contrastive loss. Among them, the calculation formula for the contrastive learning loss function is as follows:
[0104] ;
[0105] This contrastive learning loss function can ensure that is closer to the correct text T and farther from the irrelevant text , improve the logical consistency of the problem. After optimization through contrastive learning, the generated problem is the target single-step problem. For example, in a static object scenario, the visual sample is a picture showing a basket of fruits, including apples, bananas, and oranges. The text description is "There are various fruits in the basket, including apples, bananas, and oranges." Then the generated problem examples are: the confidence-driven problem (i.e., the initial single-step problem) is "Is this fruit an orange?" and "Are all the fruits in the basket round?" The problem after contrastive learning optimization (i.e., the target single-step problem) is "What is the difference in color between apples and bananas?" and "Is the orange bigger than the apple?" Additionally, in a dynamic object scenario, the visual sample is a boy kicking a football, and there is a dog running beside. The text description is "A boy is kicking a football in the park, and there is a dog running beside." Then the generated problem examples are: the confidence-driven problem is "What is this person doing?" and "Is the dog standing or running?" The problem after contrastive learning optimization is "What is the difference in the movement patterns between the boy and the dog?"
[0106] Then, based on the above-generated set of target single-step problems , construct a strategy through hierarchical reasoning to gradually expand the complexity and depth of the problem, so as to form a set of target multi-step problems with a progressive relationship . Among them, the multi-step problem is a multi-step problem generated by constructing a strategy through hierarchical reasoning from one or more single-step problems . Specifically, use the set of target single-step problems, the preset semantic expansion strategy, and the preset problem reasoning strategy to determine the set of target multi-step problems
[0107] Specifically, the preset semantic expansion strategy is to expand the problem scope by identifying the semantic boundaries of single-step problems, making it form a more challenging multi-step reasoning task. This strategy can be divided into information supplementation, context association, and complexity progression. Information supplementation is to add new reasoning elements based on the core information of single-step problems to make the problem more hierarchical. Context association is to make the problem progress around the existing context rather than directly combining multiple problems. Complexity progression is to ensure that the problem generation is a natural progressive process rather than arbitrary splicing. The formula for semantic expansion of single-step problems is , where is the new problem after semantic expansion of the single-step problem , S represents additional context information, and the context information can come from visual features, text descriptions, or external knowledge
[0108] The preset question reasoning strategy is to guide the gradual escalation of questions through multi-level reasoning to form multi-step questions with a logical progressive relationship. This strategy can be divided into hierarchical reasoning, context adjustment, and adaptive difficulty control. Hierarchical reasoning is to gradually increase new reasoning requirements on the basis of primary reasoning, such as causal analysis, reasoning chain extension, etc. Context adjustment is to ensure that the generated questions are coherent in context and do not produce abrupt breakpoints. Adaptive difficulty control is to dynamically adjust the difficulty of questions according to the model's answering ability. The formula for question reasoning for single-step questions is , where L represents the reasoning level (such as direct reasoning, indirect reasoning, hypothetical reasoning, etc.), is the finally generated multi-step question (i.e., the target multi-step question). In addition, the following examples illustrate the generated multi-step questions.
[0109] In a specific implementation manner, for example, the single-step question is "Is there a book on the table?", after semantic expansion, it becomes "How many books are there on the table?". In the process of question reasoning, the question for the first-step direct reasoning is "There are 3 books on the table.", the question for the second-step indirect reasoning is "What are the types of books on the table?", and the question for the third-step reasoning deepening is "If 2 more magazines are placed, how many are there in total?". Then the finally generated multi-step question is "How many books are there on the table? What are their respective types? If 2 more magazines are added, how many are there in total?". In another specific implementation manner, for example, the single-step question is "What will happen when the ball falls to the ground?", after semantic expansion, it becomes "How does the ball fall to the ground?". In the process of question reasoning, the question for the first-step direct reasoning is "The ball falls off the table.", the question for the second-step indirect reasoning is "When the ball falls on floors of different materials, how does it bounce?", and the question for the third-step reasoning deepening is "If the ball is made of rubber, what will be its bounce height?". Then the finally generated multi-step question is "What will happen after the ball falls off the table? If it falls on floors of different materials, how will it bounce? What if it is a rubber ball?".
[0110] Next, design a hierarchical progressive reasoning training system, or progressive thinking heuristic fine-tuning. Through "progressive cognitive guidance" and "multi-perspective thinking modeling" (i.e., the preset model reasoning strategy), help the model start from simple questions and gradually improve its reasoning ability so that it can handle more complex logical tasks. Specifically, "progressive cognitive guidance" gradually increases the question difficulty by setting training stages to ensure that the model's reasoning ability is improved from shallow to deep. Therefore, in the training stage the model needs to first pass through and then enter , , progressing step by step until having the ability of . Among them, is direct reasoning, that is, the model only needs to obtain the answer based on the input information without involving complex logical relationships. is multi-step reasoning, that is, the model needs to establish connections between multiple known conditions for reasoning and calculation. is hypothetical reasoning, that is, the model needs to consider different hypothetical conditions and conduct logical deductions based on them. is adversarial reasoning, that is, the model needs to face information loss, uncertainty, and interference factors and still be able to infer the optimal solution. In addition, "multi-perspective thinking modeling" is to train the model to understand problems from different perspectives and establish a more robust reasoning structure, including structured thinking paths. Specifically, for each problem, multiple possible reasoning paths are constructed to let the model learn different problem-solving methods. Causal relationship modeling enables the model to not only know "what" but also "why", strengthening the causal reasoning ability. Dynamic complexity adjustment enables the dynamic adjustment of the problem complexity as the model training progresses, so that the model continuously accepts challenges. And in the process of fine-tuning the initial vision-language model based on the target single-step problem set, target multi-step problem set, and preset model reasoning strategy, a training sample set is generated, which can be expressed by the formula , where C represents the causal reasoning relationship. In addition, a hierarchical mind map can be generated for each problem to help the model learn the reasoning path. The mind map structure will change continuously with the progress of model training, being relatively simple in the initial stage and more complex in the later stage. Therefore, after the progressive thinking heuristic fine-tuning is completed, the first fine-tuned model and the training sample set will be obtained. In addition, the following example is used to illustrate the process of fine-tuning the initial vision-language model based on the target single-step problem set, target multi-step problem set, and preset model reasoning strategy.
[0111] For example, if the visual sample, that is, the picture content is a basket containing 3 apples and 5 bananas, and the text description is "There are 3 apples and 5 bananas in the basket." At the stage, the question is "How many fruits are there in total in the basket?", and the thinking path is: Apples (3) + Bananas (5) = 8. At this time, the model prediction is 8, and this prediction result is correct; at the stage, the question is "The number of apples in the basket is 3, and the total number of apples and bananas is 8. How many bananas are there in the basket?", and the thinking path is to set the number of bananas as x. According to the question, 3 + x = 8, and the solution is x = 5. At this time, the model prediction is 5, and this prediction result is correct; at the At the stage, the question is "If 2 more oranges are added, how many fruits are there in the basket now?", and the thinking path is that it is known that apples (3) + bananas (5) = 8. Assuming an additional 2 oranges are added, calculate the new total: 8 + 2 = 10. At this time, the model prediction is 10, and this prediction result is correct; at the
[0112] Next, perform adaptive hybrid masking training on the first fine-tuned model obtained above, that is, fine-tune the model based on the constructed hybrid masking strategy to obtain the second fine-tuned model. Specifically, assume that the visual sample V consists of multiple regions and the text description T consists of multiple words . First, calculate the importance of each region and word. The importance of visual information is , and the importance of language information is , where is the adaptive weight coefficient. Also, assume the initial visual masking probability is , and the text masking probability is . Then the masking range is . Among them, are the dynamically adjusted thresholds respectively. It should be noted that in the initial stage of training, , strong masking to enhance the model's completion ability; in the middle stage of training, , appropriately reduce the masking to maintain the challenge of the model; in the later stage of training, , low masking to let the model adapt to full-data inference. Thus, the hybrid masking strategy is obtained, and the model is fine-tuned based on the hybrid masking strategy to obtain the second fine-tuned model. The following example illustrates the process of fine-tuning the model using the hybrid masking strategy.
[0113] For example, the current task is fruit basket reasoning. The visual sample input is a basket containing 3 apples and 5 bananas, and the text description is "There are 3 apples and 5 bananas in the basket." In the initial stage of training, the visual mask occludes some of the apples and bananas, and the model still needs to infer the total number of fruits. The text mask removes "3 apples", and the model needs to fill in the missing information based on the visual information. The question is "How many fruits are there in the basket in total?" Then the model's reasoning path is to see some fruits, identify the missing content, complete the reasoning, and calculate the total number 8. In the middle stage of training, the visual mask occludes a small part of the apples and randomly blurs the background at the same time. The text mask removes the word "bananas" to let the model learn multi-modal alignment reasoning. The question is "Given that there are 8 fruits in the basket in total, and 3 of them are apples, how many bananas are there?" Then the model's reasoning path is to combine the visual information, reverse-infer the missing number of bananas, and get the answer 5. In the later stage of training, the visual mask is a very small part of the mask, mainly used to fine-tune the model stability. The text mask is also a very small part of the mask to ensure that the model can reason normally on the complete data. The question is "If 2 more oranges are added, how many fruits are there in the basket now?" Then the model's reasoning path is to calculate the original quantity 8, add the oranges 2, and get the answer 10.
[0114] Finally, progressive enhancement training is carried out. Specifically, let the teacher model and the student model . This student model is the second fine-tuned model. Distillation training is carried out using the KL divergence loss, and the distillation loss formula is . In the formula, represents the self-knowledge distillation loss, represents the KL divergence, represents the result predicted by the model, represents 's probability distribution over , represents 's probability distribution over . After distillation, a distilled model is obtained. Then, calculate the difficulty score of the sample set, and the calculation formula is . In the formula, represents the confidence of the model in the question , represents the loss of the model on this question, and is the adaptive weight. After obtaining the difficulty score, determine the corresponding sample difficulty, and iteratively train the distilled model based on the training sample set and the sample difficulty. Specifically, when the sample difficulty is low, reduce the training frequency; when the sample difficulty is medium, train normally; when the sample difficulty is high, increase the training weight. After iteratively training the distilled model, the target vision-language model can be obtained to perform the preset model reasoning operation using the target vision-language model.
[0115] Correspondingly, referring to Figure 3 As shown, an embodiment of the present application provides a model inference device for a vision - language model, including:
[0116] A single - step question generation module 11, configured to determine a target entity set based on a semantic guidance network, an initial vision - language model, visual samples, and a first text description corresponding to the visual samples, and determine entities whose uncertainty meets a preset threshold condition from the target entity set by using the result of classification prediction of the target entity set, and generate an initial single - step question set based on the entities whose uncertainty meets the preset threshold condition; the single - step question is a single logical relationship question;
[0117] A single - step question optimization module 12, configured to optimize the initial single - step question set based on the initial single - step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual samples, so as to obtain a target single - step question set;
[0118] A first model fine - tuning module 13, configured to determine a target multi - step question set by using the target single - step question set, a preset semantic expansion strategy, and a preset question reasoning strategy, and determine a training sample set and a first fine - tuned model based on the target single - step question set, the target multi - step question set, a preset model reasoning strategy, and the initial vision - language model; the multi - step question is multiple questions with a logical progressive relationship based on the single - step question;
[0119] A second model fine - tuning module 14, configured to determine a hybrid masking strategy based on the visual samples, the first text description, a preset visual masking probability, and a preset text masking probability, and fine - tune the first fine - tuned model by using the hybrid masking strategy to obtain a second fine - tuned model;
[0120] A model inference module 15, configured to determine the second fine - tuned model as a student model, and distill the student model by using a preset teacher model to obtain a distilled model, determine the sample difficulty of the training sample set, and perform iterative training on the distilled model based on the training sample set and the sample difficulty, so as to trigger model inference by using the obtained target vision - language model.
[0121] As can be seen from the above, in the present application, first, by combining a semantic guidance network, an initial vision-language model, visual samples, and corresponding first text descriptions, a target entity set is determined. The target entities are classified and predicted to find entities with uncertainty reaching a preset threshold, so as to generate an initial single-step question set with a single logical relationship. Then, the initial single-step questions are optimized using the initial single-step question set, the first text description, and a second text description unrelated to the visual samples to obtain a target single-step question set. Subsequently, based on the target single-step question set, a preset semantic expansion and question reasoning strategy is used to obtain a target multi-step question set with a logical progressive relationship. In combination with a preset model reasoning strategy, the initial vision-language model is fine-tuned to determine a training sample set and obtain a first fine-tuned model. Then, according to the visual samples, the first text description, and preset visual and text masking probabilities, a hybrid masking strategy is determined. This strategy is used to fine-tune the first fine-tuned model again to obtain a second fine-tuned model. Finally, the second fine-tuned model is used as a student model, and a distillation model is obtained by distillation using a preset teacher model. The sample difficulty of the training sample set is determined, and based on this, the distillation model is iteratively trained to trigger model reasoning using the obtained target vision-language model. In this way, the present application can improve the model reasoning ability of the vision-language model and, to a certain extent, enhance the user experience.
[0122] In some specific embodiments, the single-step question generation module 11 specifically includes:
[0123] A cross-modal parsing unit, configured to use the initial vision-language model to parse the visual samples and the first text descriptions corresponding to the visual samples to obtain corresponding cross-modal representations;
[0124] An entity extraction unit, configured to use the cross-modal representations to extract an initial entity set from the visual samples and the first text descriptions through a semantic guidance network to complete the entity extraction operation;
[0125] A first importance determination unit, configured to determine a first importance degree and a second importance degree of the initial entity set; the first importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the visual information, and the second importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the text description;
[0126] An importance score determination unit, configured to determine the importance scores of the initial entities in the initial entity set based on the first importance degree, the second importance degree, a preset visual weight control parameter, and a preset text weight control parameter;
[0127] An entity screening unit, configured to determine, in the initial entity set, the initial entities with importance scores greater than a preset importance threshold as target entities, so as to construct a target entity set based on each of the target entities.
[0128] In some specific embodiments, the single-step question generation module 11 specifically includes:
[0129] A classification prediction unit, configured to perform classification prediction on the target entity set and obtain a classification prediction result;
[0130] A confidence screening unit, configured to determine corresponding category information from the classification prediction result and determine a first target confidence level in the category information where the category probability value is greater than a preset probability threshold;
[0131] An uncertainty determination unit, configured to determine the uncertainty result of each target entity in the target entity set based on a preset constant value and the first target confidence level, and use the uncertainty result to determine entities in the target entity set whose uncertainty meets a preset threshold condition;
[0132] An index determination unit, configured to determine the target importance score and target uncertainty of the entities whose uncertainty meets the preset threshold condition;
[0133] A weight determination unit, configured to determine an entity weight result based on the target importance score, the target uncertainty, a preset importance control coefficient, and a preset uncertainty control coefficient;
[0134] A question determination unit, configured to correspondingly adjust the question type of the generated single-step questions by using the entity weight result, so as to determine an initial single-step question set based on the adjusted question type.
[0135] In some specific embodiments, the single-step question optimization module 12 specifically includes:
[0136] A negative sample construction unit, configured to construct positive samples by using the initial single-step question set and the first text description corresponding to the visual sample, and construct negative samples by using the initial single-step question set and a second text description that has no preset corresponding relationship with the visual sample;
[0137] A first similarity determination unit, configured to determine a first semantic similarity between the initial single-step question set and the first text description in the positive samples based on the embedding vectors of the initial single-step question set and the first text description;
[0138] A second similarity determination unit, configured to determine a second semantic similarity between the initial single-step question set and the second text description in the negative samples based on the embedding vectors of the initial single-step question set and the second text description;
[0139] A first loss value determination unit, configured to determine a target loss value based on the first semantic similarity, the second semantic similarity, and a preset contrastive learning loss function;
[0140] A problem optimization unit for optimizing the initial single-step problem set by using the target loss value to obtain a target single-step problem set.
[0141] In some specific embodiments, the first model fine-tuning module 13 specifically includes:
[0142] A first model fine-tuning unit for fine-tuning the initial vision-language model by using the target single-step problem set, the target multi-step problem set, and a preset model inference strategy to obtain a first fine-tuned model;
[0143] A sample set determination unit for determining a training sample set based on the target single-step problem set and the target multi-step problem set during the process of fine-tuning the initial vision-language model.
[0144] In some specific embodiments, the second model fine-tuning module 14 specifically includes:
[0145] An information partitioning unit for partitioning the visual sample into several regions and partitioning the first text description into several words;
[0146] A second importance determination unit for determining the visual information importance based on the several regions and a first adaptive weight coefficient, and determining the text information importance based on the several words and a second adaptive weight coefficient;
[0147] A strategy determination unit for determining a mixed masking strategy based on the visual information importance, a preset visual masking probability, the text information importance, and a preset text masking probability;
[0148] A second model fine-tuning unit for fine-tuning the first fine-tuned model by using the mixed masking strategy to obtain a second fine-tuned model.
[0149] In some specific embodiments, the model inference module 15 specifically includes:
[0150] A second loss value determination unit for determining a second target confidence level of the distilled model for predicting each current problem in the training sample set, and determining a target loss value of the distilled model on each current problem;
[0151] A model iteration unit for determining the sample difficulty of the training sample set based on the second target confidence level and the target loss value, and iteratively training the distilled model based on the training sample set and the sample difficulty to obtain target model parameters;
[0152] A model inference unit, configured to adjust the current model parameters of the distillation model by using the target model parameters, and obtain a target vision-language model, so as to trigger a preset model inference operation by using the target vision-language model.
[0153] Further, an embodiment of the present application also discloses an electronic device. Figure 4 FIG. 20 is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be construed as any limitation on the scope of use of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the model inference method for a vision-language model disclosed in any of the foregoing embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0154] In this embodiment, the power supply 23 is used to provide operating voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and specific limitations are not imposed thereon here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitations are imposed here.
[0155] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method may be temporary storage or permanent storage.
[0156] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it may be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the model inference method for a vision-language model executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program that can be used to complete other specific tasks.
[0157] Further, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the model inference method for a vision-language model disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details are not described herein again.
[0158] In this specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0159] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0160] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0161] Finally, it should also be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
[0162] The above has introduced the technical solution provided by this application in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A model inference method for a vision-language model, characterized in that Including: Determine a target entity set based on a semantic guidance network, an initial vision-language model, visual samples, and a first text description corresponding to the visual samples, and determine entities in the target entity set that meet the preset threshold condition of uncertainty using the result of classification prediction of the target entity set. Generate an initial single-step question set based on the entities that meet the preset threshold condition of uncertainty; The single-step question is a single logical relationship question; Optimize the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual samples to obtain a target single-step question set; Determine a target multi-step question set using the target single-step question set, a preset semantic expansion strategy, and a preset question reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step question set, the target multi-step question set, a preset model reasoning strategy, and the initial vision-language model; The multi-step question is multiple questions with a logical progression relationship based on the single-step question; Determine a hybrid masking strategy based on the visual samples, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tune the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model; Determine the second fine-tuned model as the student model, and distill the student model using a preset teacher model to obtain a distilled model. Determine the sample difficulty of the training sample set, and iteratively train the distilled model based on the training sample set and the sample difficulty to trigger model reasoning using the obtained target vision-language model; Among them, inputting the visual samples and the first text description corresponding to the visual samples into the initial vision-language model, and determining the target entity set based on the cross-modal representation output by the initial vision-language model through the semantic guidance network includes: parsing the visual samples and the first text description corresponding to the visual samples using the initial vision-language model to obtain the corresponding cross-modal representation; extracting the initial entity set in the visual samples and the first text description using the cross-modal representation through the semantic guidance network to complete the entity extraction operation; determining the first importance degree and the second importance degree of the initial entity set; The first importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the visual information, and the second importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the text description; determining the importance score of each initial entity in the initial entity set based on the first importance degree, the second importance degree, a preset visual weight control parameter, and a preset text weight control parameter; In the initial entity set, determine the initial entities with the importance score greater than the preset importance threshold as the target entities, and construct a target entity set based on each of the target entities; Determining a training sample set and a first fine-tuned model based on the target single-step question set, the target multi-step question set, a preset model inference strategy, and the initial vision-language model includes: fine-tuning the initial vision-language model using the target single-step question set, the target multi-step question set, and the preset model inference strategy to obtain a first fine-tuned model; during the process of fine-tuning the initial vision-language model, determining the training sample set based on the target single-step question set and the target multi-step question set.
2. The model inference method for a vision-language model according to claim 1, wherein The process of determining entities with uncertainties meeting a preset threshold condition from the target entity set using the classification prediction results of the target entity set and generating an initial single-step question set based on the entities with uncertainties meeting the preset threshold condition includes: Performing classification prediction on the target entity set and obtaining classification prediction results; Determining corresponding category information from the classification prediction results and determining a first target confidence level with a category probability value greater than a preset probability threshold in the category information; Determining uncertainty results of each target entity in the target entity set based on a preset constant value and the first target confidence level, and using the uncertainty results to determine entities with uncertainties meeting the preset threshold condition from the target entity set; Determining the target importance score and target uncertainty of the entities with uncertainties meeting the preset threshold condition; Determining an entity weight result based on the target importance score, the target uncertainty, a preset importance control coefficient, and a preset uncertainty control coefficient; Using the entity weight result to correspondingly adjust the question type of the generated single-step question, and determining the initial single-step question set based on the adjusted question type.
3. The model inference method for a vision-language model according to claim 1, wherein Optimizing the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample to obtain a target single-step question set includes: Constructing a positive sample using the initial single-step question set and the first text description corresponding to the visual sample, and constructing a negative sample using the initial single-step question set and the second text description that has no preset corresponding relationship with the visual sample; In the positive sample, determining a first semantic similarity between the initial single-step question set and the first text description based on the embedding vectors of the initial single-step question set and the first text description; In the negative sample, determining a second semantic similarity between the initial single-step question set and the second text description based on the embedding vectors of the initial single-step question set and the second text description; Determining a target loss value based on the first semantic similarity, the second semantic similarity, and a preset contrastive learning loss function; Using the target loss value to optimize the initial single-step question set to obtain a target single-step question set.
4. The model inference method for a vision-language model according to claim 1, wherein Determining a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and using the hybrid masking strategy to fine-tune the first fine-tuned model to obtain a second fine-tuned model includes: Divide the visual sample into several regions and divide the first text description into several words; Determine the importance of visual information based on the several regions and the first adaptive weight coefficient, and determine the importance of text information based on the several words and the second adaptive weight coefficient; Determine a hybrid masking strategy based on the importance of visual information, a preset visual masking probability, the importance of text information, and a preset text masking probability; Fine-tune the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model.
5. The model inference method for a vision-language model according to any one of claims 1 to 4, characterized in that The determining the sample difficulty of the training sample set and iteratively training the distillation model based on the training sample set and the sample difficulty to trigger model inference using the obtained target vision-language model includes: Determine the second target confidence level of the distillation model's prediction of each current question in the training sample set, and determine the target loss value of the distillation model on each current question; Determine the sample difficulty of the training sample set based on the second target confidence level and the target loss value, and iteratively train the distillation model based on the training sample set and the sample difficulty to obtain target model parameters; Adjust the current model parameters of the distillation model using the target model parameters, and obtain a target vision-language model to trigger a preset model inference operation using the target vision-language model.
6. A model inference device for a vision-language model, characterized in that, Includes: A single-step question generation module, configured to determine a target entity set based on a semantic guidance network, an initial vision-language model, a visual sample, and a first text description corresponding to the visual sample, and determine entities whose uncertainty satisfies a preset threshold condition from the target entity set based on the result of classifying and predicting the target entity set, and generate an initial single-step question set based on the entities whose uncertainty satisfies the preset threshold condition; The single-step question is a single logical relationship question; A single-step question optimization module, configured to optimize the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that does not have a preset corresponding relationship with the visual sample to obtain a target single-step question set; A first model fine-tuning module, configured to determine a target multi-step question set using the target single-step question set, a preset semantic expansion strategy, and a preset question reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step question set, the target multi-step question set, a preset model reasoning strategy, and the initial vision-language model; The multi-step question is multiple questions with a logical progressive relationship based on the single-step question; A second model fine-tuning module, configured to determine a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tune the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model; A model inference module, configured to determine the second fine-tuned model as a student model, and distill the student model using a preset teacher model to obtain a distilled model, determine the sample difficulty of the training sample set, and iteratively train the distilled model based on the training sample set and the sample difficulty, so as to trigger model inference using the obtained target vision-language model; Wherein, the single-step question generation module specifically includes: A cross-modal parsing unit, configured to parse a visual sample and a first text description corresponding to the visual sample using an initial vision-language model to obtain corresponding cross-modal representations; An entity extraction unit, configured to extract an initial entity set in the visual sample and the first text description using the cross-modal representation through a semantic guidance network to complete the entity extraction operation; A first importance determination unit, configured to determine a first importance degree and a second importance degree of the initial entity set; the first importance degree is the importance degree of the initial entity predicted by the initial vision-language model in visual information, and the second importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the text description; An importance score determination unit, configured to determine the importance scores of the initial entities in the initial entity set based on the first importance degree, the second importance degree, a preset visual weight control parameter, and a preset text weight control parameter; An entity screening unit, configured to determine, in the initial entity set, the initial entities with importance scores greater than a preset importance threshold as target entities, so as to construct a target entity set based on each of the target entities; Wherein, the first model fine-tuning module specifically includes: A first model fine-tuning unit, configured to fine-tune the initial vision-language model using the target single-step question set, the target multi-step question set, and a preset model inference strategy to obtain a first fine-tuned model; A sample set determination unit, configured to determine a training sample set based on the target single-step question set and the target multi-step question set during the process of fine-tuning the initial vision-language model.
7. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to execute the computer program to implement the model inference method for a vision-language model according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, For storing a computer program; wherein, when the computer program is executed by the processor, the model inference method for a vision-language model according to any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Question and answer method and device based on multi-modal information and application of question and answer method and device
CN117828142A
Video language task execution method and device, video language task model training method and device, equipment and medium
CN117876940A