Model reasoning method and device for visual language model, equipment and medium

By introducing semantic guidance networks and model distillation technology into visual language models, the problem of insufficient understanding and processing capabilities of visual language models in complex tasks is solved, and more efficient visual and text information fusion and reasoning capabilities are achieved.

CN120046743AActive Publication Date: 2025-05-27SHANDONG INSPUR SCI RES INST CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510535090.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

When handling complex visual and text information fusion tasks, existing visual language models cannot accurately and in-depth understanding and processing, resulting in poor accuracy and completeness of output results, and automated data construction faces the problem of modal inconsistency.

Method used

By determining the target entity set based on the semantic guidance network and the initial visual language model, the initial single-step problem set is generated, and the model's inference ability is improved through optimization and fine-tuning processes, and finally obtaining the target visual language model through model distillation and iterative training.

Benefits of technology

It improves the model reasoning ability of the visual language model, enhances the understanding and processing ability of visual and text information, and improves the accuracy and completeness of the output results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046743A_ABST
    Figure CN120046743A_ABST
Patent Text Reader

Abstract

The invention discloses a visual language model-oriented model reasoning method and device, equipment and a medium, and relates to the field of model reasoning, and the method comprises the steps: determining an initial single-step problem set based on a semantic guidance network, an initial visual language model, a visual sample and text description, and optimizing the initial single-step problem set to obtain a target single-step problem set; determining a target multi-step question set by utilizing the target single-step question set, a preset semantic extension strategy and a preset question reasoning strategy; determining a training sample set and a first fine-tuned model based on a target single-step question set, a target multi-step question set and the initial visual language model; performing fine tuning on the first fine-tuned model by using a mixed mask strategy to obtain a second fine-tuned model; and distilling the second fine-tuned model, training the obtained distillation model by using the training sample set, and triggering model reasoning by using the obtained target visual language model. Therefore, the model reasoning capability of the visual language model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model reasoning, and in particular to a model reasoning method, device, equipment and medium for a visual language model. Background Art

[0002] In the field of visual language models (VLMs), with the continuous development of science and technology, their applications in joint visual and text information processing tasks have become more and more extensive, such as image caption generation and visual question answering, showing great potential.

[0003] However, current VLMs still face many severe challenges in practical applications. The existing way of constructing instruction datasets is difficult to meet large-scale needs. Methods based on self-supervised learning still have obvious limitations in multimodal semantic understanding and cross-modal collaboration. This makes it impossible for the model to accurately and deeply understand and process complex visual and text information fusion tasks, resulting in poor accuracy and completeness of the output results. In addition, attempts to automate data construction are plagued by modal inconsistency. This inconsistency makes it difficult for the model to handle fine-grained visual language understanding tasks, and it is impossible to make detailed and accurate associations and interpretations of visual and text information, which in turn affects the performance of the model in complex tasks.

[0004] Therefore, how to improve the model reasoning ability of visual language models is a technical problem that needs to be solved urgently. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a model reasoning method, device, equipment and medium for a visual language model, which can improve the model reasoning ability of the visual language model. The specific scheme is as follows: In a first aspect, the present application provides a model reasoning method for a visual language model, comprising: A target entity set is determined based on a semantic guidance network, an initial visual language model, a visual sample, and a first text description corresponding to the visual sample, and entities whose uncertainties meet a preset threshold condition are determined from the target entity set using a result of classification prediction of the target entity set, and an initial single-step question set is generated based on the entities whose uncertainties meet the preset threshold condition; a single-step question is a single logical relationship question; Optimizing the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample to obtain a target single-step question set; Determine a target multi-step problem set by using the target single-step problem set, a preset semantic expansion strategy, and a preset problem reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step problem set, the target multi-step problem set, a preset model reasoning strategy, and the initial vision-language model; a multi-step problem is a plurality of problems with a logical progressive relationship based on a single-step problem. Determine a hybrid masking strategy based on the visual samples, the first text description, a preset visual masking probability, and a preset text masking probability, and use the hybrid masking strategy to fine-tune the first fine-tuned model to obtain a second fine-tuned model. Determine the second fine-tuned model as the student model, and use a preset teacher model to distill the student model to obtain a distilled model, determine the sample difficulty of the training sample set, and iteratively train the distilled model based on the training sample set and the sample difficulty to trigger model reasoning by using the obtained target vision-language model.

[0006] Optionally, the determining the target entity set based on the semantic guidance network, the initial vision-language model, the visual samples, and the first text description corresponding to the visual samples includes: Use the initial vision-language model to parse the visual samples and the first text description corresponding to the visual samples to obtain corresponding cross-modal representations. Extract an initial entity set in the visual samples and the first text description by using the cross-modal representations through the semantic guidance network to complete the entity extraction operation. Determine a first importance degree and a second importance degree of the initial entity set; the first importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the visual information, and the second importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the text description. Determine the importance score of each initial entity in the initial entity set based on the first importance degree, the second importance degree, a preset visual weight control parameter, and a preset text weight control parameter. In the initial entity set, determine the initial entities with the importance score greater than a preset importance threshold as target entities, and construct a target entity set based on each target entity.

[0007] Optionally, in the process of determining the entities whose uncertainty meets the preset threshold condition from the target entity set by using the classification prediction result of classifying the target entity set and generating an initial single-step problem set based on the entities whose uncertainty meets the preset threshold condition, it includes: Perform classification prediction on the target entity set and obtain a classification prediction result. Determine the corresponding category information from the classification prediction results, and determine the first target confidence level in the category information where the category probability value is greater than the preset probability threshold; Determine the uncertainty results of each target entity in the target entity set based on the preset constant value and the first target confidence level, and use the uncertainty results to determine the entities in the target entity set whose uncertainty meets the preset threshold condition; Determine the target importance score and target uncertainty of the entities whose uncertainty meets the preset threshold condition; Determine the entity weight result based on the target importance score, the target uncertainty, the preset importance control coefficient, and the preset uncertainty control coefficient; Use the entity weight result to correspondingly adjust the problem type of the generated single-step problem, so as to determine the initial single-step problem set based on the adjusted problem type.

[0008] Optionally, optimizing the initial single-step problem set based on the initial single-step problem set, the first text description, and the second text description that has no preset corresponding relationship with the visual sample to obtain the target single-step problem set, including: Construct positive samples using the initial single-step problem set and the first text description corresponding to the visual sample, and construct negative samples using the initial single-step problem set and the second text description that has no preset corresponding relationship with the visual sample; In the positive samples, determine the first semantic similarity between the initial single-step problem set and the first text description based on the embedding vectors of the initial single-step problem set and the first text description; In the negative samples, determine the second semantic similarity between the initial single-step problem set and the second text description based on the embedding vectors of the initial single-step problem set and the second text description; Determine the target loss value based on the first semantic similarity, the second semantic similarity, and the preset contrastive learning loss function; Optimize the initial single-step problem set using the target loss value to obtain the target single-step problem set.

[0009] Optionally, the method for determining the training sample set and the first fine-tuned model based on the target single-step problem set, the target multi-step problem set, the preset model inference strategy, and the initial vision-language model includes: Fine-tune the initial vision-language model using the target single-step problem set, the target multi-step problem set, and the preset model inference strategy to obtain the first fine-tuned model; During the process of fine-tuning the initial vision-language model, determine the training sample set based on the target single-step problem set and the target multi-step problem set.

[0010] Optionally, determining a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tuning the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model, includes: Dividing the visual sample into a plurality of regions, and dividing the first text description into a plurality of words; Determining the importance degree of visual information based on the plurality of regions and a first adaptive weight coefficient, and determining the importance degree of text information based on the plurality of words and a second adaptive weight coefficient; Determining a hybrid masking strategy based on the importance degree of visual information, the preset visual masking probability, the importance degree of text information, and the preset text masking probability; Fine-tuning the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model.

[0011] Optionally, determining the sample difficulty of the training sample set, and iteratively training the distillation model based on the training sample set and the sample difficulty, so as to trigger model inference using the obtained target vision-language model, includes: Determining a second target confidence level of the distillation model for predicting each current problem in the training sample set, and determining a target loss value of the distillation model on each current problem; Determining the sample difficulty of the training sample set based on the second target confidence level and the target loss value, and iteratively training the distillation model based on the training sample set and the sample difficulty to obtain target model parameters; Adjusting the current model parameters of the distillation model using the target model parameters, and obtaining a target vision-language model, so as to trigger a preset model inference operation using the target vision-language model.

[0012] In a second aspect, the present application provides a model inference device for a vision-language model, including: A single-step problem generation module, configured to determine a target entity set based on a semantic guidance network, an initial vision-language model, a visual sample, and a first text description corresponding to the visual sample, and determine an entity whose uncertainty satisfies a preset threshold condition from the target entity set based on a classification prediction result of the target entity set, and generate an initial single-step problem set based on the entity whose uncertainty satisfies the preset threshold condition; a single-step problem is a single logical relationship problem; A single-step problem optimization module, configured to optimize the initial single-step problem set based on the initial single-step problem set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample, so as to obtain a target single-step problem set; The first model fine-tuning module is used to determine a target multi-step problem set by using the target single-step problem set, a preset semantic extension strategy, and a preset problem reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step problem set, the target multi-step problem set, a preset model reasoning strategy, and the initial vision-language model; a multi-step problem is multiple problems with a logical progressive relationship based on a single-step problem. The second model fine-tuning module is used to determine a mixed masking strategy based on the visual samples, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tune the first fine-tuned model by using the mixed masking strategy to obtain a second fine-tuned model. The model inference module is used to determine the second fine-tuned model as the student model, and distill the student model by using a preset teacher model to obtain a distilled model, determine the sample difficulty of the training sample set, and perform iterative training on the distilled model based on the training sample set and the sample difficulty to trigger model inference by using the obtained target vision-language model.

[0013] In a third aspect, the present application provides an electronic device, including: A memory for storing a computer program; A processor for executing the computer program to implement the foregoing model inference method for a vision-language model.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing model inference method for a vision-language model is implemented.

[0015] In this application, a target entity set is determined based on a semantic guidance network, an initial vision-language model, visual samples, and a first text description corresponding to the visual samples. Entities that meet the preset threshold condition for uncertainty are determined from the target entity set by using the result of classifying and predicting the target entity set. An initial single-step question set is generated based on the entities that meet the preset threshold condition for uncertainty; a single-step question is a single logical relationship question. The initial single-step question set is optimized based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual samples to obtain a target single-step question set. A target multi-step question set is determined by using the target single-step question set, a preset semantic expansion strategy, and a preset question reasoning strategy. A training sample set and a first fine-tuned model are determined based on the target single-step question set, the target multi-step question set, a preset model reasoning strategy, and the initial vision-language model; a multi-step question is multiple questions with a logical progressive relationship based on the single-step question. A hybrid masking strategy is determined based on the visual samples, the first text description, a preset visual masking probability, and a preset text masking probability. The first fine-tuned model is fine-tuned by using the hybrid masking strategy to obtain a second fine-tuned model. The second fine-tuned model is determined as the student model, and the student model is distilled by using a preset teacher model to obtain a distilled model. The sample difficulty of the training sample set is determined, and the distilled model is iteratively trained based on the training sample set and the sample difficulty to trigger model reasoning by using the obtained target vision-language model. As can be seen from the above, in this application, first, a target entity set is determined by combining a semantic guidance network, an initial vision-language model, visual samples, and the corresponding first text description. The target entities are classified and predicted to find entities with uncertainties reaching the preset threshold, and an initial single-step question set with a single logical relationship is generated based on these entities. Next, the initial single-step questions are optimized by using the initial single-step question set, the first text description, and the second text description that has no relation to the visual samples to obtain a target single-step question set. Subsequently, based on the target single-step question set, a target multi-step question set with a logical progressive relationship is obtained by using the preset semantic expansion and question reasoning strategies. In combination with the preset model reasoning strategy, the initial vision-language model is fine-tuned to determine the training sample set and obtain a first fine-tuned model. Then, according to the visual samples, the first text description, and the preset visual and text masking probabilities, a hybrid masking strategy is determined. The first fine-tuned model is fine-tuned again by using this strategy to obtain a second fine-tuned model. Finally, the second fine-tuned model is used as the student model, and a distilled model is obtained by distilling with the preset teacher model. The sample difficulty of the training sample set is determined, and the distilled model is iteratively trained based on this to trigger model reasoning by using the obtained target vision-language model. In this way, this application can improve the model reasoning ability of the vision-language model and enhance the user experience to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0017] Figure 1 It is a flowchart of a model inference method for a vision-language model disclosed in this application; Figure 2 It is a flowchart of a specific model inference method for a vision-language model disclosed in this application; Figure 3 It is a schematic structural diagram of a model inference device for a vision-language model disclosed in this application; Figure 4 It is a structural diagram of an electronic device disclosed in this application. Detailed implementation manners

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0019] Currently, VLMs face many severe challenges in practical applications. The existing construction methods of instruction datasets are difficult to meet large-scale requirements. The self-supervised learning-based methods still have obvious limitations in multi-modal semantic understanding and cross-modal collaboration. This makes the model unable to accurately and deeply understand and process complex visual and text information fusion tasks, resulting in poor accuracy and integrity of the output results. Moreover, the attempts of automated data construction are troubled by modal inconsistencies. Such inconsistencies make it difficult for the model to handle fine-grained vision-language understanding tasks, unable to make detailed and accurate associations and interpretations of visual and text information, and thus affecting the performance of the model in complex tasks. For this reason, this application provides a model inference method, device, device and medium for vision-language models, which can improve the model inference ability of vision-language models.

[0020] See Figure 1 As shown, the embodiments of the present invention disclose a model inference method for a vision-language model, including:

[0021] Step S11: Determine a target entity set based on a semantic guidance network, an initial vision - language model, visual samples, and first - text descriptions corresponding to the visual samples, and determine entities that meet the preset threshold condition of uncertainty from the target entity set using the result of classification prediction for the target entity set. Generate an initial single - step question set based on the entities that meet the preset threshold condition of uncertainty; a single - step question is a single - logical - relationship question.

[0022] In this embodiment, first, the initial vision - language model is used to parse the visual samples and the first - text descriptions corresponding to the visual samples. Among them, the initial vision - language model has the ability to process visual and text information. Through its internal algorithms and structures, it can identify and understand visual elements such as objects, attributes, and spatial relationships in visual samples, and at the same time, it can also extract and analyze semantic information in the first - text descriptions. During the parsing process, the initial vision - language model fuses visual information and text information to obtain corresponding cross - modal representations. This cross - modal representation integrates information from both the visual and text dimensions, providing certain basic data for subsequent entity extraction.

[0023] Next, the semantic guidance network (i.e., SGN, Semantics - Guided Neural Network) uses the above - obtained cross - modal representation to extract an initial entity set from the visual samples and text descriptions to complete the entity - extraction operation. The semantic guidance network can extract entities from cross - modal information. It can identify entities in visual samples and text descriptions based on the feature information in the cross - modal representation. These entities may be objects, people, scenes, etc., which are the core elements constituting the visual scene and text description. During the extraction process, the semantic guidance network deeply analyzes the cross - modal representation, and uses specific algorithms and rules to screen out entities that meet certain criteria, thereby forming an initial entity set.

[0024] Then, determine the first importance degree and the second importance degree of the initial entity set. It should be noted that the first importance degree is the importance degree of the initial entity predicted by the initial vision - language model in visual information. This is obtained by analyzing the visual samples through the initial vision - language model and evaluating factors such as the prominence of each initial entity in the visual scene and the degree of association with other visual elements. The second importance degree is the importance degree of the initial entity predicted by the initial vision - language model in the text description. The initial vision - language model analyzes factors such as the mention frequency of each entity in the text and the detail level of the description to determine its importance in the text description.

[0025] Furthermore, the importance scores of the initial entities in the initial entity set are determined based on the first importance level, the second importance level, a preset visual weight control parameter, and a preset text weight control parameter. Specifically, to comprehensively consider the influence of both visual and text aspects on entity importance, a preset visual weight control parameter and a preset text weight control parameter are introduced. These two parameters are used to adjust the relative weights of visual information and text information when determining the importance scores of entities. Through a specific calculation formula, the first importance level is multiplied by the preset visual weight control parameter, and the second importance level is multiplied by the preset text weight control parameter, and then the two results are added together to obtain the importance score of each initial entity. This approach can flexibly balance the contributions of visual and text information to entity importance according to actual needs.

[0026] In the initial entity set, the initial entities with importance scores greater than a preset importance threshold are determined as target entities, and a target entity set is constructed based on the target entities. That is to say, only those initial entities whose importance scores exceed the preset importance threshold will be recognized as target entities, and these target entities constitute the target entity set. The target entity set contains the most critical information elements in the visual samples and text descriptions and is the core basis for subsequent reasoning and question generation.

[0027] Next, a classification prediction is performed on the target entity set, and a classification prediction result is obtained. The initial vision-language model classifies each target entity in the target entity set to determine its category. For example, a target entity may be determined to be in the "fruit" category or the "animal" category, etc. Moreover, the classification prediction result includes the category information of each target entity and the corresponding category probability value.

[0028] In a specific implementation manner, the corresponding category information is determined from the classification prediction result, and a first target confidence level with a category probability value greater than a preset probability threshold is determined in the category information. That is, only when the category probability value of a certain target entity is greater than the preset probability threshold is the category considered to be a reliable classification of the target entity, and the corresponding confidence level is the first target confidence level.

[0029] Then, in this embodiment, the uncertainty results of the target entities in the target entity set are determined based on a preset constant value and the first target confidence level, and entities whose uncertainty meets the preset threshold condition are determined from the target entity set. Generally, the uncertainty result can be obtained by subtracting the first target confidence level from the preset constant value. Then, the uncertainty result of each target entity is compared with the preset threshold condition to screen out the entities whose uncertainty meets the preset threshold condition. Additionally, these entities have relatively high uncertainty in the classification prediction, which may be due to reasons such as unclear visual features or ambiguous text descriptions.

[0030] Further, in this embodiment, the target importance score and target uncertainty of the entity whose uncertainty meets the preset threshold condition are determined. For these selected entities, their target importance score and target uncertainty are determined again. The determination method of the target importance score is similar to the method of determining the importance score of the initial entity before, but it is calculated for these specific entities. The target uncertainty is the uncertainty result obtained from the above calculation.

[0031] After obtaining the target importance score and target uncertainty, the entity weight result is determined based on the target importance score, target uncertainty, preset importance control coefficient, and preset uncertainty control coefficient. The preset importance control coefficient and preset uncertainty control coefficient are used to adjust the relative importance of the target importance score and target uncertainty when determining the entity weight result. Through a specific calculation formula, the target importance score is multiplied by the preset importance control coefficient, the target uncertainty is multiplied by the preset uncertainty control coefficient, and then these two results are added together to obtain the entity weight result of each entity whose uncertainty meets the preset threshold condition.

[0032] Finally, the problem type of the generated single-step question is adjusted accordingly using the entity weight result, so as to determine the initial single-step question set based on the adjusted problem type. That is, according to the entity weight result, the probability of different types of single-step questions that may be generated is adjusted. In this way, the generated initial single-step question set can more specifically focus on those entities with higher uncertainty and importance, thereby better improving the reasoning ability of the model and the understanding ability of key information.

[0033] Step S12: Optimize the initial single-step question set based on the initial single-step question set, the first text description, and the second text description that has no preset corresponding relationship with the visual sample to obtain the target single-step question set.

[0034] In this embodiment, first, positive samples are constructed using the initial single-step question set and the first text description corresponding to the visual sample. It should be noted that the initial single-step question set is generated around the visual sample and the first text description, and there is an internal logical connection and semantic association between them. Therefore, combining the initial single-step question with the first text description forms positive samples with correct semantic corresponding relationships. At the same time, negative samples are constructed using the initial single-step question set and the second text description that has no preset corresponding relationship with the visual sample. The second text description has no direct corresponding relationship with the visual sample. Combining the initial single-step question with the second text description forms negative samples with semantic mismatches. Assuming the second text description is "There is a book on the table", then "How many kinds of fruits are there in the basket" and "There is a book on the table" form a negative sample.

[0035] Specifically, in the positive samples, the first semantic similarity between the initial single-step problem set and the first text description is determined based on the embedding vectors of the initial single-step problem set and the first text description. The embedding vector is to transform text information into a representation in the vector space. By calculating the similarity between the embedding vectors of the initial single-step problem and the first text description, the semantic proximity between them can be measured. The commonly used similarity calculation method is to calculate the cosine similarity, which represents the similarity by calculating the cosine value of the angle between two vectors. In the negative samples, the second semantic similarity between the initial single-step problem set and the second text description is also determined based on the embedding vectors of the initial single-step problem set and the second text description.

[0036] Furthermore, the target loss value is determined based on the first semantic similarity, the second semantic similarity, and a preset contrastive learning loss function. The purpose of the contrastive learning loss function is to make the semantic similarity between positive samples as high as possible and the semantic similarity between negative samples as low as possible. Through this loss function, the initial vision-language model will learn to closely associate the initial single-step problem with the correct text description and distinguish it from irrelevant text descriptions. In addition, the target loss value reflects the quality of the current initial single-step problem set in semantic matching. The smaller the loss value, the higher the matching degree between the initial single-step problem set and the correct text description, and the better the discrimination from irrelevant text descriptions.

[0037] Next, in this embodiment, the initial single-step problem set is optimized using the target loss value to obtain the target single-step problem set. The optimization process usually adjusts the parameters of the model through the backpropagation algorithm to gradually reduce the target loss value. In each iteration, the model calculates the gradient based on the target loss value and updates the relevant parameters. As the iteration progresses, the semantic similarity between the initial single-step problem set and the first text description will continuously increase, and the semantic similarity between the initial single-step problem set and the second text description will continuously decrease. Finally, when the target loss value converges to a small value, the optimized single-step problem set, that is, the target single-step problem set, is obtained.

[0038] Step S13: Use the target single-step problem set, a preset semantic expansion strategy, and a preset problem reasoning strategy to determine the target multi-step problem set, and determine the training sample set and the first fine-tuned model based on the target single-step problem set, the target multi-step problem set, a preset model reasoning strategy, and the initial vision-language model; the multi-step problem is multiple problems with a logical progressive relationship based on the single-step problem.

[0039] In this embodiment, first, the target single-step question set is processed using a preset semantic extension strategy to generate a more complex intermediate question structure. The preset semantic extension strategy aims to explore the potential semantic associations in the single-step questions and expand them into a more in-depth and broad question statement. This extension can be derived based on the information contained in the visual sample and the logic of the language expression.

[0040] Next, combined with the preset question reasoning strategy, the semantically expanded questions are further transformed into a target multi-step question set. The preset question reasoning strategy will construct a sequence of questions with progressive levels based on the nature and logical relationship of the questions. For example, it gradually transitions from simple factual questions to complex questions that require causal analysis and hypothesis reasoning. Through such a strategy, each multi-step question is based on a single-step question and is logically progressive, which increases the complexity of the problem and the challenge to the model's reasoning ability.

[0041] After determining the target multi-step question set, the initial visual language model is fine-tuned using the target single-step question set, the target multi-step question set, and the preset model reasoning strategy to obtain the first fine-tuned model. During the fine-tuning process, the model will learn and adjust parameters for the target single-step question set and the target multi-step question set based on the preset model reasoning strategy. For example, the preset model reasoning strategy may specify the reasoning path and calculation method that the model should use when dealing with different types of questions. For single-step questions, the model learns to directly extract key information from visual samples and text descriptions to answer; for multi-step questions, the model learns to deduce and draw conclusions step by step according to the logical progression of the questions. In this process, the parameters of the model will be continuously updated to better meet the needs of answering these questions.

[0042] At the same time, in the process of fine-tuning the initial visual language model, the training sample set is determined based on the target single-step problem set and the target multi-step problem set. These sample sets can provide learning materials for the model, so that the model can master the reasoning skills from simple to complex problems in continuous training, thereby improving its performance in visual language reasoning tasks, and finally obtaining the first fine-tuned model with improved reasoning ability.

[0043] Step S14: determining a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and using the hybrid masking strategy to fine-tune the first fine-tuned model to obtain a second fine-tuned model.

[0044] In this embodiment, first, the visual sample is divided into several regions, and the division method can be determined according to specific visual features and task requirements. For example, a regular grid division can be adopted to evenly divide the image into small blocks of the same size, or based on the semantic information of the image, an image segmentation algorithm can be used to divide the visual sample into regions with different semantics, such as dividing a landscape picture into regions such as sky, grassland, and trees. At the same time, the first text description is divided into several words. This step is relatively intuitive and usually performs word segmentation according to the grammar rules of natural language.

[0045] Next, the visual information importance is determined based on several regions and the first adaptive weight coefficient. Among them, the first adaptive weight coefficient is dynamically adjusted according to the characteristics of the visual sample and the situation of model training. For each divided visual region, its importance in the entire visual sample can be calculated. This can be achieved through various methods. For example, using an attention mechanism, the model will automatically learn the importance weights of each region, or it can be calculated based on the visual features of the region, such as color contrast, texture complexity, etc.

[0046] Meanwhile, the text information importance is determined based on several words and the second adaptive weight coefficient. The second adaptive weight coefficient is also adaptively adjusted. For each word, its importance can be determined according to factors such as its semantic importance in the text description and its degree of association with the visual sample. For example, in the text describing a picture, the core object names and key action words often have a higher text information importance, while the importance of some function words, auxiliary words, etc. is relatively low.

[0047] Furthermore, a mixed masking strategy is determined based on the visual information importance, the preset visual masking probability, the text information importance, and the preset text masking probability. The preset visual masking probability and the preset text masking probability are used to control the proportion of masking. Specifically, for visual regions, according to their visual information importance and the preset visual masking probability, it is determined which regions need to be masked. Regions with lower importance may be more likely to be masked to prompt the model to learn to reason from a wider range of visual information. For text words, similarly, according to their text information importance and the preset text masking probability, it is determined which words need to be masked. Such a mixed masking strategy can simultaneously consider the information of both the visual and text modalities, enabling the model to pay more attention to important information during the learning process and improving the generalization ability and reasoning ability of the model.

[0048] Finally, the first fine-tuned model is fine-tuned using the hybrid masking strategy to obtain the second fine-tuned model. During the fine-tuning process, the model masks the input visual samples and the first text description according to the hybrid masking strategy, and then tries to recover the masked parts from the remaining information. By continuously adjusting the model's parameters, the model can better understand and process visual and text information, thereby improving its performance in visual language reasoning tasks. After this round of fine-tuning, the model can better adapt to the complex situations in actual applications, further enhancing its reasoning ability and generalization ability.

[0049] Step S15: Determine the second fine-tuned model as the student model, and use a preset teacher model to distill the student model to obtain a distilled model. Determine the sample difficulty of the training sample set, and iteratively train the distilled model based on the training sample set and the sample difficulty to trigger model reasoning using the obtained target visual language model.

[0050] In this embodiment, the second fine-tuned model is determined as the student model, and a preset teacher model is used to perform distillation operation on it to obtain a distilled model. It should be noted that the core idea of model distillation is to transfer the knowledge learned by the teacher model to the student model, so that the student model can achieve performance similar to that of the teacher model while having a smaller scale and lower computational cost. The preset teacher model is usually a model that has been fully trained on a large amount of data and has high performance.

[0051] During the distillation process, the teacher model generates soft labels for each question in the training sample set. These soft labels contain the probability distribution information of the teacher model for each possible answer. The student model then tries to learn these soft labels and optimizes its own parameters by minimizing the difference from the output of the teacher model. Specifically, a loss function such as KL divergence can be used to measure the difference between the outputs of the student model and the teacher model, and the parameters of the student model are continuously adjusted to make the student model gradually approach the performance of the teacher model, thereby obtaining the distilled model.

[0052] After obtaining the distilled model, determine the second target confidence level of the distilled model for predicting each current question in the training sample set. The second target confidence level reflects the certainty of the distilled model about its own prediction results and can be calculated through the probability distribution output by the distilled model. For example, if the prediction result of the distilled model for a certain question is concentrated on a certain answer and the probability of this answer is very high, then it can be considered that the distilled model has a high confidence level in this prediction; on the contrary, if the probability distribution is relatively dispersed, then the confidence level is low.

[0053] Meanwhile, determine the target loss value of the distillation model for each current problem. The target loss value measures the difference between the prediction result of the distillation model and the true label. Common loss functions, such as the cross-entropy loss function, can be used for calculation. Then, based on the second target confidence and the target loss value, determine the sample difficulty of the training sample set. Generally speaking, when the prediction confidence of the distillation model for a certain problem is low and the target loss value is large, it means that this problem is relatively difficult for the distillation model, that is, the sample difficulty is high. On the contrary, when the prediction confidence is high and the target loss value is small, the sample difficulty is low.

[0054] Then, perform iterative training on the distillation model based on the training sample set and the sample difficulty. During the iterative training process, the training samples can be weighted according to the sample difficulty. For problems with higher sample difficulty, higher weights can be given, enabling the model to pay more attention to these difficult samples during training, thereby improving the model's performance on complex problems. Moreover, in each round of iteration, the model will perform forward propagation based on the training sample set, calculate the prediction result and the target loss value, and then update the model's parameters through the backpropagation algorithm. As the iteration progresses, the performance of the model will continuously improve, and the target loss value will gradually decrease.

[0055] Finally, in this embodiment, obtain the target model parameters based on the iterative training. The target model parameters are the parameter combinations that enable the model to achieve better performance on the training sample set after multiple rounds of iterative training. Use the target model parameters to adjust the current model parameters of the distillation model to obtain the target vision-language model. Further, trigger the preset model inference operation using the target vision-language model. In practical applications, the target vision-language model can receive new visual samples and text descriptions, and perform vision-language inference tasks, such as answering questions about the content of pictures, generating text descriptions related to pictures, etc., to provide accurate and efficient inference results for users.

[0056] As can be seen from the above, in the present application, first, by combining the semantic guidance network, the initial vision-language model, the visual samples, and the corresponding first text description, a target entity set is determined. The target entities are classified and predicted to find entities with uncertainty reaching a preset threshold, and an initial single-step question set with a single logical relationship is generated accordingly. Next, the initial single-step questions are optimized using the initial single-step question set, the first text description, and a second text description unrelated to the visual samples to obtain a target single-step question set. Subsequently, based on the target single-step question set, a preset semantic expansion and question reasoning strategy is used to obtain a target multi-step question set with a logical progressive relationship. In combination with a preset model reasoning strategy, the initial vision-language model is fine-tuned to determine a training sample set and obtain a first fine-tuned model. Then, according to the visual samples, the first text description, and the preset visual and text masking probabilities, a hybrid masking strategy is determined. This strategy is used to fine-tune the first fine-tuned model again to obtain a second fine-tuned model. Finally, the second fine-tuned model is used as the student model, and a distillation model is obtained by distillation using a preset teacher model. The difficulty of the samples in the training sample set is determined, and the distillation model is iteratively trained based on this to trigger model reasoning using the obtained target vision-language model. In this way, the present application can improve the model reasoning ability of the vision-language model and enhance the user experience to a certain extent.

[0057] Next, in combination with Figure 2 the schematic diagram shown below, the technical solution of the embodiment of the present application will be specifically described.

[0058] Specifically, first, the visual samples (denoted by V) and the text description (i.e., the first text description, denoted by T) are parsed. The VLM (i.e., the vision-language model) is used to parse the objects, attributes, and relationships in the image, and combined with the text information, a cross-modal information representation is formed. Then, through the semantic guidance network, the initial entity set in the visual samples and the text description is extracted using the cross-modal representation. Let the initial entity set be , then the importance score of each entity can be expressed by the formula: . In the formula for the importance score, represents the importance degree of the entity predicted by the VLM in the visual information, while represents the importance degree of the entity in the text description. and are hyperparameters used to control the weights of the visual and text information. After obtaining the importance score of each entity, entities in that are greater than this threshold are screened out using the threshold as the core targets for question generation.

[0059] Next, the vision-language model classifies and predicts the screened entities. Let the confidence of the highest category be , then the uncertainty measure . Among them, reflects the confidence of the model in the most likely class of the entity, representing the degree of uncertainty of the model.

[0060] Then, entities with relatively high uncertainty ( ) are selected , and questions are generated for these entities. Among them, is a preset threshold condition. The ways of question design can include: object recognition questions, such as "What is this object?"; property recognition questions, such as "What color is this item?"; quantity reasoning questions, such as "How many similar items are there here?"; relationship understanding questions, such as "What is this person doing?"

[0061] According to the combined weight of entity importance and uncertainty , the generation probabilities of different types of questions are balanced. The calculation formula for the combined weight is . Among them, controls the influence based on entity importance; controls the influence based on uncertainty.

[0062] Furthermore, to optimize the quality of question generation, through cross-modal contrastive learning, the model is made to learn to understand the same semantic concept in different contexts. Specifically, positive and negative samples are first constructed. Let be the question generated for , and the original text description is T. Then the positive sample is , and the negative sample is , where is irrelevant text. Then, the semantic similarity between the question and the text T is calculated. The calculation formula for the similarity is , where represents the extracted embedding vector. Then, a contrastive learning loss function (i.e., ContrastiveLoss) is used for training to optimize the contrast loss. Among them, the calculation formula for the contrastive learning loss function is as follows:

[0063] ;

[0064] This contrastive learning loss function can ensure that is closer to the correct text T and farther from the irrelevant text , improve the logical consistency of the questions. After being optimized through contrastive learning, the generated questions are the target single-step questions. For example, in a static object scenario, the visual sample is a picture showing a basket of fruits, including apples, bananas, and oranges. The text description is "There are various fruits in the basket, including apples, bananas, and oranges." Then the generated question examples are: The confidence-driven question (i.e., the initial single-step question) is "Is this fruit an orange?" and "Are all the fruits in the basket round?" The questions after contrastive learning optimization (i.e., the target single-step questions) are "What is the difference in color between apples and bananas?" and "Is an orange bigger than an apple?" Additionally, in a dynamic object scenario, the visual sample is a boy kicking a football, and there is a dog running beside. The text description is "A boy is kicking a football in the park, and there is a dog running beside." Then the generated question examples are: The confidence-driven questions are "What is this person doing?" and "Is the dog standing or running?" The questions after contrastive learning optimization are "What is the difference in the movement patterns between the boy and the dog?"

[0065] Then, based on the above-generated set of target single-step questions , construct a strategy through hierarchical reasoning to gradually expand the complexity and depth of the questions, so as to form a set of target multi-step questions with a progressive relationship . Among them, the multi-step questions are multi-step questions generated by constructing a strategy through hierarchical reasoning from one or more single-step questions . Specifically, determine the set of target multi-step questions by using the set of target single-step questions, the preset semantic expansion strategy, and the preset question reasoning strategy.

[0066] Specifically, the preset semantic expansion strategy is to expand the scope of the question by identifying the semantic boundaries of the single-step questions, making it form a more challenging multi-step reasoning task. This strategy can be divided into information supplementation, context association, and complexity progression. Information supplementation is to add new reasoning elements based on the core information of the single-step questions to make the questions more hierarchical. Context association is to make the questions progress around the existing context rather than directly combining multiple questions. Complexity progression is to ensure that the question generation is a natural progressive process rather than arbitrary splicing. The formula for semantic expansion of single-step questions is , where is the new question after semantic expansion of the single-step question , S represents additional context information, and the context information can come from visual features, text descriptions, or external knowledge.

[0067] The preset question reasoning strategy is to guide the gradual escalation of questions through multi-level reasoning to form multi-step questions with a logical progressive relationship. This strategy can be divided into hierarchical reasoning, context adjustment, and adaptive difficulty control. Hierarchical reasoning is to gradually increase new reasoning requirements on the basis of primary reasoning, such as causal analysis, reasoning chain expansion, etc. Context adjustment is to ensure that the generated questions are coherent in context and do not produce abrupt breaks. Adaptive difficulty control is to dynamically adjust the difficulty of questions according to the answering ability of the model. The formula for question reasoning for single-step questions is , where L represents the reasoning level (such as direct reasoning, indirect reasoning, hypothetical reasoning, etc.), is the finally generated multi-step question (i.e., the target multi-step question). In addition, the following examples illustrate the generated multi-step questions.

[0068] In a specific implementation manner, for example, the single-step question is "Is there a book on the table?", after semantic expansion, it becomes "How many books are there on the table?". In the process of question reasoning, the question for the first-step direct reasoning is "There are 3 books on the table.", the question for the second-step indirect reasoning is "What are the types of books on the table?", and the question for the third-step reasoning deepening is "If 2 more magazines are placed, how many are there in total?". Then the finally generated multi-step question is "How many books are there on the table? What are their respective types? If 2 more magazines are added, how many are there in total?". In another specific implementation manner, for example, the single-step question is "What will happen when the ball falls to the ground?", after semantic expansion, it becomes "How does the ball fall to the ground?". In the process of question reasoning, the question for the first-step direct reasoning is "The ball falls off the table.", the question for the second-step indirect reasoning is "When the ball falls on floors of different materials, how does it bounce?", and the question for the third-step reasoning deepening is "If the ball is rubber, what will be its bounce height?". Then the finally generated multi-step question is "What will happen after the ball falls off the table? If it falls on floors of different materials, how will it bounce? What if it is a rubber ball?".

[0069] Next, design a hierarchical progressive reasoning training system, or progressive thinking heuristic fine-tuning. Through "progressive cognitive guidance" and "multi-perspective thinking modeling" (i.e., the preset model reasoning strategy), help the model start from simple questions and gradually improve its reasoning ability so that it can handle more complex logical tasks. Specifically, "progressive cognitive guidance" gradually increases the question difficulty by setting training stages to ensure that the reasoning ability of the model is improved from shallow to deep. Therefore, in the training stage the model needs to first pass through and then enter , , progressing step by step until it has the ability to . Among them, is direct reasoning, that is, the model only needs to obtain the answer based on the input information without involving complex logical relationships. is multi-step reasoning, that is, the model needs to establish connections between multiple known conditions and perform reasoning calculations. is hypothetical reasoning, that is, the model needs to consider different hypothetical conditions and perform logical deductions based on them. is adversarial reasoning, that is, the model needs to face information loss, uncertainty, and interference factors and still be able to infer the optimal solution. In addition, "multi-perspective thinking modeling" is to train the model to understand problems from different perspectives and establish a more robust reasoning structure, including structured thinking paths. Specifically, for each problem, multiple possible reasoning paths are constructed to let the model learn different problem-solving methods. Causal relationship modeling enables the model to not only know "what" but also "why", strengthening the causal reasoning ability. Dynamic complexity adjustment enables the dynamic adjustment of the complexity of the problem as the model training progresses, so that the model is constantly challenged. And in the process of fine-tuning the initial vision-language model based on the target single-step problem set, target multi-step problem set, and preset model reasoning strategy, a training sample set is generated, which can be expressed by the formula , where C represents the causal reasoning relationship. In addition, a hierarchical mind map can be generated for each problem to help the model learn the reasoning path. The mind map structure will change continuously with the progress of model training, being relatively simple in the initial stage and more complex in the later stage. Therefore, after the progressive thinking heuristic fine-tuning is completed, the first fine-tuned model and the training sample set will be obtained. In addition, the following example illustrates the process of fine-tuning the initial vision-language model based on the target single-step problem set, target multi-step problem set, and preset model reasoning strategy.

[0070] For example, if the visual sample, or the picture content, is a basket containing 3 apples and 5 bananas, and the text description is "There are 3 apples and 5 bananas in the basket." At the stage, the question is "How many fruits are there in total in the basket?", and the thinking path is: Apples (3) + Bananas (5) = 8. At this time, the model prediction is 8, and this prediction result is correct; at the stage, the question is "The number of apples in the basket is 3, and the total number of apples and bananas is 8. How many bananas are there in the basket?", and the thinking path is to set the number of bananas as x. According to the question, 3 + x = 8, and the solution is x = 5. At this time, the model prediction is 5, and this prediction result is correct; at the At the [[[STAGE]]], the question is "If 2 more oranges are added, how many fruits are there in the basket now?", and the thinking path is that it is known that apples (3) + bananas (5) = 8. Assuming the newly added oranges are (2), calculate the new total: 8 + 2 = 10. Then the model prediction at this time is 10, and this prediction result is correct; at At the [[[STAGE]]], the question is "There are 3 apples and some bananas in the basket, and the total is 8. How many bananas are there?", and the thinking path is to set the number of bananas as y. According to the question, 3 + y = 8, and the solution is y = 5. Then the model prediction at this time is 5, and this prediction result is correct.

[0071] Next, perform adaptive hybrid masking training on the first fine-tuned model obtained above, that is, fine-tune the model based on the constructed hybrid masking strategy to obtain the second fine-tuned model. Specifically, assume that the visual sample V consists of multiple regions and the text description T consists of multiple words . First, calculate the importance of each region and word. The importance of visual information is , and the importance of language information is , where is the adaptive weight coefficient. Also, assume that the initial visual masking probability is , and the text masking probability is , then the masking range is . Among them, are the dynamically adjusted thresholds respectively. It should be noted that in the initial stage of training, , strong masking to enhance the model's completion ability; in the middle stage of training, , appropriately reduce the masking to maintain the model's challenge; in the later stage of training, , low masking to let the model adapt to full-data reasoning. Thus, the hybrid masking strategy is obtained, and the model is fine-tuned based on the hybrid masking strategy to obtain the second fine-tuned model. The following example illustrates the process of fine-tuning the model using the hybrid masking strategy.

[0072] For example, the current task is fruit basket reasoning. The visual sample input is a basket containing 3 apples and 5 bananas, and the text description is "There are 3 apples and 5 bananas in the basket." In the initial stage of training, the visual mask occludes some of the apples and bananas, and the model still needs to infer the total number of fruits. The text mask removes "3 apples", and the model needs to fill in the missing information based on the visual information. The question is "How many fruits are there in the basket in total?" Then the reasoning path of the model is to see some fruits, identify the missing content, complete the reasoning, and calculate the total number 8; in the middle stage of training, the visual mask occludes a small part of the apples and randomly blurs the background at the same time. The text mask removes the word "bananas" to let the model learn multi-modal alignment reasoning. The question is "Given that there are 8 fruits in the basket in total, and 3 of them are apples, how many bananas are there?" Then the reasoning path of the model is to combine the visual information, reverse-infer the missing number of bananas, and get the answer 5; in the later stage of training, the visual mask is a very small part of the mask, mainly used to fine-tune the model stability, and the text mask is also a very small part of the mask to ensure that the model can reason normally on the complete data. The question is "If 2 more oranges are added, how many fruits are there in the basket now?" Then the reasoning path of the model is to calculate the original quantity 8, add the oranges 2, and get the answer 10.

[0073] Finally, progressive enhancement training is carried out. Specifically, let the teacher model and the student model . This student model is the second fine-tuned model, and KL divergence loss is used for distillation training. The distillation loss formula is . In the formula, represents the self-knowledge distillation loss, represents the KL divergence, represents the result predicted by the model, represents for probability distribution, represents for probability distribution. After distillation, the distilled model is obtained. Then, calculate the difficulty score of the sample set, and the calculation formula is . In the formula, represents the confidence of the model in the question , represents the loss of the model on this question, is the adaptive weight. After obtaining the difficulty score, determine the corresponding sample difficulty, and perform iterative training on the distilled model based on the training sample set and the sample difficulty. Specifically, when the sample difficulty is low, reduce the training frequency; when the sample difficulty is medium, train normally; when the sample difficulty is high, increase the training weight. After performing iterative training on the distilled model, the target vision-language model can be obtained to perform the preset model reasoning operation using the target vision-language model.

[0074] Correspondingly, referring to Figure 3 As shown, an embodiment of the present application provides a model inference device for a vision-language model, including: A single-step question generation module 11, configured to determine a target entity set based on a semantic guidance network, an initial vision-language model, a visual sample, and a first text description corresponding to the visual sample, and determine an entity whose uncertainty meets a preset threshold condition from the target entity set by using the result of classifying and predicting the target entity set, and generate an initial single-step question set based on the entity whose uncertainty meets the preset threshold condition; the single-step question is a single logical relationship question; A single-step question optimization module 12, configured to optimize the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample, so as to obtain a target single-step question set; A first model fine-tuning module 13, configured to determine a target multi-step question set by using the target single-step question set, a preset semantic expansion strategy, and a preset question reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step question set, the target multi-step question set, a preset model inference strategy, and the initial vision-language model; the multi-step question is multiple questions with a logical progression relationship based on the single-step question; A second model fine-tuning module 14, configured to determine a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tune the first fine-tuned model by using the hybrid masking strategy to obtain a second fine-tuned model; A model inference module 15, configured to determine the second fine-tuned model as a student model, and distill the student model by using a preset teacher model to obtain a distilled model, determine the sample difficulty of the training sample set, and perform iterative training on the distilled model based on the training sample set and the sample difficulty, so as to trigger model inference by using the obtained target vision-language model.

[0075] As can be seen from the above, in the present application, first, by combining a semantic guidance network, an initial vision-language model, visual samples, and corresponding first text descriptions, a target entity set is determined. The target entities are classified and predicted to find entities with uncertainty reaching a preset threshold, thereby generating an initial single-step problem set with a single logical relationship. Next, the initial single-step problems are optimized using the initial single-step problem set, the first text description, and a second text description independent of the visual samples to obtain a target single-step problem set. Subsequently, based on the target single-step problem set, a preset semantic expansion and problem reasoning strategy are used to obtain a target multi-step problem set with a logical progression relationship. In combination with a preset model reasoning strategy, the initial vision-language model is fine-tuned to determine a training sample set and obtain a first fine-tuned model. Then, according to the visual samples, the first text description, and preset visual and text masking probabilities, a hybrid masking strategy is determined. This strategy is used to fine-tune the first fine-tuned model again to obtain a second fine-tuned model. Finally, the second fine-tuned model is used as a student model, and a distillation model is obtained by distillation using a preset teacher model. The sample difficulty of the training sample set is determined, and based on this, the distillation model is iteratively trained to trigger model reasoning using the obtained target vision-language model. In this way, the present application can improve the model reasoning ability of the vision-language model and enhance the user experience to a certain extent.

[0076] In some specific embodiments, the single-step problem generation module 11 specifically includes: A cross-modal parsing unit for parsing visual samples and a first text description corresponding to the visual samples using the initial vision-language model to obtain corresponding cross-modal representations; An entity extraction unit for extracting an initial entity set in the visual samples and the first text description using the cross-modal representations through a semantic guidance network to complete the entity extraction operation; A first importance determination unit for determining a first importance degree and a second importance degree of the initial entity set; the first importance degree is the importance degree of the initial entity predicted by the initial vision-language model in visual information, and the second importance degree is the importance degree of the initial entity predicted by the initial vision-language model in the text description; An importance score determination unit for determining an importance score of each initial entity in the initial entity set based on the first importance degree, the second importance degree, a preset visual weight control parameter, and a preset text weight control parameter; An entity screening unit for determining, in the initial entity set, the initial entities with importance scores greater than a preset importance threshold as target entities, and constructing a target entity set based on each of the target entities.

[0077] In some specific embodiments, the single-step problem generation module 11 specifically includes: A classification prediction unit for classifying and predicting the target entity set and obtaining a classification prediction result; A confidence screening unit for determining corresponding category information from the classification prediction result and determining a first target confidence level in the category information where the category probability value is greater than a preset probability threshold; An uncertainty determination unit for determining the uncertainty results of each target entity in the target entity set based on a preset constant value and the first target confidence level, and using the uncertainty results to determine entities in the target entity set whose uncertainty meets a preset threshold condition; An index determination unit for determining the target importance score and target uncertainty of the entities whose uncertainty meets the preset threshold condition; A weight determination unit for determining an entity weight result based on the target importance score, the target uncertainty, a preset importance control coefficient, and a preset uncertainty control coefficient; A problem determination unit for correspondingly adjusting the problem type of the generated single-step problem using the entity weight result to determine an initial single-step problem set based on the adjusted problem type.

[0078] In some specific embodiments, the single-step problem optimization module 12 specifically includes: A negative sample construction unit for constructing positive samples using the initial single-step problem set and the first text description corresponding to the visual sample, and constructing negative samples using the initial single-step problem set and a second text description that has no preset corresponding relationship with the visual sample; A first similarity determination unit for determining a first semantic similarity between the initial single-step problem set and the first text description based on the embedding vectors of the initial single-step problem set and the first text description in the positive samples; A second similarity determination unit for determining a second semantic similarity between the initial single-step problem set and the second text description based on the embedding vectors of the initial single-step problem set and the second text description in the negative samples; A first loss value determination unit for determining a target loss value based on the first semantic similarity, the second semantic similarity, and a preset contrastive learning loss function; A problem optimization unit for optimizing the initial single-step problem set using the target loss value to obtain a target single-step problem set.

[0079] In some specific embodiments, the first model fine-tuning module 13 specifically includes: A first model fine-tuning unit, configured to fine-tune the initial vision-language model by using the target single-step question set, the target multi-step question set, and a preset model inference strategy, so as to obtain a first fine-tuned model; A sample set determination unit, configured to determine a training sample set based on the target single-step question set and the target multi-step question set during the process of fine-tuning the initial vision-language model.

[0080] In some specific embodiments, the second model fine-tuning module 14 specifically includes: An information partitioning unit, configured to partition the visual sample into a plurality of regions, and partition the first text description into a plurality of words; A second importance determination unit, configured to determine visual information importance based on the plurality of regions and a first adaptive weight coefficient, and determine text information importance based on the plurality of words and a second adaptive weight coefficient; A strategy determination unit, configured to determine a hybrid masking strategy based on the visual information importance, a preset visual masking probability, the text information importance, and a preset text masking probability; A second model fine-tuning unit, configured to fine-tune the first fine-tuned model by using the hybrid masking strategy to obtain a second fine-tuned model.

[0081] In some specific embodiments, the model inference module 15 specifically includes: A second loss value determination unit, configured to determine a second target confidence level of the distilled model for predicting each current question in the training sample set, and determine a target loss value of the distilled model on each current question; A model iteration unit, configured to determine the sample difficulty of the training sample set based on the second target confidence level and the target loss value, and perform iterative training on the distilled model based on the training sample set and the sample difficulty to obtain target model parameters; A model inference unit, configured to adjust the current model parameters of the distilled model by using the target model parameters, and obtain a target vision-language model, so as to trigger a preset model inference operation by using the target vision-language model.

[0082] Furthermore, an embodiment of the present application also discloses an electronic device, Figure 4It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment. The content in the figure should not be regarded as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the model inference method for the visual language model disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0083] In this embodiment, the power supply 23 is used to provide working voltages for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application requirements, and no specific limitation is made here.

[0084] In addition, as a carrier for resource storage, the memory 22 can be a read-only memory, a random access memory, a magnetic disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0085] Among them, the operating system 221 is used to manage and control each hardware device and the computer program 222 on the electronic device 20, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the model inference method for the visual language model executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0086] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the model inference method for the visual language model disclosed above. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0087] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the apparatuses disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0088] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0089] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0090] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

[0091] The technical solutions provided in this application have been introduced in detail above. Specific examples are used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, based on the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A model reasoning method for a visual language model, characterized in that: include: Determine a target entity set based on a semantic guidance network, an initial visual language model, a visual sample, and a first text description corresponding to the visual sample, determine entities whose uncertainties meet a preset threshold condition from the target entity set using a result of classification prediction of the target entity set, and generate an initial single-step question set based on the entities whose uncertainties meet the preset threshold condition; Single-step problems are problems with a single logical relationship; Optimizing the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample to obtain a target single-step question set; Determine a target multi-step question set using the target single-step question set, the preset semantic expansion strategy, and the preset question reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step question set, the target multi-step question set, the preset model reasoning strategy, and the initial visual language model; a multi-step question is a plurality of questions having a logically progressive relationship based on a single-step question; Determining a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tuning the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model; The second fine-tuned model is determined as a student model, and the student model is distilled using a preset teacher model to obtain a distilled model, the sample difficulty of the training sample set is determined, and the distilled model is iteratively trained based on the training sample set and the sample difficulty to trigger model reasoning using the obtained target visual language model.

2. The model reasoning method for visual language model according to claim 1, characterized in that: The step of determining a target entity set based on a semantic guidance network, an initial visual language model, a visual sample, and a first text description corresponding to the visual sample includes: Parsing the visual sample and the first text description corresponding to the visual sample using the initial visual language model to obtain a corresponding cross-modal representation; Extracting an initial entity set from the visual sample and the first text description using the cross-modal representation through a semantic guidance network to complete an entity extraction operation; Determine a first importance and a second importance of the initial entity set; the first importance is the importance of the initial entity predicted by the initial visual language model in the visual information, and the second importance is the importance of the initial entity predicted by the initial visual language model in the text description; Determine the importance score of each initial entity in the initial entity set based on the first importance, the second importance, a preset visual weight control parameter, and a preset text weight control parameter; In the initial entity set, the initial entities whose importance scores are greater than a preset importance threshold are determined as target entities, so as to construct a target entity set based on each of the target entities.

3. The model reasoning method for visual language model according to claim 1, characterized in that: The process of determining entities whose uncertainties satisfy a preset threshold condition from the target entity set by using the result of classification prediction of the target entity set, and generating an initial single-step question set based on the entities whose uncertainties satisfy the preset threshold condition, includes: Performing classification prediction on the target entity set and obtaining classification prediction results; Determining corresponding category information from the classification prediction result, and determining a first target confidence level in the category information whose category probability value is greater than a preset probability threshold; Determine the uncertainty result of each target entity in the target entity set based on a preset constant value and the first target confidence, and determine an entity whose uncertainty satisfies a preset threshold condition from the target entity set using the uncertainty result; Determine a target importance score and a target uncertainty of an entity whose uncertainty satisfies a preset threshold condition; Determine an entity weight result based on the target importance score, the target uncertainty, a preset importance control coefficient, and a preset uncertainty control coefficient; The entity weight results are used to adjust the question types of the generated single-step questions accordingly, so as to determine an initial single-step question set based on the adjusted question types.

4. The model reasoning method for visual language model according to claim 1, characterized in that: The step of optimizing the initial single-step question set based on the initial single-step question set, the first text description, and the second text description that has no preset corresponding relationship with the visual sample to obtain a target single-step question set includes: constructing a positive sample using the initial single-step question set and the first text description corresponding to the visual sample, and constructing a negative sample using the initial single-step question set and a second text description that has no preset corresponding relationship with the visual sample; In the positive sample, determining a first semantic similarity between the initial single-step question set and the first text description based on the embedding vectors of the initial single-step question set and the first text description; In the negative sample, determining a second semantic similarity between the initial single-step question set and the second text description based on the embedding vectors of the initial single-step question set and the second text description; Determining a target loss value based on the first semantic similarity, the second semantic similarity, and a preset contrastive learning loss function; The initial single-step problem set is optimized using the target loss value to obtain a target single-step problem set.

5. The model reasoning method for visual language model according to claim 1, characterized in that: The determining of the training sample set and the first fine-tuned model based on the target single-step question set, the target multi-step question set, the preset model reasoning strategy, and the initial visual language model includes: Fine-tune the initial visual language model using the target single-step question set, the target multi-step question set, and a preset model reasoning strategy to obtain a first fine-tuned model; In the process of fine-tuning the initial visual language model, a training sample set is determined based on the target single-step question set and the target multi-step question set.

6. The model reasoning method for visual language model according to claim 1, characterized in that: The step of determining a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and fine-tuning the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model includes: Dividing the visual sample into a plurality of regions and dividing the first text description into a plurality of words; Determine the importance of visual information based on the plurality of regions and the first adaptive weight coefficient, and determine the importance of text information based on the plurality of words and the second adaptive weight coefficient; Determining a hybrid masking strategy based on the visual information importance, the preset visual masking probability, the text information importance, and the preset text masking probability; The first fine-tuned model is fine-tuned using the hybrid mask strategy to obtain a second fine-tuned model.

7. The model reasoning method for a visual language model according to any one of claims 1 to 6, characterized in that: The determining of the sample difficulty of the training sample set, and iteratively training the distillation model based on the training sample set and the sample difficulty, so as to trigger model reasoning using the obtained target visual language model, includes: Determine a second target confidence level of the distillation model for predicting each current problem in the training sample set, and determine a target loss value of the distillation model for each current problem; Determining the sample difficulty of the training sample set based on the second target confidence and the target loss value, and iteratively training the distillation model based on the training sample set and the sample difficulty to obtain target model parameters; The target model parameters are used to adjust the current model parameters of the distillation model, and a target visual language model is obtained, so as to trigger a preset model reasoning operation using the target visual language model.

8. A model reasoning device for a visual language model, characterized in that: include: A single-step question generation module is used to determine a target entity set based on a semantic guidance network, an initial visual language model, a visual sample, and a first text description corresponding to the visual sample, and to determine entities whose uncertainties meet a preset threshold condition from the target entity set using a result of classification prediction of the target entity set, and to generate an initial single-step question set based on the entities whose uncertainties meet the preset threshold condition; Single-step problems are problems with a single logical relationship; a single-step question optimization module, configured to optimize the initial single-step question set based on the initial single-step question set, the first text description, and a second text description that has no preset corresponding relationship with the visual sample, so as to obtain a target single-step question set; A first model fine-tuning module is used to determine a target multi-step question set by using the target single-step question set, a preset semantic extension strategy, and a preset question reasoning strategy, and determine a training sample set and a first fine-tuned model based on the target single-step question set, the target multi-step question set, the preset model reasoning strategy, and the initial visual language model; Multi-step problems are multiple problems that have a logical progressive relationship based on single-step problems; A second model fine-tuning module, configured to determine a hybrid masking strategy based on the visual sample, the first text description, a preset visual masking probability, and a preset text masking probability, and to fine-tune the first fine-tuned model using the hybrid masking strategy to obtain a second fine-tuned model; A model reasoning module is used to determine the second fine-tuned model as a student model, and use a preset teacher model to distill the student model to obtain a distilled model, determine the sample difficulty of the training sample set, and iteratively train the distilled model based on the training sample set and the sample difficulty to trigger model reasoning using the obtained target visual language model.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the model reasoning method for a visual language model as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program; wherein, when the computer program is executed by a processor, the model reasoning method for a visual language model as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Question and answer method and device based on multi-modal information and application of question and answer method and device

    CN117828142A

  • Video language task execution method and device, video language task model training method and device, equipment and medium

    CN117876940A

  • Visual language model instruction fine tuning method and device

    CN117975475A

  • Registration knowledge distillation method, system and equipment for visual language model

    CN118334463A

  • Visual text question and answer method based on entity alignment and cross-modal reasoning

    CN118428479A