Visual reasoning adjustment method based on mixed training and electronic equipment

Through a hybrid training method combined with supervised fine-tuning and reinforcement learning, the visual language model is optimized, and the problems of model generalization ability and data efficiency in visual inference tasks are solved, achieving stronger cross-domain adaptability and efficient data utilization.

CN120509455APending Publication Date: 2025-08-19BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

Patent Information

Application Number
CN202510605504.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

In existing visual inference tasks, neural symbolic methods are complex and costly, and the supervised fine-tuning method relies on a large amount of high-quality annotation data, resulting in limited generalization capabilities of the model and limited application.

Method used

A visual inference adjustment method based on hybrid training is adopted, combined with supervised fine-tuning and reinforcement learning, and the visual language model is optimized through group relative strategy optimization algorithms, reducing dependence on labeled data, and improving the generalization ability of the model.

Benefits of technology

It significantly improves the generalization ability and data efficiency of the model in cross-domain tasks, avoids overfitting and cognitive rigidity, and reduces the dependence on large-scale annotation data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509455A_ABST
    Figure CN120509455A_ABST
Patent Text Reader

Abstract

The invention discloses a visual reasoning adjustment method based on mixed training and electronic equipment, and belongs to the technical field of data processing.The method comprises the steps that a visual reasoning data set is determined, and the supervision fine tuning reasoning ability of a visual language model is activated based on the visual reasoning data set, so that the visual language model can complete a target reasoning process; and utilizing a group relative strategy optimization algorithm to optimize the supervised fine tuning reasoning capability of the visual language model. According to the method, the dependence on a large amount of annotated data is reduced, and the data efficiency is improved; and through a dynamic optimization mechanism of reinforcement learning, the method can better adapt to cross-domain tasks, and the generalization ability of the model is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data processing technology, and in particular relates to a visual reasoning adjustment method and electronic device based on hybrid training. Background Art

[0002] Currently, research on visual reasoning tasks primarily focuses on neural-symbolic methods and supervised fine-tuning (SFT). However, the implementation of neural-symbolic methods is often complex, requiring the design of complex symbolic reasoning rules or program generation mechanisms, which increases model development and maintenance costs. SFT methods require a large amount of high-quality chained thought (CoT) annotated data, which is expensive and time-consuming to obtain, limiting the widespread application of these models.

[0003] In response to the above problems, the present application proposes a visual reasoning adjustment method and electronic device based on hybrid training. Summary of the Invention

[0004] In order to address the deficiencies of the prior art, the present application provides a visual reasoning adjustment method and electronic device based on hybrid training, which can solve the problems of complex implementation, high cost and limited application in the visual reasoning method of the prior art.

[0005] The technical effects to be achieved by this application are achieved through the following solutions:

[0006] In a first aspect, the present application provides a visual reasoning adjustment method based on hybrid training, comprising:

[0007] Determining a visual reasoning dataset, and activating a supervised fine-tuning reasoning capability of a visual language model based on the visual reasoning dataset, so that the visual language model can complete a target reasoning process;

[0008] The supervised fine-tuning reasoning capability of the visual language model is optimized using a group relative strategy optimization algorithm.

[0009] In some embodiments, the visual reasoning dataset includes multiple training samples, each training sample consists of an input image x, an input question q, a reasoning step r and a final answer a, and each training sample is represented as (x, q, r, a); the reasoning step r is the intermediate reasoning process of the visual language model when answering questions.

[0010] In some embodiments, activating the supervised fine-tuning reasoning capability of the visual language model based on the visual reasoning dataset includes:

[0011] Determine the objective function, which is as follows:

[0012]

[0013] in, represents the objective function, which aims to maximize the likelihood probability of the generated sequence;

[0014] Represents the expected value, which represents the Calculate the average loss of samples sampled in ;

[0015] x represents the input image;

[0016] q indicates input question;

[0017] r represents the reasoning step;

[0018] a indicates the final answer;

[0019] represents the training dataset containing the visual reasoning dataset;

[0020] T represents the total length of the generated sequence;

[0021] π θ represents the generation strategy with parameter θ;

[0022] y t is the tth token of the generated sequence;

[0023] y <t is all tokens before the t-th token in the generated sequence;

[0024] t represents a positive integer and 1≤t≤T.

[0025] In some embodiments, after determining the objective function, the method further includes:

[0026] Maximize the objective function by gradient descent Likelihood probability, based on the input image x, input question q and y <t Predict the current token y t , and optimize the model parameters of the visual language model by minimizing the objective function.

[0027] In some embodiments, optimizing the supervised fine-tuning reasoning capability of the visual language model using a group-relative strategy optimization algorithm includes:

[0028] Determine the action group sampling, for each input state s = (x, q), from the generation strategy π θ Sampling a set of actions {a1,…,a i ,…,a G}, where x represents the input image, q represents the input question, i and G are positive integers, and 1 <i<G。

[0029] In some embodiments, the sampling action a is performed according to verifiable criteria. i Assign a reward function that includes an accuracy reward R Acc (a i ).

[0030] In some embodiments, the accuracy reward R Acc (a i )for:

[0031]

[0032] Among them, a pred is the predicted answer; a gt is the true answer; ∈1 is the tolerance threshold used to determine whether it is a complete match, and ∈2 is the upper bound of the partial reward used to determine whether it is completely wrong;

[0033] The reward mechanism is divided into three cases: If |a pred -a gt |<∈1×|a gt |, then the accuracy reward is 1; if ∈1×|a gt |≤|a pred -a gt |<∈2×|a gt |, the accuracy reward transitions smoothly between 0 and 1; if |a pred -a gt |≥∈2×|a gt |, then the accuracy reward is 0;

[0034] or,

[0035] Accuracy reward R Acc (a i )for:

[0036]

[0037] Among them, a pred is the predicted answer, a gt is the real answer.

[0038] In some embodiments, the reward function also includes a format reward R Format (a i ), format reward R Format (a i ) Generate a response according to the predefined template and get a full format reward, otherwise the format reward is 0.

[0039] In some embodiments, the sampling action a is performed according to verifiable criteria. i Assign reward functions, including:

[0040] Sampling action a according to verifiable criteria i Allocation Format Reward R Format (a i ) and accuracy reward R Acc (a i ).

[0041] In some embodiments, the visual reasoning adjustment method based on hybrid training further includes:

[0042] By calculating each sampled action a i Corresponding relative advantage A i , based on relative advantage A i Optimizing the supervised fine-tuning reasoning capability of the vision-language model.

[0043] In a second aspect, the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the aforementioned methods when executing the computer program.

[0044] In a third aspect, the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement any of the aforementioned methods.

[0045] The embodiments of the present application provide a visual reasoning adjustment method and electronic device based on hybrid training. This method reduces dependence on large amounts of labeled data and improves data efficiency. Through the dynamic optimization mechanism of reinforcement learning (RL), it can better adapt to cross-domain tasks and significantly improve the generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the embodiments of the present application or the existing technical solutions, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 Flowchart of a visual reasoning adjustment method based on hybrid training in one embodiment of the present application;

[0048] Figure 2 This is a schematic block diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0050] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which the present application belongs. The "first", "second" and similar words used in one or more embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.

[0051] In related technologies, neural symbolic methods aim to improve the interpretability and modularity of models by combining symbolic reasoning with neural networks. Such methods usually rely on program generation or rule reasoning, and can demonstrate strong reasoning capabilities in certain specific tasks. However, such methods are highly dependent on program generation, which makes it difficult for them to respond flexibly when faced with complex or changing visual reasoning tasks, especially when the task type or data distribution changes, and their generalization ability is poor. Due to the limitations of symbolic reasoning, neural symbolic methods perform poorly when processing large-scale, multimodal data and are difficult to adapt to complex real-world scenarios.

[0052] Supervised fine-tuning methods are based on the visual language model (VLM) and enhance the model's reasoning ability through end-to-end training. In recent years, supervised fine-tuning methods based on Chain-of-Thought (CoT) have been widely used, helping models better understand complex tasks by providing labeled data for step-by-step reasoning. Since the SFT method relies on fixed training data, the model is prone to overfitting on specific tasks, resulting in poor performance when facing new tasks or cross-domain tasks, and limited generalization ability. In order to improve the generalization ability of the model, the SFT method usually requires the design of complex data mixing strategies, which increases the complexity and uncertainty of training. In addition, although the SFT method performs well on specific tasks, its reasoning ability is limited by the scope of the training data, making it difficult to cope with complex, unseen reasoning tasks.

[0053] Therefore, it is necessary to adopt the visual reasoning adjustment method based on hybrid training provided by this application to overcome the shortcomings of the above two methods. This application solves the key problems of existing methods in practical applications and shows stronger generalization ability and higher data efficiency in visual reasoning tasks.

[0054] The visual reasoning adjustment method based on hybrid training (ReasonRFT) proposed in this application can address the limitations of the SFT method in generalization ability and data efficiency.

[0055] ReasonRFT introduces a two-stage training framework that combines the advantages of supervised fine-tuning and reinforcement learning. It not only activates the model's reasoning potential through SFT, but also further improves the model's generalization ability and data efficiency through the reinforcement learning algorithm of the group relative policy optimization algorithm (or group relative reward policy optimization (GRPO)).

[0056] Overall, ReasonRFT reduces its reliance on large amounts of labeled data by generating multiple reasoning-response pairs, improving data efficiency. Through the dynamic optimization mechanism of reinforcement learning, ReasonRFT is better adapted to cross-domain tasks and significantly improves the model's generalization capabilities. Furthermore, the introduction of reinforcement learning enables the model to dynamically adjust its reasoning strategy during training, avoiding the overfitting and cognitive rigidity common in SFT methods.

[0057] Various non-limiting embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0058] First, refer to Figure 1 , the visual reasoning adjustment method based on hybrid training of this application is described in detail:

[0059] In a first aspect, the present application provides a visual reasoning adjustment method based on hybrid training, comprising:

[0060] S1: Determine a visual reasoning dataset, and activate the supervised fine-tuning reasoning capability of the visual language model based on the visual reasoning dataset, so that the visual language model can complete the target reasoning process;

[0061] S2: Optimizing the supervised fine-tuning reasoning capability of the visual language model using a group-relative strategy optimization algorithm.

[0062] The visual reasoning adjustment method based on hybrid training in this application reduces the dependence on large amounts of labeled data and improves data efficiency. Through the dynamic optimization mechanism of reinforcement learning, it can better adapt to cross-domain tasks, significantly improve the generalization ability of the model, and can dynamically adjust the reasoning strategy during training, avoiding the overfitting and cognitive rigidity problems in common methods.

[0063] This application significantly improves the generalization and reasoning performance of the Visual Language Model (VLM) in complex visual reasoning tasks through a two-stage training framework. Specifically, the Reason-RFT of this application introduces a two-stage training framework for visual reasoning.

[0064] In the first stage, through supervised fine-tuning (SFT) with chain-of-thought (CoT) reasoning, high-quality visual reasoning datasets are used to activate the model's domain-specific reasoning ability, that is, to activate the supervised fine-tuning reasoning ability of the visual language model, enabling it to decompose complex tasks and generate a logically clear reasoning process.

[0065] In the second stage, the reasoning capability is further enhanced through the Group Relative Policy Optimization (GRPO) algorithm, enabling Reason-RFT to achieve excellent generalization capabilities by pushing the reasoning limits of the model.

[0066] In some embodiments, the visual reasoning dataset includes multiple training samples, each training sample consists of an input image x, an input question q, a reasoning step r and a final answer a, and each training sample is represented as (x, q, r, a); the reasoning step r is the intermediate reasoning process of the visual language model when answering questions.

[0067] For example, the reasoning step r is the intermediate reasoning process of the model when answering a question, usually presented in natural language, to help the model understand how to derive the final answer from the input image and question.

[0068] In some embodiments, activating the supervised fine-tuning reasoning capability of the visual language model based on the visual reasoning dataset includes:

[0069] Determine the objective function, which is as follows:

[0070]

[0071] in, represents the loss function for supervised fine-tuning, which aims to maximize the likelihood of the generated sequence;

[0072] Represents the expected value, which represents the Calculate the average loss of samples sampled in ;

[0073] x represents the input image; that is, the visual input of the visual language model;

[0074] q represents the input question; that is, the text input of the visual language model;

[0075] r represents the reasoning step, that is, the intermediate reasoning process;

[0076] a indicates the final answer;

[0077] Represents a training dataset containing a visual reasoning dataset; that is, a training dataset containing samples (x, q, r, a);

[0078] T represents the total length of the generated sequence;

[0079] π θ represents a generation strategy with parameter θ; that is, a parameterized probability distribution of the visual language model;

[0080] y t is the tth token of the generated sequence;

[0081] y <t is all the tokens before the t-th token in the generated sequence; that is, the historical context;

[0082] t represents a positive integer and 1≤t≤T.

[0083] For example, by minimizing this objective function, the visual language model can learn to generate coherent reasoning steps and accurate final answers, thereby improving the initial reasoning ability of multimodal tasks.

[0084] The training goal is to optimize the model by maximizing the likelihood of generating a reasoning step r and a final answer a. Specifically, the model needs to generate the correct reasoning step r and final answer a based on the input image x and the input question q. During training, the model gradually decomposes complex tasks and generates a logically clear reasoning process, thereby activating its domain-specific reasoning capabilities.

[0085] For example, after supervised fine-tuning, the visual language model is able to generate logical reasoning steps r and the final answer a.

[0086] In some embodiments, after determining the objective function, the method further includes:

[0087] Maximize the objective function by gradient descent Likelihood probability, based on the input image x, input question q and y <t Predict the current token y t , and optimize the model parameters of the visual language model by minimizing the objective function.

[0088] For example, this stage maximizes the objective function by gradient descent method. Likelihood probability, using the training data set The model is trained on the (x,q,r,a) samples in the teacher-forcing autoregressive sequence, where each step is based on the input image x, the input question q and the historical output y. <t Predict the current token y t The model parameters are optimized by minimizing the objective function. This process enables the model to imitate the reasoning path annotated by humans, providing a high-quality initialization strategy for subsequent reinforcement learning and reducing the difficulty of exploration in the RL stage.

[0089] GRPO updates the policy by comparing the relative advantages of a set of sampled actions, reducing the computational overhead required in traditional reinforcement learning methods (such as PPO) while improving the generalization ability of the model.

[0090] In some embodiments, optimizing the supervised fine-tuning reasoning capability of the visual language model using a group-relative strategy optimization algorithm includes:

[0091] Determine the action group sampling, for each input state s = (x, q), from the generation strategy π θ Sampling a set of actions {a1,…,a i ,…,a G}, where x represents the input image, q represents the input question, i and G are positive integers, and 1 <i<G。

[0092] Exemplarily, the above actions are multiple inference-response pairs generated by the model, ensuring diverse responses and avoiding premature convergence. The sampling process is performed using the model's current policy to ensure diverse responses.

[0093] In some embodiments, the sampling action a is performed according to verifiable criteria. i Assign a reward function that includes an accuracy reward R Acc (a i ).

[0094] In some embodiments, the accuracy reward R Acc (a i )for:

[0095] Among them, a pred is the predicted answer; a gt is the true answer (groundtruth); ∈1 is the tolerance threshold used to judge whether it is a complete match, ∈2 is the upper bound of the partial reward used to judge whether it is completely wrong;

[0096] The reward mechanism is divided into three cases: If |a pred -a gt |<∈1×|a gt |, then the accuracy reward is 1, indicating that the predicted answer is exactly the same as the true answer; if ∈1×|a gt |≤|a pred -a gt |<∈2×|a gt |, indicating that there is a slight deviation between the predicted answer and the true answer, the accuracy reward transitions smoothly between 0 and 1; if |a pred -a gt |≥∈2×|a gt |, indicating that the predicted answer deviates too much from the true answer, and the accuracy reward is 0;

[0097] or,

[0098] Accuracy reward R Acc (a i )for:

[0099]

[0100] Among them, a pred is the predicted answer, a gt is the true answer (groundtruth). If the model's predicted answer is exactly the same as the groundtruth, the reward is 1; otherwise, the reward is 0. This binary reward mechanism ensures the accuracy of the model in tasks that require clear answers.

[0101] In some embodiments, the reward function also includes a format reward R Format (a i ), format reward R Format (a i ) Generate a response according to the predefined template and get a full format reward, otherwise the format reward is 0.

[0102] For example, the format reward requires the model to generate responses strictly according to the predefined template, and the reasoning process must be included in <think> and< / think>The final answer is contained between the tags <answer> and< / answer> Responses that strictly follow the format will receive a full reward, otherwise the reward is 0. This design ensures that the responses generated by the model are structured and interpretable.

[0103] In some embodiments, the sampling action a is performed according to verifiable criteria. i Assign reward functions, including:

[0104] Sampling action a according to verifiable criteria i Allocation Format Reward R Format (a i ) and accuracy reward R Acc (a i ).

[0105] For example, the accuracy reward R Acc (a i ) can also be of function type;

[0106] Function-based accuracy rewards: Applicable to spatial transformation tasks, these rewards require the model to predict a sequence of transformation functions. The answer to such tasks is typically a sequence consisting of multiple transformation operations, necessitating a reward mechanism that assesses the degree of sequence alignment. Function-based accuracy rewards are designed to assess the alignment of the transformation sequence generated by the model with the true sequence, while also supporting rewards for partially correct responses, ensuring a fair assessment of the model's performance on complex tasks.

[0107] Specifically, the function type accuracy reward R acc (a i ) can be calculated as follows:

[0108]

[0109] Among them, len represents the length, T pred is the transformed sequence predicted by the model, T gt is the true transformation sequence (groundtruth), is a fully matched subsequence (the function name f, object o, and attribute value v all match), is a partial matching subsequence (the function name f and the object o both match, or the function name f and the attribute value v both match), is a subsequence that only matches the function name f, α and β are weighting coefficients used to adjust the contribution of partial matches (for example, α = 0.5, β = 0.25). The reward mechanism is divided into four cases: If a step in the predicted sequence is completely consistent with the true sequence (function, object and value all match), it is counted Get full score reward; if a step in the predicted sequence is partially consistent with the true sequence (function and object match, or function and value match), it is counted Get partial reward; if only the function matches at a step in the prediction sequence, it is counted Finally, the total reward is calculated through weighted summation to ensure that the performance of the model in generating complex sequences is fairly evaluated.

[0110] In some embodiments, the visual reasoning adjustment method based on hybrid training further includes:

[0111] By calculating each sampled action a i Corresponding relative advantage A i , based on relative advantage A i Optimizing the supervised fine-tuning reasoning capability of the vision-language model.

[0112] In order to comprehensively evaluate the performance of this application in visual reasoning tasks, this application designed a series of experiments covering three major categories of tasks: visual counting, structural perception, and spatial transformation. The experimental setup includes the selection of datasets, the definition of evaluation indicators, and the specific process of training and testing. In terms of datasets, this application uses CLEVR-Math and Super-CLEVR for visual counting tasks, GeoMath (including Geo170K and Math360K) and Geometry3K for structural perception tasks, and TRANCE for spatial transformation tasks. CLEVR-Math contains 35,000 training samples and 1,000 test samples. Super-CLEVR is a cross-domain test set containing 1,000 samples to evaluate the generalization ability of the model. GeoMath contains 4,500 training samples and 820 test samples. Geometry3K is a cross-domain test set containing 800 samples to evaluate the performance of the model on geometric problems. The TRANCE dataset contains 60,000 training samples and 6,000 test samples, and generates cross-domain test samples by rendering data from different perspectives (such as left and right perspectives) to evaluate the robustness of the model to changes in perspective. In terms of evaluation indicators, this application uses overall accuracy as the main evaluation indicator. For numerical answers, the correctness is verified by mathematical equivalence. For multiple-choice questions, the correctness is verified by string matching. For function type sequences, the correctness is verified by multi-level step-by-step evaluation. In terms of models and training, this application uses Qwen2-VL-2B and Qwen2-VL-7B as the basic models, which are implemented based on Open-R1 and vLLM frameworks. The experiments were conducted on a server cluster equipped with 8×A800 GPUs to ensure efficient training and testing.

[0113] Baseline Methods: We compared various training strategies, including Zero-Shot, ANS-SFT, COT-SFT, Reason-RFT-Zero, and Reason-RFT. Zero-Shot performs inference directly without fine-tuning; ANS-SFT performs supervised fine-tuning based on answer generation; COT-SFT performs supervised fine-tuning based on chained reasoning; Reason-RFT-Zero trains directly using reinforcement learning (GRPO) without a cold start phase; and Reason-RFT trains COT-SFT with reinforcement learning.

[0114] Evaluation Metrics: This application uses overall accuracy as the primary evaluation metric to measure the model's performance across various tasks. For numerical answers, we verify their correctness through mathematical equivalence; for multiple-choice questions, we verify their correctness through string matching; and for function-type sequences, we verify their correctness through multi-level step-by-step evaluation.

[0115] Reason-RFT demonstrates exceptional data efficiency. On the TRANCE dataset, the 2B model achieved 70% of the performance of Reason-RFT-Zero using only 3% of the training data (1,600 samples), and 82.5% of the performance using 9% of the data. The 7B model achieved 92% of the performance of Reason-RFT-Zero using only 3% of the training data, demonstrating its strong capabilities in few-shot learning scenarios. This efficient data utilization capability gives Reason-RFT a significant advantage in practical applications, especially in scenarios where data acquisition costs are high or the amount of data is limited. By combining the advantages of supervised fine-tuning and reinforcement learning, Reason-RFT is able to quickly improve model performance with a small amount of data, reducing its reliance on large-scale labeled data.

[0116] Compared with existing visual reasoning methods, the method of this application has the following three advantages:

[0117] (1) Efficient performance improvement: We achieve state-of-the-art performance in multiple visual reasoning tasks, surpassing existing open source and proprietary models;

[0118] (2) Enhanced generalization capability: It maintains strong performance across diverse tasks and domains, significantly outperforming traditional SFT and RL methods;

[0119] (3) High data efficiency: It performs well in few-shot learning scenarios, achieving 95% of the baseline performance using less than 20% of the data.

[0120] This application significantly improves the performance of the visual language model (VLM) in visual reasoning tasks by combining the advantages of supervised fine-tuning (SFT) and reinforcement learning (GRPO). Compared with traditional SFT methods, Reason-RFT performs well in generalization ability, data efficiency and cross-domain adaptability. Specifically, Reason-RFT has achieved state-of-the-art performance in tasks such as visual counting, structural perception and spatial transformation, especially demonstrating strong generalization ability in cross-domain tasks, and can effectively deal with complex problem types that have not been seen before. In addition, Reason-RFT shows excellent data efficiency in few-sample learning scenarios, and can achieve performance close to that of full-scale data training using only a small amount of data, reducing dependence on large-scale labeled data. Through a dynamic reinforcement learning mechanism, Reason-RFT avoids the common problems of overfitting and cognitive rigidity in traditional SFT methods, and significantly improves the adaptability and reasoning ability of the model in practical applications. These advantages make Reason-RFT of great value in promoting multimodal research and practical applications of visual reasoning tasks.

[0121] It should be noted that the method of one or more embodiments of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and completed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of one or more embodiments of the present application, and the multiple devices will interact with each other to complete the described method.

[0122] It should be noted that the above description is of specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0123] Based on the same inventive concept, corresponding to any of the above embodiments and methods, the present application also discloses an electronic device;

[0124] Specifically, Figure 2 The following is a schematic diagram of the hardware structure of an electronic device that implements a hybrid training-based visual reasoning adjustment method provided in this embodiment. The device may include a processor 410, a memory 420, an input / output interface 430, a communication interface 440, and a bus 450. The processor 410, the memory 420, the input / output interface 430, and the communication interface 440 are interconnected within the device via the bus 450.

[0125] The processor 410 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0126] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 420 can store an operating system and other application programs. When the technical solutions provided in the embodiments of the present application are implemented through software or firmware, the relevant program codes are stored in the memory 420 and are called and executed by the processor 410.

[0127] The input / output interface 430 is used to connect an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0128] The communication interface 440 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (e.g., USB, network cable, etc.) or a wireless method (e.g., mobile network, Wi-Fi, Bluetooth, etc.).

[0129] The bus 450 comprises a pathway for transmitting information between the various components of the device (eg, the processor 410 , the memory 420 , the input / output interface 430 , and the communication interface 440 ).

[0130] It should be noted that although the above device only shows the processor 410, the memory 420, the input / output interface 430, the communication interface 440, and the bus 450, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of the present application, and does not necessarily include all the components shown in the figure.

[0131] The electronic device of the above embodiment is used to implement the corresponding visual reasoning adjustment method based on hybrid training in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0132] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, one or more embodiments of the present application also provide a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the visual reasoning adjustment method based on hybrid training as described in any of the above embodiments.

[0133] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0134] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the visual reasoning adjustment method based on hybrid training as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0135] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present application (including the claims) is limited to these examples. In line with the present application, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of different aspects of one or more embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.

[0136] In addition, to simplify the description and discussion, and in order not to make one or more embodiments of the present application difficult to understand, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making one or more embodiments of the present application difficult to understand, and this also takes into account the following fact, that is, the details of the implementation of these block diagram devices are highly dependent on the platform on which one or more embodiments of the present application will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present application, it will be apparent to those skilled in the art that one or more embodiments of the present application can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0137] Although the present invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may utilize the embodiments discussed.

[0138] The one or more embodiments of the present application are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the one or more embodiments of the present application shall be included within the scope of protection of the present application.

Claims

1. A visual reasoning adjustment method based on hybrid training, characterized in that: include: Determining a visual reasoning dataset, and activating a supervised fine-tuning reasoning capability of a visual language model based on the visual reasoning dataset, so that the visual language model can complete a target reasoning process; The supervised fine-tuning reasoning capability of the visual language model is optimized using a group relative strategy optimization algorithm.

2. The visual reasoning adjustment method based on hybrid training according to claim 1, characterized in that: The visual reasoning dataset includes multiple training samples, each of which consists of an input image x, an input question q , reasoning step r and the final answer a, each training sample is represented as (x, q, r, a); the reasoning step r is the intermediate reasoning process of the visual language model when answering questions.

3. The visual reasoning adjustment method based on hybrid training according to claim 1 or 2, characterized in that: Activating the supervised fine-tuning reasoning capability of the visual language model based on the visual reasoning dataset includes: Determine the objective function, which is as follows: in, represents the objective function, which aims to maximize the likelihood probability of the generated sequence; Represents the expected value, which represents the Calculate the average loss of samples sampled in ; x represents the input image; q indicates input question; r represents the reasoning step; a indicates the final answer; represents the training dataset containing the visual reasoning dataset; T represents the total length of the generated sequence; π θ represents the generation strategy with parameter θ; y t is the tth token of the generated sequence; y <t is all tokens before the t-th token in the generated sequence; t represents a positive integer and 1≤t≤T.

4. The visual reasoning adjustment method based on hybrid training according to claim 3, characterized in that: After determining the objective function, it also includes: Maximize the objective function by gradient descent Likelihood probability, based on the input image x, input question q and y <t Predict the current token y t , and optimize the model parameters of the visual language model by minimizing the objective function.

5. The visual reasoning adjustment method based on hybrid training according to claim 1, characterized in that: The method of optimizing the supervised fine-tuning reasoning capability of the visual language model using a group-relative strategy optimization algorithm includes: Determine the action group sampling, for each input state s = (x, q), from the generation strategy π θ Sampling a set of actions {a1,…,a i ,…,a G }, where x represents the input image, q represents the input question, i and G are positive integers, and 1 <i<G。 6. The visual reasoning adjustment method based on hybrid training according to claim 5, characterized in that: Sampling action a according to verifiable criteria i Assign a reward function that includes an accuracy reward R Acc (a i ).

7. The visual reasoning adjustment method based on hybrid training according to claim 6, characterized in that: Accuracy reward R Acc (a i )for: Among them, a pred is the predicted answer; a gt is the true answer; ∈1 is the tolerance threshold used to determine whether it is a complete match, and ∈2 is the upper bound of the partial reward used to determine whether it is completely wrong; The reward mechanism is divided into three cases: If |a pred -a gt |<∈1×|a gt |, then the accuracy reward is 1; if ∈1×|a gt |≤|a pred -a gt |<∈2×|a gt |, the accuracy reward transitions smoothly between 0 and 1; if |a pred -a gt |≥∈2×|a gt |, then the accuracy reward is 0; or, Accuracy reward R Acc (a i )for: Among them, a pred is the predicted answer, a gt is the real answer.

8. The visual reasoning adjustment method based on hybrid training according to claim 6 or 7, characterized in that: The reward function also includes a format reward R Format (a i ), format reward R Format (a i ) Generate a response according to the predefined template and get a full format reward, otherwise the format reward is 0.

9. The visual reasoning adjustment method based on hybrid training according to claim 8, characterized in that: Sampling action a according to verifiable criteria i Assign reward functions, including: Sampling action a according to verifiable criteria i Allocation Format Reward R Format (a i ) and accuracy reward R Acc (a i ).

10. The visual reasoning adjustment method based on hybrid training according to claim 8, characterized in that: The visual reasoning adjustment method based on hybrid training also includes: By calculating each sampled action a i Corresponding relative advantage A i , based on relative advantage A i Optimizing the supervised fine-tuning reasoning capability of the vision-language model.

11. An electronic device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the visual reasoning adjustment method based on hybrid training as described in any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Method and system for synthesizing reasoning data

    CN119476479A

Cited By

  • Chain thinking enhanced multi-modal spatial reasoning method for highway video data

    CN120822627A

  • A chain thought enhancement multi-modal spatial reasoning method for highway video data

    CN120822627B

  • Unified supervision fine tuning and reinforcement learning training method based on dynamic weight fusion

    CN121145972A

  • Unified supervision fine-tuning and reinforcement learning training method based on dynamic weight fusion

    CN121145972B

  • Training method, device and storage medium of visual language model

    CN122574397B