A semi-offline policy reinforcement learning method for visual language slow thinking reasoning
By employing a semi-offline reinforcement learning approach that combines online visual understanding with offline inference trajectory generation, the shortcomings of slow-thinking reasoning ability and illusion problems in visual language models are addressed, enabling efficient learning and generalization of visual language models in complex tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
- Filing Date
- 2025-06-27
- Publication Date
- 2026-06-02
Smart Images

Figure CN120781972B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and in particular to a semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language. Background Technology
[0002] In recent years, large-scale language models (LLMs), such as the OpenAI o series and the Deepin R1 model, have made breakthroughs in the field of text reasoning due to their strong "slow thinking" capabilities (i.e., modeling complex reasoning processes). This progress has prompted the academic community to attempt to transfer "slow thinking" capabilities to visual-language models (LVLMs) to solve complex problems such as mathematical geometry and multidisciplinary multimodal problems.
[0003] The shortcomings of existing technologies are mainly reflected in three aspects. First, previous distillation strategies primarily guide the model to memorize specific reasoning patterns, making it difficult to acquire true "slow thinking" reasoning abilities. Second, online policy reinforcement learning methods are limited by the initial policy distribution of existing LVLMs, making it difficult to break through existing behavioral patterns. Finally, directly using external model trajectories, LVLMs lack sufficient perception of the visual features involved, easily leading to serious degradation phenomena such as optimization direction conflicts and visual "illusions." Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the existing technology by providing a semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language, so as to solve or partially solve the problems of illusion and difficulty in learning new reasoning methods in visual language models during the reasoning process.
[0005] The objective of this invention can be achieved through the following technical solutions:
[0006] One aspect of the present invention provides a semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language, comprising the following steps:
[0007] Obtain a multimodal training dataset consisting of multiple samples, where each sample includes image information, text question information, and standard answer information;
[0008] Using the image information of the samples as input to a trainable visual language model, the input images are perceived and represented using an online strategy to extract the corresponding visual understanding information.
[0009] The visual understanding information is concatenated with the textual question information of the sample to form a composite input that includes an image structure description and the requirements of the task to be answered. Using a large inference language model, a multi-step logical reasoning trajectory is generated in an offline strategy.
[0010] Based on the logical reasoning trajectory, a reward signal based on the result is calculated. By sampling multiple logical reasoning trajectories, the trajectory reward for obtaining the correct answer is evenly distributed in the visual description layer as a reward label for visual understanding information.
[0011] Based on visual understanding information with reward annotations, logical reasoning trajectories, and their reward signals, the visual language model is optimized for policy gradients to achieve semi-offline policy reinforcement learning.
[0012] As a preferred technical solution, the visual understanding information includes a visual description of the overall structure, local elements, object attributes, and their interrelationships of the image.
[0013] As a preferred technical solution, the logical reasoning trajectory includes restating and understanding the problem, logical reasoning and formula derivation of partial conclusions, correction of errors, and presentation of the final answer.
[0014] As a preferred technical solution, the process of calculating the reward signal based on the result based on the logical reasoning trajectory includes the following steps:
[0015] Regular expressions are used to extract the location of the answer from the logical reasoning trajectory, and the equivalence of mathematical expressions is used to evaluate the difference between the logical reasoning trajectory and the standard answer. If the reasoning trajectory is consistent with the standard answer, the reward is assigned as 1; otherwise, it is assigned as 0.
[0016] As a preferred technical solution, the method of evenly distributing the trajectory reward for obtaining the correct answer in the visual description layer, and the reward annotation as visual understanding information, includes the following steps:
[0017] Multiple logical reasoning trajectories are sampled based on visual understanding information;
[0018] Calculate the reward signal corresponding to each logical reasoning trajectory;
[0019] The average value of the reward signal corresponding to each logical reasoning trajectory is calculated and used as the reward for visual understanding information.
[0020] As a preferred technical solution, the process of performing policy gradient optimization on the visual language model includes the following steps:
[0021] Based on multiple logical reasoning trajectories corresponding to visual understanding information and the reward signals corresponding to the logical reasoning trajectories, samples with high-quality visual understanding and reasoning accuracy are selected using a preset visual reward threshold to construct an offline policy dataset.
[0022] Based on the offline policy dataset, the visual language model is optimized for policy gradient using a policy gradient optimization algorithm with reward signals.
[0023] As a preferred technical solution, the multimodal training dataset includes samples from the disciplines of higher mathematics, physics, engineering, and financial management.
[0024] As a preferred technical solution, during the policy gradient optimization process of the visual language model, the visual encoder weights are frozen.
[0025] In another aspect, an electronic device is provided, comprising: one or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for performing the aforementioned semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language.
[0026] In another aspect, the present invention provides a computer-readable storage medium comprising one or more programs executable by one or more processors of an electronic device, said one or more programs including instructions for performing the aforementioned semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language.
[0027] Compared with the prior art, the present invention has at least one of the following beneficial effects:
[0028] (1) Reduce visual illusion problem: Based on logical reasoning trajectory, the present invention calculates the reward signal based on the result. By sampling multiple logical reasoning trajectories, the reward of the trajectory that obtains the correct answer is evenly distributed in the visual description layer as a reward label for visual understanding information. The visual language model is optimized by strategy gradient. By establishing the above-mentioned visual description reward feedback mechanism, the reasoning reward is evenly fed back to the visual description for reward distribution, which strengthens the influence of visual understanding in the reasoning process and significantly reduces the visual illusion problem.
[0029] (2) Strong learning ability: This invention uses the image information of the sample as the input of the trainable visual language model, perceives and represents the input image in an online strategy, extracts the corresponding visual understanding information, and concatenates the visual understanding information with the text question information of the sample to form a composite input including the image structure description and the requirements of the task to be answered. Using the reasoning large language model, multi-step logical reasoning trajectory is generated in an offline strategy. By combining the online strategy visual understanding of LVLM with the offline strategy reasoning trajectory generation method of external LLM, the visual language model can learn new reasoning methods that exceed its current capabilities.
[0030] (3) Strong generalization ability and high data utilization: This invention has achieved leading results on multiple public benchmark tasks through experimental verification. It surpasses closed-source large models such as GPT-4.1 on highly difficult competition mathematics and competition physics datasets, demonstrating strong generalization ability and high data utilization efficiency. Attached Figure Description
[0031] Figure 1 This is a flowchart of a semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language, as shown in the embodiment.
[0032] Figure 2 This is a schematic diagram of a semi-offline policy reinforcement learning approach for slow-thinking reasoning in visual language, as illustrated in the embodiment.
[0033] Figure 3 These are the prompt words used for visual understanding sampling in the embodiments;
[0034] Figure 4 These are the prompt words used for inference sampling in the embodiments;
[0035] Figure 5 This is a schematic diagram of the electronic device in the embodiment. Detailed Implementation
[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0037] Example 1
[0038] To address the aforementioned problems in existing technologies, this embodiment provides a semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language, aiming to solve the problem of insufficient reasoning ability of current large-scale visual language models (LVLMs) in complex multimodal tasks. The method mainly involves a scalable semi-offline policy reinforcement learning (SOPHIA) framework, which aims to systematically improve the visual slow-thinking reasoning ability of LVLMs while overcoming the shortcomings of existing methods in visual understanding consistency and reasoning generalization ability.
[0039] See Figure 1 and Figure 2 The core technical process includes four main steps: visual understanding adoption, inference sampling, reward propagation, and policy optimization. The entire process is driven collaboratively by an online policy visual understanding model and an offline policy reasoning large language model (LVLM), which are responsible for visual perception and slow-thinking reasoning respectively. The inference capability of the LVLM is efficiently optimized through a reward propagation and allocation mechanism. The method mainly includes the following steps:
[0040] Step S1, Dataset Construction: Obtain a multimodal training dataset including multiple samples.
[0041] Prepare a multimodal training dataset containing images, text questions and their standard answers. This dataset covers multiple disciplines such as advanced mathematics, physics, engineering, and financial management to demonstrate the integration of complex visual scenes and high-order reasoning tasks.
[0042] To ensure data authenticity and diversity, SOPHIA supports both private, high-quality datasets and public datasets such as MathV360K. In practice, data sampling must not only ensure that each question is closely related to a specific image, but also prevent overlap with benchmark data. Keyword filtering and large-scale model-assisted screening are used to eliminate evaluation bias caused by "data leakage."
[0043] Step S2, Visual Understanding Sampling: Using the image information of the samples as input to a trainable visual language model, the input image is perceived and represented using an online strategy to extract the corresponding visual understanding information.
[0044] A trainable LVLM is used to perform detailed perception and representation of the input image using an online policy. Through carefully designed cue engineering, the model is prompted to describe the spatial layout, semantic relationships between objects, and fine-grained low-level visual details (such as cue words) in the image as comprehensively as possible. Figure 3 (As shown). The visual description generated by the model typically includes information such as the overall structure of the image, local features, object attributes, and their interrelationships.
[0045] For example, for an image containing multiple geometric shapes, the model will not only identify the position and size of each shape, but also describe their relative arrangement, overlap, and background color and texture features.
[0046] This visual understanding process employs a fully online sampling strategy, with the LVLM itself outputting data from the current training epoch. This ensures that the visual description is tied to the model's actual perceptual capabilities, reducing the interference of "illusions" or irrelevant content on the inference results.
[0047] Step S3, Inference Sampling: Visual understanding information is concatenated with the textual question information of the sample to form a composite input including image structure description and the requirements of the task to be solved. Using the inference big language model, a multi-step logical reasoning trajectory is generated in an offline strategy.
[0048] During the inference sampling phase, SOPHIA first concatenates the visual description of the image obtained in step S2 with the original question, and then formats it into a composite input suitable for language model understanding through cue engineering. This input not only includes a detailed description of the image structure, but also embeds the specific task requirements to be solved (such as...). Figure 4 (As shown). Subsequently, this combined input is fed into a finely tuned inference language model to generate slow-thinking, multi-step logical reasoning trajectories using an offline strategy.
[0049] After receiving the aforementioned clues, the inference model can simulate a human-like problem-solving process, including "image reading," encompassing phenomena analysis, fact summarization, intermediate calculations, and logical deduction. The sampled inference trajectory typically includes several paragraphs, formula derivations, partial conclusions, and the final answer. This offline strategy inference method overcomes the limitations of LVLM's self-sampling capability. It is important to emphasize that because the visual description is generated by the current LVLM, the inference trajectory always remains consistent with the model's actual visual perception level, reducing optimization direction conflicts caused by external trajectory perception mismatches.
[0050] It should be noted that the inference trajectory sampler (i.e., the inference model) can be replaced with any other accessible LLM, including other open-source or closed-source models. Furthermore, the reward allocation method can be extended to segmented / hierarchical rewards or a weighted mechanism based on multimodal content. This invention framework can be extended to fields such as cross-modal retrieval and visual recognition.
[0051] Step S4, Reward Propagation: Based on the logical reasoning trajectory, calculate the reward signal based on the result. By sampling multiple logical reasoning trajectories, the reward for the trajectory that yields the correct answer is evenly distributed in the visual description layer as a reward label for visual understanding information.
[0052] SOPHIA strengthens the causal link between visual understanding and reasoning trajectories through reward propagation. Specifically, it first automates the evaluation of the final output of each reasoning trajectory. Considering that most reasoning trajectories are long and complex, SOPHIA adopts "outcome-based reward," that is, it only assigns a reward signal of 0 or 1 to the final answer.
[0053] Specifically, regular expressions are used to extract the location of the answer from the reasoning trajectory, and the equivalence of mathematical expressions is used to evaluate the difference between the reasoning trajectory and the standard answer. If it matches the standard answer, a reward of 1 is assigned; otherwise, a reward of 0 is assigned.
[0054] Subsequently, a reward propagation mechanism based on reasoning results was designed for each visual description: for the same set of images and questions, multiple reasoning trajectories were sampled, and the rewards for all trajectories that obtained the correct answer were evenly distributed in the visual description layer as reward annotations for visual understanding.
[0055] Specifically, each visual understanding description generates several inference trajectories. These inferred trajectories are then assigned rewards according to the previously described method. After receiving these rewards, the reward for this visual understanding is considered to be the average of the rewards for the trajectories generated by that visual understanding.
[0056] During subsequent trajectory selection, by setting an appropriate visual reward threshold (e.g., accuracy greater than α), SOPHIA can further filter out samples that possess both high-quality visual understanding and inference accuracy, and prioritize their use for offline policy optimization. This reward mechanism not only ensures the correctness of the inference results but also estimates the rationality of the visual information within them.
[0057] Step S5, Policy Optimization: Based on visual understanding information with reward annotations, logical reasoning trajectories, and their reward signals, the policy gradient of the visual language model is optimized to achieve semi-offline policy reinforcement learning.
[0058] First, SOPHIA organizes the visual descriptions, inference trajectories, and corresponding reward signals obtained in step S4 into a standard offline policy dataset. Then, the model employs a policy gradient optimization algorithm with reward signals to efficiently iteratively update the parameters of the visual-language model. During the actual optimization process, SOPHIA also incorporates techniques such as freezing the visual encoder weights and prioritizing the shortest trajectory to improve model stability and training efficiency. Experimental results show that SOPHIA achieves significant performance improvements on multiple public and private multimodal inference benchmarks, especially in highly challenging scenarios such as mathematics and geometry, far surpassing similar open-source and closed-source large models.
[0059] Specifically, the policy gradient optimization algorithm with reward signal is implemented using the method described in "Richard S. Sutton and Andrew G. Barto. 2018. Reinforcement Learning: An Introduction. Bradford Book, Cambridge, MA, USA."
[0060] The semi-offline policy behavior model, once trained, can be used to process input images and text questions, and output answers to the text questions.
[0061] To demonstrate the effectiveness of this method, experiments were conducted on several mainstream public multimodal inference benchmarks (such as MMMU Pro, MathVision, OlympiadBench, etc.) and mathematical inference datasets, covering LVLMs of scales of 8 billion and 38 billion. See Table 1 for the experimental results:
[0062] Table 1 Experimental Results
[0063]
[0064] Experimental results show that SOPHIA significantly improves inference performance on mainstream tasks, especially the InternVL3.0-38B model, which achieves the best overall performance among open-source models after training with SOPHIA. SOPHIA achieves leading results on multiple public benchmark tasks, with an average improvement of 8.5%, and surpasses closed-source large models such as GPT-4.1 on challenging datasets such as MathVision and OlympiadBench, demonstrating strong generalization ability and high data utilization efficiency.
[0065] In summary, the main improvements made to this method are as follows:
[0066] (1) Combination mechanism of semi-offline policy behavior model: By combining the online policy visual understanding of LVLM with the offline policy reasoning trajectory generation method of external LLM, a semi-offline policy behavior model combining online policy visual understanding and offline policy reasoning is constructed. Through offline policy trajectory sampling, LVLM can learn new reasoning methods beyond its current capabilities.
[0067] (2) Visual-reasoning reward feedback mechanism: The reward distribution method of average feedback of reasoning reward to visual description strengthens the influence of visual understanding in the reasoning process and significantly reduces the visual "illusion" problem.
[0068] (3) Offline policy optimization method based on visual and reasoning rewards: improve the visual slow thinking reasoning ability of LVLM and ensure efficient and low bias policy updates.
[0069] (4) No manual annotation or closed-source model support is required, making it easy to apply on a large scale and highly scalable.
[0070] This embodiment uses a semi-offline strategy approach, integrating policy visual understanding and off-policy reasoning trajectories to overcome the "illusion" problem caused by perceptual mismatch, enhance the linkage between vision and reasoning, improve the generalization ability of multimodal reasoning, and ultimately realize a scalable and data-efficient LVLM slow thinking reasoning training paradigm.
[0071] Example 2
[0072] Based on Embodiment 1, this embodiment provides an electronic device, including: one or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for executing the semi-offline policy reinforcement learning method for slow thinking reasoning of visual language as described in Embodiment 1.
[0073] like Figure 5At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1 The method described herein. Of course, in addition to software implementation, this invention does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.
[0074] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0075] Example 3
[0076] This embodiment provides a computer-readable storage medium including one or more programs executable by one or more processors of an electronic device, the one or more programs including instructions for performing a semi-offline policy reinforcement learning method for slow-thinking reasoning of visual language as described in Embodiment 1.
[0077] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0078] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language, characterized in that, Includes the following steps: Obtain a multimodal training dataset consisting of multiple samples, where each sample includes image information, text question information, and standard answer information; Using the image information of the samples as input to a trainable visual language model, the input images are perceived and represented using an online strategy to extract the corresponding visual understanding information. The visual understanding information is concatenated with the textual question information of the sample to form a composite input that includes an image structure description and the requirements of the task to be answered. Using a large inference language model, a multi-step logical reasoning trajectory is generated in an offline strategy. Based on the logical reasoning trajectory, a reward signal based on the result is calculated. By sampling multiple logical reasoning trajectories, the trajectory reward for obtaining the correct answer is evenly distributed in the visual description layer as a reward label for visual understanding information. Based on visual understanding information with reward annotations, logical reasoning trajectories, and their reward signals, policy gradient optimization is performed on the visual language model to achieve semi-offline policy reinforcement learning. The visual understanding information includes visual descriptions of the overall structure, local features, object attributes, and their interrelationships of an image. The aforementioned logical reasoning process includes restating and understanding the problem, logical reasoning and formula derivation of partial conclusions, correction of errors, and the final answer.
2. The semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language according to claim 1, characterized in that, The process of calculating the reward signal based on the result according to the logical reasoning trajectory includes the following steps: Regular expressions are used to extract the location of the answer from the logical reasoning trajectory, and the equivalence of mathematical expressions is used to evaluate the difference between the logical reasoning trajectory and the standard answer. If the reasoning trajectory is consistent with the standard answer, the reward is assigned as 1; otherwise, it is assigned as 0.
3. The semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language according to claim 1, characterized in that, The method of evenly distributing the trajectory reward for obtaining the correct answer in the visual description layer, and using the reward annotation as visual understanding information, includes the following steps: Multiple logical reasoning trajectories are sampled based on visual understanding information; Calculate the reward signal corresponding to each logical reasoning trajectory; The average value of the reward signal corresponding to each logical reasoning trajectory is calculated and used as the reward for visual understanding information.
4. The semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language according to claim 1, characterized in that, The process of performing policy gradient optimization on the visual language model includes the following steps: Based on multiple logical reasoning trajectories corresponding to visual understanding information, and the reward signals corresponding to the logical reasoning trajectories, samples are filtered using a preset visual reward threshold to construct an offline strategy dataset. Based on the offline policy dataset, the visual language model is optimized for policy gradient using a policy gradient optimization algorithm with reward signals.
5. The semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language according to claim 1, characterized in that, The multimodal training dataset includes samples from the disciplines of advanced mathematics, physics, engineering, and financial management.
6. The semi-offline policy reinforcement learning method for slow-thinking reasoning in visual language according to claim 1, characterized in that, During the policy gradient optimization process of the visual language model, the visual encoder weights are frozen.
7. An electronic device, characterized in that, include: One or more processors and a memory, wherein the memory stores one or more programs, the one or more programs including instructions for executing a semi-offline policy reinforcement learning method for slow-thinking reasoning of visual language as described in any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, Includes one or more programs that are executed by one or more processors of an electronic device, said one or more programs including instructions for performing a semi-offline policy reinforcement learning method for slow-thinking reasoning of visual language as described in any one of claims 1-6.