Task processing method and device based on multi-source visual reasoning
By introducing single-source reward as a dynamic anchor point in multi-source visual reasoning, quantifying information gain, and adaptively adjusting the contribution weight of visual sources, the problem of performance fluctuation in multimodal scenarios is solved, achieving more stable and efficient multi-source visual reasoning, and improving the security and robustness of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 启元实验室
- Filing Date
- 2026-04-23
- Publication Date
- 2026-05-29
AI Technical Summary
Existing multi-source visual reasoning technologies lack a quantitative evaluation mechanism for information gain and interference in multimodal scenarios, leading to performance fluctuations and affecting the security and robustness of the system.
By constructing a multi-source visual dataset, extracting multi-source visual samples, using single-source rewards as dynamic anchors, calculating multi-source rewards and single-source rewards, combining a mixed advantage function and a verifiable reward function for reinforcement learning, adjusting the normalized statistics of the multi-source advantage function, establishing an explicit quantification mechanism, and adaptively adjusting the contribution weights of different visual sources.
It achieves more stable and efficient optimization in multi-source visual reasoning, improves the system's reliability and generalization ability, ensures noise suppression in modal conflict, and enhances synergy in modal complementarity.
Smart Images

Figure CN122116076A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, specifically to a task processing method and apparatus based on multi-source visual reasoning. Background Technology
[0002] Currently, multi-source visual reasoning technologies are mainly based on reinforcement learning and verifiable reward (RLVR) frameworks. While they perform well in single-modal tasks, they face significant challenges in multi-source input scenarios (such as infrared, depth, multi-view, and text). These modalities differ fundamentally in physical characteristics and semantic representation, posing a fundamental challenge to existing fusion reasoning methods.
[0003] Existing methods generally involve simply overlaying data from multiple sources, lacking a quantitative evaluation mechanism for information gain and interference. This often leads to unpredictable performance fluctuations in actual deployments, severely impacting the security and robustness of the system. Summary of the Invention
[0004] Based on this, this application provides a task processing method and apparatus based on multi-source visual reasoning, which achieves multi-source visual reasoning with good security and robustness.
[0005] According to one aspect of this application, a task processing method based on multi-source visual reasoning is proposed, comprising: constructing a multi-source visual dataset based on multi-source visual samples, wherein the multi-source visual samples include task samples, corresponding multiple visual data samples, and labeled task processing results, and the multiple visual data samples come from multiple visual sources; extracting the current multi-source visual sample from the multi-source visual dataset; inputting the current multi-source visual sample into a pre-constructed initial task processing model to generate a reasoning trajectory and a multi-source joint prediction result, and calculating the multi-source reward and multiple single-source rewards for the current multi-source visual sample, wherein the multi-source reward represents the multiple visual sources of the current multi-source visual sample. The reward value of the multi-source joint inference trajectory generated jointly by multiple visual data samples is the model output corresponding to the multi-source joint inference trajectory. The single-source reward is the reward value of the single-source inference trajectory generated by the visual data sample corresponding to each visual source of the current multi-source visual sample. Based on the multi-source reward, multiple single-source rewards, multi-source joint prediction results, and labeled task processing results, reinforcement learning is performed on the initial task processing model based on the pre-set hybrid advantage function and the pre-set verifiable reward function to obtain the target multi-source visual inference model. The task to be processed is input into the target multi-source visual inference model to obtain the multi-source joint prediction result.
[0006] According to some embodiments, the current multi-source visual sample is input into a pre-built initial task processing model to generate inference trajectories and multi-source joint prediction results, and multi-source rewards and multiple single-source rewards for the current multi-source visual sample are calculated. This includes: inputting the current multi-source visual sample into the pre-built initial task processing model, and using multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample to jointly generate multi-source joint inference trajectories and multi-source joint prediction results; calculating multi-source rewards based on the multi-source joint inference trajectories and multi-source joint prediction results; inputting the current multi-source visual sample into the pre-built initial task processing model, and using visual data samples corresponding to each visual source of the current multi-source visual sample to generate multiple single-source inference trajectories and their corresponding single-source prediction results; calculating the single-source reward for each single-source inference trajectory to obtain multiple single-source rewards.
[0007] According to some embodiments, based on multi-source rewards, multiple single-source rewards, multi-source joint prediction results, and labeled task processing results, reinforcement learning is performed on an initial task processing model based on a pre-set mixture dominance function and a pre-set verifiable reward function to obtain a target multi-source visual inference model. This includes: calculating a verifiable reward value based on the multi-source joint prediction results, labeled task processing results, and the verifiable reward function; calculating a mixture statistic for multi-source visual samples based on the verifiable reward value, multi-source rewards, multiple single-source rewards, and the mixture dominance function; and performing reinforcement learning on the initial task processing model based on the mixture statistic and a pre-set reinforcement learning framework to obtain the target multi-source visual inference model.
[0008] According to some embodiments, the current multi-source visual samples are input into a pre-built initial task processing model, and multiple visual data samples corresponding to multiple visual sources of the current multi-source visual samples are used to jointly generate a multi-source joint inference trajectory and a multi-source joint prediction result. This includes: extracting a first text feature based on the task sample in the current multi-source visual samples; extracting a first visual feature based on the multiple visual data samples in the current multi-source visual samples; fusing the first visual feature and the first text feature to obtain a first fused feature; and inputting the first fused feature and a preset prompt word template into the initial task processing model to obtain the multi-source joint inference trajectory and the multi-source joint prediction result.
[0009] According to some embodiments, the current multi-source visual samples are input into a pre-built initial task processing model, and multiple single-source inference trajectories and their corresponding single-source prediction results are generated using the visual data samples corresponding to each visual source of the current multi-source visual samples. The process includes: S1: Extracting second text features based on the task samples in the current multi-source visual samples; S2: Extracting the current visual data sample from multiple visual data samples in the current multi-source visual samples; S3: Extracting the visual features of the current visual data sample as the second visual features; S4: Fusing the second text features and the second visual features to obtain the second fused features; S5: Inputting the second fused features and a preset prompt word template into the initial task processing model to obtain the single-source inference trajectory; S6: Repeating steps S2-S5 until multiple visual data samples are traversed to obtain multiple single-source inference trajectories.
[0010] According to some embodiments, the annotation task processing results include visual localization results for visual localization tasks and / or visual question answering results for visual question answering tasks; the initial task processing model is pre-built through the following steps: extracting a first training dataset for visual localization tasks from a multi-source visual dataset; fine-tuning a pre-selected visual language model using the first training dataset to obtain the initial task processing model; and / or extracting a second training dataset for visual question answering tasks from a multi-source visual dataset; fine-tuning a pre-selected visual language model using the second training dataset to obtain the initial task processing model.
[0011] According to some embodiments, the verifiable reward function includes an intersection-union reward function and a format reward function; or the verifiable reward function includes an accuracy reward function and a format reward function.
[0012] According to one aspect of this application, a task processing apparatus based on multi-source visual reasoning includes: a multi-source data unit, configured to construct a multi-source visual dataset based on multi-source visual samples, wherein the multi-source visual samples include task samples, corresponding multiple visual data samples, and labeled task processing results, and the multiple visual data samples come from multiple visual sources; a current data unit, configured to extract current multi-source visual samples from the multi-source visual dataset; and a reward calculation unit, configured to input the current multi-source visual samples into a pre-constructed initial task processing model, generate inference trajectories and multi-source joint prediction results, and calculate the multi-source reward and multiple single-source rewards for the current multi-source visual samples, wherein the multi-source reward is the reward for the current multi-source visual samples. The reward value of the multi-source joint inference trajectory is generated jointly by multiple visual data samples corresponding to a visual source. The multi-source joint prediction result is the model output corresponding to the multi-source joint inference trajectory. The single-source reward is the reward value of the single-source inference trajectory generated by the visual data sample corresponding to each visual source in the current multi-source visual sample. The reinforcement learning unit is used to perform reinforcement learning on the initial task processing model based on the multi-source reward, multiple single-source rewards, multi-source joint prediction result, and labeled task processing result, and based on the pre-set mixture advantage function and the pre-set verifiable reward function to obtain the target multi-source visual inference model. The model inference unit is used to input the task to be processed into the target multi-source visual inference model to obtain the multi-source joint prediction result.
[0013] According to one aspect of this application, an electronic device is provided, comprising: one or more processors; a storage device for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the method as described above.
[0014] According to one aspect of this application, a computer-readable medium is provided that stores a computer program or instructions thereon, which, when executed by a processor, implement the method as described above.
[0015] Through the embodiments provided in this application, during the training phase, multi-source visual input is received as multi-source visual samples to construct a multi-source visual dataset; current multi-source visual samples are extracted from the multi-source visual dataset; for the current multi-source visual samples, single-source inference trajectories and multi-source joint inference trajectories are generated respectively through the same initial task processing model, and multiple corresponding single-source rewards and multi-source rewards are calculated simultaneously; reinforcement learning is performed on the initial task processing model using multiple single-source rewards, multi-source rewards, verifiable reward functions, and mixed advantage functions. During the reinforcement learning process, single-source rewards are used as dynamic anchors to adjust the normalized statistics of the multi-source advantage function to obtain the target multi-source visual inference model. The trained model can adaptively adjust the contribution weights of different visual sources, effectively quantifying the information gain from single-source to multi-source inference, thereby achieving more stable and efficient optimization in reinforcement learning training; during the application phase, the task to be processed is input into the trained target multi-source visual inference model to obtain multi-source joint prediction results, achieving multi-source visual inference with good security and robustness. Attached Figure Description
[0016] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application.
[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings, without exceeding the scope of protection claimed by this application.
[0018] Figure 1 A flowchart of a task processing method based on multi-source visual reasoning provided in an embodiment of this application;
[0019] Figure 2 The flowchart of the task processing method based on multi-source visual reasoning provided in the embodiments of this application is as follows: inputting the current multi-source visual sample into the pre-constructed initial task processing model, generating the reasoning trajectory and multi-source joint prediction result, and calculating the multi-source reward and multiple single-source rewards of the current multi-source visual sample. Figure 3 A flowchart illustrating the reinforcement learning process of the initial task processing model for the task processing method based on multi-source visual reasoning provided in the embodiments of this application; Figure 4 A schematic diagram illustrating the integration process of single-source rewards and multi-source rewards in a hybrid distribution in the task processing method based on multi-source visual reasoning provided in the embodiments of this application; Figure 5A schematic diagram illustrating the construction process of the initial task processing model for the task processing method based on multi-source visual reasoning provided in this application embodiment; Figure 6 A block diagram of a task processing device based on multi-source visual reasoning provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0022] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0023] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0024] It should be understood that although the terms first, second, third, etc., may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component discussed below may be referred to as the second component without departing from the teachings of this application. As used herein, the term "and / or" includes all combinations of any one and more of the associated listed items.
[0025] Existing multi-source visual reasoning methods generally simply superimpose multi-source data, lacking a quantitative evaluation mechanism for information gain and interference, leading to the following core problems: 1) Modal conflicts are difficult to reconcile: Existing technologies generally employ simple data overlay strategies when processing multi-source data, lacking a scientific evaluation mechanism for information gain and interference effects. When a certain modality contains key signals for the task, the introduction of other modalities can have a negative impact. Taking the nighttime pedestrian detection task in the LLVIP dataset as an example, infrared images can clearly present thermal radiation information, while RGB images generate a lot of noise under low light conditions. Traditional multi-source fusion methods often fail to effectively identify the dominant modality, thus reducing performance and even falling below the results of single-source inference.
[0026] 2) Lack of Dynamic Interaction Modeling: The normalization process of existing RLVR frameworks relies solely on the statistical characteristics of multi-source trajectories, completely ignoring the crucial reference information provided by single-source inference results. This leads to blind and unstable gradient update directions. This blindness in optimization is directly reflected in the drastic fluctuations during training, resulting in poor model convergence stability and significantly reduced reliability of inference results. The lack of single-source benchmarks makes multi-source optimization lack a clear direction, making it difficult to establish an effective collaborative mechanism between modalities.
[0027] 3) Limited Generalization Ability: In high-risk scenarios such as autonomous driving and medical diagnosis, the reliability of inference results is often crucial. However, due to the aforementioned shortcomings, existing multi-source methods frequently experience unpredictable performance fluctuations in practical deployments. For example, in autonomous driving scenarios under adverse weather conditions, the reliability of different sensor modalities changes dynamically, but traditional methods lack the ability to evaluate modal quality in real time and cannot adjust the fusion strategy according to environmental changes, severely impacting the system's safety and robustness.
[0028] Based on this, this application proposes a task processing method and apparatus based on multi-source visual reasoning.
[0029] For specific implementation details, please refer to the following examples.
[0030] Figure 1 A flowchart illustrating a task processing method based on multi-source visual reasoning provided in an embodiment of this application. Figure 1 As shown, the method includes steps S110-S150.
[0031] In step S110, a multi-source visual dataset is constructed based on the multi-source visual samples. The multi-source visual samples include task samples, corresponding multiple visual data samples, and labeled task processing results. The multiple visual data samples come from multiple visual sources.
[0032] Multi-source visual samples include task samples (e.g., question q), multiple visual data samples corresponding to the task samples, and the results of the labeled task processing.
[0033] Multiple visual data samples come from multiple visual sources, such as RGB, infrared, and depth.
[0034] According to the example embodiment, multiple visual data samples are represented in the form of multiple visual sources i. At the same time, a separate data loading pipeline is built for each source to ensure that different modalities can be accessed individually or jointly to obtain the corresponding visual data samples.
[0035] The image data in the multi-source visual samples can be acquired directly through multi-source visual acquisition devices or collected from existing resources, and this application does not impose any restrictions on this.
[0036] Accordingly, the task samples and annotation task processing results in the multi-source visual samples can be collected from existing resources or annotated using expert experience and knowledge, and this application does not impose any restrictions on this.
[0037] The collected multi-source visual samples constitute a multi-source visual dataset.
[0038] In step S120, the current multi-source visual sample is extracted from the multi-source visual dataset.
[0039] According to the example embodiment, a multi-source visual sample is selected from the multi-source visual dataset as the current multi-source visual sample.
[0040] In step S130, the current multi-source visual sample is input into the pre-built initial task processing model to generate inference trajectories and multi-source joint prediction results. The multi-source reward and multiple single-source rewards for the current multi-source visual sample are calculated. The multi-source reward is the reward value of the multi-source joint inference trajectory jointly generated by multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample. The multi-source joint prediction result is the model output corresponding to the multi-source joint inference trajectory. The single-source reward is the reward value of the single-source inference trajectory generated by the visual data samples corresponding to each visual source of the current multi-source visual sample.
[0041] The inference trajectory (denoted as the multi-source joint inference trajectory) is jointly generated using multiple visual data samples corresponding to all visual sources in the current multi-source visual sample to obtain the multi-source reward. For each visual data sample in the current multi-source visual sample, an inference trajectory (denoted as the single-source inference trajectory) and the corresponding single-source reward are generated, thereby obtaining the single-source inference trajectory and its single-source reward for each visual data sample.
[0042] In step S140, based on multi-source rewards, multiple single-source rewards, multi-source joint prediction results, and labeled task processing results, reinforcement learning is performed on the initial task processing model based on a pre-set hybrid advantage function and a pre-set verifiable reward function to obtain the target multi-source visual reasoning model.
[0043] Based on the pre-set hybrid advantage function and verifiable reward function, the initial task processing model is reinforced by using the multi-source rewards and multiple single-source rewards corresponding to the current multi-source visual samples to obtain the target multi-source visual inference model.
[0044] According to the example implementation, repeat steps S120-S140 until the training termination condition is met, and output the training result of the last iteration as the target multi-source visual reasoning model.
[0045] The conditions for ending the training can be set according to the actual situation, and this application does not impose any restrictions on this.
[0046] Reinforcement learning can be implemented based on existing reinforcement learning frameworks. This application does not impose specific restrictions on reinforcement learning frameworks, and the appropriate framework can be selected as needed.
[0047] The initial task processing model can be selected according to the actual situation, such as Qwen2.5-VL or similar multimodal visual language models, or it can be a model fine-tuned based on multimodal visual language models for tasks related to multi-source visual datasets.
[0048] This application establishes an explicit quantification mechanism for multi-source information gain by introducing single-source rewards as dynamic anchors, providing an interpretable and adaptive optimization paradigm for multimodal reasoning. It can dynamically adjust the fusion strategy according to the actual contributions between modalities, significantly enhancing the reliability and generalization ability of the system while ensuring performance improvement, thus providing an interpretable and adaptive optimization paradigm for multimodal reasoning.
[0049] Specifically, by using single-source rewards as anchors, the differences between trajectories are precisely utilized, thereby improving performance through precise information gain from single-source rewards to multi-source rewards.
[0050] In step S150, the task to be processed is input into the target multi-source visual reasoning model to obtain the multi-source joint prediction result.
[0051] During the usage phase, the task to be processed is input into the target multi-source visual reasoning model, and the model outputs the multi-source joint prediction results.
[0052] The tasks to be processed are related to multi-source visual datasets, such as visual localization tasks and visual question answering tasks.
[0053] In the training phase of this application, multi-source visual input is received as multi-source visual samples to construct a multi-source visual dataset. Current multi-source visual samples are extracted from the dataset. For each current multi-source visual sample, a single-source inference trajectory and a multi-source joint inference trajectory are generated using the same initial task processing model, while simultaneously calculating multiple single-source and multi-source rewards. Reinforcement learning is performed on the initial task processing model using multiple single-source and multi-source rewards, as well as a verifiable reward function and a hybrid advantage function. During reinforcement learning, the single-source reward is used as a dynamic anchor point to adjust the normalized statistics of the multi-source advantage function, resulting in a target multi-source visual inference model. The trained model can adaptively adjust the contribution weights of different visual sources, effectively quantifying the information gain from single-source to multi-source inference, thereby achieving more stable and efficient optimization during reinforcement learning training. In the application phase, the task to be processed is input into the trained target multi-source visual inference model to obtain multi-source joint prediction results, achieving multi-source visual inference with good security and robustness.
[0054] According to some embodiments, refer to Figure 2 In step S130, the current multi-source visual sample is input into the pre-built initial task processing model to generate inference trajectory and multi-source joint prediction result, and the multi-source reward and multiple single-source rewards of the current multi-source visual sample are calculated. Specifically, this can be achieved through steps S210-S240.
[0055] In step S210, the current multi-source visual sample is input into the pre-built initial task processing model, and multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample are used to jointly generate multi-source joint inference trajectory and multi-source joint prediction result.
[0056] The current multi-source visual samples are input into the initial task processing model. During inference, multiple visual data samples corresponding to multiple visual sources of the multi-source visual samples are used for multi-source joint inference. The initial task processing model does not process the visual data samples of each visual source separately and then merge the results. Instead, it performs cross-source feature alignment, comparison and fusion internally to generate multi-source joint inference trajectories and multi-source joint prediction results.
[0057] In step S220, multi-source rewards are calculated based on the multi-source joint inference trajectory and the multi-source joint prediction results.
[0058] Based on multi-source empirical samples Calculate multi-source rewards .
[0059] .
[0060] in, This represents the sampling index in reinforcement learning. This represents the current capacity of the buffer in reinforcement learning. This includes multi-source joint inference trajectories and multi-source joint prediction results.
[0061] In step S230, the current multi-source visual samples are input into the pre-built initial task processing model, and multiple single-source inference trajectories and their corresponding single-source prediction results are generated using the visual data samples corresponding to each visual source of the current multi-source visual samples.
[0062] The current multi-source visual samples are input into the initial task processing model. During inference, multiple single-source joint inferences are performed using the visual data samples corresponding to each visual source of the multi-source visual samples. The initial task processing model processes the visual data samples of each visual source separately, generating single-source inference trajectories and single-source prediction results for each visual source.
[0063] In step S240, the single-source reward for each single-source inference trajectory is calculated, resulting in multiple single-source rewards.
[0064] For each single-source inference trajectory, based on single-source empirical samples Calculate single-source reward .
[0065] .
[0066] in, This represents the sampling index in reinforcement learning. This represents the current capacity of the buffer in reinforcement learning. This includes the single-source inference trajectory and single-source prediction results of the currently calculated visual source.
[0067] According to some embodiments, refer to Figure 3 In step S140, based on multi-source rewards, multiple single-source rewards, multi-source joint prediction results, and labeled task processing results, reinforcement learning is performed on the initial task processing model based on a pre-set hybrid advantage function and a pre-set verifiable reward function to obtain the target multi-source visual reasoning model. This can be specifically achieved through steps S310-S330.
[0068] In step S310, a verifiable reward value is calculated based on the multi-source joint prediction results, the labeled task processing results, and the verifiable reward function.
[0069] Calculate verifiable reward value r ( q,i ).
[0070] in, q This represents the task sample of the current multi-source visual samples. iThis represents multiple visual data samples corresponding to multiple visual sources.
[0071] Verifiable rewards are a key component of reinforcement learning, used to adjust the model’s preferences and check whether the content and format of the multi-source joint prediction results match the correct answer (i.e., the labeled task processing results).
[0072] In step S320, the mixture statistics of the multi-source visual samples are calculated based on the verifiable reward value, multi-source reward, multiple single-source rewards, and mixture dominance function.
[0073] Single-source rewards and multi-source rewards Combined calculation of mixed statistics Used for dominance function normalization: .
[0074] in, The function representing the arithmetic mean. This represents the standard deviation function.
[0075] In step S330, based on the mixed statistics and a preset reinforcement learning framework, reinforcement learning is performed on the initial task processing model to obtain the target multi-source visual reasoning model.
[0076] The policy gradient optimization is performed. Based on the preset reinforcement learning framework, the policy gradient of the initial task processing model is updated using the hybrid advantage function to achieve optimization of information gain perception. The reinforcement learning result is denoted as the target multi-source visual reasoning model.
[0077] Specifically, such as Figure 4 As shown, Figure 4 This demonstrates the integration process of single-source and multi-source rewards in a mixed distribution. Single-source trajectories do not directly participate in policy updates but indirectly guide the optimization direction by influencing the calculation of the mean and standard deviation. When multi-source rewards outperform the single-source benchmark, the model is encouraged to explore beyond single-source cues; conversely, when modal conflicts exist, the multi-source learning process is regularized, causing it to tend towards more reliable single-source behavior.
[0078] In this application's reinforcement learning process, single-source rewards are used as dynamic anchors to adjust the normalized statistics of the multi-source advantage function. This design enables the model to adaptively adjust the contribution weights of different visual sources, effectively quantifying the information gain from single-source to multi-source inference, thereby achieving more stable and efficient optimization during reinforcement learning training. Through unbiased gradient estimation and information gain regularization, optimization stability is ensured. This theoretical framework provides an interpretable optimization path for multi-source fusion, ensuring that the model automatically suppresses noise during modal conflicts and enhances synergistic effects during modal complementarity.
[0079] According to some embodiments, in step S210, the current multi-source visual sample is input into the pre-built initial task processing model, and multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample are used to jointly generate multi-source joint inference trajectory and multi-source joint prediction result, which can be specifically implemented through steps S211-S214.
[0080] In step S211, the first text features are extracted based on the task samples in the current multi-source visual samples.
[0081] Extract the text features from the task samples and use them as the first text features.
[0082] According to the example embodiment, this step can be implemented using a text encoder.
[0083] In step S212, the first visual feature is extracted based on multiple visual data samples in the current multi-source visual samples.
[0084] Visual features are extracted from multiple visual data samples and used as the first visual feature.
[0085] According to the example embodiment, this step can be implemented using a visual encoder.
[0086] In step S213, the first visual feature and the first text feature are fused to obtain the first fused feature.
[0087] This application does not restrict the method of feature fusion, and the method can be selected according to the actual situation.
[0088] According to the example implementation, the fusion of the first visual feature and the first text feature can be achieved by concatenation, and the fusion result is denoted as the first fused feature. The fusion of multiple features within the first visual feature can be performed using a cross-attention mechanism or by concatenation.
[0089] In step S214, the first fusion feature and the preset prompt word template are input into the initial task processing model to obtain the multi-source joint inference trajectory and the multi-source joint prediction result.
[0090] The preset prompt word templates include prompt word templates that guide the generation of reasoning trajectories.
[0091] The first fused feature and the preset prompt word template are input into the initial task processing model. The model generates multi-source joint inference trajectory and multi-source joint prediction result based on the input data.
[0092] According to some embodiments, in step S230, the current multi-source visual sample is input into the pre-built initial task processing model, and multiple single-source inference trajectories and their corresponding single-source prediction results are generated using the visual data samples corresponding to each visual source of the current multi-source visual sample. Specifically, this can be achieved through steps S1-S6.
[0093] S1: Extract the second text features based on the task samples in the current multi-source visual samples.
[0094] Extract the text features from the task samples as the second text features.
[0095] According to the example embodiment, this step can be implemented using a text encoder.
[0096] S2: Extract the current visual data sample from multiple visual data samples in the current multi-source visual sample.
[0097] According to an example embodiment, one of a plurality of visual data samples is selected as the current visual data sample.
[0098] S3: Extract the visual features of the current visual data sample as the second visual feature.
[0099] Extract the visual features of the current visual data sample as the second visual feature.
[0100] According to the example embodiment, this step can be implemented using a visual encoder.
[0101] S4: Combine the second text feature and the second visual feature to obtain the second fused feature.
[0102] This application does not restrict the method of feature fusion, and the method can be selected according to the actual situation.
[0103] According to the example implementation, the fusion of the second visual feature and the third text feature can be achieved by concatenation, and the fusion result is denoted as the second fused feature.
[0104] S5: Input the second fusion feature and the preset prompt word template into the initial task processing model to obtain the single-source inference trajectory.
[0105] The preset prompt word templates include prompt word templates that guide the generation of reasoning trajectories.
[0106] The fused second fusion feature and the preset prompt word template are input into the initial task processing model. The model generates a single-source inference trajectory based on the input data. The model output also includes the single-source joint prediction result.
[0107] S6: Repeat steps S2-S5 until multiple visual data samples are traversed to obtain multiple single-source inference trajectories.
[0108] According to some embodiments, the annotation task processing results include visual positioning results for visual positioning tasks and / or visual question answering results for visual question answering tasks.
[0109] Based on the above embodiments, referring to Figure 5 In step S120, the initial task processing model can be pre-constructed through steps S510-S520 and / or steps S530-S540.
[0110] In step S510, a first training dataset for the visual localization task is extracted from the multi-source visual dataset.
[0111] The first training dataset is formed by extracting multi-source visual samples from the annotation task processing results in the multi-source visual dataset to form the visual localization results for the visual localization task.
[0112] In step S520, the pre-selected visual language model is fine-tuned using the first training dataset to obtain the initial task processing model.
[0113] The selected multimodal visual language model was fine-tuned using the first training dataset to enable it to acquire visual localization capabilities. The resulting model is denoted as the initial task processing model.
[0114] In step S530, a second training dataset for the visual question answering task is extracted from the multi-source visual dataset.
[0115] Extract multi-source visual samples from the multi-source visual dataset to form the visual question answering results for the visual question answering task, and construct the second training dataset.
[0116] In step S540, the pre-selected visual language model is fine-tuned using the second training dataset to obtain the initial task processing model.
[0117] The selected multimodal visual language model was fine-tuned using the second training dataset to enable it to perform visual question answering. The resulting model is denoted as the initial task processing model.
[0118] According to some embodiments, the reward functions that can be verified include the intersection-union ratio reward function and the format reward function.
[0119] For visual localization tasks, verifiable reward functions include the intersection-union reward function and the format reward function.
[0120] According to the example embodiment, it can be verified that the reward function is the sum of the intersection-union ratio reward function and the format reward function.
[0121] According to some embodiments, the reward functions can be verified to include an accuracy reward function and a format reward function.
[0122] For visual question answering tasks, verifiable reward functions include accuracy reward functions and format reward functions.
[0123] According to the example implementation, it can be verified that the reward function is the sum of the accuracy reward function and the format reward function.
[0124] According to experiments, the task processing method based on multi-source visual reasoning provided in this application achieves an average performance improvement of 3.2% and 4.9% under the GRPO and DAPO frameworks, respectively, with multi-source reasoning performance approaching or exceeding the single-source limit. The method provided in this application only requires adding single-trajectory sampling, with negligible increase in computational overhead, outperforming traditional multi-trajectory methods and significantly reducing the complexity and technical risks of engineering implementation. It demonstrates excellent cross-modal generalization ability, applicable to multiple modalities such as infrared, depth, multi-view, and text, without requiring adjustment of the model structure. This generalization stems from the core design of the method, which does not rely on specific modal features or model structures, but rather achieves adaptive adjustment between modalities through a single-source anchoring mechanism.
[0125] To further illustrate the task processing method based on multi-source visual reasoning provided in this application, two specific embodiments are given.
[0126] In one specific embodiment, image data from an RGB camera and an infrared sensor are simultaneously received and transmitted to a preprocessing unit for size normalization and channel alignment, serving as the visual data sample for the current multi-source visual samples. Simultaneously, pre-trained Qwen2.5-VL-3B weights are loaded to construct a two-stream feature extraction network, which serves as the initial task processing model.
[0127] The implementation steps include: 1) Single-source anchor point generation stage: RGB images and infrared images are input into the model separately as current visual data samples to generate corresponding single-source inference trajectories. At this time, RGB images are difficult to identify distant pedestrians under low light conditions (reward value 0.32), while infrared images clearly show the thermal radiation characteristics of pedestrians (reward value 0.90).
[0128] 2) Multi-source fusion optimization stage: The dual-modal images are jointly input, with an initial multi-source reward of 0.74. The mixture statistics are calculated using single-source anchor advantage normalization.
[0129] 3) Gradient correction process: When the information gain ΔIG = 0.74 - 0.90 = -0.16 < 0 is detected, the conflict suppression mechanism is activated, and the gradient update weight of the RGB mode is reduced.
[0130] 4) Results output: The traditional GRPO method suffers from missed detections due to its over-reliance on RGB images, while the method provided in this application accurately focuses on the pedestrian area in the infrared image by using anchor points, thus improving the detection box positioning accuracy.
[0131] In this embodiment, the traditional GRPO method relies excessively on RGB images, resulting in its inability to accurately detect clearly visible human targets in infrared images under low-light conditions, and the predicted bounding boxes exhibit significant bias. In contrast, the method provided in this application, through a single-source anchoring mechanism, can adaptively focus on key information in infrared images, generating accurate human detection boxes.
[0132] In another embodiment, an RGB-D camera is used to simultaneously acquire color images and depth information, which are then transmitted via a multimodal data bus.
[0133] The implementation steps include: 1) Modal characteristic analysis: Depth images provide spatial geometric information (reward 0.60), while RGB images provide apparent texture features (reward 0.85). Single-source anchor points play a dominant role in identifying RGB modalities.
[0134] 2) Multi-source fusion optimization stage: The dual-modal images are jointly input, with an initial multi-source reward of 0.80. The mixture statistics are calculated using single-source anchor advantage normalization.
[0135] 3) Adaptive weight allocation: Calculate the information gain ΔIG = 0.80 - 0.85 = -0.05, trigger the regularization term to moderately suppress deep features, and increase the weight of RGB features.
[0136] 4) Visualization of the reasoning process: When answering the question of "object distance", traditional methods are affected by depth and produce incorrect inferences. However, the model attention of the method provided in this application is shifted from the depth background region to the spatial relationship features in the RGB image, and the RGB information is correctly used to generate accurate answers, which verifies the effectiveness of the anchor point mechanism.
[0137] In this embodiment, the traditional GRPO model, due to its over-reliance on depth information, produces incorrect answers due to insufficient understanding of spatial relationships. In contrast, the method provided in this application explicitly models multi-source information gain, fully utilizing the geometric cues provided by RGB images to correctly infer the spatial relationships between objects. These qualitative results strongly validate the advantages of the method provided in this application when processing multi-source data with significantly different modal characteristics, demonstrating its ability to identify key information sources and its anti-interference performance in practical applications.
[0138] The following describes an apparatus embodiment of this application, which can be used to perform the method embodiment of this application. For details not disclosed in the apparatus embodiment of this application, please refer to the method embodiment of this application.
[0139] Figure 6 A block diagram of a task processing apparatus based on multi-source visual reasoning according to an exemplary embodiment is shown.
[0140] Figure 6 The apparatus shown can perform the task processing method based on multi-source visual reasoning according to the embodiments of this application.
[0141] like Figure 6 As shown, a task processing device based on multi-source visual reasoning may include: See Figure 6 Referring to the preceding description, the multi-source data unit 610 is used to construct a multi-source visual dataset based on multi-source visual samples. The multi-source visual samples include task samples, corresponding multiple visual data samples, and labeled task processing results. The multiple visual data samples come from multiple visual sources.
[0142] The current data unit 620 is used to extract the current multi-source visual sample from the multi-source visual dataset.
[0143] The reward calculation unit 630 is used to input the current multi-source visual sample into the pre-built initial task processing model, generate inference trajectories and multi-source joint prediction results, and calculate the multi-source reward and multiple single-source rewards for the current multi-source visual sample. The multi-source reward is the reward value of the multi-source joint inference trajectory jointly generated by multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample. The multi-source joint prediction result is the model output corresponding to the multi-source joint inference trajectory. The single-source reward is the reward value of the single-source inference trajectory generated by the visual data sample corresponding to each visual source of the current multi-source visual sample.
[0144] The reinforcement learning unit 640 is used to perform reinforcement learning on the initial task processing model based on multi-source rewards, multiple single-source rewards, multi-source joint prediction results, and labeled task processing results, and based on a pre-set hybrid advantage function and a pre-set verifiable reward function, to obtain the target multi-source visual reasoning model.
[0145] The model inference unit 650 is used to input the task to be processed into the target multi-source visual inference model to obtain the multi-source joint prediction result.
[0146] The device performs functions similar to those described above; other functions are described in the preceding descriptions and will not be repeated here.
[0147] This application discloses an electronic device, including: a processor; and a memory storing a computer program, which, when executed by the processor, causes the processor to execute the above-described instruction generation method.
[0148] For example, refer to Figure 7 , Figure 7 The illustrated electronic device 700 includes a processor 701 and a memory 703. The processor 701 and the memory 703 are connected, for example, via a bus 702. Optionally, the electronic device 700 may also include a transceiver 704. It should be noted that in practical applications, the transceiver 704 is not limited to one type, and the structure of this electronic device 700 does not constitute a limitation on the embodiments of this application.
[0149] Processor 701 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in this application. Processor 701 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0150] Bus 702 may include a pathway for transmitting information between the aforementioned components. Bus 702 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 702 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0151] The memory 703 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other storage medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0152] The memory 703 is used to store application code that executes the solution of this application, and its execution is controlled by the processor 701. The processor 701 is used to execute the application code stored in the memory 703 to implement the content shown in the foregoing method embodiments.
[0153] Figure 7 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0154] This application discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, causes the processor to execute an instruction generation method.
[0155] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0156] The above are only some embodiments of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A task processing method based on multi-source visual reasoning, characterized in that, include: A multi-source visual dataset is constructed based on multi-source visual samples, wherein the multi-source visual samples include task samples, multiple corresponding visual data samples, and labeled task processing results, and the multiple visual data samples come from multiple visual sources; Extract the current multi-source visual sample from the multi-source visual dataset; The current multi-source visual sample is input into a pre-built initial task processing model to generate an inference trajectory and a multi-source joint prediction result. The multi-source reward and multiple single-source rewards for the current multi-source visual sample are calculated. The multi-source reward is the reward value of the multi-source joint inference trajectory jointly generated by multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample. The multi-source joint prediction result is the model output corresponding to the multi-source joint inference trajectory. The single-source reward is the reward value of the single-source inference trajectory generated by the visual data sample corresponding to each visual source of the current multi-source visual sample. Based on the multi-source reward, the multiple single-source rewards, the multi-source joint prediction result, and the labeled task processing result, reinforcement learning is performed on the initial task processing model based on a pre-set hybrid advantage function and a pre-set verifiable reward function to obtain the target multi-source visual reasoning model. The task to be processed is input into the target multi-source visual reasoning model to obtain the multi-source joint prediction result.
2. The method according to claim 1, characterized in that, The current multi-source visual samples are input into a pre-built initial task processing model to generate inference trajectories and multi-source joint prediction results. The multi-source reward and multiple single-source rewards for the current multi-source visual samples are calculated, including: The current multi-source visual sample is input into the pre-built initial task processing model, and the multi-source joint inference trajectory and the multi-source joint prediction result are jointly generated using multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample. Calculate the multi-source reward based on the multi-source joint inference trajectory and the multi-source joint prediction result; The current multi-source visual samples are input into the pre-built initial task processing model, and multiple single-source inference trajectories and their corresponding single-source prediction results are generated using the visual data samples corresponding to each visual source of the current multi-source visual samples. Calculate the single-source reward for each of the single-source inference trajectories to obtain multiple single-source rewards.
3. The method according to claim 1, characterized in that, Based on the multi-source rewards, the multiple single-source rewards, the multi-source joint prediction results, and the labeled task processing results, reinforcement learning is performed on the initial task processing model based on a pre-set hybrid advantage function and a pre-set verifiable reward function to obtain a target multi-source visual reasoning model, including: Calculate the verifiable reward value based on the multi-source joint prediction results, the labeled task processing results, and the verifiable reward function; The mixture statistics of the multi-source visual samples are calculated based on the verifiable reward value, the multi-source reward, the multiple single-source rewards, and the mixture dominance function. Based on the mixed statistics, reinforcement learning is performed on the initial task processing model using a preset reinforcement learning framework to obtain the target multi-source visual reasoning model.
4. The method according to claim 2, characterized in that, The current multi-source visual samples are input into a pre-constructed initial task processing model, and multiple visual data samples corresponding to multiple visual sources of the current multi-source visual samples are used to jointly generate the multi-source joint inference trajectory and the multi-source joint prediction result, including: Extract the first text features based on the task samples in the current multi-source visual samples; Based on multiple visual data samples in the current multi-source visual samples, extract the first visual feature; The first visual feature and the first text feature are fused to obtain the first fused feature; The first fusion feature and the preset prompt word template are input into the initial task processing model to obtain the multi-source joint inference trajectory and the multi-source joint prediction result.
5. The method according to claim 2, characterized in that, The current multi-source visual samples are input into a pre-constructed initial task processing model, and multiple single-source inference trajectories and their corresponding single-source prediction results are generated using the visual data samples corresponding to each visual source of the current multi-source visual samples, including: S1: Extract the second text features based on the task samples in the current multi-source visual samples; S2: Extract the current visual data sample from multiple visual data samples in the current multi-source visual sample; S3: Extract the visual features of the current visual data sample as the second visual feature; S4: Fuse the second text feature and the second visual feature to obtain the second fused feature; S5: Input the second fusion feature and the preset prompt word template into the initial task processing model to obtain the single-source inference trajectory; S6: Repeat steps S2-S5 until the multiple visual data samples are traversed to obtain multiple single-source inference trajectories.
6. The method according to claim 1, characterized in that, The annotation task processing results include visual positioning results for visual positioning tasks and / or visual question answering results for visual question answering tasks. The initial task processing model is pre-built through the following steps: Extract a first training dataset for the visual localization task from the multi-source visual dataset; Using the first training dataset, the pre-selected visual language model is fine-tuned to obtain the initial task processing model; and / or Extract a second training dataset for the visual question answering task from the multi-source visual dataset; Using the second training dataset, the pre-selected visual language model is fine-tuned to obtain the initial task processing model.
7. The method according to claim 1, characterized in that, The verifiable reward function includes the intersection-union reward function and the format reward function; or The verifiable reward functions include an accuracy reward function and a format reward function.
8. A task processing device based on multi-source visual reasoning, characterized in that, include: A multi-source data unit is used to construct a multi-source visual dataset based on multi-source visual samples, wherein the multi-source visual samples include task samples, multiple corresponding visual data samples, and labeled task processing results, and the multiple visual data samples come from multiple visual sources. The current data unit is used to extract the current multi-source visual sample from the multi-source visual dataset; The reward calculation unit is used to input the current multi-source visual sample into a pre-built initial task processing model, generate an inference trajectory and a multi-source joint prediction result, and calculate the multi-source reward and multiple single-source rewards for the current multi-source visual sample. The multi-source reward is the reward value of the multi-source joint inference trajectory jointly generated by multiple visual data samples corresponding to multiple visual sources of the current multi-source visual sample. The multi-source joint prediction result is the model output corresponding to the multi-source joint inference trajectory. The single-source reward is the reward value of the single-source inference trajectory generated by the visual data samples corresponding to each visual source of the current multi-source visual sample. The reinforcement learning unit is used to perform reinforcement learning on the initial task processing model based on the multi-source reward, the multiple single-source reward, the multi-source joint prediction result, and the labeled task processing result, and on a pre-set hybrid advantage function and a pre-set verifiable reward function, to obtain the target multi-source visual reasoning model. The model inference unit is used to input the task to be processed into the target multi-source visual inference model to obtain the multi-source joint prediction result.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the method as described in any one of claims 1-7.