Visual language model target detection capability enhancement method

By constructing an inference-based object detection dataset and using an inference-based reinforcement learning method, the object detection capability of the visual language model is improved, solving the problems of insufficient detection performance and generalization ability in existing technologies, and achieving higher accuracy and faster object detection.

CN120976676APending Publication Date: 2025-11-18HONGLONG TECH (HANGZHOU) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511097208.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing visual language models suffer from significant performance gaps compared to dedicated models in object detection tasks, and traditional supervised fine-tuning methods fail to effectively utilize the model's reasoning and generalization capabilities.

Method used

By constructing an inference-based object detection dataset, and employing an inference-based reinforcement learning approach, we designed complex semantic labels and reward mechanisms, including format rewards and ODLength rewards. We then combined these with the GRPO algorithm to optimize the policy network parameters, thereby improving the model's object detection capabilities.

Benefits of technology

It significantly improves the target detection accuracy and generalization ability of visual language models in complex scenes, reduces redundant predictions, and enhances the adaptability and recognition accuracy of the models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976676A_ABST
    Figure CN120976676A_ABST
Patent Text Reader

Abstract

The invention discloses a visual language model target detection capability enhancing method, which comprises the following steps of: firstly, constructing an inference type target detection data set containing complex semantic tags such as attributes, interaction, orientation, negative and hard negative samples; and secondly, under a GRPO reinforcement learning framework, guiding the VLM to generate a reasoning process through a specific cue word, and then outputting a detection result. The present invention employs a composite reward function to evaluate a plurality of candidate outputs generated by the model, the function including a format reward that ensures that the output follows a preset thinking and answer structure, and an innovative ODLength reward. According to the ODLength reward, the average precision mean value is combined with a length penalty term, and redundant prediction is effectively restrained. And finally, updating the model strategy network according to the total reward value. According to the method, the target detection precision and generalization ability of the VLM in a complex reasoning scene can be remarkably improved, and the reasoning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, and in particular to a visual language model target detection capability enhancement method based on reasoning type reinforcement learning, which is suitable for improving the target detection precision and generalization capability of a visual language model (VLM) in a complex scene. BACKGROUND

[0002] In recent years, Vision-Language Models (VLM) have shown great potential in various tasks due to their strong cross-modal understanding capabilities. VLMs can simultaneously process and understand image and text information, making them highly concerned in complex visual tasks such as Open-Vocabulary Object Detection that require deep semantic understanding.

[0003] However, the VLMs in the prior art still face the following prominent technical problems when applied to the target detection task: (1) The performance gap between target detection and specialized models is significant: Although VLMs have the ability to understand complex text descriptions, their accuracy in visual grounding is significantly inferior to traditional models designed specifically for target detection tasks. For example, on the public COCO dataset, the mean average precision (mAP) of an advanced VLM (such as Qwen2.5-VL-3B) is only 14.2, while the performance of specialized detection models can be as high as 55.0. This huge performance gap limits the application of VLMs in practical scenarios that require high-precision detection.

[0004] (2) The limitations of traditional fine-tuning methods result in the failure to realize the potential of the model: Currently, the mainstream method to improve the target detection capability of VLMs is Supervised Fine-Tuning (SFT). This method relies on a large amount of "image-text-rectangle" labeled data, simplifying the training process to pattern fitting on labeled data. This approach has two main defects: Failure to effectively utilize the reasoning capabilities of the model: The SFT method tends to have the model learn direct mapping relationships, without effectively stimulating and utilizing the logical reasoning capabilities inherent in VLMs. Especially when dealing with detection tasks that require complex semantic understanding (such as attribute reasoning, spatial relationships, negative descriptions, etc.), the model only mechanically performs feature matching, rather than truly "thinking" and positioning, resulting in its inability to handle complex detection instructions.

[0005] Insufficient generalization ability and prone to overfitting: The SFT method is highly dependent on the quantity and quality of labeled data, which can easily lead to overfitting of the model on the training data and insufficient generalization ability. When faced with complex real-world scenes that are unseen or whose distribution is inconsistent with the training data, its detection performance will drop significantly.

[0006] In summary, how to effectively improve the target detection accuracy of VLM in complex scenarios and overcome the problems of insufficient inference ability and poor generalization ability caused by traditional supervised fine-tuning methods are technical challenges that urgently need to be solved in the field of artificial intelligence. Summary of the Invention

[0007] This invention primarily addresses the technical problems of insufficient target detection capability and limited fine-tuning effect in existing technologies, and provides a method for enhancing the target detection capability of visual language models based on inference-based reinforcement learning, which can effectively improve the target detection capability of visual language models.

[0008] The present invention addresses the aforementioned technical problems primarily through the following technical solution: a method for enhancing the object detection capability of a visual language model, comprising the following steps: S1: Data preparation: Collect image-text pairs, label instance-level bounding boxes and complex semantic descriptions, and generate an inference-based object detection dataset containing positive and negative samples; S2: Initialize the visual language model; S3: Add prompts for reasoning tasks, requiring the visual language model to first output the reasoning process, then output the answer, and output the target detection results in the answer in the specified JSON format; S4: Sample N candidate outputs; typically N is set to 8; the sampling method generates N data points based on the temperature parameter in the model being 0.9. S5: Calculate the format reward and ODLength reward, and calculate the total reward value; S6: Calculate the objective function value of GRPO based on the total reward value and update the policy network parameters.

[0009] Preferably, in step S1, the types of complex semantic tags annotated include: (1) Attribute reasoning: the target object in the image and its attributes; such as "a coffee cup made of metal, a black car"; (2) Multi-object interaction: The interaction relationship between several objects in an image; such as "a person riding a bicycle"; (3) Positional relationship: The relative or absolute positional relationship of the target object in the image; such as "the person on the left side of the picture, the mobile phone on the table"; (4) Negative description: The label contains a negative description of “not” or “no”, which points to certain objects in the image, such as “not a red car”; (5) Hard negative samples: objects that do not exist in the image; In the training dataset, objects that do not appear in the actual scene are deliberately added as hard negative samples. For example, in a set of indoor scene images, objects that are obviously not part of the scene, such as "spaceship", are added. In this way, the model can better learn the boundary conditions of the existence and non-existence of objects and improve its ability to accurately judge objects in real scenes.

[0010] Traditional object detection datasets typically only contain text labels at the object level, such as people and vehicles. These simple text labels are overly simplistic and lack the ability to detect objects in complex reasoning scenarios. They don't require complex reasoning to directly detect the corresponding object, thus failing to activate the reasoning advantages of multimodal large language models (VLMs) during reinforcement learning, thereby hindering the improvement of VLM's object detection performance. This solution designs a reasoning-based object detection dataset, constructing complex semantic target text labels from five dimensions of text semantics. This allows the model to learn to understand complex semantic target text labels through step-by-step reasoning before performing detection, thereby improving the model's object detection capabilities.

[0011] As a preferred option, the format reward for the i-th candidate output is calculated as follows: For the i-th candidate output, if the text output by the visual language model contains the inference and response start and end marks specified in the prompt words (such as "..."), <think>< / think> <answer>< / answer> If the visual language model satisfies the inference format, then the visual language model is determined to satisfy the inference format. For the i-th candidate output, if the object detection result in the answer part of the visual language model follows the JSON format specified in the prompt words, such as a list of JSON formats containing bbox_2d (boundary box coordinates) and label (object category): [ {"bbox_2d": [x1, y1, x2, y2], "label": "Target category 1"}, {"bbox_2d": [x3, y3, x4, y4], "label": "Target Category 2"}, ] Then it is determined that the visual language model meets the target detection result format; If the visual language model satisfies both the inference format and the object detection result format, the format reward for the i-th candidate output is 1 point; otherwise, it is 0 points.

[0012] As a preferred option, the ODLength reward R of the i-th candidate output is...i_OD Calculated in the following way: R i_OD =[min(1,L gt / L i_pred )]•mAP(b i_pred ,b gt ); In the formula, L gt L represents the number of real targets in the image, for each candidate output. gt They're all the same, L i_pred b is the number of targets predicted by the i-th candidate output of the visual language model. i_pred b is the bounding box predicted by the i-th candidate output of the visual language model. gt The predicted bounding box is the actual bounding box in the image; mAP is a function that calculates the average precision between the predicted bounding box and the actual bounding box.

[0013] Traditional mAP-based reward functions place too little emphasis on recall calculations, easily leading to redundant predictions and a "reward deception" phenomenon. This means that during training, the model can maximize the mAP reward by generating numerous redundant bounding boxes for all target objects in the graph, even without correctly detecting the required target labels. However, such models perform poorly in real-world applications. To address this issue and suppress redundant predictions, this solution introduces a length penalty term, defined as the ratio of the true number of targets to the predicted number of targets (not exceeding 1). When the predicted number of targets exceeds the true number (L... pred >L gt If the length penalty term is less than 1, a penalty is applied to redundant predictions; if the number of predictions is equal to or less than the actual number (L... pred ≤ L gt If the length penalty term is 1, then no penalty is applied.

[0014] Multiplying mAP by the length penalty term yields the ODLength reward function value (ranging from 0 to 1). Based on mAP, the length penalty term suppresses redundant predictions. For example, if the model generates a large number of irrelevant bounding boxes (L... pred >L gt Even if mAP is high, the final reward will be weakened due to the decrease in length penalty, thus avoiding the phenomenon of "reward cheating".

[0015] When calculating the mAP value between the predicted bounding box and the ground truth bounding box, the target bounding box b predicted by the model is first... pred Sort by confidence level in descending order, and calculate the relationship between each value and the true bounding box b. gtThe Intersection over Union (IoU) ratio is used to determine valid matches based on a preset IoU threshold (e.g., 0.5-0.95). True positives (TP) and false positives (FP) are counted, and a precision-recall (PR) curve is generated. The average precision (AP) for each class is calculated using 11-point interpolation or full integration. Finally, the average AP for all classes is taken to obtain the mean AP (mAP). Since conventional mAP calculations require a corresponding confidence level (range 0-1) for each bounding box in the model's prediction results, and the results generated by the VLM model do not include confidence levels, this invention uses 1.0 as the confidence level for each predicted bounding box when calculating the mAP value.

[0016] As a preferred option, the total reward value R of the i-th candidate output is... i This is the sum of the format reward and the ODLength reward for the i-th candidate output. The total reward value from N samples will be used together in subsequent calculations of advantage A. i And the gradient is updated to update the network parameters.

[0017] During each training step, the current model is used with the temperature set to 0.9 (as mentioned above) to generate N answers, or N candidate detection results, based on the question. Because the model exhibits relatively high randomness in outputting answers when the temperature is 0.9, the N output answers are all different. Then, GRPO needs to calculate the total reward value for each of the N answers, that is, to calculate the total reward value for each of the N answers sequentially.

[0018] As a preferred option, the update process for the policy network parameters is as follows: First, calculate the advantage value of each candidate output based on the N candidate outputs obtained from sampling. The advantage value of the i-th candidate output is denoted as A. i The formula for calculating the dominance value is: A i =(R i -mean{R1,R2,…,R N}) / std{R1,R2,…,R N}; Where R i Let be the total reward value of the i-th candidate output. This formula measures the relative quality of each candidate output compared to other outputs. mean is the average value, and std is the standard deviation. Next, the objective function of GRPO (Group Relative Policy Optimization) is calculated based on all the advantage values ​​to obtain the objective function value. Then, the gradient is calculated based on the objective function value to update the network parameters according to the optimization algorithm.

[0019] The objective function of GRPO is calculated as follows: ; s1=π θ (o i |q) / π θ_old (o i |q); s2=clip(π θ (o i |q) / π θ_old (o i |q),1+ε,1-ε); In the formula, q represents the input question (prompt) given to the model; o i Representation of the policy network π θ The i-th of the N candidate outputs generated for problem q; N represents the number of outputs from the policy network π. θ The total number of candidate outputs sampled for each question q; π θ This represents the policy network currently being optimized, with parameter θ; the policy generates a series of possible output responses o based on q. i ;π θ_old This represents the old policy network prior to this iteration update; π ref Let π represent a reference policy network; in the formula, π is calculated by... θ With reference policy network π ref The KL divergence between the two states can be penalized by applying a penalty term to prevent the current policy from deviating too far from the initial state, thereby stabilizing the training process; A i Indicates the i-th candidate output o i The advantage value; s1 is the probability ratio of importance sampling, used to correct for the distribution difference between the old and new policies; s2 is the clipped version of the probability ratio s1, whose value is restricted to the range of [1−ε,1+ε], in order to prevent the single-step policy update from being too large, thereby ensuring the stability of training; ε is a hyperparameter used to define the clipping range in s2, usually a small value (e.g., 0.2); β is the coefficient of the KL divergence penalty term, also a hyperparameter used to control the strength of the regularization term, usually 0-0.04; Indicates the current policy π θ The KL divergence between the current policy and the reference policy is used as a regularization term to penalize the difference between the current policy and the reference policy; Let be the mathematical expectation, representing the expectation for all strategies from the old policy π. θ_old The average of the output set obtained from the sampling is performed.

[0020] Compared to supervised fine-tuning (SFT) of conventional large models, the GRPO reinforcement learning algorithm differs significantly. SFT typically fine-tunes the model on large amounts of supervised data, aiming to minimize a predefined loss function to fit the labeled data patterns. In contrast, GRPO, through a reinforcement learning framework, dynamically generates multiple candidate outputs during training, optimizes the policy based on relative scores of the reward function (such as answer correctness and logical coherence), and maintains stability through KL divergence constraints. GRPO achieves policy generalization through within-group comparisons and dynamic reward signals, while SFT is limited by data coverage and label quality.

[0021] The substantial effects of this invention are: 1. Enhanced generalization ability: For object detection tasks and open object detection tasks requiring complex reasoning, the model exhibits stronger adaptability. Compared with other models, it can identify targets more accurately and stably, reducing false positives and false negatives. 2. Suppression of redundant predictions: Through the ODLength reward function, the model output length is reduced by 40%, and redundant predictions are reduced by 67%, thus maintaining higher accuracy while achieving faster inference speed. Attached Figure Description

[0022] Figure 1 This is a flowchart of a method for enhancing the object detection capability of a visual language model according to the present invention. Detailed Implementation

[0023] The technical solution of the present invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings.

[0024] Example: A method for enhancing the object detection capability of a visual language model, such as... Figure 1 As shown, it includes the following steps: S1: Data preparation: Collect image-text pairs, label instance-level bounding boxes and complex semantic descriptions, and generate an inference-based object detection dataset containing positive and negative samples; S2: Initialize the visual language model; S3: Add prompts for reasoning tasks, requiring the visual language model to first output the reasoning process, then output the answer, and output the target detection results in the answer in the specified JSON format; S4: Sample N candidate outputs; N is set to 8; the sampling method is to generate N data points based on the temperature parameter in the model being 0.9; S5: Calculate the format reward and ODLength reward, and calculate the total reward value; S6: Calculate the objective function value of GRPO based on the total reward value and update the policy network parameters.

[0025] In step S1, the types of complex semantic tags annotated include: (1) Attribute reasoning: the target object in the image and its attributes; such as "a coffee cup made of metal, a black car"; (2) Multi-object interaction: The interaction relationship between several objects in an image; such as "a person riding a bicycle"; (3) Positional relationship: The relative or absolute positional relationship of the target object in the image; such as "the person on the left side of the picture, the mobile phone on the table"; (4) Negative description: The label contains a negative description of “not” or “no”, which points to certain objects in the image, such as “not a red car”; (5) Hard negative samples: objects that do not exist in the image; In the training dataset, objects that do not appear in the actual scene are deliberately added as hard negative samples. For example, in a set of indoor scene images, objects that are obviously not part of the scene, such as "spaceship", are added. In this way, the model can better learn the boundary conditions of the existence and non-existence of objects and improve its ability to accurately judge objects in real scenes.

[0026] Traditional object detection datasets typically only contain text labels at the object level, such as people and vehicles. These simple text labels are overly simplistic and lack the ability to detect objects in complex reasoning scenarios. They don't require complex reasoning to directly detect the corresponding object, thus failing to activate the reasoning advantages of multimodal large language models (VLMs) during reinforcement learning, thereby hindering the improvement of VLM's object detection performance. This solution designs a reasoning-based object detection dataset, constructing complex semantic target text labels from five dimensions of text semantics. This allows the model to learn to understand complex semantic target text labels through step-by-step reasoning before performing detection, thereby improving the model's object detection capabilities.

[0027] The format reward for the i-th candidate output is calculated as follows: For the i-th candidate output, if the text output by the visual language model contains the inference and response start and end marks specified in the prompt words (such as "..."), <think>< / think> <answer>< / answer> If the visual language model satisfies the inference format, then the visual language model is determined to satisfy the inference format. For the i-th candidate output, if the object detection result in the answer part of the visual language model follows the JSON format specified in the prompt words, such as a list of JSON formats containing bbox_2d (boundary box coordinates) and label (object category): [ {"bbox_2d": [x1, y1, x2, y2], "label": "Target category 1"}, {"bbox_2d": [x3, y3, x4, y4], "label": "Target Category 2"}, ] Then it is determined that the visual language model meets the target detection result format; If the visual language model satisfies both the inference format and the object detection result format, the format reward for the i-th candidate output is 1 point; otherwise, it is 0 points.

[0028] The reward R for the ODLength of the i-th candidate output i_OD Calculated in the following way: R i_OD =[min(1,L gt / L i_pred )]•mAP(b i_pred ,b gt ); In the formula, L gt L represents the number of real targets in the image. i_pred b is the number of targets predicted by the i-th candidate output of the visual language model. i_pred b is the bounding box predicted by the i-th candidate output of the visual language model. gt The predicted bounding box is the actual bounding box in the image; mAP is a function that calculates the average precision between the predicted bounding box and the actual bounding box.

[0029] Traditional mAP-based reward functions place too little emphasis on recall calculations, easily leading to redundant predictions and a "reward deception" phenomenon. This means that during training, the model can maximize the mAP reward by generating numerous redundant bounding boxes for all target objects in the graph, even without correctly detecting the required target labels. However, such models perform poorly in real-world applications. To address this issue and suppress redundant predictions, this solution introduces a length penalty term, defined as the ratio of the true number of targets to the predicted number of targets (not exceeding 1). When the predicted number of targets exceeds the true number (L... pred >L gt If the length penalty term is less than 1, a penalty is applied to redundant predictions; if the number of predictions is equal to or less than the actual number (L... pred ≤ L gt If the length penalty term is 1, then no penalty is applied.

[0030] Multiplying mAP by the length penalty term yields the ODLength reward function value (ranging from 0 to 1). Based on mAP, the length penalty term suppresses redundant predictions. For example, if the model generates a large number of irrelevant bounding boxes (L... pred >L gt Even if mAP is high, the final reward will be weakened due to the decrease in length penalty, thus avoiding the phenomenon of "reward cheating".

[0031] When calculating the mAP value between the predicted bounding box and the ground truth bounding box, the target bounding box b predicted by the model is first... pred Sort by confidence level in descending order, and calculate the relationship between each value and the true bounding box b. gt The Intersection over Union (IoU) ratio is used to determine valid matches based on a preset IoU threshold (e.g., 0.5-0.95). True positives (TP) and false positives (FP) are counted, and a precision-recall (PR) curve is generated. The average precision (AP) for each class is calculated using 11-point interpolation or full integration. Finally, the average AP for all classes is taken to obtain the mean AP (mAP). Since conventional mAP calculations require a corresponding confidence level (range 0-1) for each bounding box in the model's prediction results, and the results generated by the VLM model do not include confidence levels, this invention uses 1.0 as the confidence level for each predicted bounding box when calculating the mAP value.

[0032] The total reward value R of the i-th candidate output i This is the sum of the format reward and the ODLength reward for the i-th candidate output. The total reward value from N samples will be used together in subsequent calculations of advantage A. i And the gradient is updated to update the network parameters.

[0033] The update process for policy network parameters is as follows: First, calculate the advantage value of each candidate output based on the sampled N (N=8) candidate outputs. The advantage value of the i-th candidate output is denoted as A. i The formula for calculating the dominance value is: A i =(R i -mean{R1,R2,…,R N}) / std{R1,R2,…,R N}; Where R i Let be the total reward value of the i-th candidate output. This formula measures the relative quality of each candidate output compared to other outputs. mean is the average value, and std is the standard deviation. Next, the objective function of GRPO is calculated based on all the advantage values ​​to obtain the objective function value. Then, the gradient is calculated based on the objective function value, and the network parameters are updated based on the optimization algorithm. The learning rate is 1e-6, and the training time is 100-3000 steps.

[0034] The objective function of GRPO is calculated as follows: ; s1=π θ (o i |q) / π θ_old (o i|q); s2=clip(π θ (o i |q) / π θ_old (o i |q),1+ε,1-ε); In the formula, q represents the input question (prompt) given to the model; o i Representation of the policy network π θ The i-th of the N candidate outputs generated for problem q; N represents the number of outputs from the policy network π. θ The total number of candidate outputs sampled for each question q; π θ This represents the policy network currently being optimized, with parameter θ; the policy generates a series of possible output responses o based on q. i ;π θ_old This represents the old policy network prior to this iteration update; π ref Let π represent a reference policy network; in the formula, π is calculated by... θ With reference policy network π ref The KL divergence between the two states can be penalized by applying a penalty term to prevent the current policy from deviating too far from the initial state, thereby stabilizing the training process; A i Indicates the i-th candidate output o i The advantage value; s1 is the probability ratio of importance sampling, used to correct for the distribution difference between the old and new policies; s2 is the clipped version of the probability ratio s1, whose value is restricted to the range of [1−ε,1+ε], in order to prevent the single-step policy update from being too large, thereby ensuring the stability of training; ε is a hyperparameter used to define the clipping range in s2, usually a small value (e.g., 0.2); β is the coefficient of the KL divergence penalty term, also a hyperparameter used to control the strength of the regularization term, usually 0-0.04; Indicates the current policy π θ The KL divergence between the current policy and the reference policy is used as a regularization term to penalize the difference between the current policy and the reference policy; Let represent the mathematical expectation, and let represent the expectation for all strategies from the old strategy π. θ_old The average of the output set obtained from the sampling is performed.

[0035] Compared to supervised fine-tuning (SFT) of conventional large models, the GRPO reinforcement learning algorithm differs significantly. SFT typically fine-tunes the model on large amounts of supervised data, aiming to minimize a predefined loss function to fit the labeled data patterns. In contrast, GRPO, through a reinforcement learning framework, dynamically generates multiple candidate outputs during training, optimizes the policy based on relative scores of the reward function (such as answer correctness and logical coherence), and maintains stability through KL divergence constraints. GRPO achieves policy generalization through within-group comparisons and dynamic reward signals, while SFT is limited by data coverage and label quality.

[0036] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

[0037] Although this document uses terms such as ODLength reward and format reward frequently, the possibility of using other terms is not excluded. These terms are used merely for the convenience of describing and explaining the essence of this invention; interpreting them as any additional limitation would contradict the spirit of this invention.

Claims

1. A method for enhancing the object detection capability of a visual language model, characterized in that, Includes the following steps: S1: Data preparation: Collect image-text pairs, label instance-level bounding boxes and complex semantic descriptions, and generate an inference-based object detection dataset containing positive and negative samples; S2: Initialize the visual language model; S3: Add prompts for reasoning tasks, requiring the visual language model to first output the reasoning process, then output the answer, and output the target detection results in the answer in the specified JSON format; S4: Sample N candidate outputs; S5: Calculate the format reward and ODLength reward, and calculate the total reward value; S6: Calculate the objective function value of GRPO based on the total reward value and update the policy network parameters.

2. The method for enhancing the target detection capability of a visual language model according to claim 1, characterized in that, In step S1, the types of complex semantic tags annotated include: (1) Attribute reasoning: the target object in the image and its attributes; (2) Multi-object interaction: The interaction relationship between several objects in an image; (3) Orientation: The relative or absolute positional relationship of the target object in the image; (4) Negative description: The label contains a negative description such as "not" or "no"; (5) Hard negative samples: objects that do not exist in the image.

3. The method for enhancing the target detection capability of a visual language model according to claim 1 or 2, characterized in that, The format reward for the i-th candidate output is calculated as follows: For the i-th candidate output, if the text output by the visual language model contains the inference and response start and end marks specified in the prompt words, then the visual language model is determined to satisfy the inference format. For the i-th candidate output, if the object detection result in the answer part of the visual language model follows the JSON format specified in the prompt words, then the visual language model is determined to satisfy the object detection result format. If the visual language model satisfies both the inference format and the object detection result format, the format reward for the i-th candidate output is 1 point; otherwise, it is 0 points.

4. A method for enhancing the target detection capability of a visual language model according to claim 1 or 2, characterized in that, The reward R for the ODLength of the i-th candidate output i_OD Calculated in the following way: R i_OD =[min(1,L gt / L i_pred )]•mAP(b i_pred ,b gt ): In the formula, L gt L represents the number of real targets in the image. i_pred b is the number of targets predicted by the i-th candidate output of the visual language model. i_pred b is the bounding box predicted by the i-th candidate output of the visual language model. gt It is the actual rectangular frame in the image; mAP is a function that calculates the average accuracy between the predicted bounding box and the true bounding box.

5. The method for enhancing the target detection capability of a visual language model according to claim 3, characterized in that, The total reward value R of the i-th candidate output i It is the sum of the format reward and the ODLength reward of the i-th candidate output.

6. The method for enhancing the target detection capability of a visual language model according to claim 3, characterized in that, The update process for policy network parameters is as follows: First, calculate the advantage value of each candidate output based on the N candidate outputs obtained from sampling. The advantage value of the i-th candidate output is denoted as A. i The formula for calculating the dominance value is: A i =(R i -mean{R1,R2,…,R N }) / std{R1,R2,…,R N }; Where R i Let be the total reward value of the i-th candidate output, mean be the average value, and std be the standard deviation value; Next, the objective function of GRPO is calculated based on all the advantage values ​​to obtain the objective function value. Then, the gradient is calculated based on the objective function value to update the network parameters based on the optimization algorithm.

7. The method for enhancing the target detection capability of a visual language model according to claim 6, characterized in that, The objective function of GRPO is calculated as follows: ; s1=π θ (or i |q) / π θ_old (or i |q); s2=clip(π θ (the i |q) / π θ_old (the i |q),1+ε,1-ε); In the formula, q represents the input problem given to the model; o i Representation of the policy network π θ The i-th of the N candidate outputs generated for problem q; N represents the number of outputs from the policy network π. θ The total number of candidate outputs sampled for each question q; π θ This represents the policy network currently being optimized, with parameters θ and π. θ_old This represents the old policy network prior to this iteration update; π ref Represents a reference policy network; A i Indicates the i-th candidate output o i The advantage value; s1 is the probability ratio of importance sampling; s2 is the version after cropping the probability ratio s1; ε is a hyperparameter used to define the clipping range in s2; β is the coefficient of the KL divergence penalty term, which is also a hyperparameter. Represents the current policy π θ KL divergence between the reference strategy and the reference strategy; Let be the mathematical expectation, representing the expectation for all strategies from the old policy π. θ_old The average of the output set obtained from the sampling is performed.

Citation Information

Cited By

  • Illusion relieving method and system fusing two-channel process reward and layered punishment

    CN122088564A

  • Adaptive-reflective-based tool-augmented imaging optical system automatic design method

    CN122413982A