Reinforcement learning framework-based availability generalization reasoning method and system, computer equipment and medium
By employing an availability generalization reasoning method based on a reinforcement learning framework, and utilizing a binocular stereo vision camera and a thought chain reasoning mechanism, the problem of insufficient out-of-domain generalization ability of multimodal large language models is solved, thereby improving the reliability and accuracy of robot operation in unstructured environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-12
AI Technical Summary
Existing availability reasoning techniques, supported by multimodal large language models, suffer from insufficient out-of-domain generalization ability, making it difficult to adapt to complex scenarios and unseen objects in unstructured environments. They also lack explicit logical deduction mechanisms, resulting in insufficient reliability and accuracy of robot operations.
We employ an availability generalization reasoning method based on a reinforcement learning framework. By integrating high-precision binocular stereo vision cameras to acquire multimodal environmental data, we introduce availability prior knowledge and construct task input representations. We then utilize a thought chain reasoning mechanism for iterative optimization and generalization enhancement to generate high-precision availability region predictions.
It improves the reliability and interpretability of the robot's availability region reasoning in unstructured environments, enhances its adaptability to unseen objects and complex scenes, and enables more accurate and interpretable operations.
Smart Images

Figure CN122021940A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot intelligent control technology, and in particular to an availability generalization reasoning method, system, computer device, and medium based on a reinforcement learning framework. Background Technology
[0002] In the field of autonomous robot operation, availability ( Availability reasoning, as a key technology bridging a robot's perception capabilities and physical manipulation behaviors, aims to enable robots to autonomously identify potential physical areas in an object or environment where certain actions can be performed, based on perceived information. With the widespread application of robots in complex scenarios such as industrial automation and home services, the reliability and generalization ability of availability reasoning technology have become crucial foundations for achieving intelligent robot operation.
[0003] In the development of availability reasoning technology, several methods have emerged, including manually defined methods based on geometric rules, pixel-level recognition methods based on deep learning, and, in recent years, reasoning enhancement methods combining multimodal large language models. Invention patent (application publication number CN119693768A) discloses an attribute prediction method based on a multimodal large language model using a multimodal thought chain. This method improves the contextual understanding and logical consistency of attribute prediction tasks by constructing a mask generator and a scene graph parser, combined with hierarchical thought chain reasoning and logic checking mechanisms. While this technical solution demonstrates certain advancements in attribute recognition and relationship inference, its direct application to robotic availability reasoning tasks still has significant limitations. Specifically, existing methods mostly focus on static matching of attributes and object categories, lacking an explicit reasoning mechanism for the essential logic of availability. This makes it difficult to adapt to the dynamic changes in availability regions between different objects and the need for cross-category generalization. At the same time, when faced with challenges such as diverse object shapes and complex scene interference in open environments, the reasoning process of existing methods lacks interpretability and has limited adaptability to unseen object types or interaction scenarios. As a result, the accuracy and reliability of availability reasoning in actual robot operations cannot meet the needs of practical applications.
[0004] Therefore, although existing availability reasoning techniques have certain semantic guidance capabilities with the support of multimodal large language models, they still suffer from insufficient out-of-domain generalization ability, which restricts the further improvement of robots' autonomous operation capabilities in unstructured environments. Summary of the Invention
[0005] To address the aforementioned shortcomings or deficiencies, this invention provides an availability generalization reasoning method, system, computer device, and medium based on a reinforcement learning framework, which can solve the technical problem of insufficient out-of-domain generalization ability in existing technologies supported by multimodal large language models.
[0006] This invention provides an availability generalization reasoning method based on a reinforcement learning framework. This method is based on a robot equipped with a perception system, which integrates a binocular stereo vision camera supporting high-precision depth information perception, and includes: The system acquires multimodal environmental data collected by the perception system and performs preprocessing on the multimodal environmental data, which includes visual image data captured by a binocular stereo vision camera and task description text.
[0007] Based on preprocessed visual image data and task description text, we introduce availability prior knowledge to construct an availability reasoning task input representation.
[0008] The availability reasoning task is input into a pre-defined large language model, and availability reasoning calculations are performed through the large language model to generate an initial availability region prediction.
[0009] Based on the reinforcement learning framework, the initial availability region prediction is iteratively optimized and generalized using the thinking chain reasoning mechanism to output the target availability region information.
[0010] According to a second aspect, this invention provides an availability generalization reasoning system based on a reinforcement learning framework. This availability generalization reasoning system is based on a robot equipped with a perception system, which integrates a binocular stereo vision camera supporting high-precision depth information perception, and includes: The multimodal environment data acquisition module is configured to acquire multimodal environment data collected by the robot's perception system and perform preprocessing on the multimodal environment data, which includes visual image data captured by a binocular stereo vision camera and task description text.
[0011] The Availability Input Representation Construction Module is configured to incorporate prior availability knowledge to construct an availability reasoning task input representation based on preprocessed visual image data and task description text.
[0012] The initial availability region prediction module is configured to input the availability reasoning task input representation into a preset large language model, perform availability reasoning calculations through the large language model, and generate an initial availability region prediction.
[0013] The availability reasoning output module is configured to use a reinforcement learning framework and a thought chain reasoning mechanism to iteratively optimize and generalize the initial availability region prediction in order to output the target availability region information.
[0014] According to a third aspect, the present invention provides a computer device comprising: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to execute any of the availability generalization reasoning methods based on the reinforcement learning framework in the embodiments of the present invention.
[0015] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute any of the availability generalization inference methods based on reinforcement learning frameworks in the embodiments of the present invention.
[0016] The present invention provides an availability generalization inference method based on a reinforcement learning framework. This method is achieved through four core steps: multimodal data acquisition and preprocessing driven by dedicated hardware, availability task input representation construction, initial inference using a large language model, and iterative optimization using reinforcement learning. Specifically, a perception system integrating a binocular stereo vision camera supporting high-precision depth information perception acquires multimodal environmental data and performs preprocessing operations to extract structured features from raw visual and textual information containing rich 3D spatial information. Availability prior knowledge is introduced based on the preprocessed visual image data and task description text to construct an availability inference task input representation, integrating domain-specific logical knowledge into the inference front end. The availability inference task input representation is input into a pre-set large language model, and availability inference calculations are performed to generate an initial availability region prediction, enabling preliminary region localization based on semantic understanding. Finally, based on a reinforcement learning framework, a thought chain inference mechanism is used to iteratively optimize and generalize the initial availability region prediction, outputting target availability region information to achieve continuous calibration and improved adaptability of the prediction results.
[0017] In this technical solution, the present invention addresses the opaque reasoning process issue described in the background art by introducing prior knowledge of availability to construct a task input representation and explicitly guiding the reasoning logic using a thought chain reasoning mechanism. This achieves interpretability and logical traceability of the availability derivation process, thus solving the black box problem of reasoning caused by the lack of an explicit logical derivation mechanism in existing technologies. Regarding the insufficient out-of-domain generalization ability, the invention introduces precise depth priors provided by a binocular stereo vision camera and combines iterative optimization and generalization enhancement processing based on a reinforcement learning framework to establish a dynamic calibration mechanism for initial predictions. This enables the model to adapt more quickly and accurately to the reasoning needs of unseen objects or complex scenes, solving the problems of strong dependence on training data distribution and weak cross-scene transfer ability in existing technologies. In particular, the binocular stereo vision camera provides crucial three-dimensional geometric constraints for the spatial localization of availability regions, improving the physical rationality and operational feasibility of judging the availability function of objects in unstructured environments. Therefore, the technical solution of the present invention solves the technical problem of insufficient out-of-domain generalization ability in the existing technology under the support of multimodal large language models, and improves the reliability, interpretability and generalization operation ability of the robot in reasoning about the availability region of objects. Attached Figure Description
[0018] Figure 1 This is a flowchart of an availability generalization reasoning method based on a reinforcement learning framework according to an embodiment of the present invention; Figure 2 This diagram illustrates the framework of a reward-design ablation experiment according to an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of an availability generalization reasoning system based on a reinforcement learning framework according to an embodiment of the present invention; Figure 4 This is a block diagram of a computer device for implementing embodiments of the present invention. Detailed Implementation
[0019] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0020] During the development of this invention, the inventors, through extensive experiments and data analysis, revealed the intrinsic connection between the structured chain-of-thought mechanism and cross-scenario generalization ability in availability reasoning based on a large language model (LLM): by introducing a chain-of-thought reasoning mechanism, common-sense knowledge (such as "handle-like structure adaptation and grasping") contained in the large language model can be transformed into a logical deduction chain for availability reasoning, thereby significantly improving the model's adaptability to unseen objects or scenarios. Based on this relationship, the inventors innovatively proposed this technical solution, utilizing a multimodal large language model, through the chain-of-thought reasoning mechanism, combined with a reinforcement learning framework, to achieve iterative optimization and out-of-domain generalization of availability reasoning, embodying the core concept of "perception-cognition-reinforcement learning collaborative driving."
[0021] Specifically, through comparative experiments, the invention team discovered that traditional deep learning-based affordance reasoning methods suffer from insufficient out-of-domain generalization ability: their reasoning process relies on the types of objects covered by the training data, and when faced with unseen objects (such as irregularly shaped vases not included in the training set), they cannot effectively transfer knowledge, and the reasoning process lacks interpretability. These technical shortcomings result in low operational reliability of robots in unstructured environments and difficulty in adapting to diverse scenarios.
[0022] Therefore, according to the first aspect, this invention provides an availability generalization reasoning method based on a reinforcement learning framework, which can be applied to autonomous operation tasks in the field of industrial automation and service robots (hereinafter referred to as "the system"). This system can be deployed locally or collaboratively in the cloud on a robot control unit or edge computing server to complete the intelligent identification and reasoning tasks of object availability regions.
[0023] Specifically, this system can be deployed in various hardware environments, including but not limited to: industrial robot controllers, service robot embedded systems, and cloud-based GPU (Graphics Processing Unit) server clusters. This flexible deployment architecture allows the system to meet the high real-time and reliability requirements of industrial scenarios while adapting to the resource-constrained and dynamic computing needs of service scenarios. In terms of its operational mechanism, the system achieves decoupling and efficient collaboration among its core components—multimodal data acquisition, availability inference task processing, and reinforcement learning optimization—through modular design.
[0024] like Figure 1 As shown, this method is based on a robot equipped with a perception system. The perception system integrates a binocular stereo vision camera that supports high-precision depth information perception and may include: Step S110: Acquire multimodal environmental data collected by the sensing system and perform preprocessing on the multimodal environmental data.
[0025] The multimodal environment data includes visual image data captured by a binocular stereo vision camera and task description text. The visual image data refers to a two-dimensional pixel array containing the target object and its surrounding environment, which is collected by an optical sensor. The task description text refers to a sequence of instructions expressed in natural language that specifies the operation actions and their constraints.
[0026] Specifically, the availability reasoning and generation process of this invention mainly includes two stages: preliminary reasoning and refined segmentation. In the preliminary reasoning stage, the system can segment the network through instances (such as...). The system performs object instance segmentation on visual image data, extracting pixel-level mask regions of target objects. Simultaneously, a semantic parsing module (such as BERT) performs intent recognition on the task description text, extracting semantic features such as action type (e.g., "grab" or "place") and associated objects. The extracted visual and semantic features are then fused and input into the Affordance-R1 inference framework based on a Large Language Model (LLM). The Policy Model in this framework first processes the visual image into visual tokens using its image encoder, and simultaneously converts the task description text into text tokens using a text tokenizer; these tokens are then input into the LLM. The LLM performs affordance inference computation based on the input multimodal information, outputting bounding boxes containing affordance regions, their center points, and interpretable chain-of-thought text. In the fine-grained segmentation and mask generation stage, the target bounding box and center point output by LLM are used as visual cues and input into Segment Anything Model 2 (SAM2). This model performs pixel-level segmentation on the original image, ultimately generating a high-precision affordance mask. This mask provides direct visual priors and spatial guidance for the robot's subsequent operations.
[0027] in, The full English name is (Mask Region Convolutional Neural Network). This network generates pixel-level segmentation masks while performing object detection by adding a mask prediction branch. It is often used to extract the precise contour regions of target objects in images, such as the instance segmentation processing of visual image data to extract the mask regions of target objects as described in this embodiment of the invention. BERT stands for Bidirectional Encoder Representations from Transformers. This pre-trained language model, through its Transformer encoder architecture and bidirectional contextual understanding capabilities, is able to perform deep semantic encoding of text.
[0028] For example, when processing the "Grasp the box" instruction, the binocular camera captures images of the desktop environment, and the instance segmentation network initially locates the "box" target; the semantic parsing module extracts the "grasp" intent. This information is then input into the Affordance-R1 framework, where its policy model and LLM work together to infer the bounding box and center point coordinates of the box, which best suits the grasping region. Finally, the SAM2 model, based on these coordinate cues, segments a precise mask of the region on the original image, successfully guiding the robot to complete the grasping operation.
[0029] Step S120: Based on the preprocessed visual image data and task description text, introduce prior knowledge of availability to construct the input representation of the availability reasoning task.
[0030] In this context, the input representation for the affordance reasoning task refers to the sequence of embedded vectors formed by fusing visual features, text features, and prior knowledge, which can be directly processed by the Large Language Model (LLM).
[0031] Specifically, the system can first construct a structured Chain-of-Thought (CoT) prompt template based on available prior knowledge. This template typically includes explicit guidance for the reasoning stages, for example, using... <think> ...< / think> , <rethink> ...< / rethink> , <answer> ...< / answer>The system uses labels to structure the reasoning process, explicitly guiding the LLM to derive affordance logic. Then, through a cross-modal encoder (such as a Transformer-based fusion network), visual features (masked regions) extracted from instance segmentation models (such as Mask R-CNN), textual features (action intent) extracted from semantic parsing models (such as BERT), and the aforementioned thought chain prompt template are aligned and fused in a multimodal manner. This process involves mapping feature vectors from different modalities to the same semantic space and performing weighted concatenation through an attention mechanism, ultimately generating a unified-dimensional input tensor. For example, for the task of "safely grasping containers with handles," the system generates the following prompt template: <think> The target object is a container with a handle, and the grasping action is usually adapted to the handle-like structure.< / think> <rethink> Ensure that there is enough space in the handle to accommodate the mechanical fingers.< / rethink> <answer>Then, the feature vector encoded from the template text is concatenated with the visual feature vector of the container mask region and the semantic feature vector of the "grabbing" action, ultimately forming a dimension of The input representation tensor is used as the input to a large language model.
[0032] Step S130: Input the availability reasoning task into the preset large language model, perform availability reasoning calculation through the large language model, and generate an initial availability region prediction.
[0033] The initial availability region prediction refers to the structured data output by the model, which includes spatial coordinates (such as bounding boxes) and availability labels (such as "crawlerable region").
[0034] Specifically, the system can perform encoding-decoding computations on the input representation using multi-layered Transformer blocks of a large language model. First, it generates thought chain reasoning text (e.g., "The object's handle is the grasping area"), and then parses the coordinate and label information within the text. For example, the system uses the Qwen-2.5 model (a large-scale pre-trained language model developed by Alibaba Group) to process the input and outputs the thought chain: " <think> The handle is the gripping area.< / think> <rethink> The handle structure can accommodate mechanical claws.< / rethink> <answer> Coordinates: [x1, y1, x2, y2], Tags: crawling< / answer> The model extracts initial predictions from these predictions. In more complex scenarios, the model may output multiple availability region predictions. For example, for a container with both a handle and a lid, the lid region may be output as an "openable region" and the handle as a "gripable region" under the "open" command.
[0035] Step S140: Based on the reinforcement learning framework, the initial availability region prediction is iteratively optimized and generalized using the thought chain reasoning mechanism to output the target availability region information.
[0036] Among them, iterative optimization and generalization enhancement processing refers to ranking multiple candidate predictions by reward scores through reinforcement learning strategies, dynamically calibrating the prediction results, and improving the model's generalization ability for inference on unseen objects or complex scenes.
[0037] Specifically, the system can fine-tune the policy model using a grouped relative policy optimization algorithm. During optimization, for a given input, the policy model generates a set containing multiple candidate availability region predictions and their corresponding thought chain texts. Each candidate answer is evaluated using a carefully designed hybrid reward function that comprehensively measures the quality across three dimensions: (1) Format reward: Evaluate whether candidate answers follow the preset thought chain structure (such as containing explicit `, <rethink> 、 <answer>(Labels and their logical coherence) (2) Perceived reward: assess the degree of spatial conformity between the predicted availability region (boundary box) and the actual annotation or environmental constraints, usually calculated based on the generalized intersection-union ratio; (3) Cognitive reward: The semantic rationality of the thought chain text and the accuracy of the availability label are evaluated, usually based on the semantic similarity between it and the prior knowledge of the task.
[0038] The overall reward score for each candidate answer is the weighted sum of the three rewards mentioned above. The goal of the optimization algorithm is to drive the policy model to output a prediction with a higher overall reward through iterative training. For example, in one training iteration, the system generates eight candidate predictions for the same input. After evaluation by the mixed reward function, the prediction with the highest overall score (e.g., 0.92 points) is selected as the final target availability region information. After optimization using this reinforcement learning framework, the proposed method achieves a generalized intersection-union (OU) ratio of 67.41% on the availability reasoning benchmark dataset (e.g., [database name]), outperforming the unoptimized baseline model and validating its effectiveness in improving prediction accuracy and generalization ability.
[0039] In other embodiments, such as Figure 2 This invention demonstrates the reinforcement learning training and update process framework based on Group Relative Policy Optimization (GRPO). This framework specifically illustrates the complete process of how the policy inference model iteratively self-optimizes during the training phase using multi-dimensional reward signals and reference model constraints. Figure 2 As shown, the process begins with a bimodal input of instructions and images: after receiving this input, the policy reasoning model performs structured availability reasoning, generating a thought chain text containing three stages: "thinking," "reflection," and "answer," along with corresponding availability region predictions. Simultaneously, a fixed pre-trained reference model produces a reference output under the same input; the difference between the two output distributions is measured using relative entropy to stabilize training and prevent mode collapse. Subsequently, the system finely evaluates the output of the policy reasoning model through a carefully designed hybrid reward function: a format reward checks whether the output strictly follows the "thinking-reflection-answer" stage structure; a perceptual reward calculates the spatial matching degree based on the intersection-union ratio (IUU) between the predicted bounding box and the ground truth annotations; and a cognitive reward evaluates its semantic rationality by comparing the Manhattan distance between the reasoning text and the semantic embedding vectors of prior knowledge. These discrete reward signals, together with continuous relative entropy constraints, constitute the optimization objective, driving the grouped relative policy optimization algorithm to update the parameters of the policy reasoning model. Through repeated iterations of this closed-loop process, the policy reasoning model gradually learns to generate outputs with more standardized formats, more accurate spatial positioning, and more reasonable semantic reasoning. For example, in the early stages of training, the model might output multiple inaccurate bounding boxes and ambiguous thought chains for the instruction "push open the left cabinet door." After multiple rounds of optimization in this process, the model can not only stably focus the prediction box on the cabinet door handle area, but its generated thought chains (such as "the cabinet door handle is a common point of force application, and the left side is a hinge, so the force should be applied by translating to the right") have also significantly improved in terms of rationality and interpretability. This optimization and update mechanism is the core technical guarantee for this solution to achieve high-precision and high-generalization availability reasoning.
[0040] Therefore, according to the above implementation method, the system is achieved collaboratively through four core steps: multimodal data acquisition and preprocessing driven by dedicated hardware, construction of availability task input representation, initial inference of a large language model, and iterative optimization using reinforcement learning. Specifically, a perception system integrating a binocular stereo vision camera supporting high-precision depth information perception acquires multimodal environmental data and performs preprocessing operations to extract structured features from raw visual and textual information containing rich 3D spatial information; availability prior knowledge is introduced based on the preprocessed visual image data and task description text to construct availability inference task input representation operations, which integrate domain-specific logical knowledge into the inference front end; the availability inference task input representation is input into a pre-set large language model and availability inference calculations are performed to generate initial availability region prediction operations, which achieve preliminary region localization based on semantic understanding; and the initial availability region prediction is iteratively optimized and generalized using a thought chain inference mechanism based on a reinforcement learning framework to output target availability region information operations, which achieve continuous calibration and adaptive improvement of the prediction results.
[0041] Specifically, in this implementation, to address the opaque reasoning process issue mentioned in the background technology, prior knowledge of availability is introduced to construct a task input representation. A thought chain reasoning mechanism is used to explicitly guide the reasoning logic, achieving interpretability and logical traceability of the availability derivation process. This solves the black-box problem of reasoning caused by the lack of an explicit logical derivation mechanism in existing technologies. Regarding insufficient out-of-domain generalization ability, precise depth priors provided by a binocular stereo vision camera are introduced. Combined with iterative optimization and generalization enhancement processing based on a reinforcement learning framework, a dynamic calibration mechanism for initial predictions is established. This enables the model to adapt more quickly and accurately to the reasoning needs of unseen objects or complex scenes, solving the problems of strong dependence on training data distribution and weak cross-scene transfer ability in existing technologies. In particular, the binocular stereo vision camera provides crucial three-dimensional geometric constraints for the spatial localization of availability regions, improving the physical rationality and operational feasibility of judging object availability functions in unstructured environments. Therefore, the technical solution of this implementation solves the technical problem of insufficient out-of-domain generalization ability of existing technologies under the support of multimodal large language models, and improves the reliability, interpretability and generalization operation ability of robots in reasoning about the availability region of objects in unstructured environments.
[0042] In other embodiments, Table 1 below shows the experimental results of a systematic evaluation of the availability generalization inference method based on the reinforcement learning framework provided by this invention on various typical operational tasks. By comparing the task success rates of the same method in "seen" and "unseen" scenarios, this table quantitatively illustrates the significant effect of this invention in improving the model's generalization ability across scenarios, objects, and tasks.
[0043] The tasks listed in Table 1 are explained below: Human-like grasping refers to the task of a robot mimicking human hand movements to grasp objects.
[0044] Open-Vocabulary Pick & Place: This refers to the task of a robot picking up and placing objects based on natural language instructions (the vocabulary is not limited to the fixed set seen during training).
[0045] Hinge Object Open & Close: This refers to the task of opening or closing objects with hinge structures (such as doors, cabinets, etc.).
[0046] Guide Rail Objects Pull & Push: This refers to the task of performing linear motion operations such as pushing and pulling on objects mounted on guide rails.
[0047] Using Tools (Tool Use Task): This refers to a task in which a robot uses tools (such as screwdrivers, pliers, etc.) to complete a specific operation.
[0048] Long-horizon manipulation refers to complex operational tasks consisting of multiple sub-steps that require long-term planning.
[0049] Cross-Embodiment: refers to the task of transferring operational skills between robot platforms of different shapes or configurations.
[0050] Cross-Environments (Cross-Environment Generalization): refers to tasks that perform the same operation under different physical environments or scenario conditions.
[0051] As shown in Table 1, in the "Human-like Grasping" task, the method achieved a 100% success rate when dealing with objects already seen during training; while when handling unseen objects, the success rate remained at 75%, verifying that the method can effectively transfer the grasping strategy to new objects through availability prior knowledge and thought chain reasoning. In the "Open-Vocabulary Pick & Place" task, the method achieved a 60% success rate in both seen and unseen scenarios, demonstrating its stability in zero-shot operations based on semantic understanding. Particularly noteworthy is that in tasks involving complex interaction mechanisms, such as "Hinge Object Open & Close" and "Guide Rail Objects Pull & Push," the method achieved success rates of 90% and 75% respectively in unseen scenarios, indicating that the reinforcement learning iterative optimization mechanism introduced in this invention can effectively calibrate reasoning for complex functional structures. In the "Cross-Embodiment" and "Cross-Environments" tasks, which demonstrate high-order generalization capabilities, the method's success rate in unseen scenarios (80% and 70%, respectively) is comparable to or even better than in seen scenarios, strongly demonstrating the effectiveness of the "perception-cognition-reinforcement learning collaborative driving" architecture in adapting to different robot platforms and environmental changes. Although there is significant room for improvement in the success rate (40%) in unseen scenarios of high-difficulty, combinatorial tasks such as "Using Tools," the overall data clearly shows that this invention significantly enhances the generalization performance and practical reliability of the availability reasoning model in unstructured real-world environments through structured thought chain reasoning and multi-reward reinforcement learning optimization.
[0052] In other embodiments, Table 2 below shows a comparison of the semantic understanding and localization performance of different advanced models in complex object manipulation scenarios. Table 2 provides a detailed comparison of the generalized intersection-over-union (gIoU) and central intersection-over-union (cIoU) metrics of G-DINO (DETR with Improved deNoising anchOrboxes for Open-Set Grounded Object Detection), LISA (Language-Instructed Segment Anything Model), GLaMM (Grounded Large Multimodal Model), Vision-Reasoner (Visual Reasoning Unified Multimodal Model), and the core model of this solution, Affordance-RI (Affordance Reasoning Iterative Optimization Framework), under a zero-shot setting, on three benchmark datasets with different difficulty levels: 3DOI (3D Object Interaction Affordance Dataset), HANDAL-easy (HANDAL Real-World Manipulable Object Dataset - Easy Difficulty Subset), and HANDAL-hard (Real-World Manipulable Object Dataset - Hard Difficulty Subset). The data clearly demonstrates that the Affordance-RI model used in this solution achieves leading performance across the vast majority of metrics, especially in the more challenging HANDAL-hard scenario, where its gIoU and cIoU reach 40.7 and 37.9 respectively, outperforming other comparative models. For example, in simulating a robot retrieving a fragile vase from a cluttered storage cabinet (corresponding to the HANDAL-easy scenario), the table data shows that traditional models such as G-DINO have a gIoU of only 3.6, meaning that the predicted grasping area has very little overlap with the actual operable area, easily leading to task failure; while Vision-Reasoner, which is also based on a large language model for visual reasoning, has a gIoU of 29.6, showing a significant improvement. However, the Affordance-RI model in this scheme achieved a gIoU of 43.1 and a cIoU of 38.7 in this scenario, which means that it can not only locate the availability region more accurately, but also the center point of the predicted region is more consistent with the center of the real region, which is crucial for robot operations that require high-precision alignment.Furthermore, when operating objects with complex hinges or rails in a simulated home environment (such as pushing open a stuck drawer or closing a heavy cabinet door, corresponding to the HANDAL-hard scenario), the Affordance-RI model maintained high performance with a gIoU of 40.7 and a cIoU of 37.9, while the performance of other models showed a significant decline. This comparative result strongly demonstrates that the reinforcement learning iterative optimization and structured thought chain reasoning mechanism deeply integrated in this solution can effectively improve the model's deep understanding of the functional structure of objects and the accurate inference of spatial relationships. This enables the model to exhibit excellent robustness and generalization ability when facing real-world challenges such as variable object appearances, complex scenes, and difficult interactive tasks, providing crucial technical support for robots to achieve safe, reliable, and accurate operation in unstructured environments.
[0053] In other embodiments, Table 3 below shows a set of core hyperparameters and their typical values used in the reinforcement learning training framework adopted in this invention for Grouped Relative Policy Optimization (GRPO) of the Policy Model. These hyperparameters collectively define the key configurations in the model training process, directly affecting optimization efficiency, convergence stability, and the performance of final availability inference. Specifically, the Batch Size is set to 8, and the Experience pool is set to 16, which ensures that each parameter update is based on an appropriate amount and variety of samples, balancing training efficiency and stability. The Kullback-Leibler Divergence Weight (KL divergence weight) is set to 0.05 to constrain the difference between the output distribution of the policy model and the pre-trained reference model, effectively preventing mode collapse or catastrophic forgetting during the optimization process. The optimizer chosen is AdamW (Adaptive Moment Estimation with Weight Decay Regularization), and its learning rate is set to... It also incorporates a weight decay factor. This helps achieve stable and efficient gradient updates when fine-tuning large parameter models such as large language models. Of particular note is the weight configuration of the components in the hybrid reward function: the format reward weight, perception reward weight, and recognition reward weight are set to 1, 2, and 1, respectively. This configuration means that during training, the model gives the highest priority to the spatial alignment of the predicted region with the ground truth label (perception reward), which is highly consistent with the core goal of the availability reasoning task—generating accurate operable regions—while also taking into account the standardization of the output structure (format reward) and the rationality of the semantics (recognition reward). For example, in a training task for "grabbing a transparent glass," it is precisely the prominent role of the perception reward weight that drives the model to continuously correct its predicted bounding box during iterations, ultimately making it closely fit the real outline of the handle, while the format and recognition rewards ensure that the output thought chain text structure is clear and conforms to the common sense that "grabbing should act on the handle." Furthermore, the settings for the maximum response length (2048) and the maximum prompt length (1300) ensure the complete expression of complex thought chain reasoning, while the maximum and minimum number of pixels (max / min pixels) regulate the resolution range of the input image. These carefully designed and experimentally verified hyperparameter combinations together constitute an important technical foundation for achieving efficient and stable training and ultimately obtaining a high-performance availability reasoning model.
[0054] In some embodiments, the sensing system is configured with an image acquisition device; acquiring multimodal environmental data acquired by the sensing system and performing preprocessing on the multimodal environmental data, including: Instance segmentation processing is performed on the visual image data acquired by the image acquisition device to extract at least one mask region of the target object.
[0055] Among them, instance segmentation refers to a computer vision task that uses a pixel-level classification network to distinguish each target object in an image from the background and other objects and generate an accurate contour; the mask region refers to a binary matrix representing the pixel position occupied by the target object, where the target object region has a value of 1 and the background region has a value of 0.
[0056] Specifically, the system can utilize instance segmentation models based on convolutional neural networks (such as...) This is a computer vision model capable of simultaneously performing object detection and pixel-level segmentation. It uses ResNet-50 (a deep residual network, belonging to the convolutional neural network architecture, whose core innovation lies in the introduction of a residual learning mechanism. This model contains 50 deep layers and effectively solves the gradient vanishing and network degradation problems in deep neural network training through cross-layer identity mapping, thus enabling stable training of extremely deep networks and improving feature extraction capabilities) as the backbone network. It performs feature extraction, region proposal, and mask prediction on the input visual image data, ultimately outputting the mask region corresponding to each detected target object. For example, when processing a table scene image containing a cup and a plate, the instance segmentation model successfully identifies and segments three independent objects: the cup body, the handle, and the plate body, generating a mask region of size [size missing] for each object. A binary mask matrix for pixels.
[0057] Semantic parsing is performed on the task description text to extract the semantic features of the action intent corresponding to the task instructions.
[0058] Semantic parsing refers to the computational process of using natural language processing models to understand the deeper meaning of text and transform it into a structured semantic representation; action intention semantic features refer to the vectorized representations of operational actions (such as "grab" or "place") and target objects (such as "cup" or "handle") abstracted from the task description text.
[0059] Specifically, the system can encode task description text using a pre-trained language model (such as BERT-base, a pre-trained language model based on the Transformer architecture, belonging to the standard-scale version of the Bidirectional Encoder Representations from Transformers series). It leverages the Transformer architecture to capture contextual dependencies between words and extracts fixed-dimensional semantic feature vectors through specific classification heads or pooling layers. For example, when the system parses the task description text "Please safely grasp the handle of the red mug," the semantic parsing module outputs a 768-dimensional semantic feature vector, which encodes the action type "grab," the target object "mug," its attribute "red," and the component "handle," representing a composite intent.
[0060] Cross-modal association is performed between the masked region and the semantic features of the action intent. A cross-modal attention mechanism is used to align the visual features and semantic features in the vector space to construct the input representation for the affordance reasoning task.
[0061] Cross-modal association refers to the operation of aligning and fusing data features from different modalities (visual and textual) at the semantic level to form a unified joint representation rich in multi-source information.
[0062] Specifically, the system can use cross-modal attention mechanisms or multimodal fusion networks (such as the CLIP encoder structure) to concatenate or weight-fuse the visual feature vectors of the masked region with the semantic feature vectors of the action intent, generating a tensor that combines spatial and semantic information as the input representation. For example, the system inputs the masked region of the cup handle (visual features, dimension 512) and the semantic features of "grabbing the cup handle" (textual features, dimension 768) into a bilinear fusion layer, outputting a 1280-dimensional fused feature vector as the input representation for the affordance reasoning task. CLIP ( The contrastive language-image pre-trained encoder structure is an overall architecture within the CLIP multimodal pre-trained model proposed by OpenAI. It consists of two parallel encoder modules that independently process image and text inputs, achieving cross-modal semantic alignment through a contrastive learning objective. This structure is the core component for aligning images and text in a unified semantic space.
[0063] Therefore, according to the above implementation method, the system can effectively align and fuse the spatial information of objects in the visual scene with the semantic information of natural language task instructions, providing a unified input representation with complete information and clear structure for subsequent availability reasoning tasks, and laying the data foundation for accurate reasoning.
[0064] In some embodiments, availability prior knowledge refers to the essential logical knowledge pre-stored in the large language model regarding the adaptation relationship between object functional attributes and action intentions; based on preprocessed visual image data and task description text, availability prior knowledge is introduced to construct the input representation of the availability reasoning task, including: Based on prior knowledge of availability, a thought chain prompt template is constructed to guide the large language model in making availability logic deductions.
[0065] Among them, the thought chain prompt template refers to one that includes labels for fixed reasoning stages (such as...). <think> 、< / think> , <rethink> 、< / rethink> , <answer> 、< / answer> A structured text framework is used to force the model to output thoughts, reflections, and conclusions step by step.
[0066] Specifically, the system can generate prompt text containing multi-stage reasoning instructions by parsing the action intent (such as "safely grasp") in the task description text and combining it with prior availability knowledge (such as "the grasping action is adapted to handle-shaped, raised, or flat edge structures"). For example, for the "safely grasp a cup" task, the system constructs the following prompt template: "Please reason in the following format:" <think> ...< / think> <rethink> ...< / rethink> <answer> Coordinates: [x1, y1, x2, y2], Label: XXX< / answer> The process involves three phases: the thinking phase guides the model to analyze the structural features of the cup, the reflection phase verifies the stability of the grasping area, and the answer phase outputs the final coordinates and labels.
[0067] By integrating visual image data, task description text, and thought chain prompt templates, an input representation for the affordance reasoning task is generated.
[0068] In this context, the input representation for the availability reasoning task refers to a unified tensor or embedding vector that integrates visual features, semantic features, and reasoning frameworks, serving as the multimodal input for a large language model.
[0069] Specifically, the system can use a cross-modal encoder (such as a Transformer-based fusion network) to concatenate and weight the pixel features of the image, the semantic features of the text, and the instruction features of the prompt template, generating a dimension of... The input representation. For example, the system will... The image of the cup is encoded as pixels. The visual feature sequence encodes the task text into... The text feature sequence is concatenated with the instruction vector of the prompt template, and then a final dimension of [dimensionality missing] is generated through a cross-modal attention layer. The input representation.
[0070] Therefore, according to the above implementation method, the system can explicitly guide the reasoning logic of the large language model through a structured thought chain prompt template, and achieve efficient integration of visual, text and reasoning instructions, providing an input representation that combines semantic guidance and spatial awareness for availability region prediction.
[0071] In some embodiments, the large language model is configured with a dedicated affordance inference module; the step of performing affordance inference computation through the large language model to generate an initial affordance region prediction includes: The input representation of the availability reasoning task is derived hierarchically through an availability-specific reasoning module.
[0072] Hierarchical reasoning derivation refers to a systematic reasoning process that decomposes the availability reasoning task into multiple sub-tasks with logically progressive relationships, and executes these sub-tasks in sequence to obtain the final reasoning result.
[0073] Specifically, the system can sequentially execute subtasks such as object functional attribute recognition and action adaptation region recommendation through a pre-defined inference pipeline. The output of each subtask serves as the input for the next subtask, forming a complete inference chain. For example, the system first identifies the "handle-like structure" and "cylindrical container" attributes of a water cup, then recommends the handle region based on the rule of "grasping action adapting to the handle-like structure," and finally outputs a grasping region suggestion with coordinates.
[0074] After the action adaptation region recommendation phase is completed, a reflection and verification subprocess is triggered to check the logical consistency between the recommended region and the availability functional attributes.
[0075] Among them, the reflection and verification subprocess refers to an automated quality inspection mechanism that performs logical review and consistency verification on the preliminary reasoning results, aiming to discover and correct contradictions or errors in the reasoning process.
[0076] Specifically, the system can use a rule engine or constraint satisfaction check module to verify whether the recommended area meets the geometric constraints, physical feasibility, and other conditions defined in the availability functional attribute definition. For example, the system checks whether the recommended "pouring area" is indeed located at the opening of the container and has sufficient space to accommodate the liquid flow. If the recommended area is found to be located at the bottom of the cup, it is marked as logically inconsistent.
[0077] Based on the output of the hierarchical reasoning and reflection verification subprocess, an initial availability region prediction containing regional coordinate information and availability labels is generated.
[0078] Among them, regional coordinate information refers to the spatial location data of the availability region expressed in pixel coordinates or normalized coordinates; availability labels refer to standardized identifiers used to describe the functional attributes of the region.
[0079] Specifically, the system can convert bounding boxes or polygon vertex sequences into a standardized coordinate format using a coordinate encoder, and map semantic descriptions to a predefined set of labels using a label mapping table. For example, the system outputs "coordinates: The tag "grip" indicates that the recommended gripping area is the 25%~75% width range and the 35%~90% height range of the target object.
[0080] The hierarchical reasoning process includes an object functional attribute identification stage and an action adaptation region recommendation stage. In the object functional attribute identification stage, the structural features of the target object are analyzed and associated with the available functional attributes that match the action intention in the task description text. In the action adaptation region recommendation stage, based on the identified available functional attributes, the physical regions in the visual image data that can perform the action intention are located and recommended.
[0081] Specifically, the system can extract geometric features (such as edges and curvature) of an object through a convolutional neural network and associate these features with action keywords (such as "grab" and "place") in the task description through an attention mechanism. For example, the system recognizes the association between the "ring structure" of scissors and the intention of "grabbing" and then recommends the ring area at the handle of the scissors as the grasping area.
[0082] Therefore, according to the above implementation method, the system can achieve accurate identification of availability areas and ensure logical consistency through a structured hierarchical reasoning and reflective verification mechanism, providing reliable spatial guidance information for robot operation.
[0083] In some embodiments, the reinforcement learning framework is configured with a hybrid reward function, and the thought chain reasoning mechanism is configured to generate a structured reasoning chain containing outputs of a thinking phase, a reflection phase, and a response phase. Based on the reinforcement learning framework, the thought chain reasoning mechanism is used to iteratively optimize and generalize the initial availability region prediction to output target availability region information, including: Multiple candidate answers are generated using a reinforcement learning framework. Each candidate answer contains thought content, reflection content, and answer content generated based on the thought chain reasoning mechanism.
[0084] Among them, candidate answers refer to multiple sets of alternative reasoning results generated by sampling the same input through a policy network, and each set of results contains a complete thinking-reflection-response reasoning chain.
[0085] Specifically, the system can perform multiple forward inferences in parallel on the same input sample through a grouping generation strategy. Each time, different random seeds are used to generate independent inference paths, forming a set of candidate answers with diverse outputs. For example, the system generates 8 candidate answers for the task of "safely grabbing a cup". Each answer contains a complete inference chain, including a thinking stage (analyzing the cup structure), a reflection stage (verifying the stability of the grabbing point), and an answering stage (outputting the grabbing coordinates).
[0086] Grouping relative strategy optimization calculations are performed on multiple candidate answers, and the reasoning quality of each candidate answer is evaluated and ranked based on a hybrid reward function.
[0087] Among them, group relative policy optimization calculation refers to a reinforcement learning algorithm that divides candidate answers into the same group and optimizes the policy through relative comparison within the group rather than absolute scores.
[0088] Specifically, the system can calculate the relative advantage score of each candidate answer within its group using the GRPO (Group Relative Policy Optimization) algorithm, and construct a multi-objective optimization function by combining policy probability loss and KL divergence penalty term. For example, the system groups eight candidate answers together, calculates the comprehensive score of each answer based on a hybrid reward function, and then uses a formula... Update the policy parameters, where This indicates relative advantage within the group.
[0089] Based on the evaluation and ranking results, the candidate answer with the highest comprehensive reward score is selected as the target availability area information.
[0090] The comprehensive reward score refers to the final evaluation value obtained by weighting and merging the scores of three dimensions: format reward, perceived reward, and cognitive reward.
[0091] Specifically, the system can calculate the overall score using a linear weighting method: The weighting coefficients satisfy For example, the system awards candidate answer A 0.9 points for format, 0.8 points for perception, and 0.95 points for cognition, weighted accordingly. After weighting, the overall score was 0.875, ranking first in the group and thus selected as the final output.
[0092] Therefore, according to the above implementation method, the system can fully utilize generative diversity through a grouping relative optimization mechanism, significantly improving training efficiency while ensuring inference quality, and... Achieved 67.41% on the dataset. The accuracy is 9.73 percentage points higher than the traditional PPO (Proximal Policy Optimization) algorithm.
[0093] In some embodiments, the reinforcement learning framework is further configured with a grouping relative policy optimization algorithm, the step of performing grouping relative policy optimization calculations on multiple candidate answers including: Calculate the strategy probability loss term for each candidate answer.
[0094] The policy probability loss term refers to the loss function component used to measure the difference in probability distribution between the output of the current policy model and the output of the reference policy model. Its goal is to minimize the policy update magnitude to maintain training stability.
[0095] Specifically, the system can calculate the difference between the probability of the current policy model generating a candidate response sequence and the probability of the reference model generating the same sequence using the negative log-likelihood loss function. For example, the system calculates the log probability of the candidate response sequence as follows: The log probability of the sequence corresponding to the reference model is The strategy probability loss term is then... .
[0096] Calculate the relative entropy penalty term for the model output distribution corresponding to each candidate answer.
[0097] Among them, the relative entropy penalty term of the model output distribution refers to the term obtained through KL ( The divergence is used as the basis for calculating the relative entropy penalty term of the model output distribution to prevent the policy model from deviating too far from the reference model during reinforcement learning training. It measures the difference between the current policy model and the reference model in the output distribution and is a regularization term used to prevent the policy update from deviating too far from the initial model.
[0098] Specifically, the system can calculate the difference in output probability distributions between the current policy model and the reference model on the vocabulary using the KL divergence formula, where the reference model is typically the initial pre-trained model. For example, if the system measures the output probability of the current policy model on the keyword "crawl" to be 0.85, and the corresponding probability of the reference model is 0.92, then the KL divergence penalty term value is... .
[0099] The relative advantage score of each candidate answer within the group is determined based on the reward value calculated using the hybrid reward function.
[0100] Among them, the relative advantage score within a group refers to the standardized indicator that converts the original reward value into a comparable relative score through within-group standardization, and is used to eliminate the difference in reward dimensions between different tasks.
[0101] Specifically, the system can standardize the reward by calculating the difference between the reward value of the current candidate answer and the average reward value of other answers in the same group, and then dividing by the standard deviation of the rewards within the group. For example, the system calculates the average reward for a group containing 8 candidate answers. Standard deviation If a candidate's answer has a reward value of 0.92, then its relative advantage score is... .
[0102] Based on the policy probability loss term, the relative entropy penalty term of the model output distribution, and the relative advantage score of the groups, a total loss function for group relative policy optimization is constructed.
[0103] Among them, the total loss function for grouped relative policy optimization refers to a multi-objective optimization function that integrates policy loss, distribution constraints, and relative advantage scores, and is used to balance policy improvement and training stability.
[0104] Specifically, the system can construct the total loss function using a weighted summation method: ,in For example, system settings hyperparameters*. The total loss function value calculated based on the aforementioned example values is: .
[0105] The total loss function is optimized by minimizing the grouping relative strategy, and the model parameters of the large language model are updated.
[0106] Among them, model parameter update refers to the training process of adjusting the weights of the neural network through the gradient backpropagation algorithm so that the total loss function value gradually converges to a local optimum.
[0107] Specifically, the system can use the Adam optimizer (an adaptive moment estimation optimization algorithm, a widely used first-order gradient optimization method in deep learning) to perform gradient descent optimization on the total loss function with a learning rate of 0.0001, updating all model parameters in each iteration. For example, after 500 steps of training, the total loss function value changes from the initial... convergence to Corresponding to On the dataset The indicator improved from 60.58% to 67.41%.
[0108] Therefore, according to the above implementation method, the system can achieve efficient and stable policy learning through the grouped relative policy optimization algorithm, significantly improving training efficiency while ensuring the quality of model output, as shown in Table 2 above, under the complete reward configuration. The indicator reached 67.41%, an increase of 6.83 percentage points from the baseline.
[0109] In some embodiments, the hybrid reward function includes format reward, perceptual reward, and cognitive reward; the step of evaluating and ranking the reasoning quality of each candidate answer based on the hybrid reward function includes: Calculate the format reward score for candidate responses, which is determined based on the completeness of the structured labels and the correctness of the pairings in the candidate responses.
[0110] Among them, the integrity and correctness of structured labels refer to the technical requirement that candidate answers must contain complete three-stage label pairs of thinking-reflection-response, and that the label opening and closing symbols must be correctly matched.
[0111] Specifically, the system can use regular expression matching algorithms to detect whether candidate response text contains complete... <think>< / think> , <rethink>< / rethink> , <answer>< / answer> The system generates tag pairs and assigns a score to each complete pair of tags. For example, a candidate answer containing all three complete tag pairs receives a maximum score of 3 points, a missing tag loses 1 point, and a tag with an unclosed tag receives 0 points.
[0112] The perceived reward score of the candidate answer is calculated. The perceived reward score is determined based on the spatial intersection-union ratio of the coordinate information of the availability region in the candidate answer and the preset reference region.
[0113] Among them, the spatial intersection-union ratio refers to the degree of overlap between the predicted area and the actual labeled area, and its value range is between 0 and 1.
[0114] Specifically, the system can calculate the matching degree by measuring the intersection-union ratio (IUU) between the predicted bounding box and the ground truth bounding box, using the following formula: For example, the system calculates the prediction region. With real area of The value is 0.72. When the threshold is set to 0.7, a full score of 1 point is assigned.
[0115] The cognitive reward score of the candidate answer is calculated. The cognitive reward score is determined based on the cosine similarity between the word vectors of the availability label semantics in the candidate answer and the preset reference label.
[0116] Among them, word vector cosine similarity refers to the degree of semantic association between two words by measuring the cosine distance in the word vector space.
[0117] Specifically, the system can convert labeled words into vector representations using pre-trained language models such as Word2Vec or BERT, and then calculate the cosine similarity between the vectors. For example, after converting the predicted label "grasp" and the reference label "hold" into 300-dimensional vectors, the system calculates a cosine similarity of 0.89, and assigns a full score of 1 when the threshold is set to 0.85.
[0118] The overall reward score for a candidate answer is calculated by weighting the format reward score, perception reward score, and cognitive reward score.
[0119] Weighted sum refers to a scoring method that linearly combines reward scores from different dimensions according to preset weights.
[0120] Specifically, the system can be implemented using the formula: "To perform calculations, where" For example, the system sets the weight to... When a candidate's answer scores 3, 1, and 1 points respectively in the three sections, the overall score is: .
[0121] Multiple candidate answers are ranked based on their overall reward scores.
[0122] The sorting operation refers to arranging candidate answers from highest to lowest based on their overall scores in order to determine the optimal solution selection order.
[0123] Specifically, the system can use a quicksort algorithm to sort the candidate answer set in descending order of their overall scores. For example, the system can sort the overall scores of 8 candidate answers... After sorting, an ordered sequence is obtained. .
[0124] Therefore, according to the above implementation method, the system can comprehensively evaluate the quality of candidate answers through a multi-dimensional reward mechanism and achieve objective and fair ranking selection based on quantitative scores, providing precise optimization guidance for reinforcement learning training.
[0125] Figure 3 This is a structural block diagram of an availability generalization reasoning system based on a reinforcement learning framework according to an embodiment of the present invention.
[0126] like Figure 3 As shown, this availability generalization reasoning system based on a reinforcement learning framework is based on a robot equipped with a perception system. The perception system integrates a binocular stereo vision camera that supports high-precision depth information perception, and includes: The multimodal environment data acquisition module 210 is configured to acquire multimodal environment data collected by the robot perception system and perform preprocessing on the multimodal environment data, which includes visual image data captured by a binocular stereo vision camera and task description text.
[0127] The Availability Input Representation Construction Module 220 is configured to incorporate prior availability knowledge to construct an availability reasoning task input representation based on preprocessed visual image data and task description text.
[0128] The initial availability region prediction module 230 is configured to input the availability reasoning task input representation into a preset large language model, perform availability reasoning calculation through the large language model, and generate an initial availability region prediction.
[0129] The availability reasoning output module 240 is configured to perform iterative optimization and generalization enhancement processing on the initial availability region prediction based on the reinforcement learning framework and using the thinking chain reasoning mechanism to output the target availability region information.
[0130] The specific functions and examples of each module and submodule of the device in this embodiment of the invention can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0131] According to embodiments of the present invention, the above-described method of the present invention can be applied to a computer device and a readable storage medium.
[0132] Figure 4 A schematic block diagram of an example computer device 600 that can be used to implement embodiments of the present invention is shown. The computer device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The computer device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0133] like Figure 4 As shown, the computer device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the computer device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0134] Multiple components in computer device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows computer device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0135] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as an availability generalization inference method based on a reinforcement learning framework. For example, in some embodiments, an availability generalization inference method based on a reinforcement learning framework can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the computer device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the availability generalization inference method based on a reinforcement learning framework described above can be performed. Alternatively, in other embodiments, computing unit 601 may be configured by any other suitable means (e.g., by means of firmware) to perform an availability generalization reasoning method based on a reinforcement learning framework.
[0136] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0137] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0138] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0139] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0140] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0141] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0142] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0143] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this invention should be included within the scope of protection of this invention.< / answer> < / rethink> < / answer>
Claims
1. An availability generalization reasoning method based on a reinforcement learning framework, characterized in that, The method is based on a robot equipped with a perception system, which integrates a binocular stereo vision camera supporting high-precision depth information perception, including: The system acquires multimodal environmental data collected by the perception system and performs preprocessing on the multimodal environmental data, which includes visual image data captured by the binocular stereo vision camera and task description text. Based on the preprocessed visual image data and task description text, availability prior knowledge is introduced to construct the availability reasoning task input representation; The availability reasoning task is input into a preset large language model, and availability reasoning calculations are performed through the large language model to generate an initial availability region prediction. Based on the reinforcement learning framework, the initial availability region prediction is iteratively optimized and generalized using the thought chain reasoning mechanism to output the target availability region information.
2. The method according to claim 1, characterized in that, The sensing system is equipped with an image acquisition device; the acquisition of multimodal environmental data collected by the sensing system and the preprocessing of the multimodal environmental data include: Instance segmentation processing is performed on the visual image data acquired by the image acquisition device to extract at least one mask region of the target object; Perform semantic parsing processing on the task description text to extract the semantic features of the action intent corresponding to the task instructions; The masked region is associated with the semantic features of the action intent across modalities, and the visual features and semantic features are aligned in the vector space through a cross-modal attention mechanism to construct the input representation of the availability reasoning task.
3. The method according to claim 1, characterized in that, The aforementioned availability prior knowledge refers to the essential logical knowledge pre-stored in the large language model regarding the adaptation relationship between the functional attributes of objects and the intention of actions; The step of constructing an input representation for the ability-to-reason task based on the preprocessed visual image data and task description text, by introducing prior knowledge of availability, includes: Based on the aforementioned prior knowledge of availability, a thought chain prompt template is constructed, which is used to guide the large language model to perform availability logic deduction. The visual image data, the task description text, and the thought chain prompt template are integrated to generate the input representation of the availability reasoning task.
4. The method according to claim 3, characterized in that, The large language model is equipped with a dedicated affordance inference module; the step of performing affordance inference calculations through the large language model to generate initial affordance region predictions includes: The availability-specific reasoning module performs hierarchical reasoning derivation on the availability reasoning task input representation. After the action adaptation region recommendation stage is completed, a reflection and verification subprocess is triggered to check the logical consistency between the recommended region and the availability functional attribute. Based on the output of the hierarchical reasoning derivation and the reflection verification subprocess, the initial availability region prediction containing regional coordinate information and availability labels is generated; The hierarchical reasoning process includes an object functional attribute identification stage and an action adaptation region recommendation stage. In the object functional attribute identification stage, the structural features of the target object are analyzed and associated with available functional attributes that match the action intent in the task description text. In the action adaptation region recommendation stage, based on the identified available functional attributes, physical regions in the visual image data that can perform the action intent are located and recommended.
5. The method according to claim 1, characterized in that, The reinforcement learning framework is configured with a hybrid reward function, and the thought chain reasoning mechanism is configured to generate a structured reasoning chain containing outputs of a thinking phase, a reflection phase, and a response phase. The step of iteratively optimizing and generalizing the initial availability region prediction using the thought chain reasoning mechanism, based on the reinforcement learning framework, to output target availability region information, includes: Multiple candidate answers are generated through the reinforcement learning framework, and each candidate answer includes thinking content, reflection content and answer content generated based on the thought chain reasoning mechanism. Grouping relative strategy optimization calculations are performed on the multiple candidate answers, and the reasoning quality of each candidate answer is evaluated and ranked based on the hybrid reward function; Based on the evaluation and ranking results, the candidate answer with the highest comprehensive reward score is selected as the target availability region information.
6. The method according to claim 5, characterized in that, The reinforcement learning framework is also configured with a grouping relative policy optimization algorithm. The step of performing grouping relative policy optimization calculation on the multiple candidate answers includes: Calculate the strategy probability loss term for each candidate answer; Calculate the relative entropy penalty term for the model output distribution corresponding to each candidate answer; Based on the reward value calculated by the hybrid reward function, the relative advantage score of each candidate answer within the group is determined. Based on the policy probability loss term, the model output distribution relative entropy penalty term, and the group relative advantage score, construct the group relative policy optimization total loss function; The model parameters of the large language model are updated by optimizing the total loss function by minimizing the grouping relative strategy.
7. The method according to claim 6, characterized in that, The hybrid reward function includes format reward, perceptual reward, and cognitive reward; the step of evaluating and ranking the reasoning quality of each candidate answer based on the hybrid reward function includes: Calculate the format reward score for the candidate answer, which is determined based on the completeness and correctness of the structured tags in the candidate answer; Calculate the perceived reward score of the candidate answer, which is determined based on the spatial intersection-union ratio of the availability region coordinates in the candidate answer and a preset reference region; Calculate the cognitive reward score of the candidate answer, which is determined based on the cosine similarity between the word vectors of the availability tag semantics and the preset reference tag in the candidate answer; The comprehensive reward score of the candidate answer is calculated based on the weighted sum of the format reward score, the perception reward score, and the cognitive reward score. The candidate answers are ranked based on the comprehensive reward score.
8. An availability generalization reasoning system based on a reinforcement learning framework, characterized in that, The availability generalization reasoning system is based on a robot equipped with a perception system, which integrates a binocular stereo vision camera supporting high-precision depth information perception, including: The multimodal environment data acquisition module is configured to acquire multimodal environment data collected by the robot perception system and perform preprocessing on the multimodal environment data, which includes visual image data captured by the binocular stereo vision camera and task description text. The availability input representation construction module is configured to introduce availability prior knowledge to construct availability reasoning task input representation based on the preprocessed visual image data and task description text; The initial availability region prediction module is configured to input the availability reasoning task input representation into a preset large language model, perform availability reasoning calculations through the large language model, and generate an initial availability region prediction. The availability reasoning result output module is configured to perform iterative optimization and generalization enhancement processing on the initial availability region prediction based on a reinforcement learning framework and using a thought chain reasoning mechanism to output target availability region information.
9. A computer device, characterized in that, include: At least one processor; and a memory that is communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, Computer instructions are used to cause a computer to perform the method according to any one of claims 1-7.