Robotic dexterous manipulation system, method, apparatus and media for transparent objects

By constructing a noisy hand-object interaction point cloud and performing multimodal feature fusion, the problem of robot dexterity under occlusion and transparent object interference was solved, achieving stable operation and high adaptability, and improving the robot's ability to understand the shape and pose of transparent objects.

CN121613825BActive Publication Date: 2026-05-01UNIV OF SCI & TECH OF CHINA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UNIV OF SCI & TECH OF CHINA
Filing Date
2026-01-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies suffer from reduced reliability of 3D representations for robot dexterity when faced with occlusion and interference from transparent objects. Furthermore, relying on intermediate perception optimization methods can easily introduce errors and makes it difficult to adapt to dynamic scenarios.

Method used

A point cloud data construction module is used to generate a noisy hand-object interaction point cloud. High-dimensional latent features are extracted through a point cloud feature extraction module and a query prediction module. Combined with a perception coding module and a multimodal feature fusion module, self-attention and cross-attention fusion of multi-source perception features are achieved to generate differentiated motion commands to control the robotic arm and dexterous hand.

Benefits of technology

Stable robot dexterity was achieved even with incomplete perception information, enhancing operational stability and adaptability under occluded and transparent objects, avoiding errors from intermediate perception completion modules, and improving the model's generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121613825B_ABST
    Figure CN121613825B_ABST
Patent Text Reader

Abstract

The application discloses a kind of robot dexterous operation systems, methods, equipment and media for transparent object, belong to robot dexterous operation and three-dimensional perception field, system includes: point cloud data construction module, can be constructed and noisy to obtain noisy hand-object interaction point cloud by physical simulation;Point cloud feature extraction module can extract high-dimensional latent features from noisy hand-object interaction point cloud;Query prediction module can infer object shape and pose by query point generation and decoder decoding;Perception coding module can extract features from global point cloud, hand-object interaction point cloud and tactile information to obtain multi-source perception features;Multi-modal feature fusion module can realize feature integration through self-attention mechanism and cross-attention mechanism calculation, and output fusion features;Action generation strategy module can generate differentiated action instructions for robot arm and dexterous hand based on fusion features.The system can achieve higher adaptability and stronger generalization capability in transparent object operation task.
Need to check novelty before this filing date? Find Prior Art

Description

Dexterous operating systems, methods, devices and media for robots targeting transparent objects Technical Field

[0001] This invention relates to the field of robot dexterity manipulation and three-dimensional perception technology, and in particular to a robot dexterity operating system and method for transparent objects. Background Technology

[0002] Dexterous manipulation is a core capability for robots to efficiently perform daily services and industrial tasks. Vision-based end-to-end strategies, especially those utilizing 3D point clouds, have shown potential in various manipulation tasks due to their ability to provide more robust spatial information. However, the limited perception capabilities of visual information regarding contact states and interaction details restrict operational reliability. To address this, a visual-tactile fusion imitation learning framework has been introduced, supplementing contact information with tactile sensors to improve operational accuracy.

[0003] Nevertheless, the self-occlusion and depth perception ambiguity caused by transparent objects during dexterous hand operations still reduce the reliability of 3D representation. Existing methods that rely on intermediate perception optimization are prone to introducing errors and are difficult to adapt to dynamic scenes. Therefore, it is an issue that needs to be addressed to provide a system and method that can directly address perception ambiguity issues such as occlusion and interference from transparent objects, and thus achieve stable robot dexterity operations even with incomplete perception information.

[0004] In view of this, the present invention is hereby proposed. Summary of the Invention

[0005] The purpose of this invention is to provide a robot dexterous operating system and method for transparent objects, which can directly address perceptual ambiguity issues such as occlusion and interference from transparent objects without relying on independent perceptual completion. This enables stable robot dexterous operation even with incomplete perceptual information, thereby solving the aforementioned technical problems in the prior art.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] A dexterous operating system for robots oriented towards transparent objects, comprising:

[0008] The point cloud data construction module can construct a robot's dexterous grasping scene through physical simulation, sample from it and add noise to generate a noisy hand-object interaction point cloud;

[0009] The point cloud feature extraction module, connected to the point cloud data construction module, can preprocess the noisy hand-object interaction point cloud into local feature blocks, perform random masking on the local feature blocks, and extract high-dimensional latent features through a self-attention mechanism encoder.

[0010] The query prediction module, connected to the point cloud feature extraction module, can mark the object point cloud in the noisy hand-object interaction point cloud as positive sample query points and other point clouds in the environment as negative sample query points. The decoder combines the high-dimensional latent features to output the object membership confidence of the query point to infer the object shape and pose.

[0011] The perception encoding module, connected to the point cloud feature extraction module, can extract features from the global point cloud acquired by the robot's camera, the noisy hand-object interaction point cloud, and the robot's tactile information, respectively, and obtain and output multi-source perception features.

[0012] The multimodal feature fusion module, connected to the perceptual coding module, can use self-attention and cross-attention mechanisms to optimize the multi-source perceptual feature representation layer by layer and output fused features.

[0013] The motion generation strategy module, connected to the multimodal feature fusion module, can generate differentiated motion commands based on the fused features, and output motion sequences that control the robot's robotic arm and dexterous hand respectively.

[0014] A method for a robot dexterous operating system for transparent objects as described in this invention includes:

[0015] The system's point cloud data construction module constructs a smart grasping scene through physical simulation, samples it from the scene and adds noise to generate a noisy hand-object interaction point cloud.

[0016] The point cloud feature extraction module of the system preprocesses the noisy hand-object interaction point cloud into local feature blocks. After performing random masking on the local feature blocks, high-dimensional latent features are extracted through a self-attention mechanism encoder.

[0017] The system's query prediction module marks the object point cloud in the noisy hand-object interaction point cloud as a positive sample query point and other point clouds in the environment as negative sample query points. The decoder combines the high-dimensional latent features to output the object membership confidence of the query point to infer the object's shape and pose.

[0018] The system's perception coding module extracts features from the robot's global point cloud, hand-object interaction point cloud, and tactile information acquired by the robot's camera, respectively, and outputs multi-source perception features.

[0019] The system's multimodal feature fusion module uses self-attention and cross-attention mechanisms to optimize multi-source perception feature representations layer by layer and outputs fused features.

[0020] The system's motion generation strategy module generates differentiated motion commands based on fused features, and outputs motion sequences to control the robotic arm and dexterous hand respectively.

[0021] A processing apparatus, comprising:

[0022] At least one memory for storing one or more programs;

[0023] At least one processor is capable of executing one or more programs stored in the memory, such that when the processor executes one or more programs, the processor can implement the method of the present invention.

[0024] A readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the methods described in this invention.

[0025] Compared with the prior art, the beneficial effects of the robot dexterous operating system, method, device, and medium for transparent objects provided by the present invention include:

[0026] By combining point cloud data construction, point cloud feature extraction, query prediction, perception encoding, multimodal feature fusion, and motion generation strategy modules, a point cloud reconstruction pre-training and end-to-end visual-tactile fusion motion generation strategy are achieved, forming a robot control system for dexterous manipulation of transparent objects. This system addresses the limited ability of models to understand the shape and pose of transparent objects under sparse observation and perceptual ambiguity by recovering the complete geometric structure of the object from noisy and masked interactive point clouds. Furthermore, by decomposing tactile information into position point clouds and array force signals and performing fine-grained encoding and multi-level attention fusion, differentiated collaborative control of the robotic arm and dexterous hand is provided, effectively enhancing operational stability and adaptability under complex interferences such as occlusion and reflection. The end-to-end motion generation mechanism eliminates the need for intermediate perception completion modules, ensuring higher adaptability and stronger generalization capabilities in transparent object manipulation tasks. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 is a schematic diagram of the structure of a robot dexterous operating system for transparent objects provided in an embodiment of the present invention.

[0029] Figure 2 is a flowchart of the 3D point cloud reconstruction pre-training process in the system provided by the embodiment of the present invention.

[0030] Figure 3 is a schematic diagram of the motion generation process of visual-tactile fusion in the system provided by the embodiment of the present invention.

[0031] Figure 4 is a flowchart of the 3D point cloud reconstruction pre-training process in the system provided by the embodiment of the present invention.

[0032] Figure 5 is a flowchart of the motion generation process of visual-tactile fusion in the system provided by the embodiment of the present invention. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them, and do not constitute a limitation on the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the protection scope of the present invention.

[0034] First, the following explanations are provided for the terms that may be used in this article:

[0035] The term "and / or" means that either or both can be achieved simultaneously. For example, X and / or Y means that it includes both "X" or "Y" as well as the three cases of "X and Y".

[0036] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0037] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0038] Unless otherwise explicitly specified or limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this document according to the specific circumstances.

[0039] The terms “center,” “longitudinal,” “lateral,” “length,” “width,” “thickness,” “up,” “down,” “front,” “back,” “left,” “right,” “vertical,” “horizontal,” “top,” “bottom,” “inner,” “outer,” “clockwise,” and “counterclockwise” indicate the current orientation or positional relationship, and are only for the convenience and simplification of description, and do not explicitly or implicitly suggest that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this document.

[0040] The technical solution provided by this invention will be described in detail below. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they shall be performed according to conventional conditions in the art or conditions recommended by the manufacturer. Reagents or instruments used in the embodiments of this invention whose manufacturers are not specified are all conventional products that can be purchased commercially.

[0041] As shown in Figure 1, this invention provides a robot dexterous operating system for transparent objects, suitable for humanoid arm-hand robots operating in sparse observation conditions, enabling motion generation through visual-tactile fusion. Specifically, the system includes: a point cloud data construction module, a point cloud feature extraction module, a query prediction module, a perception encoding module, a multimodal feature fusion module, and a motion generation strategy module.

[0042] As shown in Figure 2, the point cloud data construction module, point cloud feature extraction module, and query prediction module in the above system constitute a 3D point cloud reconstruction pre-training framework. This framework provides the entire system with fundamental perception capabilities, enabling the robot to accurately understand the geometric characteristics of objects from sparse observations. This framework achieves this goal through the collaborative work of the following modules:

[0043] The point cloud data construction module can construct a robot's dexterous grasping scenario through physical simulation, sample from it, and add noise to generate a noisy hand-object interaction point cloud. Specifically, it uses a physical simulation engine to build a robot's dexterous grasping environment covering everyday life and professional experimental scenarios, including objects of various types and shapes, ensuring the scenario generalization and object diversity of the dataset. By uniformly sampling the dexterous hand and object surfaces, it generates interactive point cloud data that completely records the hand-object contact state, and performs multi-parameter random noise addition processing on the object point cloud to simulate sensor noise interference in the real environment. The generated point cloud data is then filtered, retaining only the dexterous hand point cloud and discarding the robotic arm point cloud, to guide subsequent models to focus on learning the object reconstruction rules of hand-object interaction configuration.

[0044] Preferably, the aforementioned point cloud data construction module includes: a scene construction submodule, a point cloud sampling submodule, a noise processing submodule, and a data filtering submodule; wherein,

[0045] The scene building submodule can build robot dexterous grasping scenes containing daily necessities and chemical laboratory instruments through a physical simulation engine, and select a variety of objects covering daily necessities and professional experimental instruments.

[0046] The point cloud sampling submodule is communicatively connected to the scene building submodule and can uniformly sample the surface of the robot's dexterous hand and the surface of objects in the scene to generate a complete hand-object interaction point cloud.

[0047] The noise-adding submodule is communicatively connected to the point cloud sampling submodule. It can perform random noise-adding processing on the object point cloud in the hand-object interaction point cloud generated by the point cloud sampling submodule. The noise-adding parameters include direction, percentage, and noise level. Each noise-adding parameter is set in a variety of ways to fully simulate the real noise environment. The noise-adding hand-object interaction point cloud is then filtered to obtain a noisy hand-object interaction point cloud that retains only the dexterous hand point cloud and discards the robotic arm point cloud. This ensures that the model's learning focus is on the object reconstruction information related to hand-object interaction.

[0048] The point cloud feature extraction module, communicatively connected to the point cloud data construction module, preprocesses the noisy hand-object interaction point cloud into local feature blocks. After performing random masking on the local feature blocks, a self-attention mechanism encoder extracts high-dimensional latent features. Specifically, the noisy hand-object interaction point cloud is normalized to eliminate the influence of scale differences. A sampling algorithm selects a center point from the normalized noisy hand-object interaction point cloud. A neighborhood aggregation algorithm is used to aggregate related points around the center point into local feature blocks. Random masking is performed on the local feature blocks to force the model to learn the correlation between global and local features. The visible local feature blocks are fused with the center point position and converted into a labeled sequence. Learnable clustering labels are introduced, and a self-attention mechanism encoder extracts high-dimensional latent features from the labeled sequence. The self-attention mechanism encoder mines complex correlations in the point cloud and outputs high-dimensional latent features that combine local details and global context.

[0049] Preferably, the point cloud feature extraction module described above includes:

[0050] The module comprises a point cloud preprocessing submodule, a local feature block construction submodule, a mask processing submodule, and a feature encoding submodule; among which,

[0051] The point cloud preprocessing submodule is connected to the point cloud data construction module and can normalize the noisy hand-object interaction point cloud to unify the data scale standard.

[0052] The local feature block construction submodule is communicatively connected to the point cloud preprocessing submodule. It can select the center point from the normalized noisy hand-object interaction point cloud through a sampling algorithm, and aggregate the related points around the center point into a local feature block using a neighborhood aggregation algorithm.

[0053] The masking submodule is connected to the local feature block construction submodule and can perform masking on local feature blocks at random ratios, retaining some visible local feature blocks to enhance the model's learning of global-local feature correlation.

[0054] The feature encoding submodule is communicatively connected to the masking submodule. It can input the preserved visible local feature blocks into the pre-trained encoder, fuse the center point position embedding, and convert it into a label sequence that can be processed by the self-attention mechanism encoder. It introduces learnable clustering labels, mines complex relationships in the point cloud through the self-attention mechanism encoder, and outputs high-dimensional latent features.

[0055] The query prediction module, connected to the point cloud feature extraction module, can mark the object point cloud in the noisy hand-object interaction point cloud as positive sample query points and other point clouds in the environment as negative sample query points. The decoder combines the high-dimensional latent features to output the object membership confidence of the query points to infer the object shape and pose. Specifically, the original object point cloud in the noisy hand-object interaction point cloud is marked as a positive sample query point, and other point clouds in the environment are marked as negative sample query points. The query point position is embedded by the query point encoder and fused with the high-dimensional latent features extracted by the point cloud feature extraction module before being input into the decoder. The binary prediction structure of the decoder outputs the object membership confidence of each query point. A classification loss is constructed based on the object membership confidence of each query point. The training objective is to minimize the classification loss and the feature similarity loss corresponding to the cluster label, so that the point cloud feature extraction module and the query prediction module can infer the shape and spatial pose of the object under noise and masking conditions.

[0056] Preferably, the query prediction module described above includes:

[0057] The module includes a query point generation submodule, a decoding submodule, and a loss optimization submodule; among which,

[0058] The query point generation submodule can mark the object point cloud in the noisy hand-object interaction point cloud constructed by the point cloud data construction module as positive sample query points, and mark other points in the environment and random points generated based on the extended range of the interaction point cloud as negative sample query points, and the negative sample points include the dexterous hand point cloud.

[0059] The decoding submodule is communicatively connected to the query point generation submodule and the point cloud feature extraction module, respectively. It can obtain the position embedding of the query point through the encoder, fuse the position embedding with the high-dimensional latent features output by the point cloud feature extraction module and input it into the decoder. The decoder outputs the object membership confidence of each query point through the binary prediction structure.

[0060] The loss optimization submodule, communicatively connected to the decoding submodule, constructs a classification loss based on the object membership confidence of each query point. With the training objective of minimizing the classification loss and the feature similarity loss corresponding to the cluster labels, it optimizes encoder performance, enabling the point cloud feature extraction module and query prediction module to robustly infer the object's shape and spatial pose under noisy and masked conditions. This ensures the robustness of the point cloud feature extraction module and query prediction module in inferring the object's shape and spatial pose under complex noise and masked conditions.

[0061] As shown in Figure 3, the perception encoding module, multimodal feature fusion module, and action generation strategy module in the above system constitute a multimodal motion generation framework. This framework, based on the point cloud feature extraction module in the aforementioned 3D point cloud reconstruction pre-training framework, extracts and integrates multimodal information to achieve differentiated action prediction between the robotic arm and the dexterous hand, ensuring operational efficiency and accuracy. This framework is specifically implemented through the following modules:

[0062] The aforementioned perception encoding module, communicatively connected to the point cloud feature extraction module, can extract features from the global point cloud and hand-object interaction point cloud acquired by the robot's camera, as well as the tactile information acquired by the robot, to obtain and output multi-source perception features. Specifically, it performs targeted encoding processing on the global point cloud, hand-object interaction point cloud, and tactile information respectively. For the global point cloud, it is first unified to the robot's base coordinate system, and after registration and fine-tuning to eliminate errors, the data format is cropped and sampled to be regularized, retaining the core information and encoding it as global features. For the hand-object interaction point cloud, the pre-trained encoder of the point cloud feature extraction module is reused to extract clustering features, and then the fully connected layer connected to the pre-trained encoder adjusts the feature dimension and distribution to meet the requirements of downstream tasks. For tactile information, the spatial position of the touch point is calculated based on the joint motion information and supplemented into the interaction point cloud. The tactile force signal is restructured and encoded as tactile force features, and the touch point position is encoded separately to provide a contact posture reference.

[0063] Preferably, the above-mentioned perceptual coding module includes:

[0064] The system includes a global point cloud encoding submodule, an interactive point cloud encoding submodule, and a tactile information encoding submodule; among which,

[0065] The global point cloud encoding submodule can unify the global point cloud to the robot base coordinate system, perform registration and fine-tuning through the iterative nearest point algorithm to eliminate errors, and after cropping and sampling, it is normalized to a fixed number of points, retaining position and color information and encoding it as global point cloud features;

[0066] The interactive point cloud encoding submodule is communicatively connected to the point cloud feature extraction module and can use a pre-trained encoder to extract the clustering features of the hand-object interaction point cloud as the hand-object interaction point cloud features.

[0067] By adapting the network to adjust the feature dimensions and distribution, it can be adapted to downstream multimodal fusion tasks;

[0068] The tactile information encoding submodule can calculate the spatial position of the touch point based on the robot's joint motion information and supplement it to the hand-object interaction point cloud. After the robot's tactile force signal is restructured into an image, it is encoded into tactile force features by a tactile force encoder composed of a convolutional neural network. The touch point position of the tactile force signal is separately encoded into tactile position features, providing a reference for contact posture control.

[0069] The multimodal feature fusion module, communicatively connected to the perceptual encoding module, uses self-attention and cross-attention mechanisms to progressively optimize multi-source perceptual feature representations and output fused features. Specifically, it learns the internal associations of visual modalities through the self-attention mechanism to generate preliminary fused features. Then, using these preliminary fused features as keys and values, and tactile force features as queries, the cross-attention mechanism modulates visual features with tactile information. This process is repeated to progressively optimize feature representations and output fused features. A deep fusion architecture is constructed through the attention mechanism to achieve the organic integration of visual and tactile features. First, global point cloud features, hand-object interaction point cloud features, and tactile position features are input into the self-attention module to learn the internal associations of visual modalities and output preliminary fused features focusing on global motion. Then, using these preliminary fused features as keys and values, and tactile force features as queries, the cross-attention module modulates visual features with tactile information and outputs cross-modal fused features focusing on fine interaction. Through multiple rounds of iterative optimization of feature representation, the self-attention module outputs features to adapt to the robotic arm's motion guidance, and the cross-attention module outputs features to adapt to the dexterous hand's fine control, thereby achieving differentiated feature output.

[0070] Preferably, the above-mentioned multimodal feature fusion module includes:

[0071] The system comprises a visual modality fusion submodule, a cross-modal modulation submodule, and a feature optimization submodule; among which,

[0072] The aforementioned visual modality fusion submodule is communicatively connected to the perception encoding module. It can receive global point cloud features, hand-object interaction point cloud features, and tactile position features. Through the self-attention module set in the visual modality fusion submodule, it learns the internal correlation of visual modalities and outputs preliminary fusion features to adapt to the robot's robotic arm action guidance.

[0073] The cross-modal modulation submodule is communicatively connected to the visual modal fusion submodule and the perceptual encoding module, respectively. It can use the preliminary fusion features as keys and values, use the tactile force features output by the perceptual encoding module as queries, and input the cross-attention module set in the cross-modal modulation submodule to modulate the visual features with tactile information, and output cross-modal fusion features adapted to the robot's dexterous hand for fine control.

[0074] The feature optimization submodule is connected to both the visual modality fusion submodule and the cross-modality modulation submodule. It can repeatedly invoke these submodules for corresponding modulation processing, progressively optimizing both the initial full-scale fusion features and the cross-modality fusion features. The optimized features yield globally oriented fusion features and finely detailed interactive fusion features. This enhances the differentiation between global motion-oriented and locally finely detailed interactive features, providing support for subsequent differentiated action generation.

[0075] The motion generation strategy module, communicatively connected to the multimodal feature fusion module, generates differentiated motion commands based on fused features, outputting motion sequences to control the robot's robotic arm and dexterous hand respectively. Specifically, based on differentiated fused features, targeted generation strategies are adopted for the motion characteristics of the robotic arm and dexterous hand. For the robotic arm, a direct prediction strategy is used, outputting a low-dimensional end-effector pose sequence based on globally guided fused features through a fully connected network, ensuring high motion efficiency. For the dexterous hand, a conditional diffusion strategy is used, constrained by refined interactive fused features and the robot's own state, performing multi-step optimization on random initial signals to generate high-dimensional joint motion sequences, ensuring operational accuracy. By balancing efficiency and accuracy through differentiated strategies, collaborative linkage between the robotic arm and dexterous hand is achieved.

[0076] Preferably, the above-mentioned action generation strategy module includes:

[0077] The module includes a robotic arm motion prediction submodule and a dexterous hand motion prediction submodule; among which...

[0078] The robotic arm motion prediction submodule is communicatively connected to the visual modality fusion submodule of the multimodal feature fusion module. It can use a prediction network composed of multi-layer sensing mechanisms to directly predict the pose sequence of the robotic arm end effector based on the optimized global guidance fusion features, and output a low-dimensional motion sequence to control the robotic arm, thus meeting the real-time and high-efficiency requirements of the overall motion.

[0079] The dexterous hand motion prediction submodule is communicatively connected to the cross-modal modulation submodule of the multimodal feature fusion module. It can use a conditional denoising diffusion model as the prediction network, and perform multi-step denoising processing on random Gaussian noise based on the optimized fine interactive fusion features and the robot's body state, thereby generating a high-dimensional motion sequence to control the dexterous hand and ensuring the accuracy and flexibility of fine operations.

[0080] By combining the differentiated design and collaborative work of the 3D point cloud reconstruction pre-training framework and the multimodal fusion motion generation framework composed of the above two types of sub-modules, the robot's motion efficiency and dexterity hand operation accuracy are balanced, ensuring that the robot can achieve integrated operation with rapid response and precise execution under sparse observation conditions.

[0081] This invention also provides a method for a robot dexterous operating system for transparent objects as described above, comprising:

[0082] The system's point cloud data construction module constructs a smart grasping scene through physical simulation, samples it from the scene and adds noise to generate a noisy hand-object interaction point cloud.

[0083] The point cloud feature extraction module of the system preprocesses the noisy hand-object interaction point cloud into local feature blocks. After performing random masking on the local feature blocks, high-dimensional latent features are extracted through a self-attention mechanism encoder.

[0084] The system's query prediction module marks the object point cloud in the noisy hand-object interaction point cloud as a positive sample query point and other point clouds in the environment as negative sample query points. The decoder combines the high-dimensional latent features to output the object membership confidence of the query point to infer the object's shape and pose.

[0085] The system's perception coding module extracts features from the robot's camera-acquired global point cloud, noisy hand-object interaction point cloud, and tactile information, respectively, and outputs multi-source perception features.

[0086] The system's multimodal feature fusion module uses self-attention and cross-attention mechanisms to optimize multi-source perception feature representations layer by layer and outputs fused features.

[0087] The system's motion generation strategy module generates differentiated motion commands based on fused features, and outputs motion sequences to control the robotic arm and dexterous hand respectively.

[0088] The present invention further provides a processing apparatus, comprising:

[0089] At least one memory for storing one or more programs;

[0090] At least one processor is capable of executing one or more programs stored in the memory, such that when the processor executes one or more programs, the processor can implement the methods described above.

[0091] The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the above-described method.

[0092] In summary, the operating system and method of the present invention have at least the following advantages:

[0093] Significantly Enhanced Perceptual Robustness and Geometric Understanding: Through an innovative point cloud reconstruction pre-training mechanism, this system enables the encoder to robustly infer the complete geometry and pose of objects from incomplete and noisy point clouds. This pre-training process employs random masking and diverse noise addition strategies to force the model to learn the correlation between global and local features, thereby improving perceptual stability in challenging scenarios such as transparent objects and partial occlusion. The pre-trained feature extraction module can directly reconstruct the essential geometric information of objects from sparse point clouds containing dexterous hands. This design replaces the independent perceptual completion module in traditional solutions, effectively reducing error propagation and computational latency. Furthermore, since the pre-training data extensively covers objects in various scenarios such as homes and laboratories, it ensures the model has strong generalization capabilities for unknown environments.

[0094] Multimodal deep fusion and arm-hand collaborative optimization: In terms of technical implementation, the system adopts a cascaded structure of self-attention and cross-attention modules, integrating global point cloud features, interactive point cloud features, tactile position information, and force information layer by layer, ultimately forming a feature representation that combines overall environmental cognition with fine interactive details. For the different motion characteristics of the robotic arm and dexterous hand, the system designs different policy networks for them respectively. This differentiated prediction strategy ensures the natural coordination of arm-hand movements in both time and space.

[0095] High efficiency and real-time performance of the end-to-end architecture: Leveraging the integrated design of pre-training and motion generation, this invention forms an end-to-end framework that eliminates the need for intermediate perception completion modules, demonstrating superior system efficiency. This architecture can directly generate control commands based on multimodal perception input, avoiding redundant processing steps, significantly reducing system complexity, and meeting the stringent requirements for rapid response in dynamic environments.

[0096] To more clearly demonstrate the technical solution and its effects provided by the present invention, the following detailed description of the solution provided by the embodiments of the present invention is provided with reference to specific examples.

[0097] Example 1

[0098] As shown in Figure 1, this embodiment provides a robot dexterous operating system for transparent objects, including: a point cloud data construction module, a point cloud feature extraction module, a query prediction module, a perception encoding module, a multimodal feature fusion module, and an action generation strategy module; wherein the point cloud data construction module, the point cloud feature extraction module, and the query prediction module constitute a 3D point cloud reconstruction pre-training framework (see Figure 2), and the perception encoding module, the multimodal feature fusion module, and the action generation strategy module constitute a multimodal fusion motion generation framework (see Figure 3).

[0099] Referring to Figure 2, the specific implementation of the above-mentioned 3D point cloud reconstruction pre-training framework is as follows:

[0100] This 3D point cloud reconstruction pre-training framework demonstrates robust understanding of object shape and pose under sparse observation conditions. This embodiment uses common objects from everyday life scenarios as the operational objects. Through the collaborative linkage of a point cloud data construction module, a point cloud feature extraction module, and a query prediction module for constructing a dexterous hand interactive dataset, it achieves the technical goal of accurately reconstructing the complete geometric shape of objects from noisy sparse point clouds.

[0101] As shown in Figure 4, the specific implementation process of this framework is as follows:

[0102] First, the point cloud data construction module completes the generation and optimization of training data.

[0103] A high-fidelity robot dexterous grasping scenario was built using a physics simulation engine. The types, sizes, and spatial positions of objects in the scenario were diverse, ensuring the dataset covered the interaction needs of different forms and operational scenarios. Within the constructed scenario, the surfaces of the dexterous hand and the object were uniformly sampled to generate complete hand-object interaction point cloud data that records the hand-object contact state and spatial pose relationship.

[0104] To simulate sensor noise interference in real-world operating environments, multi-parameter random noise addition processing was performed on the object point cloud in the hand-object interaction point cloud. The noise addition direction covered multiple dimensions, including radial and tangential, and the noise addition ratio and noise intensity were set in a manner consistent with the noise distribution characteristics of real sensors, fully reproducing the noise scenarios in actual applications.

[0105] By retaining only the point cloud data related to the dexterous hand and discarding the point cloud of the robotic arm to obtain the noisy hand-object interaction point cloud, the model is guided to focus on the core information of the hand-object interaction configuration, and to concentrate on learning the mapping rules of inferring the shape of the object from the interaction posture, thus avoiding the interference of irrelevant data on the training effect.

[0106] Subsequently, the point cloud feature extraction module extracts and enhances the features of the noisy hand-object interaction point cloud.

[0107] First, the noisy hand-object interaction point cloud is preprocessed by normalization to eliminate interference from different object sizes and scene scales, providing standardized input data for subsequent feature extraction. Then, a suitable number of center points are selected from the normalized noisy hand-object interaction point cloud using a sampling algorithm. Next, a neighborhood aggregation algorithm is used to aggregate the associated points around each center point to form structured local feature blocks, achieving hierarchical representation of the point cloud data.

[0108] To enhance the model's ability to learn the correlation between global and local features, the generated local feature blocks are randomly masked, retaining only some visible local feature blocks. This forces the model to fill in local features with global context information even when some local information is missing, thereby improving the robustness and generalization ability of feature extraction.

[0109] For visible local feature blocks that are not masked, they are input into a pre-trained encoder, which simultaneously fuses the embedding information of the center point's position, transforming the structured point cloud data (i.e., local feature blocks) into a labeled sequence that can be processed by a self-attention mechanism encoder. To achieve effective aggregation of global features, learnable clustering labels are introduced. Through the self-attention mechanism encoder, the dependencies between local features and the mapping relationships between global and local features in the point cloud data are deeply mined. Finally, high-dimensional latent features with both local detail accuracy and global context integrity are output, providing core feature support for subsequent object shape and pose inference.

[0110] Finally, the model training and optimization are completed through the query prediction module.

[0111] The query prediction module takes the query point classification task as its core and constructs a set of positive and negative sample query points: the object point cloud in the noisy hand-object interaction point cloud is defined as a positive sample query point, and other points in the environment and random points generated based on the extended range of the interaction point cloud are defined as negative sample query points.

[0112] To ensure the accuracy of sample labeling, distance verification is performed on the generated negative sample points. If the minimum distance between a negative sample point and a positive sample point is less than the average distance between positive sample points, the negative sample point is reclassified as a positive sample to avoid affecting the model training effect due to sample labeling errors.

[0113] After the query point encoder obtains the location embedding information, it is fused with the high-dimensional latent features output by the point cloud feature extraction module, and then input into the decoder. The decoder's binary prediction structure outputs the confidence score of each query point belonging to the target object.

[0114] The 3D point cloud reconstruction pre-training framework employs a dual-loss collaborative optimization strategy during training. On one hand, classification loss ensures the accuracy of query point classification results and strengthens the ability to recognize object boundaries and shapes. On the other hand, feature similarity loss enhances the consistent representation of object features and improves resistance to noise and masking interference. Through the collaborative optimization of dual losses, it can still accurately and robustly infer the complete shape and spatial pose of objects under complex sparse observation conditions with noise interference and local masking, fully verifying the effectiveness of the 3D point cloud reconstruction pre-training framework of this invention.

[0115] Referring to Figure 3, the multimodal fusion motion generation framework described above can realize differentiated motion generation strategies for visual-tactile fusion, as detailed below:

[0116] This multimodal fusion motion generation framework can verify the enhancement effect of deep integration of multimodal information on the precision operation of robots. In this embodiment, a round-bottomed beaker in a chemical laboratory scenario is used as a typical operation object. Through the collaborative work of the perception encoding module, the multimodal feature fusion module and the motion generation strategy module, efficient collaborative control of the robotic arm and the dexterous hand is achieved, realizing the technical goal of "global efficient positioning of the robotic arm + local fine operation of the dexterous hand".

[0117] As shown in Figure 5, the specific implementation process is as follows:

[0118] First, complete the preprocessing and integration of multi-source sensing data.

[0119] Global point cloud data of the laboratory scene is acquired using the robot's depth camera and uniformly converted to the robot's base coordinate system. A registration and fine-tuning algorithm is then used to eliminate hand-eye calibration errors and sensor measurement errors, improving the coordinate accuracy of the global point cloud data. The registered global point cloud is then cropped and sampled to standardize the data format and number of points, while retaining the core information of point cloud position and color, forming a global point cloud that accurately represents the overall scene layout.

[0120] Define a clipping box within a specific range in the dexterous hand base coordinate system, extract the local point cloud of the hand-object interaction area from the global point cloud, i.e., the hand-object interaction point cloud, and remove the color dimension to focus on geometric and interactive information.

[0121] When the robot's tactile sensors detect a contact force signal, the robot calculates the three-dimensional spatial position of the contact point based on the joint motion information of the robot's robotic arm and dexterous hand. This position information is then added to the hand-object interaction point cloud to enrich the semantic information and interaction details of the interaction point cloud, providing a more comprehensive perceptual input for subsequent multimodal fusion.

[0122] Secondly, the perceptual coding module is used to extract targeted features from multi-source perceptual data.

[0123] For global point clouds, the input is a global point cloud encoder constructed by fully connected layers, which extracts global features that can characterize the global scene layout and object spatial distribution. For hand-object interaction point clouds, the pre-trained encoder optimized in the pre-training stage is reused to extract clustering features. The pre-trained encoder is then connected to a fully connected layer to adjust the feature dimension and distribution to meet the needs of downstream multimodal fusion tasks.

[0124] For tactile perception information, a multi-dimensional processing strategy is adopted: the touch point location information is treated as a tactile point cloud, and a tactile point cloud encoder composed of a fully connected network is used to encode it separately into tactile position features, providing a precise reference for contact posture control; the tactile force signal is structurally reshaped, converted into a format suitable for feature extraction, and then upsampled at an appropriate multiple. Finally, a tactile force encoder composed of a convolutional neural network extracts tactile force features that characterize contact force and contact state. Through the above targeted encoding processing, effective representation of global visual information, local interactive visual information, and tactile information is achieved, laying a solid foundation for multimodal deep fusion.

[0125] Subsequently, the visual and tactile features are organically integrated through the multimodal feature fusion module.

[0126] The multimodal feature fusion module adopts an attention mechanism to build a deep fusion architecture. First, the global point cloud features, hand-object interaction point cloud features and tactile position features are input into the self-attention module built into the multimodal feature fusion module. By mining the spatial correlation and semantic dependency between features within the visual modality, the internal optimization and preliminary fusion of visual features are achieved, and the preliminary fusion features focused on global motion planning are output.

[0127] Subsequently, using the initial fusion feature as the association benchmark (key and value) and the tactile force feature as the query signal, the cross-attention module built into the multimodal feature fusion module is input. Through the dynamic modulation of visual features by tactile information, the feature representation related to fine operation is strengthened, and the cross-modal fusion feature focusing on local fine interaction is output.

[0128] The above fusion process can be iterated in multiple rounds, continuously improving the representation ability and relevance of features through layer-by-layer optimization, and finally outputting two types of differentiated fusion features: features optimized by the self-attention module focus on the global scene and overall motion efficiency, which is suitable for the action guidance of robotic arms; features optimized by the cross-attention module integrate the fine interactive information of tactile force, which is adapted to the fine control needs of dexterous hands.

[0129] Finally, differentiated action instructions are output through the action generation strategy module.

[0130] To address the low-dimensional motion space and high-efficiency positioning characteristics of robotic arms, a prediction network composed of multi-layer sensing mechanisms is adopted. Based on the fusion features focused on the global context, a continuous end-efficiency pose sequence is directly output to ensure the real-time performance and efficiency of the robotic arm's motion, meeting the time requirements of the overall operation process.

[0131] To address the high-dimensional motion space and precise control requirements of dexterous hands, a conditional denoising diffusion model is employed as the prediction network. This model uses cross-modal fusion features focusing on fine-grained interactions and robot state information as constraints to perform multi-step optimization on random initial signals, progressively generating high-dimensional joint motion sequences that meet the demands of precision operations. This ensures the accuracy and flexibility of the dexterous hand's manipulation. By employing differentiated motion generation strategies for robotic arms and dexterous hands, a dynamic balance between global motion efficiency and local operational accuracy is achieved, enabling precise tasks such as grasping, moving, and placing round-bottomed beakers.

[0132] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0133] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. A dexterous operating system for robots oriented towards transparent objects, characterized in that, include: The point cloud data construction module can construct a robot's dexterous grasping scene through physical simulation, sample from it and add noise to generate a noisy hand-object interaction point cloud; the point cloud feature extraction module, connected to the point cloud data construction module, can preprocess the noisy hand-object interaction point cloud into local feature blocks, perform random masking on the local feature blocks, and extract high-dimensional latent features through a self-attention mechanism encoder. The query prediction module, connected to the point cloud feature extraction module, can mark the object point cloud in the noisy hand-object interaction point cloud as positive sample query points and other point clouds in the environment as negative sample query points. The decoder combines the high-dimensional latent features to output the object membership confidence of the query point to infer the object shape and pose. The query prediction module includes a query point generation submodule, a decoding submodule, and a loss optimization submodule. The query point generation submodule marks object point clouds in the noisy hand-object interaction point cloud constructed by the point cloud data construction module as positive sample query points, and marks other points in the environment and random points generated based on the extended range of the interaction point cloud as negative sample query points, with the negative sample points including the dexterous hand point cloud. The decoding submodule, connected to both the query point generation submodule and the point cloud feature extraction module, obtains the position embedding of the query points through the encoder, fuses the position embedding with the high-dimensional latent features output by the point cloud feature extraction module, and inputs it into the decoder. The decoder outputs the object membership confidence of each query point through its binary prediction structure. The loss optimization submodule, connected to the decoding submodule, optimizes the loss based on the object membership of each query point. The classification loss is constructed using body membership confidence, with the training objective being to minimize the classification loss and the feature similarity loss corresponding to the cluster labels. This enables the point cloud feature extraction module and the query prediction module to infer the shape and spatial pose of objects under noise and masking conditions. The perception encoding module, connected to the point cloud feature extraction module, can extract features from the global point cloud, hand-object interaction point cloud, and tactile information acquired by the robot's camera, respectively, to obtain and output multi-source perception features. The multi-modal feature fusion module, connected to the perception encoding module, can use self-attention and cross-attention mechanisms to optimize the multi-source perception feature representation layer by layer and output fused features. The action generation strategy module, connected to the multi-modal feature fusion module, can generate differentiated action commands based on the fused features and output action sequences that control the robot's robotic arm and dexterous hand respectively.

2. The robot dexterous operating system for transparent objects according to claim 1, characterized in that, The point cloud data construction module includes: a scene construction submodule, a point cloud sampling submodule, and a noise processing submodule. The scene construction submodule uses a physics simulation engine to build a robot dexterous grasping scene containing everyday items and chemical laboratory instruments. The point cloud sampling submodule, connected to the scene construction submodule, uniformly samples the surfaces of the robot's dexterous hand and objects in the scene to generate a complete hand-object interaction point cloud. The noise processing submodule, also connected to the point cloud sampling submodule, performs random noise processing on the object point clouds generated by the scene construction submodule. Noise parameters include direction, percentage, and noise level, with diverse settings to simulate a real noise environment. The noise-added hand-object interaction point cloud is then filtered to obtain a noisy hand-object interaction point cloud that retains only the dexterous hand point cloud and discards the robotic arm point cloud.

3. The dexterous operating system for robots oriented towards transparent objects according to claim 1 or 2, characterized in that, The point cloud feature extraction module includes: a point cloud preprocessing submodule, a local feature block construction submodule, a masking submodule, and a feature encoding submodule. The point cloud preprocessing submodule, connected to the point cloud data construction module, normalizes the noisy hand-object interaction point cloud. The local feature block construction submodule, also connected to the point cloud preprocessing submodule, selects a center point from the normalized noisy hand-object interaction point cloud using a sampling algorithm and aggregates related points around the center point into local feature blocks using a neighborhood aggregation algorithm. The masking submodule, connected to the local feature block construction submodule, masks the local feature blocks at random ratios, retaining some visible local feature blocks. The feature encoding submodule, connected to the masking submodule, inputs the retained visible local feature blocks into a pre-trained encoder, fuses the center point position embedding, converts it into a labeled sequence that can be processed by a self-attention mechanism encoder, introduces learnable clustering labels, and then mines the complex relationships in the labeled sequence through the self-attention mechanism encoder to output high-dimensional latent features.

4. The dexterous operating system for robots oriented towards transparent objects according to claim 1, characterized in that, The perception encoding module includes: a global point cloud encoding submodule, an interactive point cloud encoding submodule, and a tactile information encoding submodule. The global point cloud encoding submodule unifies the global point cloud to the robot's base coordinate system, performs registration fine-tuning using an iterative nearest-point algorithm, and after cropping and sampling, normalizes it to a fixed number of points, preserving position and color information and encoding it as global point cloud features. The interactive point cloud encoding submodule, connected to the point cloud feature extraction module, uses a pre-trained encoder from the point cloud feature extraction module to extract clustering features of the hand-object interaction point cloud as hand-object interaction point cloud features. The tactile information encoding submodule calculates the spatial position of the touch points based on the robot's joint motion information and supplements it to the hand-object interaction point cloud. It then restructures the robot's tactile force signal into an image and encodes it as tactile force features using a tactile force encoder composed of a convolutional neural network, and separately encodes the touch point position of the tactile force signal as tactile position features.

5. The robot dexterous operating system for transparent objects according to claim 4, characterized in that, The multimodal feature fusion module includes: a visual modality fusion submodule, a cross-modal modulation submodule, and a feature optimization submodule. The visual modality fusion submodule, connected to the perception encoding module, receives global point cloud features, hand-object interaction point cloud features, and tactile position features. Through a self-attention module within this submodule, it learns the internal relationships within the visual modalities and outputs preliminary fused features adapted to guide the robot's arm movements. The cross-modal modulation submodule, connected to both the visual modality fusion submodule and the perception encoding module, uses the preliminary fused features as keys and values. The tactile force features output by the perception encoding module are used as a query and input into the cross-attention module set within the cross-modal modulation submodule to modulate the visual features with tactile information, outputting cross-modal fusion features adapted for the robot's dexterous hand fine control. The feature optimization submodule is connected to the visual modal fusion submodule and the cross-modal modulation submodule respectively, and can repeatedly call the visual modal fusion submodule and the cross-modal modulation submodule to perform corresponding modulation processing, optimizing the two types of feature representations, namely the full preliminary fusion feature and the cross-modal fusion feature, layer by layer, and obtaining the globally guided fusion feature and the fine interactive fusion feature after optimization.

6. The dexterous operating system for robots oriented towards transparent objects according to claim 5, characterized in that, The motion generation strategy module includes a robotic arm motion prediction submodule and a dexterous hand motion prediction submodule. The robotic arm motion prediction submodule, connected to the visual modality fusion submodule of the multimodal feature fusion module, uses a prediction network composed of multilayer perception mechanisms to directly predict the robotic arm's end-effector pose sequence based on optimized global guidance fusion features, outputting a low-dimensional motion sequence to control the robotic arm. The dexterous hand motion prediction submodule, connected to the cross-modal modulation submodule of the multimodal feature fusion module, uses a conditional denoising diffusion model as the prediction network. It performs multi-step denoising processing on random Gaussian noise, based on optimized fine interactive fusion features and the robot's body state, to generate a high-dimensional motion sequence to control the dexterous hand.

7. A method for a robot dexterous operating system for transparent objects as described in any one of claims 1-6, characterized in that, include: The system's point cloud data construction module constructs a smart grasping scene through physical simulation, samples from it, and adds noise to generate a noisy hand-object interaction point cloud. The system's point cloud feature extraction module preprocesses the noisy hand-object interaction point cloud into local feature blocks. After performing random masking on the local feature blocks, a self-attention mechanism encoder extracts high-dimensional latent features. The system's query prediction module marks the object point cloud in the noisy hand-object interaction point cloud as positive sample query points and other point clouds in the environment as negative sample query points. The decoder combines the high-dimensional latent features to output the object membership confidence of the query points to infer the object's shape and pose. The system's perception encoding module extracts features from the robot's global point cloud, the noisy hand-object interaction point cloud, and the robot's tactile information to obtain and output multi-source perception features. The system's multi-modal feature fusion module uses self-attention and cross-attention mechanisms to optimize the multi-source perception feature representation layer by layer, outputting fused features. The system's motion generation strategy module generates differentiated motion commands based on fused features, and outputs motion sequences to control the robotic arm and dexterous hand respectively.

8. A processing apparatus, characterized in that, include: At least one memory for storing one or more programs; At least one processor is capable of executing one or more programs stored in the memory, such that when the one or more programs are executed by the processor, the processor can perform the method of claim 7.

9. A readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement the method of claim 7.

Citation Information

Patent Citations

  • Mechanical arm test tube grabbing method based on transparent object depth completion

    CN116385518A

  • Robot motion planning method and device based on multi-modal information fusion

    CN119550335A

  • Three-dimensional point cloud reconstruction method and application thereof in target detection

    CN120953516A