Robot dynamic grabbing method and system based on multi-modal data fusion
Through multimodal data fusion and reinforcement learning algorithms, robots can accurately identify and capture objects in dynamic environments, solving the problem of failed grabbing in dynamic environments in the existing technology, and achieving efficient dynamic grabbing effect.
Patent Information
- Application Number
- CN202510446797.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing robot crawling methods mainly rely on static environments and cannot effectively deal with object recognition and crawling planning in dynamic environments, especially when objects move, deform or are blocked, resulting in crawling failure or inefficiency.
Through the robot collecting visual data and tactile data of dynamic objects for multimodal fusion, the object position and attitude are identified by the region-proposed object detection algorithm, and combined with the reinforcement learning algorithm to generate the grab path, realizing the state prediction and grabbing of dynamic objects.
It improves the success rate and efficiency of the robot in a dynamic environment, enhances its ability to adapt to environmental changes, accurately identify and predict the moving trajectory of objects, and improves the accuracy and flexibility of grasping.
Smart Images

Figure CN120347735A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robot control, and specifically relates to a robot dynamic grasping method and system based on multi-modal data fusion. Background Art
[0002] Currently, with the rapid development of robot technology, robots are increasingly widely used in many fields, such as logistics, manufacturing, medical treatment, etc. As an important part of them, robot grasping technology can achieve automated operations and task executions, improving production efficiency. However, most of the existing grasping methods are designed based on static environments and mainly rely on the static features of objects for learning and recognition. These methods perform well in fixed environments, but their object recognition and grasping planning capabilities in dynamic environments are significantly insufficient.
[0003] In many actual application scenarios, objects will move, deform due to external forces, or the recognition difficulty increases due to the occlusion of other objects. Traditional grasping methods usually cannot handle these complex changes, resulting in grasping failures or low efficiency. In addition, existing methods often only focus on single-modal data and cannot fully utilize the information of multiple perception signals, thus limiting the accuracy and efficiency of recognition. Summary of the Invention
[0004] In order to solve the problem that in the process of a robot grasping an object, due to the movement, deformation or occlusion of objects in the environment, the existing grasping methods can only learn grasping in a static environment and have insufficient object recognition and grasping planning capabilities in a dynamic environment, the present invention proposes a robot dynamic grasping method based on multi-modal data fusion, including:
[0005] Collecting visual data and tactile data of a dynamic object to be grasped by the robot, and performing multi-modal fusion on the visual data and the tactile data to obtain multi-modal fusion data of the dynamic object to be grasped;
[0006] According to the multi-modal fusion data, using a region proposal-based object detection algorithm to perform object recognition on the dynamic object to be grasped, and obtaining the position and pose information of the dynamic object to be grasped;
[0007] Based on the position and pose information, performing state prediction on the dynamic object to be grasped to obtain state prediction information of the dynamic object to be grasped;
[0008] According to the state prediction information of the dynamic object to be grasped, using a reinforcement learning algorithm to generate a grasping path of the robot.
[0009] Optionally, the performing multi-modal fusion on the visual data and the tactile data to obtain multi-modal fusion data of the dynamic object to be grasped includes:
[0010] Extract features from the visual data through the visual feature selection method to obtain visual feature data;
[0011] Extract features from the tactile data through a one-dimensional convolutional neural network to obtain tactile feature data;
[0012] Adopt an adversarial training mechanism to perform cross-modal mapping on the visual feature data and the tactile feature data to obtain a fused feature vector after mapping;
[0013] Output the multi-modal fusion data of the dynamic object to be grasped according to the fused feature vector after mapping;
[0014] Among them, the visual data includes one or more of the following: color image data, depth image data, optical flow data, and ambient light data;
[0015] The tactile data includes one or more of the following: robotic force feedback data, robotic torque data, grasping pressure data, object surface friction data, and object hardness data.
[0016] Optionally, the expression of the adversarial training mechanism is as follows:
[0017] F fused = G(F visual ; θ g ) + βH(F haptic ; θ h );
[0018] In the formula,
[0019]
[0020] Among them, F fused represents the fused feature vector after mapping; G represents the generator model; F visual represents the visual feature data; θ g represents the parameters of the generator model; β represents the dynamic weight factor; H represents the non-linear transformation model; F haptic represents the tactile feature data; θ h represents the parameters of the non-linear transformation model; Similarity represents the similarity score.
[0021] Optionally, according to the multi-modal fusion data, use a region proposal-based object detection algorithm to perform object recognition on the dynamic object to be grasped to obtain the position and pose information of the dynamic object to be grasped, including:
[0022] Generate candidate region proposals through a selective search algorithm according to the multi-modal fusion data;
[0023] According to the candidate region proposal, use an object detection algorithm to perform object recognition on the dynamic object to be grasped, and obtain the classification result and bounding box coordinate information of the dynamic object to be grasped;
[0024] According to the classification result and bounding box coordinate information of the dynamic object to be grasped, perform pose estimation on the dynamic object to be grasped, and obtain the position and pose information of the dynamic object to be grasped.
[0025] Optionally, the performing state prediction on the dynamic object to be grasped based on the position and pose information to obtain the state prediction information of the dynamic object to be grasped includes:
[0026] Use an autoencoder to perform state feature extraction on the position and pose information to obtain the fused feature representation of the dynamic object to be grasped;
[0027] According to the fused feature representation of the dynamic object to be grasped, use a long short-term memory network combined with a conditional random field to perform trajectory prediction on the dynamic object to be grasped, and obtain the motion trajectory prediction information of the dynamic object to be grasped;
[0028] According to the motion trajectory prediction information of the dynamic object to be grasped, use a deep reinforcement learning algorithm to perform state prediction on the dynamic object to be grasped, and obtain the state prediction information of the dynamic object to be grasped.
[0029] Optionally, the performing trajectory prediction on the dynamic object to be grasped according to the fused feature representation of the dynamic object to be grasped by using a long short-term memory network combined with a conditional random field to obtain the motion trajectory prediction information of the dynamic object to be grasped includes:
[0030] Perform context feature transformation on the fused feature representation of the dynamic object to be grasped to obtain feature sequence data;
[0031] Use a long short-term memory network to perform implicit prediction on the feature sequence data and output the hidden state sequence of the feature sequence data;
[0032] Perform post-processing on the hidden state sequence through a conditional random field to obtain the motion trajectory prediction information of the dynamic object to be grasped.
[0033] Optionally, the generating the grasping path of the robot according to the state prediction information of the dynamic object to be grasped by using a reinforcement learning algorithm includes:
[0034] According to the state prediction information of the dynamic object to be grasped and the pre-collected environmental information of the dynamic object to be grasped, perform scene perception on the dynamic object to be grasped, and obtain the scene state description information of the dynamic object to be grasped;
[0035] According to the described scenario status description information, a dynamic grasping policy network is constructed by using the deep deterministic policy gradient algorithm to obtain a dynamic grasping policy;
[0036] According to the dynamic grasping policy, a path generation algorithm based on imitation learning is used for path generation to obtain the grasping path of the dynamic object to be grasped.
[0037] Based on the same inventive concept, the present invention also provides a robot dynamic grasping system based on multi-modal data fusion, including:
[0038] A multi-modal fusion module for collecting visual data and tactile data of the dynamic object to be grasped through a robot, and performing multi-modal fusion on the visual data and the tactile data to obtain multi-modal fusion data of the dynamic object to be grasped;
[0039] A target recognition module for performing target recognition on the dynamic object to be grasped according to the multi-modal fusion data by using a region proposal-based object detection algorithm to obtain the position and attitude information of the dynamic object to be grasped;
[0040] A state prediction module for performing state prediction on the dynamic object to be grasped based on the position and attitude information to obtain state prediction information of the dynamic object to be grasped;
[0041] A path generation module for generating a grasping path of the robot according to the state prediction information of the dynamic object to be grasped by using a reinforcement learning algorithm.
[0042] Optionally, the multi-modal fusion module includes:
[0043] A visual feature extraction sub-module for extracting features from the visual data by using a visual feature selection method to obtain visual feature data;
[0044] A tactile feature extraction sub-module for extracting features from the tactile data by using a one-dimensional convolutional neural network to obtain tactile feature data;
[0045] A cross-modal mapping sub-module for performing cross-modal mapping on the visual feature data and the tactile feature data by using an adversarial training mechanism to obtain a mapped fusion feature vector;
[0046] A data fusion sub-module for outputting the multi-modal fusion data of the dynamic object to be grasped according to the mapped fusion feature vector;
[0047] Wherein, the visual data includes one or more of the following: color image data, depth image data, optical flow data, and ambient light data;
[0048] The haptic data includes one or more of the following: robotic force feedback data, robotic torque data, grasping pressure data, object surface friction data, and object hardness data.
[0049] Optionally, the expression of the adversarial training mechanism is as follows:
[0050] F fused = G(F visual ; θ g ) + βH(F haptic ; θ h );
[0051] In the formula,
[0052]
[0053] where F fused represents the fused feature vector after mapping; G represents the generator model; F visual represents the visual feature data; θ g represents the parameters of the generator model; β represents the dynamic weight factor; H represents the non - linear transformation model; F haptic represents the haptic feature data; θ h represents the parameters of the non - linear transformation model; Similarity represents the similarity score.
[0054] Optionally, the target recognition module includes:
[0055] A candidate region generation sub - module, configured to generate candidate region proposals according to the multi - modal fusion data through a selective search algorithm;
[0056] A target detection sub - module, configured to perform target recognition on the dynamic object to be grasped according to the candidate region proposals by using a target detection algorithm, and obtain the classification result and bounding box coordinate information of the dynamic object to be grasped;
[0057] An attitude estimation sub - module, configured to perform attitude estimation on the dynamic object to be grasped according to the classification result and bounding box coordinate information of the dynamic object to be grasped, and obtain the position and attitude information of the dynamic object to be grasped.
[0058] Optionally, the state prediction module includes:
[0059] A feature representation sub - module, configured to use an auto - encoder to extract state features from the position and attitude information, and obtain the fused feature representation of the dynamic object to be grasped;
[0060] A trajectory prediction sub-module, configured to perform trajectory prediction on the to-be-grasped dynamic object by using a long short-term memory network combined with a conditional random field according to the fused feature representation of the to-be-grasped dynamic object, so as to obtain motion trajectory prediction information of the to-be-grasped dynamic object;
[0061] A state estimation sub-module, configured to perform state prediction on the to-be-grasped dynamic object by using a deep reinforcement learning algorithm according to the motion trajectory prediction information of the to-be-grasped dynamic object, so as to obtain state prediction information of the to-be-grasped dynamic object.
[0062] Optionally, the trajectory prediction sub-module includes:
[0063] A feature conversion unit, configured to perform context feature conversion on the fused feature representation of the to-be-grasped dynamic object to obtain feature sequence data;
[0064] A hidden prediction unit, configured to perform hidden prediction on the feature sequence data by using a long short-term memory network and output a hidden state sequence of the feature sequence data;
[0065] A post-processing unit, configured to perform post-processing on the hidden state sequence through a conditional random field to obtain motion trajectory prediction information of the to-be-grasped dynamic object.
[0066] Optionally, the path generation module includes:
[0067] Perform scene perception on the to-be-grasped dynamic object according to the state prediction information of the to-be-grasped dynamic object and the pre-collected environment information of the to-be-grasped dynamic object, so as to obtain scene state description information of the to-be-grasped dynamic object;
[0068] Construct a dynamic grasping policy network by using a deep deterministic policy gradient algorithm according to the scene state description information to obtain a dynamic grasping policy;
[0069] Perform path generation by using a path generation algorithm based on imitation learning according to the dynamic grasping policy to obtain a grasping path of the to-be-grasped dynamic object.
[0070] On the other hand, the present invention further provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected through a bus;
[0071] The memory is configured to store one or more programs;
[0072] When the one or more programs are executed by the at least one processor, the method for dynamically grasping a robot based on multi-modal data fusion as described above is implemented.
[0073] On the other hand, the present invention also provides a computer device-readable storage medium with an execution program stored thereon. When the execution program is executed, a robot dynamic grasping method based on multi-modal data fusion as described above is implemented.
[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0075] The present invention provides a robot dynamic grasping method and system based on multi-modal data fusion, including: collecting visual data and tactile data of a dynamic object to be grasped by a robot, and performing multi-modal fusion on the visual data and the tactile data to obtain multi-modal fusion data of the dynamic object to be grasped; according to the multi-modal fusion data, using a region proposal-based object detection algorithm to perform object recognition on the dynamic object to be grasped to obtain the position and pose information of the dynamic object to be grasped; based on the position and pose information, performing state prediction on the dynamic object to be grasped to obtain the state prediction information of the dynamic object to be grasped; according to the state prediction information of the dynamic object to be grasped, using a reinforcement learning algorithm to generate the grasping path of the robot; through multi-modal fusion of the collected visual data and tactile data, the robot can obtain comprehensive environmental information, increasing the robot's adaptability to environmental changes; by predicting the state of the object to be grasped, the robot can predict the movement trajectory of the object, which is beneficial to improving the grasping success rate of the robot; therefore, the method of the present invention can significantly improve the grasping effect and efficiency of the robot in a dynamic environment. Description of the Drawings
[0076] Figure 1 It is a flowchart showing a robot dynamic grasping method based on multi-modal data fusion provided by the present invention;
[0077] Figure 2 It is a flowchart showing the multi-modal fusion process of the robot in a robot dynamic grasping method based on multi-modal data fusion provided by the present invention;
[0078] Figure 3 It is a flowchart showing the object recognition process of the robot in a robot dynamic grasping method based on multi-modal data fusion provided by the present invention;
[0079] Figure 4 It is a flowchart showing the state prediction process of the robot in a robot dynamic grasping method based on multi-modal data fusion provided by the present invention;
[0080] Figure 5 It is a flowchart showing the path generation process of the robot in a robot dynamic grasping method based on multi-modal data fusion provided by the present invention;
[0081] Figure 6 Schematic diagram of the structural composition of a robot dynamic grasping system based on multi-modal data fusion provided by the present invention;
[0082] Figure 7 Schematic diagram of the structure of an electronic device provided by the present invention. Specific embodiments
[0083] The present invention proposes a robot dynamic grasping method, system, device and medium based on multi-modal data fusion. The following further elaborates on the specific embodiments of the present invention with reference to the accompanying drawings.
[0084] Embodiment 1:
[0085] The present invention provides a robot dynamic grasping method based on multi-modal data fusion. The process schematic diagram is as Figure 1 shown and includes:
[0086] Step 1: Collect visual data and tactile data of the dynamic object to be grasped by the robot, and perform multi-modal fusion on the visual data and tactile data to obtain multi-modal fusion data of the dynamic object to be grasped;
[0087] Step 2: According to the multi-modal fusion data, use a region proposal-based object detection algorithm to perform object recognition on the dynamic object to be grasped to obtain the position and pose information of the dynamic object to be grasped;
[0088] Step 3: Based on the position and pose information, perform state prediction on the dynamic object to be grasped to obtain state prediction information of the dynamic object to be grasped;
[0089] Step 4: According to the state prediction information of the dynamic object to be grasped, use a reinforcement learning algorithm to generate the grasping path of the robot.
[0090] Generally, in the research and application of robotics, object grasping is a crucial step. However, most of the existing grasping methods mainly focus on the recognition and operation of objects in a static environment and usually rely only on single-modal data (such as visual information). For example, traditional visual recognition systems mainly obtain image data through cameras and then extract object features for target recognition. This single visual modality often faces problems in dynamic scenes, including object occlusion, lighting changes, and state changes caused by object movement. These factors may lead to grasping failures or misidentifications, seriously affecting the working efficiency and success rate of robots in practical applications. To solve these problems, this application considers collecting multi-modal data, that is, simultaneously obtaining the tactile data and visual data of dynamic objects. Through this multi-modal data collection, it helps the robot to more comprehensively and accurately understand the characteristics and current state of the object, thereby improving its grasping performance in a dynamic environment. For example, the visual data collected by the above robot may include one or more of the following: color image data, depth image data, optical flow data, and ambient light data; in this example, the visual data can be captured by a high-resolution camera or a depth sensor (such as an RGB-D camera), and these devices can provide the image and depth information of the object, enabling the robot to identify features such as the shape, color, and size of the object;
[0091] For example, the tactile data collected by the above robot may include one or more of the following: robotic force feedback data, robotic torque data, grasping pressure data, object surface friction data, and object hardness data; in this example, the tactile data can be obtained through tactile sensors installed on the robot arm, and these sensors can sense information such as the force feedback, surface texture, and temperature of the object, helping the robot to evaluate the state and texture of the object when actually contacting the object;
[0092] In addition to tactile and visual modality data, it is also possible to consider collecting auditory data (through additional microphones or acoustic sensors, the robot can collect audio information in the environment, which is beneficial for identifying the collision sound or mechanical movement sound of the object, thereby providing auxiliary information for grasping decisions), inertial data (with the help of sensors such as accelerometers and gyroscopes, the robot can obtain its own position and attitude changes), and other modality data to enhance the robot's perception ability and grasping effect.
[0093] The multi-modal data collected through the above steps can improve the robot's perception and all-round information acquisition ability. To further realize the efficient utilization of these multi-modal data, the following multi-modal data fusion methods can be considered. Specifically:
[0094] In one implementation, such as Figure 2As shown, the process of performing multi-modal fusion on visual data and tactile data in the above step 1 to obtain multi-modal fusion data of the dynamic object to be grasped may include:
[0095] Extract features from the visual data through a visual feature selection method to obtain visual feature data;
[0096] Extract features from the tactile data through a one-dimensional convolutional neural network to obtain tactile feature data;
[0097] Adopt an adversarial training mechanism to perform cross-modal mapping on the visual feature data and the tactile feature data to obtain a fused feature vector after mapping;
[0098] Output the multi-modal fusion data of the dynamic object to be grasped according to the fused feature vector after mapping;
[0099] In this implementation, when performing visual data processing, the visual feature selection method adopted can filter out low-quality light deviation data for different types of visual inputs (such as RGB images, depth images, or optical flow data), focus on edge information and shape contours, significantly reduce redundant features, and select the most representative feature data. This not only improves the efficiency of data processing but also enhances the expression ability of key features for dynamic objects. The result can help the robot more accurately perceive the shape, trajectory, and environmental changes of objects. Moreover, this feature extraction can reduce noise interference and ultimately improve the accuracy of the target detection algorithm. Especially in the scenario of dynamically grasping objects, redundant visual information is eliminated through the visual feature selection method, and only key feature data is retained. This selection not only optimizes computing resources but also ensures the real-time performance of the grasping operation. When performing tactile data processing, the one-dimensional convolutional neural network adopted can extract key time-series features from tactile data with compact computing resources. When one-dimensional convolution processes feedback data (such as force feedback and pressure data), by identifying the laws of change of tactile sensors over time, the original tactile data is refined into sequence features of surface friction, hardness, and dynamic object contact. This feature extraction method can not only process highly dynamic unstructured data but also support information synchronization between touch and vision, establishing deep semantic associations for subsequent fusion operations. And in this implementation, by adopting a one-dimensional convolutional neural network, it is possible to focus on the feature extraction of time-series signals, capturing the subtle changes of tactile sensors at the moment of object contact, such as a sudden increase in pressure, changes in friction, or differences in object hardness. In the scenario of object grasping, the advantage of using one-dimensional convolution is that it can directly process time-series inputs, has high computing efficiency and is easy to parallelize. Through this design, tactile data is not only converted into high-dimensional semantic information but also can express the dynamic object contact state in the real-time scene, which is of great significance for solving the problem of temporal synchronization in the combination of multi-modal data.In this implementation method, when performing cross-modal mapping, an adversarial training mechanism is introduced. Through the mutual game between the generator and the discriminator, visual features and tactile features are mapped into a shared latent semantic space. Visual data and tactile data essentially belong to heterogeneous data and have natural inter-modal differences. Through the adversarial network, a shared feature space of the two modal data can be gradually found in the game framework of a generator and a discriminator, enabling visual information and tactile information to complement and align with each other semantically in a high-dimensional space. For example, the features of an object's contour in visual data can correspond to the hardness or friction features in tactile data, and the generator forms this correspondence during the mapping process. The adversarial design of introducing the discriminator ensures the accuracy of feature mapping and avoids information loss or fuzzification problems that may occur during the multi-modal data fusion process. This mapping method significantly improves the expressive ability of the fused feature vector, enabling the fused data to have both the global perception of vision and the fine discrimination of touch. This can not only solve the problem of different modal data characteristics but also maximize the retention of the implicit correlation between modalities. For example, by extracting the mapping weights between vision and touch through the generator model, both fine-grained target identification and cross-modal coordination can be ensured. This design can avoid information loss and improve the expressiveness of multi-modal data fusion, enabling the robot to efficiently extract the global features of dynamic objects in a complex environment. When performing multi-modal fusion data output based on the fused feature vector after mapping, a semantic alignment mechanism can be established to ensure that the final integration result of the data can accurately reflect the complex dynamics of the target object. For example, visually, it can more accurately reflect the motion trajectory or deformation of the target object; tactilely, it can reflect subtle changes such as real-time contact force and dynamic friction, thereby improving the overall grasping success rate.
[0100] For example, the expression of the above adversarial training mechanism can be as follows:
[0101] F fused = G(F visual ; θ g ) + βH(F haptic ; θ h );
[0102] In the formula,
[0103]
[0104] where, F fused represents the fused feature vector after mapping; G represents the generator model; F visual represents the visual feature data; θ g represents the parameters of the generator model; β represents the dynamic weight factor; H represents the non-linear transformation model; F haptic represents the tactile feature data; θ hRepresents the parameters of the non - linear transformation model; Similarity represents the similarity score. In this example, by fusing the generator model and the non - linear transformation model, the feature vectors are adjusted and optimized during the fusion process of multi - modal data, so as to achieve more accurate and robust dynamic environment target perception. And by introducing a dynamic weight factor β (corrected by the similarity score between visual features and tactile features), when the visual and tactile features have a high similarity, the weight factor β tends to a higher value, indicating that the fused features after mapping combine the common characteristics of the two to a greater extent and enhance the compatibility of information; while when the similarity is low, β reduces the fusion dependence between the two, and the important information unique to either modality will not be ignored. This dynamic adjustment enables the generator to more accurately map two different - modality data into a shared semantic space, thus retaining the unique features in each modality that are crucial for understanding the overall dynamic behavior.
[0105] Through the above - mentioned steps of fusing multi - modal data, the visual features and tactile features of the dynamic object to be grasped can be extracted. It is possible to consider using an adversarial training mechanism to achieve an effective mapping between visual feature data and tactile feature data, forming a fused feature vector that can reflect the object state. Specifically:
[0106] In one implementation, as Figure 3 shown, the process of obtaining the position and pose information of the dynamic object to be grasped by using the region - proposal - based object detection algorithm according to the multi - modal fusion data in step 2 above can include:
[0107] According to the multi - modal fusion data, generate candidate region proposals through the selective search algorithm;
[0108] According to the candidate region proposals, use the object detection algorithm to perform object recognition on the dynamic object to be grasped, and obtain the classification result and bounding box coordinate information of the dynamic object to be grasped;
[0109] According to the classification result and bounding box coordinate information of the dynamic object to be grasped, perform pose estimation on the dynamic object to be grasped, and obtain the position and pose information of the dynamic object to be grasped;
[0110] In this implementation, the selective search algorithm adopted generates candidate region proposals on the multi-modal fusion data, which will reduce the computational burden in complex scenarios and preferentially select regions that may contain the target object, thereby optimizing the subsequent object detection process. This algorithm adopts a method based on image segmentation, which can significantly reduce the number of candidate regions while maintaining a high proposal quality, enabling the object detection algorithm to focus on the most relevant parts. This focusing strategy is beneficial to improving the computational efficiency, reducing the risk of misidentification, and accelerating the overall processing speed. When performing object recognition, the object detection algorithm (such as Fast RCNN) adopted uses these candidate regions to perform accurate object recognition and classification, generating classification results and bounding box coordinate information for each potential object. The core of this process lies in being able to accurately identify the type of the object, so as to quickly provide the necessary data for the grasping strategy. Especially in a dynamic environment, objects may move in different postures and speeds. Therefore, the object detection algorithm needs to have high precision and real-time response capabilities. Once the target is correctly identified, it will further provide a reliable basis for the subsequent grasping operation, laying a foundation for the robot to achieve more automated grasping tasks. When the robot grasps a dynamic object, since it is necessary to ensure that the robot can not only identify the object but also grasp it in the correct way, it is necessary to perform pose estimation on the dynamic object (three-dimensional key point detection, global minimization method, or deep learning-based pose regression, etc. can be considered). Because the position and shape of the object often change during actual operation, accurate pose information can help the robot adjust its grasping strategy, thereby minimizing the grasping error rate and improving the grasping success rate. Therefore, this implementation does not stay at the theoretical or single-modal data analysis level, but breaks through the limitations of traditional technologies in dealing with dynamic environments, makes full use of various types of information to enhance the discriminative power and execution ability, demonstrating the improvement of the new multi-modal data fusion and real-time response capabilities, ensuring that the robot can achieve efficient and accurate grasping in a changing environment.
[0111] Through the above steps, object recognition can be performed based on multi-modal fusion data to obtain the position and pose information of the dynamic object to be grasped. For further research on the dynamic characteristics and behavior prediction of the object, it can be considered to combine long short-term memory networks and conditional random fields to predict the motion trajectory of the dynamic object, enabling the robot to accurately identify the target to be grasped in a complex environment. Specifically:
[0112] In one implementation, as Figure 4 shown, the process of predicting the state of the dynamic object to be grasped based on the position and pose information in step 3 above to obtain the state prediction information of the dynamic object to be grasped may include:
[0113] Use an autoencoder to extract the state features of the position and attitude information, and obtain the feature representation of the dynamic object to be grasped after fusion;
[0114] According to the feature representation of the dynamic object to be grasped after fusion, use a long short-term memory network combined with a conditional random field to predict the trajectory of the dynamic object to be grasped, and obtain the motion trajectory prediction information of the dynamic object to be grasped;
[0115] According to the motion trajectory prediction information of the dynamic object to be grasped, use a deep reinforcement learning algorithm to predict the state of the dynamic object to be grasped, and obtain the state prediction information of the dynamic object to be grasped;
[0116] In this implementation, by predicting the state of the dynamic object to be grasped, accurate dynamic behavior understanding and control are achieved. The key lies in the effective combination of an autoencoder, a long short-term memory network (LSTM), a conditional random field, and deep reinforcement learning. In this implementation, the autoencoder is used to extract state features from the position and pose information. It converts high-dimensional input data into low-dimensional feature representations through the encoder, extracts the most critical information from them, and eliminates redundancy. This processing method not only effectively reduces the complexity of the input data but also enhances the concentration and effectiveness of the subsequent model in feature representation, making the state features more representative and practical, thus laying a foundation for subsequent data analysis and processing. After extracting the fused features, the LSTM is used in combination with the conditional random field for trajectory prediction. The LSTM network focuses on the processing of time-series data and can capture long-term and short-term dependencies in dynamic changes, enabling the model to remember previous states and foresee future trajectories. At the same time, combining the conditional random field provides reinforcement of context information, making the trajectory prediction more coherent and accurate. The application of this method improves the prediction ability for complex dynamic behaviors, not only better reflecting the laws of object movement but also providing a practical trajectory reference for dynamic grasping. After trajectory prediction, deep reinforcement learning is introduced for state prediction. By interacting with the environment and optimizing the decision-making process through a feedback mechanism, it can evaluate the impact and results of different strategies on the object to be grasped. This makes state prediction not just a static result but a dynamic adaptive process. Specifically, the deep reinforcement learning model can continuously learn and adjust strategies, and based on the current state and predicted trajectory, it can update the grasping strategy in real time, thus more effectively coping with dynamic changes. This adaptive ability significantly improves the operation ability of the robot in a complex environment, making state prediction not only reflect existing laws but also be predictive and flexible, significantly enhancing the success rate and efficiency of grasping. In a real environment, the state of an object changes in many ways, and a single technical means cannot meet the actual needs. However, this implementation provides an effective solution for the dynamic grasping task through a multi-level and multi-dimensional processing idea, highlighting the necessity of the collaborative effect of different algorithms and achieving accurate capture and classification of the behaviors of dynamic objects.
[0117] Specifically, the process of using the long short-term memory network in combination with the conditional random field to predict the trajectory of the dynamic object to be grasped based on the fused feature representation of the dynamic object to be grasped and obtaining the motion trajectory prediction information of the dynamic object to be grasped in the above implementation can include:
[0118] Perform context feature transformation on the fused feature representation of the dynamic object to be grasped to obtain feature sequence data;
[0119] Use a long short-term memory network to perform implicit prediction on the feature sequence data and output the hidden state sequence of the feature sequence data;
[0120] Perform post-processing on the hidden state sequence through a conditional random field to obtain the motion trajectory prediction information of the dynamic object to be grasped;
[0121] In this implementation method, during context feature transformation, by transforming the fused feature representation, the context information of the time series data is extracted. This process can effectively capture the correlation between features, convert fixed features into temporal feature sequences, enabling the subsequent model to better understand the change trend of dynamic objects in the time dimension. This transformation not only enhances the timeliness and importance of the feature data but also provides more valuable input information for the deep learning model, improving the overall performance ability of the model. After obtaining the feature sequence, use a long short-term memory network to perform implicit prediction on the data. Due to the design of its forget gate, input gate, and output gate, LSTM can effectively learn and remember long-term dependencies and is more capable of processing data with longer time series compared to traditional recurrent neural networks. Therefore, when dealing with the motion of the object to be grasped, LSTM can effectively predict the future motion trend based on the previous state and output the hidden state sequence, which provides a reliable basis for trajectory prediction. The progress of this stage provides in-depth insights into the behavior pattern learning of dynamic objects, enabling the model to better adapt to changing motion situations; in this implementation method, through the introduction of a conditional random field (CRF), post-processing is performed on the basis of the LSTM hidden state sequence to obtain more accurate trajectory prediction. The strength of CRF lies in its ability to consider the context relationship between feature sequences. By modeling the front and back states, it can effectively correct the possible prediction deviation in LSTM and enhance the coherence and accuracy of trajectory prediction. Through the correlation analysis of feature transfer and output sequence, the predicted trajectory is not only affected by a single output but can better reflect the integrity and consistency of the object's motion state. This post-processing method can significantly improve the stability and resolvability of motion trajectory prediction and ensure the application reliability of the model in various scenarios; therefore, in this implementation method, through the combination of LSTM and CRF, the entire prediction process becomes more fluent in a dynamic scenario, can adjust and optimize the prediction results in real time, and demonstrates extremely high adaptability and practicality. Especially in a dynamic environment, accurate prediction of the motion trajectory of the dynamic object to be grasped is achieved.
[0122] Through the above steps, the motion trajectory prediction of the dynamic object is achieved. To further optimize the grasping strategy, it is possible to consider combining the environmental information of the dynamic object for scene perception and generating a suitable grasping path for the robot through a reinforcement learning algorithm. Specifically:
[0123] In one implementation, as Figure 5 shown, the process of generating the grasping path of the robot by using the reinforcement learning algorithm according to the state prediction information of the dynamic object to be grasped in step 4 above may include:
[0124] Performing scene perception on the dynamic object to be grasped according to the state prediction information of the dynamic object to be grasped and the pre-collected environmental information of the dynamic object to be grasped, so as to obtain the scene state description information of the dynamic object to be grasped;
[0125] Constructing a dynamic grasping policy network by using the deep deterministic policy gradient algorithm according to the scene state description information to obtain a dynamic grasping policy;
[0126] Generating a path by using a path generation algorithm based on imitation learning according to the dynamic grasping policy to obtain the grasping path of the dynamic object to be grasped;
[0127] In this implementation, when performing scene perception, the robot compares and analyzes the state prediction information of the dynamic object to be grasped and the pre-collected environmental information. This process can use computer vision techniques such as object detection and image segmentation to identify and extract the feature information of objects in the environment, generating a detailed description of the scene state. This perception ability enables the robot to obtain rich environmental information, ensuring that the grasping decision can be based on the actual situation of the current environment, thereby improving the reliability and success rate of grasping. Through a comprehensive understanding of the environment, it can better adapt to dynamically changing conditions and reduce the grasping risks brought about by environmental changes. After obtaining the scene state description information, the model uses the Deep Deterministic Policy Gradient (DDPG) algorithm to construct a dynamic grasping policy network. The advantage of the DDPG algorithm is that it can handle continuous action spaces and optimize the policy through deep learning. Specifically, the algorithm explores to obtain experience in the environment and uses the policy network and value network to select and evaluate actions, effectively improving the efficiency of policy optimization. This process enables the grasping policy to flexibly respond to different grasping environments and object states, elevating the quality of grasping decisions to a new level and ensuring that the robot can better execute complex grasping tasks in a changing environment. And by introducing imitation learning in the path generation process, the core of imitation learning is to improve the learning efficiency by imitating the decision-making process of experts. Specifically, the path generation algorithm learns the expert-level path selection in the training set and uses it as a reference for generating the grasping path. Such a design can not only shorten the process of the robot learning the optimal path but also greatly reduce the unnecessary time and resource waste caused by incorrect exploration. This strategy enables the robot to not only learn how to grasp under ideal conditions but also master effective grasping strategies in complex environments by combining with the experience of experts. Although some of the technical means in this implementation have been widely used in the industry, the comprehensive application of these technologies in the specific scenario of dynamic object grasping, through the combination of scene perception and deep learning, combines the timeliness and accuracy of scene perception with the flexibility of dynamic grasping strategies, forming an overall dynamic grasping solution. The application of DDPG ensures the continuity and real-time response of policy optimization, thus enabling it to cope with the ever-changing dynamic environment, while imitation learning provides an efficient reference for path generation, reducing the learning time and ensuring an efficient grasping process.
[0128] In summary, in the process of a robot grasping an object, due to the movement, deformation, or occlusion of objects in the environment, existing grasping methods can only learn grasping in a static environment and have insufficient object recognition and grasping planning capabilities in the face of a dynamic environment. To address this problem, the present invention proposes a robot dynamic grasping method based on multi-modal data fusion. By performing multi-modal fusion of visual data and tactile data, the robot can obtain richer and more comprehensive information, not limited to static features. This fusion enables the robot to more accurately identify and locate the dynamic object to be grasped, thus solving the problem that static grasping methods cannot be effectively applied under dynamic conditions. After obtaining the multi-modal fusion data, the object detection algorithm based on region proposal further improves the accuracy of object recognition, enabling the robot to clearly extract the position and pose information of the object even in a complex and dynamic environment. Due to the complexity of the dynamic environment, this step is crucial for ensuring the success of grasping. Through state prediction of the obtained pose information, the present application can evaluate the motion trajectory and change trend of the object in real time, providing dynamic support for the grasping strategy. In addition, by introducing a reinforcement learning algorithm, the robot can adaptively generate the optimal grasping path in a dynamic environment. This mechanism not only makes up for the deficiencies of traditional static methods but also provides flexibility and real-time performance for the robot's decision-making, ensuring that it can quickly respond when facing changing objects.
[0129] Embodiment 2:
[0130] Based on the same inventive concept, the present invention also provides a robot dynamic grasping system based on multi-modal data fusion. The schematic structural composition diagram is as Figure 6 shown, including:
[0131] A multi-modal fusion module for collecting visual data and tactile data of the dynamic object to be grasped by the robot and performing multi-modal fusion on the visual data and tactile data to obtain multi-modal fusion data of the dynamic object to be grasped;
[0132] An object recognition module for performing object recognition on the dynamic object to be grasped according to the multi-modal fusion data by using an object detection algorithm based on region proposal to obtain the position and pose information of the dynamic object to be grasped;
[0133] A state prediction module for performing state prediction on the dynamic object to be grasped based on the position and pose information to obtain state prediction information of the dynamic object to be grasped;
[0134] A path generation module for generating a grasping path of the robot according to the state prediction information of the dynamic object to be grasped by using a reinforcement learning algorithm.
[0135] In one implementation, the above multi-modal fusion module may include:
[0136] A visual feature extraction sub-module, which is used to extract features from visual data through a visual feature selection method to obtain visual feature data;
[0137] A tactile feature extraction sub-module, which is used to extract features from tactile data through a one-dimensional convolutional neural network to obtain tactile feature data;
[0138] A cross-modal mapping sub-module, which is used to perform cross-modal mapping on the visual feature data and the tactile feature data by adopting an adversarial training mechanism to obtain a fused feature vector after mapping;
[0139] A data fusion sub-module, which is used to output multi-modal fusion data of the dynamic object to be grasped according to the fused feature vector after mapping;
[0140] Among them, the visual data includes one or more of the following: color image data, depth image data, optical flow data, and ambient light data;
[0141] The tactile data includes one or more of the following: robotic force feedback data, robotic torque data, grasping pressure data, object surface friction data, and object hardness data.
[0142] Exemplarily, the expression of the above adversarial training mechanism can be as follows:
[0143] F fused = G(F visual ; θ g ) + βH(F haptic ; θ h );
[0144] In the formula,
[0145]
[0146] Among them, F fused represents the fused feature vector after mapping; G represents the generator model; F visual represents the visual feature data; θ g represents the parameters of the generator model; β represents the dynamic weight factor; H represents the non-linear transformation model; F haptic represents the tactile feature data; θ h represents the parameters of the non-linear transformation model; Similarity represents the similarity score.
[0147] In one implementation manner, the above target recognition module may include:
[0148] A candidate region generation sub-module, which is used to generate candidate region proposals according to the multi-modal fusion data through a selective search algorithm;
[0149] A target detection sub-module, which is used to perform target recognition on the dynamic object to be grasped according to the candidate region proposals by using a target detection algorithm, and obtain the classification result and bounding box coordinate information of the dynamic object to be grasped;
[0150] An attitude estimation sub-module, which is used to perform attitude estimation on the dynamic object to be grasped according to the classification result and bounding box coordinate information of the dynamic object to be grasped, and obtain the position and attitude information of the dynamic object to be grasped.
[0151] In one implementation, the above-mentioned state prediction module may include:
[0152] A feature representation sub-module, which is used to extract state features from the position and attitude information by using an autoencoder, and obtain the fused feature representation of the dynamic object to be grasped;
[0153] A trajectory prediction sub-module, which is used to perform trajectory prediction on the dynamic object to be grasped according to the fused feature representation of the dynamic object to be grasped by using a long short-term memory network combined with a conditional random field, and obtain the motion trajectory prediction information of the dynamic object to be grasped;
[0154] A state estimation sub-module, which is used to perform state prediction on the dynamic object to be grasped according to the motion trajectory prediction information of the dynamic object to be grasped by using a deep reinforcement learning algorithm, and obtain the state prediction information of the dynamic object to be grasped.
[0155] In this implementation, the above-mentioned trajectory prediction sub-module may include:
[0156] A feature conversion unit, which is used to perform context feature conversion on the fused feature representation of the dynamic object to be grasped, and obtain feature sequence data;
[0157] An implicit prediction unit, which is used to perform implicit prediction on the feature sequence data by using a long short-term memory network, and output the hidden state sequence of the feature sequence data;
[0158] A post-processing unit, which is used to perform post-processing on the hidden state sequence by using a conditional random field, and obtain the motion trajectory prediction information of the dynamic object to be grasped.
[0159] In one implementation, the above-mentioned path generation module may include:
[0160] Perform scene perception on the dynamic object to be grasped according to the state prediction information of the dynamic object to be grasped and the pre-collected environmental information of the dynamic object to be grasped, and obtain the scene state description information of the dynamic object to be grasped;
[0161] Construct a dynamic grasping policy network by using a deep deterministic policy gradient algorithm according to the scene state description information, and obtain a dynamic grasping policy;
[0162] According to the dynamic grasping strategy, a path generation algorithm based on imitation learning is used to generate a path, and a grasping path for the dynamic object to be grasped is obtained.
[0163] Embodiment 3:
[0164] As Figure 7 shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected through a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and this data can be called and / or modified when the instructions are executed.
[0165] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a robot dynamic grasping method based on multi-modal data fusion in the above embodiments.
[0166] Embodiment 4:
[0167] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in an electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides storage space, and this storage space stores the operating system of the terminal. Moreover, in this storage space, there is also stored one or more instructions suitable for being loaded and executed by a processor, and these instructions can be one or more executable programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a method for dynamic grasping of a robot based on multi-modal data fusion in the above embodiments can be implemented.
[0168] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0169] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0170] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or one block or a plurality of blocks in the flow Figure 1 one process or a plurality of processes and / or Figure 1 blocks. The steps for realizing the functions specified in one block or a plurality of blocks
[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit the scope of its protection. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: after reading the present invention, those skilled in the art can still make various changes, modifications or equivalent replacements to the specific implementation manners of the application, but these changes, modifications or equivalent replacements are all within the scope of the protection of the claims pending for approval of the application.
Claims
1. A robot dynamic grasping method based on multi-modal data fusion, characterized in that, Including: Collecting visual data and tactile data of a dynamic object to be grasped by a robot, and performing multimodal fusion on the visual data and the tactile data to obtain multimodal fusion data of the dynamic object to be grasped; According to the multimodal fusion data, using a region proposal-based object detection algorithm to perform object recognition on the dynamic object to be grasped, and obtaining the position and pose information of the dynamic object to be grasped; Based on the position and pose information, performing state prediction on the dynamic object to be grasped, and obtaining the state prediction information of the dynamic object to be grasped; According to the state prediction information of the dynamic object to be grasped, using a reinforcement learning algorithm to generate a grasping path of the robot.
2. The method according to claim 1, wherein The performing multimodal fusion on the visual data and the tactile data to obtain multimodal fusion data of the dynamic object to be grasped includes: Performing feature extraction on the visual data through a visual feature selection method to obtain visual feature data; Performing feature extraction on the tactile data through a one-dimensional convolutional neural network to obtain tactile feature data; Adopting an adversarial training mechanism to perform cross-modal mapping on the visual feature data and the tactile feature data to obtain a fused feature vector after mapping; According to the fused feature vector after mapping, outputting the multimodal fusion data of the dynamic object to be grasped; Wherein, the visual data includes one or more of the following: color image data, depth image data, optical flow data, and ambient light data; The tactile data includes one or more of the following: robot force feedback data, robot torque data, grasping pressure data, object surface friction data, and object hardness data.
3. The method according to claim 2, wherein The expression of the adversarial training mechanism is as follows: F fused = G(F visual ; θ g ) + βH(F haptic ; θ h ); In the formula, Among them, F fused represents the fused feature vector after mapping; G represents the generator model; F visual represents the visual feature data; θ g represents the parameters of the generator model; β represents the dynamic weight factor; H represents the non-linear transformation model; F haptic represents the tactile feature data; θ h represents the parameters of the non-linear transformation model; Similarity represents the similarity score.
4. The method according to claim 1, characterized in that The performing object recognition on the dynamic object to be grasped according to the multimodal fusion data by using a region proposal-based object detection algorithm to obtain the position and pose information of the dynamic object to be grasped includes: According to the multimodal fusion data, generating candidate region proposals through a selective search algorithm; According to the candidate region proposals, using an object detection algorithm to perform object recognition on the dynamic object to be grasped, and obtaining the classification result and bounding box coordinate information of the dynamic object to be grasped; According to the classification result and bounding box coordinate information of the dynamic object to be grasped, performing pose estimation on the dynamic object to be grasped, and obtaining the position and pose information of the dynamic object to be grasped.
5. The method according to claim 1, characterized in that, The performing state prediction on the dynamic object to be grasped based on the position and pose information to obtain the state prediction information of the dynamic object to be grasped includes: Using an autoencoder to perform state feature extraction on the position and pose information to obtain a fused feature representation of the dynamic object to be grasped; According to the fused feature representation of the dynamic object to be grasped, using a long short-term memory network combined with a conditional random field to perform trajectory prediction on the dynamic object to be grasped, and obtaining the motion trajectory prediction information of the dynamic object to be grasped; According to the motion trajectory prediction information of the dynamic object to be grasped, using a deep reinforcement learning algorithm to perform state prediction on the dynamic object to be grasped, and obtaining the state prediction information of the dynamic object to be grasped.
6. The method according to claim 5, wherein Based on the feature representation after fusing the dynamic object to be grasped, using a long short-term memory network combined with a conditional random field to predict the trajectory of the dynamic object to be grasped, and obtaining the motion trajectory prediction information of the dynamic object to be grasped, including: Performing context feature transformation on the feature representation after fusing the dynamic object to be grasped to obtain feature sequence data; Using a long short-term memory network to perform implicit prediction on the feature sequence data and outputting the hidden state sequence of the feature sequence data; Performing post-processing on the hidden state sequence through a conditional random field to obtain the motion trajectory prediction information of the dynamic object to be grasped.
7. The method according to claim 1, wherein Based on the state prediction information of the dynamic object to be grasped, using a reinforcement learning algorithm to generate the grasping path of the robot, including: Performing scene perception on the dynamic object to be grasped according to the state prediction information of the dynamic object to be grasped and the pre-collected environmental information of the dynamic object to be grasped, and obtaining the scene state description information of the dynamic object to be grasped; According to the scene state description information, using a deep deterministic policy gradient algorithm to construct a dynamic grasping policy network and obtaining a dynamic grasping policy; According to the dynamic grasping policy, using a path generation algorithm based on imitation learning to generate a path and obtaining the grasping path of the dynamic object to be grasped.
8. A robot dynamic grasping system based on multi-modal data fusion, characterized in that, Including: A multi-modal fusion module, configured to collect visual data and tactile data of a dynamic object to be grasped by a robot, and perform multi-modal fusion on the visual data and the tactile data to obtain multi-modal fusion data of the dynamic object to be grasped; A target recognition module, configured to perform target recognition on the dynamic object to be grasped according to the multi-modal fusion data by using a region proposal-based target detection algorithm, and obtain the position and attitude information of the dynamic object to be grasped; A state prediction module, configured to perform state prediction on the dynamic object to be grasped based on the position and attitude information, and obtain the state prediction information of the dynamic object to be grasped; A path generation module, configured to generate the grasping path of the robot by using a reinforcement learning algorithm according to the state prediction information of the dynamic object to be grasped.
9. An electronic device, characterized in that, Including: At least one processor and a memory; The memory and the processor are connected by a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the method for dynamically grasping a robot based on multi-modal data fusion according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, There is an execution program stored thereon, and when the execution program is executed, the method for dynamically grasping a robot based on multi-modal data fusion according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Mechanical arm active grabbing device and method based on multi-model fusion
CN109129474A
Cross-modal image generation method and device based on audio-tactile signal fusion
CN113627482A
Robot based on visual touch fusion and grabbing system and method thereof
CN114700947A
Visual capture detection method and system based on self-supervised representation learning
CN114820796A
Robot grabbing method and device based on multi-source information fusion
CN115256377A
Cited By
Robot motion state prediction method and device, electronic equipment and storage medium
CN121132612A
Robot motion state prediction method and device, electronic equipment, and storage medium
CN121132612B
Mechanical arm grabbing method and system based on multi-modal information fusion
CN121403370A
Chemical emergency manipulator control system and method based on intelligent perception
CN121848412A