Grabbing detection method and system based on state space model

By combining feature encoding of the state-space model with the latent distribution module, the problem of low accuracy in grasping detection in unstructured scenarios is solved, achieving efficient grasping detection and improving feature utilization and computational efficiency.

CN120862702AActive Publication Date: 2025-10-31HUNAN UNIV
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202511373802.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-10-31
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing grasping detection methods cannot accurately identify the pose and position of target objects in unstructured scenarios, resulting in low grasping detection accuracy. Furthermore, existing networks do not fully extract effective grasping features, showing strong local feature attention but insufficient global feature attention.

Method used

A grasping detection method based on a state space model is adopted. By combining feature encoding, feature fusion, latent distribution module and grasping parameter estimation module, forward and backward state space models are established. The latent distribution module is used to represent the continuous distribution of grasping posture, thereby improving feature utilization and the accuracy of global correlation information.

Benefits of technology

It improves the accuracy of grasping and detection, reduces computational costs, and accelerates the convergence speed of neural networks by utilizing the parallel computing capabilities of GPUs, thereby enhancing the accuracy and efficiency of grasping and detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120862702A_ABST
    Figure CN120862702A_ABST
Patent Text Reader

Abstract

The invention discloses a grabbing detection method and system based on a state space model, and the method comprises the steps: obtaining scene point clouds containing a target object, and carrying out the marking of the scene point clouds, and obtaining a label data set; building a grabbing detection model; training the capture detection model, calculating the total loss according to the prediction result of the capture detection model and the label data set, adjusting the weight of the capture detection model, and minimizing the total loss until a set iteration stop condition is reached; and evaluating the trained capture detection model by using the test set, storing model parameters with the best performance as a final capture detection model, and inputting scene point cloud containing a target object in reality into the final capture detection model to obtain a detection result. According to the method, global feature expression is established by adopting the state space model, unstructured scene influences such as target object backgrounds, postures and position randomness are reduced, the understanding of the network on grabbing tasks is improved in combination with the probability model, and higher-precision grabbing detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of grasping detection technology, and in particular to a grasping detection method and system based on a state-space model. Background Technology

[0002] Grasping and detection, a fundamental task in robot operation, is crucial for accurately locating the graspable position of a target object and generating a stable grasping posture, which directly impacts the robot's ability to complete subsequent assembly, sorting, and other tasks. A typical grasping and detection system includes a perception module, an information processing module, a communication module, and an execution module. The perception module collects image information from the scene, the information processing module processes the scene information and outputs the grasping posture, the execution module performs the grasping operation, and the communication module handles information transmission between the various modules. With the development of intelligent manufacturing, robotic automation systems have received significant attention and research, and are gradually replacing manual labor as the primary workforce in industrial scenarios such as autonomous assembly and automated parts sorting in production and manufacturing.

[0003] With the deepening of research on grasping detection, autonomous transfer and loading / unloading have been widely applied in structured scenarios such as neatly arranged objects or those fixed by clamps. However, when the target object's posture and position are random, its types are mixed, and the background is cluttered, existing grasping detection methods cannot achieve accurate identification, leading to the need to add redundant workstations to adjust the target object's posture and position. Therefore, how to improve the success rate of grasping detection in unstructured scenarios remains an important research problem. Point cloud data can provide structural information for analysis compared to two-dimensional images, but as sparse data, point cloud features are more dispersed and require neural networks with stronger understanding capabilities. Existing networks such as PointNet++ extract features from the scene by grouping and aggregating point sets, which has a strong ability to focus on local features. However, the layer-by-layer sampling and aggregation method does not pay enough attention to global features, limiting the effective feature extraction and expression for grasping, thus resulting in low grasping detection accuracy. Therefore, how to optimize the feature extraction stage of grasping detection and enhance the network's focus on features effective for grasping detection remains a research area that needs attention. Summary of the Invention

[0004] This invention provides a grasping and detection method and system based on a state-space model to solve the technical problems mentioned in the background.

[0005] To achieve the above objectives, the technical solution of the present invention is implemented as follows: This invention provides a grasping and detection method based on a state-space model, comprising the following steps: S1. Obtain several scene point clouds containing the target object, label each scene point cloud separately to obtain a set of labeled data with the target object's grasping posture, and divide the set of labeled data into a training set and a test set according to a set ratio. S2. Build a grasping detection model based on a state space model. The grasping detection model includes a feature encoding module, a feature fusion module, a latent distribution module, and a grasping parameter estimation module connected in sequence. S3. Train the crawling detection model using the training set, set the total loss function, calculate the total loss based on the prediction results of the crawling detection model and the label data set, adjust the weights of the crawling detection model based on the total loss, minimize the total loss, until the set iteration stopping condition is reached. S4. Evaluate the trained grasping detection model using the test set, and save the model parameters that perform best on the test set as the final grasping detection model. Input the real-world scene point cloud containing the target object into the final grasping detection model to obtain the detection results.

[0006] In another aspect, the present invention provides a grasping and detection system based on a state-space model, including a device end, which performs grasping and detection on a target object according to a grasping and detection method.

[0007] The beneficial effects of this invention are: 1. High accuracy in capture and detection This invention establishes forward and backward state space models using a state space model, providing sufficient feature data for subsequent grasping detection. The excellent long-range modeling capability of the state space model ensures the accuracy of tasks relying on globally related information, such as graspable position identification and collision detection, during the grasping detection process, thus greatly improving the accuracy of grasping detection. The latent distribution module replaces the one-to-one mapping relationship of the deterministic model, representing the continuous distribution of grasping postures through latent distribution, effectively improving feature utilization and increasing grasping detection accuracy.

[0008] 2. Low computational cost for capture and detection. This invention utilizes a state-space model to construct a grasping and detection model. The state-space model possesses long-range modeling capabilities comparable to the Transformer, while avoiding redundant variables in the computational process, thus saving computational costs. Furthermore, the linear expression of the state-space model effectively leverages the parallel computing power of GPUs (graphics cards), accelerating the convergence speed of the neural network. Attached Figure Description

[0009] Figure 1 This is a structural block diagram of the capture and detection model in this invention. Detailed Implementation

[0010] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many other different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.

[0011] To address the following issues in intelligent manufacturing tasks such as autonomous assembly and automatic parts sorting: a) Existing grasping and detection networks do not sufficiently extract effective grasping features, and incomplete global feature extraction leads to low grasping and detection accuracy; b) Deterministic models map features as a single output, while grasping posture parameters are spatially continuous, resulting in incomplete feature utilization and affecting grasping and detection results.

[0012] To address problem a, this invention proposes a grasping detection model based on a state-space model. By introducing linear state-space modeling, it sequentially models and dynamically fuses the local grasping feature sequences perceived at multiple scales (i.e., the features used by the grasping parameter estimation module), enhancing the spatiotemporal correlation between local regions and improving the responsiveness of grasping feature representation to key regions. To address problem b, this invention uses latent variable distribution to represent the phenomenon of continuous distribution in the effective grasping posture space. By fitting the conditional distribution relationship between scene features and grasping posture parameters through a latent distribution module (a probabilistic model), it realizes the transformation from deterministic mapping to probabilistic modeling, thereby discovering more potential high-quality grasping postures and improving the utilization rate of features in the grasping detection model.

[0013] The specific solution of the present invention is described below: Reference Figure 1 This application provides a grasping and detection method based on a state-space model, comprising the following steps: S1. Obtain several scene point clouds containing the target object, with a scene point cloud length of [length missing]. With a dimension of 3, each scene point cloud is labeled to obtain a set of labeled data with the target object's grasping posture, and the set of labeled data is divided into a training set and a test set according to a set ratio; S2. Build a grasping detection model based on a state space model. The grasping detection model includes a feature encoding module, a feature fusion module, a latent distribution module, and a grasping parameter estimation module connected in sequence. S3. Train the crawling detection model using the training set, set the total loss function, calculate the total loss based on the prediction results of the crawling detection model and the label data set, adjust the weights of the crawling detection model based on the total loss, minimize the total loss, until the set iteration stopping condition is reached. S4. Evaluate the trained grasping detection model using the test set, and save the model parameters that perform best on the test set as the final grasping detection model. Input the real-world scene point cloud containing the target object into the final grasping detection model to obtain the detection results.

[0014] In some embodiments, S1 specifically includes the following steps: S11. Obtain several scene point clouds containing the target object, and label the target in each scene point cloud. The labeling content includes the grasping posture quality score, gripper contact point pairs, gripper rotation angle, and grasping posture tolerance. Obtain a labeled data set with the grasping posture of the target object, expressed by the following formula: ; ; ; in, This represents a set of labeled data with the target object's grasping posture. , , , These represent the grasping mass fraction, gripper contact point pairs, gripper rotation angle, and grasping posture tolerance, respectively. Grasping posture tolerance is the probability that the gripper, in its current posture, forms an effective closure of the contact surface with the target object. Represents the set of real numbers; Represents the number of points in the scene's point cloud; S12. According to a set ratio (e.g., 9:1), collect the tagged data set with the target object's grasping posture. It is divided into training set and test set.

[0015] In some embodiments, the feature encoding module comprises a PointNet++ deep learning model based on 3D point cloud data processing, including point set grouping blocks and feature clustering blocks connected in sequence, wherein the point set grouping blocks are used to encode features according to the input point cloud. The distance between each point in the input point cloud and the anchor points obtained by random sampling affects the input point cloud. The points are grouped, and the feature clustering block is used to merge the features of each group of points, and the feature of the current point set is used as the feature value of the current anchor point.

[0016] The feature encoding module takes a scene point cloud with a sampling length of N as input, performs point cloud grouping and feature aggregation, and outputs a sample length of N. The feature length is Feature data, in this embodiment , , .

[0017] In some embodiments, the feature fusion module comprises a normalization layer, a forward state space model, a backward state space model, and a projection layer; The forward state space model consists of sequentially connected one-dimensional convolutional layers. and state space model layer One-dimensional convolutional layer Input dimension is The output dimension is State-space model layer Input dimension is The output dimension is ; The backward state space model consists of sequentially connected one-dimensional convolutional layers. and state space model layer The projection layer contains a fully connected layer. (One-dimensional convolutional layer) Input dimension is The output dimension is State-space model layer Input dimension is The output dimension is The input dimension of the projection layer is The output dimension is ; The input to the feature fusion module is feature data. The scene features are modeled, and effective feature information for crawling is extracted. The output sample length is [missing information]. The feature length is High-dimensional feature data (i.e., feature information) ); In some embodiments, the latent distribution module includes a prior distribution calculation block, a posterior distribution calculation block, and a latent variable sampling block; The prior distribution computation block comprises three sequentially connected prior distribution computation sub-modules. Each prior distribution computation sub-module includes a sequentially connected one-dimensional convolutional layer 1, a batch normalization layer 1, and a mapping layer 1. The input and output dimensions of the one-dimensional convolutional layer 1 are both... In this example The value is 256; the mapping layer is a one-dimensional convolutional layer. The input dimension is The output is a prior distribution expressed in terms of mean and variance, with dimension 1. ; The posterior distribution computation block includes sequentially connected one-dimensional convolutional layers. The system consists of three posterior distribution computation submodules and a second mapping layer. Each posterior distribution computation submodule includes a second one-dimensional convolutional layer and a second batch normalization layer. The one-dimensional convolutional layer... The input is another contact point that is used as the gripper contact point pair, containing the current sampling point. Gripper rotation angle and gripping posture tolerance Tag data set The high-dimensional feature data output by the feature fusion module has an input dimension of [missing information]. The output dimension is The second mapping layer is a one-dimensional convolutional layer. The input dimension is The output is the posterior distribution with dimension . ; The latent distribution module takes high-dimensional feature data (i.e., feature information) as input. This involves capturing the complex mapping relationship between high-dimensional feature data and the continuous distribution of the grasping posture space through the latent variable space, with an output sample length of [missing information]. , length is The latent variables, in this example .

[0018] In some embodiments, the grasping parameter estimation module includes sequentially connected one-dimensional convolutional layers. Dropout layer, one-dimensional convolutional layer One-dimensional convolutional layer Input dimension is The output dimension is One-dimensional convolutional layer Input dimension is The output dimension is 5.

[0019] In some embodiments, S3 specifically includes the following steps: S31. Randomly select a scene point cloud containing the target object from the training set, and use sampling methods to sample N points from the selected scene point cloud to obtain the input point cloud. ; S32, Input point cloud The input is fed into the feature encoding module, where point set grouping blocks calculate the input point cloud. The distance between each point in the input point cloud and the anchor points obtained by random sampling is calculated, and the results are used to adjust the input point cloud. The points are grouped; then the feature clustering blocks merge the features of each group of points, and the current point set feature is used as the current anchor point feature value to obtain the feature data. ; S33, Transfer feature data The input is fed into the feature fusion module for feature modeling. The forward state space model and backward state space model in the feature fusion module capture the correlation of point cloud features at different distances, and output (high-dimensional) feature information. ; S33 is expressed by a formula, as follows: ; ; ; ; In the formula, The feature data output by the feature encoding module. This indicates that it belongs to one layer. This indicates the output of a single layer; Representing the state-space model layer The output; Representing the state-space model layer The output; This indicates a feature dimension concatenation operation. Indicates a fully connected layer. This represents the feature data output by the feature fusion module; S34. Transfer feature information The set of labeled data with the target object's grasping posture is input into the latent distribution module, and the prior distribution calculation block calculates the feature information. The mapping is represented by a prior distribution in terms of mean and variance, and the posterior distribution is used to compute block-by-block feature information. The dataset consists of a label set with the target object's grasping posture, and the concatenated result is mapped to a posterior distribution represented by mean and variance. The prior and posterior distributions are aligned using relative entropy loss KL (Kullback-Leibler Divergence), and the prior distribution is constrained by the posterior distribution to make the prior distribution more closely match the mapping relationship from features to the true label. The latent variable sampling block samples the prior distribution through reparameterization to generate latent variable z. S34 is expressed by a formula, as follows: ; ; ; ; ; In the formula, and These represent the prior distribution computation submodule in the prior distribution computation block and the posterior distribution computation submodule in the posterior distribution computation block, respectively. This represents the output of the prior distribution calculation submodule; This represents the latent variable sampling operation for reparameterization. and Let these represent the prior and posterior distributions in terms of mean and variance, respectively. This represents the output of the posterior distribution calculation submodule; S35, The parameter estimation module receives feature information. The model calculates the latent variable z and outputs the prediction results of the gripping detection model. The prediction results include the predicted gripper contact point pairs, the predicted gripper rotation angle, the predicted gripping quality score, and the predicted gripping posture tolerance. The predicted gripper contact point pairs consist of two contact points, one of which is the current sampling point, and the other is the contact point predicted by the gripping detection model.

[0020] The input to the parameter estimation module is high-dimensional feature data (i.e., feature information). The model maps the latent variables back to the grasping posture using high-dimensional features as a reference. The output includes the predicted gripper contact point pair, the predicted gripper rotation angle, the predicted grasping quality score, and the predicted grasping posture tolerance. The predicted gripper contact point pair consists of two contact points, one of which is the current sampling point, and the other is the contact point predicted by the grasping detection model.

[0021] In some embodiments, the total loss function is specifically as follows: ; in, Indicates the total loss. , , , , These represent contact point loss, gripper rotation loss, gripping mass loss, gripping tolerance loss, and relative entropy loss KL, respectively. , , , , These are weighting coefficients; , , , , The specific calculation formula is as follows: ; ; ; ; ; in, This represents the mean squared error loss function; Represents the cross-entropy loss function; , , , These represent the predicted gripper contact point pairs, the predicted gripper rotation angle, the predicted gripping quality fraction, and the predicted gripping posture tolerance, respectively. Mean squared error loss function Cross-entropy loss function The calculation formulas are as follows: ; ; Where n represents the current number of training samples; U represents the first input parameter, which is the label data. Inside , , , One of the parameters in; U i This represents the i-th training sample of the first input parameter; This represents the second input parameter, which is... , , , One of the parameters inside, This represents the i-th training sample of the second input parameter. Represents a logarithmic function.

[0022] This invention establishes forward and backward state space models using a state space model, providing sufficient feature data for subsequent grasping detection. The excellent long-range modeling capability of the state space model ensures the accuracy of tasks relying on globally related information, such as graspable position identification and collision detection, during the grasping detection process, thus greatly improving the accuracy of grasping detection. The latent distribution module replaces the one-to-one mapping relationship of the deterministic model, representing the continuous distribution of grasping postures through latent distribution, effectively improving feature utilization and increasing grasping detection accuracy.

[0023] Furthermore, this invention utilizes a state-space model to construct a grasping and detection model. The state-space model possesses long-range modeling capabilities comparable to the Transformer, while avoiding redundant variables in the computational process, thus saving computational costs. Moreover, the linear expression of the state-space model effectively leverages the parallel computing power of the GPU, accelerating the convergence speed of the neural network.

[0024] In another aspect, the present invention provides a grasping and detection system based on a state-space model, including a device end, which performs grasping and detection on a target object according to a grasping and detection method.

[0025] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A grasping and detection method based on a state-space model, characterized in that, Includes the following steps: S1. Obtain several scene point clouds containing the target object, label each scene point cloud separately to obtain a set of labeled data with the target object's grasping posture, and divide the set of labeled data into a training set and a test set according to a set ratio. S2. Build a grasping detection model based on a state space model. The grasping detection model includes a feature encoding module, a feature fusion module, a latent distribution module, and a grasping parameter estimation module connected in sequence. S3. Train the crawling detection model using the training set, set the total loss function, calculate the total loss based on the prediction results of the crawling detection model and the label data set, adjust the weights of the crawling detection model based on the total loss, minimize the total loss, until the set iteration stopping condition is reached. S4. Evaluate the trained grasping detection model using the test set, and save the model parameters that perform best on the test set as the final grasping detection model. Input the real-world scene point cloud containing the target object into the final grasping detection model to obtain the detection results.

2. The grasping and detection method based on a state-space model according to claim 1, characterized in that, S1 specifically includes the following steps: S11. Obtain several scene point clouds containing the target object, and label the target in each scene point cloud. The labeling content includes the grasping posture quality score, gripper contact point pairs, gripper rotation angle, and grasping posture tolerance. Obtain a labeled data set with the grasping posture of the target object, expressed by the following formula: ; ; ; in, This represents a set of labeled data with the target object's grasping posture. , , , These represent the grasping mass fraction, gripper contact point pairs, gripper rotation angle, and grasping posture tolerance, respectively. Represents the set of real numbers; Represents the number of points in the scene's point cloud; S12. Collect the tagged data containing the target object's grasping posture according to the set ratio. It is divided into training set and test set.

3. The grasping and detection method based on a state-space model according to claim 2, characterized in that, The feature encoding module includes point set grouping blocks and feature clustering blocks connected in sequence, wherein the point set grouping blocks are used to encode the input point cloud. The distance between each point in the input point cloud and the anchor point obtained by random sampling affects the input point cloud. The points are grouped, and the feature clustering block is used to merge the features of each group of points, and the feature of the current point set is used as the feature value of the current anchor point.

4. The grasping and detection method based on a state-space model according to claim 3, characterized in that, The feature fusion module consists of a normalization layer, a forward state space model, a backward state space model, and a projection layer. The forward state space model consists of sequentially connected one-dimensional convolutional layers. and state space model layer ; The backward state space model consists of sequentially connected one-dimensional convolutional layers. and state space model layer The projection layer contains a fully connected layer.

5. The grasping and detection method based on a state-space model according to claim 4, characterized in that, The latent distribution module includes a prior distribution calculation block, a posterior distribution calculation block, and a latent variable sampling block. The prior distribution computation block comprises three sequentially connected prior distribution computation sub-modules. Each prior distribution computation sub-module includes a sequentially connected one-dimensional convolutional layer 1, a batch normalization layer 1, and a mapping layer 1; the mapping layer 1 is a one-dimensional convolutional layer. ; The posterior distribution computation block includes sequentially connected one-dimensional convolutional layers. The system consists of three posterior distribution computation submodules and a second mapping layer. Each posterior distribution computation submodule includes a second one-dimensional convolutional layer and a second batch normalization layer. The second mapping layer is a one-dimensional convolutional layer. .

6. The grasping and detection method based on a state-space model according to claim 5, characterized in that, The grasping parameter estimation module includes sequentially connected one-dimensional convolutional layers. Random deactivated layer, one-dimensional convolutional layer .

7. The grasping and detection method based on a state-space model according to claim 6, characterized in that, S3 specifically includes the following steps: S31. Randomly select a scene point cloud containing the target object from the training set, and use sampling methods to sample from the selected scene point cloud. From these points, we obtain the input point cloud. ; S32, Input point cloud The input is fed into the feature encoding module, where point set grouping blocks calculate the input point cloud. The distance between each point in the input point cloud and the anchor points obtained by random sampling is calculated, and the results are used to adjust the input point cloud. The points are grouped; then the feature clustering blocks merge the features of each group of points, and the current point set feature is used as the current anchor point feature value to obtain the feature data. ; S33, Transfer feature data The input is fed into the feature fusion module for feature modeling. The forward state space model and backward state space model in the feature fusion module capture the correlation of point cloud features at different distances and output feature information. ; S34. Transfer feature information The set of labeled data with the target object's grasping posture is input into the latent distribution module, and the prior distribution calculation block calculates the feature information. The mapping is represented by a prior distribution in terms of mean and variance, and the posterior distribution is used to compute block-by-block feature information. The dataset contains labeled data with the target object's grasping posture, and the stitched results are mapped to a posterior distribution expressed in terms of mean and variance; the latent variable sampling block samples the prior distribution through reparameterization to generate the latent variable z. S35, The parameter estimation module receives feature information. The model calculates the latent variable z and outputs the prediction results of the gripping detection model. The prediction results include the predicted gripper contact point pairs, the predicted gripper rotation angle, the predicted gripping quality score, and the predicted gripping posture tolerance. The predicted gripper contact point pairs consist of two contact points, one of which is the current sampling point, and the other is the contact point predicted by the gripping detection model.

8. The grasping and detection method based on a state-space model according to claim 7, characterized in that, S33 is expressed by a formula, as follows: ; ; ; ; In the formula, The feature data output by the feature encoding module. This indicates that it belongs to one layer. This indicates the output of a single layer; Representing the state-space model layer The output; Representing the state-space model layer The output; This indicates a feature dimension concatenation operation. Indicates a fully connected layer. This represents the feature data output by the feature fusion module; S34 is expressed by a formula, as follows: ; ; ; ; ; In the formula, and These represent the prior distribution computation submodule in the prior distribution computation block and the posterior distribution computation submodule in the posterior distribution computation block, respectively. This represents the output of the prior distribution calculation submodule; This represents the latent variable sampling operation for reparameterization. and Let represent the prior and posterior distributions expressed in terms of mean and variance, respectively. This represents the output of the posterior distribution calculation submodule.

9. The grasping and detection method based on a state-space model according to claim 8, characterized in that, The total loss function is as follows: ; in, Indicates the total loss. , , , , These represent contact point loss, gripper rotation loss, gripping mass loss, gripping tolerance loss, and relative entropy loss KL, respectively. , , , , These are weighting coefficients; , , , , The specific calculation formula is as follows: ; ; ; ; ; in, This represents the mean squared error loss function; Represents the cross-entropy loss function; , , , These represent the predicted gripper contact point pairs, the predicted gripper rotation angle, the predicted gripping quality fraction, and the predicted gripping posture tolerance, respectively. Mean squared error loss function Cross-entropy loss function The calculation formulas are as follows: ; ; Where n represents the current number of training samples; This indicates the first input parameter, which is the label data. Inside , , , One of the parameters in; This represents the i-th training sample of the first input parameter; This represents the second input parameter, which is... , , , One of the parameters inside, This represents the i-th training sample of the second input parameter. Represents a logarithmic function.

10. A grasping and detection system based on a state-space model, characterized in that, Includes a device end, which performs grasping and detection on a target object according to the grasping and detection method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Intelligent agent training method and device, computer equipment and storage medium

    CN113919482A

  • Double-stage mechanical arm grabbing planning method and system based on multiplexing structure

    CN114800511A

  • Control method and system based on DO model and time delay feedforward model

    CN118466198A

  • 3D point cloud target detection method based on state space model

    CN119741697A

  • Grabbing attitude generation method and system based on multi-modal large model

    CN120588235A