Target grasping point generation method based on semi-supervised generative reasoning network

By constructing a target grab point generation method based on a semi-supervised generative inference network, and combining target detection and grab inference networks, the problem of high cost of labeled data in supervised learning is solved. This method enables fast and accurate generation of target grab points with limited labeled data, reduces computational cost and improves detection accuracy.

CN119850911BActive Publication Date: 2025-10-28TIANJIN BINGGUANJIA TECHNOLOGY DEVELOPMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411903089.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-28
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

In existing technologies, waste sorting and detection methods mainly rely on supervised learning, which requires a large amount of labeled data, resulting in high data labeling costs and scarce labeled data. Semi-supervised learning methods are rarely used, and there is a lack of efficient training methods.

Method used

We employ a semi-supervised generative inference network, combining an object detection network and a grasping inference network. By constructing a channel-switched generative residual grasping inference network CE-GR-NET, and utilizing a semi-supervised VAE to add equality constraints of deconvolution representation during the training phase, we reduce the number of labeled samples and achieve fast and accurate generation of object grasping points.

Benefits of technology

Under limited labeled data conditions, the system achieves accurate generation of target capture configurations in images, reduces the number of labeled samples required, improves the efficiency and accuracy of capture detection, and has low computational cost and high speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850911B_ABST
    Figure CN119850911B_ABST
Patent Text Reader

Abstract

This invention discloses a method for generating target grabbing points based on a semi-supervised generative inference network, comprising: acquiring multi-source solid waste images and constructing datasets: an image dataset X0 labeled with the presence or absence of recyclable targets, an image dataset X1 containing recyclable targets and grabbing configuration labels with those targets, and an image dataset X2 containing recyclable targets but without grabbing configuration labels with those targets; constructing a network model, including an object detection network and a grabbing inference network, wherein the grabbing inference network adopts VQ-VAE; training the object detection network using X0, and training the grabbing inference network using X1 and X2 in a semi-supervised manner; for a new multi-source solid waste image, using the trained object detection network to determine whether there are recyclable targets; if there are recyclable targets, using the trained grabbing inference network to generate the corresponding grabbing configuration. This invention can accurately generate grabbing configurations for targets in images under the condition of limited label data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robotic arm grasping data processing, specifically relating to a method for generating target grasping points based on a semi-supervised generative inference network. Background Technology

[0002] With the rapid development and widespread application of artificial intelligence, deep learning technology can assist people in garbage sorting. This not only replaces manual sorting in harsh environments but also significantly improves the accuracy of garbage detection and sorting precision. Manually labeling datasets consumes significant human and material resources, and current mainstream garbage detection methods are primarily based on supervised learning, requiring large amounts of labeled data. Semi-supervised methods, however, require only a small amount of labeled data to achieve high accuracy, making them an effective means of solving the problems of high data labeling costs and scarce labeled data. Semi-supervised learning methods are less commonly used in the field of garbage detection, and there is still a need to develop efficient semi-supervised training methods. Summary of the Invention

[0003] This invention provides a target grabbing point generation method based on a semi-supervised generative inference network, which can accurately generate the grabbing configuration of the target in the image under the condition of limited label data.

[0004] To achieve the above technical objectives, the present invention adopts the following technical solution:

[0005] A method for generating target grabbing points based on a semi-supervised generative inference network includes:

[0006] Step 1, acquire multi-source solid waste images and construct a dataset, including: (1) an image dataset labeled as having recyclable targets. (2) Image datasets with retrievable targets and labeled with target grabbing configurations. (3) Image datasets containing retrievable targets but without target-specific tags for capture configuration.

[0007] Among them, w m Let l represent the m-th image sample in dataset X0. m Represents image sample w m The label used to indicate whether a recyclable target exists is N0, where N0 represents the number of samples in dataset X0. Let y represent the nth image sample in dataset X1 containing a recyclable target. n Represents image samples The capture configuration labels for recyclable targets, where N1 represents the number of samples in dataset X1. N represents the nth image sample in dataset X2 that contains a recyclable target, and N2 represents the number of samples in dataset X2.

[0008] Step 2: Construct the network model, including the object detection network and the grasping and inference network;

[0009] Step 3: Train the object detection network using dataset X0, and train the grasping inference network using datasets X1 and X2 in a semi-supervised manner.

[0010] Step 4: For the newly arrived multi-source solid waste image, use the trained object detection network to determine whether there are recyclable targets; if there are recyclable targets, then use the trained grasping inference network to generate the corresponding grasping configuration.

[0011] Furthermore, the target detection network employs a binary classification neural network.

[0012] Furthermore, the crawling inference network includes three convolutional layers in the encoding stage, three convolutional transpose layers in the decoding stage, and five residual layers between the encoding and decoding stages.

[0013] Furthermore, the encoding stage includes two different branches, which respectively input RGB color images and depth images; on the two branches: (1) firstly, convolution operation is performed on the respective input images in the convolutional layer, then normalization operation is performed in the BN layer, and then the data in the two channels below the specified threshold is replaced with the data in the other channel, and then the data after channel swapping is activated; (2) the second channel swapping and the third channel swapping are performed according to the same process as (1); (3) the feature maps obtained from the two branches are spliced ​​along the channel dimension, and then sent to the residual layer for processing.

[0014] Furthermore, the step of using datasets X1 and X2 and training the crawling inference network in a semi-supervised manner includes:

[0015] Step a1: Initialize the parameter set of the crawling inference network. Among them, C k Representing the k-th category, P(C k ) represents the distribution model corresponding to the k-th category, where K is the number of distribution models, and μ k and δ k These represent the mean and variance of the k-th distribution model, respectively.

[0016] Step a2, update each distribution model and its mean:

[0017]

[0018] In the formula, M k This represents the number of image samples in the k-th category; Indicates data without labels Under the condition that it belongs to the kth category C k The posterior probability;

[0019] Step a3, maximize the following log-likelihood function:

[0020]

[0021] In the formula, L(ψ) is the likelihood function, which represents the joint probability of all image samples in datasets X1 and X2 given the model parameter set ψ; Indicates labeled data The joint probability is determined by the model parameter set ψ;

[0022] Indicates unlabeled data The marginal probability is obtained by summing over all possible categories, i.e.

[0023]

[0024] Step a4: Iteratively update the parameter set ψ until convergence.

[0025] Furthermore, the crawling and inference network employs a vector quantization variational autoencoder, i.e., VQ-VAE, which includes encoder P. ε (z e |x,y), quantization layer Q φ (z e )→z q Decoder P ξ (x|y,z q ) and detector P ζ (y|x); ε represents the encoder parameters, ξ represents the decoder parameters, ζ represents the detector parameters, x represents the image sample, y represents the label corresponding to image sample x, and z represents the image sample x. e The latent features obtained by the encoder; z q For z e The latent features obtained after quantization;

[0026] The method of using datasets X1 and X2 and training the crawling and inference network in a semi-supervised manner includes:

[0027] First, the encoder, quantization layer, and decoder in the crawling and inference network are pre-trained using the unlabeled dataset X2 and based on the following objective function:

[0028]

[0029] In the formula, U(ε,ξ,y;x,γ) is the pre-training objective function, γ is a Lagrange multiplier and γ>0. It is the loss function; K(P) ζ(y|x)) is the KL divergence, used to constrain the complexity of the distribution; E P(x,y) logP ζ (y|x) is the core objective, used to directly optimize model accuracy, while KL divergence is an indirect constraint, serving as an auxiliary term.

[0030] The encoder is then connected to the detector after passing through a quantization layer;

[0031] Then fix the parameters of the encoder and quantization layer, and train the parameters of the detector using the labeled dataset X1 and based on the following classification loss function;

[0032] E P(x,y) logP ζ (y|x)

[0033] In the formula, P(x,y) represents the joint probability distribution of image sample x and label y; in the objective function, the joint distribution P(x,y) is used to calculate the expectation E. P(x,y) logP ζ (y|x) represents the model's average classification performance on the data distribution; the main role of P(x,y) is to serve as the expected weight distribution, hoping that the model can maximize the conditional log-likelihood on the true distribution P(x,y), that is, optimize the model P by maximizing this expectation. ζ (y|x) to more accurately predict the label y given x;

[0034] The encoder, after passing through the quantization layer, is connected to the detector and trained to form the overall grasping and inference network.

[0035] Furthermore, the image samples w in dataset X0 m For RGB images; image samples x from datasets X1 and X2. n Includes RGB color images and depth images.

[0036] Furthermore, the crawling configuration includes: crawling center, crawling angle, crawling width, crawling quality score, and crawling target category.

[0037] This invention presents a target grabbing point generation method based on a semi-supervised generative inference network. By fusing a target detection network and a grabbing inference network, a channel-switched generative residual grabbing inference network CE-GR-NET is constructed, enabling rapid and accurate generation of grabbing points for solid waste. In the training phase of the grabbing inference network, a semi-supervised VAE is used. After estimating the posterior distribution of the data, an equality constraint of deconvolution representation is added to guide the learning of the latent representation, making the semi-supervised learning of image data more effective and reducing the number of labeled samples. Attached Figure Description

[0038] Figure 1This is an overall flowchart of the target grabbing point generation method based on a semi-supervised generative inference network according to an embodiment of the present invention.

[0039] Figure 2 These are the three specific locations for channel switching in this embodiment of the invention.

[0040] Figure 3 This is a diagram illustrating the training strategy of a semi-supervised crawling and detection network according to an embodiment of the present invention.

[0041] Figure 4 This is a visualization result of a self-built dataset crawling experiment in an embodiment of the present invention. Detailed Implementation

[0042] The method described in this embodiment can be tested using the Python programming language, or it can be applied in engineering using the C / C++ programming language. This invention does not impose any specific limitations.

[0043] This embodiment provides a target grabbing point generation method based on a semi-supervised generative inference network. For the overall framework and processing flow, please refer to [link to documentation]. Figure 1 This includes the following steps:

[0044] Step 1: Construct the following dataset required for training the embodiment of this invention based on the multi-target grasping dataset:

[0045] (1) Image datasets that have been labeled to indicate whether recyclable targets exist. Among them, w m Let w represent the m-th image sample in dataset X0, and let w be the image sample m. m For RGB images, l n Represents image sample w m The label used to indicate whether a recyclable target exists is N0, where N0 represents the number of samples in dataset X0.

[0046] (2) Image datasets with retrievable targets and labeled with target grabbing configurations. in, Let x represent the nth image sample in dataset X1 containing a recyclable target, and let x be the image sample x. n Includes RGB color images and depth images, y n Represents image samples The capture configuration label for recyclable targets, where N1 represents the number of samples in dataset X1;

[0047] (3) Image datasets containing retrievable targets but without target-specific tags for capture configuration. in, This represents the nth image sample in dataset X2 containing a recyclable target, and image sample x... nIncludes RGB color images and depth images, where N2 represents the number of samples in dataset X2.

[0048] Step 2: Integrate the object detection network and the grasping inference network to construct the generative residual grasping inference network CE-GR-NET based on channel switching.

[0049] The target detection network is used to detect whether there is a specific target, recyclable solid waste, in an RGB image. Specifically, it can be implemented using a conventional binary classification neural network.

[0050] If the object detection network determines that there is recyclable solid waste in the RGB image, it will pass the RGB image and the corresponding depth image to the grasping inference network. The grasping inference network will generate a pair of grasping points for each pixel of the image and obtain the grasping quality, grasping angle and required gripper width image of each pair of grasping points, thereby predicting the best grasping point and generating a grasping rectangle.

[0051] The crawling and inference network comprises three convolutional layers in the encoding stage, three transposed convolutional layers in the decoding stage, and five residual layers between the encoding and decoding stages. The convolutional layers extract features from the input image, extracting high-level features through nonlinear transformations, and then input the output of the convolutional layers into the five residual layers. In neural networks, each layer transforms the output of the previous layer. As the number of network layers increases, the accuracy also increases. However, such transformations cause the input information to be continuously compressed and transformed, leading to information loss and gradient vanishing, making information transmission very difficult and prone to gradient vanishing problems, making the model difficult to train. In the residual layers, the output value is the input value of the previous layer plus a residual term, which is the difference between the input and output value of the previous layer. The introduction of residual layers allows information to be transmitted more smoothly in the network, improving the gradient vanishing problem. After passing through these convolutional and residual layers, the image size is reduced to 56×56. To facilitate the interpretation and preservation of the spatial features of the image after the convolution operation, the image is then upsampled using the convolution transpose operation, so that the size of the output image is the same as that of the input image, which is 224×224.

[0052] In the encoding stage of the inference network, a channel swapping method is used to enhance the network's ability to perceive images of different modalities. The encoding stage includes two different branches, which take RGB color images and depth images as inputs, respectively. On the two branches: (1) First, convolution operations are performed on the respective input images in the convolutional layer, then normalization operations are performed in the BN layer, and then the data of the two channels are swapped, and then the data after channel swapping is activated; (2) The second and third channel swaps are performed in the same way as (1); (3) The feature maps obtained from the two branches are concatenated along the channel dimension and then sent to the residual layer for processing. See the three specific locations of channel swapping for details. Figure 2 .

[0053] The grabbing configuration for the i-th object is defined as follows:

[0054] G i =(x,y,θ) i W i Q, C)

[0055] Where (x,y) is the grab center, θ i The angle of capture, that is, the amount of rotation required to capture an object, is expressed as... The value of W; i This indicates the capture width, which is [0, W]. max Values ​​within the range, W max Q represents the actual maximum width of the gripper; Q represents the gripping quality score, which is a score between 0 and 1, with a higher success rate as it is closer to 1; C is the object category number generated along with the gripping quality score Q, which is a score between 0 and 13, with each score corresponding to an object category.

[0056] For a specific grab rectangle, Redmon et al.'s method is used.

[221] The proposed evaluation criteria: A candidate grab rectangle is considered valid only if it meets the following conditions:

[0057] 1) The angle between the predicted grab rectangle and the actual grab rectangle is less than 30°;

[0058] 2) Predicted grab rectangle G i Compared to a real grab rectangle The IoU score between them is greater than 0.25:

[0059]

[0060] For objects containing D = {D1…D2} n The dataset consists of input scene images I = {I}. 1 …I N} and each object D i Crawling configuration set The model is trained end-to-end by minimizing the negative log-likelihood of a grasping configuration set conditioned on the input image scene I, thereby learning the mapping function γ(I,D)=G. i Defined as:

[0061]

[0062] The model uses the Adam optimizer for parameter updates. During training, it employs the standard backpropagation algorithm and mini-batch stochastic gradient descent technique, with a learning rate set to 10. -3Eight samples are randomly selected for training each time. A random seed is used during training to ensure that the same random split is obtained when training is run at different times, thus eliminating the differences introduced by randomness.

[0063] To mitigate the vanishing gradient problem and smooth the Huber loss, the network loss is defined as follows:

[0064]

[0065] Among them, G i It is the crawl configuration generated by the crawl inference network. It is a genuine capture of configuration tags.

[0066] Step 3: In the model training phase of Step 2, a semi-supervised VAE is used. After estimating the posterior distribution of the data, an equality constraint for deconvolution representation is added to guide the learning of the latent representation, making semi-supervised learning of image data more effective and reducing the number of labeled samples. See [link to semi-supervised grasping and detection network training strategy] for details. Figure 3 Specifically as follows,

[0067] In real life, labeled data is often scarce, with most data being unlabeled, because collecting a large number of labels is time-consuming and labor-intensive. The idea behind semi-supervised learning is to make good use of the supervision information from a small portion of labeled data obtained based on prior knowledge, while also fully capturing structural and other information from unlabeled data to train a good model.

[0068] In general, the amount of unlabeled data is much larger than the amount of labeled data. Therefore, in this embodiment of the invention, the number of samples N1 in dataset X1 is much larger than the number of samples N2 in dataset X2. Although unlabeled data lacks labels, some useful information can be obtained by modeling its distribution. For semi-supervised generative models, considering unlabeled data may affect the mean μ, variance δ, and decision boundary of the distribution model. Updating the model allows it to converge, ultimately yielding a new generative model. For example, the update method for a model with a distribution size K of 2 is as follows:

[0069] First, initialize the parameters ξ = {P(C1), P(C2), μ1, μ2, δ1, δ2} using random initialization, and calculate the posterior probability P of the unlabeled data. ξ (C1|x u Next, update the model:

[0070]

[0071] Update P(C2) and μ in the same way. 2 The above calculations are performed iteratively until convergence.

[0072] If only labeled data is available, then maximizing the following log-likelihood function is sufficient to solve the problem, and a closed-form solution exists:

[0073]

[0074] To consider unlabeled data, we need to maximize the following log-likelihood function:

[0075]

[0076] P in the formula ξ (x u The calculation is performed according to the following formula:

[0077] P ξ (x u ) = P ξ (x|C1)P(C1)+P ξ (x|C2)P(C2)

[0078] Because the addition of unlabeled data causes the maximum log-likelihood function in the above equation to have no closed-form solution, the parameters need to be updated iteratively. That is, completing the above P(C1) and μ1 once will increase the log-likelihood function slightly, eventually leading to convergence.

[0079] The effectiveness of deep generative models in acquiring data distributions has led to the popularity of semi-supervised models based on deep generative models, with many generative semi-supervised methods based on VAEs. A typical VAE consists of an encoder network P. ε (z|x) and a decoder network P ξ The semi-supervised VAE consists of (x|z) layers. The encoder network encodes the input x into a latent representation z, and the decoder network reconstructs x from z. The basic idea of ​​a semi-supervised VAE is to add a detector on top of the latent representation. Therefore, a semi-supervised VAE typically consists of three main components: the encoder network P... ε (z|x,y), decoder network P ξ (x|y,z) and detector P ε (y|x). In the model training phase, this invention employs semi-supervised VQ-VAE. After estimating the posterior distribution of the data, it adds equality constraints of deconvolution representation to guide the learning of latent representation, making semi-supervised learning of image data more effective and reducing the number of labeled samples.

[0080] Variational autoencoders (VQ-VAEs) are among the most popular deep generative models. A key step in their development is calculating P... ε (x), the relevant calculation is:

[0081]

[0082] in It means sg[z e ] and z q The L2 norm, sg[z e The symbol ] indicates that the gradient update operation is stopped, ensuring that the quantization error is used only to update the vector z in the codebook. q This does not affect the backpropagation of the encoder parameters. L(ε,ξ;x) is decomposed using ELBO and defined as:

[0083]

[0084] In the above formula, P ε The (z|x) term acts as an encoder, extracting latent features from the observed data x. P is obtained by minimizing the quantization error. ε The (z|x) term can be used to calculate the approximate true posterior distribution P. ξ (x|z). The loss function can be rewritten as follows:

[0085]

[0086] The first term is the reconstruction error (RE); the second term is the encoder output z. e And quantization representation z q The mean square error between them is used to measure z e Pull towards z q Here, sg indicates stopping backpropagation to prevent codebook updates; the third term reverses z. q Pull towards z e However, it stops the gradient update of the encoder parameters. This term is used to make the choice of quantization vector more stable, where β is a tradeoff coefficient that balances the effects of the second and third terms.

[0087] To achieve semi-supervised learning, a detector K(P) is introduced into the above equation. ζ (y|x)) yields:

[0088]

[0089] When faced with a dataset containing partially labeled data, the classification loss E based on the label information is... P(x,y) logP ζ (y|x) is added to the objective function.

[0090] Step 4: Anchor boxes are used on the output feature map to indicate the possible locations of the target. Based on the anchor boxes, the final output vector of class probability, confidence score, and bounding box is generated, thereby realizing the generation of the grab point.

[0091] To facilitate understanding of the technical effects of this invention, a comparison of the application of this invention and conventional methods is provided below:

[0092] The proposed CE-GR-NET network has over 1.9 million parameters, making it relatively small in size compared to other crawling networks. Therefore, compared to other architectures using similar mastery prediction techniques with millions of parameters and complex structures, this network has lower computational costs and faster speed. Table 1 shows a comparison of accuracy and crawling inference speed with other crawling networks on the Cornell crawling dataset, demonstrating the effectiveness of the proposed method.

[0093] Table 1 Comparison of Web Scraping Performance

[0094]

[0095]

[0096] The network was trained on the Cornell crawl dataset for object detection. The results are shown in Table 2. Precision represents the proportion of true positives in the prediction results, while Recall represents the proportion of all positives that were correctly predicted. mAP50 is the average precision of all images in each class when IoU is set to 0.5, and then the average is calculated for all classes. mAP50-95 represents the average mAP at different IoU thresholds (from 0.5 to 0.95, with a step size of 0.05).

[0097] Table 2 Detection results from the Cornell crawl dataset

[0098]

[0099] As shown in Table 2, the object detection network has high precision, recall, and mAP50 for the above six types of objects. One important reason is that the Cornell crawl dataset is based on crawling objects of a single object category with clear visual features.

[0100] In this embodiment of the invention, an image scene in the MSSW dataset may contain multiple targets to be captured. In this case, the target detection network first performs detection according to the dataset's primary classification. If a target category exists, the corresponding depth and color images are sent to the capture detection network. The capture detection network generates capture configurations for all targets in the image. When the center of a capture configuration falls within the target detection result area, the corresponding capture configuration is converted into a rectangular representation. The classification result of the target detection network is a secondary classification result for subsequent retrieval processing. See the capture results below. Figure 4 .

[0101] Regarding model accuracy, this embodiment uses a self-built MSSW dataset for repeated training five times. The model with the highest crawling accuracy in the validation set is selected as the base prediction model for subsequent steps. If multiple models have the same crawling accuracy, the one with the highest classification accuracy is selected. On the self-built dataset, the accuracy of the image-based model is 96.03%, the accuracy of the object-based model is 94.40%, and the object classification accuracy is 97.87%. The average accuracy comparison between the method of this invention and some generative crawling network models is shown in Table 3.

[0102] Table 3 Comparison of the average accuracy of the method of the present invention with other models.

[0103]

[0104] Note: Bold numbers indicate the best performance results.

[0105] The semi-supervised learning method used employs the VQ-VAE model as a feature extractor. The VQ-VAE model consists of an encoder, a quantization layer, and a decoder. The training strategy for the semi-supervised grasping and detection network is as follows: Figure 3 As shown in the diagram. Specifically, unlabeled color images are first input into the VQ-VAE model for pre-training, representing high-dimensional image information as discrete low-dimensional vectors. Then, the pre-trained VQ-VAE model is connected to the grasping and detection network CE-GR-NET. Afterward, a small amount of labeled data is input for training, during which the pre-trained weights of the encoder and quantization layers are frozen, and the weights of the decoder are initialized.

[0106] The comparative experiments included three training methods: Method 1: conventional training with a training set to test set ratio of 9:1; Method 2: semi-supervised training with a training set to test set ratio of 5:5; and Method 3: conventional training with a training set to test set ratio of 5:5. Method 1 served as the upper limit for semi-supervised training, while Method 3 served as the baseline for semi-supervised training. Eight sets of comparative experiments were conducted, with identical training and test sets in each set. The capture detection accuracy comparisons are shown in Table 4, where Method 1 achieved an average accuracy of 95.92%, Method 2 90.98%, and Method 3 86.83%. It is evident that the semi-supervised training method used in this section outperforms the conventional training method under the same experimental conditions. Although its accuracy is lower than Method 1, it significantly reduces the number of training samples required, achieving high detection accuracy while reducing the cost of labeled samples.

[0107] Table 4 Comparison of Accuracy of Different Training Methods

[0108]

[0109] The above embodiments are preferred embodiments of this application. Those skilled in the art can make various changes or improvements based on them. Without departing from the overall concept of this application, these changes or improvements should fall within the scope of protection claimed in this application.

Claims

1. A method for generating target grasping points based on a semi-supervised generative inference network, characterized in that, include: Step 1, acquire multi-source solid waste images and construct a dataset, including: (1) an image dataset labeled as having recyclable targets. (2) An image dataset with retrievable targets and labeled grabbing configurations. (3) Image datasets with retrievable targets but without target-specific tags for capture configuration. ; in, Represents the dataset The first in Image samples, Represents image samples Tags used to indicate the presence of recyclable targets. Represents the dataset The number of samples; Represents the dataset The Middle An image sample containing a recyclable target. Represents image samples Capture and configure tags for recyclable targets. Represents the dataset The number of samples, Represents the dataset The Middle An image sample containing a recyclable target. Represents the dataset The number of samples; Step 2: Construct the network model, including the object detection network and the grasping and inference network; The grasping and inference network includes three convolutional layers in the encoding stage, three convolutional transpose layers in the decoding stage, and five residual layers between the encoding and decoding stages. The encoding stage includes two different branches, which take RGB color images and depth images as inputs, respectively. On the two branches: (b1) First, convolution operations are performed on the respective input images in the convolutional layers, then normalization operations are performed in the BN layer, and then data in the two channels below a specified threshold are replaced with data in the other channel, and then the data after channel swapping is activated; (b2) The second and third channel swaps are performed according to the same process as (b1); (b3) The feature maps obtained from the two branches are concatenated along the channel dimension and then sent to the residual layers for processing. Step 3, using the dataset Train the object detection network using the dataset and The grasping and reasoning network is trained using a semi-supervised approach. The crawling and inference network employs a vector quantization variational autoencoder, or VQ-VAE, which includes an encoder. Quantization layer decoder and detector ; Indicates the parameters of the encoder, Indicates the parameters of the decoder. The parameters representing the detector, Represents image samples, Represents image samples The corresponding tags The latent features obtained by the encoder; for The latent features obtained after quantization; The use of datasets and The grasping and reasoning network is trained using a semi-supervised approach, including: First, use an unlabeled dataset. The encoder, quantization layer, and decoder in VQ-VAE are pre-trained based on the following objective function: ; In the formula, For the pre-trained objective function, For Lagrange multipliers and , It is a loss function; It is the KL divergence, used to constrain the complexity of the distribution; The core objective is to directly optimize model accuracy, while KL divergence is an indirect constraint, serving as an auxiliary term. The encoder is then connected to the detector after passing through a quantization layer; Then fix the parameters of the encoder and quantization layer, and use a labeled dataset. The detector parameters are trained based on the following classification loss function; ; In the formula, P(x,y) represents the joint probability distribution of image sample x and label y; in the objective function, the joint distribution P(x,y) is used to calculate the expectation. This represents the model's average classification performance on the data distribution; P(x,y) is the expected weight distribution, which aims to maximize the conditional log-likelihood on the true distribution P(x,y), i.e., to optimize the model by maximizing this expectation. To more accurately predict the label y given x; The encoder, after passing through a quantization layer, is connected to the detector and trained to form the overall grasping and inference network. Step 4: For the newly arrived multi-source solid waste image, use the trained object detection network to determine whether there are recyclable targets; if there are recyclable targets, then use the trained grasping inference network to generate the corresponding grasping configuration.

2. The target grabbing point generation method based on semi-supervised generative inference network according to claim 1, characterized in that, The target detection network uses a binary classification neural network.

3. The target grabbing point generation method based on semi-supervised generative inference network according to claim 1, characterized in that, The use of datasets and The grasping and reasoning network is trained using a semi-supervised approach, including: Step a1: Initialize the parameter set of the crawling inference network. ;in, Representing the Categories Representing the Distribution models corresponding to each category The number of distribution models, and Representing the first The mean and variance of each distribution model; Step a2, update each distribution model and its mean: , ; In the formula, Indicates the first Number of image samples in each category; Indicates data without labels Under the condition that it belongs to the first categories The posterior probability; Step a3, maximize the following log-likelihood function: ; In the formula, Let be the likelihood function, representing the likelihood of a given set of model parameters. In the case of dataset and The joint probability of all image samples; Indicates labeled data The joint probability is given by the model parameter set. Decide; Indicates unlabeled data The marginal probability is obtained by summing over all possible categories, i.e. ; Step a4: Update the parameter set iteratively. Until convergence.

4. The target grabbing point generation method based on a semi-supervised generative inference network according to claim 1, characterized in that, Dataset Image samples RGB images; dataset and dataset Image samples Includes RGB color images and depth images.

5. The target grabbing point generation method based on a semi-supervised generative inference network according to claim 1, characterized in that, The crawling configuration includes: crawling center, crawling angle, crawling width, crawling quality score, and crawling target category.

Citation Information

Patent Citations

  • Target detection method and system

    WO2019223582A1