Semi-supervised anaphora grabbing detection method and system
Through the teacher-student network framework and dynamic pseudo-label generation, the problem of expensive and inefficient training data in robot reference crawling tasks is solved, and more efficient data utilization and crawling accuracy are achieved.
Patent Information
- Application Number
- CN202510172348.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The prior art requires a large amount of expensive and time-consuming labeled data in robot reference grab tasks, and the existing semi-supervised methods have problems of low training efficiency and insufficient perception and grabbing accuracy.
Generate from easy to difficult input perturbation learning process and dynamic threshold pseudo-labels to build a teacher-student network framework, and normalize structures through course enhancement and geometric consistency, optimize model parameters, and improve data validity and grabbing accuracy.
The data effectiveness and target perception accuracy of robot reference crawling tasks are improved, and the training efficiency is improved, especially the performance under the limited annotation data set.
Smart Images

Figure CN120347729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of robotic referential grasping, and particularly to a semi-supervised referential grasping detection method and system. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Referential grasping is the most important skill in robotic interaction operations, enabling the robot to understand language expressions, visually locate objects, and reason about grasping postures. In a cluttered home environment, language expressions can provide more precise object information to assist in positioning, such as geometric shape, color, and location. Different from classical vision-based grasping, language-conditioned grasping is a goal-oriented task that requires the robot to understand the conceptual semantics of natural language expressions and information from the vision-based scene, process multi-modal inputs, and reason about the unique goal desired by humans.
[0004] Recent research has utilized the cosine similarity between candidate region embeddings and language embeddings to retrieve targets. However, this method largely depends on the ability of the candidate extraction algorithm. There is also research that mainly focuses on the disambiguation of natural language and object bases but ignores the reasoning of grasping postures.
[0005] Benefiting from the progress of vision-language pre-training, contrastive language-image pre-training models, coarse-to-fine fusion modules, and large-scale datasets are used to jointly optimize the segmentation and grasping detection tasks, while completing object segmentation and grasping under end-to-end language conditions. To bridge the semantic gap between free-form language instructions and various visual objects, current methods collect large-scale image-text-pose pairs to train the model. However, obtaining the dataset is usually expensive and time-consuming because a large amount of manpower is required to annotate the positions of all objects, master the postures in the images, and match them with different language instructions respectively. It takes about five minutes to annotate an image of an object, and an effective dataset requires about tens of thousands of image-text-pose pairs.
[0006] In the context of the progress of data-driven learning, semi-supervised learning (SSL) provides a feasible solution to this challenge by training labeled and unlabeled data to improve data efficiency and model performance. Semi-supervised learning can combine a large amount of unlabeled data and a small amount of labeled data to improve the generalization ability of the model. Since unlabeled data in the real world is abundant and easy to collect, SSL has been widely explored in various application fields, such as semantic segmentation, vision-language reasoning, and grasping detection.
[0007] In some semi-supervised frameworks, the core of consistency regularization is that the model's predictions should be consistent under different perturbations of the input or model parameters. Therefore, most studies design data augmentation strategies and pseudo-label correction to improve prediction consistency. In particular, random sampling augmentation is used to reduce the distortion of the data distribution. However, the learning feature from easy to difficult is ignored, which can reduce the early distribution distortion. The MeanTeacher framework is used to filter out low-uncertainty pixels in the grasping position heatmap, but the consistency between the region and the visual position is not considered. The ST++ framework is used to alleviate incorrect pseudo-labels, generating pseudo-labels using a self-training pipeline and curriculum learning, and retraining reliable unlabeled samples. Attention-based consistency loss and adaptive pseudo-label weighting strategies are used for the degradation of pseudo-label noise in semi-supervised vision-language bases.
[0008] However, the above semi-supervised frameworks cannot be well directly transferred to the referential grasping task because: (1) Excessive distorted data perturbations lead to slow and unstable convergence of the multimodal input in the early stage, resulting in reduced training efficiency. (2) Existing methods do not consider the inconsistency between the perceived position and the grasping position, leading to uncoordinated hand-eye and decreased grasping accuracy.
[0009] In summary, for traditional referential grasping models, their training requires a large amount of expensive and time-consuming labeled training data, resulting in insufficient data validity; for existing semi-supervised methods, they use fixed strong augmentation and fixed thresholds to generate pseudo-labels, having problems of low training efficiency and insufficient perception and grasping accuracy. Summary of the Invention
[0010] To solve the above problems, the present invention proposes a semi-supervised referential grasping detection method and system, considering the learning process of input perturbation from easy to difficult and dynamic threshold pseudo-label generation, effectively improving the data validity of referential grasping and the accuracy of target perception and grasping.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] In the first aspect, the present invention provides a semi-supervised referential grasping detection method, including:
[0013] Obtaining labeled sample data and unlabeled sample data;
[0014] After augmenting the labeled sample data and unlabeled sample data using a preset strong augmentation method, obtaining strongly augmented labeled and unlabeled sample data, and after augmenting the unlabeled sample data using a preset weak augmentation method, obtaining weakly augmented sample data;
[0015] Construct a semi-supervised object detection model composed of a teacher network and a student network, and train the semi-supervised object detection model; where the training process includes:
[0016] Use the student network to predict the strongly augmented sample data, and construct a supervised loss function according to the prediction results of the student network; use the teacher network to predict the weakly augmented sample data, and obtain pseudo-labels after masking dynamic filtering and region consistency processing of the prediction results of the teacher network. Construct an unsupervised loss function according to the prediction results of the student network and the pseudo-labels;
[0017] Train and optimize the student network according to the supervised loss function and the unsupervised loss function, and update the teacher network according to the trained student network;
[0018] For the image-text pair to be tested, use the trained semi-supervised object detection model to obtain the detection results of the grasping target and the grasping pose.
[0019] As an alternative implementation, the process of obtaining pseudo-labels after masking dynamic filtering and region consistency processing of the prediction results of the teacher network includes: The prediction results of the teacher network include predicted segmentation masks Predicted grasping quality Predicted angle And predicted width Set the maximum threshold c max 、Minimum threshold c min And convergence threshold c base , Determine the filtering threshold c of positive samples h And the filtering threshold c of negative samples l Are:
[0020]
[0021] Filter the predicted segmentation mask to obtain the segmentation mask pseudo-label S p , And use element-wise product operation to obtain the pseudo-labels Q p ,Φ p ,
[0022]
[0023] Use min-max normalization to adjust the pseudo-label of the predicted grasping quality to the range of [0-1];
[0024]
[0025] Among them, τ is the number of training iterations; Γ is the total number of training steps, κ is the step size of the aging strategy, and m is the number of pixels; Is the value of the pseudo-label of the predicted grasping quality at pixel i; The value of the pseudo-label for predicting the grasping quality at pixel j.
[0026] As an alternative implementation, the supervised loss function is:
[0027]
[0028] where, is the strong augmentation function; is the training loss function; B is the number of strong augmented sample data; is the student network; is the image in the labeled samples; is the language description in the labeled samples; is the segmentation mask in the labeled samples; is the grasping map in the labeled samples.
[0029] As an alternative implementation, the unsupervised loss function is:
[0030]
[0031] where, is the weak augmentation function; is the training loss function; is the image in the unlabeled samples; is the language description in the unlabeled samples; is the teacher network; B is the number of strong augmented sample data; is the student network; is the strong augmentation function; φ(·) is the function for generating pseudo-labels.
[0032] As an alternative implementation, the training loss function is the sum of the segmentation error and the grasping error ;
[0033] The segmentation error is:
[0034] The grasping error is:
[0035] where, N is the number of pixels; is the value of the ground truth of the segmentation mask at pixel i in the labeled sample data; is the value of the grasping map M at pixel position i.
[0036] As an alternative implementation, the grasping pose G * =(x* , y * , θ * , w * , q * ) are:
[0037]
[0038] Among them, x * , y * , θ * , w * , q * are the optimal grasping abscissa value, ordinate value, angle, width, and confidence respectively; Q is the grasping quality map predicted by the referring grasping model; (x, y) are the abscissa and ordinate of the quality map; Φ is the grasping angle predicted by the referring grasping model; is the grasping width predicted by the referring grasping model.
[0039] In a second aspect, the present invention provides a semi-supervised referring grasping detection system, including:
[0040] An acquisition module, configured to acquire labeled sample data and unlabeled sample data;
[0041] An enhancement module, configured to enhance the labeled sample data and unlabeled sample data by using a preset strong enhancement method to obtain strongly enhanced sample data with labels and without labels, and enhance the unlabeled sample data by using a preset weak enhancement method to obtain weakly enhanced sample data;
[0042] A training module, configured to construct a semi-supervised object detection model composed of a teacher network and a student network, and train the semi-supervised object detection model; among them, the training process includes:
[0043] Predict the strongly enhanced sample data by using the student network, and construct a supervised loss function according to the prediction results of the student network; predict the weakly enhanced sample data by using the teacher network, and obtain pseudo-labels after masking dynamic filtering and region consistency processing of the prediction results of the teacher network, and construct an unsupervised loss function according to the prediction results of the student network and the pseudo-labels;
[0044] Train and optimize the student network according to the supervised loss function and the unsupervised loss function, and update the teacher network according to the trained student network;
[0045] A detection module, configured to use the trained semi-supervised object detection model for the image-text pair to be tested to obtain the detection results of the grasping target and the grasping posture.
[0046] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in the first aspect is completed.
[0047] In a fourth aspect, the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method described in the first aspect is completed.
[0048] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0050] For the robot's referential grasping task, the present invention proposes a semi-supervised referential grasping detection method and system based on curriculum enhancement and geometric consistency, which not only considers the learning process of input perturbation from easy to difficult, but also considers dynamic threshold pseudo-label generation and perception-grasp alignment, effectively improving the data effectiveness of referential grasping and the accuracy of target perception and grasping, and improving the training efficiency under a limited labeled data set.
[0051] The curriculum scheduling strategy proposed by the present invention aims to reduce the number of unreliable samples. The perturbation of easy-difficult inputs enables the student network to generate more effective positive and negative samples, promoting the convergence of weak-strong consistency learning and semi-supervised learning; at the same time, simple uniform sampling is used to gradually increase more augmented and difficult samples.
[0052] During the prediction training process, the present invention designs a geometric consistency normalization structure, combines pseudo-label and consistency regularization techniques to improve the training efficiency of unlabeled samples. To reduce the perturbation of the fixed reliability threshold in the initial stage to unreliable pseudo-masks, a dynamic feasibility allocation strategy is introduced to select high-quality positive and negative mask pixels. After selecting the true pixel samples through the dynamic feasibility allocation strategy, the pseudo-mask region is regarded as a prior for guiding the alignment of the grasping region. The inconsistent grasping confidence is filtered out by element-wise product operation, and finally the predicted grasping quality is adjusted to the range of [0-1] using min-max normalization. The pseudo-labels thus formed are used to optimize the model parameters of the student network, improving the training efficiency under a limited labeled data set.
[0053] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided drawings.
[0055] Figure 1 It is the overall architecture diagram of the semi-supervised anaphora capture detection method provided in Embodiment 1 of the present invention;
[0056] Figure 2 It is the schematic diagram of the data augmentation strategy based on the curriculum plan provided in Embodiment 1 of the present invention;
[0057] Figure 3 It is the schematic diagram of the dynamic feasibility adjustment based on geometric consistency provided in Embodiment 1 of the present invention;
[0058] Figure 4 It is the flowchart of training, deployment and inference provided in Embodiment 1 of the present invention. Detailed implementation manners
[0059] The following further describes the present invention in conjunction with the accompanying drawings and embodiments.
[0060] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0061] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0062] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0063] Term explanation:
[0064] Semi-supervised: A machine learning method that combines the characteristics of supervised learning and unsupervised learning, using a small amount of labeled data and a large amount of unlabeled data for training. Its goal is to utilize the information of unlabeled data to improve the performance of the model when labeled data is scarce or expensive.
[0065] Referential grasping: Given a scene image and a natural language description of a target, segment the pixels of a specific target object or region in the image according to the language description, and give the optimal grasping pose of the target.
[0066] Embodiment 1
[0067] This embodiment provides a semi-supervised referential grasping detection method, including:
[0068] Obtain labeled sample data and unlabeled sample data;
[0069] After enhancing the labeled sample data and unlabeled sample data using a preset strong enhancement method, obtain strongly enhanced labeled and unlabeled sample data. After enhancing the unlabeled sample data using a preset weak enhancement method, obtain weakly enhanced sample data;
[0070] Construct a semi-supervised object detection model composed of a teacher network and a student network, and train the semi-supervised object detection model; where the training process includes:
[0071] Use the student network to predict the strongly enhanced sample data, and construct a supervised loss function according to the prediction results of the student network; use the teacher network to predict the weakly enhanced sample data, and obtain pseudo-labels after mask dynamic filtering and region consistency processing of the prediction results of the teacher network. Construct an unsupervised loss function according to the prediction results of the student network and the pseudo-labels;
[0072] Train and optimize the student network according to the supervised loss function and the unsupervised loss function, and update the teacher network according to the trained student network;
[0073] For the image-text pair to be measured, use the trained semi-supervised object detection model to obtain the detection results of the grasping target and the grasping pose.
[0074] In this embodiment, first construct the SSRefGrasp semi-supervised training dataset to verify the performance of object segmentation and grasping detection under language conditions. SSRefGrasp matches the text mask pairs of OCID-Ref and the grasping poses of OCID-grasp in the ARID20 scene, and finally contains 34,713 image-text-mask-pose pairs, with an average text length of 7.9.
[0075] These datasets were collected in indoor cluttered scenarios, including three spatial situations: free, touch, and stacked. SSRefGrasp is divided into a training set, a validation set, and a test set. Here, the instances in the validation set are described with different language expressions to verify the robustness of language diversity. The test set contains different scenarios to verify the robustness of visual perception.
[0076] In this embodiment, a semi-supervised object detection model composed of a teacher network and a student network is constructed. Semi-supervised learning can extract cognitive knowledge from supervised and unsupervised behaviors. Therefore, this embodiment adopts a teacher-student knowledge distillation structure as the semi-supervised framework, as Figure 1 shown. The student network and the teacher network adopt the same referential grasping model, parameterized by θ S and θ T respectively. Generally, the referential grasping model consists of an image-text encoder and a multimodal decoder.
[0077] This embodiment uses a burn-in training strategy to supervise the training of labeled data to initialize the parameters of the teacher network and provide stable learning guidance for the student network.
[0078] In the semi-supervised stage, the teacher network is frozen and used to generate pseudo-labels; the student network is trained with labeled data and unlabeled data; gradient backpropagation is performed in the student network and stopped in the teacher network.
[0079] The parameters of the teacher network are updated by an exponential moving average strategy; that is:
[0080]
[0081] where γ represents the number of training iterations; θ represents the parameters of the referential grasping model, T represents the teacher network; and S represents the student network.
[0082] The update metric α for knowledge distillation from the student network to the teacher network is calculated by a cosine function:
[0083]
[0084] where τ is the number of training iterations; α0 and α1 represent the starting value and the ending value, which are set to 0.5 and 0.9996 by default, respectively; κ represents the last step of the burn-in training strategy; and Γ is the total number of training iterations.
[0085] In this embodiment, a curriculum-based weak-to-strong data augmentation strategy is designed. Weak-strong data augmentation is an input perturbation strategy widely used in semi-supervised learning to improve consistency regularization. This embodiment proposes a curriculum-based weak-strong data augmentation strategy as shown in Figure 2 shown, and designs K width-increasing operations Among them, it includes eight data augmentation operations: random cropping, saturation adjustment, brightness adjustment, contrast adjustment, random rotation, pepper and salt, synonym replacement, and horizontal flipping.
[0086] For the weak augmentation method, a standard pixel cropping operation is selected as the weak augmentation method. Specifically, the image, segmentation, and grasping pose map are randomly cropped within the range of [0, 50] pixels.
[0087] For the strong augmentation method, it includes color jittering, rotation, and pepper and salt noise, and at the same time, synonym replacement and horizontal flipping operations are introduced; among them, small-amplitude random rotation is used to prevent the spatial position from being destroyed; random cropping, color jittering, random rotation, and pepper and salt are used to enhance the robustness to visual interference; synonym replacement is used to improve the understanding of language expressions; the horizontal flipping operation is used to enhance the correlation between the visual position and the position words.
[0088] The fixed augmentation method is widely used in semi-supervised learning to perform all However, the strong augmentation with excessive distortion in the early stage will seriously damage the distribution consistency and generate more unstable and unreliable samples. Therefore, in this embodiment, the student network is allowed to learn the consistency of input perturbations from easy to difficult. As Figure 2 shown, the curriculum scheduling strategy of this embodiment aims to reduce the number of unreliable samples. The perturbations of easy-to-difficult inputs enable the student network to generate more effective positive and negative samples, promoting the convergence of weak-strong consistency learning and semi-supervised learning.
[0089] In this embodiment, during the process of selecting curriculum augmentation operations, simple uniform sampling is adopted to gradually increase more augmentations and difficult samples;
[0090]
[0091] where τ and Γ represent the current step size and the total training step size; κ represents the steps of the aging strategy; K is the number of addition operations; is the floor operation; is the selected strong augmentation operation.
[0092] Based on the semi-supervised object detection model constructed above and the weak-to-strong data augmentation strategy based on the curriculum plan, the obtained labeled sample data and unlabeled sample data are augmented using the strong augmentation method, and the unlabeled sample data is augmented using the weak augmentation method, thereby obtaining labeled and unlabeled strongly augmented sample data, as well as weakly augmented sample data. Further, the student network is predicted and trained based on the strongly augmented sample data, and the teacher network is predicted and trained based on the weakly augmented sample data.
[0093] During the prediction training process, in this embodiment, a geometric consistency normalization structure is designed, combined with pseudo-label and consistency regularization techniques to improve the training efficiency of unlabeled samples, as Figure 3 shown. In the teacher network-student network framework, the teacher network uses weakly augmented samples to predict the segmentation mask, grasping quality, angle, and width, and obtains pseudo-labels after mask dynamic filtering and regional consistency processing of the prediction results; in the weak-to-strong augmentation, the prediction results of the student network have strict geometric consistency with the pseudo-labels of the teacher network. Therefore, the quality of the pseudo-labels determines the effect of teacher-student distillation.
[0094] Specifically:
[0095] To reduce the perturbation of the fixed reliability threshold on unreliable pseudo-masks in the initial stage, this embodiment introduces a dynamic viability assignment strategy (DRA) to select high-quality positive and negative mask pixels.
[0096] First, set the maximum threshold c max , minimum threshold c min and convergence threshold c base ; in the early stage of training, the filtering threshold c h of positive samples and the filtering threshold c l of negative samples are respectively close to the maximum thresholds c max and c min , which can make the pseudo-labels contain as many true positive and true negative samples as possible; as the teacher network obtains more stable knowledge, c h and c l gradually approach c base , increasing the positive and negative samples until c h =c l =c base .
[0097] The dynamic viability assignment process is expressed as:
[0098]
[0099] where τ is the number of training iterations.
[0100] Secondly, through the dynamic viability assignment strategy, filter the predicted segmentation mask, select the true pixel samples, and obtain the segmentation mask pseudo-label; regard the pseudo-mask area as the prior for guiding the alignment of the grasping area, and use the element-wise product operation to filter out inconsistent grasping confidences, thereby obtaining the pseudo-labels of the predicted grasping quality, predicted angle, and predicted width; defined as:
[0101]
[0102] where S p is the predicted segmentation mask pseudo-label; The segmentation mask predicted by the teacher network; Q p , Φ p , are the pseudo-labels for the predicted grasping quality, predicted angle, and predicted width, respectively; are the grasping quality, angle, and width predicted by the teacher network, respectively.
[0103] Finally, the pseudo-label of the predicted grasping quality is adjusted to the range of [0-1] using min-max normalization;
[0104]
[0105] where m is the number of pixels; is the value of the grasping quality pseudo-label at pixel position i; is the value of the grasping quality pseudo-label at pixel position j; i and j represent pixel positions; Q p is the grasping quality pseudo-label.
[0106] In summary, the pseudo-labels composed of S p and are used to optimize the model parameters θ of the student network S .
[0107] In this embodiment, the specific training process of the above semi-supervised framework is as follows.
[0108] As Figure 4 shown, professional annotators only need to provide a small number of image-text-mask pose pairs, non-experts need to collect a large number of image-text pairs, use the teacher-student knowledge distillation structure as a semi-supervised learning framework, and at the same time design an enhancement strategy based on curriculum scheduling and geometric consistency regularization to implement network training, and finally deploy the trained model to the robot.
[0109] Set the grasping pose G and segmentation mask S of the supervision target; given the segmentation mask S GT , the segmentation error is calculated using the binary cross-entropy loss
[0110]
[0111] where N is the number of pixels; is the value of the segmentation mask ground truth at pixel position i in the labeled data.
[0112] Given the true grasping pose the grasping error is calculated using the SmoothL1 loss
[0113]
[0114] where the grasping map To grab the value of Figure M at pixel position i.
[0115] Thus, the training loss function representing the grasping model is obtained:
[0116] In this embodiment, the student network predicts strongly augmented sample data, and a supervised loss function is constructed based on the prediction results of the student network (including the segmentation mask, grasping quality, angle, and width)
[0117] where, is the strong augmentation function; B represents the number of strongly augmented sample data; is the student network; is the image in the labeled sample; is the language description in the labeled sample; is the segmentation mask in the labeled sample; is the grasping figure in the labeled sample.
[0118] The teacher network predicts weakly augmented sample data in sequence, and at the same time uses the reliability function φ(·) to generate pseudo-labels; an unsupervised loss function is constructed based on the prediction results of the student network and the pseudo-labels
[0119]
[0120] where, is the weak augmentation function; is the image in the unlabeled sample; is the language description in the unlabeled sample; is the teacher network.
[0121] Thus, the overall optimization objective is:
[0122]
[0123] where λ is the hyperparameter weight of the unsupervised loss function.
[0124] In this embodiment, during the inference process, the segmentation mask S is calculated, and a threshold λ is set s , to obtain the predicted mask S * ; then the optimal grasping pose g is obtained through the following formula * =(x * ,y * ,θ * ,w * ,q * );
[0125]
[0126] Among them, x * , y * , θ * , w * , q * are respectively the optimal grasping abscissa value, ordinate value, angle, width, and confidence; Q is the grasping quality map predicted by the referring grasping model; (x, y) are the abscissa and ordinate of the quality map; Φ is the grasping angle predicted by the referring grasping model; is the grasping width predicted by the referring grasping model.
[0127] To perform object operations, this embodiment uses the hand-eye transformation matrix T rc and the image-camera transformation matrix T ci to calculate the operation pose g r = T rc (T ci (g * )) in the robot coordinate system, where g * is the optimal grasping pose in the image coordinate system, and g r is the grasping pose in the robot coordinate system; finally, the robot calculates the joint angles through inverse kinematics and drives the arm to grasp the target.
[0128] Embodiment 2
[0129] This embodiment provides a semi-supervised referring grasping detection system, including:
[0130] An acquisition module, configured to acquire labeled sample data and unlabeled sample data;
[0131] An enhancement module, configured to enhance the labeled sample data and unlabeled sample data by using a preset strong enhancement method to obtain strongly enhanced labeled and unlabeled sample data, and enhance the unlabeled sample data by using a preset weak enhancement method to obtain weakly enhanced sample data;
[0132] A training module, configured to construct a semi-supervised object detection model composed of a teacher network and a student network, and train the semi-supervised object detection model; among them, the training process includes:
[0133] Use the student network to predict the strongly enhanced sample data, and construct a supervised loss function according to the prediction results of the student network; use the teacher network to predict the weakly enhanced sample data, and obtain pseudo-labels after mask dynamic filtering and region consistency processing of the prediction results of the teacher network, and construct an unsupervised loss function according to the prediction results of the student network and the pseudo-labels;
[0134] Train and optimize the student network according to the supervised loss function and the unsupervised loss function, and update the teacher network according to the trained student network;
[0135] The detection module is configured to use the trained semi - supervised object detection model for the image - text pair to be tested, and obtain the detection results of the grasping target and the grasping pose.
[0136] As an alternative implementation, the process of obtaining the pseudo - labels after mask dynamic filtering and region consistency processing of the teacher network prediction results includes: The teacher network prediction results include predicted segmentation masks Predicting the grasping quality Predicting the angle And the predicted width Set the maximum threshold c max , the minimum threshold c min And the convergence threshold c base , and determine the filtering threshold c of the positive samples h And the filtering threshold c of the negative samples l As:
[0137]
[0138] Filter the predicted segmentation mask to obtain the segmentation mask pseudo - label S p , and use the element - wise product operation to obtain the pseudo - labels Q p , Φ p ,
[0139]
[0140] Use min - max normalization to adjust the pseudo - label of the predicted grasping quality to the range of [0 - 1];
[0141]
[0142] Where τ is the number of training iterations; Γ is the total number of training steps, κ is the step size of the aging strategy, m is the number of pixels; Is the value of the pseudo - label of the predicted grasping quality at pixel i; Is the value of the pseudo - label of the predicted grasping quality at pixel j.
[0143] As an alternative implementation, the supervised loss function Is:
[0144]
[0145] Where Is the strong augmentation function; Is the training loss function; B is the number of strong augmentation sample data; Is the student network; Is the image in the labeled samples; is the language description in the labeled sample; is the segmentation mask in the labeled sample; is the grasping image in the labeled sample.
[0146] As an alternative embodiment, the unsupervised loss function is:
[0147]
[0148] wherein, is the weak augmentation function; is the training loss function; is the image in the unlabeled sample; is the language description in the unlabeled sample; is the teacher network; B is the number of strongly augmented sample data; is the student network; is the strong augmentation function; φ(·) is the function for generating pseudo labels.
[0149] As an alternative embodiment, the training loss function is the sum of the segmentation error and the grasping error ;
[0150] The segmentation error is:
[0151] The grasping error is:
[0152] wherein, N is the number of pixels; is the value of the segmentation mask ground truth at pixel i in the labeled sample data; is the value of the grasping image M at pixel position i.
[0153] As an alternative embodiment, the grasping pose G * =(x * , y * , θ * , w * , q * ) is:
[0154]
[0155] wherein, x * , y * , θ * , w * , q *They are respectively the optimal grasping abscissa value, ordinate value, angle, width, and confidence level; Q represents the grasping quality map predicted by the grasping model; (x, y) are the abscissa and ordinate of the quality map; Φ represents the grasping angle predicted by the grasping model; represents the grasping width predicted by the grasping model.
[0156] It should be noted here that the above modules correspond to the steps described in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0157] In more embodiments, there is also provided:
[0158] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in Embodiment 1 is completed. For the sake of brevity, it will not be elaborated here.
[0159] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0160] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random memory. For example, the memory may also store information about the device type.
[0161] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.
[0162] The method in Embodiment 1 can be directly embodied as being executed by a hardware processor, or by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0163] A computer program product includes a computer program. When the computer program is executed by the processor, the method described in Embodiment 1 is implemented and completed.
[0164] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which are executed in a device on a target real or virtual processor to perform the processes / methods described above. Generally, program modules include routines, programs, libraries, objects, classes, components, data structures, etc. that perform specific tasks or implement specific abstract data types. In various embodiments, the functions of program modules can be combined or divided as needed. The machine-executable instructions for program modules can be executed within local or distributed devices. In a distributed device, program modules can be located in local and remote storage media.
[0165] The computer program code for implementing the method of the present invention can be written in one or more programming languages. This computer program code can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program code is executed by the computer or other programmable data processing device, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the computer, partially on the computer, as a stand-alone software package, partially on the computer and partially on a remote computer, or entirely on a remote computer or server.
[0166] In the context of the present invention, the computer program code or related data can be carried by any suitable carrier so that a device, apparatus, or processor can perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, etc. Examples of signals can include electrical, optical, radio, sound, or other forms of propagated signals, such as carrier waves, infrared signals, etc.
[0167] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with this embodiment can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0168] Although the specific implementation manners of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Those skilled in the art should understand that based on the technical solution of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A semi-supervised anaphora grabbing detection method, characterized in that, Including: Obtain labeled sample data and unlabeled sample data; After enhancing the labeled sample data and the unlabeled sample data using a preset strong enhancement method, obtain strongly enhanced labeled and unlabeled sample data. After enhancing the unlabeled sample data using a preset weak enhancement method, obtain weakly enhanced sample data; Construct a semi-supervised object detection model composed of a teacher network and a student network, and train the semi-supervised object detection model; among them, the training process includes: Use the student network to predict the strongly enhanced sample data, and construct a supervised loss function according to the prediction results of the student network; use the teacher network to predict the weakly enhanced sample data, and obtain pseudo-labels after mask dynamic filtering and region consistency processing of the prediction results of the teacher network. Construct an unsupervised loss function according to the prediction results of the student network and the pseudo-labels; Train and optimize the student network according to the supervised loss function and the unsupervised loss function, and update the teacher network according to the trained student network; For the image-text pair to be measured, use the trained semi-supervised object detection model to obtain the detection results of the grasping target and the grasping posture.
2. The semi-supervised anaphoric grab detection method according to claim 1, characterized in that, The process of obtaining pseudo-labels after mask dynamic filtering and region consistency processing of the prediction results of the teacher network includes: The teacher network prediction results include predicted segmentation masks predicted grasping quality predicted angle and predicted width Set the maximum threshold c max , the minimum threshold c min and the convergence threshold c base , and determine the filtering threshold c h for positive samples and the filtering threshold c l for negative samples as follows: Filter the predicted segmentation mask to obtain the segmentation mask pseudo-label S p , and use the element-wise product operation to obtain the pseudo-labels Q, Φ of the predicted grasping quality, predicted angle, and predicted width p , Φ p , Use min-max normalization to adjust the pseudo-labels of the predicted grasping quality to the range of [0-1]; where τ is the number of training iterations; Γ is the total training step length, κ is the step length of the aging strategy, and m is the number of pixels; is the value of the pseudo-label for predicting the grasping quality at pixel i; is the value of the pseudo-label for predicting the grasping quality at pixel j.
3. The semi-supervised anaphora capture detection method according to claim 1, characterized in that, The supervised loss function is as follows: Among them, is the strong enhancement function; is the training loss function; B is the number of strongly enhanced sample data; is the student network; is the image in the labeled samples; is the language description in the labeled samples; is the segmentation mask in the labeled samples; is the grasping map in the labeled samples.
4. The semi-supervised anaphoric grab detection method according to claim 1, wherein The unsupervised loss function is as follows: Among them, is a weak augmentation function; is the training loss function; is the image in the unlabeled samples; is the language description in the unlabeled samples; is the teacher network; B is the number of strongly augmented sample data; is the student network; is a strong augmentation function; φ(·) is the function used to generate pseudo-labels.
5. A semi-supervised anaphoric grab detection method according to claim 3 or 4, characterized in that The training loss function is the sum of the segmentation error and the grasping error ; Segmentation error is Grasping error is as follows: where N is the number of pixels; is the value of the segmentation mask ground truth at pixel i in the labeled sample data; is the value of the grasping map M at pixel position i.
6. The semi-supervised anaphora capture detection method according to claim 1, wherein Grasping posture G * =(x * , y * , θ * , w * , q * ) is as follows: Among them, x * , y * , θ * , w * , q * are respectively the optimal grasping abscissa value, ordinate value, angle, width and confidence; Q is the grasping quality map predicted by the grasping model; (x, y) are the abscissa and ordinate of the quality map; Φ is the grasping angle predicted by the grasping model; is the grasping width predicted by the grasping model.
7. A semi-supervised anaphoric grab detection system, characterized in that, Including: An acquisition module configured to obtain labeled sample data and unlabeled sample data; An enhancement module configured to, after enhancing the labeled sample data and the unlabeled sample data using a preset strong enhancement method, obtain strongly enhanced labeled and unlabeled sample data, and after enhancing the unlabeled sample data using a preset weak enhancement method, obtain weakly enhanced sample data; A training module configured to construct a semi-supervised object detection model composed of a teacher network and a student network, and train the semi-supervised object detection model; among them, the training process includes: Use the student network to predict the strongly enhanced sample data, and construct a supervised loss function according to the prediction results of the student network; use the teacher network to predict the weakly enhanced sample data, and obtain pseudo-labels after mask dynamic filtering and region consistency processing of the prediction results of the teacher network. Construct an unsupervised loss function according to the prediction results of the student network and the pseudo-labels; Train and optimize the student network according to the supervised loss function and the unsupervised loss function, and update the teacher network according to the trained student network; A detection module configured to, for the image-text pair to be measured, use the trained semi-supervised object detection model to obtain the detection results of the grasping target and the grasping posture.
8. An electronic device, characterized in that, Including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the method described in any one of claims 1-6 is completed.
9. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the method described in any one of claims 1-6 is completed.
10. A computer program product, characterized in that, Including a computer program, when the computer program is executed by the processor, the method described in any one of claims 1-6 is implemented.
Citation Information
Patent Citations
End-to-end semi-supervised target detection method and device for image and readable medium
CN116630745A
Semi-supervised medical image segmentation method based on hybrid-decoupling training
CN116935054A
Liver tumor image segmentation method based on bounding box weak supervision
CN117788825A
Semi-supervised target detection method and device, computer equipment and storage medium
CN117934893A
Uncertainty perception semi-supervised pancreas segmentation method based on evidence learning
CN118674927A
Cited By
Weak supervision accessibility positioning robot grabbing method based on self-adaption during testing
CN121733548A
Robot grasping method based on test-time adaptive weakly supervised affordance localization
CN121733548B