Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

65 results about "3D pose estimation" patented technology

3D pose estimation is the problem of determining the transformation of an object in a 2D image which gives the 3D object. One of the requirements of 3D pose estimation arises from the limitations of feature-based pose estimation. There exist environments where it is difficult to extract corners or edges from an image. To circumvent these issues, the object is dealt with as a whole in noted techniques through the use of free-form contours.

System for estimating a three dimensional pose of one or more persons in a scene

A system for estimating a three dimensional pose of one or more persons in a scene is disclosed herein. The system includes one or more cameras and a data processor configured to execute computer executable instructions. The computer executable instructions include: (i) receiving the one or more images of the scene from the one or more cameras; (ii) extracting features from the one or more images of the scene for providing inputs to a three dimensional pose estimation neural network; (iii) generating, by using the three dimensional pose estimation neural network, vertices of a canonical human mesh model for the one or more images of the scene; (iv) retrieving a weight matrix from an annotation server that corresponds to a desired keypoint set; and (v) generating the desired keypoint set by multiplying the retrieved weight matrix with the vertices of the canonical human mesh model.
Owner:BERTEC CORP

Three-dimensional video highlight from a camera source

According to an aspect, a method includes generating a three-dimensional (3D) video segment from two-dimensional (2D) video content captured by a camera system, including obtaining, from a 3D pose estimation engine, 3D movement data of an object detected in the 2D video content, and generating an animated object based on the 3D movement data such that a movement of the animated object corresponds to a movement of the object in the 2D video content. The method includes generating 3D video content from the 3D video segment and transmitting the 3D video content to a user device for display.
Owner:GOOGLE LLC

Dynamic Gaussian digital human image rendering method and device, equipment and storage medium

The invention discloses a dynamic Gaussian digital human image rendering method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: determining a target multi-view image frame from a multi-view RGB video image sequence according to a preset attitude, and constructing a deformable parameterized model according to the target multi-view image frame; constructing an attitude space driving attitude corresponding to the multi-view RGB video image sequence based on a three-dimensional attitude estimation technology, and generating a two-dimensional position map according to the attitude space driving attitude and the deformable parameterized model; performing three-dimensional Gaussian binding on the deformable parameterized model to obtain a local attribute of the three-dimensional Gaussian; training is carried out according to the two-dimensional position map, and a target StyleUNet neural network is obtained; and predicting the new attitude through the target StyleUNet neural network to obtain a dynamic Gaussian digital human image. In this way, the high-fidelity drivable high-frequency detail digital human image can be automatically rendered.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Model training method and apparatus, three-dimensional pose estimation method and apparatus, medium, and electronic device

Provided are a model training method and apparatus, a three-dimensional pose estimation method and apparatus, and an electronic device, relating to the field of computer vision. The method comprises: acquiring at least one training sample, and inputting each training sample into an initial three-dimensional pose estimation model for training to obtain a trained three-dimensional pose estimation model. Each round of training comprises: inputting a sample image of the training sample into a current three-dimensional pose estimation model to obtain the three-dimensional Gaussian mixture representation of each key point on a preset three-dimensional coordinate system; for each key point, obtaining the two-dimensional Gaussian mixture representations of the key point on three mutually perpendicular two-dimensional planes in the preset three-dimensional coordinate system; for each key point, determining the loss value of the key point on each two-dimensional plane; and adjusting model parameters of the current three-dimensional pose estimation model on the basis of the loss value of each key point on each two-dimensional plane. The embodiments of the present application effectively improve the accuracy of three-dimensional pose estimation.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Three-dimensional attitude estimation method and device, storage medium and electronic equipment

The invention discloses a three-dimensional attitude estimation method and device, a storage medium and electronic equipment, and relates to the technical field of three-dimensional attitude estimation, and the method comprises the steps: carrying out the joint estimation of a multi-view image of a target object, and obtaining a joint feature map and a joint heat map of each view; performing cross-view-angle alignment analysis on the multi-view-angle joint heat map to obtain an initial three-dimensional joint coordinate of each joint point; carrying out deformable cross-attention feature aggregation on the multi-view joint feature map to obtain multi-view fusion features of the joints; and performing three-dimensional attitude prediction based on the initial three-dimensional joint coordinates of the joints and the multi-view fusion features to obtain target three-dimensional joint coordinates of the joints. According to the invention, the estimation accuracy of the three-dimensional coordinates of the joint points during three-dimensional attitude estimation can be effectively improved.
Owner:SHENZHEN TCL NEW-TECH CO LTD

Vision-based three-dimensional human pose estimation system and method for ergonomic risk assessment

ActiveUS12511929B1Image enhancementImage analysisErgonomic riskVision based
Disclosed herein are vision-based three-dimensional (3D) pose estimation system and method for ergonomic risk assessment. An example system may comprise a computing device configured to obtain a monocular video capturing motions of a subject performing at least one working activity for a selected duration of time, perform a whole-body two dimensional (2D) pose estimation based at least on extracted frames of the monocular video, perform a whole-body 3D pose estimation based at least on the whole-body 2D pose estimation, calculate joint angles based at least on the whole-body 3D pose estimation, determine a posture score for each identified joint in each frame of the monocular video, and determine an ergonomic risk level of each identified joint based at least upon the posture score.
Owner:VELOCITYEHS HOLDINGS INC

3D pose estimation based on color and depth features

An electronic device and method for 3D pose estimation based on color and depth features is provided. The electronic device acquires color image data that includes an object and depth image data associated with the object. The electronic device generates a set of color feature maps based on the color image data and a set of depth feature maps based on the depth image data. The electronic device generates a set of feature fusion maps based on the set of color feature maps and the set of depth feature maps. The electronic device further generates 3D feature data of the object based on projection of each feature included in each feature fusion map to a 3D point of a 3D volume. The electronic device computes a 3D pose of the object based on application of a pose-estimation neural network on the 3D feature data and controls rendering of the 3D pose.
Owner:SONY GROUP CORP

Method and system for sequence-aware estimation of ultrasound probe pose in laparoscopic ultrasound procedures

The present teaching relates to estimating 3D pose of an ultrasound probe. An ultrasound image is acquired by an ultrasound probe deployed at a three-dimensional (3D) probe pose during a medical procedure directed to a target organ. A mask for each two-dimensional (2D) anatomical structure is identified and a corresponding label is estimated from the ultrasound image to generate an ASM / label pair. If a sequence of prior 3D probe poses does not exist, an ASM / label representation for the ultrasound image is generated based on the ASM / label pairs from the ultrasound image and used to estimate the 3D pose of the probe via an ASM-pose mapping model. If the sequence of prior 3D probe poses exists, the 3D probe pose is predicted based on the sequence of prior 3D probe poses and an ASM / label representation is generated based on a virtual ultrasound image created based on the predicted 3D probe pose.
Owner:EDDA TECHNOLOGY INC

Training method and estimation method of single-view multi-person three-dimensional attitude estimation model

The invention discloses a training method and an estimation method of a single-view multi-person three-dimensional attitude estimation model, and the training method of the single-view multi-person three-dimensional attitude estimation model comprises the steps: collecting multi-person attitude pictures; processing the multi-person posture picture by adopting structured posture representation to obtain a joint point set corresponding to the multi-person posture picture; generating a mask value according to a preset mask strategy, and performing mask processing on the joint point set by using the mask value to obtain a mask attitude tensor; and pre-training the backbone network for three-dimensional attitude estimation by using the mask attitude tensor to obtain a trained backbone network. According to the method, attitude mask training is introduced, so that the model can learn and adapt to various possible shielding conditions, and the robustness and estimation precision of the model in a complex motion environment are improved.
Owner:CENT SOUTH UNIV

3D human body posture estimation method based on double-flow space-time diagram network

The invention discloses a 3D human body posture estimation method based on a double-flow space-time diagram network. Collecting a plurality of human body posture maps and marking three-dimensional coordinates of key points, and further constructing a training set of three-dimensional human body posture estimation; establishing a three-dimensional human body posture detection network, and inputting the training set into the three-dimensional human body posture detection network for training to obtain a trained three-dimensional human body posture detection network; and obtaining a to-be-detected human body posture image, and inputting the to-be-detected human body posture image into the trained three-dimensional human body posture detection network for posture detection to obtain a three-dimensional posture estimation result of the human body. The method has a time form and a space form, the features of Transform and GCN can be aggregated in a simple and effective mode in the space-time dimension, and comprehensive modeling of local and global features of the human skeleton is achieved.
Owner:ZHEJIANG UNIV OF SCI & TECH

A method and system for estimating three-dimensional human posture with motion prompts

The present invention provides a method and system for three-dimensional human pose estimation with action prompts, comprising: extracting pose position features, pose sequence features, and action features from a two-dimensional pose sequence using a pose encoder; obtaining text prompt features based on the pose position features; aligning the text prompt features with the action features to extract action category information; selecting pose prompt features corresponding to the action category information, combining them based on the correlation between the pose prompt features and the pose sequence features to obtain enhanced pose sequence features; and linearly mapping the enhanced pose sequence features to obtain a three-dimensional pose estimate. During the pose estimation process, the present invention mines prior information related to the action, introduces multimodal information of action-related text and pose features, and addresses depth ambiguity. Its plug-and-play module compactly extracts input data features, saving network model parameters. It significantly improves the accuracy of pose estimation for self-occlusion and complex actions, enhancing flexibility and scalability.
Owner:SHANGHAI JIAOTONG UNIV

Joint modeling method, application method and device of gesture recognition and pose estimation

The application discloses a gesture recognition and posture estimation joint modeling method, application method and device, relates to the technical field of perception interaction, and the modeling method comprises the following steps: performing two-dimensional posture estimation on a historical gesture image region to obtain a two-dimensional key point sequence, and inputting the two-dimensional key point sequence and an initial gesture type label into a three-dimensional posture estimation network to obtain an initial three-dimensional key point sequence; based on a video frame sequence, a first predicted gesture type label is acquired; based on the initial three-dimensional key point sequence, a second predicted gesture type label is acquired; the two predicted gesture type labels are fused to obtain a next round gesture type label; the next round gesture type label is taken as the initial gesture type label, the initial three-dimensional key point sequence is taken as the two-dimensional key point sequence, and the initial three-dimensional key point sequence acquisition step is returned; when a set iteration number is reached or a convergence condition is met, a joint modeling model is obtained. The application can realize more efficient and more accurate gesture recognition and three-dimensional posture estimation.
Owner:CHINA MARITIME POLICE ACADEMY

Multi-view real-time human limb driving method and device and electronic equipment

The invention relates to a multi-view real-time human limb driving method and device and electronic equipment. The method comprises the following steps: acquiring a human body action multi-view image; performing two-dimensional human body posture estimation on the multi-view-angle image to obtain thermodynamic diagram coordinates and feature diagrams of two-dimensional human body key points of each view angle; obtaining the view angle weight of the view angle image where each key point is located by adopting an attention mechanism according to the feature map of each view angle; according to thermodynamic diagram coordinates and view angle weights of the two-dimensional human body key points, triangularizing the two-dimensional human body key points in each view angle image by using a deep learning model to obtain three-dimensional human body key point space coordinates; and driving the virtual human based on the three-dimensional human body key point space coordinates. The weight of each view angle can be dynamically optimized, view angle noise is suppressed, the accuracy of three-dimensional attitude estimation is effectively improved, and a real-time driving effect can be achieved.
Owner:MIGU CO LTD +1

A 3D Human Pose Estimation Method, System, Electronic Device and Storage Medium

This application belongs to the field of computer vision and discloses a 3D human pose estimation method, system, electronic device and storage medium. The method includes: obtaining a 2D key point sequence, generating high-dimensional features for each frame through pose space embedding, and applying a random mask and position encoding to form a sparse temporal Token sequence; using a pose temporal interaction module based on Transformer to extract global temporal features; at the same time, adopting joint-level spatial embedding, a spatial interaction module and hierarchical convolution for the current frame and its neighboring frames to capture local temporal details and obtain local features; after comprehensively integrating global and local features through an adaptive fusion strategy, inputting them into a pose optimization network (a contrastive learning module for 2D pose, global 3D pose and central frame 3D pose) to generate a 3D human pose, then calculating the overall loss and using a sparse-dense training strategy for optimization, and finally outputting the 3D human pose. This method can significantly improve the accuracy of 3D pose estimation.
Owner:NANCHANG UNIV

Method, apparatus, and system for three-dimensional pose estimation

The latest progress in neural networks has shown significant advancements in human pose estimation tasks. Pose estimation can be divided into monocular 2D pose estimation, multi-view 3D pose estimation, and single-view 3D pose estimation. Among them, 3D pose has recently received increasing attention for applications in AR / VR, gaming, and human-computer interaction. However, current academic benchmarks for human 3D pose estimation only consider the performance of its relative pose. Root localization over time, in other words, the "trajectory" of the entire body in 3D space, is not well considered. Applications such as motion capture require not only the accurate relative pose of the body but also the root position of the entire body in 3D space. Therefore, this paper describes an effective monocular full 3D pose recovery model from 2D pose inputs that can be applied to the above applications.
Owner:SONY GROUP CORP

A multi-person 3D pose estimation method and system based on explicit limb representation

The present application belongs to the field of 3D human posture estimation, and particularly relates to a multi-person 3D posture estimation method and system based on explicit limb representation, comprising: acquiring a multi-person scene image, inputting the multi-person scene image into a trained multi-person 3D posture estimation model to obtain a key point heat map, a limb orientation vector field, a limb relative depth map and a root key point absolute depth map, and post-processing the key point heat map, the limb orientation vector field, the limb relative depth map and the root key point absolute depth map to obtain an estimated multi-person 3D posture; the present application encodes the limb information of a person into a two-dimensional limb orientation vector field and a limb relative depth map to ensure that the matching between candidate key points can directly utilize the most intuitive limb information, so as to improve the phenomenon of incorrect matching of candidate key points of different human bodies and improve the quality of the matched posture.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

A method and system for single-view animal 3D pose estimation

The application discloses a kind of single-view animal 3D posture estimation method and system, comprising: backbone network receives 2D image data, the backbone network is trained, obtains trained backbone network;The output of backbone network is connected with 3D posture estimation network, constructs weakly supervised learning module, uses 3D labeled data to train weakly supervised learning module, forms initial model, uses initial model to predict 2D unlabeled data, obtains pseudo-labeled data set, and real 3D labeled data is merged to form training set, and weakly supervised learning module is trained;Real-time acquisition single-view animal image is input into trained weakly supervised learning module and carries out single-view animal 3D posture estimation;The application has the advantages that based on a large number of 2D unlabeled data, a small amount of 3D labeled data, under the premise of small dependence on labeled data, simultaneously not depending on artificial experience, effectively improve the performance upper limit on 3D posture estimation accuracy.
Owner:JIANGXI ACAD OF FORESTRY

A method, apparatus and equipment for multi-view, multi-person 3D pose estimation

ActiveCN116403243BRealize differentiated matchingAccurately definedBiometric pattern recognitionThree-dimensional object recognitionPattern recognitionHuman body
This invention relates to the field of pose estimation technology, specifically to a multi-view, multi-person 3D pose estimation method, apparatus, and device. The method includes: acquiring image information and depth information of the target using multi-view acquisition technology; segmenting the person in a single view based on the image information and extracting 2D semantic features; reconstructing the point cloud based on the depth information to obtain a 3D point cloud in a preset world coordinate system; segmenting the point cloud for different human figures based on the 3D point cloud and calculating the point cloud boundaries as human body boundaries; mapping the 2D semantic features to the 3D point cloud to obtain 3D features; and classifying different human figures and constructing 3D poses based on the 3D features and human body boundaries. This technical solution can segment the point cloud for different human figures based on the 3D point cloud, accurately define the joint positions of different figures, and achieve differentiation and matching of different figures. Furthermore, based on discrete point cloud data, it has better flexibility and fault tolerance.
Owner:BEIJING INSTITUTE OF PETROCHEMICAL TECHNOLOGY

Incremental 2D-to-3D pose lifting for fast and accurate human pose estimation

Techniques related to 3D pose estimation from a 2D input image are discussed. Such techniques include incrementally adjusting an initial 3D pose generated by applying a lifting network to a detected 2D pose in the 2D input image by projecting each current 3D pose estimate to a 2D pose projection, applying a residual regressor to features based on the 2D pose projection and the detected 2D pose, and combining a 3D pose increment from the residual regressor to the current 3D pose estimate.
Owner:INTEL CORP

Incremental 2d-to-3d pose lifting for fast and accurate human pose estimation

Techniques related to 3D pose estimation from a 2D input image are discussed. Such techniques include incrementally adjusting an initial 3D pose generated by applying a lifting network to a detected 2D pose in the 2D input image by projecting each current 3D pose estimate to a 2D pose projection, applying a residual regressor to features based on the 2D pose projection and the detected 2D pose, and combining a 3D pose increment from the residual regressor to the current 3D pose estimate.
Owner:INTEL CORP

Human pose estimation method based on adaptive multi-view feature fusion

PendingCN122336798AData setLearning network
This invention discloses a human pose estimation method based on adaptive multi-view feature fusion. Specifically, it involves: acquiring synchronized multi-view images and preprocessing them to form a dataset, which is then divided into a training set, a validation set, and a test set; designing a multi-view feature fusion network, including a 2D pose estimation backbone network, an appearance embedding network, a geometric embedding network, a weight prediction and learning network, and a multi-view feature weighted fusion network; training and validating the multi-view feature fusion network using the training and validation sets; inputting the test set into the trained multi-view feature fusion network for testing; and outputting a pose estimation image. This method can enhance the features of occluded views by borrowing features from visible views, while adaptively learning view weights to prevent low-quality views from contaminating global features, significantly improving the accuracy and robustness of 3D pose estimation.
Owner:XIAN UNIV OF TECH

3D human body posture estimation method and system based on meta-learning

The invention discloses a 3D human body posture estimation method and system based on meta-learning, and belongs to the technical field of computer vision and deep learning, and the method comprises the steps: constructing a dual-task framework of a 3D human body posture estimation main task and a self-supervised 2D posture reordering auxiliary task, and combining a mechanism of internal circulation fine tuning and external circulation optimization of meta-learning, and rapid adaptation of the model to new target domain data in a test stage is realized. According to the method, the problems of slow convergence and poor generalization ability of a traditional 3D attitude estimation model in a cross-domain scene are solved, and the 3D attitude estimation precision can be improved only by a small amount of iteration under the condition of no target domain labeling through the combination of the self-supervised auxiliary task and meta-learning.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

3d pose estimation by 2d camera

A system and method for obtaining the 3D pose of an object using 2D images from a 2D camera and a learning-based neural network. The neural network extracts multiple features from the 2D image of the object and generates a generated heatmap for each of the extracted features, which uses color representation to identify the probability of the location of feature points on the object. The method provides a feature point image including each feature point from the respective heatmaps on the 2D image, and estimates the 3D pose of the object by comparing the feature point image with a 3D virtual CAD model of the object.
Owner:FANUC LTD

Zero-sample unknown object 3D pose estimation method based on two-stage RGB-D fusion

The invention discloses a zero-sample unknown object 3D pose estimation method based on two-stage RGB-D fusion. Comprising the following steps: giving a 3D model of an unknown object, and rendering the 3D model of the unknown object under different viewpoints to generate a series of RGB-D template images; based on a feature embedded network PoseFusion, texture features and geometric features are extracted from an RGB template image and a Depth template image through a modified ResNet50 network, and then a series of template scene features are generated through two-stage fusion of the texture features and the geometric features. Similarly, a real RGB-D image obtained by shooting an unknown object is also embedded into a real scene feature through PoseFusion. And finally, matching the real scene features with the series of template scene features by calculating occlusion local similarity, so that the 3D pose of the unknown object is determined as the matched template 3D pose.
Owner:ZHEJIANG UNIV

A multi-view three-dimensional pose reconstruction method

This invention provides a multi-view 3D pose reconstruction method, comprising: constructing a virtual simulation environment; setting the parameters of the virtual simulation environment using a domain randomization strategy; acquiring multi-view images and extracting 2D joint coordinates and confidence scores; normalizing the 2D joint coordinates based on the parameters of each virtual camera to obtain normalized coordinates; concatenating the normalized coordinates with the confidence scores and inputting them into a shared-weight attention encoder; learning the joint topology within a single view through self-attention and outputting view features; concatenating the rotation matrix and translation vector of each view and mapping them to an explicit extrinsic semantic vector via a multilayer perceptron; inputting the view features and extrinsic semantic vector together into a global attention module; performing geometric consistency evaluation through an attention mechanism; and outputting 3D joint coordinates. This method solves the problems of weak cross-view generalization ability, large interference from differences in camera intrinsic parameters, and insufficient robustness against occlusion in existing technologies for multi-view 3D pose estimation.
Owner:BEIJING JINGCAI INTELLIGENT TECH CO LTD

Deep augmented human pose and size estimation using occlusion-aware neural networks

Methods and systems are disclosed for estimating 3D poses and dimensions of vehicle occupants using neural networks. Images of the interior of a vehicle are captured and monocular depth maps are generated. Both the depth maps and the images are fed into a 3D pose estimation network as a combined four-channel RGBD input. The neural network can include an occlusion-aware masking layer that generates an occlusion score for keypoints associated with an occupant. The occlusion score helps the network adjust the weighting of the depth information such that keypoints with higher occlusion scores (indicating that they are more likely to be hidden or partially occluded) receive lower weights. A scaling function that estimates a scale factor integrates the depth information with the occlusion-aware mask to estimate the absolute depth position of the keypoints.
Owner:NVIDIA CORP

Machine learning model training apparatus and method, and three-dimensional pose estimation apparatus

PCT designated stage expiredWO2024255510A9Machine learningPattern recognitionMachine learning
The present disclosure relates to the technical field of computers, and relates to a machine learning model training apparatus and method, and a three-dimensional pose estimation apparatus. The training apparatus comprises at least one processor, the at least one processor being configured to: acquire a plurality of two-dimensional images and a machine learning model to be trained, the plurality of two-dimensional images comprising a target, and the machine learning model comprising a first three-dimensional pose estimation module and a second three-dimensional pose estimation module; according to any one two-dimensional image, use the first three-dimensional pose estimation module to estimate a first three-dimensional pose of the target; according to at least two two-dimensional images, use the second three-dimensional pose estimation module to estimate a second three-dimensional pose of the target, the viewing angles of the at least two two-dimensional images being different from each other; and training at least one of the first three-dimensional pose estimation module or the second three-dimensional pose estimation module according to the difference between the first three-dimensional pose and the second three-dimensional pose.
Owner:BOE TECHNOLOGY GROUP CO LTD +1

Pose estimation method, apparatus, device, and medium

The present disclosure provides a pose estimation method, which relates to the field of artificial intelligence. The method comprises: collecting N body images of a user at the same time from N perspectives, the N perspectives comprising an upper perspective, the collection position of the upper perspective being higher than the height of the user, and N being greater than or equal to 3; inputting the N body images into N two-dimensional pose estimation models respectively to obtain N two-dimensional pose data, wherein the N two-dimensional pose estimation models have a one-to-one mapping relationship with the N perspectives; and inputting the N two-dimensional pose data into a three-dimensional pose estimation model to obtain three-dimensional pose data of the user. The present disclosure also provides a pose estimation device, equipment, storage medium and program product.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

3D pose estimation based on color and depth features

PCT designated stage expiredWO2025146667A1Image enhancementImage analysisColor imageNerve network
An electronic device and method for 3D pose estimation based on color and depth features is provided. The electronic device acquires color image data that includes an object and depth image data associated with the object. The electronic device generates a set of color feature maps based on the color image data and a set of depth feature maps based on the depth image data. The electronic device generates a set of feature fusion maps based on the set of color feature maps and the set of depth feature maps. The electronic device further generates 3D feature data of the object based on projection of each feature included in each feature fusion map to a 3D point of a 3D volume. The electronic device computes a 3D pose of the object based on application of a pose-estimation neural network on the 3D feature data and controls rendering of the 3D pose.
Owner:SONY GROUP CORP

Robot sucker grabbing detection network and method based on multi-scale attention

The invention belongs to the technical field of robot control, and discloses a robot suction cup grabbing detection network and method based on multi-scale attention, and the network comprises a multi-scale attention module, a multi-scale fusion module and an improved suction cup grabbing evaluation module. The multi-scale attention module is based on a HarDNet-68 backbone network, introduces a frequency domain channel attention mechanism, combines multi-scale depth separable convolution and expansion convolution, and enhances the multi-scale feature extraction capability; the multi-scale fusion module adopts a two-stage progressive fusion strategy, effectively integrates shallow details and deep semantic features, and generates a capture quality map and an object center map; the improved evaluation module improves the precision of side grabbing detection and 3D attitude estimation through multi-level threshold and truth value supervision. According to the method, the multi-scale object detection and grabbing success rate is remarkably improved in a complex industrial scene, and high robustness and real-time performance are achieved.
Owner:ROBOTICS RESEARCH CENTER OF YUYAO CITY +1