Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

2245 results about "Pose" patented technology

In computer vision and robotics, a typical task is to identify specific objects in an image and to determine each object's position and orientation relative to some coordinate system. This information can then be used, for example, to allow a robot to manipulate an object or to avoid moving into the object. The combination of position and orientation is referred to as the pose of an object, even though this concept is sometimes used only to describe the orientation. Exterior orientation and translation are also used as synonyms of pose.

Heavy-load robot motion trail method and system based on machine learning

The invention relates to the technical field of robot control, and discloses a heavy-load robot motion trail method and system based on machine learning. The method comprises the steps that historical movement track data of the heavy-load robot in a working scene are collected, and the data comprise a joint position sequence, an end effector pose sequence and environment obstacle distribution information; the data is preprocessed, track features are extracted, a space-time correlation matrix is constructed, and the matrix is used for representing the dynamic coupling relation between joint movement and the tail end pose; training a trajectory prediction model containing a long and short-term memory network and an attention mechanism based on the matrix, and generating a collaborative mapping relation between a joint position and a tail end pose; obtaining a current task target pose sequence and an environment constraint condition in real time, and outputting a candidate track set meeting dynamic constraint through a model; and adopting a multi-objective optimization algorithm to screen candidate tracks, generating an optimal track instruction and issuing the optimal track instruction to an execution mechanism. The method adapts to the complex characteristics and variable working conditions of the heavy-load robot, and the track adaptability is improved.
Owner:NINGBO WELLLIH ROBOTS TECH CO LTD

Target structure automatic detection method and device, equipment and medium

The invention relates to the technical field of intelligent manufacturing, and discloses a target structure automatic detection method, device and equipment and a medium, and the method comprises the steps: obtaining scanning path planning data of a target detection structure, driving an ultrasonic probe to execute surrounding scanning motion, and collecting an ultrasonic image sequence and spatial pose data; dynamically adjusting a pressure application angle and a scanning speed based on force feedback information, fusing spatial pose data and an image sequence to perform three-dimensional reconstruction, constructing a three-dimensional geometric model of a target detection structure, extracting feature distribution data by applying an intelligent analysis model, generating a feature decision set, and performing feature extraction; and mapping the feature decision set to a three-dimensional coordinate system to construct an analysis report containing the feature type marks and the topological relation. Through fusion of force control scanning, image reconstruction and intelligent analysis, standardization and intelligence of a detection process are realized, image consistency and structure identification precision are improved, manual dependence is reduced, and comprehensiveness and reliability of lesion identification are enhanced.
Owner:SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD

Mechanical arm grabbing method and system based on multi-modal information fusion

The invention provides a mechanical arm grabbing method and system based on multi-modal information fusion, and belongs to the technical field of robot intelligent control. Comprising the steps that the conversion relation between a camera coordinate system and a mechanical arm base coordinate system is established through camera calibration, a deep learning neural network is used for conducting grabbing pose estimation on an obtained RGB-D image, and multiple candidate grabbing poses are determined; analyzing a natural language instruction input by a user based on a multi-modal large model, and recognizing a target object region from the RGB-D image by combining a target detection and image segmentation technology; based on the obtained candidate grabbing poses and the target object area, an optimal grabbing pose is screened through a scoring mechanism and mapped to a mechanical arm base coordinate system; and then a dynamic grabbing path is generated by adopting an imitation learning algorithm, and the mechanical arm is controlled to execute grabbing operation. Through multi-modal semantic understanding, accurate grabbing of the mechanical arm in a complex environment can be achieved.
Owner:SHANDONG UNIV

Action control method and device based on physical reference, equipment and medium

The invention relates to the technical field of robot visual perception and motion control, and discloses a motion control method and device based on physical reference, equipment and a medium, and the method comprises the steps: obtaining instruction information, a multi-view image and movable assembly pose information; processing the multi-view image according to the instruction information to generate target segmentation information; generating a scale normalization point cloud and a model estimation baseline; determining a physical reference baseline and generating a scale calibration factor; converting the scale normalization point cloud into a physical space point cloud by using a scale calibration factor; extracting a three-dimensional relative position of the target object relative to the movable component in combination with the target segmentation information; an action instruction is generated based on the multi-modal input. According to the method, physical scale alignment of the point cloud is realized through physical reference baseline calibration, so that a visual reconstruction result has real space significance, an accurate action instruction is generated, and the robot space understanding and operation precision is improved.
Owner:SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD

Ring main unit inspection robot autonomous navigation method and system based on SLAM

The invention discloses a ring main unit inspection robot autonomous navigation method and system based on SLAM, particularly relates to the technical field of robot autonomous navigation and intelligent inspection, and is used for solving the problem of positioning drift caused by repeated features of an existing ring main unit scene. Semantic feature analysis and topological constraints are introduced into an SLAM processing flow, acquired image data and point cloud data are processed through a deep learning model, objects such as an electrical cabinet, a corridor channel and a cable trench are identified, and a semantic feature set with category labels and spatial position information is generated; and constructing a topological graph containing node spacing, connectivity and directivity constraints based on the semantic features, adding the topological graph as a constraint factor into SLAM back-end optimization, and performing joint optimization in combination with vision, a laser odometer and inertial prior information, thereby avoiding only depending on repeated geometric feature positioning, and improving the positioning accuracy. The problems of loopback misjudgment and drifting caused by feature confusion are reduced, and the pose resolving stability in the ring main unit environment is improved.
Owner:STATE GRID HUBEI ELECTRIC POWER CO XIAOGAN POWER SUPPLY CO

Target pose sensing estimation method and system for mechanical arm grabbing operation and robot

The invention belongs to the technical field of artificial intelligence, and provides a target pose sensing estimation method and system for mechanical arm grabbing operation and a robot. According to the technical scheme, multi-modal data of a target object in the mechanical arm grabbing process under a historical language instruction is obtained; based on the obtained multi-modal data of the target object in the historical mechanical arm grabbing process under the historical language instruction, the end-to-end large model is trained, and a trained large model is obtained; according to real-time environment information and task requirements, the weight of each sub-target in multi-target optimization in the large model is dynamically adjusted to obtain an inference model, an action sequence is generated based on a vector field predicted by the inference model, the generation process is optimized, and meanwhile, multi-modal information is fused for real-time decision making; and the action sequence of the mechanical arm executing the grabbing task in the current environment is determined. And the grabbing efficiency and safety of the mechanical arm are improved.
Owner:STATE GRID INTELLIGENCE TECHNOLOGY CO LTD

Visual prediction method for space-time manifold and implicit state deduction of robot

The invention discloses a visual prediction method for space-time manifold and implicit state deduction of a robot, and relates to the technical field of robot space visual perception. The motion state and the visual image of the robot are collected; calculating carrier pose offset and generating a reverse compensation matrix; modeling an image target into a visual cone probability manifold, and extracting an explicit manifold observation vector; monitoring the confidence coefficient in real time through an observation quality evaluation network, and activating an implicit deduction mode during observation degradation; reading historical time sequence characteristics by using an improved Transform model, and performing prediction and deduction on a future nonlinear motion state of the target; performing coordinate correction and probability field decoding on a prediction result in combination with the reverse compensation matrix, and outputting a three-dimensional prediction trajectory and a spatial covariance ellipsoid; according to the method, the problem that traditional visual tracking fails under the conditions of violent shaking of a carrier and target shielding is solved, and continuous and foresight physical scale positioning and safety decision making of a robot on a dynamic target are achieved.
Owner:SUZHOU MENGWU INTELLIGENT TECHNOLOGY CO LTD

Control method and device based on master-slave cooperation, equipment and medium

The invention relates to the technical field of robot control, and discloses a master-slave cooperation-based control method, device, equipment and medium, and the method comprises the steps: collecting a spatial pose, a contact force and a dual-view image, and constructing a multi-modal demonstration data set; executing time alignment and normalization to form a standardized input sequence; action features, force features and visual features are extracted through the action modal coding network, the mechanical modal coding network and the visual coding network; fusing to form a multi-modal tensor sequence, and inputting the multi-modal tensor sequence into a Transform decoder to generate a joint time sequence representation; outputting a training action prediction vector from the joint time sequence representation through an action predictor, constructing a supervision loss function in combination with expert demonstration annotation to update each network, and obtaining an optimization control model; and processing the real-time input control driving execution mechanism based on the optimization control model. According to the method, the control model is trained and optimized through multi-modal sensing and time sequence modeling, the joint control instruction is generated in deployment, and stable execution of a complex interaction task is achieved.
Owner:SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD

Target positioning and capturing method and system based on multi-modal semantics

The invention discloses a target positioning and capturing method and system based on multi-modal semanteme, and relates to the technical field of image recognition and mechanical control, and the method comprises the steps: obtaining a two-dimensional image and a natural language interaction instruction; visual-language feature alignment processing is carried out, if semantic ambiguity exists in a natural language interaction instruction in the alignment processing process, a reverse question statement is generated based on a generative interaction mechanism, and a unique target is determined; collecting a multi-view two-dimensional image of a target, and fusing to obtain a three-dimensional point cloud model; inputting a spatial pose generation network to obtain candidate six-degree-of-freedom spatial poses; performing semantic common sense filtering and geometric interference filtering in sequence; calculating a posture score of the filtered candidate six-degree-of-freedom space posture, and selecting an optimal target three-dimensional posture; generating a grabbing operation parameter sequence; and executing the grabbing operation parameter sequence. According to the method provided by the invention, the problem of unsuccessful grabbing caused by semantic ambiguity and shielding is solved.
Owner:XIAN UNIV OF POSTS & TELECOMM

Semantic aerial view visual relocation method and device in non-exposed scene, electronic equipment, storage medium and program product

The invention provides a semantic aerial view visual relocation method and device in a non-exposed scene, electronic equipment, a storage medium and a computer program product. The method comprises the following steps: acquiring a multi-view image sequence under a non-exposed scene (such as a tunnel, an underground pipe gallery or an underground parking lot); semantic recognition is carried out based on a pre-trained semantic target detection model, and spatial consistency semantic features are extracted through a semantic-geometric dual-channel fusion mechanism combining a semantic mask and geometric constraints; the method comprises the following steps of: realizing three-dimensional reconstruction by using a voxel micro-renderable modeling method (VGGT), and generating a dense three-dimensional semantic point cloud fusing semantics and a geometric structure; two-dimensional semantics are mapped to a three-dimensional space through a projection and back projection relation, and point cloud semantics are endowed; main structure planes such as the ground, the left wall surface and the right wall surface are extracted, and a two-dimensional semantic aerial view with semantic annotation is generated; and pose estimation is carried out based on a reciprocal matching strategy guided by a semantic mask, so that visual repositioning with high precision, high robustness and semantic interpretability is realized. The method breaks through the problems of low precision, sparse features and poor semantic consistency of traditional visual repositioning in a non-exposed environment, and can be widely applied to the fields of intelligent transportation, underground inspection and unmanned system positioning.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Multi-modal large model three-dimensional perception method based on Riemannian manifold priori guidance

The invention discloses a multi-modal large model three-dimensional perception method based on Riemannian manifold prior guidance, and relates to a computer vision technology. The method comprises the following steps: firstly, constructing a multi-modal large model and point cloud sensing network coordinated cross attention reinforcement learning joint training framework, so that the large model obtains a three-dimensional scene space topology understanding capability; secondly, a conformal property on a Riemannian manifold is utilized to construct a three-dimensional contour and topological association, and multi-scale invariance of homotopy mapping learning contours is introduced, so that the challenges of shape diversity and great size difference of objects are solved; in addition, a lightweight three-dimensional attention gate filtering mechanism is introduced into a point cloud encoder, and more effective global and local point cloud geometric semantic association is established. And finally, when the motion pose quality is evaluated, fusing prior knowledge of the three-dimensional physical relationship to form mixed physical measurement so as to overcome label noise caused by single evaluation scale. And universal understanding and high-precision perception of the robot on complex three-dimensional environments and objects can be realized.
Owner:XIAMEN UNIV

Robot arm control method, device, equipment, medium and product

The invention discloses a robot arm control method and device, equipment, a medium and a product, and the method comprises the steps: obtaining a pre-training dynamic model and a real visual depth map of a real robot arm visual angle, and the pre-training dynamic model is used for completing a specified task; determining a joint control value according to the real vision depth map and a pre-training kinetic model; hybrid action control is carried out based on the joint control value, and remote target navigation is carried out on the response action of the real robot arm through a visual navigation model; and the controlled simulation target image and the actual image are obtained, feature matching and closed-loop estimation are carried out based on the simulation target image and the actual image, and pose error compensation is carried out on the action of the real robot arm. The motion is observed and deduced through a real visual depth map, then real and simulated mixed motion control is carried out to reduce a visual and dynamic gap, and finally pose error compensation is carried out. And the pose error of the arm is reduced, and a start pose guarantee is provided for downstream control.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

System and method for calibration of humanoid robots

The present disclosure provides a method for calibrating a humanoid robot, comprising obtaining a humanoid robot with original kinematic biasing values, controlling the humanoid robot through predetermined poses, capturing image data of body parts using vision sensors mounted on the humanoid robot while moving through the poses, determining revised kinematic biasing values by processing the image data using a bipedal spatial perception model trained using synthetic image data containing keypoints, and replacing the original kinematic biasing values with the revised kinematic biasing values. The bipedal spatial perception model processes captured image data to generate observed keypoint locations on robot components, which are compared with kinematic-based locations from joint encoder measurements to minimize discrepancies through optimization algorithms.
Owner:FIGURE AI INC

CNN and Transform fused self-supervised monocular depth estimation system and method

The invention relates to the technical field of computer vision, and particularly discloses a CNN and Transform fused self-supervised monocular depth estimation system and method. According to the system, local representation is enhanced through a multi-scale feature fusion mechanism, a cross-regional attention network is constructed to realize global context association, and fine reduction of a fine structure of a complex scene is realized. Firstly, based on DCB, multilayer expansion convolution is adopted to expand a receptive field, multi-scale pixel features are fused, and local details of a key area are enhanced; secondly, capturing fine-grained local information of the image by using parallel local convolution of ELGF, and acquiring long-distance dependency by using a self-attention mechanism, thereby realizing collaborative modeling of local information and global dependency, and remarkably enhancing feature expression ability; and finally, estimating a relative pose between adjacent images through a pre-trained ResNet18-based lightweight encoder, and constructing reprojection loss to optimize depth prediction. Experiments show that the model constructed by the method provided by the invention reaches 0.102 and 4.430 in AbsRel and RMSE indexes respectively, and is obviously superior to the existing mainstream method.
Owner:SHANGHAI DIANJI UNIV

Task-oriented grabbing method and system for cross-level constraint reasoning

The invention discloses a task-oriented grabbing method and system based on cross-level constraint reasoning, and belongs to the technical field of robot grabbing control. The method comprises the following steps: aligning a text instruction with an input image through a mask alignment module, generating a target area mask by utilizing SAM-Clip, and generating a target point cloud by taking the target area mask as spatial prior; and further analyzing an internal physical structure of the target by using VLM, guiding to generate a 6-DOF grabbing attitude, performing collision detection and quality scoring by combining high-level function and bottom-level geometric prior, and outputting an optimal grabbing attitude. According to the method, the problems that task understanding and scene perception are disjointed, and the grabbing posture lacks constraint are solved, and the grabbing success rate and the intelligent level of the robot in the open environment are remarkably improved.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

Robot scanning path adjusting method and device, equipment and medium

The invention relates to the technical field of robot control, and discloses a robot scanning path adjusting method, device, equipment and medium, and the method comprises the steps: controlling a robot equipped with a sensor to move along a preliminary path, collecting a preliminary two-dimensional image of a target object, processing the image, extracting the feature points of the target object, and combining the spatial information of the sensor when the image is collected; the method comprises the steps of generating a primary three-dimensional point cloud, determining a plurality of ideal sensor postures based on the primary three-dimensional point cloud, determining a new sensor position of each posture on a contactable surface to generate a re-planned scanning path, controlling a robot to move along the re-planned scanning path, adjusting the sensor postures, and collecting a final two-dimensional image of a target object. According to the method, the scanning path is optimized in combination with the initial three-dimensional point cloud local surface normal direction, so that the attitude of the sensor is matched with the geometrical characteristics of the target surface, the problem of three-dimensional reconstruction surface breakage caused by inter-frame pose mismatching under free breathing is solved, and the accuracy of autonomous ultrasonic scanning of the robot is improved.
Owner:SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD

Machine-learned model architecture for predicting future object state

Predicting a future state, such as a future position and / or orientation (i.e., pose), of an object may comprise classifying, by a first machine-learned model, a lane the object may occupy and classifying, by a second machine-learned model, a target pose the object may occupy. A third machine-learned model may determine an offset from the target pose that may be used to determine a predicted (future) pose of the object by applying the offset to the target pose.
Owner:ZOOX INC

Thick plate weld groove orthogonal section contour three-dimensional point cloud reconstruction method and system

The invention belongs to the technical field of visual identification, and particularly discloses a thick plate weld groove orthogonal section contour three-dimensional point cloud reconstruction method and system. Comprising the following steps: dividing a three-dimensional point cloud of a groove contour into a plurality of parts according to geometric characteristics of a weld groove; receiving designation of a target part, extracting a point cloud subset at the left junction of the target part, and further fitting to obtain an edge line; confirming a point closest to the three-dimensional point cloud of the current frame on the edge line; calculating a tangent vector of the edge line at the nearest point as a normal vector of an orthogonal section; calculating the distance between each point in the three-dimensional point cloud of the current frame and the normal vector of the orthogonal section, and screening out all points which do not exceed a preset threshold value; and calculating a projection point of each screened point on the orthogonal section to jointly form a reconstructed thick plate weld groove orthogonal section contour. According to the method, the coordinate system orthogonally perpendicular to the welding seam in real time is constructed, the three-dimensional point cloud is projected to the orthogonally perpendicular plane, the orthogonal section contour of the groove of the welding seam is obtained, and the measurement error caused by the pose of the sensor is eliminated.
Owner:TIANJIN UNIV

Indoor scene three-dimensional reconstruction method based on deep fusion and confidence modeling

The invention discloses an indoor scene three-dimensional reconstruction method based on deep fusion and confidence modeling, and belongs to the technical field of computer vision. According to the method, composite data frames such as a color image, a depth map, an IMU (Inertial Measurement Unit) and a camera attitude are comprehensively utilized to carry out regional three-dimensional reconstruction on an indoor scene: firstly, the scene is divided into a smooth region (such as a wall, a ground, a ceiling, a glass plane, a mirror surface, a blackboard and other planes) and a complex curved surface region; aiming at the smooth area, adopting geometric prior guide plane fitting provided by a visual large model, and combining sensor attitude information to quickly reconstruct a regular plane model; for a complex curved surface area, a multi-frame point cloud fusion strategy is adopted to accumulate different view angle information, and a deep residual error refining network is utilized to recover curved surface details, so that a high-precision curved surface model is obtained. According to the method, the three-dimensional structure of the indoor scene can be efficiently reconstructed, the global framework of the smooth area is reserved, and the details of the surface of a complex object are depicted in detail.
Owner:CHONGQING UNIV OF EDUCATION +1

Control method and system for collaborative trajectory planning of container inspection robot

The invention discloses a control method for collaborative trajectory planning of a container inspection robot. The control method comprises the following steps: S1, constructing a global task and a map; s2, obtaining real-time state sensing and positioning of the inspection robot (6); s3, establishing a unified three-dimensional cooperative control decision algorithm, and cooperatively controlling the height of an electric rod on the robot and the pitch angle (pitch), the yaw angle (yaw) and the roll angle (roll) of a holder based on the real-time pose of the robot, the distance between the ideal track and the target container and the standard size of the pre-modeled container; s4, the robot runs according to the planned track to execute the task and collects data; s5, performing global state management and safety strategy of the electric rod of the robot; the method has the beneficial effects that a detection blind area is eliminated, and three-dimensional cooperative control is realized; a camera can lock a target in a three-dimensional space at an optimal pose, image quality consistency is ensured, and jitter is inhibited through active holder movement; and full-process automation and safety protection are realized.
Owner:SHANDONG GUOCHUANG INTELLIGENT ROBOT RESEARCH INSTITUTE CO LTD

Indoor SLAM map construction method and system

The invention provides an indoor SLAM map construction method and system, and relates to the technical field of image processing, and the construction method comprises the steps: building a unified global space reference through cross-device space-time calibration, supplementing the visual blind area information of a robot in real time through the global static features extracted by an environment camera, and obtaining the visual blind area information of the robot; and meanwhile, dynamic interference features are accurately eliminated in combination with continuous frame analysis, and finally, the pose of the robot is continuously corrected by taking global features as anchor point constraints in back-end optimization. The complete technology chain enables the constructed indoor three-dimensional map to have higher integrity, precision and consistency, significantly improves the positioning robustness and navigation reliability of the robot in an environment with dense goods shelves and dynamic activities, reduces the performance dependence on a robot body sensor and the scene reconstruction cost, and improves the reliability of the robot. And a stable and efficient environment sensing basis is provided for automatic operation of industries such as storage and archive management.
Owner:ZHEJIANG BEITAI INTELLIGENT TECH CO LTD

Object three-dimensional reconstruction method, device and system based on deep learning

The invention discloses an object three-dimensional reconstruction method, device and system based on deep learning. The reconstruction method comprises the following steps: acquiring a multi-view color image of an object through a controllable image acquisition device; reconstructing a sparse three-dimensional point cloud by using a motion recovery structure method and obtaining a camera pose; initializing parameters of the three-dimensional Gaussian sputtering model based on the sparse point cloud and performing training optimization; a target object semantic segmentation data set is constructed, and a low-rank adaptive technology is adopted to finely segment all models; generating prompts through an open vocabulary detection model at each view angle, obtaining an accurate segmentation mask, and optimizing a three-dimensional segmentation weight by adopting a joint loss function fusing color consistency loss and edge perception loss; and finally outputting the color three-dimensional point cloud of the target object. According to the method, the original image is segmented, so that the influence of the quality of the rendered image is avoided; the segmentation precision of the model in a specific scene is improved through field adaptive fine tuning; and the accuracy of the segmentation boundary is ensured by adopting a double-loss joint optimization mechanism.
Owner:HUNAN AGRI UNIV

High-precision map reconstruction method and system based on monocular vision, medium and equipment

The invention belongs to the technical field of robot positioning and three-dimensional mapping, and discloses a high-precision map reconstruction method and system based on monocular vision, a medium and equipment. Scene image data are acquired through a monocular image acquisition module, after feature extraction and matching are performed on each frame of image, attitude information acquired by an inertial measurement unit is fused, and pose calculation of a robot is completed through a sparse vision SLAM system; meanwhile, an image dense depth map is generated by a monocular depth estimation model based on an attention mechanism, and the image dense depth map is converted into a single-frame color dense point cloud in combination with image RGB color information. According to the invention, a loose coupling fusion strategy is adopted to carry out spatial registration on a single-frame colored dense point cloud and a robot pose at a corresponding moment, multi-frame data fusion is completed through point cloud splicing and optimization, and a globally consistent three-dimensional dense point cloud map is constructed to realize scene modeling. The method has the characteristics of high robustness and high reconstruction precision, and can be effectively applied to three-dimensional map construction in an outdoor complex environment.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Automatic high-precision three-dimensional dense reconstruction method for box girder reinforcement cage

The invention belongs to the technical field of steel reinforcement framework three-dimensional reconstruction, and particularly discloses an automatic high-precision three-dimensional dense reconstruction method for a box girder steel reinforcement framework, which comprises the following steps of: dividing a scanning area into a plurality of sub-areas according to a depth image by moving an inspection robot along the steel reinforcement framework, and performing multi-angle local scanning in each sub-area; constructing a local point cloud model by combining camera poses during acquisition, further calculating a relative spatial transformation relationship by using an overlapping region between adjacent local models, constructing a pose map and implementing global optimization, and uniformly adjusting the poses of the local models in a world coordinate system; and finally, all corrected local point clouds are fused to generate a complete global three-dimensional model, so that collaborative reconstruction operation of domain-divided scanning, segmentation reconstruction and global optimization is realized, frame-by-frame accumulation of errors in traditional scanning-while-splicing reconstruction is effectively blocked, long-distance drift is remarkably inhibited, and the precision of three-dimensional reconstruction of the reinforcement cage is greatly improved.
Owner:CHINA TIESIJU CIVIL ENGINEERING GROUP CO LTD +1

Complex stacking scene robot disordered grabbing method based on three-dimensional vision

The invention discloses a complex stacking scene robot disordered grabbing method based on three-dimensional vision, and the method comprises the steps: firstly obtaining an original point cloud through a depth camera, and obtaining a high-quality point cloud after statistics and radius filtering, moving least square smoothing, background plane removal and grid simplification; then, hand-eye calibration of participation eyes inside the camera outside the hand is completed, and alignment of vision and a robot coordinate system is achieved; single workpieces are separated through Euclidean clustering; in the key point extraction stage, geometric significance and super voxel density screening are integrated, weighted three-dimensional SIFT and multi-scale SHOT descriptors are adopted for feature expression, and rough registration and symmetric ICP fine registration are combined to obtain the six-degree-of-freedom pose of a workpiece; and based on the self-adaptive curvature and the force closing score, a grabbing posture is optimized, a collision-free path is planned by using an RRT algorithm, and grabbing is executed. The method does not need pre-modeling, can adapt to shielding, symmetry and density change scenes, is low in deployment cost, is simple to maintain, and is suitable for industrial disordered feeding and discharging and logistics sorting.
Owner:JIANGSU UNIV OF TECH

Scene three-dimensional reconstruction and vector information extraction method and device based on vehicle return data, equipment and storage medium

The invention discloses a scene three-dimensional reconstruction and vector information extraction method and device based on vehicle return data, equipment and a storage medium, and the method comprises the steps: carrying out the preprocessing of a time sequence image returned by a vehicle and corresponding pose data, and obtaining a semantic mask and a camera pose of each frame of image in the time sequence image; based on the semantic mask and the camera pose, performing three-dimensional geometric reconstruction on the static scene area in the time sequence image to obtain a static three-dimensional scene model; a dynamic target is separated from the time sequence image according to the semantic mask, three-dimensional motion modeling and independent three-dimensional reconstruction are carried out on the dynamic target, and a dynamic target three-dimensional model is generated; and generating a dynamic three-dimensional scene model and vector labeling information based on the static three-dimensional scene model and the dynamic target three-dimensional model. According to the method, the limitation of automatic driving shadow mode data is overcome, the three-dimensional reconstruction of the scene and the automatic generation of the true value information are realized, the data acquisition cost is further reduced, and the efficiency of automatic driving research and development is improved.
Owner:FOSS (HANGZHOU) INTELLIGENT TECH CO LTD

VLA model training method and device

The invention provides a training method and device for a VLA model, and relates to the technical field of intelligent robots with bodies, and the method comprises the steps that the VLA model outputs a joint angle sequence of a mechanical arm of a robot with a body based on visual input and language input; the differentiable forward kinematics module outputs a real-time pose of an end effector of the mechanical arm in a task space; based on joint space loss formed by the joint angle sequence and the joint angle in the teaching data, task space loss formed by the real-time pose and the tail end pose in the teaching data, and task space constraint loss, a multi-objective loss function is constructed, and total loss is calculated; calculating the gradient of the total loss to the joint angle sequence and the real-time pose through back propagation, and optimizing preset parameters of the VLA model; the above steps are repeatedly executed until the total loss meets the expectation, and a trained target VLA model is obtained; the technical problems that when a VLA model is trained in a joint space, the hardware coupling performance is high, and the generalization ability is insufficient are solved.
Owner:ANHUI KAIYANG TECHNOLOGY CO LTD +1

Three-dimensional space data correction method for virtual reality

The invention provides a virtual reality three-dimensional space data correction method, which comprises the following steps of: firstly, acquiring an image, depth and inertia measurement data of a virtual reality terminal, and a pose, a speed and a force feedback signal of an interaction body, and carrying out time sequence alignment under a unified time reference; an interaction event is identified through collision detection, and parameters such as a contact point, a normal direction, a relative speed, a penetration depth and a contact duration are extracted. Then, physical consistency constraints including no penetration, contact stability, friction consistency, momentum conservation and energy non-increase are constructed, and geometric consistency, depth consistency and physical consistency residual functions are established in the local contact area; and solving surface normal, depth, scale, pose, calibration parameters and other corrections through a robust nonlinear optimization method so as to update the point cloud, the grid and the calibration model. According to the method, the consistency and the stability of geometric modeling and physical interaction in the virtual reality scene can be remarkably improved, and the immersion and the interaction precision are improved.
Owner:LILU (SHANGHAI) CULTURE TECH CO LTD

AUV (Autonomous Underwater Vehicle) three-dimensional pose joint estimation method and system based on multi-modal layering

The invention relates to the technical field of underwater positioning, in particular to an AUV (Autonomous Underwater Vehicle) three-dimensional pose joint estimation method and system based on multi-modal layering, and the method comprises the steps: distributing a historical observation sequence to corresponding independent convolutional neural network branches according to modals, carrying out the local time sequence feature extraction through each branch, and carrying out the local time sequence feature extraction; outputting local time sequence characteristics of each mode; performing hierarchical feature fusion based on local time sequence features of each mode, inputting global fusion features into a double-branch regression head, respectively decoding through a position regression branch and an attitude regression branch, and outputting a three-dimensional position increment and an attitude increment; and constructing a loss function by using the three-dimensional position increment and the attitude increment, training to obtain an optimal model, and deploying the optimal model to an AUV platform, thereby breaking through the limitation that the traditional method only focuses on a two-dimensional plane or separately estimates the attitude, completely covering the full-space positioning demand of the AUV three-dimensional maneuvering task, and improving the integrity and consistency of the attitude estimation.
Owner:OCEAN UNIV OF CHINA

Three-dimensional geographic environment real-time intelligent deduction method based on multi-modal large model

The invention discloses a three-dimensional geographic environment real-time intelligent deduction method based on a multi-modal large model, and relates to the technical field of computer vision. The method is used for solving the technical problem of unified modeling and real-time deduction of a dynamic object and a static environment in a three-dimensional geographical environment. The method comprises the following steps: firstly, resolving a camera pose through a motion recovery structure algorithm to generate a sparse point cloud, and separating a dynamic foreground object in a monitoring video to extract motion features; thirdly, initializing a three-dimensional Gaussian distribution set based on the sparse point cloud, extracting semantic features through a visual encoder, and mapping the semantic features to corresponding Gaussian distribution; thirdly, a topological graph structure of Gaussian distribution is constructed, motion features are used as initial excitation, and coordinate offset and appearance variation of each distribution are iteratively updated through message passing calculation; finally, Gaussian distribution attributes are dynamically updated, a continuous deduction image sequence is synthesized through micro-rasterization rendering, and high-reality real-time simulation of dynamic evolution of the three-dimensional geographical environment is achieved.
Owner:LIAONING HONGTU CHUANGZHAN SURVEYING & MAPPING CO