Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

537 results about "2d images" patented technology

Semantic simultaneous localization and mapping method and system based on Gaussian splashing

The invention relates to a semantic simultaneous localization and mapping method and system based on Gaussian splashing. The method comprises the following steps: firstly, collecting a frame of RGB-D image, modeling a scene into a 3D semantic Gaussian field containing a plurality of 3D semantic gausses according to the RGB-D image, and rendering the 3D semantic gausses by using a tile rasterization technology to obtain 2D image plane gausses; rendering results of RGB color, depth and semantic features are extracted from the 2D image plane in a Gaussian mode, and the semantic features are decoded into semantic tags; constructing a mapping and tracking loss function by using an RGB color rendering result, a depth rendering result, a semantic tag and a truth value, and jointly optimizing a camera pose and a semantic Gaussian field based on a tracking stage and a mapping stage; and repeating the steps for each new frame of RGB-D image to complete the construction of the incremental semantic Gaussian map. Compared with the prior art, the method has the advantages of realizing robust camera tracking, real-time high-quality rendering, accurate 3D semantic reconstruction and the like.
Owner:TONGJI UNIV

Generating 2d image of 3D scene with conditioning signal

A computer-implemented method for generating a 2D image of a 3D scene. The method comprises obtaining arrangement data comprising a layout of the 3D scene and at least one conditioning signal. Each conditioning signal has a type among a predetermined set of at least two types. The method comprises applying a machine-learning function to the obtained arrangement data and viewpoint. The function comprises a scene encoder and a generative image model. The scene encoder takes as input the obtained arrangement data and viewpoint and outputting a scene encoding tensor. The generative image model takes as input the scene encoding tensor outputted by the scene encoder and outputting the generated 2D image. Such a generating method forms an improved solution for controllably generating a 2D image of a 3D scene.
Owner:DASSAULT SYSTEMES SA

Determining feature poses of electric vehicles to automatically charge electric vehicles

The invention is notably directed to a computer-implemented method for automatically charging an electric vehicle via an end effector (10) of a robotic arm (40) of an automated vehicle charging robot. The end effector is assumed to be structured so as to be able to connect to a charge port (220) of a vehicle. In addition, the automated vehicle charging robot further includes a camera system (102) having depth sensing capability. The method comprises the following steps. First, a reference position of a reference feature (210) of the vehicle is estimated thanks to the camera system. Next, a pose of the charge port of the vehicle is determined based on the estimated reference position. The robotic arm is subsequently instructed to actuate the end effector, based on the determined pose of the charge port, to connect the end effector to the charge port with a view to charging the vehicle. The reference position is estimated as follows. Both a 2D image and a depth image of a surface portion of the vehicle are obtained. This surface portion includes the reference feature, i.e., the feature of interest. Contour points of the reference feature are then extracted from the 2D image obtained. The 3D coordinates of the extracted contour points are subsequently reconstructed based on the depth image obtained. A geometric object (such a 2D plane) is then matched to the reconstructed 3D coordinates, e.g., by fitting the geometric object to the reconstructed 3D coordinates. Eventually, the reference position of the reference feature is determined based on the matched geometric object. The invention is further directed to related automated vehicle charging robots and computer program products.
Owner:EMBOTECH AG

System and method for dynamic generation and rendering of threedimensional objects from two-dimensional images

A computer-implemented process for creating 3D objects from 2D images includes receiving an input image, conditioning a generative model using image-derived, voxelized three-dimensional features, generating, by a transformer-based rectified-flow generative model parameterized as a base network optionally coupled to one or more low-rank adapter modules activatable at inference, a volumetric latent of the target object, the volumetric latent including a sparse, feature-augmented volumetric lattice obtained by transporting an initial random sample toward a learned manifold via a rectified-flow sampling process, decoding the volumetric latent by mapping the volumetric latent to a feature-bearing sparse volumetric field consistent with the volumetric lattice, decoding the field to a continuous implicit surface function, and extracting a watertight mesh by isosurface extraction, estimating camera-pose parameters by render-and-compare alignment between silhouettes rendered from the mesh and silhouettes of the input image, and, performing style-preserving inverse rendering on the mesh that updates UV-space albedo and material maps.
Owner:ARTLABS US INC

Depth estimation using odometry and hand tracking

A head-worn augmented reality (AR) device system includes cameras, display devices, and processors, along with a memory that stores specific instructions. When these instructions are executed by the processors, they enable the device to perform several operations. First, the device accesses a two-dimensional (2D) camera image taken by its camera. The device then generates a first set of three-dimensional (3D) tracked points using the device's odometry system applied to this 2D image. Optionally, a second set of tracked 3D points is created based on one or more images captured by the camera. These 3D points are projected onto the 2D camera image to create a sparse depth image. Finally, this 2D camera image, along with the newly formed depth image, is fed into a first machine learning model to generate a metric depth estimation.
Owner:SNAP INC

Methods and systems for dynamic inspection of transport structures

Methods and systems for assessing, determining, or quantifying structural properties of a transport structure are provided, including methods and systems for capturing, using first and second image capture sensors of an inspection system, a plurality of 2-dimensional (2D) images of an inspection area of the inspection system; detecting, using an AI engine, a transport structure in a first image from the plurality of 2D images; extracting, using the AI engine and based on the first image, a second image from the plurality of 2D images; generating, using the AI engine and based on the first image and the second image, a computing model representing the transport structure; analyzing, using the AI engine, the computing model thereby generating analysis data; generating, using a data processing unit, a report that indicates the analysis data; and initiate formatting, using the data processing unit, for display on a graphical interface, the report.
Owner:IVISYS SWEDEN AB

Three-dimensional target detection method based on cone point cloud clustering and 2D detection

The invention discloses a three-dimensional target detection method based on view cone point cloud clustering and 2D detection, and the method comprises the steps: generating a plurality of target bounding boxes with categories according to a 2D image, converting the target bounding boxes to an image coordinate system, and carrying out the screening of 3D point clouds, so as to obtain view cone point clouds; performing a point cloud clustering operation on the view cone point cloud, and calculating a result of the point cloud clustering operation to obtain a 3D frame meeting a preset requirement; and performing IOU calculation of different visual angles on the 3D frames to merge the 3D frames of the same object under different visual angles so as to obtain a 3D bounding box. According to the method, the problem of geometric information loss can be effectively avoided, and the accuracy of the finally obtained 3D bounding box is ensured; the problem that global search is needed due to the fact that 3D point cloud is directly adopted is avoided, the complexity of the calculation process is reduced, and the overall calculation efficiency is effectively improved; the preset frame requirement can be set according to the requirement of a use scene, so that different requirements are met, and the wide use range is ensured.
Owner:城市之光(深圳)无人驾驶有限公司

Laser radar point cloud densification method and device fusing image information and medium

The invention relates to a laser radar point cloud densification method and device fusing image information, and a medium. The method comprises the following steps: fixing a visual camera and a laser radar on a rigid tool, and carrying out joint calibration and data alignment; a laser radar is used to collect three-dimensional point cloud data, and a visual camera is used to shoot a scene; performing super-pixel segmentation on the image by using an image brightness linear iterative clustering algorithm; mapping the three-dimensional point cloud data into a segmented two-dimensional image superpixel pattern spot region by using a jointly calibrated camera model conversion relationship to realize feature clustering of the original three-dimensional point cloud of the laser radar; performing curved surface fitting on a clustered result by using a random sampling consistency algorithm; and linear interpolation is carried out on the fitted curved surface, dense points which are not covered by the original point cloud of the laser radar are supplemented and generated, and densification of the point cloud of the laser radar is realized. According to the method, the point cloud densification precision and practicability are improved, and technical support is provided for automatic driving, robot navigation and three-dimensional reconstruction.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Mechanical arm fine positioning and closed-loop calibration method and system based on heterogeneous visual fusion

The invention discloses a mechanical arm fine positioning and closed-loop calibration method and system based on heterogeneous visual fusion, and the method comprises the steps: heterogeneous data synchronous collection and cross-modal joint calibration, semantic segmentation based on prompt learning and high-precision mask generation. Point cloud refining and edge noise elimination in a view cone space; and differential compensation and closed-loop feedback control based on a golden model. A heterogeneous visual fusion framework is constructed, and 3D point cloud noise is constrained by using a sub-pixel edge of a high-resolution 2D image, so that the positioning precision breaks through a sub-millimeter level; a segmentation large model with zero sample generalization ability and a prompt learning mechanism are introduced, so that the system can adapt to various workpieces under a complex background without repeated training; refining point cloud and a standard three-dimensional gold model are registered and fed back to the mechanical arm in a closed loop mode, so that the system can compensate and calibrate drift and accumulative errors in real time, and long-term stable operation is achieved.
Owner:NINGBO XINGBOYUAN INTELLIGENT TECHNOLOGY CO LTD

Photorealistic content generation from animated content by neural radiance field diffusion guided by vision-language models

A method implemented by a computing device. The method includes obtaining one or more of animated video content (X), a text prompt (Y), and view information (V); generating photorealistic two dimensional (2D) image frames ({circumflex over (X)}) based on the animated video content, the text prompt, and the view information using a vision-language model; and rendering photorealistic three dimensional (3D) image frames (X) based on the photorealistic 2D image frames using a 3D representation model.
Owner:HUAWEI TECH CO LTD

Method and system for determining the position and / or the orientation of tools of construction machines in a construction area

The invention relates to a method for determining the position and / or orientation of tools of construction machines in a construction area, comprising the steps of: attaching at least one stereo camera to at least one construction machine or to a movable carrier, such that the stereo camera can detect at least one detection region within the construction area, preferably a working region of the construction machine, providing at least one reference point, the position of which is known relative to a predefined 3D world coordinate system and which can be recognised in the images of the stereo camera, in the detection region of the at least one stereo camera, calibrating the at least one stereo camera on the basis of the detection of the at least one reference point and triangulation in order to ascertain a transformation between the 3D world coordinate system and a 2D image coordinate system of the at least one stereo camera, preparing a virtual digital 3D terrain model of the construction area by means of digital image processing, in particular stereoscopy, and image recognition on the basis of the detection of the construction area by means of the at least one stereo camera, and representing predetermined working positions and detected actual tool positions and / or actual tool inclinations of the tool of the at least one construction machine in the virtual digital 3D terrain model.
Owner:BAUER MASCH GMBH

Semantic information-based lunar surface target three-dimensional point cloud detection and positioning method and system

The invention discloses a lunar surface target three-dimensional point cloud detection and positioning method and system based on semantic information, relates to the technical field of three-dimensional object detection, and receives three-dimensional point cloud data and a two-dimensional image detection result, the two-dimensional image detection result is obtained by detecting acquired image data based on a pre-established two-dimensional target detection model, performing ground segmentation on the three-dimensional point cloud data to obtain non-ground point cloud data, mapping the non-ground point cloud data into a two-dimensional grid with a preset size to obtain a two-dimensional clustering result, and performing clustering on the non-ground point cloud data to obtain a clustering result; mapping the two-dimensional clustering result to the three-dimensional point cloud data to obtain a plurality of point cloud clusters; the points of each point cloud cluster are projected into an image coordinate system, dynamic weighted scores are calculated based on the number and confidence of the projection points of each point cloud cluster falling into the detection frame and the geometrical characteristics of the point cloud clusters, the point cloud cluster with the highest dynamic weighted score is selected for distance calculation, and the distance of the point cloud cluster with the highest dynamic weighted score is calculated; and obtaining a lunar surface target three-dimensional point cloud detection and positioning result.
Owner:DEEP SPACE EXPLORATION LABORATORY

Automatic rigging with 2d supervised learning

PendingUS20260080601A1AnimationMedicineAnimation
According to one aspect of the present disclosure, a method of training a deformation prediction model is provided. In some implementations, a method includes obtaining a neutral expression three-dimensional (3D) mesh and a set of facial action coding system (FACS) weights, wherein the set of FACS weights represent a target facial pose or a target facial expression. The method further includes obtaining a predicted 3D mesh from the deformation prediction model, wherein the predicted mesh is arranged to at least partially mimic the target facial pose or target facial expression, rendering a two-dimensional (2D) image from the predicted mesh, and adjusting the deformation prediction model based on one or more 2D loss functions, the one or more 2D loss functions being based on comparison of the 2D image with a groundtruth 2D image obtained from a pre-trained 2D animation model.
Owner:ROBLOX CORP

Method and system for detection of road objects using 2d image sign sightings

The disclosure provides a method, a system, and a computer program product for object detection using 2D sighting data of the object obtained using a 2D sensor. The method comprises obtaining 2D sighting data of the object using the 2D sensor. Further, the method comprises determining position candidate data for the object based on (i) 2D centroid data associated with the 2D sighting data of the object and (ii) 3D centroid data determined using projection data of one or more skew lines associated with the 2D centroid data. The position candidate data is then filtered based on (i) offset data associated with vector offset between the 2D centroid data and the 3D centroid data, (ii) postprocessing data, and (iii) scaling factor data associated with the position candidate data. Further, the detection data for the object is outputted based on the filtered position candidate data.
Owner:HERE GLOBAL BV

Joint semantic segmentation method for camera and laser radar in cross-country environment

The invention discloses a joint semantic segmentation method for a camera and a laser radar in an off-road environment, and relates to the technical field of computer vision. According to the method, the 2D image and the 3D point cloud data are respectively processed by constructing the double-branch basic segmentation network, so that efficient processing and fine segmentation of the point cloud are realized; a multi-scale one-way knowledge distillation framework is designed, feature fusion of a common-view area is realized through a channel self-attention mechanism, and knowledge migration from an image mode to a point cloud mode is realized by adopting a dynamic temperature adjustment and Logit standardization technology; in the training stage, multi-modal data is utilized, and high-precision segmentation can be completed only through point cloud input in the reasoning stage. According to the method, the semantic segmentation precision and robustness in the cross-country environment are remarkably improved, the segmentation performance of the occlusion region and the sparse feature region is effectively improved, and meanwhile, the computing resource consumption is greatly reduced.
Owner:CHONGQING UNIV

Method and apparatus for three-dimensional reconstruction of a scene, electronic device, and storage medium

The present application provides a method and apparatus for three-dimensional (3D) reconstruction of a scene, an electronic device, and a storage medium. The method includes: obtaining a two-dimensional (2D) image and 3D point cloud data of a scene to be reconstructed, where the 2D image is a foreground image including a specified object; identifying an object in the 2D image, obtaining 2D data of the object, and performing 3D reconstruction based on the 2D data to obtain first feature data; performing point cloud segmentation on the 3D point cloud data to obtain object 3D point cloud data and background 3D point cloud data; performing 3D reconstruction based on the object 3D point cloud data to obtain second feature data, and performing 3D reconstruction based on the background 3D point cloud data to obtain third feature data; fusing the first feature data and the second feature data to obtain fusion feature data; and rendering the third feature data and the fusion feature data into a virtual space based on a position relationship between the object and a background to realize 3D reconstruction of the scene to be reconstructed. The method can improve the accuracy of 3D reconstruction of the scene.
Owner:SAMSUNG ELECTRONICS CO LTD

A method for line of sight estimation based on 2D data

This invention provides a method for gaze estimation based on 2D data, belonging to the field of CV algorithms. On a single eye's 2D image data, 50 key points and the pupil center key point are labeled. Then, based on the top, bottom, left, and rightmost points from these 50 key points, the gaze point (gaze point) of the pupil center under normal eye-viewing conditions is estimated. The offset of the labeled gaze point from the current pupil center is calculated. This solves the difficulties in acquiring gaze estimation data and the tedious work of converting 3D data to 2D data. Thus, using 2D data achieves the effect of a model trained on 3D data, with low cost; the dataset can be created independently using only ordinary 2D image data.
Owner:SHENZHEN TIANSHUANG TECH CO LTD

Neural rendering for inverse graphics generation

Approaches are presented for training an inverse graphics network. An image synthesis network can generate training data for an inverse graphics network. In turn, the inverse graphics network can teach the synthesis network about the physical three-dimensional (3D) controls. Such an approach can provide for accurate 3D reconstruction of objects from 2D images using the trained inverse graphics network, while requiring little annotation of the provided training data. Such an approach can extract and disentangle 3D knowledge learned by generative models by utilizing differentiable renderers, enabling a disentangled generative model to function as a controllable 3D “neural renderer,” complementing traditional graphics renderers.
Owner:NVIDIA CORP

Simulation of viewpoint capture from environment rendered with ground truth heuristics

Aspects of this technical solution can generate, according to one or more first environment metrics, a three-dimensional (3D) model including a first surface corresponding to one or more physical ways through a physical environment, the one or more first environment metrics indicative of boundaries of the one or more physical ways, generate, according to one or more second environment metrics, one or more geometric two-dimensional (2D) objects on the first surface, the second environment metrics indicative of the one or more physical ways, identify, according to one or more viewpoint metrics indicative of cameras of a physical object configured to move along the one or more physical ways, one or more viewpoints oriented to capture corresponding portions of the 3D model, and render, from the one or more corresponding portions of the 3D model, one or more 2D images each corresponding to respective ones of the viewpoints.
Owner:TESLA INC

Target grabbing method and device and robot

The embodiment of the invention provides a target grabbing method and device and a robot. In the embodiment of the invention, the 2D image acquisition equipment such as a 2D camera is used for assisting the mechanical arm of the robot to realize target grabbing, and compared with a 3D camera, the hardware cost is effectively reduced. According to the embodiment of the invention, the first relative pose of the reference object coordinate system relative to the mechanical arm base coordinate system when the target object is in the reference pose and the teaching pose of the mechanical arm when the target object is in the reference pose are taken as the reference; and in combination with a second relative pose of the reference object coordinate system relative to the base coordinate system of the mechanical arm when the target object is in the non-reference pose, the target pose of the mechanical arm capable of successfully grabbing the target object in the non-reference pose can be determined, so that the target object in the non-reference pose can be grabbed, namely disordered grabbing is realized. In this way, the situation that the grabbing speed is affected by the scanning speed of the 3D camera can be avoided, and therefore the target grabbing efficiency is effectively improved.
Owner:HANGZHOU HIKROBOT TECH CO LTD

Image-based tooth identification using digital dental models

Systems and methods for identifying teeth in a patient image are provided. For example, a computer-implemented method can include, by one or more processors, accessing a 2D image including a depiction of a patient's teeth, where a plurality of the depicted patient's teeth are annotated with tooth identifiers according to a first tooth identification scheme. The computer-implemented method can further include transmitting, to a server computing device, the 2D image, and receiving from the server computing device, a second tooth identification scheme for the patient's teeth in the 2D image. The second tooth identification scheme can be generated by accessing a 3D model of the patient's teeth, projecting the 3D model onto the 2D image, comparing the projection to the 2D image to determine a probability parameter for each of one or more teeth in the 2D image, and determining, based on the probability parameters, the second tooth identification scheme.
Owner:ALIGN TECHNOLOGY INC

A low-cost high-efficiency three-dimensional face information acquisition system and method

The application discloses a low-cost and high-efficiency three-dimensional face information acquisition system and method, belongs to the technical field of three-dimensional data acquisition, and comprises the steps of determining a shooting position, arranging a left RGB camera, a right RGB camera and a middle RGB camera, arranging a left 3D module, arranging a right 3D module, jointly calibrating, obtaining internal parameters and external parameters, calibrating the left RGB camera, the right RGB camera and the middle RGB camera, collecting images, obtaining a plurality of 2D images and a plurality of 3D point clouds, performing face segmentation on each 2D image by using a semantic segmentation model, filtering background point clouds, performing face key point recognition on each 2D image, fusing point clouds, generating triangular mesh data, selecting an optimal view angle, pasting texture on the triangular mesh according to the selected view angle, and obtaining three-dimensional face information. The application has the effect of more rapidly and accurately collecting three-dimensional face information.
Owner:QINGDAO XIAOYOU INTELLIGENT TECH CO LTD

Pallet pose recognition method for unmanned forklift

The invention provides a pallet pose recognition method for an unmanned forklift, and the method comprises the steps: firstly obtaining depth point cloud data, and converting the depth point cloud data into 2D image data; then constructing a pallet coarse positioning model, processing the 2D image data through the pallet coarse positioning model, obtaining 2D bounding box data, and obtaining a pallet index matrix through the 2D bounding box data; then acquiring a plate instance point cloud based on the plate index matrix, and acquiring a left support leg point cloud, a middle support leg point cloud and a right support leg point cloud through a classification cutting strategy and the plate instance point cloud; according to the method, the 3D depth point cloud is converted into the 2D binary image, a complex 3D point cloud recognition problem is converted into a more efficient and mature 2D image target detection problem, and the template matching problem of a pure point cloud scheme is replaced by a lightweight neural network, so that the optimal balance between the speed and the precision is realized.
Owner:GUANGDONG JATEN ROBOT & AUTOMATION

Solder defect detection method, detection device, computer equipment and storage medium

ActiveCN116309397B3d imageEngineering
This application discloses a method, apparatus, computer equipment, and computer-readable storage medium for detecting solder defects. The method includes: acquiring an original 3D image of a solder area; converting the original 3D image into a 2D image; processing the 2D image to obtain outer edge information of the solder area; dividing the solder area into different detection zones in the original 3D image based on the outer edge information; determining whether solder defects exist within the detection zones; and if solder defects exist within the detection zones, determining the information of the solder defects. This method, by acquiring an original 3D image of a solder area, dividing the solder area based on the outer edge information of the 2D image converted from the original 3D image, and then determining the solder defect information in different zones, achieves convenient, rapid, and accurate location of defect locations, improves detection efficiency, and compared to manual inspection, requires less manpower, is lower in cost, and has higher reliability.
Owner:BEIJING LUSTER LIGHTTECH

Annotation of 3D models with signs of use visible in 2D images

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for annotation of 3D models with signs of use that are visible in 2D images. In one aspect, methods are performed by data processing apparatus. The methods can include projecting signs of use in a relatively larger field of view image of an instance of an object onto a 3D model of the object based on a pose of the instance in the relatively larger field of view image, and estimating a relative pose of the instance of the object in a relatively smaller field of view image based on matches between the signs of use in the relatively larger field of view image and the same signs of use in the relatively smaller field of view image.
Owner:INAIT SA

System and method for technical data protection using nerf models

PendingUS20260112106A13D modellingAlgorithmImaging data
A method for providing protection of technical data using a neural radiance field (NeRF) model includes providing a NeRF model, storing a representation of a three-dimensional (3D) model in the NeRF model, receiving a first instruction indicating a requested view of the 3D model, and generating, from the NeRF model and according to the requested view, two dimensional (2D) image data associated with the 3D model, wherein the 2D image data is generated with at least one portion of the 3D model in the requested view being obfuscated in the 2D image.
Owner:TEXTRON INNOVATIONS INC