Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

21 results about "Structure from motion" patented technology

Structure from motion (SfM) is a photogrammetric range imaging technique for estimating three-dimensional structures from two-dimensional image sequences that may be coupled with local motion signals. It is studied in the fields of computer vision and visual perception. In biological vision, SfM refers to the phenomenon by which humans (and other living creatures) can recover 3D structure from the projected 2D (retinal) motion field of a moving object or scene.

Multi-view three-dimensional reconstruction method for slender object based on curve guidance

The invention discloses a multi-view three-dimensional reconstruction method for a slender object based on curve guidance. According to the method, geometric prior and a deep learning segmentation network are fused, two-stage foreground extraction, a curve-guided motion recovery structure (SfM) and a surface optimization module capable of differential Poisson reconstruction are provided, and a set of complete processing flow is established for the problem that a slender and weak-texture object is difficult to accurately reconstruct. The method comprises the following steps: acquiring a multi-view image through a consumer-level terminal and performing frame extraction to obtain an input data set; generating a high-quality foreground mask by utilizing depth estimation and semantic segmentation; performing curve-guided SfM initialization based on the foreground skeleton curve to realize joint estimation of the camera pose and the sparse three-dimensional curve; carrying out surface geometric optimization through a micro Poisson equation and differentiable grid rendering; and designing a multi-branch loss function fusing luminosity, mask, regularization and Gaussian rendering supervision to train the model. Finally, the method can output a three-dimensional reconstruction result of the slender structure with high geometric accuracy, strong structural integrity and smooth surface, and has good robustness and application value.
Owner:NORTH CHINA ELECTRIC POWER UNIV

Method, apparatus, and storage medium for three-dimensional reconstruction of buildings based on missing point cloud data

ActiveUS12567208B2Image enhancementImage analysisPoint cloudStructure from motion
The invention provides a method, apparatus, and storage medium for reconstructing three-dimensional models of buildings based on missing point cloud data. The method includes integrating image-based point cloud generation, neural network techniques, and skeleton line extraction methods, offering a novel approach to handling missing point cloud data. The generation of point cloud data is achieved using principles of Structure from Motion based on video or panoramic image data. The point cloud is sampled and segmented using a region growing algorithm. A neural network based on PointNet is constructed, utilizing cross-entropy loss functions to assess the missing points in the point cloud. For mapping high-confidence point clouds from sampled points, Truth Points is employed to complete the entire process of real-world three-dimensional reconstruction. The integration of images into the three-dimensional scene is achieved with strict geometric relationships.
Owner:WUHAN UNIV

Text-to-three-dimensional surface generation method and system based on two-dimensional Gaussian surface element

PendingCN121259176A3D-image rendering3D modellingGeometric consistencyStructure from motion
The invention discloses a text-to-three-dimensional surface generation method and system based on a two-dimensional Gaussian surface element, and belongs to the technical field of three-dimensional reconstruction, and the method comprises the steps: representing a three-dimensional object surface as a Gaussian surface element on a local tangent plane; receiving a text description, generating normal diagrams and texture diagrams of a plurality of visual angles based on a pre-training text-to-image diffusion model, and generating a preliminary background mask at the same time; determining the center of sphere and the maximum radius according to the extrinsic parameters of the camera, and randomly generating Gaussian surface elements in the spherical range; removing the surface elements falling in the preliminary background mask area; on the basis of a two-dimensional Gaussian dot drawing renderer, the normal graph and the texture graph are rendered from multiple perspectives, and curvature consistency regular loss and surface convergence constraint loss are calculated; geometric and texture parameters of Gaussian surface elements are optimized through back propagation, iteration is carried out until convergence, and a final three-dimensional model is output. According to the method, the high-fidelity three-dimensional surface can be generated from the text without SfM (Structural Recovery Motion) initialization, and the method has relatively high geometric consistency and generation speed.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Systems and methods for reconstructing 3D objects

PendingUS20260094291A1Image enhancementBronchoscopesHigh-resolution computed tomographySubglottic hemangioma
Systems and methods are disclosed herein for using structure from motion (SfM) techniques to reconstruct three-dimensional (3D) surface models of tubular patient anatomies, such as the larynx and trachea, from clinical endoscopy videos. The disclosed methods may improve understanding of upper airway disease including vocal fold paralysis, laryngeal cancer, subglottic hemangiomas, subglottic stenosis, tracheal stenosis, tracheal cartilaginous sleeves, complete tracheal rings and tracheomalacia, and allow for quantitative analysis of complex laryngotracheal geometries. Quantitative measures of airway caliber and shape, which are critical for diagnostic purposes, may be obtained using the disclosed methods as a cost-effective and radiation-free alternative to relying on imaging studies. Results have demonstrated excellent resolution of reconstructions, when compared to high-resolution computed tomography (CT) scans (surface errors <0.300 mm).
Owner:UNIV OF WASHINGTON +2

Structure from motion enhancements using generalized camera model and motion parametrization

PCT designated stageWO2026101640A1Image enhancementImage analysisPattern recognitionStructure from motion
This disclosure provides systems, methods, and devices that utilize machine learning models to determine corresponding spatial positions and motion trajectories for images. In one aspect, a method is provided that includes receiving an image of a scene captured by a camera; determining, with a first machine learning model, a plurality of position values relative to the camera for at least a subset of pixels within the image; and training a second machine learning model based on the determined position values. The method further includes determining basis trajectories based on movement of the positions relative to previous image frames, and determining, for each of a subset of pixels, a movement trajectory relative to the previous frames as a weighted combination of the basis trajectories. These techniques can be employed as part of a structure-from-motion pipeline to estimate multi-view three-dimensional geometry, camera poses, and per-pixel object motion. Additional aspects are provided.
Owner:QUALCOMM INC

Robust motion from structure and structure from motion in videos

Systems, methods, and computer program code for motion segmentation using a motion from structure approach; and related image processing tasks. Structure-from-motion systems are also described. In some implementations feature representations for pixels of an image frame are used to extrapolate from high reliability regions to lower reliability regions, based on semantic similarity as represented by the feature representations. This can link different visible parts of the same object, e.g. for generating a segmentation map for moving objects.
Owner:DEEPMIND TECH LTD +1

Three-dimensional variable perspective radiance field viewing

PendingUS20260099990A1Image analysisImage generationPoint cloudStructure from motion
Embodiments of the present disclosure include a computer-implemented method for displaying three-dimensional viewing content based on a user perspective, the method comprising: receiving point cloud data by one or more structure from motion tools; creating a Gaussian splat scene by associating Gaussian splats with the point cloud data; generating two-dimensional raster images from the Gaussian splat scene; deriving a view matrix from a user position in a three-dimensional coordinate space, the view matrix comprising vectors indicating the user position within the Gaussian splat scene; determining a projection matrix comprising vectors for a transform of the three-dimensional coordinate space to two-dimensional screen space coordinates for display on a user interface; and projecting the one or more first two-dimensional raster images in a position in the two-dimensional screen space determined by matrix multiplication of the view matrix and projection matrix.
Owner:D3LABS INC

Method and apparatus for 3-d auto tagging

A multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. Selectable tags can be placed at locations on the object in the MVIDMR. When the selectable tags are selected, media content can be output which shows details of the object at location where the selectable tag is placed. A machine learning algorithm can be used to automatically recognize landmarks on the object in the frames of the MVIDMR and a structure from motion calculation can be used to determine 3-D positions associated with the landmarks. A 3-D skeleton associated with the object can be assembled from the 3-D positions and projected into the frames associated with the MVIDMR. The 3-D skeleton can be used to determine the selectable tag locations in the frames of the MVIDMR of the object.
Owner:FUSION INC

Scaffold engineering compliance intelligent evaluation method based on three-dimensional reconstruction technology

PendingCN121366347AThree-dimensional object recognitionStructure from motionFitting algorithm
The invention discloses a scaffold engineering compliance intelligent evaluation method based on a three-dimensional reconstruction technology. According to the scheme, an unmanned aerial vehicle aerial photography sampling plan is made according to scaffold engineering characteristics, a machine vision model is trained to detect scaffold key components, and the distribution compliance of the scaffold system key components is evaluated according to the types and the number of the key components. A scaffold three-dimensional point cloud model is generated by adopting an incremental SfM and a multi-view stereo matching algorithm MVS based on an aerial image, a point cloud model of each rod piece is obtained through segmentation, and a central characteristic line equation of the rod piece is further extracted through a cylinder fitting algorithm. According to relevant specifications and standards of scaffolds, a scaffold engineering compliance set is summarized; and calculating the spatial relationship of the characteristic line equation, and comparing the scaffold compliance set to evaluate the construction compliance of the scaffold engineering of the project under construction. According to the method, the unmanned aerial vehicle, intelligent image recognition, three-dimensional reconstruction and feature line extraction technologies are comprehensively used, and autonomous and efficient evaluation of the safety compliance of the floor type fastener scaffold is achieved.
Owner:FUJIAN UNIV OF TECH

A Gaussian splashing method for dynamic scenes based on spatiotemporal motion distillation

PendingCN122312913APattern recognitionMorphing
This invention discloses a Gaussian splashing method for dynamic scenes based on spatiotemporal motion distillation. It outputs a set of sparse point clouds from images at different viewpoints and times through a motion inference structure. The sparse point clouds are used to initialize Gaussian point attributes. The initialized Gaussian point set undergoes a fixed number of pre-training iterations to obtain a standard spatial set. Learnable motion feature representations are introduced into the Gaussian point attributes to explicitly model the spatiotemporal motion of the Gaussian points. Motion anchor points are extracted by distilling the motion information of the Gaussian points. During the iteration process, an adaptive density control mechanism continuously updates the density distribution of the Gaussian points, outputting the corresponding Gaussian model and deformation field weights for subsequent rendering evaluation. This invention fully utilizes the spatiotemporal characteristics of dynamic objects to efficiently reconstruct and render dynamic scenes, effectively solving the artifact and noise problems that occur during dynamic scene rendering, thereby outputting high-quality rendered images.
Owner:ANHUI UNIV

A plant phenotype-oriented three-dimensional reconstruction method and acquisition system

The application discloses a plant phenotype-oriented three-dimensional reconstruction method and a collection system. The method first acquires multi-view image data of a plant through a 360-degree ring collection system integrated with an RGB camera and a depth camera; then generates an accurate foreground mask by using a foreground semantic segmentation model; then combines structure from motion (SfM) and depth point cloud, and performs fusion under the guidance of the foreground mask to obtain an initial point cloud; then performs mask-guided 3D Gaussian splashing (3DGS) optimization reconstruction based on the point cloud, and the optimization process improves the identification of small plant organs by using a mask weighted loss function and a semantic-guided density control strategy; and finally derives a color point cloud from the optimized Gaussian model, which can be directly used for extraction of phenotype parameters such as plant height and leaf area. The application solves the problems of large noise and detail loss in the prior art when reconstructing plants in a complex background, and realizes fast three-dimensional reconstruction which can be directly used for phenotype analysis.
Owner:SHIHEZI UNIVERSITY +1

3-D reconstruction using augmented reality frameworks

ActiveUS12462514B2Image enhancementImage analysisStructure from motionComputer graphics (images)
System and method are provided for scaling a 3-D representation of a building structure. The method includes obtaining images of the building structure, including non-camera anchors. The method also includes identifying reference poses for images based on the non-camera anchors. The method also includes obtaining world map data including real-world poses for the images. The method also includes selecting candidate poses from the real-world poses based on corresponding reference poses. The method also includes calculating a scaling factor for a 3-D representation of the building structure based on correlating the reference poses with the selected candidate poses. Some implementations use structure from motion techniques or LiDAR, in addition to augmented reality frameworks, for scaling the 3-D representations of the building structure. In some implementations, the world map data includes environmental data, such as illumination data, and the method includes generating or displaying the 3-D representation.
Owner:HOVER INC

Structure from motion enhancements using generalized camera model and motion parametrization

PendingUS20260127750A1Image enhancementImage analysisPattern recognitionStructure from motion
This disclosure provides systems, methods, and devices that utilize machine learning models to determine corresponding spatial positions and motion trajectories for images. In one aspect, a method is provided that includes receiving an image of a scene captured by a camera; determining, with a first machine learning model, a plurality of position values relative to the camera for at least a subset of pixels within the image; and training a second machine learning model based on the determined position values. The method further includes determining basis trajectories based on movement of the positions relative to previous image frames, and determining, for each of a subset of pixels, a movement trajectory relative to the previous frames as a weighted combination of the basis trajectories. These techniques can be employed as part of a structure-from-motion pipeline to estimate multi-view three-dimensional geometry, camera poses, and per-pixel object motion. Additional aspects are provided.
Owner:QUALCOMM INC

Bridge damage positioning and quantifying method based on unmanned aerial vehicle panoramic unfolding and TCFormer driving

The invention discloses a bridge damage positioning and quantifying method based on unmanned aerial vehicle panoramic unfolding and TCFormer driving, and the method comprises the steps: 1) unmanned aerial vehicle image collection: employing a matrix type flight path planning strategy, and dividing a bridge bottom into a plurality of independent collection regions; 2) panorama construction: performing single-component three-dimensional reconstruction on the acquired image through a motion recovery structure SfM and a multi-view stereoscopic vision MVS technology; 3) performing multi-class damage segmentation: performing semantic segmentation on the panoramic image cutting area based on a lightweight TCFormer model; 4) component-level positioning: establishing a standardized component coordinate system and an index system, and dividing the components into five types; according to the method, a health evaluation system including PMCI single component scoring and PCCI whole-span comprehensive scoring is constructed, full-process automation from image acquisition, damage identification, spatial positioning, geometric quantification to health evaluation is achieved, and the efficiency, precision and standardization level of bridge detection are improved.
Owner:YANGZHOU UNIV

Systems and methods for pose estimation of a fluoroscopic imaging device and for three-dimensional imaging of body structures

Imaging systems and methods estimate poses of a fluoroscopic imaging device, which may be used to reconstruct 3D volumetric data of a target area, based on a sequence of fluoroscopic images of a medical device or points, e.g., radiopaque markers, on the medical device captured by performing a fluoroscopic sweep. The systems and methods may identify and track the points along a length of the medical device appearing in the captured fluoroscopic images. The 3D coordinates of the points may be obtained, for example, from electromagnetic sensors or by performing a structure from motion method on the captured fluoroscopic images. In other aspects, a 3D shape of the catheter is determined, then the angle at which the 3D catheter projects onto the 2D catheter in each captured fluoroscopic image is found.
Owner:COVIDIEN LP

Highway subgrade receiver collaborative metering system based on live-action three-dimensional model

PendingCN121962227AImplement deep bindingconvenient amountDetails involving processing stepsImage enhancementFeature extractionStructure from motion
The invention relates to the technical field of engineering metering, in particular to a highway subgrade receiver collaborative metering system based on a live-action three-dimensional model, and the system comprises a data collection module which is used for the data collection of multi-source equipment, so as to obtain subgrade basic data; the data preprocessing module carries out standardization processing on the roadbed basic data to obtain preprocessed data; the live-action three-dimensional model generation module generates a roadbed model through an SfM algorithm and an MVS algorithm; the metering platform performs feature extraction on the roadbed model to obtain point location engineering quantity metering data, length engineering quantity metering data and area engineering quantity metering data; and the report module is used for acquiring data of the metering platform to generate a project quantity report. According to the highway subgrade receiver collaborative metering system based on the live-action three-dimensional model, point location, linearity and area engineering quantity metering can be integrated, and the problems that a traditional metering method is low in precision and low in efficiency are solved.
Owner:GUANGXI ROAD & BRIDGE ENG GRP CO LTD

Determining a location of a shopping container within a store

PCT designated stageWO2026028019A1Image enhancementImage analysisStructure from motionComputer graphics (images)
There is provided a method for determining a location of a shopping container within a store, the method includes (i) acquiring by a first side camera associated with the shopping container, a side first side image of first content located to a first side of the shopping container; (ii) determining a pose of the first side camera, based on the first side image and by a machine learning process trained using a structure from motion model of the store, wherein the structure from motion model is generated from untagged side images acquired by side cameras of shopping containers that moved within the store during an image acquisition process, wherein the untagged side images are associated with side cameras pose information learnt during the generation of the structure from motion model; and (iii) determining the location of the shopping container based on the pose of the first camera.
Owner:SHOPIC TECH LTD

Automatic monitoring method for earthwork volume

PendingCN121861519AImage enhancementImage analysisStructure from motionData acquisition
The invention relates to the technical field of earthwork, in particular to an automatic earthwork volume monitoring method, which comprises the following steps: acquiring a high-overlapping-degree unmanned aerial vehicle aerial image sequence of a construction area; identifying the engineering vehicle in the image by using an instance segmentation model and generating a mask; after the mask is subjected to morphological expansion, an image restoration model is adopted to remove vehicle pixels and restore the terrain, and an image sequence without vehicle interference is generated; constructing a digital terrain surface model based on the sequence through a motion recovery structure and a multi-view stereoscopic vision technology; and calculating the earthwork volume variation by comparing spatial differences between different time point models. Interference of moving vehicles is eliminated from an image source, automation of the whole process from data collection to result output is achieved, and the monitoring precision and efficiency are remarkably improved.
Owner:KUNMING UNIV OF SCI & TECH

Methods for installing control points and taking photographs for ground photogrammetry used for three-dimensional measurements

PendingJP2025178991APicture interpretationStructure from motionPoint cloud
To provide methods for installing control points and taking photographs for realizing simplification, acceleration, and labor saving in a cross-sectional view creation method by three-dimensional measurements using photogrammetry at a collapse site necessary for assessing a small-scale collapse site for a disaster.SOLUTION: In photogrammetry called SfM (Structure from Motion) or the like that has been put into practical use in recent years, it is possible to automatically express, as a three-dimensional point cloud shape, a subject appearing in an overlapping portion of photographs captured in an overlapping manner from different positions. Provided is a simple working method capable of eliminating surveying required for installation of a control point and installing a multi-axis coordinate ruler by one person by using a scale engraved on the multi-axis coordinate ruler provided with a plurality of axes for the control point required for correctly giving a scale and horizontal and vertical directions to a three-dimensional point group shape. Provided is a method capable of automatically restoring a three-dimensional point group shape of a photograph photographed in an overlapping manner, and measuring a base line length capable of obtaining an appropriate overlapping degree by gait.SELECTED DRAWING: Figure 2
Owner:株式会社北斗測量設計社

Irregular large area uniform block adjustment positioning method and device of unmanned aerial vehicle image

PendingCN122156306AImage analysisStructure from motionComputer graphics (images)
Embodiments of the present disclosure provide an unmanned aerial vehicle image block adjustment positioning method and device for irregular large area uniform block. The method comprises: acquiring a set of unmanned aerial vehicle images and POS data; acquiring geographical range information of a shooting area of the set of unmanned aerial vehicle images according to the POS data; performing block on the set of unmanned aerial vehicle images based on the geographical range information of the shooting area to obtain a set of unmanned aerial vehicle image blocks; performing pose estimation on each unmanned aerial vehicle image block in the set of unmanned aerial vehicle image blocks according to a structure from motion (SfM) algorithm to obtain a set of unmanned aerial vehicle image block information; performing global coordinate conversion on the set of unmanned aerial vehicle image block information to obtain a set of rough unmanned aerial vehicle image block information in a global coordinate system; and performing global block adjustment calculation on the set of rough unmanned aerial vehicle image block information to obtain final global three-dimensional coordinates of a physical point and global unmanned aerial vehicle image pose parameter information. The present application reduces the problem of block adjustment failure in pose solution caused by error accumulation through an adaptive block method.
Owner:PERCEPTION WORLD (BEIJING) INFORMATION TECH CO LTD

Method and apparatus for 3-D auto tagging

A multi-view interactive digital media representation (MVIDMR) of an object can be generated from live images of an object captured from a camera. Selectable tags can be placed at locations on the object in the MVIDMR. When the selectable tags are selected, media content can be output which shows details of the object at location where the selectable tag is placed. A machine learning algorithm can be used to automatically recognize landmarks on the object in the frames of the MVIDMR and a structure from motion calculation can be used to determine 3-D positions associated with the landmarks. A 3-D skeleton associated with the object can be assembled from the 3-D positions and projected into the frames associated with the MVIDMR. The 3-D skeleton can be used to determine the selectable tag locations in the frames of the MVIDMR of the object.
Owner:FUSION INC