Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

201 results about "View synthesis" patented technology

Currently a study branch of Computer Science Research, aims to create new views of a specific subject starting from a number of pictures taken from given point of views. Vision Research and Artificial Intelligence fields are involved in the definition of suitable approaches to the problem.

System and method for efficient scene continuity in visual and multimedia using generative artificial intelligence

A system and method for generating multimedia artifacts with managed scene continuity in visual and multimedia using an AI-based and scene continuity aware media generation platform. The system receives a user or AI agent specification or simulation result(s), selects or trains generative models based on the specification, preprocesses relevant data, and generates scene narrative or frame-specific, sequence specific or broader continuity aware content using the selected or trained model(s). The generated content may be further enhanced using frame interpolation and view synthesis techniques to create smooth transitions or novel viewpoints or to aid in more efficient transmission or viewing or persistence of resultant content. The system enables efficient and customizable generation of high-quality scene continuity aware content for various applications in visual and multimedia production using neuro-symbolic and simulation enhanced compression, representation and generation processes.
Owner:QOMPLX INC

Virtual Walkthrough Experience Generation Based on Neural Radiance Field Model Renderings

Systems and methods for generating and providing a virtual walkthrough interface can include generating a virtual walkthrough video based on view synthesis renderings generated by neural radiance field model. The neural radiance field model can be trained based on a plurality of images of an environment and may generate the view synthesis renderings based on processing positions along a determined walkthrough path. The generated virtual walkthrough video can then be scrubbed through to provide the virtual walkthrough interface.
Owner:GOOGLE LLC

Three-dimensional scene reconstruction method, electronic equipment, storage medium and program product

The invention discloses a three-dimensional scene reconstruction method, electronic equipment, a storage medium and a program product, and relates to the technical field of data processing, and the method comprises the steps: inputting each frame of multi-view image into a depth model, and obtaining a prediction depth map of each frame of multi-view image; for any target frame image in each frame of multi-view image, optimizing a local 3D Gaussian model and camera attitude information of the target frame image based on the predicted depth map of the target frame image and the depth information of the previous frame image to obtain an updated local 3D Gaussian model and updated camera attitude information of the target frame image; and optimizing the global 3D Gaussian model to obtain an updated global 3D Gaussian model based on the updated local 3D Gaussian model and the updated camera attitude information of each frame of multi-view image, and performing three-dimensional scene reconstruction based on the updated global 3D Gaussian model. According to the invention, the accuracy of a view synthesis result based on three-dimensional scene reconstruction is improved.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Real-time dynamic three-dimensional reconstruction method and system based on fusion of patch matching and lightweight monocular depth estimation

The invention discloses a real-time dynamic three-dimensional reconstruction method and system based on patch matching and lightweight monocular depth estimation fusion, and the method comprises the steps: carrying out the PatchMatch matching of an input multi-view image pair, obtaining sparse depth information, and fusing the sparse depth information with a monocular depth estimation result; and the fused depth map is used for constructing a three-dimensional Gaussian point cloud, and high-quality three-dimensional scene reconstruction and new view rendering are realized. In the whole process, an end-to-end differentiable training framework is adopted, and the reconstruction precision and rendering efficiency of the 3D Gaussian point cloud in a dynamic scene are improved by optimizing a pixel-level Gaussian parameter regression network and a lightweight depth estimation model. The method can complete training and depth inference in a short time, is suitable for real-time or near-real-time application scenes, can be widely applied to the fields of virtual reality, augmented reality, film and television production and the like, and is particularly suitable for efficient three-dimensional reconstruction and new view angle synthesis of a target in a multi-view-angle dynamic scene.
Owner:NANJING UNIV

Non-static scene reconstruction method and system based on multi-modal occlusion perception scoring

The invention discloses a non-static scene reconstruction method and system based on multi-modal occlusion perception scoring. The method comprises the following steps: generating a geometric consistency distribution diagram, static feature points and geometric prior masks through three-dimensional reconstruction of a multi-view image; fusing the geometric prior mask and a semantic segmentation model to extract a semantic mask and a fused semantic feature map; guiding the image segmentation model to generate candidate masks based on static feature point positive point prompt and occlusion area negative frame prompt, and optimizing the consistency by using a grid complementary fusion method; constructing a multi-modal shielding scoring module, and fusing the multi-source features to output a binary static mask; and utilizing a static mask to constrain neural radiation field training, inhibiting dynamic interference and optimizing static scene reconstruction. According to the method, the static region is sensed cooperatively through multi-modal information, the robustness and accuracy of mask generation are improved, the interference of dynamic elements on neural radiation field modeling is effectively inhibited, and high-quality three-dimensional image reconstruction and new view synthesis of a non-static scene are realized.
Owner:HANGZHOU DIANZI UNIV

3D scene content generation using 2d inpainting diffusion

Provided is a general approach to inpainting 3D content by using a 2D inpainting diffusion model trained on static scenes as a generative prior. In particular, while existing methods for 3D inpainting focus on the specific task of object removal, systems and methods of the present disclosure aim to generate realistic content in any masked 3D region while preserving the complexity of the original scene at the object and scene scale. The proposed techniques can be used in various applications such as 3D object removal, 3D inpainting and novel view synthesis.
Owner:GOOGLE LLC

Robot indoor navigation method and system

The invention provides a robot indoor navigation method and system, and the method comprises the steps: obtaining an RGB image frame sequence collected by a camera, and estimating a camera pose and a sparse three-dimensional landmark point of an image frame through a depth block visual odometer; based on the camera pose and the sparse three-dimensional landmark point, dense geometric enhancement is carried out on the image frame through a deep neural network of a transformer architecture to obtain dense geometric prior information; training a neural radiation field according to the initial camera pose, the dense geometric prior information and the image frame data, and performing joint optimization on neural radiation field parameters and the camera pose based on luminosity consistency, geometric consistency and camera pose constraints; based on image frames collected in real time, novel view synthesis is carried out through the optimized nerve radiation field, and robot navigation is carried out based on a novel view. According to the scheme, accurate estimation of the camera pose and high-quality reconstruction of the three-dimensional scene can be realized, the computing resource overhead is low, and the real-time performance can be guaranteed.
Owner:HUAZHONG UNIV OF SCI & TECH

View synthesis using camera poses learned from a video

View synthesis is a computer graphics process that generates a new image of a scene from a novel (previously unseen) viewpoint of the scene. Typically, the graphics process relies on a machine learning model that has been trained with ground truth pose information. Since ground truth pose information is not readily available, some solutions rely on a Structure-from-Motion (SfM) library COLMAP to generate pose information for a given image. However, this pre-processing step is not only time-consuming but also can fail due to its sensitivity to feature extraction errors and difficulties in handling texture-less or repetitive regions. The present disclosure provides view synthesis from learned camera poses without relying on SfM pre-processing.
Owner:NVIDIA CORP

Integrated visual angle synthesis and sparse visual angle CT reconstruction method based on 3DGS

The invention discloses a 3DGS-based integrated visual angle synthesis and sparse visual angle CT reconstruction method, which comprises the following steps: establishing an internal and external parameter matrix through X-ray scanning parameters, constructing an initial voxel space and an initialized 3DGS radiation field by combining sparse visual angle projection data and based on an ACUI strategy, modeling anisotropic radiation intensity by utilizing a radiation intensity response function, and reconstructing a 3D visual angle CT model. And generating a virtual projection by adopting volume consistency rendering. And the optimized radiation field is converted into a voxel grid, density contribution is applied to carry out voxel space reconstruction, a reconstructed voxel space is obtained, and a new 3DGS radiation field is generated. And re-rendering the virtual visual angle projection based on the new 3DGS radiation field, performing iterative optimization by using a loss function, and outputting a new visual angle synthetic image and a CT reconstruction body until a preset number of iterations is reached. And an efficient process of one-time training and dual output is realized. The efficiency and precision of X-ray imaging are improved, the radiation dose of a patient is reduced, and meanwhile, a higher-quality imaging solution is provided for medical diagnosis.
Owner:GUANGDONG UNIV OF TECH

Layered densification Gaussian sputtering method based on visibility

The invention discloses a layered densification Gaussian sputtering method based on visibility, and relates to the technical field of artificial intelligence and computer vision. According to the scheme, an initial three-dimensional Gaussian primitive set is generated based on sparse multi-view observation data, and initial scene representation is established by extracting spatial distribution parameters, morphological parameters and radiation parameters; performing fusion analysis on the Gaussian primitives based on the multi-dimensional observability parameter set to obtain comprehensive observability index data; executing hierarchical clustering according to the index data, and constructing a multi-layer Gaussian structure of a significant layer, a transition layer and a background layer; performing density enhancement, geometric continuity constraint and parameter update processing on different levels of Gaussian structures, and generating a rendered image of a target view angle under a volume light traveling and transparency hybrid mechanism; according to the method, continuous reconstruction of a scene structure and accurate expression of radiation characteristics can be realized under the sparse view condition, and the geometric fidelity and rendering consistency of new view angle synthesis are improved.
Owner:HENAN JINSHU INTELLIGENT TECH CO LTD

Sparse dynamic 3D Gaussian splash method based on global-local feature extraction

The invention discloses a sparse dynamic 3D Gaussian splash method based on global-local feature extraction. According to the method, a global-local feature extraction module, an inter-frame feature flow aggregation network and depth prior are combined with a 4D pseudo pose supervision mechanism for cooperative work, and high-quality dynamic scene reconstruction and rendering are realized under the condition of a sparse camera visual angle. According to the method, the multi-view image information is compressed to the feature space for processing, so that the calculation complexity is reduced, and the geometric structure and the dynamic features of the scene are effectively reserved; the fused depth prior information and the 4D pseudo pose provide accurate geometric constraints, and the inter-frame feature flow aggregation network optimizes the time consistency of the dynamic scene. The method is particularly suitable for processing a complex dynamic scene under a sparse view angle, the sense of reality and the real-time performance of virtual view synthesis are remarkably improved, and a dynamic view with a more accurate geometric structure, higher time continuity and richer details can be generated.
Owner:HANGZHOU DIANZI UNIV

Target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment

The invention discloses a target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment. The target-level three-dimensional point cloud cross-modal semantic retrieval method comprises the steps that independent three-dimensional target point clouds are segmented from three-dimensional scene point clouds, and unique identifiers are given to the independent three-dimensional target point clouds; projecting each target point cloud to three orthogonal two-dimensional observation planes to generate a multi-view composite image; a pre-trained vision-language basic model is adopted to code and fuse the synthesized image, and a unified multi-modal feature vector is generated; a vector database in which the feature vectors are associated with their identifiers is constructed. And encoding a natural language query text into a text query vector by using the model, calculating the semantic similarity between the text query vector and the feature vector in the database, and positioning and returning the corresponding three-dimensional target point cloud. According to the method, accurate and efficient cross-modal retrieval from a natural language to a three-dimensional point cloud target is realized, the problem that a traditional method is difficult to support semantic fine-grained retrieval of the three-dimensional target point cloud is solved, and a key technical support is provided for training data management in the fields of intelligence and the like.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Three-dimensional scene reconstruction method without camera pose based on 3DGS

The invention discloses a camera-free pose three-dimensional scene reconstruction method based on 3DGS, and the method comprises the steps: firstly obtaining an ordered picture sequence shot by a camera and a depth map of the ordered picture sequence, constructing a picture, and pre-training a feature encoder for a data set; then initializing a camera pose of a first frame of picture in the ordered picture sequence and an empty training picture set, recovering point cloud information and initializing a 3D Gaussian model; randomly selecting a color image and a depth image of the image from the training image set as supervision information to iteratively optimize the 3D Gaussian model; iteratively optimizing the camera pose of the next frame of picture in the ordered picture sequence according to the optimized 3D Gaussian model; and finally, according to a picture sequence in the ordered picture sequence, alternately iteratively optimizing the 3D Gaussian model and optimizing the camera pose of a new frame of picture by adopting a progressive training mode to realize three-dimensional scene reconstruction. According to the method, three-dimensional reconstruction with higher quality can be realized, the quality of a picture synthesized by a new view angle is improved, and the error of camera pose estimation is reduced.
Owner:SOUTH CHINA UNIV OF TECH

Three-dimensional Gaussian splashing and rasterization pipeline fused view synthesis system and method

The invention discloses a three-dimensional Gaussian splashing and rasterization pipeline fused view synthesis system and method. The method comprises the following steps of: 1, acquiring two-dimensional Gaussian primitives; 2, converting the two-dimensional Gaussian primitive into two triangular grids, and converting the two triangular grids into a world coordinate system; 3, baking the spherical harmonic coefficient of the color and the opacity data in a texture form; 4, enabling each Gaussian primitive to correspond to a texture block on the texture image set, enabling a triangle to cover the texture block in a UV space, and obtaining a Gaussian weight; 5, based on the obtained Gaussian weight, simulating an alpha mixing behavior in the three-dimensional Gaussian splashing pipeline by using a time-based jitter anti-aliasing method; and 6, training a symbol vector field SDF branch while training, and mutually supervising and constraining the output of the symbol vector field and the output of the two-dimensional Gaussian. The three-dimensional Gaussian splash can be rendered in a traditional rasterization pipeline of the unreal engine 4, and the fidelity equivalent to that of a special training pipeline is kept.
Owner:XIDIAN UNIV +1

Small sample new view synthesis method based on reinitialized three-dimensional Gaussian splashing

The invention discloses a small sample new view synthesis method based on reinitialization three-dimensional Gaussian splashing. The method comprises the following steps: firstly, acquiring sparse three-dimensional point cloud and camera internal and external parameters from a training view through a motion recovery structure algorithm; then, sampling points are generated in the point cloud bounding box by adopting a spatial expansion hybrid sampling strategy, and a coarse-grained Gaussian set is constructed and optimized; thirdly, obtaining a rendered image through Gaussian initialization and splash rendering, calculating pixel importance through depth errors and transmissivity, generating a fine-grained Gaussian set through back projection, and optimizing the fine-grained Gaussian set; and finally, calculating a sampling probability based on the cross-view contribution degree, screening key Gaussian distribution, and optimizing to form a final Gaussian set, thereby realizing high-quality new view synthesis. According to the method, the problem of sparsity difference is solved by eliminating extended view dependence, the multi-view consistency and local geometric details of a new view scene are improved, and high-quality new view synthesis of sparse data is realized.
Owner:ZHEJIANG UNIV

3D video generation method, 3D video viewing method, and electronic device

Provided in the embodiments of the present application are a 3D video generation method and an electronic device. The method is applied to the electronic device. The method comprises: on the basis of captured data, acquiring 3D video material data, wherein the captured data comprises a 2D video and / or a 2D image; on the basis of the 3D video material data, analyzing a 3D video scene type; on the basis of the 3D video scene type, acquiring a new-view generation model corresponding to the 3D video scene type; and using the new-view generation model corresponding to the 3D video scene type to generate 3D video data on the basis of the 3D video material data. On the basis of the method of the embodiments of the present application, at a 3D video data generation stage, a corresponding new-view generation model is called on the basis of a 3D video scene type, and thus a more complete scene synthesis view can be acquired on the basis of a new-view synthesis algorithm that uses deep learning and artificial intelligence, such that when 3D video data is played, a user can immersively view 3D video content from different angles, thereby providing a stronger sense of 3D and richer 3D video content.
Owner:HUAWEI TECH CO LTD

Virtual walkthrough experience generation based on neural radiance field model renderings

Systems and methods for generating and providing a virtual walkthrough interface can include generating a virtual walkthrough video based on view synthesis renderings generated by neural radiance field model. The neural radiance field model can be trained based on a plurality of images of an environment and may generate the view synthesis renderings based on processing positions along a determined walkthrough path. The generated virtual walkthrough video can then be scrubbed through to provide the virtual walkthrough interface.
Owner:GOOGLE LLC

Method and Device for Learning Depth Estimation Based on View Synthesis

A method for controlling autonomous driving of a vehicle is introduced. The method may comprise, training, based on an inference depth and an inference pose, a synthetic image model for generating a synthetic image, generating, based on the synthetic image, a first virtual image to be associated with the original image, generating, based on the original image, a second virtual image, training a generative adversarial network (GAN) for determining, based on the original image, authenticity of the first virtual image and the second virtual image, training, based on the trained GAN, a depth network, wherein the trained GAN outputs a determination of the authenticity of the first virtual image, outputting, based on the trained depth network, signal, and controlling, based on the signal, autonomous driving of the vehicle.
Owner:HYUNDAI MOTOR CO LTD +1

Neural radiation field compression rendering method based on decomposition expression

The invention provides a neural radiation field compression rendering method based on decomposition expression, and relates to the technical field of computer graphics, and the method comprises the steps: inputting a multi-view image and camera internal and external parameters, carrying out the ray sampling, generating a space sampling point, carrying out the mixed feature coding of the space sampling point, and obtaining a multi-view image; obtaining three-dimensional voxel features, aligned and fused three-plane features and position codes, and splicing the three features to form fused features; inputting the fusion features into a factorization neural BRDF rendering network to predict volume density, geometric latent features, material parameters and reflection features, and improving the view angle color under the compression condition through a BRDF modulator; a ray weight is predicted through a plane-ray combined modeling module, and a compressed neural radiation field model is obtained through combined optimization of miniaturized body rendering, weighted reconstruction loss and compression constraint and is used for target view angle image rendering; according to the method, the storage overhead of the neural radiation field model is reduced, and meanwhile, the synthesis quality and rendering consistency of the new view angle under different compression ratios are improved.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Gaussian depth regularization new viewpoint synthesis method based on hierarchical mask

The invention discloses a Gaussian depth regularization new viewpoint synthesis method based on layered masks. Firstly, a layered depth optimization mechanism with foreground strong supervision and background weak constraint is constructed, and multi-view depth continuity and texture consistency collaborative constraint is introduced to realize stable optimization of three-dimensional Gaussian parameters. According to the method, monocular depth noise propagation is effectively inhibited, three-dimensional reconstruction precision and visual consistency are improved, training efficiency and engineering feasibility are considered, and the method is suitable for application scenes such as virtual reality, digital twinning and three-dimensional content generation; the innovation point of the invention is that the mask-guided hierarchical supervision and multi-view collaborative optimization mechanism solves the problem that the existing 3DGS lacks structural constraints, and has clear technical improvement; the method is based on a mature depth estimation and segmentation model, is high in technical maturity, and has an implementation basis.
Owner:NANJING UNIV OF POSTS & TELECOMM

View synthesis using a multiview image generative machine learning model

Disclosed are systems, apparatuses, processes, and computer-readable media for generating images based on sparse images using a multiview image generative machine learning model. A disclosed method includes determining, using a machine learning (ML) model, a plurality of 3D features from one or more two-dimensional (2D) images based on respective pose information associated with each 2D image of the one or more 2D images, wherein the respective pose information is relative to a target pose of a target 2D image; combining, using the ML model, the plurality of 3D features into a plurality of multiview features; and generating the target 2D image including the target pose based on the plurality of 3D features and the plurality of multiview features.
Owner:QUALCOMM INC

Unsupervised volumetric animation

Unsupervised volumetric 3D animation (UVA) of non-rigid deformable objects without annotation learns the 3D structure and dynamics of the object only from a single view red / green / blue (RGB) video and decomposes the single view RGB video into semantically meaningful portions that can be tracked and animated. Using a 3D automatic decoder framework, the UVA model learns 3D geometry and partial decomposition of underlying objects from still or video images in a fully unsupervised manner via a microperspective n-point (PnP) algorithm pairing with a keypoint estimator. This allows the UVA model to perform 3D segmentation, 3D keypoint estimation, novel view synthesis, and animation. The UVA model may obtain animatable 3D objects from a single or several images. The UVA method also characterizes a space in which all objects are represented in their normative, animated ready form. Applications include creating a shot from an image or video of a social media application.
Owner:SNAP INC

Ray tracing volumetric particles for real-time novel view synthesis

Approaches presented herein provide for efficient rendering of high quality, novel views of a scene, in this case achieved through a combination of volumetric particle representations and ray tracing. An object can be represented using a set of volumetric particles (e.g., 3D distributions) that are aligned to the underlying structure or geometry of the object. Volumetric particles can be encapsulated in a bounding mesh or proxy geometry that can be used to efficiently compute ray-particle intersections. For a view to be rendered, ray tracing can be performed to determine an intersection of the rays with the proxy geometry. When a hit is determined, the precise intersection location with the volumetric particle is computed and the value of the distribution returned for that ray. If a ray passes through multiple semi-transparent volumetric particles then the color value is determined based upon the values returned from those particles.
Owner:NVIDIA CORP

View synthesis from images and / or video using model-based inpainting

Systems and techniques for image processing are described. For example, a computing device can receive a pose of a camera with a field-of-view (FOV) of a scene and can determine poses of the camera and / or pose(s) of other camera(s). The computing device can obtain first camera layers associated with the camera. The computing device can obtain, based on the pose of the camera and pixels within the first camera layers, a mask corresponding to the pose of the camera. The computing device can generate composited layers and can determine pixels for regions of the image using a first model based on the mask, the plurality of composited layers, the pose of the camera, and a plurality of images of the scene. The computing device can generate a final image of the scene corresponding to the pose of the camera based on providing the determined pixels to the regions.
Owner:QUALCOMM INC

Conditional image generation method based on point rasterization and related device

The invention provides a condition image generation method based on point rasterization and a related device, and the method comprises the steps: carrying out the coloring processing of a target point cloud set according to an environment image corresponding to a target moment; performing foreground and background separation processing on the target point cloud set after coloring processing to determine a dynamic foreground region and a static background region in the target point cloud set; complementing the incomplete area in combination with a historical point cloud set in the time window; projecting the complemented target point cloud set to an image coordinate system of the vehicle by using a perspective projection mapping relation to obtain a first projection image; and rendering the grating area in the first projection image to obtain a point rasterization condition image. According to the method, frame-by-frame analysis and image plane mapping are carried out on original laser radar point cloud data, point rasterization processing is combined, a geometric condition image with pixel-level precision is constructed, and structural prior support is provided for subsequent new view synthesis and controllable video generation.
Owner:BEIHANG UNIV

Nerve radiation field two-stage three-dimensional reconstruction method based on sign distance function

The invention discloses a two-stage three-dimensional reconstruction method for a neural radiation field based on a symbolic distance function. A more accurate and real object surface grid is generated from a two-dimensional image. The method comprises the following steps of: dividing a reconstruction process into two stages, and respectively modeling the color and the volume density of an object by using two different networks; in the first stage, a zero-order set of a symbol distance field is used for representing the surface of an object, and rough grids are preliminarily extracted; in the second stage, loss is optimized continuously, and the vertex position and the surface density are adjusted step by step to achieve refinement of the grid surface; multi-resolution hash coding is utilized to accelerate training; three-dimensional position and view vector information are added into each layer of a multi-layer perceptron, supervision network training is displayed through multi-view geometric constraints, and reconstruction quality is improved. According to the method, the three-dimensional reconstruction process is divided into two stages, a good view synthesis effect is kept by using implicit representation and geometric constraint, and the precision and speed of model training and reconstruction are improved.
Owner:BEIJING UNIV OF TECH

Layered view synthesis system and method

A method of computer-implemented synthesized view image generation and a synthesized view image generation system provide layered view synthesis. The method includes receiving an input image having a plurality of pixels having color values; generating a dilated depth map by dilating a depth map associated with the input image, the depth map with depth values respectively associated with each pixel in the input image; determining an inpainting mask using the dilated depth map; performing an inpainting operation based on the inpainting mask and the input image to generate a background image; and rendering a synthesized view image using the background image, the input image, and the dilated depth map.
Owner:LEIA INC

View synthesis with learned gaussing splatting and weighted sum rendering

A system generates initial Gaussian elements defined by parameter sets that include, for each Gaussian element a spherical harmonics (SH) coefficient array, a learnable parameter vector, and a learnable weight vector. The system performs a training process comprising rasterizing current Gaussian elements to generate a rendered image of the scene as viewable from a current camera position, wherein for each Gaussian element of the current Gaussian elements that intersects the camera ray, the system determines an opacity value for a location based on a view-dependent scaling value that depends on the current camera position, a position vector, and the learnable parameter vector.
Owner:QUALCOMM INC