Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

108 results about "View synthesis" patented technology

Currently a study branch of Computer Science Research, aims to create new views of a specific subject starting from a number of pictures taken from given point of views. Vision Research and Artificial Intelligence fields are involved in the definition of suitable approaches to the problem.

System and method for efficient scene continuity in visual and multimedia using generative artificial intelligence

ActiveUS20250378537A1Image enhancementPattern recognitionGenerative process
A system and method for generating multimedia artifacts with managed scene continuity in visual and multimedia using an AI-based and scene continuity aware media generation platform. The system receives a user or AI agent specification or simulation result(s), selects or trains generative models based on the specification, preprocesses relevant data, and generates scene narrative or frame-specific, sequence specific or broader continuity aware content using the selected or trained model(s). The generated content may be further enhanced using frame interpolation and view synthesis techniques to create smooth transitions or novel viewpoints or to aid in more efficient transmission or viewing or persistence of resultant content. The system enables efficient and customizable generation of high-quality scene continuity aware content for various applications in visual and multimedia production using neuro-symbolic and simulation enhanced compression, representation and generation processes.
Owner:QOMPLX INC

Integrated visual angle synthesis and sparse visual angle CT reconstruction method based on 3DGS

The invention discloses a 3DGS-based integrated visual angle synthesis and sparse visual angle CT reconstruction method, which comprises the following steps: establishing an internal and external parameter matrix through X-ray scanning parameters, constructing an initial voxel space and an initialized 3DGS radiation field by combining sparse visual angle projection data and based on an ACUI strategy, modeling anisotropic radiation intensity by utilizing a radiation intensity response function, and reconstructing a 3D visual angle CT model. And generating a virtual projection by adopting volume consistency rendering. And the optimized radiation field is converted into a voxel grid, density contribution is applied to carry out voxel space reconstruction, a reconstructed voxel space is obtained, and a new 3DGS radiation field is generated. And re-rendering the virtual visual angle projection based on the new 3DGS radiation field, performing iterative optimization by using a loss function, and outputting a new visual angle synthetic image and a CT reconstruction body until a preset number of iterations is reached. And an efficient process of one-time training and dual output is realized. The efficiency and precision of X-ray imaging are improved, the radiation dose of a patient is reduced, and meanwhile, a higher-quality imaging solution is provided for medical diagnosis.
Owner:GUANGDONG UNIV OF TECH

Layered densification Gaussian sputtering method based on visibility

The invention discloses a layered densification Gaussian sputtering method based on visibility, and relates to the technical field of artificial intelligence and computer vision. According to the scheme, an initial three-dimensional Gaussian primitive set is generated based on sparse multi-view observation data, and initial scene representation is established by extracting spatial distribution parameters, morphological parameters and radiation parameters; performing fusion analysis on the Gaussian primitives based on the multi-dimensional observability parameter set to obtain comprehensive observability index data; executing hierarchical clustering according to the index data, and constructing a multi-layer Gaussian structure of a significant layer, a transition layer and a background layer; performing density enhancement, geometric continuity constraint and parameter update processing on different levels of Gaussian structures, and generating a rendered image of a target view angle under a volume light traveling and transparency hybrid mechanism; according to the method, continuous reconstruction of a scene structure and accurate expression of radiation characteristics can be realized under the sparse view condition, and the geometric fidelity and rendering consistency of new view angle synthesis are improved.
Owner:HENAN JINSHU INTELLIGENT TECH CO LTD

Target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment

The invention discloses a target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment. The target-level three-dimensional point cloud cross-modal semantic retrieval method comprises the steps that independent three-dimensional target point clouds are segmented from three-dimensional scene point clouds, and unique identifiers are given to the independent three-dimensional target point clouds; projecting each target point cloud to three orthogonal two-dimensional observation planes to generate a multi-view composite image; a pre-trained vision-language basic model is adopted to code and fuse the synthesized image, and a unified multi-modal feature vector is generated; a vector database in which the feature vectors are associated with their identifiers is constructed. And encoding a natural language query text into a text query vector by using the model, calculating the semantic similarity between the text query vector and the feature vector in the database, and positioning and returning the corresponding three-dimensional target point cloud. According to the method, accurate and efficient cross-modal retrieval from a natural language to a three-dimensional point cloud target is realized, the problem that a traditional method is difficult to support semantic fine-grained retrieval of the three-dimensional target point cloud is solved, and a key technical support is provided for training data management in the fields of intelligence and the like.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Three-dimensional Gaussian splashing and rasterization pipeline fused view synthesis system and method

The invention discloses a three-dimensional Gaussian splashing and rasterization pipeline fused view synthesis system and method. The method comprises the following steps of: 1, acquiring two-dimensional Gaussian primitives; 2, converting the two-dimensional Gaussian primitive into two triangular grids, and converting the two triangular grids into a world coordinate system; 3, baking the spherical harmonic coefficient of the color and the opacity data in a texture form; 4, enabling each Gaussian primitive to correspond to a texture block on the texture image set, enabling a triangle to cover the texture block in a UV space, and obtaining a Gaussian weight; 5, based on the obtained Gaussian weight, simulating an alpha mixing behavior in the three-dimensional Gaussian splashing pipeline by using a time-based jitter anti-aliasing method; and 6, training a symbol vector field SDF branch while training, and mutually supervising and constraining the output of the symbol vector field and the output of the two-dimensional Gaussian. The three-dimensional Gaussian splash can be rendered in a traditional rasterization pipeline of the unreal engine 4, and the fidelity equivalent to that of a special training pipeline is kept.
Owner:XIDIAN UNIV +1

Small sample new view synthesis method based on reinitialized three-dimensional Gaussian splashing

The invention discloses a small sample new view synthesis method based on reinitialization three-dimensional Gaussian splashing. The method comprises the following steps: firstly, acquiring sparse three-dimensional point cloud and camera internal and external parameters from a training view through a motion recovery structure algorithm; then, sampling points are generated in the point cloud bounding box by adopting a spatial expansion hybrid sampling strategy, and a coarse-grained Gaussian set is constructed and optimized; thirdly, obtaining a rendered image through Gaussian initialization and splash rendering, calculating pixel importance through depth errors and transmissivity, generating a fine-grained Gaussian set through back projection, and optimizing the fine-grained Gaussian set; and finally, calculating a sampling probability based on the cross-view contribution degree, screening key Gaussian distribution, and optimizing to form a final Gaussian set, thereby realizing high-quality new view synthesis. According to the method, the problem of sparsity difference is solved by eliminating extended view dependence, the multi-view consistency and local geometric details of a new view scene are improved, and high-quality new view synthesis of sparse data is realized.
Owner:ZHEJIANG UNIV

Neural radiation field compression rendering method based on decomposition expression

The invention provides a neural radiation field compression rendering method based on decomposition expression, and relates to the technical field of computer graphics, and the method comprises the steps: inputting a multi-view image and camera internal and external parameters, carrying out the ray sampling, generating a space sampling point, carrying out the mixed feature coding of the space sampling point, and obtaining a multi-view image; obtaining three-dimensional voxel features, aligned and fused three-plane features and position codes, and splicing the three features to form fused features; inputting the fusion features into a factorization neural BRDF rendering network to predict volume density, geometric latent features, material parameters and reflection features, and improving the view angle color under the compression condition through a BRDF modulator; a ray weight is predicted through a plane-ray combined modeling module, and a compressed neural radiation field model is obtained through combined optimization of miniaturized body rendering, weighted reconstruction loss and compression constraint and is used for target view angle image rendering; according to the method, the storage overhead of the neural radiation field model is reduced, and meanwhile, the synthesis quality and rendering consistency of the new view angle under different compression ratios are improved.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Gaussian depth regularization new viewpoint synthesis method based on hierarchical mask

The invention discloses a Gaussian depth regularization new viewpoint synthesis method based on layered masks. Firstly, a layered depth optimization mechanism with foreground strong supervision and background weak constraint is constructed, and multi-view depth continuity and texture consistency collaborative constraint is introduced to realize stable optimization of three-dimensional Gaussian parameters. According to the method, monocular depth noise propagation is effectively inhibited, three-dimensional reconstruction precision and visual consistency are improved, training efficiency and engineering feasibility are considered, and the method is suitable for application scenes such as virtual reality, digital twinning and three-dimensional content generation; the innovation point of the invention is that the mask-guided hierarchical supervision and multi-view collaborative optimization mechanism solves the problem that the existing 3DGS lacks structural constraints, and has clear technical improvement; the method is based on a mature depth estimation and segmentation model, is high in technical maturity, and has an implementation basis.
Owner:NANJING UNIV OF POSTS & TELECOMM

View synthesis with learned gaussing splatting and weighted sum rendering

A system generates initial Gaussian elements defined by parameter sets that include, for each Gaussian element a spherical harmonics (SH) coefficient array, a learnable parameter vector, and a learnable weight vector. The system performs a training process comprising rasterizing current Gaussian elements to generate a rendered image of the scene as viewable from a current camera position, wherein for each Gaussian element of the current Gaussian elements that intersects the camera ray, the system determines an opacity value for a location based on a view-dependent scaling value that depends on the current camera position, a position vector, and the learnable parameter vector.
Owner:QUALCOMM INC

Three-Dimensional Diffusion Models

Provided are systems and methods to perform novel view synthesis of a three-dimensional (3D) scene with a machine-learned diffusion model. Example implementations of the proposed models may be referred to as “3D Diffusion Models” or 3DiM. The models described herein can be or include an image-to-image diffusion model that takes one or more (e.g., a single) reference views and one or more (e.g., a single) relative poses as input and generates the target view. Thus, the machine-learned diffusion models described herein can perform novel view synthesis from as few as a single image.
Owner:GOOGLE LLC

System and method for efficient scene continuity in visual and multimedia using generative artificial intelligence

ActiveUS12499515B2Image enhancementPattern recognitionGenerative process
A system and method for generating multimedia artifacts with managed scene continuity in visual and multimedia using an AI-based and scene continuity aware media generation platform. The system receives a user or AI agent specification or simulation result(s), selects or trains generative models based on the specification, preprocesses relevant data, and generates scene narrative or frame-specific, sequence specific or broader continuity aware content using the selected or trained model(s). The generated content may be further enhanced using frame interpolation and view synthesis techniques to create smooth transitions or novel viewpoints or to aid in more efficient transmission or viewing or persistence of resultant content. The system enables efficient and customizable generation of high-quality scene continuity aware content for various applications in visual and multimedia production using neuro-symbolic and simulation enhanced compression, representation and generation processes.
Owner:QOMPLX INC

Self-supervised depth for volumetric rendering regularization

An example method includes generating embeddings of image data that includes multiple images, where each image has a different viewpoints of a scene, generating a latent space and a decoder, wherein the decoder receives embeddings as input to generate an output viewpoint, for each viewpoint in the image data, determining a volumetric rendering view synthesis loss and a multi-view photometric loss, and applying an optimization algorithm to the latent space and the decoder over a number of epochs until the volumetric rendering view synthesis loss is within a volumetric threshold and the multi-view photometric loss is within a multi-view threshold.
Owner:MASSACHUSETTS INST OF TECH +1

A method and related apparatus for generating conditional images based on point rasterization

The present application provides a kind of conditional image generation method and related device based on point rasterization, according to the environment image corresponding to target time, target point cloud set is colored processing;The target point cloud set after coloring processing is carried out front and background separation processing, to determine the dynamic foreground region and static background region therein;Combining the historical point cloud set in time window, the incomplete area is completed;Using perspective projection mapping relationship, the target point cloud set after completion is projected to the image coordinate system of vehicle, to obtain first projection image;The raster area in first projection image is rendered, to obtain the conditional image of point rasterization.It is analyzed frame by frame to original laser radar point cloud data and image plane mapping, and combined with point rasterization processing, constructs the geometric conditional image with pixel level precision, provides structure priori support for subsequent new view synthesis and controllable video generation.
Owner:BEIHANG UNIV

Ray tracing volumetric particle for real-time novel view synthesis

To provide a method for efficient rendering of high-quality, novel views of a scene.SOLUTION: A method is realized through a combination of volumetric particle representations and ray tracing. An object can be represented using a set of volumetric particles (e.g., 3D distributions) that are aligned to an underlying structure or geometry of the object. The volumetric particles can be encapsulated in a bounding mesh or proxy geometry that can be used for efficiently computing ray-particle intersections. For a view to be rendered, the ray tracing can be performed to determine an intersection of the rays with the proxy geometry. When a hit is determined, a precise intersection location with the volumetric particle is computed and a value of a distribution returns for that ray. If a ray passes through multiple semi-transparent volumetric particles, then a color value is determined based upon values returned from those particles.SELECTED DRAWING: Figure 5
Owner:NVIDIA CORP

Conditioned generative model for consistent novel views synthesis

Described is a computer apparatus (1400) configured to: obtain a reference image (202) of a 3D scene from a reference viewpoint; obtain a reference point map (203) corresponding to the reference image (202), the reference point map (203) comprising the coordinates of a plurality of pixels of the reference image (202); obtain a target point map (204) corresponding to the rendered image (206), the target point map (204) comprising the coordinates of a plurality of pixels of the rendered image (206); and generate the rendered image (206) in dependence on the reference image (202), the reference point map (203) and the target point map (204). In this way, the rendering process may be guided by the positions in the point maps.
Owner:HUAWEI TECH CO LTD +1

Unsupervised monocular image-based online 3D scene reconstruction method and device

The present disclosure provides a monocular image-based unsupervised three-dimensional scene online reconstruction method and device. The method of the present disclosure comprises: obtaining a static background mask and a mask of each dynamic object instance of a current frame monocular image through semantic segmentation, obtaining a surround view synthesis image through a static multi-view generator, determining the motion parameters of each dynamic object instance through motion modeling, obtaining a multi-view depth map based on the surround view synthesis image through a depth estimation network obtained through self-supervised training, obtaining a local depth map of each dynamic object instance through a local depth estimation network obtained through self-supervised training, and obtaining a complete 3D scene representation through the multi-view depth map, the local depth map of each dynamic object instance, and the motion parameters thereof. The present disclosure can avoid true value dependence, effectively reduce hardware cost, and at the same time improve the reliability and robustness of monocular image three-dimensional scene reconstruction.
Owner:BEIJING TRUNK TECHNOLOGY CO LTD

Neural view synthesis using tiled multiplane images

A method for generating tiled multiplane images from a source image is disclosed. The method includes obtaining color and depth images. The method also includes extracting a feature map using a first neural network. The method also includes generating masks for a tile using a second neural network based on a corresponding tile of the feature map and corresponding sections of the color and depth images. The method also includes computing depths of a planes corresponding to the tile based on the masks. The method also includes generating a per-tile multiplane image for the tile based on the masks. The method also includes rendering an image using per-tile multiplane images and depths. A system for generating tiled multiplane images from a source image is also disclosed.
Owner:META PLATFORMS TECHNOLOGIES LLC

A new view synthesis method, system and device

The application relates to a new view synthesis method, system and device, obtains images and camera poses under multiple viewing angles, encodes to generate input labels to be processed; encodes to generate target labels based on target viewing angle poses; inputs the input labels of each viewing angle into an encoder composed of multiple layers of Transformer blocks, and the output of each layer of the Transformer blocks is saved in a key-value cache; inputs the target labels into a decoder for processing, and outputs label representations of target viewing angle images, the decoder has the same number of layers of Transformer blocks as the encoder, performs self-attention processing on input data as a query, uses the key-value cache of the corresponding layer in the encoder as a key and a value, performs cross-attention query operation, and the result is transmitted into a feedforward neural network to obtain a final output; and the label representations of the target viewing angle images are converted into RGB images under the target viewing angle. Compared with the prior art, the application can more accurately learn scene semantics of input views and rendering rules of target views, and generate high-quality new views.
Owner:FUDAN UNIVERSITY

Deblurred view synthesis method and device, electronic equipment and storage medium

The embodiment of the invention discloses a deblurred view synthesis method and device, electronic equipment and a storage medium, and relates to the technical field of computer vision and three-dimensional reconstruction, and the method comprises the steps: obtaining an initial camera pose through three-dimensional reconstruction software, and employing a cubic Bezier curve to model a camera motion track; dynamically predicting three-plane resolution by using a feedforward neural network according to the track length; self-adaptive three-plane downsampling is carried out based on the predicted resolution, and a clear image is generated through a volume rendering technology; and finally, obtaining a full-resolution three-plane field through a loss optimization and gradient descent algorithm, and generating a clear image corresponding to the pose of the target camera according to the learnable parameters and the three-plane field. According to the method, the problems of low sampling efficiency, long training time and rendering quality degradation under a long track when facing camera motion tracks with different lengths in the prior art are solved, the training efficiency is remarkably improved, the memory occupation is reduced, and the generalization ability and reconstruction quality of the model are enhanced.
Owner:SHENZHEN TIANHAI CHENGUANG TECH CO LTD

Platform for enabling multiple users to generate and use neural radiance field models

Systems and methods for enabling users to generate and utilize neural radiance field models can include obtaining user image data and training one or more neural radiance field models based on the user image data. The systems and methods can include obtaining user images based on a determination that the user images depicted objects of a particular object type. The trained neural radiance field models can then be utilized for view synthesis image generation of the particular user objects.
Owner:GOOGLE LLC

Automatic correction of view foreshortening in cardiac echo using view synthesis

Systems and methods for automatic correction of view foreshortening in cardiac echo using view synthesis. A trained image segmentation method is used to infer a 2D pose from the input image. The pose is transferred to a new view representing a non-foreshortened view plane. A non-foreshortened image is then rendered.
Owner:SIEMENS MEDICAL SOLUTIONS USA INC

View synthesis system and method using depth map

In multiview image generation and display, a computing device can synthesize view images of a multiview image of a scene from a color image and a depth map. Each view image can include color values at respective pixel locations. The computing device can render the view images of the multiview image on a multiview display. Synthesizing a view image can include, for a pixel location in the view image, the following operations. The computing device can cast a ray from the pixel location toward the scene in a direction corresponding to a view direction of the view image. The computing device can determine a ray intersection location at which the ray intersects a virtual surface specified by the depth map. The computing device can set a color value of the view image at the pixel location to correspond to a color of the color image at the ray intersection location.
Owner:LEIA INC

A new view synthesis method based on gaussian probability distribution and feature regularization

The application provides a new view synthesis method based on Gaussian probability distribution and feature regularization, relates to the technical field of new view synthesis, and comprises the following steps: extracting a feature tensor from a preprocessed input image, performing Gaussian probability calculation based on the mean and standard deviation of the feature tensor; introducing an elastic net regularization loss to constrain the feature tensor, performing mean calculation on the Gaussian probability of the feature tensor, substituting the calculation result into a loss function, minimizing the loss function, obtaining an optimized Gaussian point set, performing multi-scale densification and pruning, and generating a new view. The application solves the technical problem that, due to overfitting in a sparse scene, the number of Gaussian points is too small and the transparency is low, which further affects the accuracy of view synthesis, and improves the modeling effect and the accuracy of view synthesis in a sparse view angle by introducing Gaussian probability mapping and feature regularization, thereby increasing efficiency and improving precision.
Owner:BEIJING INSTITUTE OF SURVEYING AND MAPPING

Initial noise optimization method of new view angle synthesis model based on diffusion model

The invention discloses an initial noise optimization method of a new view angle synthesis model based on a diffusion model, belongs to the diffusion model, and can be used in the field of new view angle synthesis. According to the optimization method, firstly, a machine learning network is provided, and the network is composed of an encoder and a decoder; then training an encoder-decoder network according to initial noise-optimized noise pair acquired by a new view angle synthesis model based on a diffusion model; and finally, inserting the trained encoder-decoder network into the new visual angle synthesis model. The invention aims to change initial noise into optimized noise through network action, so that a model generation result is improved on the premise of not finely tuning a model.
Owner:NANJING TECH UNIV

Differentiable real-time radiance field rendering for large scale view synthesis

A method that includes obtaining images of an environment that are captured by one or more image capture devices, determining intrinsic parameters and extrinsic parameters of the one or more image capture devices that are associated with each of the images, creating a differentiable radiance field associated with the environment, and generating a three-dimensional representation of the environment. The three-dimensional representation contains one or more portions of the environment uncaptured in the images.
Owner:OPAL AI INC

A method and system for depth map super-resolution with view synthesis fusion

The application belongs to the technical field of image processing, and provides a depth map super-resolution method and system fusing view synthesis, comprising: acquiring a low-resolution depth map; obtaining a high-low resolution depth map according to the acquired low-resolution depth map and an optimized super-resolution network; in the application, a color picture of a target view point obtained by view synthesis from a high-resolution depth true value map is used as a true value of a color image; the super-resolution network is optimized by comparing the difference between the true value of the color image and the color picture of the target view point generated by the network reconstructed depth map predicted to obtain the optimized super-resolution network, the problem that the high-resolution color image is only used to extract features and fuse the features of the depth map is solved, and the precision of the super-resolution network and the effect of the depth map super-resolution are improved.
Owner:SHANDONG UNIV

A real-time dynamic three-dimensional reconstruction method and system based on patch matching and lightweight monocular depth estimation fusion

The application discloses a kind of real-time dynamic three-dimensional reconstruction method and system based on patch matching and lightweight monocular depth estimation fusion, the method is carried out PatchMatch matching to input multi-view image pair, obtains sparse depth information, and is fused with monocular depth estimation result;The depth map after fusion is used to construct three-dimensional Gaussian point cloud, realize high-quality three-dimensional scene reconstruction and new view rendering.The whole process uses end-to-end differentiable training framework, by optimizing pixel-level Gaussian parameter regression network and lightweight depth estimation model, the reconstruction accuracy and rendering efficiency of 3D Gaussian point cloud in dynamic scene are improved.The application can complete training and depth inference in a short time, suitable for real-time or near real-time application scenarios, can be widely used in virtual reality, augmented reality, film and television production and other fields, especially suitable for efficient three-dimensional reconstruction and new view synthesis of target in multi-view dynamic scene.
Owner:NANJING UNIV

Robustifying NeRF Model Novel View Synthesis to Sparse Data

Systems and methods for training a neural radiance field model can include the use of image patches for ground truth training. For example, the systems and methods can include generating patch renderings with a neural radiance field model, comparing the patch renderings to ground truth patches from ground truth images, and adjusting one or more parameters based on the comparison. Additionally and / or alternatively, the systems and methods can include the utilization of a flow model for mitigating and / or minimizing artifact generation.
Owner:GOOGLE LLC