Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

160 results about "View synthesis" patented technology

Currently a study branch of Computer Science Research, aims to create new views of a specific subject starting from a number of pictures taken from given point of views. Vision Research and Artificial Intelligence fields are involved in the definition of suitable approaches to the problem.

System and method for efficient scene continuity in visual and multimedia using generative artificial intelligence

ActiveUS20250378537A1Image enhancementPattern recognitionGenerative process
A system and method for generating multimedia artifacts with managed scene continuity in visual and multimedia using an AI-based and scene continuity aware media generation platform. The system receives a user or AI agent specification or simulation result(s), selects or trains generative models based on the specification, preprocesses relevant data, and generates scene narrative or frame-specific, sequence specific or broader continuity aware content using the selected or trained model(s). The generated content may be further enhanced using frame interpolation and view synthesis techniques to create smooth transitions or novel viewpoints or to aid in more efficient transmission or viewing or persistence of resultant content. The system enables efficient and customizable generation of high-quality scene continuity aware content for various applications in visual and multimedia production using neuro-symbolic and simulation enhanced compression, representation and generation processes.
Owner:QOMPLX INC

Non-static scene reconstruction method and system based on multi-modal occlusion perception scoring

The invention discloses a non-static scene reconstruction method and system based on multi-modal occlusion perception scoring. The method comprises the following steps: generating a geometric consistency distribution diagram, static feature points and geometric prior masks through three-dimensional reconstruction of a multi-view image; fusing the geometric prior mask and a semantic segmentation model to extract a semantic mask and a fused semantic feature map; guiding the image segmentation model to generate candidate masks based on static feature point positive point prompt and occlusion area negative frame prompt, and optimizing the consistency by using a grid complementary fusion method; constructing a multi-modal shielding scoring module, and fusing the multi-source features to output a binary static mask; and utilizing a static mask to constrain neural radiation field training, inhibiting dynamic interference and optimizing static scene reconstruction. According to the method, the static region is sensed cooperatively through multi-modal information, the robustness and accuracy of mask generation are improved, the interference of dynamic elements on neural radiation field modeling is effectively inhibited, and high-quality three-dimensional image reconstruction and new view synthesis of a non-static scene are realized.
Owner:HANGZHOU DIANZI UNIV

Integrated visual angle synthesis and sparse visual angle CT reconstruction method based on 3DGS

The invention discloses a 3DGS-based integrated visual angle synthesis and sparse visual angle CT reconstruction method, which comprises the following steps: establishing an internal and external parameter matrix through X-ray scanning parameters, constructing an initial voxel space and an initialized 3DGS radiation field by combining sparse visual angle projection data and based on an ACUI strategy, modeling anisotropic radiation intensity by utilizing a radiation intensity response function, and reconstructing a 3D visual angle CT model. And generating a virtual projection by adopting volume consistency rendering. And the optimized radiation field is converted into a voxel grid, density contribution is applied to carry out voxel space reconstruction, a reconstructed voxel space is obtained, and a new 3DGS radiation field is generated. And re-rendering the virtual visual angle projection based on the new 3DGS radiation field, performing iterative optimization by using a loss function, and outputting a new visual angle synthetic image and a CT reconstruction body until a preset number of iterations is reached. And an efficient process of one-time training and dual output is realized. The efficiency and precision of X-ray imaging are improved, the radiation dose of a patient is reduced, and meanwhile, a higher-quality imaging solution is provided for medical diagnosis.
Owner:GUANGDONG UNIV OF TECH

Layered densification Gaussian sputtering method based on visibility

The invention discloses a layered densification Gaussian sputtering method based on visibility, and relates to the technical field of artificial intelligence and computer vision. According to the scheme, an initial three-dimensional Gaussian primitive set is generated based on sparse multi-view observation data, and initial scene representation is established by extracting spatial distribution parameters, morphological parameters and radiation parameters; performing fusion analysis on the Gaussian primitives based on the multi-dimensional observability parameter set to obtain comprehensive observability index data; executing hierarchical clustering according to the index data, and constructing a multi-layer Gaussian structure of a significant layer, a transition layer and a background layer; performing density enhancement, geometric continuity constraint and parameter update processing on different levels of Gaussian structures, and generating a rendered image of a target view angle under a volume light traveling and transparency hybrid mechanism; according to the method, continuous reconstruction of a scene structure and accurate expression of radiation characteristics can be realized under the sparse view condition, and the geometric fidelity and rendering consistency of new view angle synthesis are improved.
Owner:HENAN JINSHU INTELLIGENT TECH CO LTD

Sparse dynamic 3D Gaussian splash method based on global-local feature extraction

The invention discloses a sparse dynamic 3D Gaussian splash method based on global-local feature extraction. According to the method, a global-local feature extraction module, an inter-frame feature flow aggregation network and depth prior are combined with a 4D pseudo pose supervision mechanism for cooperative work, and high-quality dynamic scene reconstruction and rendering are realized under the condition of a sparse camera visual angle. According to the method, the multi-view image information is compressed to the feature space for processing, so that the calculation complexity is reduced, and the geometric structure and the dynamic features of the scene are effectively reserved; the fused depth prior information and the 4D pseudo pose provide accurate geometric constraints, and the inter-frame feature flow aggregation network optimizes the time consistency of the dynamic scene. The method is particularly suitable for processing a complex dynamic scene under a sparse view angle, the sense of reality and the real-time performance of virtual view synthesis are remarkably improved, and a dynamic view with a more accurate geometric structure, higher time continuity and richer details can be generated.
Owner:HANGZHOU DIANZI UNIV

Target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment

The invention discloses a target-level three-dimensional point cloud cross-modal semantic retrieval method and device and electronic equipment. The target-level three-dimensional point cloud cross-modal semantic retrieval method comprises the steps that independent three-dimensional target point clouds are segmented from three-dimensional scene point clouds, and unique identifiers are given to the independent three-dimensional target point clouds; projecting each target point cloud to three orthogonal two-dimensional observation planes to generate a multi-view composite image; a pre-trained vision-language basic model is adopted to code and fuse the synthesized image, and a unified multi-modal feature vector is generated; a vector database in which the feature vectors are associated with their identifiers is constructed. And encoding a natural language query text into a text query vector by using the model, calculating the semantic similarity between the text query vector and the feature vector in the database, and positioning and returning the corresponding three-dimensional target point cloud. According to the method, accurate and efficient cross-modal retrieval from a natural language to a three-dimensional point cloud target is realized, the problem that a traditional method is difficult to support semantic fine-grained retrieval of the three-dimensional target point cloud is solved, and a key technical support is provided for training data management in the fields of intelligence and the like.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Three-dimensional Gaussian splashing and rasterization pipeline fused view synthesis system and method

The invention discloses a three-dimensional Gaussian splashing and rasterization pipeline fused view synthesis system and method. The method comprises the following steps of: 1, acquiring two-dimensional Gaussian primitives; 2, converting the two-dimensional Gaussian primitive into two triangular grids, and converting the two triangular grids into a world coordinate system; 3, baking the spherical harmonic coefficient of the color and the opacity data in a texture form; 4, enabling each Gaussian primitive to correspond to a texture block on the texture image set, enabling a triangle to cover the texture block in a UV space, and obtaining a Gaussian weight; 5, based on the obtained Gaussian weight, simulating an alpha mixing behavior in the three-dimensional Gaussian splashing pipeline by using a time-based jitter anti-aliasing method; and 6, training a symbol vector field SDF branch while training, and mutually supervising and constraining the output of the symbol vector field and the output of the two-dimensional Gaussian. The three-dimensional Gaussian splash can be rendered in a traditional rasterization pipeline of the unreal engine 4, and the fidelity equivalent to that of a special training pipeline is kept.
Owner:XIDIAN UNIV +1

Small sample new view synthesis method based on reinitialized three-dimensional Gaussian splashing

The invention discloses a small sample new view synthesis method based on reinitialization three-dimensional Gaussian splashing. The method comprises the following steps: firstly, acquiring sparse three-dimensional point cloud and camera internal and external parameters from a training view through a motion recovery structure algorithm; then, sampling points are generated in the point cloud bounding box by adopting a spatial expansion hybrid sampling strategy, and a coarse-grained Gaussian set is constructed and optimized; thirdly, obtaining a rendered image through Gaussian initialization and splash rendering, calculating pixel importance through depth errors and transmissivity, generating a fine-grained Gaussian set through back projection, and optimizing the fine-grained Gaussian set; and finally, calculating a sampling probability based on the cross-view contribution degree, screening key Gaussian distribution, and optimizing to form a final Gaussian set, thereby realizing high-quality new view synthesis. According to the method, the problem of sparsity difference is solved by eliminating extended view dependence, the multi-view consistency and local geometric details of a new view scene are improved, and high-quality new view synthesis of sparse data is realized.
Owner:ZHEJIANG UNIV

3D video generation method, 3D video viewing method, and electronic device

Provided in the embodiments of the present application are a 3D video generation method and an electronic device. The method is applied to the electronic device. The method comprises: on the basis of captured data, acquiring 3D video material data, wherein the captured data comprises a 2D video and / or a 2D image; on the basis of the 3D video material data, analyzing a 3D video scene type; on the basis of the 3D video scene type, acquiring a new-view generation model corresponding to the 3D video scene type; and using the new-view generation model corresponding to the 3D video scene type to generate 3D video data on the basis of the 3D video material data. On the basis of the method of the embodiments of the present application, at a 3D video data generation stage, a corresponding new-view generation model is called on the basis of a 3D video scene type, and thus a more complete scene synthesis view can be acquired on the basis of a new-view synthesis algorithm that uses deep learning and artificial intelligence, such that when 3D video data is played, a user can immersively view 3D video content from different angles, thereby providing a stronger sense of 3D and richer 3D video content.
Owner:HUAWEI TECH CO LTD

Method and Device for Learning Depth Estimation Based on View Synthesis

A method for controlling autonomous driving of a vehicle is introduced. The method may comprise, training, based on an inference depth and an inference pose, a synthetic image model for generating a synthetic image, generating, based on the synthetic image, a first virtual image to be associated with the original image, generating, based on the original image, a second virtual image, training a generative adversarial network (GAN) for determining, based on the original image, authenticity of the first virtual image and the second virtual image, training, based on the trained GAN, a depth network, wherein the trained GAN outputs a determination of the authenticity of the first virtual image, outputting, based on the trained depth network, signal, and controlling, based on the signal, autonomous driving of the vehicle.
Owner:HYUNDAI MOTOR CO LTD +1

Neural radiation field compression rendering method based on decomposition expression

The invention provides a neural radiation field compression rendering method based on decomposition expression, and relates to the technical field of computer graphics, and the method comprises the steps: inputting a multi-view image and camera internal and external parameters, carrying out the ray sampling, generating a space sampling point, carrying out the mixed feature coding of the space sampling point, and obtaining a multi-view image; obtaining three-dimensional voxel features, aligned and fused three-plane features and position codes, and splicing the three features to form fused features; inputting the fusion features into a factorization neural BRDF rendering network to predict volume density, geometric latent features, material parameters and reflection features, and improving the view angle color under the compression condition through a BRDF modulator; a ray weight is predicted through a plane-ray combined modeling module, and a compressed neural radiation field model is obtained through combined optimization of miniaturized body rendering, weighted reconstruction loss and compression constraint and is used for target view angle image rendering; according to the method, the storage overhead of the neural radiation field model is reduced, and meanwhile, the synthesis quality and rendering consistency of the new view angle under different compression ratios are improved.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Gaussian depth regularization new viewpoint synthesis method based on hierarchical mask

The invention discloses a Gaussian depth regularization new viewpoint synthesis method based on layered masks. Firstly, a layered depth optimization mechanism with foreground strong supervision and background weak constraint is constructed, and multi-view depth continuity and texture consistency collaborative constraint is introduced to realize stable optimization of three-dimensional Gaussian parameters. According to the method, monocular depth noise propagation is effectively inhibited, three-dimensional reconstruction precision and visual consistency are improved, training efficiency and engineering feasibility are considered, and the method is suitable for application scenes such as virtual reality, digital twinning and three-dimensional content generation; the innovation point of the invention is that the mask-guided hierarchical supervision and multi-view collaborative optimization mechanism solves the problem that the existing 3DGS lacks structural constraints, and has clear technical improvement; the method is based on a mature depth estimation and segmentation model, is high in technical maturity, and has an implementation basis.
Owner:NANJING UNIV OF POSTS & TELECOMM

Ray tracing volumetric particles for real-time novel view synthesis

Approaches presented herein provide for efficient rendering of high quality, novel views of a scene, in this case achieved through a combination of volumetric particle representations and ray tracing. An object can be represented using a set of volumetric particles (e.g., 3D distributions) that are aligned to the underlying structure or geometry of the object. Volumetric particles can be encapsulated in a bounding mesh or proxy geometry that can be used to efficiently compute ray-particle intersections. For a view to be rendered, ray tracing can be performed to determine an intersection of the rays with the proxy geometry. When a hit is determined, the precise intersection location with the volumetric particle is computed and the value of the distribution returned for that ray. If a ray passes through multiple semi-transparent volumetric particles then the color value is determined based upon the values returned from those particles.
Owner:NVIDIA CORP

View synthesis from images and / or video using model-based inpainting

Systems and techniques for image processing are described. For example, a computing device can receive a pose of a camera with a field-of-view (FOV) of a scene and can determine poses of the camera and / or pose(s) of other camera(s). The computing device can obtain first camera layers associated with the camera. The computing device can obtain, based on the pose of the camera and pixels within the first camera layers, a mask corresponding to the pose of the camera. The computing device can generate composited layers and can determine pixels for regions of the image using a first model based on the mask, the plurality of composited layers, the pose of the camera, and a plurality of images of the scene. The computing device can generate a final image of the scene corresponding to the pose of the camera based on providing the determined pixels to the regions.
Owner:QUALCOMM INC

Conditional image generation method based on point rasterization and related device

The invention provides a condition image generation method based on point rasterization and a related device, and the method comprises the steps: carrying out the coloring processing of a target point cloud set according to an environment image corresponding to a target moment; performing foreground and background separation processing on the target point cloud set after coloring processing to determine a dynamic foreground region and a static background region in the target point cloud set; complementing the incomplete area in combination with a historical point cloud set in the time window; projecting the complemented target point cloud set to an image coordinate system of the vehicle by using a perspective projection mapping relation to obtain a first projection image; and rendering the grating area in the first projection image to obtain a point rasterization condition image. According to the method, frame-by-frame analysis and image plane mapping are carried out on original laser radar point cloud data, point rasterization processing is combined, a geometric condition image with pixel-level precision is constructed, and structural prior support is provided for subsequent new view synthesis and controllable video generation.
Owner:BEIHANG UNIV

Layered view synthesis system and method

A method of computer-implemented synthesized view image generation and a synthesized view image generation system provide layered view synthesis. The method includes receiving an input image having a plurality of pixels having color values; generating a dilated depth map by dilating a depth map associated with the input image, the depth map with depth values respectively associated with each pixel in the input image; determining an inpainting mask using the dilated depth map; performing an inpainting operation based on the inpainting mask and the input image to generate a background image; and rendering a synthesized view image using the background image, the input image, and the dilated depth map.
Owner:LEIA INC

View synthesis with learned gaussing splatting and weighted sum rendering

A system generates initial Gaussian elements defined by parameter sets that include, for each Gaussian element a spherical harmonics (SH) coefficient array, a learnable parameter vector, and a learnable weight vector. The system performs a training process comprising rasterizing current Gaussian elements to generate a rendered image of the scene as viewable from a current camera position, wherein for each Gaussian element of the current Gaussian elements that intersects the camera ray, the system determines an opacity value for a location based on a view-dependent scaling value that depends on the current camera position, a position vector, and the learnable parameter vector.
Owner:QUALCOMM INC

Three-Dimensional Diffusion Models

Provided are systems and methods to perform novel view synthesis of a three-dimensional (3D) scene with a machine-learned diffusion model. Example implementations of the proposed models may be referred to as “3D Diffusion Models” or 3DiM. The models described herein can be or include an image-to-image diffusion model that takes one or more (e.g., a single) reference views and one or more (e.g., a single) relative poses as input and generates the target view. Thus, the machine-learned diffusion models described herein can perform novel view synthesis from as few as a single image.
Owner:GOOGLE LLC

System and method for efficient scene continuity in visual and multimedia using generative artificial intelligence

ActiveUS12499515B2Image enhancementPattern recognitionGenerative process
A system and method for generating multimedia artifacts with managed scene continuity in visual and multimedia using an AI-based and scene continuity aware media generation platform. The system receives a user or AI agent specification or simulation result(s), selects or trains generative models based on the specification, preprocesses relevant data, and generates scene narrative or frame-specific, sequence specific or broader continuity aware content using the selected or trained model(s). The generated content may be further enhanced using frame interpolation and view synthesis techniques to create smooth transitions or novel viewpoints or to aid in more efficient transmission or viewing or persistence of resultant content. The system enables efficient and customizable generation of high-quality scene continuity aware content for various applications in visual and multimedia production using neuro-symbolic and simulation enhanced compression, representation and generation processes.
Owner:QOMPLX INC

Self-supervised depth for volumetric rendering regularization

An example method includes generating embeddings of image data that includes multiple images, where each image has a different viewpoints of a scene, generating a latent space and a decoder, wherein the decoder receives embeddings as input to generate an output viewpoint, for each viewpoint in the image data, determining a volumetric rendering view synthesis loss and a multi-view photometric loss, and applying an optimization algorithm to the latent space and the decoder over a number of epochs until the volumetric rendering view synthesis loss is within a volumetric threshold and the multi-view photometric loss is within a multi-view threshold.
Owner:MASSACHUSETTS INST OF TECH +1

A method and related apparatus for generating conditional images based on point rasterization

The present application provides a kind of conditional image generation method and related device based on point rasterization, according to the environment image corresponding to target time, target point cloud set is colored processing;The target point cloud set after coloring processing is carried out front and background separation processing, to determine the dynamic foreground region and static background region therein;Combining the historical point cloud set in time window, the incomplete area is completed;Using perspective projection mapping relationship, the target point cloud set after completion is projected to the image coordinate system of vehicle, to obtain first projection image;The raster area in first projection image is rendered, to obtain the conditional image of point rasterization.It is analyzed frame by frame to original laser radar point cloud data and image plane mapping, and combined with point rasterization processing, constructs the geometric conditional image with pixel level precision, provides structure priori support for subsequent new view synthesis and controllable video generation.
Owner:BEIHANG UNIV

Ray tracing volumetric particle for real-time novel view synthesis

To provide a method for efficient rendering of high-quality, novel views of a scene.SOLUTION: A method is realized through a combination of volumetric particle representations and ray tracing. An object can be represented using a set of volumetric particles (e.g., 3D distributions) that are aligned to an underlying structure or geometry of the object. The volumetric particles can be encapsulated in a bounding mesh or proxy geometry that can be used for efficiently computing ray-particle intersections. For a view to be rendered, the ray tracing can be performed to determine an intersection of the rays with the proxy geometry. When a hit is determined, a precise intersection location with the volumetric particle is computed and a value of a distribution returns for that ray. If a ray passes through multiple semi-transparent volumetric particles, then a color value is determined based upon values returned from those particles.SELECTED DRAWING: Figure 5
Owner:NVIDIA CORP

Conditioned generative model for consistent novel views synthesis

Described is a computer apparatus (1400) configured to: obtain a reference image (202) of a 3D scene from a reference viewpoint; obtain a reference point map (203) corresponding to the reference image (202), the reference point map (203) comprising the coordinates of a plurality of pixels of the reference image (202); obtain a target point map (204) corresponding to the rendered image (206), the target point map (204) comprising the coordinates of a plurality of pixels of the rendered image (206); and generate the rendered image (206) in dependence on the reference image (202), the reference point map (203) and the target point map (204). In this way, the rendering process may be guided by the positions in the point maps.
Owner:HUAWEI TECH CO LTD +1

Ray-tracing volumetric particle for novel real-time viewing synthesis

The approaches presented here provide an efficient rendering of high-quality, novel views of a scene, in this case achieved through a combination of volumetric particle representations and ray tracing. An object can be represented using a set of volumetric particles (e.g., 3D distributions) aligned with the object's underlying structure or geometry. Volumetric particles can be encapsulated within a boundary mesh or proxy geometry, which can be used to efficiently compute ray-particle intersections. To render a view, ray tracing can be performed to determine a ray intersection with the proxy geometry. When a match is found, the precise location of the intersection with the volumetric particle is calculated, and the distribution value for that ray is returned.If a beam passes through several semi-transparent volumetric particles, the color value is determined based on the values ​​returned by these particles.
Owner:NVIDIA CORP

Aerial thermal infrared image positioning method, device and equipment based on view synthesis

The present application relates to a method, apparatus, and device for positioning aerial thermal infrared images based on view synthesis. The method comprises: constructing a query image set, a reference map set, and a ground control point set; the query image set includes multiple query images; calculating the initial pose of the query image using prior information acquired by the device sensor, rendering and synthesizing a pre-constructed three-dimensional reference model based on the initial pose and rendering software to obtain a reference image and a depth image of the reference image; matching feature points of the query image and the reference image based on a deep learning algorithm to construct a depth-based 2D-3D correspondence relationship, and using an n-point perspective algorithm to solve the motion between the three-dimensional point set and the two-dimensional point set with a 2D-3D correspondence relationship in the query image to obtain the final pose. This method can improve the accuracy of thermal infrared image positioning.
Owner:NAT UNIV OF DEFENSE TECH

Unsupervised monocular image-based online 3D scene reconstruction method and device

The present disclosure provides a monocular image-based unsupervised three-dimensional scene online reconstruction method and device. The method of the present disclosure comprises: obtaining a static background mask and a mask of each dynamic object instance of a current frame monocular image through semantic segmentation, obtaining a surround view synthesis image through a static multi-view generator, determining the motion parameters of each dynamic object instance through motion modeling, obtaining a multi-view depth map based on the surround view synthesis image through a depth estimation network obtained through self-supervised training, obtaining a local depth map of each dynamic object instance through a local depth estimation network obtained through self-supervised training, and obtaining a complete 3D scene representation through the multi-view depth map, the local depth map of each dynamic object instance, and the motion parameters thereof. The present disclosure can avoid true value dependence, effectively reduce hardware cost, and at the same time improve the reliability and robustness of monocular image three-dimensional scene reconstruction.
Owner:BEIJING TRUNK TECHNOLOGY CO LTD

Neural view synthesis using tiled multiplane images

A method for generating tiled multiplane images from a source image is disclosed. The method includes obtaining color and depth images. The method also includes extracting a feature map using a first neural network. The method also includes generating masks for a tile using a second neural network based on a corresponding tile of the feature map and corresponding sections of the color and depth images. The method also includes computing depths of a planes corresponding to the tile based on the masks. The method also includes generating a per-tile multiplane image for the tile based on the masks. The method also includes rendering an image using per-tile multiplane images and depths. A system for generating tiled multiplane images from a source image is also disclosed.
Owner:META PLATFORMS TECHNOLOGIES LLC

A new view synthesis method, system and device

The application relates to a new view synthesis method, system and device, obtains images and camera poses under multiple viewing angles, encodes to generate input labels to be processed; encodes to generate target labels based on target viewing angle poses; inputs the input labels of each viewing angle into an encoder composed of multiple layers of Transformer blocks, and the output of each layer of the Transformer blocks is saved in a key-value cache; inputs the target labels into a decoder for processing, and outputs label representations of target viewing angle images, the decoder has the same number of layers of Transformer blocks as the encoder, performs self-attention processing on input data as a query, uses the key-value cache of the corresponding layer in the encoder as a key and a value, performs cross-attention query operation, and the result is transmitted into a feedforward neural network to obtain a final output; and the label representations of the target viewing angle images are converted into RGB images under the target viewing angle. Compared with the prior art, the application can more accurately learn scene semantics of input views and rendering rules of target views, and generate high-quality new views.
Owner:FUDAN UNIVERSITY