Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1300 results about "Gaussian" patented technology

Gaussian /ˈɡaʊsiən/ is a general purpose computational chemistry software package initially released in 1970 by John Pople and his research group at Carnegie Mellon University as Gaussian 70. It has been continuously updated since then. The name originates from Pople's use of Gaussian orbitals to speed up molecular electronic structure calculations as opposed to using Slater-type orbitals, a choice made to improve performance on the limited computing capacities of then-current computer hardware for Hartree–Fock calculations. The current version of the program is Gaussian 16. Originally available through the Quantum Chemistry Program Exchange, it was later licensed out of Carnegie Mellon University, and since 1987 has been developed and licensed by Gaussian, Inc.

Three-dimensional scene reconstruction method and device based on large model geometric prior, and medium

The invention discloses a three-dimensional scene reconstruction method and device based on large model geometric prior, and a medium, and aims to solve the problems that a conventional 3DGS is liable to have artifacts and detail loss in geometric discontinuity, data redundancy and illumination variation scenes, and predicts a dense depth map and a normal map from a monocular image by using a pre-trained large model. The position and form of the Gaussian kernel are constrained as additional geometric priori; a primitive adjustment strategy based on kernel density estimation is introduced in the training stage, small Gaussian primitives with similar structures and adjacent spaces are combined into a large Gaussian primitive, the rendering quality is kept, redundancy is reduced, and the volume of the model is reduced; an exposure coefficient is adaptively estimated for each input image, an exposure compensation image loss function is constructed, and floating artifacts caused by illumination differences at shooting moments are eliminated. Experiments show that compared with the prior art, the method improves the three-dimensional reconstruction precision and real-time rendering quality of complex illumination and less-texture areas in a public data set and an unmanned aerial vehicle aerial photography scene.
Owner:NARI INFORMATION & COMM TECH

Short temporary rainfall prediction method and system based on multi-model random scheduling integration

The invention belongs to the technical field of rainfall prediction, and discloses a short and temporary rainfall prediction method based on multi-model random scheduling integration, which develops a robust training and pushing framework based on a continuous rolling prediction strategy, and decomposes long-sequence prediction into manageable stages. According to the method, training is carried out through teacher forcing and planned sampling, error propagation is relieved, and the training process is stabilized. The invention further designs asymmetric encoder-decoders (DSE and AFD) that achieve lower FLOPs than competitive baselines under standardized assessment, where DSE selectively compresses significant features and AFD stepwise reconstructs details to mitigate excessive smoothing problems. Finally, an intensity weighted Gaussian KL divergence loss function is designed, and the key problem of data balance is solved by modeling and predicting on a distribution level and endowing a large weight to a meteorological important heavy rainfall event.
Owner:YIBIN UNIV

Gaussian splatting with gradient-based pruning and semantically aware-robust optimization

PCT designated stageWO2025259820A13D modellingPattern recognitionComputer vision
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating a 3D representation of a scene. In particular, one of the methods includes obtaining a plurality of training images of a scene; initializing a three-dimensional (3D) representation of the scene, the 3D representation comprising a set of Gaussian distributions; and generating a final 3D representation of the scene, comprising, at each of a plurality of update steps, updating, using rendered images rendered using the set of Gaussians and corresponding training images, the set of Gaussian distributions. As part of the updating, gradient-based pruning, semantically-aware optimization, or both can be performed.
Owner:GDM HOLDING LLC

Methods and procedures for a one-way quantum channel authentication for secure quantum communication

The present technology pertains to systems and methods for one-way authentication of quantum channels. A transmitter generates an entangled quantum state comprising a first state and a second state, modulates the second state according to a clock-synchronized pattern, and transmits it through a quantum channel to a receiver. The first state is retained and measured at the transmitter to extract quantum-state information. Authentication is performed based on this information, without requiring feedback from the receiver. The quantum-state information may be derived using quadrature measurements, Gaussian tomography, or other statistical analyses to detect whether the second state underwent irreversible interactions such as eavesdropping or decoherence. The system enables secure unidirectional quantum authentication, reduces protocol complexity, and supports real-time anomaly detection. It is compatible with continuous-variable quantum states, quantum key distribution (QKD), and scalable communication architectures.
Owner:EIGENQ INC

Lightweight three-dimensional reconstruction method and system based on three-dimensional Gaussian model

The invention relates to the technical field of three-dimensional reconstruction, in particular to a lightweight three-dimensional reconstruction method and system based on a three-dimensional Gaussian model, and the method comprises the steps: constructing the three-dimensional Gaussian of a target scene based on a multi-view original image of the target scene through employing a 3DGS algorithm; identifying the minimum resolution of each Gaussian primitive in the three-dimensional Gaussian, setting a spherical region according to the minimum resolution, taking the intersection result of the spherical region and the Gaussian primitive as a region evaluation score, and trimming the redundancy region geometry in the three-dimensional Gaussian in combination with the opacity; and taking the maximum zoom scale of the Gaussian primitive as a spatial density proxy variable, and dynamically distributing spherical harmonic coefficient orders according to a local spatial density proxy variable so as to meet the rendering requirements of different three-dimensional Gaussian geometric regions. On the premise that the image synthesis quality is not lost, the model scale and the transmission overhead can be remarkably compressed, the resource constraint condition of the edge computing equipment can be effectively adapted, and the higher training speed and the better rendering effect are achieved.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Dynamic gaussian splatting learned from hierarchical motion model

Some embodiments of a method may include: obtaining a reference 3D Gaussian frame, a camera position C, and a time t; extracting a multi-scale feature for each 3D Gaussian of one or more 3D Gaussians using a neural network block, wherein the multi-scale feature represents multi-scale spatial information about a dynamic object or scene; predicting 3D motion based on the multi-scale features and the time t; predicting a 3D Gaussian frame for time t by manipulating the one or more 3D Gaussians in a spatial domain based on the predicted 3D motion; and outputting the 3D Gaussian frame for time t.
Owner:INTERDIGITAL VC HOLDINGS INC

Volume cloud rendering method and system based on three-dimensional Gaussian splashing

The invention discloses a volume cloud rendering method and system based on three-dimensional Gaussian splash, and the method comprises the steps: carrying out the feature matching and structure reconstruction of a volume cloud image through an SfM algorithm, obtaining a sparse point cloud in a scene, and representing the sparse point cloud as an anisotropic Gaussian ellipsoid; decomposing the illumination transmission process of the volume cloud into two parts of forward single scattering and internal multiple scattering based on a radiation transmission equation to obtain an illumination modeling formula embedded volume cloud rendering process; determining the contribution of each three-dimensional Gaussian ellipsoid to a rendering result according to the opacity index of the three-dimensional Gaussian ellipsoid, and dynamically cutting the low-contribution three-dimensional Gaussian ellipsoid by adopting a delay deletion strategy to reduce redundancy; a comprehensive loss function is constructed, so that the three-dimensional Gaussian model is focused on the volume cloud region and the structure detail reduction capability is enhanced; and outputting the optimized volume cloud rendering result and the three-dimensional Gaussian model. According to the method, the volume cloud rendering efficiency and quality are effectively improved on the premise of ensuring real-time rendering.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Three-dimensional reconstruction method and device based on aerial image

The invention discloses a three-dimensional reconstruction method and device based on an aerial image. The method comprises the following steps: collecting a plurality of aerial images of a to-be-reconstructed region and corresponding position data and attitude data; generating an initial three-dimensional point cloud according to the plurality of aerial images and the corresponding position data and attitude data; and converting the initial three-dimensional point cloud into a Gaussian original body, carrying out iteration on the Gaussian original body to obtain a Gaussian sputtering three-dimensional model, and generating a three-dimensional reconstruction result by using the Gaussian sputtering three-dimensional model. According to the scheme, through combination of traditional point cloud reconstruction and Gaussian sputtering modeling, efficient conversion from the aerial image to the high-precision three-dimensional model is realized. The method not only improves the reconstruction precision of the three-dimensional model, but also has high stability and calculation efficiency when processing complex terrains and large-scale scenes.
Owner:HANGZHOU JINGAN TECH CO LTD

Rapid high-fidelity reconstruction method for automatic driving scene

The invention discloses a rapid high-fidelity reconstruction method for an automatic driving scene, and belongs to the technical field of computer vision. The invention aims to solve the problems of excessive model parameters and low rendering efficiency in the existing automatic driving scene reconstruction process. Comprising the following steps: decomposing a dynamic driving scene into a static background and a dynamic foreground; establishing a three-dimensional Gaussian foreground model under a local coordinate system, converting the three-dimensional Gaussian foreground model to a world coordinate system, and splicing the three-dimensional Gaussian foreground model with a three-dimensional Gaussian background model under the world coordinate system to obtain a whole scenic spot cloud; carrying out densification pruning and refining pruning to obtain a final full-scene spot cloud; and an ETB-Box method and an MITS method are adopted to optimize a three-dimensional Gaussian rendering pipeline of the final point cloud of the whole scene, a three-dimensional Gaussian accurate tile corresponding to each point cloud is calculated, the three-dimensional Gaussian is rendered, and high-fidelity reconstruction of the automatic driving scene is realized. According to the invention, rapid high-fidelity reconstruction of the automatic driving scene is realized.
Owner:HARBIN INST OF TECH +1

Online performance testing method for breather valve for oil and gas storage and transportation

The invention relates to the technical field of sealing performance testing, in particular to a breather valve performance online testing method for oil and gas storage and transportation, which comprises the following steps: acquiring a vibration signal of a breather valve body, a multi-band acoustic signal of a valve port and a total pressure signal in a storage tank; performing multi-scale complex wavelet decomposition, and constructing a time frequency-energy correlation feature tensor; the total pressure signal and the time frequency-energy correlation characteristic tensor serve as a combined observation value and are input into a preset continuous Gaussian mixture hidden Markov model containing four hidden states of sealing, transient micro-leakage, continuous leakage and full-amount opening, and the posterior probability is calculated; and calculating to obtain a real-time leakage rate. According to the method, the opening pressure can be accurately determined, the real-time leakage rate can be quantitatively calculated after the leakage state is recognized, comprehensive and accurate quantitative online evaluation of core performance parameters of the breather valve is achieved, and the multi-source information fusion degree and the anti-interference capacity are improved.
Owner:TAICANG YANGHONG PETROCHEMICAL CO LTD

Point cloud reconstruction method and system based on three-dimensional Gaussian sputtering

The invention discloses a point cloud reconstruction method and system based on three-dimensional Gaussian sputtering, and is used for solving the technical problem that the structural stability of a final three-dimensional point cloud model is not good enough due to the fact that a traditional point cloud reconstruction method causes gradient propagation abnormity, and the optimization process is not convergent or falls into a local minimum value. The method comprises the following steps: firstly, acquiring a multi-view image and a reference view image, and constructing a covariance degradation risk probability graph; generating a plurality of Gaussian three-dimensional points to be regulated and controlled, performing three-dimensional point screening and fitting credibility score calculation, outputting secondary regulation three-dimensional points and corresponding scores, and constructing an initial three-dimensional point cloud model; performing secondary adjustment on the three-dimensional point by combining the multi-view image and fractional optimization to obtain a target Gaussian three-dimensional point, and updating the initial model into an intermediate model; and updating the intermediate model through a covariance updating gating mechanism based on gradient convergence dynamic monitoring, and outputting a target three-dimensional point cloud model.
Owner:FOSHAN POWER SUPPLY BUREAU GUANGDONG POWER GRID

Transient electromagnetic and magnetic method data joint inversion method

The invention discloses a transient electromagnetic and magnetic method data joint inversion method. The method comprises the following steps: acquiring transient electromagnetic observation data and magnetic method observation data; converting the transient electromagnetic observation data into transient electromagnetic moment data; based on the transient electromagnetic moment data and the magnetic method observation data, a joint inversion objective function is constructed, and the joint inversion objective function comprises a data fitting residual term and a model constraint term; performing parameter coupling on the model constraint term by using a multivariate Gaussian mixture model to obtain a coupled model constraint term; and combining the data fitting residual term and the coupled model constraint term into a final objective function, and minimizing the final objective function through an optimization algorithm to obtain an electrical parameter and a magnetic parameter of the underground medium. According to the method, transient electromagnetic and magnetic method data joint inversion is realized, the problem of low three-dimensional inversion calculation efficiency is solved by introducing transient electromagnetic moment, and electrical and magnetic parameters are coupled by utilizing rock physics constraint, so that the inversion result is more accurate, and the structure is more consistent.
Owner:INSTITUTE OF GEOLOGY AND GEOPHYSICS CHINESE ACADEMY OF SCIENCES

Three-dimensional scene reconstruction method and system based on monocular depth estimation

The invention discloses a three-dimensional scene reconstruction method and system based on monocular depth estimation, and belongs to the technical field of computer vision and three-dimensional reconstruction, and the method comprises the steps: carrying out the multi-scale feature coding of a monocular RGB image through a mixed attention depth coding module, and obtaining the hierarchical depth feature representation; carrying out autoregression depth decoding through a self-adaptive edge perception depth decoding module to generate an initial depth map; a depth confidence map is calculated through a geometric consistency constraint optimization module and is fed back to a coding module for iterative optimization, and a refined depth map is output; and three-dimensional Gaussian ellipsoid scene representation is constructed through the Gaussian ellipsoid scene reconstruction module. According to the invention, high-precision depth estimation and high-quality three-dimensional reconstruction are realized by constructing a depth-coupled closed-loop cooperative system.
Owner:HARBIN INST OF TECH

Digital human rendering method based on Gaussian splashing and multi-scale characteristic field distillation

The invention discloses a digital human rendering method based on Gaussian splashing and multi-scale characteristic field distillation, and belongs to the field of three-dimensional human body digital reconstruction. According to the method, feature extraction is carried out through a Vision Transformer encoder, based on an SMPL model, through a cross-modal parameter estimation module and dynamic human body modeling of three-dimensional Gaussian splashing and semantic feature rendering, three-dimensional Gaussian is projected to a two-dimensional image plane to calculate a covariance matrix and color mixing, and after a rendered color image and an initial feature field image are output, a three-dimensional image is obtained. A student feature map is obtained through a convolution acceleration module, a teacher feature map is obtained after feature extraction is carried out through a two-dimensional basic model, the constructed model is trained, and optimization training is completed through comprehensive total loss function calculation; according to the method, the problems of fuzzy semantics, detail missing, low rendering efficiency and inaccurate human body-scene separation in the existing method are effectively solved, so that more efficient, fine and robust monocular or multi-view human body three-dimensional reconstruction is realized.
Owner:YUNNAN UNIV

Layered densification Gaussian sputtering method based on visibility

The invention discloses a layered densification Gaussian sputtering method based on visibility, and relates to the technical field of artificial intelligence and computer vision. According to the scheme, an initial three-dimensional Gaussian primitive set is generated based on sparse multi-view observation data, and initial scene representation is established by extracting spatial distribution parameters, morphological parameters and radiation parameters; performing fusion analysis on the Gaussian primitives based on the multi-dimensional observability parameter set to obtain comprehensive observability index data; executing hierarchical clustering according to the index data, and constructing a multi-layer Gaussian structure of a significant layer, a transition layer and a background layer; performing density enhancement, geometric continuity constraint and parameter update processing on different levels of Gaussian structures, and generating a rendered image of a target view angle under a volume light traveling and transparency hybrid mechanism; according to the method, continuous reconstruction of a scene structure and accurate expression of radiation characteristics can be realized under the sparse view condition, and the geometric fidelity and rendering consistency of new view angle synthesis are improved.
Owner:HENAN JINSHU INTELLIGENT TECH CO LTD

Deep foundation pit crack feature coding detection method based on image recognition

The invention discloses a deep foundation pit crack feature coding detection method based on image recognition, and aims to solve the problems that a two-dimensional image is difficult to accurately characterize a three-dimensional crack, and the opening width and the slab staggering height are inaccurately estimated, the deep foundation pit crack feature coding detection method comprises the following steps of: calibrating and collecting through short time sequence and multi-polarization imaging, segmenting the crack to extract a center line, and dividing a left side domain and a right side domain; carrying out local reconstruction in a mask neighborhood by adopting a nerve radiation field or three-dimensional Gaussian sputtering, establishing a left implicit surface and a right implicit surface, fusing a polarization solution line, applying occupancy priori and seam-crossing repulsive potential in a gap region, combining a reference object to obtain an absolute scale, measuring a millimeter-level opening and a dislocation along a central line section, and generating a feature code; and threshold value and time sequence comparison judgment is carried out, and the technical effects of high-precision three-dimensional measurement, reliable risk assessment and traceable management of the deep foundation pit crack are achieved.
Owner:CHEM IND GEOTECHN ENG

Physical attribute inversion and three-dimensional reconstruction method based on Gaussian splashing and micro rendering

The invention discloses a physical attribute inversion and three-dimensional reconstruction method based on Gaussian splashing and micro-rendering, and belongs to the technical field of computer vision, and the method comprises the steps: generating a sparse point cloud based on a multi-view image to initialize a Gaussian ball; determining a Gaussian pivot point based on a Gaussian ball, constructing a self-adaptive tetrahedral mesh in combination with gradient optimization, extracting an explicit surface from the self-adaptive tetrahedral mesh, and sampling to obtain geometric information; predicting the physical attribute of each sampling point through an implicit texture network, and providing ambient light through an HDR ambient light map; geometric information, physical attributes and ambient light are used as input, a differentiable PBR renderer is adopted to generate a physical rendering image, multi-term loss function reverse optimization is constructed, and collaborative optimization and inversion of three-dimensional geometry and surface physical attributes are realized. According to the method, the problems of texture dislocation, thin-wall structure disappearance and high rendering calculation cost under dynamic topology are effectively solved, and a digital model with fine geometry and a realistic material can be efficiently recovered from a multi-view image.
Owner:ZHEJIANG UNIV

Laser radar camera calibration method and device and medium

The invention relates to a laser radar camera calibration method and device, and a medium. The method comprises the steps: collecting multi-frame laser radar point cloud data and synchronous corresponding image data in a construction scene; carrying out multi-frame point cloud fusion and dense reconstruction to generate a laser dense point cloud; visual sparse point cloud reconstruction is carried out, and cross-modal scale unification and space alignment are carried out on an initial visual point cloud obtained through reconstruction and the laser dense point cloud; generating a visual dense point cloud through three-dimensional Gaussian splashing; and performing registration on the laser dense point cloud and the visual dense point cloud by adopting a point-to-line iterative nearest point algorithm, and performing calculation to obtain an external parameter calibration matrix of the camera and the laser radar. Compared with the prior art, the method has the advantages of high precision, low cost, high stability and the like.
Owner:SHANGHAI TONGJI INDEPENDENT INTELLIGENT UNMANNED SYSTEMS RESEARCH INSTITUTE +1

Remote sensing image three-dimensional reconstruction method based on semantic information

The invention discloses a remote sensing image three-dimensional reconstruction method based on semantic information, and relates to the technical field of remote sensing image three-dimensional modeling, and the method comprises the steps: carrying out the processing of a remote sensing image covering a target region, obtaining point cloud data, and carrying out the semantic segmentation based on a semantic segmentation model, and determining a semantic category label; dividing voxel units based on a three-dimensional sparse point cloud distribution condition, and constructing to obtain an initial anchor point; the distribution density of the initial anchor points is adjusted in combination with the semantic category labels, and the adjusted initial anchor points are obtained; iteratively training the three-dimensional Gaussian splash model for multiple times to update the anchor points to obtain scene anchor points; and obtaining a three-dimensional reconstruction model of the target area based on scene anchor point rendering. According to the method, semantic information is introduced and a 3DGS three-dimensional modeling technology is fused, so that the densification quality of anchor points and the geometric boundary definition of the model are effectively improved, the modeling efficiency can be improved, and the structural rationality and semantic interpretation of a three-dimensional reconstruction result can be enhanced.
Owner:WUHAN UNIV

Method and system for accurately calculating axial pretightening force of ultrasonic bolt

The invention provides an accurate calculation method and system for the axial pretightening force of an ultrasonic bolt, and the method combines a high-order statistic theory with a nonlinear filtering technology, and solves the problem that the measurement precision of a conventional method is insufficient under the conditions of strong noise and nonlinearity. Particle filtering and Kalman filtering are combined, so that the problem that a traditional Kalman filtering method easily causes sensitivity of an initial value of filtering divergence is solved; a double-Gaussian attenuation model is adopted to fit an echo envelope, the problem that a single Gaussian or single index model cannot adapt to envelope distortion caused by stress change is solved, and a nonlinear state space model is established to adapt to nonlinear echo signal characteristics under dynamic stress; gaussian noise is effectively suppressed by using a third-order cumulant cross-correlation algorithm, the time difference resolution in a low signal-to-noise ratio environment is remarkably improved, and the Gaussian noise and asymmetric interference are effectively suppressed.
Owner:SHANDONG UNIV

Gaussian point cloud lossless coding and decoding method for three-dimensional reconstruction

PendingCN121126005ADigital video signal modificationLossless codingPoint cloud
The invention relates to a three-dimensional reconstruction-oriented Gaussian point cloud lossless coding and decoding method, a product and a storage medium. The method comprises the following steps: extracting a first Gaussian point cloud corresponding to anhor data in a numpy array format at a coding side; converting the first Gaussian point cloud in the numpy array format into a first Gaussian point cloud in a ply format; meanwhile, through a Morton code-based spatial sorting method, performing Morton code sorting on geometric data and attribute data of the first Gaussian point cloud in the numpy array format, and generating a Morton sequence index table; and carrying out AVS coding on the first Gaussian point cloud in the ply format to obtain compressed data in a binary code stream form. On the decoding side, AVS decoding is carried out when the compressed data in the binary code stream form and the Morton sequence index table are received, and a second Gaussian point cloud in the ply format is obtained through reduction; performing format conversion on the second Gaussian point cloud in the ply format to obtain a second Gaussian point cloud in a numpy array format; and rearranging the attribute data and the geometric data based on the Morton sequence index table to realize one-to-one correspondence of the geometric data and the attribute data between the first Gaussian point cloud and the second Gaussian point cloud before and after coding and decoding. Compared with an existing method, the data consistency before and after coding is improved.
Owner:GUANGDONG UNIV OF TECH

Dynamic digital twinning three-dimensional Gaussian splash rendering system and method

The invention discloses a dynamic digital twinning three-dimensional Gaussian splash rendering system and a dynamic digital twinning three-dimensional Gaussian splash rendering method. According to the system, a static scene is reconstructed by utilizing a three-dimensional Gaussian splashing technology through a semantic scene reconstruction module, semantic information is associated to Gaussian primitives in combination with spatial alignment and semantic injection, and a model with a semantic identifier is generated; multi-source dynamic data are fused and converted into a dynamic data field which can be accessed by a GPU through a multi-mode spatio-temporal data processing module; and through a rendering attribute dynamic modulation module, screening a target primitive according to the semantic identifier, and directly modulating internal rendering attributes such as color, transparency or shape of the target primitive based on a data field query value, thereby realizing internal rendering and visualization of dynamic data. According to the method, the problems that in the prior art, rendering reality and real-time performance are difficult to give consideration to and data and scene fusion is superficial are solved, and dynamic digital twinning rendering with high reality and strong immersion is realized.
Owner:XIAMEN UNIV ARCHITECTURAL DESIGN & RES INST CO LTD

Chip side digital bar code identification and data tracing method and system

The invention relates to a chip side digital bar code identification and data tracing method and system. The method comprises the steps of detecting a chip side contour, calculating an inclination angle and correcting a chip attitude; adjusting the fusion weight of the multi-view image and the exposure time, and screening qualified images; reflecting noise is suppressed through a self-adaptive OTSU block segmentation algorithm and dynamic Gaussian filtering, and characters are corrected to be in a horizontal state through affine transformation; an improved CRNN is adopted, sequence recognition is achieved through an improved ResNet50 network, an edge attention mechanism, Bi-LSTM and CTC decoding, and the validity of a character sequence is verified in combination with a Luhn algorithm; a chip bar code ID is used as a core to match multi-source data, and meanwhile, multi-database collaborative storage and block chain evidence storage are adopted to ensure data credibility; and finally, a quality improvement closed loop is formed through intelligent tracing and decision application. According to the invention, the chip side bar code identification accuracy is effectively improved.
Owner:NANJING YUNTONG TECH CO LTD

Self-supervised aerial view perception method fusing Gaussian spattering and time sequence modeling, electronic equipment and readable storage medium

The invention relates to the technical field of computer vision and automatic driving, in particular to a self-supervised aerial view perception method fusing Gaussian spattering and time sequence modeling, electronic equipment and a readable storage medium, and the method comprises the following steps: S1, constructing a BEV model; s2, feature grid mapping is carried out; s3, rendering a two-dimensional image; s4, performing self-supervised learning and optimization; s5, downstream task application; according to the method, BEV features are mapped into three-dimensional Gaussian parameters, end-to-end self-supervised learning is achieved through differential rendering, the method is applied to downstream tasks, and three-dimensional target detection, semantic segmentation or occupancy prediction are carried out; downstream tasks are trained and reasoned through self-supervision loss, no manual data annotation is needed, the data construction cost is remarkably reduced, and meanwhile the perception performance and generalization ability of the model in an automatic driving scene are improved.
Owner:JINING UNIV +1

Defect detection method and system based on adaptive double-domain filtering and Gaussian mixture prior constraint, medium and equipment

The invention relates to the field of computer vision, and discloses a defect detection method, system, medium and equipment based on adaptive dual-domain filtering and Gaussian mixture prior constraint, and the method comprises the steps: carrying out the multi-scale feature extraction of an ultrasonic C-scan image through a Vision Transform network after the ultrasonic C-scan image is preprocessed; respectively inputting the shallow fusion features and the deep fusion features into a frequency-space double-domain adaptive feature filtering module for filtering, inputting the filtered features into a Gaussian mixture modeling module Ada-GMM, and modeling normal feature distribution; carrying out Ada-GMM-Guided decoding, and carrying out interactive fusion on the corresponding deep semantic features and shallow texture features by adopting a deep and shallow multi-scale feature interaction mechanism; performing optimization by adopting cosine reconstruction loss, filtering consistency, entropy regularization loss and distribution alignment loss, and adaptively learning normal distribution characteristics according to an optimization process to obtain a model weight; and reasoning the input ultrasonic C-scan image by using the trained network weight to realize anomaly detection and positioning, and outputting an interpretable anomaly thermodynamic diagram.
Owner:UNIV OF CHINESE ACAD OF SCI

Camera pose estimation method based on 2D Gaussian splashing

The invention provides a camera pose estimation method based on 2D Gaussian splash. The camera pose estimation method comprises the following steps: acquiring a depth map and a normal map from a training image pose based on a pre-trained 2D Gaussian splash model; acquiring a grid ray starting point based on the depth map; generating a main ray and a hemispherical ray for each sampling point based on the normal diagram and the ray starting point; constructing a ray and image matching network and a lightweight convolutional neural network, and performing network training through a loss function; estimating the position of the camera through the trained ray and image matching network and the lightweight convolutional neural network based on the ray features of the main ray and the hemispherical ray and the image features of the query image, and constructing a rotation matrix to realize initial camera pose estimation; and optimizing the initial camera pose by adopting a depth-guided pose optimization method to obtain a final optimized camera 6D pose estimation result. The method gets rid of the dependence of an initial value, and has the advantages of high precision, strong robustness, efficient calculation and wide application.
Owner:JIANGSU UNIV

Dynamic Gaussian digital human image rendering method and device, equipment and storage medium

The invention discloses a dynamic Gaussian digital human image rendering method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: determining a target multi-view image frame from a multi-view RGB video image sequence according to a preset attitude, and constructing a deformable parameterized model according to the target multi-view image frame; constructing an attitude space driving attitude corresponding to the multi-view RGB video image sequence based on a three-dimensional attitude estimation technology, and generating a two-dimensional position map according to the attitude space driving attitude and the deformable parameterized model; performing three-dimensional Gaussian binding on the deformable parameterized model to obtain a local attribute of the three-dimensional Gaussian; training is carried out according to the two-dimensional position map, and a target StyleUNet neural network is obtained; and predicting the new attitude through the target StyleUNet neural network to obtain a dynamic Gaussian digital human image. In this way, the high-fidelity drivable high-frequency detail digital human image can be automatically rendered.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Using Game Metadata to Animate User-Generated Object in Video Game

A technique for generating, from a video from a computer game, a three-dimensional (3D) representation of space in which Gaussians represent objects in the video. Metadata from the game can be used in creating the 3D representation. User-input content such as a hand-drawn game path is inserted into the 3D representation of space and aligned and scaled. The opacity of the Gaussians in the 3D representation of space is then set to zero such that Gaussians representing objects in the video are transparent and only one or more portions of the user-input content are not transparent. The 3D representation of space is then combined with the video so that the user-input content is presented with the video and animated according to the metadata.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Sparse view angle three-dimensional reconstruction method and system based on voxel grid constraint

PendingCN121527352A3D modellingVoxelAlgorithm
The invention discloses a sparse visual angle three-dimensional reconstruction method and system based on voxel grid constraint, and belongs to the technical field of visual three-dimensional reconstruction. Constructing a three-dimensional voxel grid of a self-adaptive scene scale based on point cloud distribution, and dividing the point cloud to the corresponding voxel grid; fusing the geometric features of the fast point feature histogram and the confidence score in the grid, generating a geometric confidence comprehensive measure, and screening key points to initialize Gaussian primitives; designing a voxel grid constrained gradient clipping strategy, limiting Gaussian primitive error diffusion through a distance attenuation coefficient, and adaptively optimizing grid distribution in combination with a dynamic grid deletion and addition mechanism; and finally, carrying out iterative training by using a 3D Gaussian splash radiation field loss function to realize high-fidelity static scene reconstruction under a sparse view angle. According to the method, scene geometric priori is introduced, an optimization strategy based on voxel grid constraint is designed to effectively control excessive diffusion or drift of Gaussian primitives, generation of artifacts is reduced, and meanwhile, the situation that robustness is reduced due to the influence of priori quality is avoided.
Owner:BEIJING INST OF TECH

Object attitude estimation method based on scene-level semantic three-dimensional Gaussian splash

The invention relates to an object attitude estimation method based on scene-level semantic three-dimensional Gaussian spatter, which comprises the following steps of: firstly, constructing scene-level three-dimensional representation according to a multi-view image, and associating semantic embedding of three-dimensional Gaussian points of each target object; positioning a target mask area of a target object described by a language instruction according to the query image, and extracting each target three-dimensional Gaussian point subset through semantic matching and three-dimensional space clustering; secondly, an ICP registration method is guided through two-stage learning, and the initial 6D pose of the target object is obtained; and finally, performing cascade fine optimization by combining pose rendering and similarity comparison under disturbance to obtain a 6D pose of the target object in the target scene. According to the design scheme, target retrieval and instance extraction under open vocabularies are achieved through three-dimensional Gaussian reconstruction of semantic enhancement, and the attitude estimation precision under the conditions of shielding and low texture in a complex scene is effectively improved by combining language prompt and registration and rendering fine optimization guided by two-stage learning.
Owner:SOUTHEAST UNIV