Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

112 results about "Monocular video" patented technology

Processing monocular videos using three-dimensional gaussian splatting

The present disclosure describes techniques for processing monocular videos using three-dimensional gaussian splatting (3DGS). Spatial decomposition and temporal decomposition are performed on a monocular video to generate a plurality of clips. A first set of 3DGS representing foreground objects in each of the plurality of clips are initialized and optimized. A second set of 3DGS representing background in each of the plurality of clips are initialized and optimized. Two images are generated for each frame comprised in each of the plurality of clips based on the first set of 3DGS and the second set of 3DGS, respectively. Two images are merged to generate a resulting image for each frame in each of the plurality of clips. The resulting image accurately represents a corresponding frame in the monocular video.
Owner:LEMON INC(GB)

Three-dimensional human body posture estimation method and system

The invention discloses a three-dimensional human body posture estimation method and system, and relates to the technical field of computer vision, and the method comprises the steps: extracting two-dimensional human body posture key points from a monocular video picture sequence, and generating a two-dimensional human body posture key point sequence; projecting the two-dimensional key point sequence to a feature space through nonlinear high-dimensional mapping to generate a high-dimensional feature space matrix; and inputting the high-dimensional feature matrix into a three-dimensional human body posture recognition model fusing motion constraints and frequency division spatial-temporal features to obtain a three-dimensional human body posture key point sequence, and realizing three-dimensional human body posture estimation through three-dimensional coordinates. According to the method, the robustness and the detection precision of the monocular three-dimensional human body posture estimation method are improved. Error values of relative movement speed, skeleton length and skeleton direction of the key points are calculated, so that training is easier to converge, and the training process is more stable.
Owner:TONGJI UNIV

Monocular depth guided object level NeRF reconstruction method

The invention relates to the technical field of three-dimensional reconstruction, and discloses a monocular depth guided object-level NeRF reconstruction method, which comprises the following steps: firstly, through monocular video sequence input, generating a frame-by-frame initial depth map by using a depth estimation module, and estimating a relative camera attitude between adjacent frames through a relative attitude estimation module; calculating the absolute attitude of the camera in combination with the initial depth map and the relative camera attitude; then constructing a geometrically enhanced NeRF model, and optimizing scene representation through multi-resolution hash position coding and spherical harmonic direction coding; introducing photometric loss, depth contrast loss and density loss to jointly optimize parameters of the NeRF model, and constraining geometric reconstruction of the object by using depth information; and finally, through a four-stage iterative training strategy, alternately optimizing depth estimation, a camera attitude and a NeRF model, and generating an object-level controllable three-dimensional model. According to the method, depth information is fully utilized to optimize the NeRF training process, so that end-to-end reconstruction from a monocular video to an object-level controllable model is completed.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Monocular video scene dynamic three-dimensional reconstruction method based on optical flow

The invention discloses a monocular video scene dynamic three-dimensional reconstruction method based on optical flow, and relates to the technical field of scene reconstruction. Calculating an optical flow of each pixel in each frame of image in the monocular video, and determining a dynamic region mask of the image; according to the dynamic region mask of each frame of image, determining a plurality of dynamic object instances with consistent time and space; performing four-dimensional Gaussian sputtering conversion on each frame of image to obtain four-dimensional Gaussian distribution representation; for any one dynamic object instance in any one frame of image, acquiring other images containing the dynamic object instance in different image frames, and according to the coordinates of the Gaussian point clouds of the image and other images, generating a visual angle point cloud, which is not observed in the image, of the object instance; and according to color information of each frame of Gaussian point cloud in the four-dimensional Gaussian distribution representation, performing color rendering on each frame of Gaussian point cloud generating the visual angle point cloud to obtain a reconstructed scene. The method can accurately realize scene reconstruction.
Owner:XIAN FANGJU XINGCHEN TECHNOLOGY CO LTD

Real-time digital human generation method and system

The invention relates to a real-time digital person generation method and system, and the method comprises the following steps: obtaining a monocular video of a target person, and extracting 3DMM information in the monocular video; performing Gaussian point initialization on the monocular video according to the 3DMM information to obtain Gaussian parameters in a standard space; extracting voice audio features in the monocular video, and inputting the voice audio features into the audio-motion model to obtain a universal face key point motion sequence; converting the universal face key point motion sequence into a target face key point motion sequence through a projection algorithm; inputting the target face key point motion sequence and the Gaussian parameters in the standard space into a face key point guided Gaussian deformation network to obtain Gaussian deformation parameters; and rendering the Gaussian deformation parameter into a corresponding video frame through a Gaussian rasterizer, thereby obtaining a digital person video of the target person. Compared with the prior art, the digital human video generation accuracy and creation flexibility are improved.
Owner:SHANGHAI UNIV

Monocular video three-dimensional human body high-quality reconstruction method based on Gaussian splashing and normal perception

The invention relates to a monocular video three-dimensional human body high-quality reconstruction method based on Gaussian splashing and normal perception, and belongs to the field of computer aided design, graphics and computer vision. The method comprises the following steps: firstly, based on a monocular video frame in an input three-dimensional human body data set, guiding posture deformation through a time sequence, and carrying out whole body correlation deformation on a standard space posture to obtain an observation space human body posture; secondly, through normal perception Gaussian optimization, adaptive Gaussian density control and human body normal mapping supervision are carried out on the human body posture in the observation space, and a human body model of Gaussian representation is obtained; then, carrying out human body shadow feature learning on the input human body normal map by adopting an illumination model to obtain human body shadow features; and finally, light and shadow enhancement human body rendering is performed on the Gaussian representation human body model in combination with the human body light and shadow features and human body color information input into the monocular video, a human body rendering effect is output, and the quality and reality sense of three-dimensional human body reconstruction of the monocular video are improved.
Owner:KUNMING UNIV OF SCI & TECH

Dynamic human body reconstruction method and system based on double-motion embedding and point cloud fusion

The invention discloses a dynamic human body reconstruction method and system based on double-motion embedding and point cloud fusion, and relates to the field of digital human reconstruction processing. According to the method, the input frame sequence composed of three adjacent frames is intercepted from the monocular video and preprocessed to obtain necessary input parameters, then the dynamic human body reconstruction model is designed to process the input frame sequence and the input parameters to obtain the human body reconstruction image, the overall operation is fast and convenient, and the use effect is good. According to the dynamic human body reconstruction model designed by the invention, on one hand, a DMPF network part is adopted, dual-motion embedding is adopted to extract multi-modal motion features, and efficient fusion of 2D and 3D motion features is realized, so that richer motion information supervision is obtained, and on the other hand, a mixed point cloud encoder is adopted to fuse isolated point and overall point cloud features, so that the dynamic human body reconstruction model is more efficient in motion information supervision. Therefore, the dependency relationship of the local change of the human body surface on the global change is captured, and correct modeling of geometry and texture of the moving human body is further enhanced.
Owner:ANHUI UNIV

Three-dimensional human body and scene interaction reconstruction method and system

The invention belongs to the technical field of three-dimensional reconstruction, and particularly discloses a three-dimensional human body and scene interaction reconstruction method and system. The method is a three-dimensional human body and scene interactive reconstruction method for monocular video input, not only can quickly complete three-dimensional Gaussian reconstruction of a human body, a static scene and a moving object, but also can quickly complete the three-dimensional Gaussian reconstruction of the human body, the static scene and the moving object by introducing technical means such as posture correction, camera external parameter space regularization, joint reconstruction consistency loss and layered fusion rendering. And accurate alignment of the human body, the scene and the object in spatial positions, shielding relations and illumination styles is ensured, and dynamic interaction reconstruction conforming to real physical logic is realized. Finally, a joint rendering image with high fidelity, continuous time sequence and consistent structure can be generated, a stable and reliable three-dimensional expression capability is provided for complex human body actions, object interaction and scene understanding, and the three-dimensional interaction reconstruction quality and application value under the monocular video condition are greatly improved.
Owner:NANJING UNIV OF SCI & TECH

Methods and systems for real time video driven human 3-d posture estimation

PendingUS20250336236A1Image enhancementImage analysisEngineeringModal method
The disclosure relates generally to methods and systems for real time video driven human 3-dimensional (3-D) posture estimation during physical activities. Conventional techniques do not exploit temporal information, they do not give smooth transition of postures over time. Furthermore, the techniques that exploit the temporal information suffer from higher time requirements due to two state computations. The present disclosure solves the technical problems in the art with the methods and systems for real time video driven human 3-D posture estimation during physical activities. The present invention discloses a smart-phone camera based automatic posture monitoring system designed with an auto-encoder based architecture. The disclosed auto-encoder based cross-modal method uses monocular video (2-D image sequences) from a single low-end mobile device (for example, smart-phone camera) for estimating human 3-D posture in real time (˜5 fps) with high accuracy (less than 1 cm error per joint location).
Owner:TATA CONSULTANCY SERVICES LTD

Three-dimensional Gaussian splash reconstruction method for underwater scene

The invention discloses a three-dimensional Gaussian splash reconstruction method for an underwater scene, and belongs to the technical field of computer vision and three-dimensional reconstruction. Comprising the following steps: acquiring a monocular video frame sequence of a target underwater scene, a corresponding camera pose sequence, an initial sparse point cloud, an initial three-dimensional Gaussian point set and learnable physical parameters of an underwater imaging model; in the training process, performing weighted evaluation on a reconstruction error based on a multi-view consistency mechanism of opacity weighting, calculating an importance score of each Gaussian point, and performing densification operation on a three-dimensional Gaussian point set; adopting a staged freezing strategy to cooperatively optimize the three-dimensional Gaussian point set and underwater imaging model parameters; and performing rendering and underwater image synthesis on any new view angle camera pose based on the optimized three-dimensional Gaussian point set and underwater imaging model parameters, and outputting a new view angle synthesized image to represent a reconstruction result. According to the method, the geometric compactness, the visual fidelity and the physical interpretability of an underwater three-dimensional reconstruction result are improved.
Owner:ZHEJIANG UNIV

Aviation scene positioning and mapping method based on implicit neural rendering and optical flow assistance

The invention relates to the technical field of aviation target measurement, in particular to an aviation scene positioning and mapping method based on implicit neural rendering and optical flow assistance, and the method comprises the steps: extracting a key frame based on a monocular video stream, obtaining a dense depth value through the key frame in combination with TSDF model rendering, and estimating the posture of a camera; performing unsupervised dense scene measurement according to the adjacent key frames and the camera attitude, and performing training by adopting a loss function to obtain a depth map of the current key frame; obtaining a TSDF voxel grid through a TSDF model, fusing the depth map into the TSDF voxel grid, reconstructing a scene three-dimensional model, and forming a positioning and mapping model; an image data set is collected and preprocessed, a training data set is constructed, and a positioning and mapping model is trained; and carrying out aviation scene positioning and mapping based on the trained positioning and mapping model. Through positioning and dense mapping, the precision of positioning and reconstruction is improved, and then the perception capability of the aviation cockpit is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Monocular video-based abnormal gait assessment method and system, and computing device

The invention discloses an abnormal gait assessment method and system based on a monocular video, and a computing device. The method comprises the following steps: acquiring a human skeleton point sequence of the monocular video; performing data correction on the obtained human skeleton point sequence; and inputting the corrected human skeleton point sequence into a pre-trained space-time diagram convolutional network model, and obtaining an output result of the space-time diagram convolutional network model, the output result including an evaluation result of the gait anomaly condition of the subject in the monocular video. The method can effectively solve the problems of left and right leg confusion, easy shielding of key points and the like of skeleton point extraction in the monocular video, can enhance capture and understanding of gait space-time dynamic features, realizes accurate extraction of human gait space-time features, improves the accuracy and robustness of monocular video gait analysis, and improves the accuracy and robustness of the monocular video gait analysis. Effective gait analysis can be carried out by using a common monocular camera (such as a smart phone), the cost is lower, and the application scene is wider.
Owner:NANKAI UNIV +1

Three-dimensional model sequence generation method and related equipment

The embodiment of the invention discloses a three-dimensional model sequence generation method and related equipment. The related equipment can comprise a three-dimensional model sequence generation device, electronic equipment, a computer program product and a computer readable storage medium. According to the embodiment of the invention, feature extraction is carried out on video frames in a monocular video to obtain image features, an initial noise sequence corresponding to a three-dimensional model of a target object is generated, denoising is carried out on the initial noise sequence according to the image features to obtain a feature sequence set, and based on the frame positions of the video frames and the time distance between the video frames, the target object is obtained. Screening at least one reference feature block associated with the feature block from the feature sequence set, denoising the feature block according to the reference feature block to obtain a target feature sequence of the video frame, and generating a three-dimensional model sequence of the target object based on the target feature sequence; according to the scheme, the reference feature blocks can be screened to perform block-level cross-frame information interaction, so that the generation quality of the three-dimensional model sequence can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Sheep abnormal behavior detection method based on video analysis

The invention relates to the technical field of computer vision, in particular to a sheep abnormal behavior detection method based on video analysis, which comprises the following steps of: processing a multi-view video or monocular video data stream, performing posture recognition on each sheep in continuous video frames, establishing a skeleton movement time sequence record of each sheep in continuous time, and recording the skeleton movement time sequence record of each sheep; and generating a three-dimensional bone key point sequence of the sheep. According to the method, posture recognition is carried out on the skeleton structure of each sheep in the continuous video frames, the key point positions of the four limbs are positioned, and the motion trails in the continuous time period are established in cooperation with three-dimensional skeleton reconstruction, so that animal behaviors obtain the expression basis in the space dimension and the time dimension at the same time. And in combination with the key structure vertical displacement, step length and ground contact duration of the left and right limbs in the support phase, a parameter set with symmetry evaluation ability is constructed, and the abnormal gait state caused by small motion difference can be quantified.
Owner:SHANDONG CHAOYANG ANIMAL HUSBANDRY CO LTD

Vision-based three-dimensional human pose estimation system and method for ergonomic risk assessment

ActiveUS12511929B1Image enhancementImage analysisErgonomic riskVision based
Disclosed herein are vision-based three-dimensional (3D) pose estimation system and method for ergonomic risk assessment. An example system may comprise a computing device configured to obtain a monocular video capturing motions of a subject performing at least one working activity for a selected duration of time, perform a whole-body two dimensional (2D) pose estimation based at least on extracted frames of the monocular video, perform a whole-body 3D pose estimation based at least on the whole-body 2D pose estimation, calculate joint angles based at least on the whole-body 3D pose estimation, determine a posture score for each identified joint in each frame of the monocular video, and determine an ergonomic risk level of each identified joint based at least upon the posture score.
Owner:VELOCITYEHS HOLDINGS INC

Video processing method and device, equipment and storage medium

The embodiment of the invention discloses a video processing method and device, equipment and a storage medium, and the method comprises the steps: obtaining a three-dimensional video frame in a three-dimensional video stream received by a first terminal, and enabling the first terminal to generate the three-dimensional video frame based on a depth estimation result of a corresponding two-dimensional video frame in a two-dimensional video stream and a foreground matting result; separating foreground content from background content of the three-dimensional video frame, and updating a foreground texture patch and a background texture patch; and forming a current video frame based on the foreground texture patch and the determined background display area, and displaying the current video frame, the background display area being determined from the background texture patch based on the current pose information of the execution device. By using the method, special acquisition equipment is not needed, the effect of changing the playing visual angle of the video content along with the adjustment of the pose of the execution equipment can be realized in the visual effect only by processing the traditional monocular video frame, and the problems of cost, storage and transmission caused by the existing realization of the playing of the three-dimensional video / six-degree-of-freedom video are avoided.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Port machinery equipment moving distance detection method and system based on vision without auxiliary mark

The invention provides a vision-based auxiliary-mark-free port machinery equipment movement distance detection method and system, and relates to the field of computers.The method comprises the steps that monocular cameras are installed on lifting appliances on the two sides of a gantry crane, the gantry crane is controlled to move along a preset track, and a calibration image sequence containing ground linear features is obtained; a radial distortion parameter of the camera is calculated, a nonlinear mapping model from a pixel coordinate system to a world coordinate system is established, a ground unshielded area is selected as a dynamic monitoring area, and a texture richness thermodynamic diagram is generated; after receiving a measurement instruction of a scheduling system, initializing an optical flow accumulator, loading a current feature point set, synchronously collecting monocular video streams, and distributing the monocular video streams to at least three preprocessing threads to execute differential image enhancement; optical flow calculation is executed on all preprocessing results in parallel, a multi-mode optical flow vector field is generated, and three-level optical flow screening is implemented. The precision and the sensitivity of small-range movement measurement of the gantry crane are improved through an optical flow method, and the influence of environmental factors on a measurement result is reduced.
Owner:FUJIAN ELECTRONIC PORT CO LTD

Virtual digital human generation method and electronic equipment

The invention relates to the technical field of computer vision, and particularly provides a virtual digital human generation method and electronic equipment, and the method can comprise the steps: obtaining three-dimensional human body modeling data corresponding to each frame of image in a monocular video sequence of a target object; the three-dimensional human body modeling data comprises body posture data and identity offset data; processing the three-dimensional human body modeling data by using a pre-constructed to-be-trained virtual model to generate a virtual object image of each frame of image; the to-be-trained virtual model comprises a to-be-optimized three-plane feature body and a plurality of multi-layer neural network MLP models; optimizing the to-be-trained virtual model through the loss value between the virtual object image and each corresponding frame image, and obtaining a three-dimensional digital human model corresponding to the target object; the three-dimensional digital human model is used to generate a virtual digital human based on any drive data. According to the embodiment of the invention, the generation and rendering quality of the virtual digital human can be improved.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Multi-layer clothing reconstruction and simulation method in monocular video

The invention discloses a multi-layer clothing reconstruction and simulation method in a monocular video, and belongs to the technical field of artificial intelligence, and the method comprises the steps: S1, collecting the monocular video, carrying out the image segmentation of clothing in the video, and extracting the corresponding image label information; s2, initializing a standard digital human model through a neural network; s3, training a neural network so as to fit a digital human in the video, wherein each monocular video needs to train a neural network independently; s4, extracting a complete garment from each monocular video; s5, aligning and correcting the plurality of clothes extracted from the monocular video based on physical simulation; the invention provides a multilayer clothing reconstruction and simulation method in a monocular video. The method comprises the following steps: extracting a three-dimensional grid by using a neural network; according to the method, image segmentation labels are reversely projected to the vertexes of the grids to form a complete garment model, the interlayer penetration problem possibly occurring in the reconstruction process is solved through physical constraints, and the reality sense and usability of a final result are ensured.
Owner:FEIJIE COSI INTELLIGENT TECHNOLOGY (SHANGHAI) CO LTD

Monocular video dynamic human body reconstruction method and system based on three-dimensional gaussian splashing

This invention relates to the fields of computer vision and computer graphics, and provides a method and system for dynamic human body reconstruction from monocular video based on 3D Gaussian splashing. The method includes the following steps: data preprocessing; initialization of the normalized space 3D Gaussian; deformation of the normalized space 3D Gaussian to an intermediate pose space to obtain a non-rigidly deformable 3D Gaussian and pose-related features; transformation of the non-rigidly deformable 3D Gaussian to the observation space using linear blending skinning to obtain the observation space 3D Gaussian; decoding the viewpoint-related color based on Gaussian color features, pose-related features, and viewpoint direction; constructing a total loss function including a normal consistency regularization term to optimize the 3D Gaussian attributes and network parameters; and rendering the target human body image using a differentiable Gaussian splash rasterizer. This invention enables rapid and fully automatic reconstruction from monocular video to a high-fidelity, animable human body model, applicable to fields such as virtual reality and film production.
Owner:CHANGCHUN UNIV

Three-dimensional reconstruction method, system and device based on monocular video and medium

The invention discloses a three-dimensional reconstruction method, system and device based on a monocular video and a medium, and the method comprises the steps: carrying out the data collection processing of an indoor space through a monocular camera, obtaining a monocular video, carrying out the sliding window segmentation processing of the monocular video according to the space volume of the indoor space, and obtaining video segments, local reconstruction processing is carried out on a video clip to obtain a local reconstruction point cloud, key frame joint registration processing is carried out on the local reconstruction point cloud to obtain a registration scene frame, and global scene optimization processing is carried out on the registration scene frame according to spatial constraints to obtain a three-dimensional reconstruction result. According to the embodiment of the invention, the accuracy of three-dimensional reconstruction can be improved, and the method can be widely applied to the technical field of computer vision.
Owner:GUANGZHOU ZHONGYIYONG INTELLIGENT TECH CO LTD

Three-dimensional human body reconstruction method and system based on three-dimensional gaussian splashing

The application belongs to the field of three-dimensional vision and digitization, and relates to a three-dimensional human body reconstruction method and system based on three-dimensional Gaussian splashing. The method steps are as follows: based on a layered hash coding parameter field, the center position of each Gaussian primitive in the constructed three-dimensional Gaussian primitive set is corrected, and the color of the Gaussian primitive under the current observation angle is predicted; based on the human body posture parameters and shape parameters corresponding to the monocular video sequence, linear mixed skin transformation is performed on the corrected Gaussian primitive to map to the posture space; based on the Gaussian primitive parameters mapped to the posture space, three-dimensional Gaussian differentiable rendering is performed on the Gaussian primitive mapped to the posture space to obtain a rendering image consistent with the corresponding view angle of the monocular video; based on the constructed joint loss function, the Gaussian primitive parameters and the layered hash coding parameter field are optimized to obtain a three-dimensional human body model. The application can quickly reconstruct an animatable three-dimensional human body model with stable contours and clear textures from a monocular video.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Robot dynamic interaction skill learning method and system based on monocular video

The invention provides a robot dynamic interaction skill learning method and system based on a monocular video, and relates to the technical field of robot control and computer vision crossing. According to the method, the end-to-end, high-fidelity and high-robustness imitation learning ability of the physical robot for executing the dynamic interaction task from monocular vision observation is realized. According to the method, the objective function is optimized, so that the reconstruction result is ensured not only to be consistent with the video visually, but also to be reasonable and achievable physically, and the reconstruction result which is highly consistent physically provides high-quality input for action redirection. In the action redirection process, space-time constraints for interactive contact events are introduced, and it is ensured that the generated robot reference joint trajectory meets physical executable conditions at key interactive moments. Finally, a gap from simulation to reality is reduced by learning a training control strategy, and end-to-end simulation learning from monocular vision observation to physical robot execution is realized.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Monocular four-dimensional video generation method and device based on geometric constraint and electronic equipment

The embodiment of the invention discloses a monocular four-dimensional video generation method and device based on geometric constraints and electronic equipment. A specific embodiment of the method comprises the following steps: extracting video geometric dynamic information from a monocular video frame sequence to obtain video geometric dynamic information; performing alignment correction on the initial depth map sequence according to the static region three-dimensional point cloud to generate an aligned and corrected depth map sequence; inputting the aligned and corrected depth map sequence into a camera projection model to obtain a video frame three-dimensional point cloud sequence; generating a dynamic three-dimensional point cloud sequence and a static three-dimensional point cloud sequence; generating a dynamic initial Gaussian point cloud and a static background Gaussian point cloud; performing geometric constraint updating on the initial point cloud transformation model to obtain an updated point cloud transformation model; and generating a monocular four-dimensional video according to the static background Gaussian point cloud, the dynamic initial Gaussian point cloud and the updated point cloud transformation model. According to the embodiment, the quality and rendering efficiency of the monocular four-dimensional video can be improved.
Owner:BEIHANG UNIV

Methods, systems, equipment, and media for detecting dangerous actions based on radio frequency images.

This invention belongs to the field of vision and image processing technology, and provides a method, system, device, and medium for detecting dangerous actions based on radio frequency (RF) images. The method includes: acquiring a monocular video stream of a target scene, reconstructing a three-dimensional digital space geometric model of the target scene, and solving the pose matrix; collecting raw multidimensional feature data emitted by a signal transmitting device, mapping the raw multidimensional feature data to the three-dimensional digital space geometric model, and generating a three-dimensional scalar field data volume characterizing human radio frequency behavior disturbances; calculating a dynamic disturbance residual field based on the three-dimensional scalar field data volume, projecting and encoding the dynamic disturbance residual field to generate a 2.5D multi-channel human radio frequency behavior feature image, and identifying dangerous actions. This invention utilizes monocular vision to reconstruct a three-dimensional digital space and solve the pose of RF devices, establishing a physical space-digital space mapping relationship, and eliminating the strong dependence of traditional RF solutions on specific room layouts and multipath effects.
Owner:XI AN JIAOTONG UNIV

Digital human reconstruction method with high-fidelity triangular mesh and material texture map

The application discloses a digital human reconstruction method with high-fidelity triangular mesh and material texture mapping, and belongs to the technical field of computer graphics and digital human reconstruction. S1: performing space point sampling on each frame of picture corresponding to monocular video based on ray tracing, and deforming the sampling points to distribution under a standard posture; S2: acquiring geometric information and color information of global space points; S3: performing integration on the sampling points on each light ray to obtain volume rendering results, and completing first-stage optimization; S4: selecting a target frame, initializing a three-dimensional mesh, and generating a human body geometric surface; S5: acquiring material texture properties of the corrected human body geometric surface through a material network; S6: realizing differentiable rendering on the corrected human body geometric surface; S7: introducing an information fusion strategy to generate dense body rendering results under a virtual perspective, and supervising second-stage optimization; and S8: finally generating a digital human with a high-quality triangular mesh surface and material texture properties.
Owner:ZHEJIANG UNIV

Human body surface dynamic reconstruction method based on monocular video

PendingCN121921446AAccurately restore garment wrinklesAccurate recovery of muscle movementsImage enhancementImage analysisHuman bodyMorphing
The invention discloses a human body surface dynamic reconstruction method based on a monocular video, and the method comprises the steps: extracting a key frame of human body motion from the monocular video, and obtaining an RGB image of the key frame, a mask, parameters of an SMPL human body template, and internal and external parameters of a camera; extracting a real normal vector diagram corresponding to the key frame, and constructing a basic data set; the basic data set is used for training a multi-stage progressive human body reconstruction framework, the framework comprises human body overall motion modeling serving as a first stage, local surface dynamic deformation modeling serving as a second stage and surface appearance and illumination modeling serving as a third stage, and an optimal reconstruction model is obtained through training; and extracting a new view angle frame of human body motion from the monocular video, obtaining an RGB image of the new view angle frame, parameters of the SMPL human body template and internal and external parameters of the camera, and inputting the RGB image, the parameters of the SMPL human body template and the internal and external parameters of the camera into the optimal reconstruction model to generate a human body dynamic reconstruction result under the conditions of a new view angle, a new posture and new illumination, thereby realizing high-quality rendering output with geometric details and appearance consistency.
Owner:SOUTH CHINA UNIV OF TECH

Autodecoding latent 3D diffusion models

Systems and methods for generating static and articulated 3D assets are provided that include a 3D autodecoder at their core. The 3D autodecoder framework embeds properties learned from the target dataset in the latent space, which can then be decoded into a volumetric representation for rendering view-consistent appearance and geometry. The appropriate intermediate volumetric latent space is then identified and robust normalization and de-normalization operations are implemented to learn a 3D diffusion from 2D images or monocular videos of rigid or articulated objects. The methods are flexible enough to use either existing camera supervision or no camera information at all—instead efficiently learning the camera information during training. The generated results are shown to outperform state-of-the-art alternatives on various benchmark datasets and metrics, including multi-view image datasets of synthetic objects, real in-the-wild videos of moving people, and a large-scale, real video dataset of static objects.
Owner:SNAP INC

Multimodal feature fusion digital human video generation method and device based on time sequence position coding

The application discloses a kind of multi-modal feature fusion digital person video generation method and device based on timing position coding, comprising: constructing key point based on monocular video;Using the face Faceverse coefficient detected by Faceverse model, the face Flame coefficient of SMPLX model is fitted and replaced, the Mano hand shape detected using Hamer model, the hand representation of SMPLX model is fitted and replaced, to obtain the optimized SMPLX model;Color coding representation image and eye gaze image are obtained based on key point drawing, while based on the depth image, semantic image and normal image obtained by drawing the optimized SMPLX model;In image generation model, timing position coding for enhancing timing consistency is introduced, while based on the multi-modal feature formed by all images, a plurality of digital person images are continuously generated, and audio is added to obtain digital person video, which has wide application prospect in many fields.
Owner:ZHEJIANG UNIV

4D Gaussian structuring-based lunar surface / deep space scene target reconstruction and editing method

The invention relates to the technical field of three-dimensional graphic processing, and discloses a lunar surface / deep space scene target reconstruction and editing method based on 4D Gaussian structuring. The method is particularly suitable for dynamic target modeling and editing in deep space exploration tasks, including detectors, robots and related equipment on the lunar surface, the Mars or other celestial bodies. The method comprises the following steps: synthesizing a plurality of camera visual angles through a diffusion model based on a monocular video of an object, and reconstructing static three-dimensional Gaussian point representation in a standard space; extracting a skeleton structure of the object based on the surface grid, wherein the skeleton structure comprises a skeleton node position and a topological relation; on the basis of the skeleton structure, each Gaussian point is connected with a plurality of skeleton nodes through a linear hybrid skin mechanism, and skeleton-driven rigid deformation is achieved; non-rigid deformation is compensated through feature extraction and a regression network; and in combination with the rigid deformation and the non-rigid deformation, rendering to generate a dynamic Gaussian point model which can change along with time and can be edited. According to the method, motion modeling of an object is explicitly split into rigid motion driven by a framework and non-rigid correction, so that the definition and interpretability of motion representation are greatly improved, and particularly, higher stability and control precision are shown when challenges such as weak texture, violent illumination change and high noise are faced in a deep space environment. According to the method, object behavior modeling in a detection task is more flexible, and the method can be widely applied to scenes such as task planning, target recognition and dynamic editing.
Owner:UNIV OF SCI & TECH OF CHINA