Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

68 results about "Monocular video" patented technology

Three-dimensional human body and scene interaction reconstruction method and system

The invention belongs to the technical field of three-dimensional reconstruction, and particularly discloses a three-dimensional human body and scene interaction reconstruction method and system. The method is a three-dimensional human body and scene interactive reconstruction method for monocular video input, not only can quickly complete three-dimensional Gaussian reconstruction of a human body, a static scene and a moving object, but also can quickly complete the three-dimensional Gaussian reconstruction of the human body, the static scene and the moving object by introducing technical means such as posture correction, camera external parameter space regularization, joint reconstruction consistency loss and layered fusion rendering. And accurate alignment of the human body, the scene and the object in spatial positions, shielding relations and illumination styles is ensured, and dynamic interaction reconstruction conforming to real physical logic is realized. Finally, a joint rendering image with high fidelity, continuous time sequence and consistent structure can be generated, a stable and reliable three-dimensional expression capability is provided for complex human body actions, object interaction and scene understanding, and the three-dimensional interaction reconstruction quality and application value under the monocular video condition are greatly improved.
Owner:NANJING UNIV OF SCI & TECH

Three-dimensional Gaussian splash reconstruction method for underwater scene

The invention discloses a three-dimensional Gaussian splash reconstruction method for an underwater scene, and belongs to the technical field of computer vision and three-dimensional reconstruction. Comprising the following steps: acquiring a monocular video frame sequence of a target underwater scene, a corresponding camera pose sequence, an initial sparse point cloud, an initial three-dimensional Gaussian point set and learnable physical parameters of an underwater imaging model; in the training process, performing weighted evaluation on a reconstruction error based on a multi-view consistency mechanism of opacity weighting, calculating an importance score of each Gaussian point, and performing densification operation on a three-dimensional Gaussian point set; adopting a staged freezing strategy to cooperatively optimize the three-dimensional Gaussian point set and underwater imaging model parameters; and performing rendering and underwater image synthesis on any new view angle camera pose based on the optimized three-dimensional Gaussian point set and underwater imaging model parameters, and outputting a new view angle synthesized image to represent a reconstruction result. According to the method, the geometric compactness, the visual fidelity and the physical interpretability of an underwater three-dimensional reconstruction result are improved.
Owner:ZHEJIANG UNIV

Aviation scene positioning and mapping method based on implicit neural rendering and optical flow assistance

The invention relates to the technical field of aviation target measurement, in particular to an aviation scene positioning and mapping method based on implicit neural rendering and optical flow assistance, and the method comprises the steps: extracting a key frame based on a monocular video stream, obtaining a dense depth value through the key frame in combination with TSDF model rendering, and estimating the posture of a camera; performing unsupervised dense scene measurement according to the adjacent key frames and the camera attitude, and performing training by adopting a loss function to obtain a depth map of the current key frame; obtaining a TSDF voxel grid through a TSDF model, fusing the depth map into the TSDF voxel grid, reconstructing a scene three-dimensional model, and forming a positioning and mapping model; an image data set is collected and preprocessed, a training data set is constructed, and a positioning and mapping model is trained; and carrying out aviation scene positioning and mapping based on the trained positioning and mapping model. Through positioning and dense mapping, the precision of positioning and reconstruction is improved, and then the perception capability of the aviation cockpit is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Three-dimensional model sequence generation method and related equipment

The embodiment of the invention discloses a three-dimensional model sequence generation method and related equipment. The related equipment can comprise a three-dimensional model sequence generation device, electronic equipment, a computer program product and a computer readable storage medium. According to the embodiment of the invention, feature extraction is carried out on video frames in a monocular video to obtain image features, an initial noise sequence corresponding to a three-dimensional model of a target object is generated, denoising is carried out on the initial noise sequence according to the image features to obtain a feature sequence set, and based on the frame positions of the video frames and the time distance between the video frames, the target object is obtained. Screening at least one reference feature block associated with the feature block from the feature sequence set, denoising the feature block according to the reference feature block to obtain a target feature sequence of the video frame, and generating a three-dimensional model sequence of the target object based on the target feature sequence; according to the scheme, the reference feature blocks can be screened to perform block-level cross-frame information interaction, so that the generation quality of the three-dimensional model sequence can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Vision-based three-dimensional human pose estimation system and method for ergonomic risk assessment

ActiveUS12511929B1Image enhancementImage analysisErgonomic riskVision based
Disclosed herein are vision-based three-dimensional (3D) pose estimation system and method for ergonomic risk assessment. An example system may comprise a computing device configured to obtain a monocular video capturing motions of a subject performing at least one working activity for a selected duration of time, perform a whole-body two dimensional (2D) pose estimation based at least on extracted frames of the monocular video, perform a whole-body 3D pose estimation based at least on the whole-body 2D pose estimation, calculate joint angles based at least on the whole-body 3D pose estimation, determine a posture score for each identified joint in each frame of the monocular video, and determine an ergonomic risk level of each identified joint based at least upon the posture score.
Owner:VELOCITYEHS HOLDINGS INC

Port machinery equipment moving distance detection method and system based on vision without auxiliary mark

The invention provides a vision-based auxiliary-mark-free port machinery equipment movement distance detection method and system, and relates to the field of computers.The method comprises the steps that monocular cameras are installed on lifting appliances on the two sides of a gantry crane, the gantry crane is controlled to move along a preset track, and a calibration image sequence containing ground linear features is obtained; a radial distortion parameter of the camera is calculated, a nonlinear mapping model from a pixel coordinate system to a world coordinate system is established, a ground unshielded area is selected as a dynamic monitoring area, and a texture richness thermodynamic diagram is generated; after receiving a measurement instruction of a scheduling system, initializing an optical flow accumulator, loading a current feature point set, synchronously collecting monocular video streams, and distributing the monocular video streams to at least three preprocessing threads to execute differential image enhancement; optical flow calculation is executed on all preprocessing results in parallel, a multi-mode optical flow vector field is generated, and three-level optical flow screening is implemented. The precision and the sensitivity of small-range movement measurement of the gantry crane are improved through an optical flow method, and the influence of environmental factors on a measurement result is reduced.
Owner:FUJIAN ELECTRONIC PORT CO LTD

Virtual digital human generation method and electronic equipment

The invention relates to the technical field of computer vision, and particularly provides a virtual digital human generation method and electronic equipment, and the method can comprise the steps: obtaining three-dimensional human body modeling data corresponding to each frame of image in a monocular video sequence of a target object; the three-dimensional human body modeling data comprises body posture data and identity offset data; processing the three-dimensional human body modeling data by using a pre-constructed to-be-trained virtual model to generate a virtual object image of each frame of image; the to-be-trained virtual model comprises a to-be-optimized three-plane feature body and a plurality of multi-layer neural network MLP models; optimizing the to-be-trained virtual model through the loss value between the virtual object image and each corresponding frame image, and obtaining a three-dimensional digital human model corresponding to the target object; the three-dimensional digital human model is used to generate a virtual digital human based on any drive data. According to the embodiment of the invention, the generation and rendering quality of the virtual digital human can be improved.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

Monocular video dynamic human body reconstruction method and system based on three-dimensional gaussian splashing

This invention relates to the fields of computer vision and computer graphics, and provides a method and system for dynamic human body reconstruction from monocular video based on 3D Gaussian splashing. The method includes the following steps: data preprocessing; initialization of the normalized space 3D Gaussian; deformation of the normalized space 3D Gaussian to an intermediate pose space to obtain a non-rigidly deformable 3D Gaussian and pose-related features; transformation of the non-rigidly deformable 3D Gaussian to the observation space using linear blending skinning to obtain the observation space 3D Gaussian; decoding the viewpoint-related color based on Gaussian color features, pose-related features, and viewpoint direction; constructing a total loss function including a normal consistency regularization term to optimize the 3D Gaussian attributes and network parameters; and rendering the target human body image using a differentiable Gaussian splash rasterizer. This invention enables rapid and fully automatic reconstruction from monocular video to a high-fidelity, animable human body model, applicable to fields such as virtual reality and film production.
Owner:CHANGCHUN UNIV

Three-dimensional human body reconstruction method and system based on three-dimensional gaussian splashing

The application belongs to the field of three-dimensional vision and digitization, and relates to a three-dimensional human body reconstruction method and system based on three-dimensional Gaussian splashing. The method steps are as follows: based on a layered hash coding parameter field, the center position of each Gaussian primitive in the constructed three-dimensional Gaussian primitive set is corrected, and the color of the Gaussian primitive under the current observation angle is predicted; based on the human body posture parameters and shape parameters corresponding to the monocular video sequence, linear mixed skin transformation is performed on the corrected Gaussian primitive to map to the posture space; based on the Gaussian primitive parameters mapped to the posture space, three-dimensional Gaussian differentiable rendering is performed on the Gaussian primitive mapped to the posture space to obtain a rendering image consistent with the corresponding view angle of the monocular video; based on the constructed joint loss function, the Gaussian primitive parameters and the layered hash coding parameter field are optimized to obtain a three-dimensional human body model. The application can quickly reconstruct an animatable three-dimensional human body model with stable contours and clear textures from a monocular video.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Robot dynamic interaction skill learning method and system based on monocular video

The invention provides a robot dynamic interaction skill learning method and system based on a monocular video, and relates to the technical field of robot control and computer vision crossing. According to the method, the end-to-end, high-fidelity and high-robustness imitation learning ability of the physical robot for executing the dynamic interaction task from monocular vision observation is realized. According to the method, the objective function is optimized, so that the reconstruction result is ensured not only to be consistent with the video visually, but also to be reasonable and achievable physically, and the reconstruction result which is highly consistent physically provides high-quality input for action redirection. In the action redirection process, space-time constraints for interactive contact events are introduced, and it is ensured that the generated robot reference joint trajectory meets physical executable conditions at key interactive moments. Finally, a gap from simulation to reality is reduced by learning a training control strategy, and end-to-end simulation learning from monocular vision observation to physical robot execution is realized.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Methods, systems, equipment, and media for detecting dangerous actions based on radio frequency images.

This invention belongs to the field of vision and image processing technology, and provides a method, system, device, and medium for detecting dangerous actions based on radio frequency (RF) images. The method includes: acquiring a monocular video stream of a target scene, reconstructing a three-dimensional digital space geometric model of the target scene, and solving the pose matrix; collecting raw multidimensional feature data emitted by a signal transmitting device, mapping the raw multidimensional feature data to the three-dimensional digital space geometric model, and generating a three-dimensional scalar field data volume characterizing human radio frequency behavior disturbances; calculating a dynamic disturbance residual field based on the three-dimensional scalar field data volume, projecting and encoding the dynamic disturbance residual field to generate a 2.5D multi-channel human radio frequency behavior feature image, and identifying dangerous actions. This invention utilizes monocular vision to reconstruct a three-dimensional digital space and solve the pose of RF devices, establishing a physical space-digital space mapping relationship, and eliminating the strong dependence of traditional RF solutions on specific room layouts and multipath effects.
Owner:XI AN JIAOTONG UNIV

Digital human reconstruction method with high-fidelity triangular mesh and material texture map

The application discloses a digital human reconstruction method with high-fidelity triangular mesh and material texture mapping, and belongs to the technical field of computer graphics and digital human reconstruction. S1: performing space point sampling on each frame of picture corresponding to monocular video based on ray tracing, and deforming the sampling points to distribution under a standard posture; S2: acquiring geometric information and color information of global space points; S3: performing integration on the sampling points on each light ray to obtain volume rendering results, and completing first-stage optimization; S4: selecting a target frame, initializing a three-dimensional mesh, and generating a human body geometric surface; S5: acquiring material texture properties of the corrected human body geometric surface through a material network; S6: realizing differentiable rendering on the corrected human body geometric surface; S7: introducing an information fusion strategy to generate dense body rendering results under a virtual perspective, and supervising second-stage optimization; and S8: finally generating a digital human with a high-quality triangular mesh surface and material texture properties.
Owner:ZHEJIANG UNIV

Human body surface dynamic reconstruction method based on monocular video

PendingCN121921446AAccurately restore garment wrinklesAccurate recovery of muscle movementsImage enhancementImage analysisHuman bodyMorphing
The invention discloses a human body surface dynamic reconstruction method based on a monocular video, and the method comprises the steps: extracting a key frame of human body motion from the monocular video, and obtaining an RGB image of the key frame, a mask, parameters of an SMPL human body template, and internal and external parameters of a camera; extracting a real normal vector diagram corresponding to the key frame, and constructing a basic data set; the basic data set is used for training a multi-stage progressive human body reconstruction framework, the framework comprises human body overall motion modeling serving as a first stage, local surface dynamic deformation modeling serving as a second stage and surface appearance and illumination modeling serving as a third stage, and an optimal reconstruction model is obtained through training; and extracting a new view angle frame of human body motion from the monocular video, obtaining an RGB image of the new view angle frame, parameters of the SMPL human body template and internal and external parameters of the camera, and inputting the RGB image, the parameters of the SMPL human body template and the internal and external parameters of the camera into the optimal reconstruction model to generate a human body dynamic reconstruction result under the conditions of a new view angle, a new posture and new illumination, thereby realizing high-quality rendering output with geometric details and appearance consistency.
Owner:SOUTH CHINA UNIV OF TECH

Digital human modeling method and system combining Gaussian sputtering image and GAN model

The invention discloses a digital human modeling method and system in combination with a Gaussian sputtering image and a GAN model, and belongs to the technical field of digital human generation, and the method comprises the following steps: condition generation: extracting a human body key point sequence from a monocular video of a single person, using a human body key point sequence to drive a constructed Gaussian model to generate a Gaussian sputtering image sequence with a figure identity as a figure identity condition, and drawing based on the human body key point sequence to generate a hand image sequence, a neural semantic image sequence and an eye fixation image sequence as action conditions; and digital human modeling: a generator in the GAN model carries out digital human video frame generation by taking the human identity condition and the action condition as driving conditions of the generator in the GAN model at the same time, and a digital human video is formed. In the method and the system, character identity condition control is effectively introduced into the Gaussian sputtering image, the generalization ability of the generated model is improved, and the method and the system have wide application prospects in multiple fields.
Owner:ZHEJIANG UNIV

Large-scale monocular vision slam-gs method and system based on depth prior and subgraph management

This application discloses a large-scale monocular visual SLAM-GS method based on depth prior and subgraph management, comprising: acquiring keyframes of a monocular video stream; acquiring an original depth prior map for the keyframes; aligning the original depth prior map with the current local map to obtain an aligned depth prior map; dynamically dividing the global map into multiple subgraphs and maintaining a bounded set of active subgraphs; when creating a new subgraph, initializing the map elements of the new subgraph using keyframes in the overlapping area with the old subgraph and their aligned depth prior maps; constructing and optimizing a 3D Gaussian map based on the aligned depth prior map within the set of active subgraphs; maintaining a global pose map with subgraphs as nodes, and correcting the global cumulative error by optimizing the relative poses between subgraphs. This invention achieves integrated fusion of monocular visual SLAM and 3D Gaussian modeling in large-scale, long-term scenes.
Owner:XIAN FANGJU XINGCHEN TECHNOLOGY CO LTD

Scene reconstruction from monocular video

A technique for reconstructing a three-dimensional scene from monocular video adaptively allocates an explicit sparse-dense voxel grid with dense voxel blocks around surfaces in the scene and sparse voxel blocks further from the surfaces. In contrast to conventional systems, the two-level voxel grid can be efficiently queried and sampled. In an embodiment, the scene surface geometry is represented as a signed distance field (SDF). Representation of the scene surface geometry can be extended to multi-modal data such as semantic labels and color. Because properties stored in the sparse-dense voxel grid structure are differentiable, the scene surface geometry can be optimized via differentiable volume rendering.
Owner:NVIDIA CORP

A method and system for three-dimensional human pose estimation

The application discloses a three-dimensional human posture estimation method and system, relates to the technical field of computer vision, and comprises the following steps: extracting two-dimensional human posture key points from a monocular video picture sequence, and generating a two-dimensional human posture key point sequence; projecting the two-dimensional key point sequence to a feature space through nonlinear high-dimensional mapping, and generating a high-dimensional feature space matrix; inputting the high-dimensional feature matrix into a three-dimensional human posture recognition model which fuses motion constraints and frequency division space-time features, obtaining a three-dimensional human posture key point sequence, and realizing three-dimensional human posture estimation through three-dimensional coordinates. The method improves the robustness and detection precision of the monocular three-dimensional human posture estimation method. Error values are calculated for the relative motion speed, bone length and bone direction of the key points, so that the training is easier to converge and the training process is more stable.
Owner:TONGJI UNIV

Pet three-dimensional reconstruction and personalized interaction behavior generation system and method based on monocular video driving

The invention discloses a pet three-dimensional reconstruction and personalized interaction behavior generation system and method based on monocular video driving. According to the system, skeleton and acoustic features in a video are extracted through a multi-modal feature decoupling module, and high-fidelity appearance reconstruction and physical collision feedback are realized by using an explicit-implicit hybrid rendering technology based on discrete volume primitives (such as 3D Gaussian). Meanwhile, a style-constrained imitation learning network (GAIL) and a user feedback fine tuning mechanism (RLHF) are introduced, and the personalized exercise expression of the pet is copied from the video. According to the method, the problems of visual distortion, lack of physical feedback of interaction and rigid behavior mode of the virtual pet in the prior art are effectively solved, immersive digital twin experience of visible, touch and personality is realized, and the method has perfect end-side performance adaptation and privacy protection capabilities.
Owner:胡欣 +2

A method for travel path control based on monocular vision recognition and an unmanned laying vehicle for crack-resistant base fabric.

ActiveCN120876808BPrecisely control the direction of travelImprove laying accuracyCharacter and pattern recognitionPattern recognitionMachine vision
This invention provides a method for controlling the travel path based on monocular vision recognition and an unmanned laying vehicle for crack-resistant base fabric, belonging to the technical field of machine vision recognition. The method includes using a monocular video acquisition device to collect video information of a target reference object on the crack-resistant base fabric laying equipment and converting it into image data to determine the initial center coordinate position of the target reference object in the image; adjusting the acquisition angle of the monocular image acquisition device on the target reference object according to the initial center coordinate position of the target reference object in the image, so that the distance between the initial center coordinate position of the target reference object in the image and a preset target center coordinate position in the image meets a preset initial distance; acquiring the travel center coordinate position of the target reference object in the image in real time and calculating the deviation between the travel center coordinate position and the target center coordinate position; and controlling the travel path based on the calculated deviation. This invention can significantly improve the flatness and efficiency of crack-resistant base fabric laying and effectively reduce labor costs.
Owner:SHANXI EXPRESSWAY DEV CO LTD

Method for generating virtual digital human and electronic device

The application relates to the technical field of computer vision, and particularly provides a virtual digital person generation method and an electronic device. The method can comprise the following steps: acquiring three-dimensional human modeling data corresponding to each frame of image in a monocular video sequence of a target object; the three-dimensional human modeling data comprises body posture data and identity offset data; processing the three-dimensional human modeling data by using a pre-constructed virtual model to be trained, to generate a virtual object image of each frame of image; the virtual model to be trained comprises a three-plane feature body to be optimized and a plurality of multi-layer neural network (MLP) models; optimizing the virtual model to be trained by using a loss value between the virtual object image and each frame of image corresponding to the virtual object image, to acquire a three-dimensional digital person model corresponding to the target object; and the three-dimensional digital person model is used for generating a virtual digital person based on arbitrary driving data. The embodiment of the application can improve the generation and rendering quality of the virtual digital person.
Owner:HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD

3D multi-person human pose estimation method and system in video based on spatio-temporal attention

This invention discloses a method and system for 3D multi-person human pose estimation in videos based on spatiotemporal attention, relating to the field of pose estimation technology. This invention improves the human pose estimation scheme based on spatiotemporal feature fusion, effectively reconstructing the pose information of multiple people from monocular video sequences. The model can not only accurately estimate the 3D spatial position of each individual's joints, but also clearly reconstruct the relative positions and spatial hierarchy between people, effectively avoiding pose aliasing and computational complexity problems in multi-person scenes, and possessing excellent multi-person 3D human pose estimation capabilities. Without requiring multi-view or depth sensors, it achieves 3D multi-person human pose estimation results highly consistent with real-world scenes, verifying the effectiveness and practical value of the model in achieving high-precision 3D multi-person human pose estimation under monocular image conditions.
Owner:NORTHEASTERN UNIV CHINA

Method for disentangled reconstruction of dynamic digital human, and electronic device and storage medium

PCT designated stageWO2026143330A1Human bodyThree dimensional shape
The present invention can be applied to the technical field of computer vision and graphics. Provided are a method for disentangled reconstruction of a dynamic digital human, and an electronic device and a storage medium. The method comprises: using a technique for reconstructing three-dimensional shapes of a human body and garments from a monocular human body video, and using a representation method in which explicit geometry is combined with an implicit signed distance field (hmSDF), such that high-quality reconstruction of a dynamic disentangled digital human from a monocular video can be achieved. In the method, garments and a human body are separated and separately modeled, optimized hmSDF is used to achieve accurate segmentation of visible regions, and an SMPL model is also used to complete occluded human body regions, thereby ensuring the consistency and fidelity of the overall geometry. A linear blend skinning (LBS) deformation field and a non-rigid deformation field are used to capture human body motion and detailed variations, such that high-fidelity and continuous spatio-temporally disentangled human body and garment geometry is ultimately generated.
Owner:UNIV OF SCI & TECH OF CHINA

Automatic decoding of potential 3d diffusion models

Systems and methods are provided for generating static and articulated 3D assets, the core of which includes a 3D automatic decoder. The 3D automatic decoder framework embeds attributes learned from the target dataset into potential space, which may then be decoded into volumetric representations to render view-consistent appearances and geometries. Appropriate intermediate volume potential spaces are then identified and robust normalization and anti-normalization operations are performed to learn 3D diffusion from 2D images or monocular videos of rigid or articulated objects. These methods are sufficiently flexible, may use existing camera supervision or do not use camera information at all, but rather effectively learn camera information during training. The results generated are proved to be superior to the most advanced alternatives in various reference datasets and indices, including multi-view image datasets of synthetic objects, real field videos of moving persons, and real video datasets of large-scale static objects.
Owner:SNAP INC

Dynamic three-dimensional scene text driven editing method based on monocular video

The embodiment of the invention discloses a dynamic three-dimensional scene text-driven editing method based on a monocular video. A specific embodiment of the method comprises the following steps: acquiring a semantic feature tensor information set and user edited text information; generating each piece of target quantitative index map information; updating the dynamic three-dimensional Gaussian field; determining a Gaussian distribution information group; updating the monocular video information to obtain reference monocular video information and rendering monocular video information; and updating the Gaussian distribution information group to obtain text editing video information. According to the embodiment, the condition that the same object in the monocular video deviates in the editing areas in all the included video frames can be reduced, so that the time consumed for repairing the video frames in the monocular video is shortened, the consistency among the edited video frames is improved, the editing effect is improved, and the user experience is improved. And the time consumption of rendering and consumed computing resources are shortened.
Owner:BEIHANG UNIV

Convenient three-dimensional human body reconstruction method based on monocular video

The invention discloses a convenient three-dimensional human body reconstruction method based on a monocular video, and the method comprises the steps: constructing a parameterized human body model low-dimensional shape space, and obtaining the approximate representation of a linear model. And extracting a multi-view key frame from the input human body rotation video. And extracting a feature vector of each view key frame by using an improved ResNet-50 network fused with an attention mechanism. And defining the multi-view feature vectors as graph nodes, and constructing a graph structure. And performing cross-view depth feature fusion through the graph convolutional network to obtain global features. And mapping the global features to a parameterized human body model low-dimensional shape space. And reconstructing a complete three-dimensional human body grid by using linear model approximate representation. The method has the advantages that multi-view feature extraction and a graph structure fusion mechanism are combined, the end-to-end process from consumption-level video input to high-precision three-dimensional human body reconstruction is achieved, certain generalization ability is achieved in a real scene, and a low-cost and high-precision three-dimensional human body modeling solution can be provided for virtual fitting, digital twinning and other applications.
Owner:HIGH FASHION CHINA CO LTD

Digital human modeling method and system in combination with Gaussian splash image and visual autoregression model

The invention discloses a digital human modeling method and system combining Gaussian splash images and a visual autoregression model, and the method comprises the steps: generating a three-dimensional Gaussian splash image sequence based on a monocular video, and generating a condition image group sequence which comprises a hand model image sequence, a neural semantic key point image sequence and an eye fixation image sequence; based on a visual autoregression model containing a BSQ multi-scale image word segmentation device and a visual autoregression Transformer, introducing an action condition image guide network and a time sequence attention layer to form a digital human generation model; the digital human generation model is trained and then is used for generating a target digital human image to realize digital human modeling, so that multi-person identity, high-fidelity and time sequence continuity digital human image sequence generation can be realized, the image sequence generation speed is remarkably improved compared with a diffusion model, and the method has wide application prospects in multiple fields.
Owner:ZHEJIANG UNIV

Monocular crowd detection method and system based on bounding box height normalization

The application belongs to the technical field of computer vision, and provides a monocular person gathering detection method and system based on detection frame height normalization, which comprises the following steps: performing person target detection on the continuous frame images of a monocular video stream, and outputting the detection frame coordinates and detection confidence of each person; based on the detection frame coordinates, the corresponding foot point positions of each person in the image plane are calculated; according to the detection frame coordinates of each person, the detection frame height of each person is determined, and based on the detection frame height, the normalized distance between the foot point positions of any two persons is calculated, and an undirected connected graph is constructed based on the normalized distance; based on the detection confidence, the undirected connected graph is solved and judged by using the connected component, and the person gathering detection result is obtained by combining the time sequence and confidence fusion strategy. The scheme is light in calculation, high in real-time performance, low in deployment cost, convenient in engineering landing, and can realize stable and efficient person gathering detection on edge devices.
Owner:BEIJING EASY TIMES DIGITAL TECH

Digital human tooth optimization method based on three-dimensional deformable model and image

The invention discloses a digital human tooth optimization method based on a three-dimensional deformable model and an image, and the method comprises the steps: extracting tooth information from a multi-view image and a monocular video collected by a mobile phone, building a correct shape of a digital human tooth model, and further optimizing the correct position of the digital human tooth model under different expressions; wherein the extraction of the tooth information is based on model transformation logic and camera projection logic, and then tooth feature point information of a person under different expressions is obtained from various images; the tooth model shape is established on the basis of an optimization method, and the tooth model shape is optimized and adjusted on the basis of the extracted tooth feature point information; the position of the tooth model is established on the basis of the correct shape, and the position of the tooth model is optimized and adjusted on the basis of the extracted tooth feature point information. Compared with the prior art, the correct shape and position of the digital human tooth model can be quickly established, and the establishment speed of the digital human tooth model is greatly improved.
Owner:ZHEJIANG UNIV

A Method for Reconstructing and Predicting 3D Trajectory of Tennis Balls Based on Monocular Video

This invention relates to a method for reconstructing and predicting the 3D trajectory of a tennis ball based on monocular video. The method includes: matching key points on a 2D court with corresponding points on a 3D standard court to calculate the camera's intrinsic and extrinsic parameters, establishing a 3D world coordinate system based on the 3D standard court; using the known 3D coordinates of racket key points, combined with their 2D coordinates, calculating the racket's 3D depth and pose within the monocular video frame to determine the position of the hitting point in 3D space; and setting hitting point and landing point constraints, adjusting the input initial velocity vector and spin vector, and using physical motion equations to construct a simulated 3D trajectory passing through the hitting and landing points, then projecting the simulated 3D trajectory onto a 2D plane and fitting it with a 2D pixel trajectory to construct the optimal 3D trajectory. This method can reconstruct the 3D trajectory and output the ball's true velocity, net clearance height, and landing depth, meeting the needs of professional quantitative analysis.
Owner:ZHEJIANG UNIV

Method, device, electronic device and storage medium for equine 4d reconstruction

This application provides a method, apparatus, electronic device, and storage medium for 4D reconstruction of equines, relating to the field of image processing technology. The method includes: acquiring a monocular video sequence containing equine activity; performing motion estimation on each frame of the monocular video sequence to obtain motion parameters for each frame; wherein the motion parameters include model parameters and camera parameters of a parametric model designed based on equines, the model parameters including pose parameters, shape parameters, and global displacement parameters; selecting a frame from the monocular video sequence as a reference image, performing appearance reconstruction on the reference image to obtain 3D Gaussian attributes of 3D Gaussian units; and rendering based on the motion parameters and the 3D Gaussian attributes to obtain a target 4D video. This application can efficiently and faithfully reconstruct the 4D geometry and appearance of equines from monocular videos, and exhibits high robustness to incomplete viewpoint coverage.
Owner:SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY