Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

194 results about "Monocular image" patented technology

Three-dimensional scene reconstruction method and device based on large model geometric prior, and medium

The invention discloses a three-dimensional scene reconstruction method and device based on large model geometric prior, and a medium, and aims to solve the problems that a conventional 3DGS is liable to have artifacts and detail loss in geometric discontinuity, data redundancy and illumination variation scenes, and predicts a dense depth map and a normal map from a monocular image by using a pre-trained large model. The position and form of the Gaussian kernel are constrained as additional geometric priori; a primitive adjustment strategy based on kernel density estimation is introduced in the training stage, small Gaussian primitives with similar structures and adjacent spaces are combined into a large Gaussian primitive, the rendering quality is kept, redundancy is reduced, and the volume of the model is reduced; an exposure coefficient is adaptively estimated for each input image, an exposure compensation image loss function is constructed, and floating artifacts caused by illumination differences at shooting moments are eliminated. Experiments show that compared with the prior art, the method improves the three-dimensional reconstruction precision and real-time rendering quality of complex illumination and less-texture areas in a public data set and an unmanned aerial vehicle aerial photography scene.
Owner:NARI INFORMATION & COMM TECH

Single-image-based three-dimensional face reconstruction and editing method

The present invention provides a single-image-based three-dimensional face reconstruction and editing method, comprising: extracting features by means of a pre-trained EG3D network, so as to generate a preliminary three-dimensional face; using an inversion module to map a monocular image to a latent code space, and optimizing a latent code to generate a realistic three-dimensional face image; making use of a depth estimation technology and multi-view projection to create a pseudo multi-view image; then, by means of semantic segmentation and expression feature extraction, fusing the features to generate a comprehensive representation; according to editing requirements, optimizing and adjusting the latent code, so as to implement customized editing; and finally, a generator outputting an edited face image. The present invention can enable more efficient, flexible and high-quality three-dimensional face reconstruction and editing, thereby effectively overcoming the defects of high cost, low universality and lack of customized editing in existing face reconstruction technologies.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Method and electronic device for 3D object detection using neural networks

A method of 3D object detection using an object detection neural network includes: receiving one or more monocular images; extracting 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through a 2D feature extracting part of the object detection neural network, generating an averaged 3D voxel volume based on the 2D feature maps, extracting a 2D representation of 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through an encoder of a 3D feature extracting part of the object detection neural network, and performing 3D object detection as 2D object detection in a Bird's Eye View (BEV) plane, the 2D object detection in the BEV plane being performed by passing the 2D representation of 3D feature maps through thane outdoor object detecting part of the object detection neural network.
Owner:SAMSUNG ELECTRONICS CO LTD

Shielding area three-dimensional voxel reasoning method and system

The invention belongs to the technical field of three-dimensional mapping, and particularly relates to an occlusion area three-dimensional voxel reasoning method and system. The method comprises the following steps: S10, acquiring an RGB image containing a semantic segmentation result of an occlusion region and a corresponding depth map, and converting the depth map into an initial sparse three-dimensional voxel; and S20, constructing a three-dimensional voxel inference network, and inferring a three-dimensional voxel occupancy probability graph from the initial sparse three-dimensional voxels through the three-dimensional voxel inference network. The occlusion area can be reconstructed based on the monocular image; multi-format output and navigation interface butt joint are supported, and the method can be integrated to SLAM, 3D mapping or path search systems to serve as spatial feasibility constraints; the method can be used for querying whether any spatial position is passable; the method can be used for extracting a trafficability sub-graph from a designated area for path planning.
Owner:SHANGHAI UNIV

Three-dimensional perception robot operation knowledge distillation method based on monocular image

The invention relates to the field of robot operation and three-dimensional perception, in particular to a monocular image-based three-dimensional perception robot operation knowledge distillation method, which comprises the following steps of: establishing a strategy learning framework comprising a student model and a teacher model; constructing a strategy prediction model in the strategy learning framework; training the strategy prediction model, wherein a three-level knowledge distillation mechanism is adopted in the training process to complete knowledge migration between the teacher model and the student model; combining the three distillation losses with strategy optimization losses to form a total loss function, and performing end-to-end training on the student model until convergence to obtain a deployable monocular strategy model; the deployable monocular strategy model only retains a student model, and generates a robot operation instruction. The method has the beneficial effects that the robot under monocular RGB input has three-dimensional perception and high-precision operation capabilities while the reasoning efficiency is kept, the task success rate and generalization performance are remarkably improved, and the effectiveness and robustness of the method are verified.
Owner:ZHEJIANG UNIV OF TECH

High-precision map reconstruction method and system based on monocular vision, medium and equipment

The invention belongs to the technical field of robot positioning and three-dimensional mapping, and discloses a high-precision map reconstruction method and system based on monocular vision, a medium and equipment. Scene image data are acquired through a monocular image acquisition module, after feature extraction and matching are performed on each frame of image, attitude information acquired by an inertial measurement unit is fused, and pose calculation of a robot is completed through a sparse vision SLAM system; meanwhile, an image dense depth map is generated by a monocular depth estimation model based on an attention mechanism, and the image dense depth map is converted into a single-frame color dense point cloud in combination with image RGB color information. According to the invention, a loose coupling fusion strategy is adopted to carry out spatial registration on a single-frame colored dense point cloud and a robot pose at a corresponding moment, multi-frame data fusion is completed through point cloud splicing and optimization, and a globally consistent three-dimensional dense point cloud map is constructed to realize scene modeling. The method has the characteristics of high robustness and high reconstruction precision, and can be effectively applied to three-dimensional map construction in an outdoor complex environment.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Semantic scene completion method and device, electronic equipment and readable storage medium

The invention discloses a semantic scene completion method and device, electronic equipment and a readable storage medium. The method comprises the following steps: acquiring a binocular image; performing depth estimation processing on the binocular image to obtain a multi-scale image feature map, a disparity probability distribution map and a disparity map of the monocular image; performing conversion processing from a parallax dimension to a depth dimension on the parallax image of the monocular image to obtain a depth image of the monocular image; processing the multi-scale feature map and the depth map of the monocular image to obtain a first semantic feature map of the monocular image; performing depth feature extraction processing on the parallax probability distribution diagram of the monocular image to obtain a depth probability distribution diagram of the monocular image; and determining voxel occupation data and voxel semantic data of the semantic occupation data based on the depth map of the monocular image, the first semantic feature map and the depth probability distribution map so as to perform semantic scene completion. According to the invention, semantic scene completion can be carried out based on visual image data.
Owner:SHENZHEN SWEET POTATO ROBOT CO LTD

Methods, systems, and computer program products for generating 3D human pose and movement estimation from monocular image information

A computer-implemented method includes converting by a pose tokenizer, based on a learned codebook, pose parameters of a body into a sequence of discrete pose tokens; randomly masking a portion of the sequence of discrete pose tokens; predicting the randomly masked sequence of discrete pose tokens based on multi-scale features extracted from a monocular image by an image conditioned masked transformer; optimizing the sequence of discrete pose tokens by aligning a re-projected three-dimensional (3D) pose with an estimated two-dimensional (2D) pose; directly regressing, from the multi-scale features, a shape parameter of the body and a weak perspective camera parameter; and generating a 3D mesh reconstruction of the body based on the shape parameter and the weak perspective camera parameter.
Owner:THE UNIV OF NORTH CAROLINA AT CHAPEL HILL

Physical driving measurement method for monocular three-dimensional dynamic displacement of rotary machinery

The invention provides a physical driving measurement method for monocular three-dimensional dynamic displacement of a rotary machine, and belongs to the technical field of crossing of computer vision and industrial state monitoring. According to the method, a high-speed dynamic visual acquisition system is constructed, and a time sequence video stream is obtained by using a cooperative marker; constructing a'Gaussian + motion blur 'composite gradient model taking the motion blur width as an endogenous variable, and jointly resolving sub-pixel edge coordinates and ambiguity by adopting a nonlinear optimization algorithm; decoupling the pixel displacement of the radial X axis, the radial Y axis and the axial Z axis by using the width change and the edge displacement of the marker in combination with the geometric principles of sequential robust filtering and monocular imaging; and finally, outputting a three-way physical vibration waveform by using the calibration conversion factor. According to the method, the fuzzy influence is adaptively eliminated from a physical imaging mechanism, micron-level precision three-direction vibration synchronous measurement is realized only by a single camera, and the hardware cost and the deployment difficulty are reduced.
Owner:OCEAN UNIV OF CHINA

Oil taking port positioning method and system based on monocular vision and laser positioning

The invention provides an oil taking port positioning method and system based on monocular vision and laser positioning. The method comprises the steps that laser point cloud data and a monocular image near an oil taking port are acquired; constructing a three-dimensional model of the oil taking port based on the laser point cloud data, and projecting the three-dimensional model of the oil taking port into a two-dimensional slice image according to an acquisition position label of the monocular camera; matching the two-dimensional slice image with the monocular image to determine the two-dimensional feature position of the oil extraction port; calculating the feature distance between the oil taking port and the camera based on the matched two-dimensional feature position of the oil taking port and the focal length label and the size label of the monocular image; on the basis of the position parameters of the oil taking port in the laser point cloud data, the feature distance is combined, and initial three-dimensional coordinates of the oil taking port are generated; and unifying the pixel coordinates of the monocular vision and the laser point cloud coordinates into the same coordinate system, and carrying out preliminary three-dimensional coordinate data fusion to obtain the final three-dimensional coordinates of the oil extraction port. The positioning precision of the transformer oil taking robot on the oil taking port is improved.
Owner:HUBEI INFOTECH SYST TECH CO LTD

Underwater target identification and positioning method based on physical model and deep learning fusion

The invention discloses an underwater target identification and positioning method based on fusion of a physical model and deep learning, and belongs to the technical field of computer vision, and the method comprises the steps: carrying out the transmissivity estimation and physical restoration of a collected underwater monocular image according to an underwater light propagation model, and carrying out the image correction; key feature points are extracted, and high-quality matching point pairs are screened in combination with the joint similarity; further deriving a basic matrix through the high-quality matching point pairs meeting the epipolar geometric constraint relation, solving an essential matrix, and obtaining relative attitude parameters between the cameras in combination with weighted re-projection-LM optimization; and finally, refraction correction triangulation is carried out according to the Snell's law, pixel-level fusion is carried out after scale normalization and space alignment are carried out on the refraction correction triangulation and dense depth output by the MiDaS, three-dimensional space coordinates of the target are inverted, and high-precision recognition and positioning of the underwater target are achieved. The system is light in structure, efficient in calculation, suitable for being integrated on various underwater autonomous or remote control robot platforms and used for tasks such as target recognition, tracking and positioning.
Owner:CENT SOUTH UNIV

Miniature intelligent terminal intestinal image processing method and device, storage medium and computer equipment

The invention discloses a miniature intelligent terminal intestinal image processing method and device, a storage medium and computer equipment. The method comprises the following steps: preprocessing an original intestinal tract image set acquired by a monocular endoscope, calculating according to the preprocessed intestinal tract image set to obtain a motion calibration scale factor and a texture calibration scale factor, and adaptively fusing to obtain a conversion coefficient from a pixel to an actual size; extracting polyp region key feature points of each image of the preprocessed intestinal image set, constructing a Delaunay triangulation graph according to each polyp region key feature point, and performing non-rigid deformation matching based on a cross-graph convolutional network and an optimal transmission algorithm to obtain a matched feature point set; and sparse 3D reconstruction is carried out according to the matched feature point set to obtain a polyp point cloud, and the polyp size is calculated according to the Euclidean distance between two farthest points in the polyp point cloud and the conversion coefficient. According to the method, the problems of monocular image scale ambiguity and non-rigid deformation can be solved, and the more accurate polyp size can be obtained.
Owner:BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV +1

Large semi-trailer reversing auxiliary method, device and equipment

The invention provides a large semi-trailer reversing auxiliary method, device and equipment, and the method comprises the steps: obtaining the geometric center information of a ground marking line of a garage and the contour data of the ground marking line of the garage in real time through monocular image collection equipment; laser ranging data are obtained in real time through a laser ranging sensor; determining a first relative position relation of the vehicle relative to the garage according to the garage ground marking line contour data and the garage ground marking line geometric center information; determining a second relative position relation between the vehicle and the garage according to the laser ranging data; performing data fusion processing on the first relative position relation and the second relative position relation to obtain a target relative position between the vehicle and the garage; and performing auxiliary reversing according to the target relative position.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Monocular 3D target detection method and system based on comparative learning

The invention relates to a monocular 3D target detection method and system based on comparative learning. The method comprises the following steps: S1, performing feature extraction on a monocular image by using a backbone network to obtain a multi-scale feature map; the weight of the backbone network is obtained through a self-supervised training stage; s2, performing visual encoding by using a visual encoder to obtain an encoded visual feature; a depth predictor and a depth encoder are used in sequence to carry out depth coding on the sum to obtain coding depth features; s3, inputting the learnable object query, the coding visual features and the coding depth features into a visual-depth decoder to obtain a final query after learning; s4, mapping the final query into 3D information by adopting a detection head; and obtaining a 3D bounding box of the corresponding target. According to the method, the precision and robustness of 3D target detection are improved.
Owner:SOUTH CHINA NORMAL UNIV

Optical depth estimation using segmentation and geometric priors for interior space monitoring systems and applications

PendingUS20260094287A1Image enhancementImage analysisPattern recognitionOptical depth
In various examples, optical depth estimation for interior space monitoring systems and applications is disclosed. Absolute 3D depth estimates from monocular image data may be generated using a machine learning model using 3D geometry priors and a joint learning framework that combines depth estimation with object segmentation. Three-dimensional geometry priors provide surface-level information that enriches the model's understanding of the relevant spatial geometry and resolves scale ambiguity in monocular depth estimation within automotive in-cabin environments. The model may include a common (e.g., shared) encoder stage that outputs features extracted from an optical image sensor feed to separate decoder stages that include a depth estimation decoder and a segmentation decoder. Joint learning for depth estimation and segmentation tasks during training achieves a more nuanced understanding of the in-cabin environment, leading to significantly improved depth accuracy.
Owner:NVIDIA CORP

Monocular online reconstruction method based on three-dimensional Gaussian splashing and geometric prior

The invention discloses a monocular online reconstruction method based on three-dimensional Gaussian splashing and geometric prior. According to the method, under the condition that only a calibration-free monocular image is input, a camera and scene priori are predicted through a pre-trained visual geometric model, a direct primitive sampling strategy based on a Gaussian difference operator is combined, a three-dimensional Gaussian primitive is generated from the image and the primitive so as to expand a scene map, and the scene map is expanded. A coarse-to-fine rendering-driven joint optimization method is adopted to optimize the camera pose and the scene Gaussian map at the same time, and global consistency is ensured through online loop detection and pose map optimization. Compared with other scene reconstruction methods based on radiation field rendering and feedforward reconstruction technologies, the method has the capabilities of high-fidelity rendering and robust pose estimation in indoor and outdoor scenes.
Owner:ZHEJIANG UNIV

Depth guidance image defogging method and system for deep ground medical rescue

The invention discloses a depth guidance image defogging method and system for deep ground medical rescue, and aims to improve the visibility of medical rescue monitoring images and the accuracy and real-time performance of structural information recovery in deep ground environments such as mines. By introducing a scale collaborative attention module, collaborative learning of multi-scale features is realized, scale confusion is relieved, and the cross-scale feature collaboration capability is enhanced; and meanwhile, an error perception guide module is adopted to fuse local feature difference and global dependence, and mutual promotion optimization is performed on defogging and depth estimation tasks, so that more accurate image restoration is realized in a region with a complicated structure or non-uniform fog distribution. According to the method, close collaborative optimization and performance improvement between image defogging and depth estimation tasks can be realized, and the problems of scale confusion and limited scale collaborative capability in multi-scale feature learning can be solved to a certain extent.
Owner:YUNLONG LAKE LAB OF DEEP UNDERGROUND SCI & ENG +1

Road elevation estimation method based on monocular image

The invention discloses a pavement elevation estimation method based on a monocular image, and belongs to the technical field of automatic driving environment perception. The method comprises the following steps: preprocessing an input monocular image and extracting image features; the method comprises the following steps of: adaptively mapping image features to a bird's-eye view (BEV) space through a visual angle conversion module based on a deformable attention mechanism; and predicting the elevation value of the road surface by utilizing BEV characteristics. In the training stage, multi-frame laser radar point clouds are introduced to carry out plane fitting, and a global reference plane is generated to serve as a supervision signal, so that the robustness and convergence of the model under gradient change and complex road conditions are improved; in the reasoning stage, only monocular images are needed, a laser radar is not needed, and estimation precision and deployment efficiency are both considered. The method can effectively sense subtle fluctuation of the road surface, is suitable for complex scenes such as ramps and potholes, and provides high-quality geometric information support for automatic driving control.
Owner:UNIV OF SCI & TECH OF CHINA

Visual positioning method and system fusing semantic segmentation dynamic region elimination and Mama modeling

The invention discloses a visual positioning method and system fusing semantic segmentation dynamic region elimination and Mama modeling. The method comprises the following steps: firstly, obtaining a monocular image and preprocessing; secondly, inputting the preprocessed monocular image into a semantic segmentation model, outputting a semantic probability graph, dividing a static region and a dynamic region, and quantifying a classification probability; performing shielding or weight reduction processing on the dynamic region features of the semantic probability graph, and keeping the static region features unchanged; performing feature extraction on the semantic probability graph subjected to shielding or weight reduction processing to generate a high-dimensional spatial feature sequence; inputting the feature sequence into a time sequence modeling module, and processing the feature sequence by using a bidirectional Mama modeling module; and finally, carrying out camera pose estimation on the processed feature sequence by adopting six-degree-of-freedom representation to realize visual odometer estimation. According to the method, end-to-end training is kept, the spatial semantic information and time sequence dependency relationship can be captured at the same time, and the precision and stability of pose estimation are improved.
Owner:HANGZHOU NORMAL UNIVERSITY

Pose estimation method and device, computer device and storage medium

The application provides a pose estimation method and device, a computer device and a storage medium. The method comprises: acquiring a monocular image, a target mask image corresponding to a target to be measured in the monocular image, and a three-dimensional model of the target to be measured; sampling a plurality of initial pose data of the target to be measured under a determined view direction; determining initial position data of the target to be measured under each initial pose data based on the target mask image; wherein the initial pose data and the initial position data constitute initial pose data, and the plurality of initial pose data constitute an initial pose estimation set; selecting target quantity candidate pose data from the initial pose estimation set, adjusting the target quantity candidate pose data, and generating a plurality of intermediate pose data; constructing a loss function, and iteratively optimizing each intermediate pose data by using the loss function to generate optimized pose data; and determining pose estimation data corresponding to the target to be measured from the plurality of optimized pose data.
Owner:TSINGHUA UNIVERSITY

Image adjustment 3D content creation architecture based on 2D diffusion

The invention discloses an image adjustment 3D content creation architecture based on 2D diffusion, and relates to the technical field of diffusion models, and the architecture comprises a diffusion model which is composed of a plurality of cross attention blocks in a U-Net structure and supports effective fusion of various modes of texts, images and camera parameters. In the present invention, it is devoted to create 3D content using a potential diffusion model. The 3D geometry is not generated directly through a potential diffusion model, but a two-stage formula is employed. First, a potential diffusion model of view conditioned reflex is organized, and a multi-view image is synthesized with a monocular image and the outside of a camera as inputs. Next, a neural radiation field is trained using the synthesized multi-view image, which is easily optimized as volume rendering is differentiable. After the neural radiation field training is completed, a 3D geometric model is generated through an advancing cube algorithm applied to a density field. The framework eliminates the requirement for pairing 3D training data and does not require a large amount of computing resources.
Owner:THE INST OF AUTOMATION HEILONGJIANG ACADEMY OF SCI

Monocular image-driven efficient three-dimensional face modeling and rendering method

The invention relates to the technical field of computer vision and computer graphics, in particular to a monocular image-driven efficient three-dimensional face modeling and rendering method, which comprises the following steps of: acquiring a plurality of monocular portrait images and annotation information acquired by a camera, and establishing a portrait data set; establishing a generative model, wherein the generative model generates a portrait image from the Gaussian noise; pre-training the generated model to obtain a pre-trained 3D decoder; establishing a three-dimensional Gaussian reconstruction model, wherein the three-dimensional Gaussian reconstruction model generates a portrait image from the input image; performing joint fine tuning on the three-dimensional Gaussian reconstruction model to obtain a trained three-dimensional Gaussian reconstruction model; inputting a to-be-reconstructed monocular image into the trained three-dimensional Gaussian reconstruction model, and performing reasoning and rendering to obtain a reconstructed portrait image; according to the method, the training data collection difficulty can be reduced, and the three-dimensional face reconstruction generalization is improved.
Owner:BEIHANG UNIV

Dizziness lamp control method and system based on human body depth estimation and target detection

The application discloses a dizziness lamp control method and system based on human body depth estimation and target detection, comprising: acquiring a first data set and a second data set; training a pre-constructed depth estimation model according to the first data set, and then training a pre-constructed pedestrian detection model according to the second data set to obtain a human body depth estimation network; collecting a current monocular image, inputting the current monocular image into the human body depth estimation network to obtain human body depth information; and adjusting the brightness of the dizziness lamp according to the human body depth information. The application combines target detection and depth estimation to obtain human body depth information, can automatically and dynamically adjust the brightness of the dizziness lamp according to different target pedestrian distances, improves the single use time length of the dizziness lamp under the premise of ensuring the effect of the dizziness lamp, and can be widely applied to the technical field of light control.
Owner:GUANGZHOU CHENGZHI INTELLIGENT MACHINE TECH CO LTD

Building three-dimensional Gaussian modeling and mobile terminal display method and system based on monocular image

The invention provides a building three-dimensional Gaussian modeling and mobile terminal display method and system based on a monocular image, and the method comprises the steps: employing a monocular camera and an inertial measurement unit (IMU) which are equipped in a common Android mobile phone, and combining a cloud neural modeling algorithm based on image attitude and dense reconstruction; rapid three-dimensional reconstruction of a building external facade or an indoor space is realized; meanwhile, the dense point cloud is converted into three-dimensional Gaussian representation, lightweight loading and efficient rendering of the model at the mobile terminal are achieved, a user is further supported to conduct measurement, annotation and structured data output at the mobile terminal, and the actual application scenarios of various constructional engineering are met.
Owner:CHINA CONSTR FIFTH ENG DIV CORP LTD

Seamless splicing method for front-view real-time aerial view of auxiliary driving vehicle

The invention discloses an auxiliary driving vehicle foresight real-time aerial view seamless splicing method, which comprises the following steps: acquiring a vehicle-mounted monocular image of vehicle foresight in real time, and converting the vehicle-mounted monocular image into a real-time BEV aerial view Bc; updating the pose conversion matrix RT based on the real-time driving data; calculating the real-time BEV aerial view Bc to obtain a real-time weight map Wc under a vehicle coordinate system; based on the pose conversion matrix RT, converting the real-time BEV aerial view Bc into a global real-time BEV aerial view GBc, and converting the real-time weight map into a global real-time weight map GWc; and in a global coordinate system, seamless splicing based on the global real-time weight map is carried out on the global real-time BEV aerial view GBc and a historical global BEV aerial view, and a global BEV aerial view GBt at the current moment t is obtained. The seamless splicing method for the foresight real-time aerial view of the auxiliary driving vehicle has the advantages that the problem that the splicing effect of the foresight aerial view in a tunnel, an underground parking lot and other areas without GPS signals is poor can be solved with low computing power cost, and the like.
Owner:SCI & TECH CO LTD HEFEI INTELLIGENT VEHICLE TECH CO LTD

Depth sensor anomaly detection method and system, storage medium and electronic equipment

The invention discloses a depth sensor anomaly detection method and system, a storage medium and electronic equipment. The method comprises the steps that an absolute depth value collected by a depth sensor and a relative depth value output by a monocular depth estimation model are acquired; global sorting is carried out on the absolute depth value and the relative depth value, a sorting difference matrix between the absolute depth value and the relative depth value is further calculated, and a sorting inconsistency score vector is obtained; obtaining a first-stage detection threshold value by adopting a fused adaptive threshold value strategy based on the sorting inconsistency score vector, and forming an initial anomaly set; for each point in the non-abnormal point set, constructing a comprehensive similarity scoring matrix after fusion; and accumulating the comprehensive similarity scoring matrix according to rows to obtain a second-stage scoring vector, and determining a final refined abnormal point set. According to the invention, by introducing the relative depth information generated based on the monocular image, the robustness and environmental adaptability of the depth sensing system can be significantly improved.
Owner:HUAZHONG UNIV OF SCI & TECH +1

Stereoscopic video model training method, device, equipment, medium and program product

PendingCN121125959ASteroscopic systems3D-image renderingStereoscopic videoRadiology
The invention provides a stereoscopic video model training method and device, equipment, a medium and a program product, and relates to the technical field of computers, and the stereoscopic video model training method comprises the steps: obtaining a first prediction map and a depth map of the first prediction map according to a first stereoscopic video generation model and a first monocular map; the first stereoscopic video generation model is generated by training a first stereogram; generating a first loss function of the first encoder and the first decoder according to the first stereogram, the first monocular map, the first prediction map and a depth map of the first prediction map; and performing supervised training on the first encoder and the second encoder by using the first loss function to generate a second stereoscopic video generation model. According to the invention, the training data volume is increased, and the second stereoscopic video generation model is generated by using the first loss function, so that better scene understanding and more robust feature representation can be realized, and the video quality of the 3D video can be improved.
Owner:CHINA MOBILE COMM LTD RES INST +1

Unsupervised monocular image-based online 3D scene reconstruction method and device

The present disclosure provides a monocular image-based unsupervised three-dimensional scene online reconstruction method and device. The method of the present disclosure comprises: obtaining a static background mask and a mask of each dynamic object instance of a current frame monocular image through semantic segmentation, obtaining a surround view synthesis image through a static multi-view generator, determining the motion parameters of each dynamic object instance through motion modeling, obtaining a multi-view depth map based on the surround view synthesis image through a depth estimation network obtained through self-supervised training, obtaining a local depth map of each dynamic object instance through a local depth estimation network obtained through self-supervised training, and obtaining a complete 3D scene representation through the multi-view depth map, the local depth map of each dynamic object instance, and the motion parameters thereof. The present disclosure can avoid true value dependence, effectively reduce hardware cost, and at the same time improve the reliability and robustness of monocular image three-dimensional scene reconstruction.
Owner:BEIJING TRUNK TECHNOLOGY CO LTD

Image processing method and electronic device

The application discloses an image processing method and an electronic device, and relates to the technical field of image processing, wherein the method comprises the following steps: extracting a first geometric edge line set of an initial depth image according to the initial depth image corresponding to a target monocular image; screening geometric edge lines with gradient scores greater than a target score threshold from a depth gradient image generated according to the initial depth image and an initial edge line set of the target monocular image to determine a second geometric edge line set; and fusing the first geometric edge line set and the second geometric edge line set to determine image geometric edge lines of the target monocular image and construct a final geometric edge line set. The method solves the problems of serious texture interference, insufficient utilization of depth information, difficulty in balancing the number and purity of edge lines, and insufficient real-time performance in the prior art, and achieves the technical effects of filtering texture noise while ensuring the number and robustness of edge lines and meeting the needs of multi-sensor fusion tasks.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Method and system for 3d human pose estimation based on context and anatomical interaction

The present application belongs to the field of computer vision, and provides a three-dimensional human pose estimation method and system based on context and anatomical interaction, comprising: acquiring a monocular image and extracting a two-dimensional coordinate sequence of human joint points in the image coordinate system in the image; based on anatomical structure constraints, the two-dimensional coordinate sequence and the original image are fused to generate initial high-dimensional features that fuse visual semantics and anatomical priors; the initial high-dimensional features are input into an interactive fusion module, the local anatomical structure and the global context semantics are cooperatively modeled, and the dynamic joint dependency relationship under the current pose is captured to obtain enhanced features after deep fusion; the enhanced features are subjected to nonlinear transformation and dimension mapping, and the three-dimensional spatial coordinates of the human joint points are output. This method achieves advanced estimation accuracy and generalization ability on public benchmarks, effectively solving the ambiguity and irrationality problems in monocular three-dimensional pose estimation.
Owner:SHAANXI NORMAL UNIV