Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

32 results about "Perspective camera" patented technology

Methods, systems, and computer program products for generating 3D human pose and movement estimation from monocular image information

A computer-implemented method includes converting by a pose tokenizer, based on a learned codebook, pose parameters of a body into a sequence of discrete pose tokens; randomly masking a portion of the sequence of discrete pose tokens; predicting the randomly masked sequence of discrete pose tokens based on multi-scale features extracted from a monocular image by an image conditioned masked transformer; optimizing the sequence of discrete pose tokens by aligning a re-projected three-dimensional (3D) pose with an estimated two-dimensional (2D) pose; directly regressing, from the multi-scale features, a shape parameter of the body and a weak perspective camera parameter; and generating a 3D mesh reconstruction of the body based on the shape parameter and the weak perspective camera parameter.
Owner:THE UNIV OF NORTH CAROLINA AT CHAPEL HILL

Dynamic alignment between perspective camera and eye viewpoint in video perspective (VST) augmented reality (XR)

A method includes determining that an inter-pupil distance (IPD) between display lenses of a video perspective (VST) augmented reality (XR) device has been adjusted relative to a default IPD. The method further includes obtaining an image captured using a perspective camera of the VST XR device. The perspective camera is configured to capture an image of a three-dimensional (3D) scene. The method further includes transforming the image according to a change in the IPD relative to the default IPD to match a viewpoint of a corresponding one of the display lenses to generate a transformed image. The method further includes correcting distortion in the transformed image based on one or more lens distortion coefficients corresponding to the change in the IPD to generate a corrected image. Further, the method includes initiating presentation of the corrected image on a display panel of the VST XR device.
Owner:SAMSUNG ELECTRONICS CO LTD

Detection and classification of traffic signs using camera-radar fusion

The disclosed systems and techniques facilitate efficient detection and classification of traffic signs in driving environments. The disclosed techniques include, obtaining, using a sensing system of a vehicle a first set of perspective camera images of an environment and a second set of radar images of the environment. The techniques further include generating, using a first neural network, one or more camera features characterizing the first set of images, generating, using a second neural network, one or more radar features characterizing the second set of images, and processing the one or more camera features and the one or more radar features to obtain an identification of one or more traffic signs in the environment.
Owner:WAYMO LLC

Video see-through camera auto focus method and system for XR headsets

The present application relates to a kind of video perspective camera automatic focusing method and system for XR head-mounted display, belong to image processing technical field, solve the existing video perspective camera focusing speed slow, and the depth information of monocular image and focusing parameter cannot be effectively mapped problem.Method includes: the current image and its corresponding depth map of the video perspective camera of XR head-mounted display are collected;According to the fixation point coordinates of user in current image, the local region with fixation point coordinates as center is extracted from depth map as fixation area depth map;Fixation area depth map and fixation point coordinates are input into trained focusing parameter prediction model, and the target focusing parameter of video perspective camera is obtained;According to target focusing parameter, video perspective camera is adjusted, and single focusing is completed.Accurate mapping from image to focusing parameter is realized, and focusing speed is improved.
Owner:HANGZHOU HUIJIAN ZHILIAN TECH CO LTD

Video fusion method and system based on shadow map, and program product

The invention relates to the technical field of video fusion, and discloses a video fusion method and system based on a shadow map and a program product, and the method comprises the following steps: creating a virtual perspective camera in a three-dimensional scene, and generating the shadow map according to a visual cone of the virtual perspective camera; pixels in the three-dimensional scene are converted from a visual space coordinate system of a physical world camera to a virtual perspective camera coordinate system, NDC coordinates of the pixels are obtained, the NDC coordinates are aligned with texture coordinates of the shadow map, and sampling texture coordinates are obtained; judging whether the pixels in the visual range of the virtual perspective camera are in the visual area of the shadow map or not, and obtaining the color of the current pixel of the three-dimensional scene; and obtaining the color of the current pixel of the fused three-dimensional scene. According to the invention, the problems of shielding area processing errors and the like in the prior art are solved.
Owner:BEIJING ZHIHUI YUNZHOU TECH CO LTD

Dynamic scene online three-dimensional reconstruction system and method for humanoid robot teleoperation

The invention discloses a dynamic scene online three-dimensional reconstruction system and method for humanoid robot teleoperation. The method comprises the following steps: receiving an RGB-D video data stream collected by a first visual angle camera of a humanoid robot; performing dynamic and static separation processing on the RGB-D video data stream to obtain a mask graph comprising a static background and a dynamic object; constructing a static Gaussian map dictionary and a dynamic Gaussian map dictionary based on the mask graph; carrying out real-time asynchronous updating on the static Gaussian map dictionary and the dynamic Gaussian map dictionary by adopting a dual key frame selection strategy; and performing fusion rendering on the static Gaussian map dictionary and the dynamic Gaussian map dictionary to obtain a three-dimensional scene view for teleoperation, and completing online three-dimensional reconstruction of a dynamic scene. The method can be used for dynamic scenes where the humanoid robot performs teleoperation, such as training data acquisition in life scenes, electric power inspection, dangerous scene inspection and the like, and a real-time and immersive three-dimensional environment view is provided for a teleoperator.
Owner:HUAZHONG UNIV OF SCI & TECH +1

Systems and methods for C-arm fluoroscope camera pose refinement with secondary movement compensation

Imaging systems and methods compensate for wigwag movement of a C-arm fluoroscope to refine camera pose estimates. The methods involve computing a primary movement axis from samples of markers in fluoroscopic images of a fluoroscopic sweep of a structure of markers and processing the primary movement axis to obtain a secondary movement axis. The methods further involve aligning two-dimensional samples of each marker with the primary and secondary movement axes to obtain an aligned signal and determining a difference signal for a secondary component of the aligned signal. The difference signal is then converted to a rotation axis translation signal. The method further involves estimating a 3D position of the rotation axis. The estimated pose of the C-arm fluoroscope is then refined to compensate for the wigwag movement using the rotation axis translation signal and the estimated 3D position of the rotation axis.
Owner:COVIDIEN LP

Production line display method and device based on three-dimensional rendering technology and electronic equipment

The invention discloses a production line display method and device based on a three-dimensional rendering technology and electronic equipment, and relates to the technical field of intelligent manufacturing, and the method comprises the steps: creating L target models corresponding to a production line, the L target models being three-dimensional models corresponding to production equipment and equipment connection assemblies in the production line, the data volume of each target model is smaller than a preset data volume threshold; creating a scene container, a perspective camera and a renderer corresponding to the production line, and loading the L target models to the scene container through the renderer to obtain an initial display scene of the production line; based on the layout information of the production line, the equipment operation data and the light information, updating the initial display scene to obtain a target display scene of the production line; and dynamically displaying the target display scene of the production line through the perspective camera. The technical problems of low visualization degree and low production line management efficiency in simulation display of the industrial production line based on the prior art are solved.
Owner:CAXA TECH

A sparse-to-dense visual localization method and system based on feature gaussian splats

PendingCN122115572AReduce storage requirementsPreserve geometric richnessImage analysis3D modellingPattern recognitionHeat map
The application provides a sparse-to-dense visual positioning method and system based on feature Gaussian splash, and the method comprises the following steps: initializing a color-decoupled feature Gaussian field based on a training image set, optimizing the color-decoupled feature Gaussian field based on a query feature map set in combination with feature rendering and feature alignment loss cyclic optimization, and outputting a compact feature Gaussian scene model; screening a Gaussian landmark set in the compact feature Gaussian scene model by using a matching-oriented sampling strategy; training a scene-specific detector; extracting sparse local features of a query landmark heat map corresponding to a query image and performing sparse feature matching with the Gaussian landmark set to obtain an initial pose of a query perspective camera; based on 3D Gaussian splash, rendering a dense feature map and a depth map of the query perspective in the compact feature Gaussian scene model by using the initial pose of the query perspective camera, performing cluster-based proxy matching-based sparse-to-dense accelerated pose optimization, and obtaining accurate positioning of the query perspective camera.
Owner:WUHAN UNIV

Newborn feeding behavior image recognition and oral movement evaluation system

The invention relates to the technical field of medical image processing, in particular to a neonatal feeding behavior image recognition and oral movement evaluation system which comprises a data acquisition module, a feature extraction module, a differential geometric modeling module, a trajectory analysis module, a collaborative movement analysis module, an anomaly detection module, a result display module and an application service module. According to the system, a high-speed camera, a three-dimensional face tracking camera and an optical perspective camera work cooperatively, multi-view image data of newborn face and oral cavity movement are synchronously collected through a data collection analyzer, and a differential geometry theory is innovatively applied to oral cavity movement analysis. Facial feature point groups are regarded as Riemannian manifolds embedded in a three-dimensional Euclidean space, oral movement characteristics are accurately described through curvature feature calculation, shape operator construction and differential invariant extraction, a coupling model of lip manifolds and mandibular manifolds is systematically constructed, synchronism measurement and topological features between the manifolds are analyzed, and the oral movement characteristics are accurately described. And accurate detection of feeding abnormity is realized.
Owner:THE SEVENTH MEDICAL CENTER OF PLA GENERAL HOSPITAL

Heterogeneous binocular vision three-dimensional measurement method with telecentric lens and perspective lens

The invention provides a heterogeneous binocular stereoscopic vision three-dimensional measurement method with a telecentric lens and a perspective lens, which comprises the following steps of: firstly, constructing a heterogeneous binocular vision three-dimensional measurement system, then calibrating the heterogeneous binocular vision three-dimensional measurement system, and synchronously shooting calibration objects in different poses by using two cameras in the system, closed-form solutions of internal and external parameters of the telecentric camera and the perspective camera are respectively solved, lens distortion and consistency of relative poses among all image pairs are considered, a re-projection error optimization equation is established, and finally a calibration result is used for three-dimensional reconstruction. According to the stereoscopic vision system with the non-paired telecentric and perspective lenses, the potential of the telecentric lens in the vision system is released in 3D measurement application, and an accurate and reliable calibration result is provided.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Visual training animation production method and device, equipment and medium

The invention discloses a visual training animation production method and device, equipment and a medium, and relates to the technical field of visual training, and the visual training animation production method comprises the following steps: obtaining a plurality of display elements, and mapping the display elements to a virtual three-dimensional space provided with a perspective camera; the display elements are replaceable elements; the display elements are to-be-displayed elements stored in an associated medium of the display device and / or elements displayed in a reference area of the display device; determining a moving direction vector corresponding to the display element, and moving the display element in the virtual three-dimensional space according to a preset element moving sequence on the basis of a preset moving rule and the moving direction vector corresponding to the display element, the display elements are projected and rendered to the reference area by using the perspective camera every preset frame rendering time interval so as to complete animation production; the preset movement rule is that the display element moves between the farthest plane and the nearest plane. The convenience and universality of visual training can be improved.
Owner:SHENZHEN JIZHOU TECHNOLOGY CO LTD

Online rectification of see-through camera pair

A method includes obtaining, using multiple imaging sensors of an electronic device, a left image frame and a right image frame forming a stereo pair of image frames. The method also includes identifying, using at least one processing device of the electronic device, extrinsic parameters associated with relative positions and orientations of the imaging sensors. The method further includes performing, using the at least one processing device, an online stereo rectification of the stereo pair of image frames based on the extrinsic parameters such that epipolar lines of the left and right image frames are horizontally aligned to generate a rectified stereo pair of image frames. In addition, the method includes rendering, using the at least one processing device, one or more images for display based on the rectified stereo pair of image frames.
Owner:SAMSUNG ELECTRONICS CO LTD

Method and system for estimating temporally consistent 3D human shape and motion from monocular video

Estimating temporally consistent 3D human body shape, pose, and motion from a monocular video is a challenging task due to occlusions, poor lightning conditions, complex articulated body poses, depth ambiguity, and limited availability of annotated data. Embodiments of present disclosure provide a method for temporally consistent motion estimation from monocular video. A monocular video of person(s) is captured by a weak perspective camera and spatial features of body of the persons are extracted from each frame of the video. Then, initial estimates of body shape, body pose, and features of the weak perspective camera are obtained. The spatial features and initial estimates are then aggregated to obtain spatio-temporal features by a combination of self-similarity matrices between the spatial features, pose and the camera and self-attention maps of the camera features and the spatial features. The spatio-temporal aggregated features are then used to predict shape and pose parameters of the person(s).
Owner:TATA CONSULTANCY SERVICES LTD

Regression-based parametric model-free 3D human body posture and shape prediction method and device

The application discloses a regression-based parameterized non-model 3D human body posture and shape prediction method and device. The method comprises the following steps: processing the scaled human body image by using a non-model method, acquiring 3D vertices, image features and weak perspective camera parameters; projecting the 3D vertices to a 2D plane to acquire 2D vertices; inputting the 2D vertex coordinates and the image features into a body part perception sampling module to acquire part perception image features and initial T posture 3D vertex coordinates; connecting the part perception image features and the initial T posture 3D vertex coordinates by using a ConCat operation and inputting the connected part perception image features and the initial T posture 3D vertex coordinates into a body part decoding module to decode and regress to calculate absolute rotation and translation information of each part of the body; calculating relative rotation of each joint relative to a parent joint according to the absolute rotation information of each part of the body to obtain posture parameters; inputting the posture parameters and the 3D vertex coordinates into a shape regression module, converting the 3D vertex coordinates into standard T posture 3D vertex coordinates and inputting the standard T posture 3D vertex coordinates into a series of full connection layers to regress and calculate shape parameters.
Owner:ZHEJIANG UNIV

Method and system for rendering panoramic video

The present application discloses a method for rendering a panoramic video. The method includes obtaining a current frame of image from a video source, and generating texture map data, where the generating texture map data comprises determining a viewpoint region based on a field of view of a perspective camera, and rendering image pixels outside the viewpoint region at a lower resolution than rendering image pixels within the viewpoint region; mapping the current frame of image to a three-dimensional image based on a spherical rendering model the texture map data; and projecting the three-dimensional image onto a two-dimensional screen.
Owner:SHANGHAI BILIBILI TECH CO LTD

Dynamic alignment between see-through cameras and eye viewpoints in video see-through (VST) extended reality (XR)

PendingEP4609595A4Image enhancementSteroscopic systemsOphthalmologyPerspective camera
A method includes determining that an inter-pupillary distance (IPD) between display lenses of a video see-through (VST) extended reality (XR) device has been adjusted with respect to a default IPD. The method also includes obtaining an image captured using a see-through camera of the VST XR device. The see-through camera is configured to capture images of a three-dimensional (3D) scene. The method further includes transforming the image to match a viewpoint of a corresponding one of the display lenses according to a change in IPD with respect to the default IPD in order to generate a transformed image. The method also includes correcting distortions in the transformed image based on one or more lens distortion coefficients corresponding to the change in IPD in order to generate a corrected image. In addition, the method includes initiating presentation of the corrected image on a display panel of the VST XR device.
Owner:SAMSUNG ELECTRONICS CO LTD

Online rectification of see-through camera pair

A method includes obtaining, using multiple imaging sensors of an electronic device, a left image frame and a right image frame forming a stereo pair of image frames. The method also includes identifying, using at least one processing device of the electronic device, extrinsic parameters associated with relative positions and orientations of the imaging sensors. The method further includes performing, using the at least one processing device, an online stereo rectification of the stereo pair of image frames based on the extrinsic parameters such that epipolar lines of the left and right image frames are horizontally aligned to generate a rectified stereo pair of image frames. In addition, the method includes rendering, using the at least one processing device, one or more images for display based on the rectified stereo pair of image frames.
Owner:SAMSUNG ELECTRONICS CO LTD

Methods and devices for vertex pose adjustment using passthrough and time warp transforms for video see-through (VST) extended reality (XR)

A method includes: determining a first set of vertex adjustment values ​​for a distorted mesh at an extended reality (XR) device, and receiving image frame data of a scene captured at a first time in a first head pose using a perspective camera of the XR device. The method further includes: applying the first set of vertex adjustment values ​​for the distorted mesh to the image frame data to obtain intermediate image data, and predicting a second head pose at a second time after the first time. The method further includes: generating a second set of vertex adjustment values ​​for the distorted mesh based on the predicted second head pose, applying the second set of vertex adjustment values ​​for the distorted mesh to the intermediate image data to generate a rendered virtual frame, and displaying the rendered virtual frame by the XR device at a second time.
Owner:SAMSUNG ELECTRONICS CO LTD

Image display method, mirror image display method in extended reality space, apparatuses, electronic devices, and medium

The present disclosure relates to an image display method and apparatus, and an electronic device, a mirror image display method and apparatus in an extended reality space, an electronic device, and a medium. The image display method includes: determining a first orientation of a mirror camera; determining a first position of the mirror camera based on a position of a first-person perspective camera, a preset distance, and the first orientation; and obtaining a first image based on the first position and the first orientation, and displaying the first image on a virtual mirror.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

A 3D object detection method and system based on multi-view

The present invention discloses a multi-perspective 3D target detection method and system, which belongs to the field of computer vision technology. The method is optimized based on the original three-dimensional target detection model PETR, introduces 3DRoPE to replace the original 3D position encoding method, and sets learnable parameters for 3D position information to enhance adaptability. In addition, by fusing the pose geometry information of multi-perspective cameras into the position embedding of each perspective image, the model's ability to handle complex spatial relationships is further improved. After training with the NuScenes dataset, this model can receive real-time image input from multiple cameras and output accurate 3D target detection results, significantly improving the perception accuracy and safety in autonomous driving environments.
Owner:ZHEJIANG NORMAL UNIV

System and method for disoccluded region completion for video rendering in video see-through (VST) augmented reality (AR)

A method includes generating a virtual view image and a virtual depth map based on an image captured using a see-through camera and a corresponding depth map. The virtual view image and the virtual depth map include holes for which image data or depth data cannot be determined. The method also includes searching one or more previous images to locate a region in at least one previous image that includes missing pixels in the holes. The method further includes at least partially filling the holes in the virtual view image and the virtual depth map with image data and depth data associated with the located region to generate a filled virtual view image and a filled virtual depth map. In addition, the method includes generating a virtual view to present on a display panel of a VST AR device using the filled virtual view image and the filled virtual depth map.
Owner:SAMSUNG ELECTRONICS CO LTD

Complementary dynamic perspective camera coordination method based on polar coordinate transformation

The application discloses a kind of complementary dynamic visual angle camera coordination methods based on polar coordinate transformation, including steps: S1, target detection network, camera positioning network and cross-view target association network training;S2, top view and side view are applied to target detection network to obtain two view target positions;S3, human detection module is generated according to two view target positions Calculation generates two-view heat map;S4, according to the position of the camera positioning network identified by the photographer to the origin of the side view photographer Polar coordinate transformation is carried out on the view heat map;S5, according to the heat map after polar coordinate transformation Application search network of viewing direction determines side view angle;S6, based on the target position identified in step S3 and the side view angle determined in step S5, calculate target similarity matrix and input cross-view target association network to obtain matching matrix output, end, the application is measured with tracking algorithm Similarity of cross-view target, effectively ensure the accuracy of measurement result and the performance of cross-view association.
Owner:TIANJIN UNIV

Display device

The invention discloses a display device, and belongs to the technical field of communication. The display device comprises a color perspective camera, a gray scale tracking camera, a processor and a display assembly. Wherein the color perspective camera is used for collecting a first image, and the gray scale tracking camera is used for collecting a second image; the processor is used for acquiring a first mask of a target object according to the first image and acquiring coordinate values of a plurality of feature points of the target object according to the second image; acquiring a depth value of a depth reference point according to the coordinate values of the plurality of feature points and a first mask of the target object; obtaining a depth difference of the depth reference point relative to a target virtual plane according to the depth value of the depth reference point; according to the depth difference, determining the display transparency of at least part of the target object relative to the target virtual plane; and the display component displays the target virtual plane and the target object according to the display transparency.
Owner:VIVO MOBILE COMM CO LTD

Video see-through augmented reality

In one embodiment, a method includes capturing, by a calibration camera, a calibration pattern displayed on a display of a video see-through AR system, where the calibration camera is located at an eye position for viewing content on the video see-through AR system. The method further includes determining, based on the captured calibration pattern, one or more system parameters of the video see-through AR system that represent combined distortion caused by a display lens of the video see-through AR system and caused by a camera lens of the see-through camera; determining, based on the one or more system parameters and on one or more camera parameters that represent distortion caused by the see-through camera, one or more display-lens parameters that represent distortion caused by the display lens; and storing the one or more system parameters and the one more display-lens parameters as a calibration for the system.
Owner:SAMSUNG ELECTRONICS CO LTD

A time sequence fused three-dimensional target detection method, system, device and medium

Embodiments of the present application provide a kind of three-dimensional target detection method, system, equipment and medium of timing fusion, method includes: the view angle of multiple encoding features based on perspective camera is converted to the view angle of first encoding feature in the multiple encoding features, obtains corresponding multiple target encoding features;The multiple target encoding features and the first encoding feature are superimposed, and the camera feature of the first encoding feature corresponding time is obtained;By the camera feature of multiple view angle cameras is converted in feature space, corresponding multiple space features are obtained;The multiple space features are input into three-dimensional detection head and are processed, and three-dimensional target detection result is obtained.It aims at improving the depth perception accuracy of three-dimensional target, and then improve the accuracy of three-dimensional target detection.
Owner:CHONGQING CHANGAN TECH CO LTD

Newborn feeding behavior image recognition and oral motor assessment system

This invention relates to the field of medical image processing technology, and in particular to a system for image recognition and oral motor assessment of neonatal feeding behavior. The system includes a data acquisition module, a feature extraction module, a differential geometry modeling module, a trajectory analysis module, a cooperative motion analysis module, an anomaly detection module, a result display module, and an application service module. The system employs a high-speed camera, a 3D facial tracking camera, and an optical perspective camera working collaboratively. It simultaneously acquires multi-view image data of neonatal facial and oral motor movements through a data acquisition and analysis instrument. Innovatively, it applies differential geometry theory to oral motor analysis, treating facial feature point groups as Riemannian manifolds embedded in 3D Euclidean space. Through curvature feature calculation, shape operator construction, and differential invariant extraction, it accurately characterizes oral motor properties. The system constructs a coupled model of the lip manifold and the mandibular manifold, analyzing the synchronicity measurement and topological features between the manifolds to achieve accurate detection of feeding abnormalities.
Owner:THE SEVENTH MEDICAL CENTER OF PLA GENERAL HOSPITAL

Efficient depth-based viewpoint matching and head pose change compensation for video perspective (VST) augmented reality (XR)

A video perspective (VST) augmented reality (XR) apparatus includes a perspective camera configured to capture image frames of a three-dimensional (3D) scene; a display panel; and at least one processing device. The at least one processing device is configured to obtain an image frame, identify a depth-based transformation in the 3D space, transform the image frame into a transformed image frame based on the depth-based transformation, and initiate presentation of the transformed image frame on the display panel. The depth-based transformation provides viewpoint matching between the head pose of the VST XR device when an image frame is captured and the head pose of the VST XR device when a transformed image frame is presented, parallax correction between the head pose of the VST XR device when an image frame is captured and the head pose of the VST XR device when a transformed image frame is presented, and compensation for a change between the head pose of the VST XR device when the image frame is captured and the head pose of the VST XR device when the transformed image frame is presented.
Owner:SAMSUNG ELECTRONICS CO LTD

Weld seam scanning and tracking processing method

A weld seam scanning and tracking processing method is provided. Based on the laser scanning of the weld seam, the scanned laser data is converted into three-dimensional space point cloud data using the robot coordinate transformation relationship, and the spatial point cloud structure is used to remove noise in the single measurement data in a small range, fill or interpolate the point cloud data of the missing position, and more accurately reflect the weld seam morphology. A unique weld seam feature library is established, and the weld seam is converted from 2D image data to more intuitive 3D point cloud data. This solves the problem of inaccurate measurement results caused by optical lens distortion of the laser or perspective camera during weld seam tracking. After multiple weld bead feature recognition and correction of the segmented scanned data, the accurate weld seam morphology characteristics and position are determined through weld bead information reconstruction, eliminating the problem of poor accuracy of single measurement results. The coherent measurement processing method can accurately reflect the true situation of the weld bead and has high weld seam measurement accuracy.
Owner:XIXIAN NEW AREA URSA MAJOR INTELLIGENT TECH CO LTD