Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

244 results about "Depth perception" patented technology

Depth perception is the visual ability to perceive the world in three dimensions (3D) and the distance of an object. Depth sensation is the corresponding term for animals, since although it is known that animals can sense the distance of an object (because of their ability to move accurately, or to respond consistently, according to that distance), it is not known whether they "perceive" it in the same subjective way that humans do.

Flotation froth dynamic diagnosis and self-adaptive regulation and control system based on multi-mode depth perception and time sequence prediction

The invention discloses a flotation froth dynamic diagnosis and self-adaptive regulation and control system based on multi-mode depth perception and time sequence prediction. The flotation froth dynamic diagnosis and self-adaptive regulation and control system aims at solving the problems that in the prior art, the flotation process is not comprehensive in monitoring perception, dynamic prediction is missing, and regulation and control self-adaptability is poor. According to the invention, by deploying a multi-source heterogeneous sensor array, multi-modal data of vision, spectrum, acoustics and the like of foam are synchronously collected; and generating comprehensive foam comprehensive state characterization by using a cross-modal attention fusion network. Modeling is carried out on dynamic evolution of foam by adopting a hierarchical time sequence prediction and anomaly detection network, the future state trend is accurately predicted, and early warning of anomaly is realized. And finally, an intelligent regulation and control agent based on deep reinforcement learning is constructed, the intelligent regulation and control agent autonomously decides optimal process parameter adjustment according to the current state and future prediction, and online learning and optimization are carried out through continuous interaction with the actual process. The beneficiation recovery rate, the grade and the stability of the production process are remarkably improved, and the operation cost is reduced.
Owner:ZHEJIANG AILINGCHUANG MINING INDUSTRY TECHNOLOGY CO LTD

Depth map generation method and device based on large model, three-dimensional reconstruction method and device, electronic equipment and storage medium

The invention provides a depth map generation method and device based on a large model, a three-dimensional reconstruction method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, can be applied to real-time road scene depth perception, environment three-dimensional reconstruction and obstacle avoidance, and can be applied to real-time road scene depth perception. And virtual and real scene fusion and other scenes can be realized. The specific implementation scheme is as follows: performing visual coding on a monocular image to obtain a coded image; inputting the coded image and the target text into a pre-trained large language model for fusion to obtain fusion features; generating global guide features based on the fusion features, wherein the global guide features comprise joint semantic information of visual features and text features; adding noise to the color image of the monocular image to obtain a noise feature sequence; de-noising the noise feature sequence under the condition of the global guide feature, and generating an implicit feature matched with the joint semantic information; a depth map is generated based on the implicit features.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Naked eye 3D display optimization method based on real-time eyeball tracking

The invention discloses a naked-eye 3D display optimization method based on real-time eyeball tracking, and particularly relates to the technical field of naked-eye 3D display, and the method comprises the steps: capturing eyeball movement data in real time through a visual angle tracking sensor, and collecting illumination information in combination with an ambient light sensor; depth perception parameters are calculated based on pupil diameter variation and frequency, and a time sequence prediction type dynamic compensation coefficient is generated in combination with eyeball movement acceleration; acquiring an initial fixation point coordinate by using an improved spherical projection mapping model, and performing compensation coefficient correction to obtain a real-time coordinate; and finally, according to the real-time coordinates, dynamically adjusting the refractive index distribution of the nanostructure layer, the rotation angle of the polarizer, the focal length of the optical lens and other optical modulation parameters. The real-time distance is calculated through the binocular parallax algorithm, the depth mapping value is generated by combining the focal length of the camera and the baseline distance, the method can adapt to the illumination change and the user view angle, and the 3D display effect is optimized.
Owner:SHENZHEN EASYQUICK TECH CO LTD

Industrial robot intelligent obstacle avoidance control method and system based on visual identification

The invention discloses an industrial robot intelligent obstacle avoidance control method and system based on visual identification, and relates to the technical field of robot obstacle avoidance control, and the method comprises the steps: carrying out the calibration of a multi-mode visual sensor, obtaining an initial depth map and point cloud information, and combining the visual identification and depth compensation technology; recognizing and positioning obstacles in the operation area, and constructing a dynamic environment map layer; and the current robot state is collected, kinematics calculation and collision distance analysis are carried out, whether obstacle avoidance operation needs to be executed or not is judged, if yes, an obstacle avoidance path is generated in combination with the dynamic environment map layer and the target point location, feasibility verification is carried out after the path is generated, and the path passing the verification serves as an execution track to be issued to the control module. According to the invention, the recognition precision and depth perception integrity of the industrial robot on obstacles in a complex environment are improved, high feasibility of path planning and high-reliability obstacle avoidance capability in a dynamic environment are realized, and the intelligent decision-making level and operation safety of the system are remarkably enhanced.
Owner:JIAERXIN (JIANGSU) ENGINEERING EQUIPMENT CO LTD

Human body posture estimation method, system, equipment and medium

The invention discloses a human body posture estimation method, system and device and a medium, and relates to the technical field of computer vision, and the method comprises the steps: obtaining a to-be-estimated picture containing a human body; the method comprises the following steps of: performing feature extraction and down-sampling on a to-be-estimated picture, performing dimension raising, depth separable convolution operation, channel aggregation operation and dimension reduction on a feature map with down-sampling resolution in an output branch after down-sampling to obtain a local feature map, performing up-sampling reconstruction on the local feature map by adopting dynamic weight interpolation, and fusing an output branch which is not down-sampled to obtain a local feature map; obtaining a first fusion feature; taking the first fusion feature as an initial feature, repeating the step of obtaining the first fusion feature, obtaining a second fusion feature, carrying out fusion to obtain a dual-scale fusion feature, extracting a depth perception feature in the dual-scale fusion feature, carrying out human body posture estimation through the depth perception feature, and obtaining a human body key point heat map. According to the invention, multi-scale and global information is obtained through a lightweight structure, and accurate key point positioning is obtained.
Owner:WUXI UNIV

4D multi-target sensing and tracking method and system based on sparse representation

PendingCN121191137ABiological modelsScene recognitionSparse methodsAlgorithm
The invention relates to the technical field of target detection of automatic driving, in particular to a 4D multi-target sensing and tracking method and system based on sparse representation, which is an efficient 3D target detection algorithm, and through dynamic interaction of sparse 4D query vectors and multi-view and multi-scale features and in combination with a time sequence instance denoising and depth sensing enhancement module, the accuracy of target detection is improved. And automatic driving 3D perception with high precision and low calculation amount is realized. Benefited from a sparse query architecture, the method also inherits the advantage of high efficiency of a sparse method while keeping high precision. Through a series of targeted optimization, the generalization ability and robustness of the model in various special scenes are significantly enhanced.
Owner:HONEYCOMB (WUHAN) MICROSYSTEM TECH CO LTD

Automatic monocular depth perception calibration for camera

A system and method are provided for estimating a depth of an object within an image and training a depth estimation network. The depth estimation method includes: obtaining single-frame image data; obtaining scaling factor data based on the single-frame image data; generating scale-invariant depth data through inputting the single-frame image data into a depth estimation network; and generating metric depth data based on the scaling factor data and the scale-invariant depth data. The training method includes: inputting image data into a teacher machine learning (ML) model in order to generate metric depth data; inputting image data into a student ML model in order to generate scale-invariant depth data; and training a student network based on loss calculated using the metric depth data and the scale-invariant depth data.
Owner:FAURECIA IRYSTEC INC

Food nutrition evaluation method and system based on multi-modal fusion

The invention discloses a food nutrition evaluation method and system based on multi-modal fusion. The method comprises the steps of 1, extracting multi-scale features of a food RGB image and an RGB-D depth image in parallel through a multi-scale double-branch Transform encoder; 2, performing cross-modal fusion on the multi-scale features by using a depth perception enhanced attention fusion module; step 3, constructing nutrition prediction branches based on the fused features, and outputting five macro nutrient contents; and step 4, training the training data set on the model in an end-to-end mode, and inputting the RGB image and the RGB-D image of the food to be detected into the trained model to predict the nutrient content. According to the method, perception of volume information of different food types is enhanced by using the RGB-D depth mode corresponding to the food RGB image, visual information in the food RGB image and spatial physical characteristics in the RGB-D image are fused through the depth perception attention enhancement module, so that fusion characteristics with identification information are generated, and the evaluation accuracy is improved.
Owner:JIANGNAN UNIV

Multi-point radar camera fusion detection method based on aerial view

The invention discloses a multi-point radar camera fusion detection method based on a bird's-eye view, and the method comprises the steps: constructing an improved BEVDepth network model, employing sparse but precise radar points, introducing a radar auxiliary view conversion method RVT, converting the image features of a perspective view into a bird's-eye view BEV, making up the defects of an image in depth perception, and improving the detection precision of the bird's-eye view BEV. And more comprehensive space understanding is provided for the automatic driving system. A multi-modal deformable attention mechanism is used to further aggregate an image and a radar feature map in BEV so as to eliminate spatial dislocation between different sensor data, thereby improving perception precision. On the basis of adopting a BEVDepth model, the training of the BEVDepth model, a backbone network and the tail end of the backbone network are improved, radar data are fused, the domain invariance of the network is enhanced, and the recognition capability of the system on a 3D target is improved.
Owner:HUAIYIN INSTITUTE OF TECHNOLOGY

System for underwater depth perception having multiple image sensors

Systems described herein use either multiple image sensors or one image sensor and a complementary sensor to provide underwater depth perception. The systems include a computing system that provides the underwater depth perception based on data sensed by the multiple image sensors or the one image sensor and the complementary sensor. The systems can include a submersible device (such as a submersible mobile machine) that includes a holder configured to hold the one image sensor. The holder can be configured to hold the computing system in addition to the one image sensor. And, in some embodiments, the holder is configured to hold the complementary sensor in addition to the computing system and the one image sensor. Alternatively, in some embodiments, the holder is configured to hold the multiple image sensors. And, the holder can be configured to hold the computing system in addition to the multiple image sensors.
Owner:VOYIS IMAGING INC

3D GS cultural relic digital reconstruction method and system based on block chain

The invention discloses a 3D GS cultural relic digital reconstruction method and system based on a block chain, and the method comprises the steps: collecting the RGB image data and depth perception data of a cultural relic, eliminating the influence of different shooting conditions through an illumination separation processing technology, and building a standardized image data set; recognizing a surface area suitable for reconstruction based on image analysis, determining feature point distribution by using kernel density estimation, and generating initial three-dimensional representation through Gaussian ellipsoid fitting; performing gradient calculation and feature extraction on the depth data, and combining with Gaussian representation to form a geometric constraint mechanism; self-adaptive encryption based on visual importance is executed for a sparse region, and a layered rendering effect is achieved through opacity parameter adjustment; the rendering characteristics and the conversion relation of different view angles are analyzed, key observation points are determined through stability analysis, and a smooth multi-view-angle display sequence is constructed; and integrating multi-view rendering information to generate volumetric representation, and completing right confirmation of the high-quality three-dimensional digital model through digital signature.
Owner:HONG KONG LARGE (HANGZHOU) TECHNOLOGY INNOVATION RESEARCH INSTITUTE CO LTD +2

Ultrasonic transmission CT imaging method and system, computer and storage medium

The invention provides an ultrasonic transmission CT (Computed Tomography) imaging method and system, a computer and a storage medium, which can be used for quickly reconstructing an ultrasonic transmission image with high resolution and high contrast. The method comprises the following steps: in layer-by-layer scanning of an annular ultrasonic transducer array, calculating a TOF (Time of Flight) matrix containing transmission time of all ultrasonic paths after the annular ultrasonic transducer array rotates by a specific angle each time; after one layer of scanning is completed, the TOF matrixes of different angles are recombined into a three-dimensional matrix according to the spatial relation, and missing data are complemented; discretizing an area to be imaged into two-dimensional grids, making each grid correspond to a pixel point, counting transmission time and path length of all ultrasonic paths penetrating through the pixel point, calculating sound velocity contribution values of the paths, taking an average value as a sound velocity value of the pixel point, and generating a two-dimensional ultrasonic transmission image; a two-dimensional image is converted into a three-dimensional image through a ray casting algorithm, the visual influence of voxels is adjusted by introducing a weight factor based on depth, and the depth perception and semitransparent rendering effect of the image are enhanced.
Owner:ZHONGBEI UNIV

Phytoplankton chromatography sequence identification method and phytoplankton chromatography sequence model building method

The invention provides a phytoplankton chromatography sequence identification method and a phytoplankton chromatography sequence model building method, and belongs to the technical field of image enhancement identification. The method comprises the following steps: firstly, acquiring microscopic chromatography sequence data of phytoplankton, performing view field extraction and serialization recombination, and constructing a three-dimensional data set; then, constructing a three-dimensional recognition model containing physical perception and a sequence aggregation mechanism, extracting single-frame semantic features by the model by adopting a parameter-shared twin network, and introducing a physical definition prior module to calculate a space-frequency domain quality score of a slice; secondly, designing a deep perception sequence aggregation module, and adaptively aggregating key features of a high signal-to-noise ratio by taking definition scores as gating signals and combining spatial context information between slices; and finally, training and optimizing the model based on the image-level weak supervision label to obtain an optimal model. According to the method, the problems of information truncation and out-of-focus noise interference caused by extremely shallow depth of field of high-power microscopic imaging are solved, and full-depth-of-field stereoscopic perception can be realized under the condition that frame-by-frame fine labeling is not needed.
Owner:OCEAN UNIV OF CHINA

SLAM three-dimensional modeling method suitable for complex dynamic scene

The invention belongs to the technical field of image processing and three-dimensional modeling, and particularly relates to an SLAM three-dimensional modeling method suitable for a complex dynamic scene. By constructing a multi-modal feature hierarchical processing architecture, the system realizes technical breakthrough in three dimensions: firstly, in combination with a pixel mask pattern and a multi-view geometric method, static feature points are reserved while dynamic feature points are removed as far as possible; secondly, according to the optimal flow field mapping relation of the grid nodes, static background restoration of the shielded area is achieved by establishing the geometric constraint relation of adjacent key frames; and finally, the depth perception data and the optimized feature set are fused to construct a tightly coupled three-dimensional reconstruction model, and a closed-loop optimization framework with motion robustness is formed.
Owner:ELECTRIC POWER SCI RES INST OF STATE GRID XINJIANG ELECTRIC POWER CO LTD

View view generation method of cross-modal fusion and multi-frequency coding in ultra-low orbit scene

The invention discloses a cross-modal fusion and multi-frequency coding view angle image generation method in an ultra-low orbit scene, and solves the problem of view angle condition generation and single image three-dimensional reconstruction method based on diffusion prior in the ultra-low orbit scene of an aircraft target. Due to insufficient condition information fusion capability, limited radiation or illumination characterization and insufficient cross-view-angle geometric constraint, a problem that a new view angle image with visual coherence and geometric consistency of texture details is difficult to obtain at the same time under large-range view angle and scale change is solved; according to the method, multi-order angle information of direction distribution is effectively captured through perceptual coding of a light direction vector, then the position of a light starting point coordinate is coded, and finally the pose of a camera pose parameter is coded, so that the influence of a relative pose on depth perception and perspective distortion in an image generation process can be depicted; the three coding results are used for a decoding stage of the potential diffusion model, and geometric consistency of detail recovery under a new view angle is effectively improved.
Owner:XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI

Method for constructing and editing three-dimensional light field content

The invention discloses a three-dimensional light field content construction and editing method, which relates to the field of naked eye three-dimensional display, and comprises the following steps of: collecting multi-view images of a real scene through a camera, and constructing a unified three-dimensional space coordinate system; generating an initial three-dimensional scene model; performing adjustable content editing on each frame of multi-view image; updating the initial three-dimensional scene model through the edited image; the updated three-dimensional scene model is rendered through a shear type rendering method, and an image sequence with multi-view consistency is generated; and coding the generated image sequence into three-dimensional light field content by using a light field coding algorithm. According to the method, the real static scene can be efficiently converted into the three-dimensional light field content with a complete structure and a continuous visual angle through multi-visual-angle image acquisition, three-dimensional reconstruction and rendering processing, so that the method has stronger sense of space depth and reality, and the visual experience and the display effect of the three-dimensional light field content are remarkably improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Layout generation method and device based on multi-granularity attention, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a layout generation method, device, equipment and medium based on multi-granularity attention, and the method comprises the steps: obtaining original depth data, carrying out the standardization processing of the original depth data, and obtaining the standardized depth data; the method comprises the steps of obtaining standardized depth data, extracting a multi-scale depth perception feature set from the standardized depth data, generating a multi-granularity attention weight set based on the multi-scale depth perception feature set and the standardized depth data, performing weighted fusion processing on the multi-granularity attention weight set to generate fusion features, and generating layout data based on the fusion features. Through a depth-driven feature extraction mechanism and an attention resolution regulation and control method, granularity self-adaptive modeling of different space depth regions is realized, so that the detail retention capability of a near-end region is improved, the calculation redundancy of a far-end region is reduced, and the spatial resolution capability of layout output and the overall modeling efficiency are effectively enhanced.
Owner:PING AN TECH (SHENZHEN) CO LTD

Personalized image content generation and optimization method and system fusing generative AI

The invention provides a personalized image content generation and optimization method and system fusing generative AI, and relates to the technical field of AI. The personalized image content generation and optimization method comprises the steps of analyzing a historical interaction track and a current intention expression of a target user based on a deep perception network, and constructing a multi-dimensional behavior portrait; user groups with similar generation preferences are identified by using a group feature recursive quantization technology, and a group wisdom feature map is constructed; performing multi-level mapping analysis on the user features and the group wisdom feature map, anchoring the group affiliation relationship of the target user and concretizing the guide features; constructing a multi-level guide vector according to current intention expression and guide features, performing accurate mapping on hidden space expression of an image generation model, quantizing deviation in real time and triggering intelligent calibration; and reconstructing a group distribution structure based on the evaluation data of the target user to realize dynamic evolution of the features. According to the invention, personalized demands of users can be accurately grasped, and the accuracy of image generation and the satisfaction degree of the users are improved.
Owner:SMIC WANYE TECHNOLOGY CO LTD

Crack detection and segmentation method based on depth perception and structural feature enhancement

The invention relates to the field of image processing and target detection, in particular to a crack detection and segmentation method based on depth perception and structural feature enhancement. The method comprises the following steps: acquiring original image data containing cracks; performing multi-scale feature extraction on the image; multi-scale feature extraction is realized in combination with a C3K2SAConv module capable of switching cavity convolution; introducing a SimAM attention mechanism to enhance crack area response; a CCM crack convolution module is designed to extract a slender and fractured crack edge structure, and detection and segmentation accuracy and boundary integrity are improved; and finally completing positioning and segmentation output of the crack region. The method improves the recognition accuracy and edge integrity of the structural crack in a complex background and low contrast environment, and is suitable for application scenes of structure monitoring, intelligent maintenance, automatic driving environment perception and the like of traffic infrastructures such as roads, bridges and the like.
Owner:南宁桂电电子科技研究院有限公司 +1

Intelligent traffic cone barrel anti-collision early warning method based on binocular vision

The invention provides an intelligent traffic cone barrel anti-collision early warning method based on binocular vision, and aims to improve the safety of a road construction operation area. And accurate data are provided for subsequent detection and analysis through a binocular camera high-precision calibration technology based on circle detection. In combination with a data acquisition and enhancement technology, high-quality training data is provided for an improved RT-DETR algorithm, and the accuracy and efficiency of vehicle identification are significantly improved. A binocular vision depth perception and StrongSORT multi-target tracking algorithm is utilized to realize real-time speed measurement and distance measurement of a vehicle, and accurate input is provided for track prediction. Based on a GAN multi-modal trajectory prediction technology, a plurality of future trajectories are generated, and a multi-level early warning decision is realized through a comprehensive collision risk index (CRI). According to the method, the limitation of an existing prediction method in a complex road scene is effectively solved, the vehicle behavior intention is predicted and analyzed through the multi-modal trajectory, and reliable early warning information can be provided under various complex road conditions.
Owner:GUANGXI UNIV +1

Intelligent multi-camera tracheal intubation image processing method, device and system

The invention discloses an intelligent multi-view camera tracheal intubation image processing method, device and system, and relates to the technical field of medical instruments, the method comprises the following steps: in a tracheal intubation process, a multi-view image acquisition unit collects a plurality of images and combines the images to form an image pair; and carrying out stereo matching and depth calculation on the image pairs to generate a depth map. And performing three-dimensional reconstruction by combining the depth map and the original image to generate a three-dimensional model. And finally, identifying the target in the three-dimensional model and outputting the depth of the target as an image processing result. The technical problems that an existing multi-view camera is inaccurate in depth perception and insufficient in real-time performance during tracheal intubation image processing are solved, and the technical effect of improving the depth perception precision and the real-time processing capacity of the multi-view camera during tracheal intubation image processing is achieved.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Robot autonomous navigation method and device, robot and computer readable storage medium

The invention relates to the technical field of robot control, in particular to a robot autonomous navigation method and device, a robot and a computer readable storage medium, and the method comprises the following steps: obtaining real-time body state information, a current position, a target position and depth perception information of an external environment of the robot; encoding the depth perception information, and splicing the encoded depth perception information with the real-time body state information to obtain unified features representing the external environment and the state of the robot; inputting the unified feature and the target position into a preset navigation decision model to generate a preliminary movement speed of the robot; the linear distance between the current position and the target position is calculated in real time, the initial movement speed is adjusted based on the linear distance, and a stable expected speed parameter is obtained; and according to the stable expected speed parameter and the real-time body state information, generating an execution instruction for driving the robot to move and controlling execution, and navigating to a target position to realize autonomous navigation of the robot in a map-free scene.
Owner:PEKING UNIV

Self-service zooming method of camera system based on depth-guided image quality evaluation

The invention provides a self-service zooming method of a camera system based on depth-guided image quality evaluation. The method comprises three core links of multi-scale depth perception, depth-guided image quality evaluation and joint-driven intelligent focusing control. By fusing lightweight feature coding and a multi-scale optimization strategy, accurate depth information is generated, and a spatial perception basis is provided for subsequent processing; the depth features and the image texture features are combined, a focusing fuzzy evaluation model is constructed, and reliable feedback of the imaging quality is achieved; a key focusing area is identified based on depth and quality information, a target focal plane is decided, and a lens is driven through an intelligent control algorithm to complete rapid and accurate focusing. Through collaborative fusion of depth perception and image quality evaluation, intelligent zoom control of the camera system in a dynamic scene is realized, the imaging consistency and focusing precision are effectively improved, the method is suitable for multiple camera devices such as smart phones, security monitoring and automatic driving, and the imaging experience of users is improved.
Owner:TIANJIN UNIV

Parallax estimation method based on uncalibrated binocular image

The invention relates to the technical field of image data processing, in particular to a parallax estimation method based on an uncalibrated binocular image. The method comprises the following steps: acquiring an original binocular image pair, and preprocessing the original binocular image pair to obtain standardized binocular image data; generating a center thermodynamic diagram label of the object based on the standardized binocular image data and a preset data set; marking the standardized binocular image data and the corresponding center thermodynamic diagram label as a training sample set; training a preset SAM decoder based on the training sample set to obtain a thermodynamic diagram generation decoder; inputting the original binocular image pair into a thermodynamic diagram generation decoder to obtain a left and right view thermodynamic diagram; performing fixed threshold segmentation on the left and right view thermodynamic diagrams to obtain initial left and right mask images; according to the method, thermodynamic diagram guidance, feature fusion matching and frequency domain analysis strategies are combined, and low-cost, high-precision and high-adaptability target-level depth perception is realized under the condition of weak calibration.
Owner:SHENZHEN POLYTECHNIC

Monocular 3D object detection method for realizing depth enhancement based on visual basic model, electronic equipment and readable storage medium

The invention belongs to the technical field of computer vision, and particularly discloses a monocular 3D object detection method for realizing depth enhancement based on a visual basic model, electronic equipment and a readable storage medium, and the method comprises the steps: S1, building a data set: employing a monocular camera to collect a pavement scene, and obtaining an RGB image in the pavement scene; s2, image preprocessing: preprocessing the RGB image for subsequent feature extraction and depth estimation; s3, performing feature extraction by adopting a dual-backbone network: performing visual semantic feature extraction on the preprocessed RGB image by using DINOv2; performing depth feature extraction on the preprocessed RGB image by using a DPT head; s4, generation of depth perception query points: inputting the visual semantic features and the depth features into a DETR network to generate the depth perception query points; and S5, target detection output: using an MLP-based detection head to obtain information of the category, the size, the center point position, the depth, the 3D size and the direction of the object.
Owner:SHANGHAI UNIV

Pseudo 3D display method and system based on eye movement tracking and affine transformation

The invention discloses a pseudo 3D display method and system based on eye movement tracking and affine transformation, particularly relates to the technical field of display and man-machine interaction, and comprises the steps of 3D model preprocessing, real-time eye movement tracking, view angle parameter calculation, dynamic projection generation, affine transformation correction, display image adjustment and dynamic pseudo 3D display effect realization. The pseudo 3D effect that the visual angle is dynamically adjusted along with the eye position of the user can be achieved through a common 2D screen, special 3D display equipment is not needed, pixel-level perspective correction is achieved, the depth perception of a pseudo 3D image observed by the user is closer to a real three-dimensional space, the visual immersion and interaction naturalness are effectively improved, and the user experience is improved. The core problems that traditional pseudo 3D depth perception is fuzzy and distortion is easily generated due to view angle deviation are solved, the user immersion and interaction naturalness are remarkably improved, and the method can be widely applied to scenes such as virtual social contact, vehicle-mounted HMI, telemedicine and education simulation.
Owner:SHANGHAI ZIHAI TECHNOLOGY CO LTD

Visual training method, system and device for myopia prevention and control and correction based on naked eye 3D display and storage medium

The invention discloses a visual training method, system and device for myopia prevention and control and correction based on a naked-eye 3D display, and a storage medium, relates to the technical field of three-dimensional image generation and display control, and comprises the field of visual training of myopia prevention and control constructed on a naked-eye 3D display interface, a first visual training area and a second sensing and integrating area, in the first visual training area, through 3D interlaced pictures and videos, fine or dynamic stereoscopic vision is stimulated, front and back intersections of sight lines are adjusted, and split vision training is carried out; in the second perception and integration area, through perception, spatial positioning, binocular coordination and deep perception training, the spatial ability and response ability of eyes are stimulated; in combination with a three-dimensional display mechanism, displaying a naked eye three-dimensional image on a display, acquiring eyeball position information by adopting a human eye tracking technology, and dynamically adjusting a three-dimensional image display area; according to the method, the visual content structured organization and three-dimensional generation cooperative control is realized, and the three-dimensional display interaction matching precision is improved.
Owner:TIANJIN VISION TECHNOLOGY CO LTD

3D brain-click using binocular display

A method and system for detecting intentional selection of a user interface element using a binocular display. A first visual stimulus is presented stereoscopically to a user's eyes at a first virtual depth perceived by the user's depth perception and overlapping a first position within a field of view of the user. A second visual stimulus is presented stereoscopically to the user's eyes at a second virtual depth perceived by the user's depth perception and overlapping the first position. Neural signals are obtained from a neural signal capture device configured to detect neural activity of the user. In response to determining, based on the neural signals, that the user's eyes are focused on either the first visual stimulus or second visual stimulus, a computing system is placed into a first state or second state, respectively, associated with the first visual stimulus or second visual stimulus, respectively.
Owner:SNAP INC

Methods and systems for enhancing depth perception of a non-visible spectrum image of a scene

A method and system for providing depth perception to a two-dimensional (2D) representation of a given three-dimensional (3D) object within a 2D non-visible spectrum image of a scene is provided. The method comprises: capturing the 2D non-visible spectrum image at a capture time, by at least one non-visible spectrum sensor; obtaining 3D data regarding the given 3D object independently of the 2D non-visible spectrum image; generating one or more depth cues based on the 3D data; applying the depth cues on the 2D representation to generate a depth perception image that provides the depth perception to the 2D representation; and displaying the depth perception image.
Owner:ELBIT SYSTEMS LTD

Naked eye 3D display method based on parallax compensation

The invention provides a naked eye 3D display method based on parallax compensation, and relates to the technical field of parallax compensation, the parallax deviation value of each pixel point is calculated based on the angle relationship between a reference viewpoint and each viewing angle, the global parallax modeling and preliminary compensation of image content are realized, a neighborhood window search mechanism is introduced, and the image content is obtained. Pixels with the minimum gray gradient change and consistent directions in the local area are selected as target pixels, image synthesis and visual angle generation are carried out, local parallax optimization based on image content is achieved, when it is detected that the parallax deviation value is too large or the comfort level of a user is reduced, the system dynamically adjusts the parallax deviation value, and the visual angle is generated. And performing depth perception correction on the target pixel points exceeding the preset depth range to form a closed-loop feedback control mechanism, and performing fine parallax adjustment on the multi-view image sequence through a multi-compensation mechanism of global modeling, local optimization and user feedback driving to obtain a multi-view image sequence. On the premise that image details are not sacrificed, the problem of visual fatigue caused by too large parallax is effectively relieved.
Owner:SHANGHAI KUAILAIXIU DISPLAY TECH CO LTD