Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

201 results about "Depth perception" patented technology

Depth perception is the visual ability to perceive the world in three dimensions (3D) and the distance of an object. Depth sensation is the corresponding term for animals, since although it is known that animals can sense the distance of an object (because of their ability to move accurately, or to respond consistently, according to that distance), it is not known whether they "perceive" it in the same subjective way that humans do.

Flotation froth dynamic diagnosis and self-adaptive regulation and control system based on multi-mode depth perception and time sequence prediction

The invention discloses a flotation froth dynamic diagnosis and self-adaptive regulation and control system based on multi-mode depth perception and time sequence prediction. The flotation froth dynamic diagnosis and self-adaptive regulation and control system aims at solving the problems that in the prior art, the flotation process is not comprehensive in monitoring perception, dynamic prediction is missing, and regulation and control self-adaptability is poor. According to the invention, by deploying a multi-source heterogeneous sensor array, multi-modal data of vision, spectrum, acoustics and the like of foam are synchronously collected; and generating comprehensive foam comprehensive state characterization by using a cross-modal attention fusion network. Modeling is carried out on dynamic evolution of foam by adopting a hierarchical time sequence prediction and anomaly detection network, the future state trend is accurately predicted, and early warning of anomaly is realized. And finally, an intelligent regulation and control agent based on deep reinforcement learning is constructed, the intelligent regulation and control agent autonomously decides optimal process parameter adjustment according to the current state and future prediction, and online learning and optimization are carried out through continuous interaction with the actual process. The beneficiation recovery rate, the grade and the stability of the production process are remarkably improved, and the operation cost is reduced.
Owner:ZHEJIANG AILINGCHUANG MINING INDUSTRY TECHNOLOGY CO LTD

Naked eye 3D display optimization method based on real-time eyeball tracking

The invention discloses a naked-eye 3D display optimization method based on real-time eyeball tracking, and particularly relates to the technical field of naked-eye 3D display, and the method comprises the steps: capturing eyeball movement data in real time through a visual angle tracking sensor, and collecting illumination information in combination with an ambient light sensor; depth perception parameters are calculated based on pupil diameter variation and frequency, and a time sequence prediction type dynamic compensation coefficient is generated in combination with eyeball movement acceleration; acquiring an initial fixation point coordinate by using an improved spherical projection mapping model, and performing compensation coefficient correction to obtain a real-time coordinate; and finally, according to the real-time coordinates, dynamically adjusting the refractive index distribution of the nanostructure layer, the rotation angle of the polarizer, the focal length of the optical lens and other optical modulation parameters. The real-time distance is calculated through the binocular parallax algorithm, the depth mapping value is generated by combining the focal length of the camera and the baseline distance, the method can adapt to the illumination change and the user view angle, and the 3D display effect is optimized.
Owner:SHENZHEN EASYQUICK TECH CO LTD

Industrial robot intelligent obstacle avoidance control method and system based on visual identification

The invention discloses an industrial robot intelligent obstacle avoidance control method and system based on visual identification, and relates to the technical field of robot obstacle avoidance control, and the method comprises the steps: carrying out the calibration of a multi-mode visual sensor, obtaining an initial depth map and point cloud information, and combining the visual identification and depth compensation technology; recognizing and positioning obstacles in the operation area, and constructing a dynamic environment map layer; and the current robot state is collected, kinematics calculation and collision distance analysis are carried out, whether obstacle avoidance operation needs to be executed or not is judged, if yes, an obstacle avoidance path is generated in combination with the dynamic environment map layer and the target point location, feasibility verification is carried out after the path is generated, and the path passing the verification serves as an execution track to be issued to the control module. According to the invention, the recognition precision and depth perception integrity of the industrial robot on obstacles in a complex environment are improved, high feasibility of path planning and high-reliability obstacle avoidance capability in a dynamic environment are realized, and the intelligent decision-making level and operation safety of the system are remarkably enhanced.
Owner:JIAERXIN (JIANGSU) ENGINEERING EQUIPMENT CO LTD

Human body posture estimation method, system, equipment and medium

The invention discloses a human body posture estimation method, system and device and a medium, and relates to the technical field of computer vision, and the method comprises the steps: obtaining a to-be-estimated picture containing a human body; the method comprises the following steps of: performing feature extraction and down-sampling on a to-be-estimated picture, performing dimension raising, depth separable convolution operation, channel aggregation operation and dimension reduction on a feature map with down-sampling resolution in an output branch after down-sampling to obtain a local feature map, performing up-sampling reconstruction on the local feature map by adopting dynamic weight interpolation, and fusing an output branch which is not down-sampled to obtain a local feature map; obtaining a first fusion feature; taking the first fusion feature as an initial feature, repeating the step of obtaining the first fusion feature, obtaining a second fusion feature, carrying out fusion to obtain a dual-scale fusion feature, extracting a depth perception feature in the dual-scale fusion feature, carrying out human body posture estimation through the depth perception feature, and obtaining a human body key point heat map. According to the invention, multi-scale and global information is obtained through a lightweight structure, and accurate key point positioning is obtained.
Owner:WUXI UNIV

4D multi-target sensing and tracking method and system based on sparse representation

PendingCN121191137ABiological modelsScene recognitionSparse methodsAlgorithm
The invention relates to the technical field of target detection of automatic driving, in particular to a 4D multi-target sensing and tracking method and system based on sparse representation, which is an efficient 3D target detection algorithm, and through dynamic interaction of sparse 4D query vectors and multi-view and multi-scale features and in combination with a time sequence instance denoising and depth sensing enhancement module, the accuracy of target detection is improved. And automatic driving 3D perception with high precision and low calculation amount is realized. Benefited from a sparse query architecture, the method also inherits the advantage of high efficiency of a sparse method while keeping high precision. Through a series of targeted optimization, the generalization ability and robustness of the model in various special scenes are significantly enhanced.
Owner:HONEYCOMB (WUHAN) MICROSYSTEM TECH CO LTD

3D GS cultural relic digital reconstruction method and system based on block chain

The invention discloses a 3D GS cultural relic digital reconstruction method and system based on a block chain, and the method comprises the steps: collecting the RGB image data and depth perception data of a cultural relic, eliminating the influence of different shooting conditions through an illumination separation processing technology, and building a standardized image data set; recognizing a surface area suitable for reconstruction based on image analysis, determining feature point distribution by using kernel density estimation, and generating initial three-dimensional representation through Gaussian ellipsoid fitting; performing gradient calculation and feature extraction on the depth data, and combining with Gaussian representation to form a geometric constraint mechanism; self-adaptive encryption based on visual importance is executed for a sparse region, and a layered rendering effect is achieved through opacity parameter adjustment; the rendering characteristics and the conversion relation of different view angles are analyzed, key observation points are determined through stability analysis, and a smooth multi-view-angle display sequence is constructed; and integrating multi-view rendering information to generate volumetric representation, and completing right confirmation of the high-quality three-dimensional digital model through digital signature.
Owner:HONG KONG LARGE (HANGZHOU) TECHNOLOGY INNOVATION RESEARCH INSTITUTE CO LTD +2

Phytoplankton chromatography sequence identification method and phytoplankton chromatography sequence model building method

The invention provides a phytoplankton chromatography sequence identification method and a phytoplankton chromatography sequence model building method, and belongs to the technical field of image enhancement identification. The method comprises the following steps: firstly, acquiring microscopic chromatography sequence data of phytoplankton, performing view field extraction and serialization recombination, and constructing a three-dimensional data set; then, constructing a three-dimensional recognition model containing physical perception and a sequence aggregation mechanism, extracting single-frame semantic features by the model by adopting a parameter-shared twin network, and introducing a physical definition prior module to calculate a space-frequency domain quality score of a slice; secondly, designing a deep perception sequence aggregation module, and adaptively aggregating key features of a high signal-to-noise ratio by taking definition scores as gating signals and combining spatial context information between slices; and finally, training and optimizing the model based on the image-level weak supervision label to obtain an optimal model. According to the method, the problems of information truncation and out-of-focus noise interference caused by extremely shallow depth of field of high-power microscopic imaging are solved, and full-depth-of-field stereoscopic perception can be realized under the condition that frame-by-frame fine labeling is not needed.
Owner:OCEAN UNIV OF CHINA

View view generation method of cross-modal fusion and multi-frequency coding in ultra-low orbit scene

The invention discloses a cross-modal fusion and multi-frequency coding view angle image generation method in an ultra-low orbit scene, and solves the problem of view angle condition generation and single image three-dimensional reconstruction method based on diffusion prior in the ultra-low orbit scene of an aircraft target. Due to insufficient condition information fusion capability, limited radiation or illumination characterization and insufficient cross-view-angle geometric constraint, a problem that a new view angle image with visual coherence and geometric consistency of texture details is difficult to obtain at the same time under large-range view angle and scale change is solved; according to the method, multi-order angle information of direction distribution is effectively captured through perceptual coding of a light direction vector, then the position of a light starting point coordinate is coded, and finally the pose of a camera pose parameter is coded, so that the influence of a relative pose on depth perception and perspective distortion in an image generation process can be depicted; the three coding results are used for a decoding stage of the potential diffusion model, and geometric consistency of detail recovery under a new view angle is effectively improved.
Owner:XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI

Method for constructing and editing three-dimensional light field content

The invention discloses a three-dimensional light field content construction and editing method, which relates to the field of naked eye three-dimensional display, and comprises the following steps of: collecting multi-view images of a real scene through a camera, and constructing a unified three-dimensional space coordinate system; generating an initial three-dimensional scene model; performing adjustable content editing on each frame of multi-view image; updating the initial three-dimensional scene model through the edited image; the updated three-dimensional scene model is rendered through a shear type rendering method, and an image sequence with multi-view consistency is generated; and coding the generated image sequence into three-dimensional light field content by using a light field coding algorithm. According to the method, the real static scene can be efficiently converted into the three-dimensional light field content with a complete structure and a continuous visual angle through multi-visual-angle image acquisition, three-dimensional reconstruction and rendering processing, so that the method has stronger sense of space depth and reality, and the visual experience and the display effect of the three-dimensional light field content are remarkably improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Personalized image content generation and optimization method and system fusing generative AI

The invention provides a personalized image content generation and optimization method and system fusing generative AI, and relates to the technical field of AI. The personalized image content generation and optimization method comprises the steps of analyzing a historical interaction track and a current intention expression of a target user based on a deep perception network, and constructing a multi-dimensional behavior portrait; user groups with similar generation preferences are identified by using a group feature recursive quantization technology, and a group wisdom feature map is constructed; performing multi-level mapping analysis on the user features and the group wisdom feature map, anchoring the group affiliation relationship of the target user and concretizing the guide features; constructing a multi-level guide vector according to current intention expression and guide features, performing accurate mapping on hidden space expression of an image generation model, quantizing deviation in real time and triggering intelligent calibration; and reconstructing a group distribution structure based on the evaluation data of the target user to realize dynamic evolution of the features. According to the invention, personalized demands of users can be accurately grasped, and the accuracy of image generation and the satisfaction degree of the users are improved.
Owner:SMIC WANYE TECHNOLOGY CO LTD

Crack detection and segmentation method based on depth perception and structural feature enhancement

The invention relates to the field of image processing and target detection, in particular to a crack detection and segmentation method based on depth perception and structural feature enhancement. The method comprises the following steps: acquiring original image data containing cracks; performing multi-scale feature extraction on the image; multi-scale feature extraction is realized in combination with a C3K2SAConv module capable of switching cavity convolution; introducing a SimAM attention mechanism to enhance crack area response; a CCM crack convolution module is designed to extract a slender and fractured crack edge structure, and detection and segmentation accuracy and boundary integrity are improved; and finally completing positioning and segmentation output of the crack region. The method improves the recognition accuracy and edge integrity of the structural crack in a complex background and low contrast environment, and is suitable for application scenes of structure monitoring, intelligent maintenance, automatic driving environment perception and the like of traffic infrastructures such as roads, bridges and the like.
Owner:南宁桂电电子科技研究院有限公司 +1

Intelligent traffic cone barrel anti-collision early warning method based on binocular vision

The invention provides an intelligent traffic cone barrel anti-collision early warning method based on binocular vision, and aims to improve the safety of a road construction operation area. And accurate data are provided for subsequent detection and analysis through a binocular camera high-precision calibration technology based on circle detection. In combination with a data acquisition and enhancement technology, high-quality training data is provided for an improved RT-DETR algorithm, and the accuracy and efficiency of vehicle identification are significantly improved. A binocular vision depth perception and StrongSORT multi-target tracking algorithm is utilized to realize real-time speed measurement and distance measurement of a vehicle, and accurate input is provided for track prediction. Based on a GAN multi-modal trajectory prediction technology, a plurality of future trajectories are generated, and a multi-level early warning decision is realized through a comprehensive collision risk index (CRI). According to the method, the limitation of an existing prediction method in a complex road scene is effectively solved, the vehicle behavior intention is predicted and analyzed through the multi-modal trajectory, and reliable early warning information can be provided under various complex road conditions.
Owner:GUANGXI UNIV +1

Intelligent multi-camera tracheal intubation image processing method, device and system

The invention discloses an intelligent multi-view camera tracheal intubation image processing method, device and system, and relates to the technical field of medical instruments, the method comprises the following steps: in a tracheal intubation process, a multi-view image acquisition unit collects a plurality of images and combines the images to form an image pair; and carrying out stereo matching and depth calculation on the image pairs to generate a depth map. And performing three-dimensional reconstruction by combining the depth map and the original image to generate a three-dimensional model. And finally, identifying the target in the three-dimensional model and outputting the depth of the target as an image processing result. The technical problems that an existing multi-view camera is inaccurate in depth perception and insufficient in real-time performance during tracheal intubation image processing are solved, and the technical effect of improving the depth perception precision and the real-time processing capacity of the multi-view camera during tracheal intubation image processing is achieved.
Owner:TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH

Robot autonomous navigation method and device, robot and computer readable storage medium

The invention relates to the technical field of robot control, in particular to a robot autonomous navigation method and device, a robot and a computer readable storage medium, and the method comprises the following steps: obtaining real-time body state information, a current position, a target position and depth perception information of an external environment of the robot; encoding the depth perception information, and splicing the encoded depth perception information with the real-time body state information to obtain unified features representing the external environment and the state of the robot; inputting the unified feature and the target position into a preset navigation decision model to generate a preliminary movement speed of the robot; the linear distance between the current position and the target position is calculated in real time, the initial movement speed is adjusted based on the linear distance, and a stable expected speed parameter is obtained; and according to the stable expected speed parameter and the real-time body state information, generating an execution instruction for driving the robot to move and controlling execution, and navigating to a target position to realize autonomous navigation of the robot in a map-free scene.
Owner:PEKING UNIV

Self-service zooming method of camera system based on depth-guided image quality evaluation

The invention provides a self-service zooming method of a camera system based on depth-guided image quality evaluation. The method comprises three core links of multi-scale depth perception, depth-guided image quality evaluation and joint-driven intelligent focusing control. By fusing lightweight feature coding and a multi-scale optimization strategy, accurate depth information is generated, and a spatial perception basis is provided for subsequent processing; the depth features and the image texture features are combined, a focusing fuzzy evaluation model is constructed, and reliable feedback of the imaging quality is achieved; a key focusing area is identified based on depth and quality information, a target focal plane is decided, and a lens is driven through an intelligent control algorithm to complete rapid and accurate focusing. Through collaborative fusion of depth perception and image quality evaluation, intelligent zoom control of the camera system in a dynamic scene is realized, the imaging consistency and focusing precision are effectively improved, the method is suitable for multiple camera devices such as smart phones, security monitoring and automatic driving, and the imaging experience of users is improved.
Owner:TIANJIN UNIV

Parallax estimation method based on uncalibrated binocular image

The invention relates to the technical field of image data processing, in particular to a parallax estimation method based on an uncalibrated binocular image. The method comprises the following steps: acquiring an original binocular image pair, and preprocessing the original binocular image pair to obtain standardized binocular image data; generating a center thermodynamic diagram label of the object based on the standardized binocular image data and a preset data set; marking the standardized binocular image data and the corresponding center thermodynamic diagram label as a training sample set; training a preset SAM decoder based on the training sample set to obtain a thermodynamic diagram generation decoder; inputting the original binocular image pair into a thermodynamic diagram generation decoder to obtain a left and right view thermodynamic diagram; performing fixed threshold segmentation on the left and right view thermodynamic diagrams to obtain initial left and right mask images; according to the method, thermodynamic diagram guidance, feature fusion matching and frequency domain analysis strategies are combined, and low-cost, high-precision and high-adaptability target-level depth perception is realized under the condition of weak calibration.
Owner:SHENZHEN POLYTECHNIC

Monocular 3D object detection method for realizing depth enhancement based on visual basic model, electronic equipment and readable storage medium

The invention belongs to the technical field of computer vision, and particularly discloses a monocular 3D object detection method for realizing depth enhancement based on a visual basic model, electronic equipment and a readable storage medium, and the method comprises the steps: S1, building a data set: employing a monocular camera to collect a pavement scene, and obtaining an RGB image in the pavement scene; s2, image preprocessing: preprocessing the RGB image for subsequent feature extraction and depth estimation; s3, performing feature extraction by adopting a dual-backbone network: performing visual semantic feature extraction on the preprocessed RGB image by using DINOv2; performing depth feature extraction on the preprocessed RGB image by using a DPT head; s4, generation of depth perception query points: inputting the visual semantic features and the depth features into a DETR network to generate the depth perception query points; and S5, target detection output: using an MLP-based detection head to obtain information of the category, the size, the center point position, the depth, the 3D size and the direction of the object.
Owner:SHANGHAI UNIV

Pseudo 3D display method and system based on eye movement tracking and affine transformation

The invention discloses a pseudo 3D display method and system based on eye movement tracking and affine transformation, particularly relates to the technical field of display and man-machine interaction, and comprises the steps of 3D model preprocessing, real-time eye movement tracking, view angle parameter calculation, dynamic projection generation, affine transformation correction, display image adjustment and dynamic pseudo 3D display effect realization. The pseudo 3D effect that the visual angle is dynamically adjusted along with the eye position of the user can be achieved through a common 2D screen, special 3D display equipment is not needed, pixel-level perspective correction is achieved, the depth perception of a pseudo 3D image observed by the user is closer to a real three-dimensional space, the visual immersion and interaction naturalness are effectively improved, and the user experience is improved. The core problems that traditional pseudo 3D depth perception is fuzzy and distortion is easily generated due to view angle deviation are solved, the user immersion and interaction naturalness are remarkably improved, and the method can be widely applied to scenes such as virtual social contact, vehicle-mounted HMI, telemedicine and education simulation.
Owner:SHANGHAI ZIHAI TECHNOLOGY CO LTD

Visual training method, system and device for myopia prevention and control and correction based on naked eye 3D display and storage medium

The invention discloses a visual training method, system and device for myopia prevention and control and correction based on a naked-eye 3D display, and a storage medium, relates to the technical field of three-dimensional image generation and display control, and comprises the field of visual training of myopia prevention and control constructed on a naked-eye 3D display interface, a first visual training area and a second sensing and integrating area, in the first visual training area, through 3D interlaced pictures and videos, fine or dynamic stereoscopic vision is stimulated, front and back intersections of sight lines are adjusted, and split vision training is carried out; in the second perception and integration area, through perception, spatial positioning, binocular coordination and deep perception training, the spatial ability and response ability of eyes are stimulated; in combination with a three-dimensional display mechanism, displaying a naked eye three-dimensional image on a display, acquiring eyeball position information by adopting a human eye tracking technology, and dynamically adjusting a three-dimensional image display area; according to the method, the visual content structured organization and three-dimensional generation cooperative control is realized, and the three-dimensional display interaction matching precision is improved.
Owner:TIANJIN VISION TECHNOLOGY CO LTD

3D brain-click using binocular display

A method and system for detecting intentional selection of a user interface element using a binocular display. A first visual stimulus is presented stereoscopically to a user's eyes at a first virtual depth perceived by the user's depth perception and overlapping a first position within a field of view of the user. A second visual stimulus is presented stereoscopically to the user's eyes at a second virtual depth perceived by the user's depth perception and overlapping the first position. Neural signals are obtained from a neural signal capture device configured to detect neural activity of the user. In response to determining, based on the neural signals, that the user's eyes are focused on either the first visual stimulus or second visual stimulus, a computing system is placed into a first state or second state, respectively, associated with the first visual stimulus or second visual stimulus, respectively.
Owner:SNAP INC

Methods and systems for enhancing depth perception of a non-visible spectrum image of a scene

A method and system for providing depth perception to a two-dimensional (2D) representation of a given three-dimensional (3D) object within a 2D non-visible spectrum image of a scene is provided. The method comprises: capturing the 2D non-visible spectrum image at a capture time, by at least one non-visible spectrum sensor; obtaining 3D data regarding the given 3D object independently of the 2D non-visible spectrum image; generating one or more depth cues based on the 3D data; applying the depth cues on the 2D representation to generate a depth perception image that provides the depth perception to the 2D representation; and displaying the depth perception image.
Owner:ELBIT SYSTEMS LTD

Semantic segmentation method and device based on depth information position coding guidance

The application discloses a semantic segmentation method and device based on depth information position coding guidance, and the method comprises the following steps: acquiring an image to be processed; inputting the image to be processed into a pre-trained spatial depth perception auxiliary network to extract multi-scale features, and obtaining corresponding depth feature maps according to the multi-scale features; generating multi-scale feature vectors according to the multi-scale features and the corresponding depth feature maps; processing the multi-scale feature vectors by using a multi-scale attention mechanism and a feedforward neural network to obtain final features; performing fusion processing on the final features, and performing semantic segmentation on the processed fusion features to obtain a semantic segmentation map; in this way, the depth information is embedded in the features, the calculation cost is reduced, and the image segmentation performance is improved.
Owner:JIMEI UNIV

A dynamic surround view stitching method and system based on image overlap region feature perception

The application discloses a kind of dynamic ring vision splicing method and system based on image overlap area feature perception, method includes based on the depth information and the feature information dynamic planning splicing path;According to the splicing path of planning, image splicing is executed, including the multi-band image fusion based on depth perception.The present application extracts local and global features of the overlapping area and calculates pixel-level depth information, constructs a geometric transformation model optimized by feature-depth combination, so that panoramic stitching can accurately identify the outline of close-range objects, and dynamically plan the optimal stitching line to bypass prominent objects;By searching the distance splicing template based on the real distance and automatically unifying the focal length of each camera, the field of view range is kept consistent;Through the multi-scale depth estimation network guided by features and joint reprojection error optimization, high-precision registration can still be maintained in scenes with few feature points or low overlap areas.
Owner:NINGBO XINGBOYUAN INTELLIGENT TECHNOLOGY CO LTD

Methods and systems for determining depth perception profiles in virtual vision tests

A vision test can be performed based on real-time audio instructions in a virtual environment. An electronic device, such as a head-mounted display, can execute a visual assessment application, including generating a user interface corresponding to a three-dimensional virtual environment. The electronic device can display a plurality of visual stimuli in the user interface, and each visual stimulus can be displayed in duplication with respect to a respective target depth. The electronic device can receive one or more user responses, and each user response can indicate whether a user perceives a corresponding visual stimulus in duplication at the respective target depth. Based on the one or more user responses, the electronic device can determine a depth perception profile of the user, and the depth perception profile can include a plurality of depth perception levels corresponding to a plurality of target depths.
Owner:ZENNI OPTICAL

Binocular distance measurement method based on YOLOv8-BiFPN network model

The invention discloses a binocular ranging method based on a YOLOv8-BiFPN network model, belongs to the technical field of computer vision and target detection, and particularly relates to the binocular ranging method based on the YOLOv8-BiFPN network model. The objective of the invention is to solve the problems that in the prior art, a feature pyramid fixed weight fusion mechanism is difficult to adapt to multi-scale target detection, features are lost in a dynamic shielding scene, and algorithm calculation complexity and real-time requirements are contradictory, and especially to overcome the defect that high-precision depth perception and real-time detection in a binocular vision system are difficult to consider at the same time. The method comprises the steps of 1, obtaining a training set; 2, a YOLOv8-BiFPN network model is constructed; 3, a trained YOLOv8-BiFPN network model is obtained; and 4, calculating the distance of the target object based on the binocular camera and the trained YOLOv8-BiFPN network model.
Owner:SHENZHEN POLYTECHNIC

Intelligent equipment online language teaching translation system based on image recognition

The invention belongs to the technical field of artificial intelligence, and particularly relates to an intelligent equipment online language teaching translation system based on image recognition, which overcomes the defect of isolated recognition of characters or objects in images in the prior art, fuses visual information and accurate geographic positions, queries a scene expression rule set associated with the visual information and the accurate geographic positions, and translates the visual information and the accurate geographic positions. A scene translation data stream is generated and returned to the terminal equipment, deep perception and semantic understanding of the translation system on a real physical environment are achieved for the first time, personalized correction operation is introduced before driving display of the terminal equipment, and it is ensured that the final translation result is more accurate while the environmental context is considered. Compared with a current user individual, the method is most suitable and understandable, the interactive behaviors of the user are synchronously recorded so as to dynamically adjust the decision logic of the scene culture context, a system is helped to realize a self-evolving intelligent closed loop, and long-term and continuously evolved personalized learning experience is provided for the user.
Owner:HUNAN DIGITAL TECHNOLOGY CO LTD

Method for monitoring regional passenger flow density based on video image recognition

This invention discloses a method for monitoring regional passenger flow density based on video image recognition, belonging to the field of video passenger flow density monitoring technology. The method includes structuring a video stream to segment independent moving entities and calculating their trajectory overlap density distribution, identifying high-frequency interactive nodes and low-frequency silent regions. Based on this, a spatial pressure field model is constructed, and its gradient characteristics are used to dynamically correct the estimated entity motion velocity. The corrected velocity and density distribution are integrated to form a dynamic density field, and temporal slicing analysis is performed to extract field strength fluctuation patterns. Historical data of this pattern drives the adaptive updating of the warning threshold. This method achieves deep perception of dynamic passenger flow behavior, improving the accuracy of density monitoring and the system's adaptive warning capability.
Owner:ZHEJIANG KESHU STORE TECHNOLOGY CO LTD

Target depth estimation model training method and target depth estimation method

The disclosure provides a target depth estimation model training method and a target depth estimation method. The target depth estimation model training method comprises: obtaining a plurality of groups of training samples, each group of training samples comprising a sample image and an image label, the image label comprising a labeled depth of a target of interest in the sample image; obtaining depth prior information corresponding to the sample image, the depth prior information comprising a prior depth of a pixel in the sample image; and training a target depth prediction model using the sample image, the depth prior information and the image label. The target depth prediction model training method provided by the present scheme has more effective information of training samples than the model training method of the prior art, and thus the target depth prediction model obtained by training is more accurate and has more stable depth perception ability.
Owner:UISEE TECH BEIJING LTD

Medical system and method for setting a focus

PCT designated stageWO2026093300A1DiagnosticsMicroscopesOphthalmologyOptical axis
The present disclosure relates to a medical system (2) for imaging, having a digital microscope (4) having a focus (8) that can be adjusted along an optical axis (6) of the microscope (4), a camera (10) that is separate from the microscope (4) and designed to determine a distance (14) from the camera (10) to an object surface (16) by means of depth perception (12) and to provide it as depth information (18), and a control unit (20) that is designed to automatically set the focus (8) of the microscope (4) on the object surface (16) on the basis of the depth information (18), wherein the camera (10) has a greater field of view (S) and a greater depth of field (T) than the microscope (4). In addition, the present disclosure relates to a method for setting a focus (8) of a microscope (4) of a medical system (2) and to a use of a camera (10) with depth perception (12) of a medical system (2).
Owner:B BRAUN NEW VENTURES GMBH

Self-adaptive shadow generation method based on planar projection guidance and depth perception diffusion

The invention discloses an adaptive shadow generation method based on planar projection guidance and depth perception diffusion, and belongs to the technical field of computer vision and image synthesis. The method comprises the following steps: in a first stage, generating a hard shadow mask of a foreground object under a virtual light source through physical projection calculation, and providing geometric position and shape priori; in the second stage, multi-modal conditions such as a hard shadow mask, a background depth image and a foreground and background fusion image and noise latent variables are uniformly coded into a Token sequence, the Token sequence is input into a Diffusion Transform model, detail rendering is carried out through a self-attention mechanism, and a composite image with a realistic shadow is output. According to the method, the problems of shadow geometric distortion, inconsistent illumination, insufficient texture fitting and the like in the prior art are solved through a mixed frame combining physical guidance and neural rendering.
Owner:XIAMEN ZHENJING TECH CO LTD