Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

53 results about "3d perception" patented technology

4D multi-target sensing and tracking method and system based on sparse representation

PendingCN121191137ABiological modelsScene recognitionSparse methodsAlgorithm
The invention relates to the technical field of target detection of automatic driving, in particular to a 4D multi-target sensing and tracking method and system based on sparse representation, which is an efficient 3D target detection algorithm, and through dynamic interaction of sparse 4D query vectors and multi-view and multi-scale features and in combination with a time sequence instance denoising and depth sensing enhancement module, the accuracy of target detection is improved. And automatic driving 3D perception with high precision and low calculation amount is realized. Benefited from a sparse query architecture, the method also inherits the advantage of high efficiency of a sparse method while keeping high precision. Through a series of targeted optimization, the generalization ability and robustness of the model in various special scenes are significantly enhanced.
Owner:HONEYCOMB (WUHAN) MICROSYSTEM TECH CO LTD

Energy storage device damage diagnosis and interaction system based on penetration vision and large model

The invention provides an energy storage device damage diagnosis and interaction system based on penetration vision and a large model, and relates to the technical field of energy storage device damage diagnos.The energy storage device damage diagnosis and interaction system comprises a penetration type 3D sensing module, a periodic topology analysis module, a feature mapping and retrieval module and a large model reasoning and interaction module, and the energy storage device is scanned and reconstructed; the method comprises the following steps: constructing a time-space decoupling periodic Transform network, introducing a periodic mask matrix to force the network to pay attention to a repeatability rule of an internal structure of a battery, and identifying internal tiny deformation and structural damage by calculating topological consistency under the condition that a large amount of negative sample training is not needed; visual defect features are mapped into text embedding by utilizing a feature projection technology, a diagnostic report containing physical cause analysis and maintenance suggestions is generated by combining a retrieval enhancement generation technology and a large language model, and a user is supported to perform interactive questions and answers in a natural language. The method can solve the problems that in the prior art, three-dimensional deformation is difficult to quantify, small samples are difficult to train, and intelligent decision-making ability is lacked.
Owner:TIANFU YONGXING LAB

Image processing method and system based on four-camera cross-focal-length continuous zooming fusion

The invention relates to the technical field of multi-camera image processing and computational photography, and discloses an image processing method and system based on four-camera cross-focal-length continuous zooming fusion. The image processing method comprises the steps of system initialization, image acquisition, zoom routing, automatic ROI extraction and tracking, field-of-view cutting and geometric alignment, image fusion and binocular depth recognition. Through systematized multi-camera collaborative design, an innovative mechanism is introduced in key links such as zoom routing, geometric alignment, image fusion and depth recognition, smooth zoom, space consistency, detail fidelity, power consumption optimization and high-quality 3D perception are realized, the imaging quality is improved through the effects, the application scene is expanded, and the application prospect is wide. And a comprehensive solution is provided for mobile photography, AR and intelligent visual systems.
Owner:UNIV OF SCI & TECH OF CHINA

Scene deformation risk early warning method and device based on mobile robot

The invention relates to a scene deformation risk early warning method and device based on a mobile robot, and the method comprises the steps: obtaining perception information uploaded by at least one mobile robot, the perception information comprises real-time three-dimensional point cloud data collected by a 3D perception sensor carried by the mobile robot and a high-precision pose of the mobile robot calculated in real time based on the same frame of point cloud data; according to the pose of each mobile robot, converting the real-time three-dimensional point cloud data of at least one mobile robot into a unified global coordinate system, and performing splicing and fusion through a point cloud registration and fusion algorithm so as to incrementally construct and dynamically update a current three-dimensional morphology model; comparing the current three-dimensional shape model with a historical reference three-dimensional shape model, and calculating the deformation quantity of the target structure body in the operation scene; and when the deformation quantity exceeds a preset safety threshold value, risk early warning information is generated and output, so that the monitoring efficiency can be improved, and real-time early warning can be realized.
Owner:HANGZHOU LANXIN TECH CO LTD

Robot assembly space obstacle avoidance method and system applied to 3D perception

The invention provides a robot assembly space obstacle avoidance method and system applied to 3D perception, and relates to the technical field of robot automatic assemblation.The method comprises the steps that 3D assembly space data of a robot assembly area are obtained firstly, and the 3D assembly space data comprise assembly targets, static obstacles, dynamic obstacles and related information of a robot end effector; extracting spatial obstacle features from the data, establishing an initial obstacle association relationship, dynamically updating the initial association relationship, generating a real-time obstacle association relationship, and generating obstacle avoidance path constraints sorted according to an execution sequence according to the real-time association relationship and motion range information of an end effector; and finally, on the basis of the constraints and the three-dimensional contour information of the assembly target, generating a dynamic obstacle avoidance track of an end effector of the robot, including a continuous position coordinate sequence and a corresponding attitude angle sequence, so that efficient and safe obstacle avoidance of the robot is realized.
Owner:CHENGDU TIM WALKER TECH CO LTD

Probabilistic state simulation for end-to-end drive stack learning for autonomous and semi-autonomous machines and applications

In various examples, perception encoder uses one or more neural networks implemented using a transformer architecture, a sensor perspective encoding, a planned navigation route, and / or detected ego-motion to extract a scene embedding representing one or more aspects of an observed scene, such as visual information, motion information, ego-state of an ego-machine, a planned navigation route, and / or other types of information. The perception encoder may be used in a probabilistic state simulation stack, and / or may be used to extract and apply a scene embedding as an input for 3D perception or reconstruction tasks such as object detection and classification, semantic segmentation, depth map extraction, trajectory prediction, path planning, navigation control, and / or localization or mapping, to name a few example tasks.
Owner:NVIDIA CORP

Automatic driving perception method and device based on cross-modal distillation

The invention relates to an automatic driving perception method and device based on cross-modal distillation, and the method comprises the steps: enabling a point cloud teacher model to guide an image student model through feature extraction, dynamic attention weight calculation, deep confidence filtering and feature alignment, dual-view comparison distillation, total loss optimization, model training and deployment, and the like. And in combination with a dynamic view attention mechanism and deep confidence filtering, efficient migration of 3D knowledge to a 2D model is realized, and finally, a pure image model has a 3D perception capability close to a point cloud model. Therefore, the problems that in the prior art, a pure vision vehicle type is difficult to meet the high-precision perception requirement under the condition of no laser radar configuration, meanwhile, a fixed weight cannot adjust the learning key point according to the scene complexity, and the mapping noise interference of the same image pixel is large are solved.
Owner:CHERY AUTOMOBILE CO LTD

Homographic deformation CNN for robust 3D perception

A computer-implemented method and system relate to an image encoder that receives a digital image as input. The image encoder generates a weight map using the prior feature map. A prior feature map is generated using pixels of the digital image. A weight map is generated based on Lie data associated with the digital image. The homography is interpolated between the two planar projections of the digital image using at least the weight map and the homography matrix. The homography matrix provides a mapping between two planar projections of the digital image. A homography kernel is generated by applying homography to the convolution kernel. A homographic transform kernel is applied to the prior feature map to convolve different planar regions appearing in the digital image, and a new feature map is generated for computer vision tasks involving three-dimensional (3D) perception.
Owner:ROBERT BOSCH GMBH

Inline blade wear estimation based on processed soil surface

A method for deriving a wear state of a soil interaction component of a soil processing implement of a construction vehicle. The method comprises steps of 1.) providing a geometry model regarding an assumed shape of the soil interaction component, 2.) engaging the soil by using the soil interaction component and tracking a motion for deriving tracking data of the soil interaction component, 3.) using the tracking data and the geometry model to derive an expected 3D surface model, 4.) providing visual 3D perception data of a soil area affected by the engaging, such that the visual 3D perception data and the expected 3D surface model can be referenced to one another, and 5.) comparing the visual 3D perception data with the expected 3D surface model and, based thereof, determining a deviation of an effective shape of the soil interaction component from the assumed shape.
Owner:LEICA GEOSYST TECH

A Robot Variable Stiffness Surface Path Planning Method Based on 3D Perception and Its Application

This invention discloses a robot path planning method and application based on three-dimensional perception for variable stiffness surfaces, comprising the following steps: processing workpiece surface point cloud data; establishing a surface defect detection model to compensate for deformation deviations caused by workpiece surface defects in the point cloud data; performing path planning analysis on the point cloud model to generate an adaptive tool path. During the robot's execution of the planned path, passive analysis is performed using a variable stiffness impedance control law under energy constraints, and force feedback trajectories are collected as a reference. Combined with a quadratic programming optimization controller, force-controlled surface processing is achieved, completing the robot's planned path execution task. This invention also incorporates the planned path model to design an AI-assisted programming framework, guiding users to parameterize tasks and tools through a natural language interface and knowledge representation engine. This invention not only enhances the robot's adaptability to complex variable stiffness surfaces but also significantly improves processing consistency and efficiency, demonstrating broad application prospects.
Owner:HUNAN UNIV

A laser SLAM inspection trolley and method suitable for complex underground infrastructure environments

PendingCN122083921AEffectively deal with occlusionExcellent navigation robustnessNavigational calculation instrumentsNavigation by speed/acceleration measurementsOperational systemOdometer
This invention provides a laser SLAM inspection vehicle and method suitable for complex underground infrastructure environments. A multi-line lidar system is equipped with multiple laser emission and reception channels to generate high-density, large-area 3D point cloud data. A high-precision inertial measurement unit (IMU) provides the system with high-frequency attitude, angular velocity, and acceleration information. A wheeled odometer collects wheel speed data to provide short-term pose increment information, effectively supplementing the lidar and IMU. An edge computing device performs real-time data processing using SLAM algorithms. An embedded motion control unit runs an operating system to precisely control the chassis motors and transmit data with various subsystems. The software framework is built on the operating system and includes distributed nodes, a publish / subscribe communication model, and an open-source toolkit ecosystem. This achieves lightweight 3D perception capabilities, strong anti-interference performance, and supports fully automatic online repositioning, making it suitable for use in underground infrastructure such as subway tunnels, urban utility tunnels, and large warehousing centers.
Owner:CHONGQING UNIV

Bidirectional tracking truth value generation method and system based on MLLM and 3D perception

The invention discloses a bidirectional tracking truth value generation method and system based on MLLM and 3D perception, and the method comprises the steps: S1, inputting image data and 3D laser point cloud data which are synchronous in time sequence, processing each frame of point cloud through a 3D target detector, obtaining a 3D detection frame, and then projecting the 3D detection frame to a synchronous image, and obtaining a corresponding 2D BBOX; s2, performing forward tracking and reverse tracking on the 3D detection frame of each frame to generate a forward track and a reverse track, then performing track merging on the forward track and the reverse track, and identifying a suspicious track in the merging process; s3, for the suspicious trajectory, extracting corresponding images and 2D BBOX, constructing textualized and structured semantic descriptions for each group of images and 2D BBOX data, inputting the textualized and structured semantic descriptions to a pre-trained MLLM, outputting semantic association scores and an inference chain by the MLLM, and performing association decision on the suspicious trajectory based on the semantic association scores; and S4, outputting time sequence truth value data including the complete object ID, the 3D detection frame of each frame and the object category.
Owner:ZHUHAI KUWA TECHNOLOGY CO LTD +2

Homographically deformed CNN for robust 3D perception

A computer-implemented method and system refer to an image encoder that receives a digital image as input. The image encoder generates a weighting map using a preceding feature map. The preceding feature map is generated using pixels of the digital image. The weighting map is generated based on Lie data associated with the digital image. A homographic transformation is interpolated between two planar projections of the digital image using at least the weighting map and a homography matrix. The homography matrix provides a mapping between the two planar projections of the digital image. Homographically transformed kernels are generated by applying the homographic transformation to convolution kernels.The homographically transformed kernels are applied to the preceding feature map to perform folding on different planar areas appearing in the digital image and to generate a new feature map that is used for a computer vision task involving three-dimensional (3D) perception.
Owner:ROBERT BOSCH GMBH

Multi-view 3D perception method based on space-time modeling and context enhancement

The invention discloses a multi-view 3D sensing method based on space-time modeling and context enhancement, and relates to the technical field of automatic driving and computer vision. Comprising the following steps: acquiring a nuScenes data set, and dividing the data set into a training set and a test set according to a certain proportion; constructing a multi-view 3D perception model based on space-time modeling and context enhancement; training the model through the training set to obtain a trained multi-view 3D perception model; verifying the trained model through the test set to obtain a verified multi-view 3D perception model; and inputting a 3D perception scene to be detected into the verified multi-view 3D perception model, and outputting a 3D target detection result and a 3D occupation prediction result. According to the method, the efficient long-sequence modeling capability of the state space model and the multi-view attention context modeling capability are combined, global space-time relation capture and accurate feature improvement are both achieved, and dual optimization of performance and efficiency is achieved in a 3D perception task.
Owner:GUANGDONG TOBACCO HEYUAN CITY CO LTD

Multi-line laser three-dimensional reconstruction method based on trinocular geometric constraint

The invention provides a multi-line laser three-dimensional reconstruction method based on trinocular geometric constraint, and relates to the technical field of three-dimensional perception. Comprising the following steps: synchronously projecting specific LaS stripes to the surface of an object to be measured by using three target industrial cameras, and imaging to obtain a three-view image; preprocessing the three-view image and extracting a center line to determine an initial matching triple; performing matching candidate construction of geometric constraint on the initial matching triad to obtain a rough matching triad set; performing cross reprojection error verification and optimal matching determination on the rough matching triple set to obtain an optimal matching triple; an error function corresponding to the optimal matching triple is determined and minimized, and an optimized triple is obtained; and estimating a rigid body transformation matrix of each view corresponding to the optimized triple based on a preset circular mark point so as to realize global registration. The problems that in the prior art, multi-line laser needs to be calibrated tediously, binocular matching is not stable, and reconstruction precision and density are insufficient are solved.
Owner:SOUTHWEAT UNIV OF SCI & TECH

Intelligent glasses, control method and system

The invention provides intelligent glasses and a control method and system. The intelligent glasses at least comprise a glasses main body, an optical machine, a first camera, a second camera and a controller, the glasses main body consists of a left glasses leg, a right glasses leg and a glasses frame provided with display lenses. Acquiring a first image acquired by the first camera and a second image acquired by the second camera; calculating the distance between the target object and the intelligent glasses by using the first image and the second image to obtain a target image distance; according to the target image distance, the ray machine is controlled to project multimedia content related to the target object into the display lens, and the projection image distance of the multimedia content is equal to the target image distance. According to the scheme, the layout of the optical machine and the camera is optimized on the premise of ensuring the 3D perception precision, so that the intelligent glasses can accommodate more functional components, the functional diversity of the intelligent glasses is improved, and the user experience is improved by realizing adaptive projection of multimedia contents with different image distances.
Owner:SHENZHEN GUANGZHI TECHNOLOGY CO LTD

A three-dimensional target detection method based on sparse dynamic attention and star interaction

This invention discloses a 3D target detection method based on sparse dynamic attention and star-shaped interaction. The method obtains basic voxel features from the original LiDAR point cloud through voxelization and sparse convolution, then introduces a sparse dynamic parallel attention module. This module achieves efficient enhancement of global context and channel dimensions through dynamic attention branches and parallel channel interaction branches. A sparse star-shaped interaction module is then used to construct a star-shaped neighborhood interaction structure with a central voxel, completing local geometric modeling and nonlinear feature interaction only on non-empty voxels. Finally, keypoint sampling, RoI pooling, and a detection head output the 3D detection box, category, and confidence score. This invention, through the synergistic complementarity of SDPA and SSB, significantly improves the detection accuracy of long-distance, small-scale, and occluded targets while maintaining linear growth in computational complexity and meeting real-time requirements. It achieves balanced performance optimization across multiple categories, including vehicles, pedestrians, and cyclists, and is suitable for 3D perception scenarios with high precision and real-time requirements, such as autonomous driving.
Owner:WUXI UNIV

A method for underwater depth estimation based on cross-modal alignment and fusion of visual and sonar images

PendingCN122089803AEliminate spatial mapping biasHigh-precision depth acquisitionImage analysisCharacter and pattern recognitionPattern recognitionRgb image
This invention discloses an underwater depth estimation method based on the fusion of visual and sonar dual-modal images, belonging to the fields of underwater robotics, deep learning, and 3D perception fusion technology. Step 1: Convert the original 2D sonar image into an STR image with the same resolution and spatial alignment as the RGB image; Step 2: Construct a dual-branch encoder consisting of RGB and STR branches; Step 3: Construct a multi-scale feature reconstruction decoder to recursively decode the multi-scale fused features output in Step 2; Step 4: Construct a global context modeling module based on a lightweight visual Transformer to model long-range dependencies of the decoded high-dimensional features, and output the final underwater absolute depth map using a depth regression head. This invention addresses the technical challenges of scale ambiguity in existing underwater depth estimation methods using monocular vision, and the noise introduced by modal misalignment due to differences in imaging mechanisms during sonar-visual fusion, which limits fusion efficiency.
Owner:HARBIN INST OF TECH

Multi-unmanned aerial vehicle cooperative 3D sensing method based on channel adaptation

A multi-unmanned aerial vehicle cooperative 3D sensing method based on channel adaptation relates to the field of communication, and comprises establishing a multi-unmanned aerial vehicle cooperative 3D sensing framework based on channel adaptation, a BEV semantic coding module and a BEV semantic decoding module. And the unmanned aerial vehicle encodes and fuses the acquired image into semantic features and shares the semantic features to other unmanned aerial vehicles through a complex channel environment, and receives the decoded semantic features of the unmanned aerial vehicle to enhance the 3D perception ability of the unmanned aerial vehicle. The BEV semantic coding module is used for adjusting and adapting a process of fusing self multi-view features into BEV features according to a channel state and compressing and coding the BEV features into semantic features by the unmanned aerial vehicle at the transmitting end; and the BEV semantic decoding module is used for the receiving end unmanned aerial vehicle to decode the semantic features according to the channel state and fuse a plurality of unmanned aerial vehicle features into a perceptible feature vector. According to the method, the interference of channel noise on cooperative 3D sensing is reduced, and the cooperative 3D sensing performance of multiple unmanned aerial vehicles in a complex channel environment is improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Variable-focus binocular ranging method and system fusing instance segmentation and self-calibration

The invention belongs to the technical field of computer vision and three-dimensional perception, and particularly relates to a variable-focus binocular ranging method and system fusing instance segmentation and self-calibration, and the method comprises the specific steps: carrying out the self-calibration through employing an image collected by a binocular camera through a self-calibration method based on deep learning, and obtaining the internal and external parameters of the camera; acquiring a mask area, a category label and confidence of each detection instance in the image; the method comprises the following steps: preprocessing left and right images collected by a binocular camera, inputting the preprocessed left and right images into a stereo matching network to obtain a dense disparity map, and adjusting the dense disparity map to the size of an original image; calculating a dense depth map by using the focal length obtained by self-calibration, and converting the dense depth map into a pseudo-color depth map; extracting a depth value corresponding to a mask area pixel from the dense depth map, taking a depth median of the area as a distance estimation value of the instance, and superposing the distance estimation value in an original image collected by a camera to obtain an instance depth superposition map; and respectively converting the pseudo-color depth map and the instance depth overlay map into image messages and publishing the image messages.
Owner:BEIJING INST OF TECH

3D perception navigation module

1. Name of the product in this design: 3D Perception Navigation Module. 2. Purpose of this design: To provide robots and drones with critical spatial perception and spatial memory capabilities, and to optimize, label and reconstruct the collected 3D data to achieve a complete closed loop from front-end perception to back-end training. 3. The key design feature of this product is its shape. 4. The image or photograph that best illustrates the design's key points: a 3D model. 5. Other situations requiring explanation: The sensing lens in the stereoscopic view is made of transparent material.
Owner:SHENZHEN LIUXING TECHNOLOGY LTD

Bi-directional feature projection for 3D perception systems and applications

In various examples, bi-directional projection techniques may be used to generate enhanced Bird's-Eye View (BEV) representations. For example, a system(s) may generate one or more BEV features associated with a BEV of an environment using a projection process that associates 2D image features to one or more first locations of a 3D space. At least partially using the BEV feature(s), the system(s) may determine one or more second locations of the 3D space that correspond to one or more regions of interest in the environment. The system(s) may then generate one or more additional BEV features corresponding to the second location(s) using a different projection process that associates the second location(s) from the 3D space to at least a portion of the 2D image features. The system(s) may then generate an updated BEV of the environment based at least on the BEV feature(s) and / or the additional BEV feature(s).
Owner:NVIDIA CORP

Lightweight point cloud instance segmentation method based on knowledge distillation

The invention discloses a lightweight point cloud instance segmentation method based on knowledge distillation, and the method comprises the steps: introducing a knowledge distillation technology to achieve the lightweight of a point cloud instance segmentation model, enabling the model to serve as an intermediate technology deployed in a whole 3D perception and decision-making system, and improving the decision-making performance of a whole artificial intelligence system. By improving a distillation strategy, a matching algorithm between prediction examples is designed, and the prediction example of the teacher model and the prediction example of the student model are effectively aligned in two dimensions of semantics and space; two distillation losses are provided, instance-level point features can be decoupled, and a student model is guided to learn richer spatial expression ability; in addition, a weight map of an instance level is designed to weight distillation loss items, and key information of a point cloud space is intensified. The design enables the distillation process to better conform to the structural characteristics of the instance segmentation task, and breaks through the limitation of traditional scene-level supervision.
Owner:SOUTH CHINA UNIV OF TECH

Three-view multi-resolution three-dimensional perception method based on distance perception

The invention discloses a three-view multi-resolution three-dimensional sensing method based on distance sensing, and relates to the field of automatic driving three-dimensional environment sensing. According to the method, a near-field high-resolution and far-field low-resolution double-branch architecture is constructed, and cross-resolution initialization and a multi-scale feature fusion mechanism are combined, so that high-precision perception of a short-distance target and efficient modeling of a long-distance target are realized. In the reasoning stage, a spatial adaptive decoding strategy is adopted, high-resolution and low-resolution features are fused in a short-distance region, and low-resolution features are directly adopted in a long-distance region, so that resolution adaptive three-dimensional perception is realized. Compared with a traditional single high-resolution method, the method has the advantages that the perception precision is kept, meanwhile, the video memory occupation and the calculation overhead are remarkably reduced, the memory consumption of an experimental result table is remarkably reduced by about 26.3%, the reasoning speed is increased to the quasi-real-time level, and the method has the application value of actual deployment on a vehicle-mounted platform.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Text-to-point cloud cross-modal positioning method and system based on visual language model

This invention belongs to the field of computer vision technology. It proposes a text-to-point cloud cross-modal localization method and system based on a visual language model. The method converts 3D point clouds into a bird's-eye view image and a structured scene graph, the latter containing nodes with semantic and spatial location information. Natural language descriptions, the bird's-eye view, and the scene graph are input into a pre-trained visual-language model to generate cross-modal alignment features. Through a partial node association mechanism, the semantic information is selectively mapped to spatially matching nodes to determine the target node. Then, based on its spatial location, the original point cloud is retrieved to achieve accurate localization. This invention utilizes dual structured representation and explicit semantic-spatial alignment to effectively bridge the semantic gap between point clouds and language, improving the accuracy, robustness, and interpretability of localization. It is applicable to scenarios requiring language-guided 3D perception, such as autonomous driving and robotics.
Owner:NANKAI UNIV

Dual-camera weak texture multi-target space matching and positioning method and application system

PendingCN122089838AStrong target discrimination abilityReduce match ambiguityImage analysisCharacter and pattern recognitionPattern recognitionComputer graphics (images)
The invention provides a dual-camera weak texture multi-target space matching and positioning method and an application system, and belongs to the technical field of computer vision and three-dimensional perception. According to the method, based on a dual-camera epipolar geometric model, on the basis of obtaining a dual-camera synchronous image and performing multi-target detection, four corner points and a central point of a target outer surrounding frame are introduced to construct a multi-key-point combined epipolar constraint, and target space structure information is fully utilized. The multi-key-point combined epipolar constraint is constructed by selecting the four corner points and the center point of the outer surrounding frame of the target, the overall space structure information of the target is introduced into the dual-camera matching process, a matching constraint mechanism of multi-geometric clue fusion is formed, and compared with an epipolar constraint mode only based on a single image point, the matching constraint mechanism is more accurate. The matching ambiguity caused by appearance similarity or deformation is effectively reduced, so that the matching accuracy in a multi-target scene and the anti-interference capability of the system are improved, and the method is particularly suitable for a weak texture environment.
Owner:HANGZHOU BINGBAI INTELLIGENT TECHNOLOGY CO LTD

Target detection and semantic segmentation method, device and equipment, and storage medium

The application relates to the field of image processing and discloses a target detection and semantic segmentation method, device and equipment and a storage medium, the method comprising the following steps: acquiring a same frame monocular image photographed by multiple cameras, and performing depth labeling to obtain a depth map corresponding to the monocular image; converting the image coordinates of each pixel in the depth map to a camera coordinate system according to the camera internal parameter, to obtain pseudo point clouds of each pixel in the depth map under the camera coordinate system; converting the pseudo point clouds and corresponding image features to a preset bird's-eye view coordinate system to obtain corresponding bird's-eye view point clouds and bird's-eye view features; and performing target detection and semantic segmentation under the bird's-eye view perspective based on the bird's-eye view point clouds and the corresponding bird's-eye view features. The method can break the strong dependence on ranging sensors under the premise of ensuring safety, reduces the hardware cost, effectively utilizes the results of 2D image perception tasks in 3D perception tasks, extracts image information and transforms the image information to a 3D space, and improves the performance of a perception algorithm.
Owner:GUANGZHOU WERIDE TECH LTD CO

Parking lot license plate authentic identification method and system based on monocular 3D perception

The invention discloses a parking lot license plate authentic identification method and system based on monocular 3D perception, and the method comprises the steps: collecting an image of a driving-in vehicle based on a monocular 3D camera, carrying out the target detection and depth estimation, recognizing and positioning a license plate region Rplate and a vehicle head body region Rbody, and generating a depth map of a scene; the average or median depth Dplate of the Rplate and the average or median depth Dbody of the Rbody are extracted from the depth map; the absolute value Ddelta of the difference value between the Dplate and the Rbody is compared with a preset depth threshold value T; if Ddelta is less than T, determining that the license plate is a normal license plate; and if Ddelta is greater than or equal to T, determining that the license plate is an abnormal license plate. According to the invention, verification of physical authenticity or spatial consistency of the license plate is added while LPR identification is carried out, so that a forged license plate condition is identified, and wrong release is completely eradicated.
Owner:BEIJING SIGNALWAY TECH

System and method for the 3D thermal imaging capturing and visualization

ActiveUS12670653B2Stereoscopic videoLine sensor
A system and method for navigation in complete darkness with thermal imaging and virtual-reality headset provide adjustable-base stereopsis and maintain long-range situational awareness. Such a system includes a multiple-aperture thermal imaging subsystem with non-collinear sensors with parallel optical axes and an image processing device, resulting in a significantly more accurate depth map than perceived with a normal human stereo acuity from a pair of raw thermal images. The stereo-video presented to the user is synthetic, allowing vantage point and stereo-base adjustment; it augments natural objects' texture with the generated one to allow a 3D perception of the negative obstacles and other horizontal features that do not provide stereo cues for the horizontal binocular vision. Additional wide-field-of-view thermal sensors may be used to compare current real-world views with the predicted from the earlier captured 3D data to communicate results to the user and supplement the 3D model with the structure-from-motion algorithm.
Owner:ELPHEL INC

Small baseline light field depth estimation method based on observability driving

The invention discloses a small baseline light field depth estimation method based on observability driving. The method comprises the following steps: firstly, constructing a large / medium / small three-gear equivalent baseline set on the same light field data according to an angular step length; calculating observability scores related to imaging geometry in each file, selecting sub-view angles according to the observability scores, and constructing a weighted matching body; taking the large base line as a teacher, taking the medium / small base line as a student, and adopting curriculum type distillation training to realize migration of stable geometric prior from the large base line to the small base line; in the inference stage, small baseline matching bodies are input into a trained student model, body aggregation, probabilization and expectation regression are completed to obtain parallax, and parallax-depth mapping is performed according to baselines and focal lengths. According to the method, on the premise that external data is not introduced, the precision, boundary fidelity and robustness under the small baseline condition are improved, the view angle and the calculation overhead are controllable, and the method is suitable for three-dimensional perception application such as a light field camera, a multi-camera array and AR / VR.
Owner:SOUTH CHINA UNIV OF TECH