Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

271 results about "Monocular image" patented technology

Posture recognition algorithm for any object under monocular camera and application system

The invention provides a posture recognition algorithm for any object under a monocular camera and an application system, and the algorithm comprises the steps: S1, constructing a target three-dimensional model, carrying out the multi-view annular shooting image collection of a target, and generating a dense grid model through feature extraction, matching, posture calculation and a multi-view geometric method; s2, generating an image depth map, and predicting depth information of a target in a motion process based on a monocular image sequence; s3, extracting a target image mask, and generating a target area mask graph through an image encoder, a prompt encoder and a mask decoder; and S4, executing attitude estimation, performing attitude initialization, correction and screening by combining the three-dimensional model, the depth map and the mask map, and outputting a six-degree-of-freedom attitude result of the target. According to the method, the target is subjected to annular shooting modeling through the method based on multi-view geometry, the three-dimensional model of the target is generated, attitude estimation is achieved in combination with the image mask and the depth map, the generalization ability of an attitude estimation algorithm in an actual scene is improved, and the application range of the attitude estimation algorithm in the actual scene is widened.
Owner:HANGZHOU BINGBAI INTELLIGENT TECHNOLOGY CO LTD

Three-dimensional scene reconstruction method and device based on large model geometric prior, and medium

The invention discloses a three-dimensional scene reconstruction method and device based on large model geometric prior, and a medium, and aims to solve the problems that a conventional 3DGS is liable to have artifacts and detail loss in geometric discontinuity, data redundancy and illumination variation scenes, and predicts a dense depth map and a normal map from a monocular image by using a pre-trained large model. The position and form of the Gaussian kernel are constrained as additional geometric priori; a primitive adjustment strategy based on kernel density estimation is introduced in the training stage, small Gaussian primitives with similar structures and adjacent spaces are combined into a large Gaussian primitive, the rendering quality is kept, redundancy is reduced, and the volume of the model is reduced; an exposure coefficient is adaptively estimated for each input image, an exposure compensation image loss function is constructed, and floating artifacts caused by illumination differences at shooting moments are eliminated. Experiments show that compared with the prior art, the method improves the three-dimensional reconstruction precision and real-time rendering quality of complex illumination and less-texture areas in a public data set and an unmanned aerial vehicle aerial photography scene.
Owner:NARI INFORMATION & COMM TECH

Depth map generation method and device based on large model, three-dimensional reconstruction method and device, electronic equipment and storage medium

The invention provides a depth map generation method and device based on a large model, a three-dimensional reconstruction method and device, electronic equipment and a storage medium, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, can be applied to real-time road scene depth perception, environment three-dimensional reconstruction and obstacle avoidance, and can be applied to real-time road scene depth perception. And virtual and real scene fusion and other scenes can be realized. The specific implementation scheme is as follows: performing visual coding on a monocular image to obtain a coded image; inputting the coded image and the target text into a pre-trained large language model for fusion to obtain fusion features; generating global guide features based on the fusion features, wherein the global guide features comprise joint semantic information of visual features and text features; adding noise to the color image of the monocular image to obtain a noise feature sequence; de-noising the noise feature sequence under the condition of the global guide feature, and generating an implicit feature matched with the joint semantic information; a depth map is generated based on the implicit features.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Single-image-based three-dimensional face reconstruction and editing method

The present invention provides a single-image-based three-dimensional face reconstruction and editing method, comprising: extracting features by means of a pre-trained EG3D network, so as to generate a preliminary three-dimensional face; using an inversion module to map a monocular image to a latent code space, and optimizing a latent code to generate a realistic three-dimensional face image; making use of a depth estimation technology and multi-view projection to create a pseudo multi-view image; then, by means of semantic segmentation and expression feature extraction, fusing the features to generate a comprehensive representation; according to editing requirements, optimizing and adjusting the latent code, so as to implement customized editing; and finally, a generator outputting an edited face image. The present invention can enable more efficient, flexible and high-quality three-dimensional face reconstruction and editing, thereby effectively overcoming the defects of high cost, low universality and lack of customized editing in existing face reconstruction technologies.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Method and electronic device for 3D object detection using neural networks

A method of 3D object detection using an object detection neural network includes: receiving one or more monocular images; extracting 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through a 2D feature extracting part of the object detection neural network, generating an averaged 3D voxel volume based on the 2D feature maps, extracting a 2D representation of 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through an encoder of a 3D feature extracting part of the object detection neural network, and performing 3D object detection as 2D object detection in a Bird's Eye View (BEV) plane, the 2D object detection in the BEV plane being performed by passing the 2D representation of 3D feature maps through thane outdoor object detecting part of the object detection neural network.
Owner:SAMSUNG ELECTRONICS CO LTD

Shielding area three-dimensional voxel reasoning method and system

The invention belongs to the technical field of three-dimensional mapping, and particularly relates to an occlusion area three-dimensional voxel reasoning method and system. The method comprises the following steps: S10, acquiring an RGB image containing a semantic segmentation result of an occlusion region and a corresponding depth map, and converting the depth map into an initial sparse three-dimensional voxel; and S20, constructing a three-dimensional voxel inference network, and inferring a three-dimensional voxel occupancy probability graph from the initial sparse three-dimensional voxels through the three-dimensional voxel inference network. The occlusion area can be reconstructed based on the monocular image; multi-format output and navigation interface butt joint are supported, and the method can be integrated to SLAM, 3D mapping or path search systems to serve as spatial feasibility constraints; the method can be used for querying whether any spatial position is passable; the method can be used for extracting a trafficability sub-graph from a designated area for path planning.
Owner:SHANGHAI UNIV

Three-dimensional perception robot operation knowledge distillation method based on monocular image

The invention relates to the field of robot operation and three-dimensional perception, in particular to a monocular image-based three-dimensional perception robot operation knowledge distillation method, which comprises the following steps of: establishing a strategy learning framework comprising a student model and a teacher model; constructing a strategy prediction model in the strategy learning framework; training the strategy prediction model, wherein a three-level knowledge distillation mechanism is adopted in the training process to complete knowledge migration between the teacher model and the student model; combining the three distillation losses with strategy optimization losses to form a total loss function, and performing end-to-end training on the student model until convergence to obtain a deployable monocular strategy model; the deployable monocular strategy model only retains a student model, and generates a robot operation instruction. The method has the beneficial effects that the robot under monocular RGB input has three-dimensional perception and high-precision operation capabilities while the reasoning efficiency is kept, the task success rate and generalization performance are remarkably improved, and the effectiveness and robustness of the method are verified.
Owner:ZHEJIANG UNIV OF TECH

Scale-aware self-supervised monocular depth with sparse radar supervision

Systems and methods are provided for training a depth model to recover scale factor for self-supervised depth estimation in monocular images. Examples include deriving a depth map for an image based on a depth model. The depth map comprises depth values for pixels of the image. A first scale for the image can be estimated based on the depth values, and depth data captured by a range sensor can be received. The depth data comprises a point cloud comprising depth measures. A second scale for the point cloud can be determined based on the depth measures and a scale factor can be determined based the second scale and the first scale. The depth model can be updated based on the scale factor, wherein the depth model generates metrically accurate depth estimates based on the scale factor.
Owner:TOYOTA JIDOSHA KK

High-precision map reconstruction method and system based on monocular vision, medium and equipment

The invention belongs to the technical field of robot positioning and three-dimensional mapping, and discloses a high-precision map reconstruction method and system based on monocular vision, a medium and equipment. Scene image data are acquired through a monocular image acquisition module, after feature extraction and matching are performed on each frame of image, attitude information acquired by an inertial measurement unit is fused, and pose calculation of a robot is completed through a sparse vision SLAM system; meanwhile, an image dense depth map is generated by a monocular depth estimation model based on an attention mechanism, and the image dense depth map is converted into a single-frame color dense point cloud in combination with image RGB color information. According to the invention, a loose coupling fusion strategy is adopted to carry out spatial registration on a single-frame colored dense point cloud and a robot pose at a corresponding moment, multi-frame data fusion is completed through point cloud splicing and optimization, and a globally consistent three-dimensional dense point cloud map is constructed to realize scene modeling. The method has the characteristics of high robustness and high reconstruction precision, and can be effectively applied to three-dimensional map construction in an outdoor complex environment.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Semantic scene completion method and device, electronic equipment and readable storage medium

The invention discloses a semantic scene completion method and device, electronic equipment and a readable storage medium. The method comprises the following steps: acquiring a binocular image; performing depth estimation processing on the binocular image to obtain a multi-scale image feature map, a disparity probability distribution map and a disparity map of the monocular image; performing conversion processing from a parallax dimension to a depth dimension on the parallax image of the monocular image to obtain a depth image of the monocular image; processing the multi-scale feature map and the depth map of the monocular image to obtain a first semantic feature map of the monocular image; performing depth feature extraction processing on the parallax probability distribution diagram of the monocular image to obtain a depth probability distribution diagram of the monocular image; and determining voxel occupation data and voxel semantic data of the semantic occupation data based on the depth map of the monocular image, the first semantic feature map and the depth probability distribution map so as to perform semantic scene completion. According to the invention, semantic scene completion can be carried out based on visual image data.
Owner:SHENZHEN SWEET POTATO ROBOT CO LTD

Methods, systems, and computer program products for generating 3D human pose and movement estimation from monocular image information

A computer-implemented method includes converting by a pose tokenizer, based on a learned codebook, pose parameters of a body into a sequence of discrete pose tokens; randomly masking a portion of the sequence of discrete pose tokens; predicting the randomly masked sequence of discrete pose tokens based on multi-scale features extracted from a monocular image by an image conditioned masked transformer; optimizing the sequence of discrete pose tokens by aligning a re-projected three-dimensional (3D) pose with an estimated two-dimensional (2D) pose; directly regressing, from the multi-scale features, a shape parameter of the body and a weak perspective camera parameter; and generating a 3D mesh reconstruction of the body based on the shape parameter and the weak perspective camera parameter.
Owner:THE UNIV OF NORTH CAROLINA AT CHAPEL HILL

Endoscope monocular image depth estimation method and system

The invention relates to an endoscope monocular image depth estimation method and system, and the method comprises the steps: obtaining the front and rear frame images of an endoscope monocular; an endoscope monocular is trained by using a DA-I CGA model based on a DARES, the DA-I CGA model comprises a PoseNet module and a DAM-LoRA module, and a double-branch geometric perception module and an image-level contrast learning module are creatively introduced; and performing depth estimation on the to-be-detected endoscope monocular image through the trained DA-I CGA model. According to the method, artifacts generated in an overexposure region can be effectively solved, and the monocular depth estimation precision is further improved.
Owner:JIANGNAN UNIV

Physical driving measurement method for monocular three-dimensional dynamic displacement of rotary machinery

The invention provides a physical driving measurement method for monocular three-dimensional dynamic displacement of a rotary machine, and belongs to the technical field of crossing of computer vision and industrial state monitoring. According to the method, a high-speed dynamic visual acquisition system is constructed, and a time sequence video stream is obtained by using a cooperative marker; constructing a'Gaussian + motion blur 'composite gradient model taking the motion blur width as an endogenous variable, and jointly resolving sub-pixel edge coordinates and ambiguity by adopting a nonlinear optimization algorithm; decoupling the pixel displacement of the radial X axis, the radial Y axis and the axial Z axis by using the width change and the edge displacement of the marker in combination with the geometric principles of sequential robust filtering and monocular imaging; and finally, outputting a three-way physical vibration waveform by using the calibration conversion factor. According to the method, the fuzzy influence is adaptively eliminated from a physical imaging mechanism, micron-level precision three-direction vibration synchronous measurement is realized only by a single camera, and the hardware cost and the deployment difficulty are reduced.
Owner:OCEAN UNIV OF CHINA

Monocular image depth estimation method based on multi-scale information fusion

The invention discloses a monocular image depth estimation method based on multi-scale information fusion, and the method comprises the steps: outputting a first feature map representing local information through a CNN branch, and outputting each second feature map representing global feature correlation at a Transform branch; enhancing the input first feature map and each second feature map from three dimensions of channel, space and cross-scale by using an adaptive fine-grained channel space collaborative gating module to obtain each third feature map; using attention mechanism decoders constructed based on an up-sampling module to sequentially process the input third feature maps to obtain a monocular image depth estimation feature map; according to the invention, the problem that the processing capability of a model on a remote area in a monocular depth estimation scene cannot be improved because local features and global features cannot be effectively combined in the prior art can be solved.
Owner:YUNNAN MINZU UNIV

Monocular self-supervision depth estimation method based on credible area depth controllable generation

The invention relates to the technical field of monocular depth estimation, and discloses a monocular self-supervision depth estimation method based on credible region depth controllable generation. Firstly, a monocular image is collected and input into a depth estimation module for multi-scale feature extraction, an initial depth map of a whole scene is generated, and relative camera attitude transformation between adjacent frames is estimated through an attitude estimation module; identifying a credible area and an uncredible area according to the obtained initial depth map and relative camera attitude transformation; depth estimation of a dynamic region is guided through a pseudo depth and contrast learning method; then, constructing a multi-task joint loss function, and optimizing depth estimation of the untrusted region by utilizing depth information of the trusted region through a comparison loss mechanism; and iteratively optimizing network parameters of the depth estimation module and the attitude estimation module, and outputting a final optimized depth map. According to the method, the depth estimation precision and robustness of the model in a complex dynamic environment are improved through a credible region depth controllable generation strategy.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Oil taking port positioning method and system based on monocular vision and laser positioning

The invention provides an oil taking port positioning method and system based on monocular vision and laser positioning. The method comprises the steps that laser point cloud data and a monocular image near an oil taking port are acquired; constructing a three-dimensional model of the oil taking port based on the laser point cloud data, and projecting the three-dimensional model of the oil taking port into a two-dimensional slice image according to an acquisition position label of the monocular camera; matching the two-dimensional slice image with the monocular image to determine the two-dimensional feature position of the oil extraction port; calculating the feature distance between the oil taking port and the camera based on the matched two-dimensional feature position of the oil taking port and the focal length label and the size label of the monocular image; on the basis of the position parameters of the oil taking port in the laser point cloud data, the feature distance is combined, and initial three-dimensional coordinates of the oil taking port are generated; and unifying the pixel coordinates of the monocular vision and the laser point cloud coordinates into the same coordinate system, and carrying out preliminary three-dimensional coordinate data fusion to obtain the final three-dimensional coordinates of the oil extraction port. The positioning precision of the transformer oil taking robot on the oil taking port is improved.
Owner:HUBEI INFOTECH SYST TECH CO LTD

Underwater target identification and positioning method based on physical model and deep learning fusion

The invention discloses an underwater target identification and positioning method based on fusion of a physical model and deep learning, and belongs to the technical field of computer vision, and the method comprises the steps: carrying out the transmissivity estimation and physical restoration of a collected underwater monocular image according to an underwater light propagation model, and carrying out the image correction; key feature points are extracted, and high-quality matching point pairs are screened in combination with the joint similarity; further deriving a basic matrix through the high-quality matching point pairs meeting the epipolar geometric constraint relation, solving an essential matrix, and obtaining relative attitude parameters between the cameras in combination with weighted re-projection-LM optimization; and finally, refraction correction triangulation is carried out according to the Snell's law, pixel-level fusion is carried out after scale normalization and space alignment are carried out on the refraction correction triangulation and dense depth output by the MiDaS, three-dimensional space coordinates of the target are inverted, and high-precision recognition and positioning of the underwater target are achieved. The system is light in structure, efficient in calculation, suitable for being integrated on various underwater autonomous or remote control robot platforms and used for tasks such as target recognition, tracking and positioning.
Owner:CENT SOUTH UNIV

Fish body length estimation method and system based on global guidance semantic segmentation network

The invention provides a fish body length estimation method and system based on a global guidance semantic segmentation network, relates to the technical field of computer vision and artificial intelligence, and aims to solve the problems that in a scene that a fish body in a monocular image and complex background interference coexist, and in a multi-fish-species mixed scene, the model segmentation generalization ability in the prior art is insufficient, and the fish body length estimation accuracy is poor. And high-precision and robust mask segmentation is difficult to realize. The method comprises the following steps: acquiring a fish body monocular image, and performing distortion correction on the fish body monocular image; inputting the corrected fish body monocular image into a trained global guide semantic segmentation network for semantic segmentation to obtain a semantic segmentation mask graph of the fish body and the calibration plate, and constructing a mapping relationship between the physical size and the pixel size of the calibration plate; and performing ellipse fitting on the semantic segmentation mask graph of the fish body, calculating the pixel length of the fish body, and converting the pixel length of the fish body into the actual length of the fish body. The method solves the problems in the prior art, provides the generalization ability of the network model, and achieves the precise estimation of the length of the fish body.
Owner:SHANDONG UNIV

Decoupling representation learning and Gaussian splash-based interpretable three-dimensional reconstruction method and device, and storage medium

The invention relates to an interpretable three-dimensional reconstruction method and device based on decoupling characterization learning and Gaussian splash, and a storage medium, and the method comprises the following steps: obtaining a monocular image, extracting two-dimensional features, obtaining compression features through convolution coding, obtaining a decoupled low-dimensional potential code through full connection processing, and obtaining a three-dimensional image; respectively converting into conditional representations of a geometric reconstruction branch and an appearance reconstruction branch; based on the two-dimensional coordinates on the predefined grid, three-dimensional points in a three-dimensional space are mapped through a plurality of MLP networks, a three-dimensional point cloud is obtained, and standard deviation-mean value modulation is carried out on the intermediate features by using conditional representation of geometric reconstruction branches; based on the three-dimensional point cloud, coding and projecting the three-dimensional point cloud to a plane to serve as initial three-plane features, and based on conditional representation of appearance reconstruction branches, coding the initial three-plane features into final three-plane features through a stylized U-Net network; and obtaining three-dimensional Gaussian scatter points based on the three-plane features and the three-dimensional point cloud to realize three-dimensional reconstruction.
Owner:NINGBO DIGITAL TWIN (EASTERN UNIV OF TECH) RES INST

Slope rockfall monitoring and early warning method, device and equipment based on dynamic depth estimation and medium

The invention provides a slope rockfall monitoring and early warning method, device and equipment based on dynamic depth estimation and a medium. Relates to the technical field of civil engineering operation maintenance. The method comprises the following steps: S1, acquiring a video sequence image; s2, performing dynamic target detection on the video sequence image by adopting a frame difference method; s3, based on a pre-trained rockfall identification model, identifying rockfall in the target frame; s4, based on a depth-pro monocular image depth estimation model, performing depth estimation on the rockfall in the target frame to obtain a rockfall position distance; s5, performing feature analysis on the rockfall in combination with the depth data to obtain feature parameters; s6, performing tracking and trajectory identification on the rockfall; and S7, performing alarm reminding according to tracking and trajectory recognition results. According to the method, the extracted features are trained to identify the slope rockfall, the position and the size of the rockfall can be identified more accurately, a sharp measurement depth map is rapidly provided, and the dynamic change of the slope rockfall is captured in time.
Owner:SOUTHWEST JIAOTONG UNIV

Miniature intelligent terminal intestinal image processing method and device, storage medium and computer equipment

The invention discloses a miniature intelligent terminal intestinal image processing method and device, a storage medium and computer equipment. The method comprises the following steps: preprocessing an original intestinal tract image set acquired by a monocular endoscope, calculating according to the preprocessed intestinal tract image set to obtain a motion calibration scale factor and a texture calibration scale factor, and adaptively fusing to obtain a conversion coefficient from a pixel to an actual size; extracting polyp region key feature points of each image of the preprocessed intestinal image set, constructing a Delaunay triangulation graph according to each polyp region key feature point, and performing non-rigid deformation matching based on a cross-graph convolutional network and an optimal transmission algorithm to obtain a matched feature point set; and sparse 3D reconstruction is carried out according to the matched feature point set to obtain a polyp point cloud, and the polyp size is calculated according to the Euclidean distance between two farthest points in the polyp point cloud and the conversion coefficient. According to the method, the problems of monocular image scale ambiguity and non-rigid deformation can be solved, and the more accurate polyp size can be obtained.
Owner:BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV +1

A monocular image depth estimation method, device and equipment based on diffusion model and target prompt

The present invention provides a method, device, and apparatus for monocular image depth estimation based on a diffusion model and target prompts, and relates to the field of computer vision technology. The method comprises: inputting a sample detection image into a depth estimation model to be trained, processing the sample detection image through a pre-trained image encoder to obtain a shallow spatial representation; processing the sample detection image through a pre-trained target detection network to obtain a target detection result; processing the target detection result through a target prompt module to be trained to obtain target prompt information; inputting the shallow spatial representation and target prompt information into a denoising network to be trained to obtain multi-scale features; inputting the multi-scale features into an interactive decoder to be trained to obtain a sample depth map; and training the depth estimation model to be trained based on the label depth map and the sample depth map to obtain a trained depth estimation model for depth estimation, thereby improving the accuracy of monocular depth estimation based on the diffusion model.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Monocular depth estimation method based on multi-modal information fusion

The invention discloses a monocular depth estimation method based on multi-modal information fusion, and the method comprises the steps: collecting multi-modal data composed of a monocular image and laser point cloud data, carrying out the fusion of the multi-modal data, and constructing a monocular depth estimation model, respectively inputting the fused multi-modal data into a multi-scale convolutional layer and a semantic segmentation layer of a monocular depth estimation model, carrying out multi-scale feature extraction by using the multi-scale convolutional layer, identifying the category of an object on a monocular image by using the semantic segmentation layer, transmitting the multi-scale features to an up-sampling layer by using a residual connection layer, and carrying out multi-scale feature extraction by using the up-sampling layer; and determining the initial monocular depth of the monocular image, adjusting the initial monocular depth by using a detail optimizer in combination with the identified object category, and predicting the monocular depth of the monocular image. According to the method, the image edge detection capability, the scene adaptability and the calculation efficiency can be considered at the same time, and the accuracy and precision of monocular depth estimation are improved.
Owner:NANJING INST OF TECH

Scale recovery method based on monocular depth estimation

The invention particularly relates to a scale recovery method based on monocular depth estimation, which comprises the following steps: acquiring a to-be-processed image acquired by a monocular camera and a current height parameter of a corresponding aircraft; performing depth identification on the to-be-processed image to obtain a corresponding initial depth image; performing coordinate system conversion processing on the initial depth image under camera coordinates to obtain a corresponding relative depth image under an aircraft coordinate system; performing point cloud segmentation processing on the relative depth image, and determining a current average height parameter of the aircraft based on a point cloud segmentation result; determining a conversion factor in combination with the current average height parameter and the current height parameter; and processing the initial depth image by using the conversion factor to obtain absolute depth information corresponding to each pixel point. According to the method, the absolute depth information can be obtained only by using the monocular image, and a new solution is provided for depth information precision of a monocular vision system.
Owner:YUNNAN MINZU UNIV

A novel monocular vision 3D object detection method based on key point constraints

The present invention discloses a novel monocular vision 3D object detection method based on key point constraints. The method first performs image preprocessing and label preprocessing on the monocular image; the preprocessed monocular image extracts digital features through a convolutional neural network, and based on the digital features, information for network branch tasks is extracted, the final output result of the model is generated and decoded to obtain the center point position and length, width, and height attributes of the object; a loss function is designed according to the information of the network branch tasks, and a loss term l between the predicted 2D key points and the 2D key points calculated based on the predicted 3D positions is added to the loss function 2d‑3d , and multi-branch collaborative training and model optimization are performed on the network to obtain a trained 3D object detection model; after format conversion and compilation of the model, it is deployed in an AI computing device for online inference of the model. The present invention uses a monocular image for 3D object detection and maintains high accuracy and robustness.
Owner:DOMINANT INTELLIGENT TECH (SUZHOU) CO LTD

Information processing system, information processing method, and non-transitory recording medium

An information processing system includes circuitry to read from a memory an environment map depicting a surrounding environment in which a mobile body including first and second imaging devices to capture monocular images moves. The circuitry determines a first position of the first imaging device on the environment map based on a monocular image captured by the first imaging device, determines a second position of the second imaging device on the environment map based on a monocular image captured by the second imaging device, generates correction information for correcting a scale of the environment map based on a distance between the first position and the second position on the environment map and a reference distance between the first imaging device and the second imaging device, and determines a position of the mobile body using the environment map and the correction information.
Owner:RICOH CO LTD

Depth estimation method and system for pond monocular image

The invention belongs to the technical field of computer vision, and particularly relates to a depth estimation method and system for a pond monocular image, and the method comprises the steps: constructing a pond monocular image depth estimation model, inputting a pond image into the pond monocular image depth estimation model, and obtaining the depth mapping of a pond scene, distance information from each pixel point to the camera is obtained, a pond image depth map is generated based on the distance information, and depth estimation of the pond image is achieved. The system comprises a pond monocular image depth estimation model, wherein the pond monocular image depth estimation model comprises a U-MonoVIT encoder, a local plane guide layer and a jumping multi-scale extension self-attention mechanism. The method can effectively solve the problem that the depth estimation accuracy is reduced due to the problems of fuzzy pond images, low contrast ratio, overexposure or darkness and the like, can accurately and efficiently achieve monocular depth estimation of pond scenes, and provides key support for comprehensive exploration of pond environments and biomass estimation.
Owner:CHINA AGRI UNIV

Large semi-trailer reversing auxiliary method, device and equipment

The invention provides a large semi-trailer reversing auxiliary method, device and equipment, and the method comprises the steps: obtaining the geometric center information of a ground marking line of a garage and the contour data of the ground marking line of the garage in real time through monocular image collection equipment; laser ranging data are obtained in real time through a laser ranging sensor; determining a first relative position relation of the vehicle relative to the garage according to the garage ground marking line contour data and the garage ground marking line geometric center information; determining a second relative position relation between the vehicle and the garage according to the laser ranging data; performing data fusion processing on the first relative position relation and the second relative position relation to obtain a target relative position between the vehicle and the garage; and performing auxiliary reversing according to the target relative position.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Parallax optimization method and multi-view camera

The invention discloses a parallax optimization method and a multi-view camera, and the method comprises the steps: judging whether a binocular parallax image accords with a preset condition, and if yes, optimizing the binocular parallax image based on a monocular image obtained by a monocular camera module, and generating an optimized parallax image, judging whether the binocular parallax image meets the preset condition or not at least comprises one of the following judging conditions: judging whether the parallax sparsity in the binocular parallax image is lower than a sparse threshold value or not, and judging whether the binocular parallax image is distorted or not. Through the technical scheme in the invention, under the condition of improving part of computing power, the reliability of outputting the disparity map is improved, and the influence of illumination variation on the disparity map is reduced.
Owner:元橡科技(北京)有限公司 +1

Identifying a shake point of a tree for autonomous harvesting

An autonomous harvesting machine for orchard operating environments is described. The autonomous harvesting machine uses machine vision techniques to identify and triangulate features in the operating environment using a stream of monocular images. For instance, the harvesting machine identifies and localizes a shake point of a tree by projecting virtual rays from the pose of the identification system to the identified emergence point feature. To harvest the fruit of trees in the orchard, the harvesting machine shakes the tree at the identified shake point. Additionally, the harvesting machine autonomously navigates through the orchard using a combination high resolution spatial information based on localized features and low resolution spatial information from accessed satellite images.
Owner:BONSAI ROBOTICS INC