Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3396 results about "Rgb image" patented technology

System and Method for Multi-Modal Hyperspectral Image Generation with Cross-Modal Attention and Adaptive Quality Assurance

A system and method are disclosed for generating hyperspectral images from multi-modal sensor data including RGB, LiDAR, thermal, and near-infrared inputs. Training data includes hyperspectral images and corresponding multi-modal measurements. Spectral band grouping is performed based on correlation coefficients. A multi-modal decomposition network with cross-modal attention mechanisms generate reconstructed hyperspectral images by fusing complementary sensor information. A fine-tuning network creates reconstructed RGB images. A comprehensive quality assurance system analyzes spectral consistency, cross-modal coherence, and fusion artifacts to generate quality metrics. Missing data compensation strategies handle corrupted sensor inputs using information from other modalities. The system includes temporal integration for video sequences and multi-resolution processing for different sensor resolutions. Quality metrics guide network weight adjustments to improve reconstruction accuracy while maintaining robustness to sensor failures and environmental variations.
Owner:ATOMBEAM TECH INC

Robot grabbing posture generation method and related device

The invention discloses a robot grabbing posture generation method and a related device, and the method comprises the steps: obtaining an RGB image, a depth image and a position coordinate of a target object, and generating a spatial calibration matrix through data preprocessing; target object segmentation is carried out on the calibrated RGB matrix according to the target point coordinate sequence, a segmentation mask is generated, and object mass center coordinates are calculated; generating a three-dimensional point cloud by using the calibrated depth matrix, the segmentation mask and the camera parameters, extracting a plane point set and a non-plane point set through an RANSAC algorithm, and analyzing to obtain an axial feature vector and a plane normal vector; finally, grabbing parameters are calculated according to the feature vectors and the centroid coordinates, and a grabbing posture transformation matrix is generated through vector operation. According to the technical scheme, under the condition of not depending on a preset model library, the grabbing posture of an unknown object is generated by analyzing the geometrical characteristics of the object, the accuracy problem in the process of converting two-dimensional image information into three-dimensional grabbing parameters is solved, and the adaptability and reliability of a robot grabbing task are improved.
Owner:SUZHOU SHUTU GUCHUANG TECHNOLOGY CO LTD

Industrial robot disordered grabbing system and method based on multi-modal perception

The invention relates to the technical field of industrial robots, in particular to an industrial robot disordered grabbing system and method based on multi-modal sensing, and the system comprises a multi-modal sensing module, a pose estimation and correction module, a grabbing task planning module and a mechanical arm execution module. The multi-modal sensing module is used for acquiring RGB image data, depth image data, point cloud data and force / torque data in a working environment, and the pose estimation and correction module performs feature extraction and fusion on multi-modal data based on topology invariant mapping and performs adaptive pose estimation on a target object. The grabbing task planning module generates an optimal grabbing strategy based on deep reinforcement learning and outputs a grabbing control instruction, and the mechanical arm execution module receives the grabbing control instruction, controls a mechanical arm to execute a grabbing action and monitors a grabbing state through force / torque feedback. And the environmental adaptability and the perception robustness are improved.
Owner:GUANGDONG PLATINUM STRONTIUM TECH CO LTD

Aero-engine augmented reality virtual-real fusion method based on depth prior scene reconstruction

An aero-engine augmented reality virtual-real fusion method based on depth prior scene reconstruction comprises the following steps: firstly, collecting data of an aero-engine look-around RGB image and a depth image, calculating a camera frame pose by using an SFM algorithm and obtaining a sparse point cloud, and optimizing the camera pose by using an ICP algorithm and the point cloud converted from the depth image to obtain an aligned prior dense point cloud; gaussian points are initialized based on the priori dense point cloud and the sparse point cloud, and random generation of 3D Gaussian is limited by using the boundary of the dense point cloud as geometric constraint; through the three-dimensional reconstruction of the adjacent multi-view enhancement strategy, the spatial and visual consistency is improved, and an aero-engine model is reconstructed; then understanding the reconstructed aero-engine model based on semantic segmentation, identifying and marking core components, and finally performing virtual-real fusion visualization on the reconstructed aero-engine through augmented reality to guide training, assembling and maintenance operations. According to the invention, rapid high-quality three-dimensional reconstruction and augmented reality virtual-real fusion visualization of the aero-engine are realized.
Owner:XI AN JIAOTONG UNIV

Human shape posture recognition method and system based on image analysis

The invention relates to a human shape posture recognition method and system based on image analysis, and the method comprises the steps: inputting the reference human shape data of a monitored object, and collecting the multi-modal data of an RGB image, an infrared image and inertial measurement data in a target scene in real time; space-time alignment processing is carried out, and corresponding features are fused; image space features are extracted, time sequence modeling is carried out on inertial data, and dynamic weighted fusion of two paths of network outputs is realized through a gating mechanism; human body basic joint points are positioned, and refined posture vectors including joint angles and limb relative positions are generated; according to attitude data and environment information collected in real time, adaptively adjusting an attitude classification threshold value and a similarity measurement standard, and predicting abnormal behaviors in a future time period; and when the abnormal behavior is predicted, triggering to execute a preset safety measure. Multi-modal data can be efficiently fused, attitude features can be accurately extracted, an identification strategy can be adaptively adjusted, and human shape attitude identification with behavior trend prediction capability can be realized.
Owner:CHINA WEST NORMAL UNIVERSITY

Hyperspectral calculation imaging method and system based on multi-domain fusion of spatial, spectral and frequency domains, and medium

Disclosed in the present invention are a hyperspectral calculation imaging method and system based on multi-domain fusion of spatial, spectral and frequency domains, and a medium. The method of the present invention comprises: using a two-dimensional offline discrete cosine transform (DCT) to convert an RGB image Y into a frequency domain to obtain a frequency domain feature map Yfreq; extracting a frequency information image (I) from the frequency domain feature map Yfreq; using a two-dimensional offline inverse discrete cosine transform (IDCT) to transform the frequency information image (I) into a spatial domain to obtain a frequency information image Xfreq of the spatial domain; and fusing the frequency information image Xfreq of the spatial domain into the spatial-spectral domain features of the RGB image Y to generate a hyperspectral image (II). The present invention aims to solve the problems of poor detail information and low reconstruction accuracy of hyperspectral images in existing hyperspectral calculation imaging, and realizes high-fidelity reconstruction of a target spectrum.
Owner:HUNAN UNIV

Unmanned aerial vehicle identification early warning method and system

The invention provides an unmanned aerial vehicle identification early warning method and system. The method comprises the following steps: acquiring RGB image data, thermal radiation data and spectral data of an unmanned aerial vehicle no-fly zone through a visible light camera, an infrared thermal imager and a multispectral imager; performing transmission preprocessing on the RGB image data, the thermal radiation data and the spectral data, and adaptively adjusting multi-level fusion of a fusion weight based on real-time environmental parameters to generate target fusion data; based on a deep learning model and a tracking prediction algorithm, performing unmanned aerial vehicle identification early warning on the target fusion data, and generating early warning data; and transmitting the target fusion data and the corresponding abnormal event log to a cloud server, and updating the deep learning model by adopting the target fusion data and the abnormal event log. Through cooperative work of a multi-mode sensor, visible light, infrared, multispectral and other wave bands are covered, all-weather and full-scene unmanned aerial vehicle detection is achieved, the fusion weight is adjusted in real time based on real-time environment parameters, and the accuracy of the recognition result under the complex air situation is ensured.
Owner:GLOBAL GENERAL AVIATION (HANGZHOU) CO LTD

Multi-modal visual fusion complex scene small target detection tracking method and system

The invention discloses a multi-modal visual fusion complex scene small target detection tracking method and system, and relates to the technical field of unmanned aerial vehicle target tracking, and the method comprises the steps: employing a visible light camera, an infrared thermal imager and a laser radar sensor which are carried on an unmanned aerial vehicle platform, and synchronously collecting RGB images, thermal infrared images and point cloud data; the consistency of the multi-modal data is ensured through data preprocessing and space-time alignment; constructing a lightweight double-branch network to extract multi-scale features, generating a fusion feature map by adopting adaptive weighted fusion, and generating depth information by utilizing point cloud to assist in scale estimation; a small target detection head is designed based on the fusion feature map, and precise detection is realized in combination with a feature pyramid network, adaptive scale prediction and a context awareness suppression mechanism; furthermore, through multi-mode cooperative tracking, including target association, spatio-temporal context modeling, trajectory prediction and a re-detection mechanism, tracking continuity is ensured.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

Uncoupling robot control system and method based on multi-source visual fusion

The embodiment of the invention provides an unhooking robot control method based on multi-source visual fusion, which is applied to the technical field of robot control and comprises the following steps: acquiring an RGB image, a depth image, an infrared image and IMU data through a multi-source sensing system mounted at the tail end of a robot; carrying out feature fusion identification by adopting a double-branch neural network, and outputting the boundary contour of the lifting hook and the three-dimensional coordinates of the optimal grabbing point; the visual coordinates are unified to a robot base coordinate system through a registration correction mechanism; a Transform prediction model is constructed based on the visual and inertial signals, and future pose changes of the lifting hook are estimated; a feedforward control track is generated to counteract swing of the lifting hook, and track correction is carried out in combination with visual servo feedback; and a joint instruction is generated through path planning and inverse kinematics solution, and the mechanical arm is driven to complete precise unhooking operation. According to the method, the recognition precision, the anti-interference capability and the operation success rate of unhooking operation in complex illumination and dynamic environments are effectively improved.
Owner:ANHUI HUADIAN SUZHOU POWER GENERATION

Mechanical arm motion control method based on multi-agent cooperation

The invention discloses a mechanical arm motion control method based on multi-agent cooperation, and the method comprises the steps: firstly, receiving an RGB image through a sub-task generation agent, and generating a structured sub-task sequence according to a natural language task instruction of the RGB image; secondly, performing joint modeling on a task text and a scene image through a 3D sensing intelligent body, positioning specific coordinates of a target object in a three-dimensional space, reasoning dynamic characteristics of a current environment based on historical state information of a robot by combining an environment sensor, and generating an environment sensing vector; and finally, the action generation agent performs fusion modeling according to the subtask text, the subtask target coordinates, the current state of the robot and the environment perception vector, generates a continuous action vector, drives a mechanical arm to complete each subtask action, and constructs closed-loop feedback by a controller and a discriminator to realize task execution state judgment and automatic circulation. The precise action control instruction can be effectively generated, and the task execution fineness of the mechanical arm is remarkably improved.
Owner:CHINA JILIANG UNIV +1

Remote sensing target detection method and system for low-visibility image

The invention relates to the technical field of remote sensing monitoring, in particular to a remote sensing target detection method and system for a low-visibility image. The method comprises the following steps: acquiring multi-modal remote sensing image data; carrying out defogging enhancement processing on the low-visibility input image; normalizing the defogged RGB image and the defogged IR image, and then splicing and fusing the RGB image and the IR image; carrying out layer-by-layer coding on the multi-modal fusion image by adopting a mixed trunk structure fusing Transform, Mamba and CNN (Convolutional Neural Network); performing frequency domain decomposition on the trunk output features based on two-dimensional wavelet transform; generating an HR feature map by adaptively selecting a key region; and carrying out cross-scale aggregation on the HR feature map to obtain a detection target frame. Through the multi-modal image defogging enhancement and feature distillation mechanism, the definition and contrast of the remote sensing image in severe weather such as haze and rainy days are effectively enhanced, the shielding interference of environmental degradation on small target detection is weakened, and the stability and adaptability of the model in complex weather scenes are enhanced.
Owner:YANTAI UNIV

Grabbing attitude generation method and system based on multi-modal large model

The invention discloses a grabbing posture generation method and system based on a multi-modal large model, and the method comprises the steps: carrying out the cross-modal matching of visual features and semantic features in the multi-modal large model when a voice instruction and an RGB image are inputted, and obtaining the position information of a control function code and a target object; when an RGB image with a hand drawing instruction is input, obtaining position information of a control function code, a target object and a path point; calculating the point cloud data of the target object according to the position information of the target object and the depth information, inputting the ideal point cloud of the target object into a target recognition network model after preprocessing, carrying out the grabbing region recognition of the point cloud of the target object region, outputting a region with high grabbing confidence, and mapping a real coordinate system; constructing a point cloud bounding box, and generating a grabbing posture candidate set; the grabbing posture with the highest quality is selected as the grabbing posture of the robot by calculating the grabbing posture candidate score; and executing a target grabbing task in combination with the control function code and the grabbing path.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

System and method for reconstructing 3D scene data from 2D image data

A method and apparatus for reconstructing a three-dimensional (3D) scene from a two-dimensional (2D) input image of the scene using a fully-differentiable transformer-based encoder-decode. A 2D input image encoded into a set of image features using a pre-trained vision transformer model, wherein the vision transformer model is pre-trained with multi-view RGB image supervision and point cloud supervision. The set of image features is projected onto a 3D triplane representation using a transformer decoder to obtain output triplane tokens. A triplane representation is created from the tokens and queried. 3D point features of color and density for volumetric rendering re predicted using a multi-layer perceptron. The geometry of the generated 3D asset is represented with a surface mesh including vertices and triangular faces. A texture map by is created with a multichannel image in UV space. Multiple views of the 3D scene are simultaneously generated based on the surface mesh.
Owner:FUTUREVERSE IP LTD

Agaricus bisporus automatic picking method, device and equipment based on artificial intelligence and medium

The invention relates to the technical field of intelligent agricultural equipment, and discloses an automatic agaricus bisporus picking method, device and equipment based on artificial intelligence and a medium. The method comprises the steps that RGB images and three-dimensional point cloud data of the agaricus bisporus growth environment are collected in real time through a binocular vision camera and a depth sensor which are installed at the tail end of a picking robot, and environment illumination parameters and cultivation bed coordinate information are obtained; inputting the RGB image into a pre-trained lightweight convolutional neural network, identifying a pileus contour, a stipe position and a maturity level of the agaricus bisporus, calculating a space coordinate, a height and a growth inclination angle of the agaricus bisporus based on the three-dimensional point cloud data, and constructing an agaricus bisporus target database; a mechanical arm picking path is generated based on the agaricus bisporus target database, and path weight parameters are optimized online through a preset reinforcement learning model; according to the maturity grade and morphological characteristics of the agaricus bisporus, the clamping force and the rotating angle of the end effector are dynamically adjusted, and automatic picking of the agaricus bisporus is completed.
Owner:FUJIAN POLYTECHNIC OF WATER CONSERVANCY & ELECTRIC POWER

Diffusion-based multiple-modality image fusion

An image-guided diffusion network has two Convolution Neural Networks (CNNs). A RGB image and an IR image are concatenated with a Gaussian noise image and input to a denoising neural network that merges information from the RGB and IR images as noise is removed over many iterations. Then an enhancement neural network up-samples for Super Resolution (SR) and convolutes to generate a condition vector that controls Global Feature Modulation (GFM) at three convolution layers to generate a SRGFM enhanced fusion image. Timesteps are embedded using adaptive group normalization blocks within Adaptive Bottleneck Residual (ABR) blocks in the denoising network, which is a UNet having many levels of ABRs, and in the enhancement network before feature modulation. Global image features are detected by triple convoluting the image input to the enhancement network to generate the condition vector that controls feature modulation blocks at three layers of convolution.
Owner:HONG KONG APPLIED SCI & TECH RES INST

Defect detection method, device, computer equipment, and storage medium

Provided is a defect detection method and device, computer equipment and a storage medium. The method includes: acquiring an RGB image, a depth image and a sample label of a detection object sample; performing feature map extraction and feature map fusion on the RGB image and the depth image by a feature extraction network of the defect detection model, to obtain a fused feature map; performing defect detection based on the fused feature map by a feature reconstruction network of the defect detection model, to obtain a defect score map, wherein the defect score map being obtained by fusing a global defect score map which is generated based on a global defect detection network with a local defect score map which is generated by a local defect detection network; and updating parameters of the defect detection model based on the defect score map and the sample label.
Owner:JABIL INC

Thyroid ultrasonic robot automatic scanning method, device and equipment based on RGB image and depth information and medium

The invention relates to the technical field of computer vision, and discloses a thyroid ultrasonic robot automatic scanning method, device and equipment based on RGB images and depth information and a medium, and the method comprises the steps: obtaining image information and depth information, coding the depth information, fusing and recognizing a target scanning area, determining an initial scanning point and an initial scanning direction, controlling the scanning probe to scan and collect a real-time scanning image, analyzing the real-time scanning image to recognize a preset target and an artifact area, adjusting a scanning posture and a scanning path based on a recognition result, monitoring a continuous existence state of the preset target, and stopping scanning when the preset target is not recognized continuously. The target area is identified by fusing the multi-modal image information, the scanning posture and path are dynamically adjusted in combination with real-time image analysis, scanning termination is intelligently controlled according to the target detection result, the positioning accuracy, image quality and standardization level of ultrasonic scanning are improved, and the method is suitable for automatic ultrasonic imaging of thyroid and superficial organs.
Owner:SHENZHEN BEAUTIFUL RUBIKS CUBE ROBOT CO LTD

Operation and maintenance manipulator intelligent control method and system based on visual identification

The invention discloses an operation and maintenance manipulator intelligent control method and system based on visual identification, and relates to the technical field of intelligent manipulator control, and the method comprises the steps: collecting RGB image data and depth image data of an operation and maintenance operation area, and obtaining a standardized image matrix and a mapping relation matrix; inputting the standardized image matrix into an improved ResNet residual network model, generating a comprehensive feature descriptor, and calculating a spatial position coordinate and an attitude angle of the target equipment based on the mapping relation matrix; based on the current joint angle state of the manipulator, an improved Jacobian matrix inverse kinematics algorithm is used for solving a target angle sequence of each joint, a preset operation mode library is matched based on the comprehensive feature descriptor, and a grabbing force parameter and a motion speed parameter are determined; and converting the target angle sequence into a control instruction, and sending the control instruction to each joint driver of the manipulator to drive the manipulator to complete action planning. According to the invention, full-process automation from environment perception to task execution is realized.
Owner:AOWEI TECH (NANJING) CO LTD

Target tracking method based on fusion of RGB data and event data of visual Mama

The invention discloses an RGB data and event data fusion target tracking method based on visual Mama, relates to the technical field of computer vision, and solves the technical problem that efficient and accurate target tracking of RGB data and event data is difficult to realize. The method comprises the following steps: performing frame-level synchronization and conversion on an original RGB image of an RGB camera and an original event stream of an event camera to generate RGB input data and event input data; obtaining an RGB feature map and an event feature map through a feature extraction module of the Mamba network; performing linear transformation and deep convolution operation through an interaction module to obtain RGB features and event features; dynamically fusing the RGB features and the event features through a fusion module to obtain cross-modal fusion features; and the template features and the search features are aligned through an alignment module to generate an alignment feature graph, and target positioning and tracking are completed. According to the method, through efficient modal fusion and the state space model, the target tracking precision and robustness in a dynamic scene are remarkably improved.
Owner:UESTC (SHENZHEN) ADVANCED RES INST +1

Construction site real-time monitoring system and method based on unmanned aerial vehicle technology

The invention relates to the technical field of building construction monitoring, in particular to a construction site real-time monitoring system and method based on an unmanned aerial vehicle technology. The method specifically comprises the steps that an unmanned aerial vehicle carries a high-definition camera, an infrared thermal imager and a laser radar and collects multi-modal data of a construction site in real time; a deep learning algorithm is adopted to carry out multi-modal semantic segmentation, RGB images, infrared thermal imaging and laser point cloud projection data are fused, YOLOv8 target detection and time sequence behavior modeling technologies are combined, and abnormal risks are dynamically identified; by constructing a context-aware comprehensive risk scoring model, the risk level of a construction scene is quantified, and a graded early warning mechanism is triggered; and meanwhile, based on a three-dimensional modeling technology, tank vertex cloud data is analyzed by utilizing an LIO-SLAM algorithm, and the inclination degree of the structure is detected and is linked with emergency response. According to the invention, through multi-sensor fusion, dynamic risk assessment and closed-loop control, intelligentization and precision of construction monitoring are realized.
Owner:SHANGHAI INSTALLATION ENGINEERING GROUP CO LTD

Indoor robot navigation method based on multi-modal feature fusion

The invention relates to an indoor robot navigation method based on multi-modal feature fusion. The method comprises the following steps: constructing a semantic map based on visual observation environment information; the method comprises the following steps: acquiring an RGB image of an indoor scene object, converting the RGB image into point cloud data, and preprocessing the RGB image and the point cloud data; image multilayer semantic features and point cloud features in the RGB image and the point cloud data are extracted respectively, and initial fusion is carried out; performing weighted fusion on the multi-layer semantic features and the point cloud features of the image by adopting space-channel-cross-modal multi-attention dynamic cooperation; and predicting a long-term target in a map space from top to bottom based on the fused feature map and the semantic map, and performing navigation path planning based on the current position and the long-term target. Under low-cost hardware configuration, the robustness of environment perception, the real-time performance of decision response and the usability of system integration in a dynamic complex environment are comprehensively improved, and the method is particularly suitable for application scenes such as indoor service robots.
Owner:SOUTHWEST JIAOTONG UNIV

Intelligent discrimination method for pseudo soldering microcracks based on intelligent visual identification technology

The invention relates to an intelligent visual identification technology-based cold solder joint microcrack intelligent discrimination method, which comprises the steps of collecting an initial RGB image of a to-be-detected welding spot, carrying out two-dimensional discrete cosine transform and inverse two-dimensional discrete cosine transform on the initial RGB image to obtain an enhanced image, and fusing the enhanced image with an R channel of the initial RGB image to obtain a fused image; forming a dual-channel feature map; calculating the phase consistency of the dual-channel feature map, and obtaining a suspected candidate region of the pseudo soldering microcrack through an adaptive threshold segmentation method; acquiring an RGB image sequence of a continuous time sequence of the welding spots, and performing anomaly detection to obtain an abnormal region set; and constructing a welding spot thermal diffusion model, and inputting the geometric parameters and the environmental parameters in the abnormal region set into the welding spot thermal diffusion model to obtain a final judgment result of the pseudo soldering microcracks. According to the method, through multi-dimensional feature fusion and continuous time sequence dynamic tracking, the detection precision of the pseudo soldering microcracks is remarkably improved, the false detection rate is reduced, and the final judgment result is more accurate.
Owner:JUXIN ELECTRONICS TECH MEIZHOU CO LTD

Semi-supervised image semantic segmentation method and system based on visual basic model

The invention provides a semi-supervised image semantic segmentation method and system based on a visual basic model, and the method comprises the steps: constructing a multi-task model which comprises a visual basic model and a depth estimation basic model, and the visual basic model is connected with a task solution head, an adapter parameter efficient fine tuning module and a multi-modal cross fusion module; the task solution head comprises a semantic segmentation head and a depth estimation head; extracting semantic hierarchy features and a depth feature map of the RGB image, performing cross attention fusion on the semantic hierarchy features and the depth feature map, and inputting obtained fusion features into a semantic segmentation head and a depth estimation head respectively; semi-supervised learning is adopted to train a multi-task model, only parameters in the adapter parameter efficient fine tuning module and the multi-modal cross fusion module are trained, and a multi-task loss function is adopted. The image semantic segmentation model obtained through training can improve semantic segmentation performance, reduce training cost and is suitable for different tasks.
Owner:SHANGHAI JIAOTONG UNIV

Method and system for predicting motion track of crane based on RGBD image

The invention discloses a crane motion track prediction method based on an RGBD image, and belongs to the technical field of crane intelligent control, and the crane motion track prediction method comprises the following steps: S1, synchronously collecting multi-modal data, deploying a binocular RGBD camera array, and obtaining an RGB image and depth point cloud data of a crane operation area in real time; s2, cross-modal feature alignment and fusion are carried out; s3, track prediction and physical constraint correction; and S4, calculating the minimum safety distance between the crane and the dynamic object in real time. According to the method, a cross-modal space-time alignment mechanism is constructed by integrating the binocular RGBD camera array, the millimeter wave radar, the laser range finder and IMU data, the deformable convolutional network is adopted to dynamically compensate sensor data offset, multi-source data conflicts are eliminated, and compared with a traditional single sensor scheme, the method has the advantages that the method is simple in structure and convenient to operate. The multi-mode cooperative mechanism effectively solves the problem of sensing failure in a sensor blind area and a severe environment.
Owner:HENAN MECHANICAL & ELECTRICAL ENG COLLEGE +1

Battery replacement robot target point cloud segmentation method based on multi-scale attention aggregation

The invention discloses a multi-scale attention aggregation-based target point cloud segmentation method for a battery replacement robot, and the method comprises the steps: 1, collecting an RGB image, a depth image and three-dimensional point cloud data of a target fastener through a binocular structured light depth camera, and carrying out the fusion to generate a FastSeg3D data set; 2, a two-stage preprocessing method is provided, noise points are removed through radius filtering, and background point clusters far away from a target are removed through DBSCAN density clustering; 3, the network encoder uses a local feature aggregation module to extract geometric features, and the calculation complexity is reduced in combination with a random sampling strategy; 4, embedding a multi-scale attention aggregation module into the jump connection of the encoder and the decoder, fusing the features through a channel and a space attention unit, and achieving the self-adaptive weight weighting of the features of each layer of the encoder; and 5, recovering the resolution of the original point cloud by adopting nearest neighbor interpolation up-sampling, outputting a segmentation semantic tag, and obtaining a high-quality point cloud target. According to the method, the operation time of the battery replacement robot is shortened, and the balance problem of large-scale target point cloud segmentation speed and precision is solved.
Owner:SOUTHEAST UNIV

Construction site three-dimensional scene reconstruction method based on unmanned aerial vehicle image and monitoring video

The invention discloses a construction site three-dimensional scene reconstruction method based on an unmanned aerial vehicle image and a monitoring video. The method comprises the following steps: acquiring a construction site scene multi-view image; a sparse three-dimensional point cloud and a camera pose are generated through a feature matching and motion recovery structure algorithm, and dense reconstruction is carried out to obtain a global three-dimensional point cloud; aligning the global three-dimensional point cloud with a world coordinate system by using geographic position information; estimating the position of a shooting camera in the three-dimensional point cloud, sampling a candidate view angle and rendering a virtual RGB image; based on two-dimensional feature matching of the shot image and the virtual RGB image, internal parameters and external parameters of the shooting camera are iteratively solved through triangulation and a pose optimization algorithm; monocular depth estimation is carried out on the shot image, and the shot image is converted to a measurement scale through static region depth alignment; and projecting the depth of the shot image to the three-dimensional point cloud, updating the dynamic object in real time, and fusing the dynamic object into a complete three-dimensional scene model. The method has the advantage that the real-time three-dimensional reconstruction of the dynamic scene of the construction site is realized.
Owner:CHINA RAILWAY 24TH BUREAU GROUP CO LTD

Three-dimensional scene and object reconstruction method, system and equipment based on SDF and Gaussian field

The invention discloses a three-dimensional scene and object reconstruction method, system and device based on SDF and a Gaussian field, belongs to three-dimensional scene reconstruction in the technical field of computer vision, and aims to solve the technical problem that the three-dimensional scene reconstruction quality is not high due to the fact that geometric precision and rendering quality cannot be achieved at the same time in the three-dimensional reconstruction process in the prior art. The method comprises the following steps: initializing and updating a TSDF voxel volume by using a depth map and an RGB image to obtain a camera external parameter and a scene three-dimensional initial model; sDF prediction is carried out on the scene three-dimensional initial model through an SDF neural network, scene surface points are generated according to the light direction, and a high-precision scene geometric model is obtained; training the SDF neural network; performing grid division on a bounding box of the scene three-dimensional initial model, outputting an SDF predicted value of each grid point by using the trained SDF neural network, and extracting a refined scene grid model; and performing Gaussian rendering by using the high-precision scene geometric model to generate a final rendered image.
Owner:CHENGDU UNIV OF INFORMATION TECH

Modal sharing information layered unwrapping fusion network for RGB-T target tracking

The invention relates to the technical field of multi-modal visual tracking, and discloses a modal sharing information layered unwrapping fusion network for RGB-T target tracking, which comprises a double-flow feature extraction module which adopts a parameter sharing ResNet-50 network and is used for respectively extracting features of an RGB image and a thermal infrared image; the cross-modal attention module is connected with the double-flow feature extraction module and is used for realizing feature interaction enhancement of RGB (Red, Green, Blue) and a thermal mode through a bidirectional attention mechanism; and the layered unwrapping mining module is connected with the cross-modal attention module and is used for mining inter-modal deep complementary information through multi-level residual calculation. By introducing a lightweight attention mechanism and a modal alignment strategy, complementary information between modals is mined layer by layer and dynamic fusion is realized, so that the utilization efficiency of modal residual information is remarkably improved, and the stability and robustness of the system in a complex environment are enhanced.
Owner:SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

Dexterous hand multi-object stable grabbing method and system based on vision and touch perception

The invention provides a dexterous hand multi-object stable grabbing method and system based on vision and touch perception. The dexterous hand multi-object stable grabbing method comprises the steps that RGB images and depth maps of target objects are collected through a monocular RGB-D camera, and dense point cloud data are generated; geometric features are extracted by using a three-dimensional point cloud neural network, and composite object distribution features are analyzed in combination with visual information; an initial grasping strategy is generated based on the features, actions are executed, and the joint state and contact force data are synchronously obtained in real time through body sensing and a fingertip force sensor; in the carrying process, multi-dimensional force perception and the joint state are fused to analyze object stress distribution and center-of-gravity shift; vision and force sensing data are used as input. Vision, body sensing and force sensing data are fused, the problems of three-dimensional structure recognition and dynamic stress regulation and control in multi-object composite grabbing are solved, compared with a single sensing scheme, the grabbing planning precision and the self-adaptive capacity in a complex scene are remarkably improved, and the method is particularly suitable for stable carrying of combinations of loose objects such as dinner plates carrying tableware.
Owner:SHANGHAI JIAOTONG UNIV

Map construction method and system based on laser vision dynamic weighted fusion

The invention provides a map construction method and system based on laser vision dynamic weighted fusion, and belongs to the technical field of data processing, and the method comprises the steps: obtaining an RGB image and a depth image through a vision sensor, and obtaining point cloud data through a laser radar; aligning the depth image with the RGB image, and removing invalid depth pixels; based on an ORB-SLAM2 algorithm, performing visual SLAM, extracting features of the depth image and the RGB image, and outputting sparse visual point cloud data; based on a Gmapping algorithm, performing laser radar SLAM, and outputting a 2D occupied grid map; projecting the sparse visual point cloud data into a 2D occupation grid map, and calculating the semantic occupation probability of each grid; calculating the geometric occupancy probability of each grid through an anti-sensor model; performing weighted fusion on the semantic occupancy probability and the geometric occupancy probability, and calculating a fusion occupancy probability; and generating a fusion map according to the fusion occupation probability.
Owner:SHIHEZI UNIVERSITY