Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1920 results about "Rgb image" patented technology

Multi-modal visual fusion complex scene small target detection tracking method and system

The invention discloses a multi-modal visual fusion complex scene small target detection tracking method and system, and relates to the technical field of unmanned aerial vehicle target tracking, and the method comprises the steps: employing a visible light camera, an infrared thermal imager and a laser radar sensor which are carried on an unmanned aerial vehicle platform, and synchronously collecting RGB images, thermal infrared images and point cloud data; the consistency of the multi-modal data is ensured through data preprocessing and space-time alignment; constructing a lightweight double-branch network to extract multi-scale features, generating a fusion feature map by adopting adaptive weighted fusion, and generating depth information by utilizing point cloud to assist in scale estimation; a small target detection head is designed based on the fusion feature map, and precise detection is realized in combination with a feature pyramid network, adaptive scale prediction and a context awareness suppression mechanism; furthermore, through multi-mode cooperative tracking, including target association, spatio-temporal context modeling, trajectory prediction and a re-detection mechanism, tracking continuity is ensured.
Owner:BEIJING INSTITUTE OF GRAPHIC COMMUNICATION

Uncoupling robot control system and method based on multi-source visual fusion

The embodiment of the invention provides an unhooking robot control method based on multi-source visual fusion, which is applied to the technical field of robot control and comprises the following steps: acquiring an RGB image, a depth image, an infrared image and IMU data through a multi-source sensing system mounted at the tail end of a robot; carrying out feature fusion identification by adopting a double-branch neural network, and outputting the boundary contour of the lifting hook and the three-dimensional coordinates of the optimal grabbing point; the visual coordinates are unified to a robot base coordinate system through a registration correction mechanism; a Transform prediction model is constructed based on the visual and inertial signals, and future pose changes of the lifting hook are estimated; a feedforward control track is generated to counteract swing of the lifting hook, and track correction is carried out in combination with visual servo feedback; and a joint instruction is generated through path planning and inverse kinematics solution, and the mechanical arm is driven to complete precise unhooking operation. According to the method, the recognition precision, the anti-interference capability and the operation success rate of unhooking operation in complex illumination and dynamic environments are effectively improved.
Owner:ANHUI HUADIAN SUZHOU POWER GENERATION

Semi-supervised image semantic segmentation method and system based on visual basic model

The invention provides a semi-supervised image semantic segmentation method and system based on a visual basic model, and the method comprises the steps: constructing a multi-task model which comprises a visual basic model and a depth estimation basic model, and the visual basic model is connected with a task solution head, an adapter parameter efficient fine tuning module and a multi-modal cross fusion module; the task solution head comprises a semantic segmentation head and a depth estimation head; extracting semantic hierarchy features and a depth feature map of the RGB image, performing cross attention fusion on the semantic hierarchy features and the depth feature map, and inputting obtained fusion features into a semantic segmentation head and a depth estimation head respectively; semi-supervised learning is adopted to train a multi-task model, only parameters in the adapter parameter efficient fine tuning module and the multi-modal cross fusion module are trained, and a multi-task loss function is adopted. The image semantic segmentation model obtained through training can improve semantic segmentation performance, reduce training cost and is suitable for different tasks.
Owner:SHANGHAI JIAOTONG UNIV

Three-dimensional scene reconstruction method based on intelligent LED street lamp multi-mode sensor

The invention discloses a three-dimensional scene reconstruction method based on an intelligent LED street lamp multi-mode sensor. The three-dimensional scene reconstruction method comprises the following steps that RGB images are obtained and preprocessed; the information is input to a visual feature coding module, two-dimensional bounding box information is extracted, and an object segmentation module is guided to output a two-dimensional segmentation mask; obtaining point cloud data, projecting the point cloud data to the standardized RGB image, and screening target points in combination with the two-dimensional segmentation mask; complementing the preliminary segmentation result of the point cloud, mapping the result to a standardized RGB image, and extracting a pixel region; carrying out joint coding, implicit representation and neural decoding processing on the object-level RGB image and the point cloud complete segmentation result; and fusing into an original three-dimensional scene, and completing spatial restoration through point cloud registration, attitude optimization and semantic constraint. The invention provides an efficient three-dimensional scene reconstruction method in combination with a multi-mode sensor of an intelligent LED street lamp, and the method has high precision, real-time performance and dynamic target processing capability.
Owner:ZHEJIANG UNIV +1

Crop disease diffusion prediction method and system based on multi-modal fusion

The invention discloses a crop disease diffusion prediction method and system based on multi-modal fusion, and the method comprises the following steps: S1, collecting and preprocessing an RGB image sequence and a sensor data sequence of a crop growth environment, and generating an RGB image time sequence difference result and a sensor difference result through time difference processing; s2, mapping the RGB image time sequence difference result and the sensor difference result to a shared time sequence space through a time alignment algorithm, and generating a sensor alignment result and an RGB alignment result; s3, an FD-ViT prediction model is constructed; inputting the sensor alignment result and the RGB alignment result into an FD-ViT prediction model for prediction, and generating a prediction result; and S4, generating a disease diffusion thermodynamic diagram and early warning information according to a prediction result. According to the method, RGB image data and sensor network data are fused, a Transform-based time sequence prediction model is constructed, and early recognition and diffusion trend prediction of crop diseases are realized.
Owner:HANGZHOU DIANZI UNIV

Automobile part quality detection method and system based on artificial intelligence visual inspection

The invention discloses an automobile part quality detection method and system based on artificial intelligence visual inspection, and belongs to the field of artificial intelligence machine visual inspection, and the method comprises the steps: firstly, carrying out the registration of a collected RGB image and a depth image, and extracting a part region through a saliency detection network; two-dimensional key points are extracted based on an RGB region image and are matched with key points of a three-dimensional model, an initial three-dimensional attitude is obtained by adopting a PnP algorithm, iterative registration is performed with the three-dimensional model in combination with a point cloud generated by a depth image, and a fine three-dimensional attitude is obtained. And calculating a geometric transformation matrix from the part to a standard front view attitude according to the attitude, and performing attitude correction on the RGB and depth region image. And then matching the corrected image with a standard template image by using a feature detection and matching network so as to correct the position of the detection window. And finally, the three-dimensional size of the part is calculated in the corrected detection window in combination with the depth value, and tolerance judgment is carried out. And the precision, the robustness and the automation level of online detection of the automobile parts can be obviously improved.
Owner:XIANYANG VOCATIONAL TECHN COLLEGE

Slope geological disaster detection system based on unmanned aerial vehicle multi-sensor image fusion

The invention relates to the technical field of geological disaster detection, in particular to an unmanned aerial vehicle multi-sensor image fusion slope geological disaster detection system which comprises a data acquisition module, a manifold registration module, a feature fusion module, a weight optimization module, a fusion execution module, a disaster detection module and the like. RGB images, thermal infrared images and laser point cloud data of a slope are collected through an unmanned aerial vehicle, the surface of the slope is modeled as a Riemannian manifold, and high-precision space registration of heterogeneous data is achieved; constructing a feature manifold based on a manifold learning method, and extracting and fusing multi-scale features; evaluating the information amount of different areas by adopting a differential entropy theory, and generating a self-adaptive weight distribution diagram; performing weighted fusion on the registered multi-source data to generate a fused image; geological disaster features such as cracks, abnormal vegetation and water seepage points on the surface of the slope are recognized based on the fused image, and high-precision recognition and early warning of the geological disaster of the slope are achieved.
Owner:咸阳市公路局

Multi-modal information fused steel pipe inner surface defect area segmentation method and system

The invention provides a multi-modal information fused steel pipe inner surface defect area segmentation method and system, and the method comprises the steps: obtaining a steel pipe inner surface defect RGB image and a depth map at the same time, and carrying out the preprocessing and marking; constructing a support data set and a query data set; the support and query image feature extractor is used for respectively extracting support image multi-scale aggregation features FS, support image multi-mode semantic features IS, query image multi-scale aggregation features FQ and query image multi-mode semantic features IQ by sharing the multi-mode feature extraction backbone network; the FS, the IS, the FQ and the IQ are input into a multi-feature fusion device, a graph semantic guide module is supported to generate class guide features FA by using the FS and the MS, a similar prior feature generation module is supported to generate similar prior features FM by using the IS, the IQ and the MS, the FA, the FM and the FQ are spliced on a channel dimension, and a fusion feature graph is output and decoded by a multi-mode decoder to obtain defect area segmentation output. The method can be used for segmenting the defect area on the inner surface of the steel pipe.
Owner:UNIV OF SCI & TECH BEIJING

Wild animal detection method fusing unmanned aerial vehicle thermal infrared image and visible light image

The invention discloses a wildlife detection method fusing an unmanned aerial vehicle thermal infrared image and a visible light image, and belongs to the field of small target wildlife identification, and the method comprises the following steps: S1, obtaining a preprocessed TIR-RGB image pair set; s2, an FDM-YOLO double-source target detection model improved based on YOLOv81 is constructed, and the improved FDM-YOLO double-source target detection model is trained based on the preprocessed TIR-RGB image pair set obtained in the step S1; and S3, inputting an image acquired in real time into the improved FDM-YOLO double-source target detection model trained in the step S2, and outputting a wild animal detection result. By adopting the wild animal detection method fusing the thermal infrared image and the visible light image of the unmanned aerial vehicle, high-precision, real-time and robust detection of a small target of a wild animal in a complex field environment is realized by improving the FDM-YOLO model.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Single image super-resolution reconstruction method and system based on wavelet transform and cross-domain feature fusion

The invention discloses a single image super-resolution reconstruction method and system based on wavelet transform and cross-domain feature fusion. According to the method, firstly, a low-resolution RGB image is mapped to a high-dimensional feature space through a shallow feature extraction module; performing up-sampling and discrete wavelet decomposition on the features by using a wavelet feature mixing module to obtain multi-band features; low-frequency and high-frequency depth features are respectively extracted through a double-branch structure, cross-domain fusion is realized by means of a deformable cross attention mechanism, and the feature expression ability is enhanced in combination with residual connection; and finally, reconstructing a high-resolution image through convolution, up-sampling and regularization processing. In the training process, a pixel-level loss function is adopted to optimize network parameters, the multi-frequency-domain feature sensitivity is effectively improved, texture and structure information is balanced, the image contrast, definition and structural integrity are improved, and high-quality real-time super-resolution reconstruction can be achieved.
Owner:HUNAN UNIV

Railway tunnel foreign matter detection method based on radar point cloud and RGB image fusion

The invention discloses a railway tunnel foreign matter detection method based on radar point cloud and RGB image fusion. The method comprises the following steps: acquiring a laser radar point cloud and an RGB image in a railway tunnel; obtaining an RGB image by using an adversarial network model GAN to judge the abnormal occurrence state of the foreign matter in the railway tunnel, and if the abnormal state of the foreign matter occurs, projecting the obtained laser radar point cloud to an image plane to obtain a dense depth map; processing through a sliding window to obtain data slices which are approximately square; and carrying out data feature extraction, fusion and foreign object target detection on the data slices of the RGB image and the dense depth map by using a fusion target detection network based on the OfficientNet-FPN, and outputting a foreign object target information detection result containing a target category, a position, a confidence coefficient and a distance. According to the railway tunnel foreign matter detection method based on radar point cloud and RGB image fusion, high precision, high efficiency and high reliability of railway tunnel foreign matter detection are realized.
Owner:NANJING PIONEER AWARENESS INFORMATION TECH CO LTD

Quality detection system based on three-dimensional point cloud scanning and RGB image fusion

The invention discloses a quality detection system based on three-dimensional point cloud scanning and RGB image fusion, relates to the technical field of industrial automatic detection, and aims to solve the problem of high quality detection error rate caused by the fact that a detection result is easily interfered by visual angle change and illumination conditions in an existing detection method. According to the method, the color point cloud model under the unified coordinate system is constructed through space calibration and point pixel level registration, and the structural information and texture features of the workpiece are effectively reserved. Therefore, the problem that the quality detection error rate is high due to the fact that the detected object has curvature dramatic change, a strong reflection area or frequent shielding is solved. According to the method, the detection accuracy is improved, meanwhile, the robustness and the operation efficiency are high, the industrial workpiece detection requirements with complex curved surface structures and high surface quality requirements can be met, and the actual application requirements of an industrial site for an intelligent surface detection system are met.
Owner:HARBIN INST OF TECH

Shielding area three-dimensional voxel reasoning method and system

The invention belongs to the technical field of three-dimensional mapping, and particularly relates to an occlusion area three-dimensional voxel reasoning method and system. The method comprises the following steps: S10, acquiring an RGB image containing a semantic segmentation result of an occlusion region and a corresponding depth map, and converting the depth map into an initial sparse three-dimensional voxel; and S20, constructing a three-dimensional voxel inference network, and inferring a three-dimensional voxel occupancy probability graph from the initial sparse three-dimensional voxels through the three-dimensional voxel inference network. The occlusion area can be reconstructed based on the monocular image; multi-format output and navigation interface butt joint are supported, and the method can be integrated to SLAM, 3D mapping or path search systems to serve as spatial feasibility constraints; the method can be used for querying whether any spatial position is passable; the method can be used for extracting a trafficability sub-graph from a designated area for path planning.
Owner:SHANGHAI UNIV

Mold three-dimensional high-precision modeling method and system based on multi-modal data fusion

PendingCN121304933AImage enhancementImage analysisPoint cloudGeometric primitive
The invention relates to the technical field of three-dimensional model design of molds, and provides a mold three-dimensional high-precision modeling method and system based on multi-modal data fusion, and the method comprises the steps: collecting an RGB image, a high dynamic range image and a structured light depth image of an entity mold, and carrying out the reconstruction to obtain a globally consistent dense point cloud; denoising and semantic segmentation are carried out on the point cloud, and a cavity function area, a core function area and a runner function area are recognized; performing differential parametric modeling based on an identification result: performing geometric primitive fitting on regular geometric features, performing NURBS curved surface reconstruction on a free-form surface, and performing swept volume or rotator fitting on a flow channel to obtain a parameterized geometry, a parameterized curved surface and a parameterized channel model; and applying boundary continuity constraint to fuse the parameterized geometry, the parameterized curved surface and the parameterized channel model, and constructing a hybrid parameterized three-dimensional model. According to the method, rapid reconstruction of a high-precision, high-semantic and full-parameterized three-dimensional model of an entity mold can be realized.
Owner:SHENZHEN HENGYIYUAN PLASTIC MOULD CO LTD

Unmanned aerial vehicle target detection method and system based on visual detection algorithm

The invention discloses an unmanned aerial vehicle target detection method and system based on a visual detection algorithm. The unmanned aerial vehicle target detection method comprises the steps of generating unmanned aerial vehicle cruise route data containing a GPS coordinate sequence through a path planning algorithm; collecting multi-frame street lamp RGB image data through a visible light camera; meanwhile, single-channel infrared thermal radiation image data of the corresponding space-time position are collected through a carried infrared thermal imaging module; establishing a mapping relation between the RGB image data and the infrared thermal radiation image data to form a multi-modal original data set D; preprocessing the multi-modal original data set D to obtain a standardized multi-modal data set; according to the standardized multi-modal data set, constructing an improved UAV-YOLO model for identifying a street lamp heat source; aiming at the problems that a street lamp heat source presents small-size hot spots (10-30 pixels) in an image and is mixed with other heat sources (an automobile, an air conditioner outdoor unit and the like) in an urban environment and the distinguishing accuracy of a traditional algorithm is low under the overlook angle of an unmanned aerial vehicle, the method improves the detection precision.
Owner:SHENZHEN DEFULIAO TECH CO LTD

Unmanned aerial vehicle power distribution network equipment identification method based on mutual neighbor density peak clustering

The invention discloses an unmanned aerial vehicle power distribution network equipment identification method based on mutual neighbor density peak value clustering, and relates to the technical field of power distribution network equipment inspection, and the method comprises the steps: obtaining an original RGB image of a power distribution line; outputting the contrast characteristic coefficient of each channel; constructing a saturation retention item; calculating a contrast feature retention item, constructing a total energy function, solving an optimal channel weight by adopting a discrete search strategy, and outputting an initial grayscale image; sequentially carrying out weighted guide filtering, morphological reconstruction and super-pixel segmentation operation; performing two-dimensional wavelet decomposition on the super-pixel segmented image, and extracting a feature vector; and calculating the local density and the relative distance, selecting a clustering center, and completing sample distribution based on the shared mutual neighbor similarity to realize a power distribution network equipment identification effect. According to the method, the problems of detail loss, noise interference, edge breakage, disordered classification of multi-scale equipment and the like under complex illumination are effectively solved, and the identification precision and efficiency of unmanned aerial vehicle inspection are effectively improved.
Owner:INFORMATION & COMMNUNICATION BRANCH STATE GRID JIANGXI ELECTRIC POWER CO

Unmanned aerial vehicle and bird target identification method and system based on multiple modes

The invention discloses an unmanned aerial vehicle and bird target identification method and system based on multiple modes, and belongs to the technical field of computer vision and target identification. The method comprises the steps of collecting an RGB image, an infrared image and a continuous frame image of a monitoring area, extracting visual features, thermal radiation features and motion features after preprocessing, fusing multi-modal features by adopting an adaptive weight fusion strategy, detecting a potential target through a multi-scale feature fusion detection head, and obtaining a target detection result. The multi-mode classification module judges the target category and triggers early warning; the system comprises a data acquisition module, a preprocessing module, a feature extraction and identification module, a decision and early warning module and a database management module. Through multi-modal fusion, a self-adaptive weight strategy, multi-scale detection and multi-branch classification, the method can effectively overcome the limitation of a single modal, improves the recognition accuracy, the small target detection capability and the environmental adaptability in a complex environment, and meets the requirements of real-time recognition and early warning.
Owner:SICHUAN ZHONGKE LANGXING PHOTOELECTRIC TECH CO LTD

Double-spectrum target detection method based on physical prior constraint

The invention relates to a double-spectrum target detection method based on physical prior constraints, and belongs to the technical field of computer vision and deep learning. The method comprises the following steps: acquiring an RGB image and an IR image after registration, and extracting RGB and IR global feature vectors; constructing a double-spectrum target detection model by adopting a double-branch convolutional neural network; carrying out physical constraint fusion on the RGB and IR features to construct a bimodal fusion feature vector and training a model; wherein the model comprises a double-branch backbone network, a physical prior constraint module, a physical fusion module, a physical enhancement UNet module and a multi-scale detection head; the thermal physical constraint module establishes temperature-radiation relation constraint based on Stefan-Boltzmann law, the material analysis module identifies 12 materials and establishes physical attribute mapping, and the physical consistency verification module performs sub-pixel level registration and multi-level consistency detection; a to-be-detected image is collected and input into the trained model to obtain a detection result. According to the invention, high-precision rapid nondestructive detection of a target in a complex environment is realized.
Owner:FUZHOU UNIV

Three-dimensional scene reconstruction method and system based on monocular depth estimation

The invention discloses a three-dimensional scene reconstruction method and system based on monocular depth estimation, and belongs to the technical field of computer vision and three-dimensional reconstruction, and the method comprises the steps: carrying out the multi-scale feature coding of a monocular RGB image through a mixed attention depth coding module, and obtaining the hierarchical depth feature representation; carrying out autoregression depth decoding through a self-adaptive edge perception depth decoding module to generate an initial depth map; a depth confidence map is calculated through a geometric consistency constraint optimization module and is fed back to a coding module for iterative optimization, and a refined depth map is output; and three-dimensional Gaussian ellipsoid scene representation is constructed through the Gaussian ellipsoid scene reconstruction module. According to the invention, high-precision depth estimation and high-quality three-dimensional reconstruction are realized by constructing a depth-coupled closed-loop cooperative system.
Owner:HARBIN INST OF TECH

Mobile terminal automatic measurement method and system based on YOLO key point detection and AR platform

The invention belongs to the technical field of intelligent measurement, and discloses a mobile terminal automatic measurement method and system based on YOLO key point detection and an AR platform, and the method comprises the steps: obtaining multi-modal data, and carrying out the preprocessing; performing target detection and key point positioning based on the RGB image to obtain two-dimensional key point coordinates; based on the depth image and the pose information, a three-dimensional space coordinate system is established, the two-dimensional key point coordinates are converted into the three-dimensional space coordinate system, and in the conversion process, a self-adaptive weight fusion mechanism is adopted to perform fusion on the multi-modal data to obtain a fusion result; a nonlinear error propagation control model is adopted to suppress error accumulation in the process of converting the two-dimensional coordinates into the three-dimensional coordinates; and calculating morphological parameters of the target object based on the converted three-dimensional key point coordinates. According to the invention, the problems of low measurement efficiency, poor precision, high equipment cost, insufficient cross-platform adaptation and the like in the prior art are solved, and high-precision, real-time and low-cost automatic measurement of various types of target objects at the mobile terminal is realized.
Owner:SHANDONG UNIV

Method for estimating plant biomass based on map multi-modal feature extraction and fusion

The invention discloses a method for estimating plant biomass based on map multi-modal feature extraction and fusion. The method comprises the following steps: acquiring a plant RGB image and a multi-spectral image; inputting the preprocessed RGB image and NIR wave band image into a double-model cooperation segmentation framework to realize image segmentation; binary image features, color features, texture features, reflectivity and the like are calculated, feature splicing is carried out, and high-dimensional features are constructed; carrying out dimension reduction on the high-dimensional features; and training a deep neural network through the effective features and the biomass to realize biomass estimation. According to the method, a zero sample learning-based double-model cooperation segmentation framework is utilized to realize accurate segmentation of a single plant on the premise that a large number of training sets are not needed; multi-modal feature information is extracted based on the segmented single plant image, an improved SHAP model is introduced to reduce the feature space dimension, and the inversion precision and the operation efficiency are improved while the information effectiveness is ensured; through a high-precision deep neural network model, rapid, lossless and accurate biomass acquisition is realized.
Owner:NANJING FORESTRY UNIV

Three-dimensional defect detection method of asymmetric knowledge distillation network based on dynamic background guidance

The invention discloses a three-dimensional defect detection method of an asymmetric knowledge distillation network based on dynamic background guidance, and mainly solves the technical problems that in an existing three-dimensional defect detection model, multi-modal feature fusion is insufficient, background interference affects detection precision, and real-time performance is insufficient. The implementation scheme is as follows: 1) acquiring a data set; 2) constructing a three-dimensional defect detection model; 3) constructing a loss function; 4) training a three-dimensional defect detection model; and 5) obtaining a three-dimensional defect detection result. According to the constructed three-dimensional defect detection model, extraction of RGB image features is realized through a multi-scale feature splicing method of a two-dimensional multi-scale feature extractor, and an information recombination function from space to channel dimension is realized through a spatial recombination down-sampling module; the foreground mask dynamic extraction module is used to calculate abnormal scores only in a foreground area so as to avoid background interference, the asymmetric feature fusion module is used to realize fusion of RGB image and 3D point cloud data features, and the asymmetric knowledge distillation module is used to train a student model and a teacher model respectively so as to improve the learning efficiency. Therefore, the student model learns to return the output of the teacher model only on the defect-free image data.
Owner:CENT SOUTH UNIV

Imaging method and system based on polarization imaging lens and visual collaborative optimization, and medium

The invention provides an imaging method and system based on a polarization imaging lens and visual collaborative optimization, and a medium, and the method comprises the steps: integrating a four-way polarization filter array on the surface of an imaging sensor, and synchronously collecting a four-channel polarized light intensity image through single-frame exposure based on the imaging sensor; obtaining a three-channel RGB image, and inputting the four-channel polarized light intensity image and the three-channel RGB image into a polarization-RGB fusion neural network; polarization features and RGB features are extracted, heterogeneous information coupling optimization is carried out, and fusion features are obtained; constructing a polarization-color joint decoding network, carrying out polarization and RGB information collaborative decoding on the fusion features, training a polarization decoding process and a material classification task together based on an end-to-end joint optimization strategy, and outputting a material classification result; high-efficiency acquisition of original polarization information on a hardware level is guaranteed through the polarization filtering array; and an end-to-end joint training strategy is adopted to realize optimal decoupling and fusion of perception information.
Owner:HANGZHOU HUICUI INTELLIGENT TECH CO LTD

Multi-level re-planning-based instruction execution method and system for agent with body

The invention belongs to the field of intelligent planning control, and relates to an intelligent agent instruction execution method and system based on multi-level re-planning. The method comprises the following steps: preliminarily generating a sub-target sequence according to a natural language instruction by utilizing a high-level planning review module based on a large language model, and continuously and dynamically correcting sub-targets in combination with real-time semantic feedback of a current environment; a middle-layer semantic search module is utilized to construct a multi-layer instance semantic map, and ordered search and accurate positioning of all related instances are realized through the multi-layer instance semantic map and common sense reasoning driving; a low-layer action error correction module is used for receiving the RGB image of the current view angle of the intelligent agent in real time, the feasibility of the next preset action is judged, and if potential failure actions are found, the potential failure actions are automatically replaced with safer and more appropriate actions, so that execution failure or collision is avoided. The task success rate and the adaptive capacity of the intelligent agent for executing the natural language instruction in the complex three-dimensional environment can be improved.
Owner:PEKING UNIV

Geometric perception key point-based category-level 6D attitude estimation method

ActiveCN121304789AImage analysisBiological modelsPattern recognitionScale estimation
The invention provides a category-level 6D attitude estimation method based on geometric perception key points, and relates to the technical field of attitude estimation.The method comprises the steps that on the basis that RGB images and point cloud features are fused, a dynamic key point proposing module is designed to generate key points which are consistent in category and geometrically adaptive, so that the adaptability to intra-class deformation and weak texture targets is enhanced; secondly, introducing a spatial geometric attention module to model a spatial structure relationship between key points so as to optimize key point distribution; and finally, point cloud reconstruction and scale regression are simultaneously realized through a geometric perception reconstruction and scale estimation module under the prior condition of no CAD model, and end-to-end optimization is carried out by utilizing multi-task loss. Experimental results on REAL275, CAMERA25 and House Cat6D data sets show that the method provided by the invention is superior to the existing method in multiple indexes such as IoU and attitude precision, and shows stronger generalization and robustness in a real complex scene.
Owner:LIAO NING GONG CHENG JI SHU DA XUE E ER DUO SI YAN JIU YUAN

Unmanned aerial vehicle image small target detection optimization method based on YOLOv8

The invention discloses an unmanned aerial vehicle image small target detection optimization method based on YOLOv8. According to the method, a small target RGB image of a high-altitude scene is collected through an unmanned aerial vehicle, and an improved YOLOv8 model is input for training. The main improvement comprises the following steps: firstly, applying SPD convolution on a P2 feature layer to enhance feature expression, and fusing a result with a P3 feature layer; secondly, an Omni-Kernel module based on CSP structure optimization is introduced, and the module processes features through three parallel branches: a local branch extracts details by using 1 * 1 deep convolution, a large-scale branch expands a receptive field by using a 63 * 63 convolution kernel in combination with bar convolution, and a global branch integrates DCAM and FSAM attention mechanisms to enhance long-range dependence. In the detection stage, the preprocessed image is processed by the modules, and the target category and position are output by the detection head. According to the method, the complex environment detection precision of the unmanned aerial vehicle is remarkably improved, and real-time detection requirements are met.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Workpiece grabbing and releasing system and method based on multi-modal data fusion and Kalman filtering tracking

The invention discloses a workpiece grabbing and releasing system and method based on multi-modal data fusion and Kalman filtering tracking. Comprising a hardware grabbing and releasing module, a mechanical arm control module, a target recognition module, a dynamic tracking module and a pose estimation module. Firstly, a depth camera collects RGB images and depth point cloud data of the surface of a conveying belt, then the position and the type of a workpiece are recognized, and three-dimensional point cloud of a target workpiece is extracted from the point cloud data according to the position of the workpiece. And the pose estimation module searches the stored model library for the point cloud of the template model matched with the target workpiece, takes the point cloud as an initial iteration value of the ICP algorithm, and outputs the optimal grabbing pose of the target workpiece. And the dynamic tracking module dynamically compensates the position of the target workpiece in the pose estimation process through a Kalman filtering algorithm. And the mechanical arm control module drives the mechanical arm to perform position adjustment according to a prediction result of the dynamic tracking module, and calculates an instruction for controlling the mechanical arm to complete the grabbing action according to the optimal grabbing pose.
Owner:COLLEGE OF SCI & TECH NINGBO UNIV

Method and device for analyzing uncertainty of livestock target detection model

The invention provides an uncertainty analysis method and device for a livestock target detection model. The method comprises the following steps: acquiring an RGB image of a livestock target; constructing an uncertainty analysis model of the livestock target detection model, and calling the uncertainty analysis model to detect the livestock target in the RGB image to obtain the category of the livestock target and a detection frame coordinate; and respectively calculating a cognitive uncertainty index and a random uncertainty index of the livestock target detection model according to the category and the coordinates of the detection frame. The method is used for solving the problems that in the prior art, an uncertainty analysis method is not comprehensive enough, uncertainty result attribution is lacked, and the uncertainty evaluation effect is poor.
Owner:AEROSPACE INFORMATION RES INST CAS

Robot action reasoning method and system based on Gaussian action field

The invention relates to a robot motion reasoning method and system based on a Gaussian motion field, and the method comprises the steps: inputting a sparse and uncalibrated multi-view RGB image, extracting mixed scene features through a visual Transform backbone network, constructing a dynamic Gaussian motion field, endowing each Gaussian unit with a learnable motion attribute, and carrying out the reasoning of the motion of a robot. And realizing synchronous modeling of scene geometry and motion evolution. The system further comprises a multi-modal query module, a point cloud registration module, a diffusion model optimization module and a closed-loop control module which are respectively used for current and future scene reconstruction, mechanical arm tail end clamping jaw action estimation, action sequence optimization and real-time feedback and model updating in the execution process. According to the method, through a unified space-time modeling and closed-loop control mechanism, the robot operation problem in dynamic, shielding and uncalibrated environments is effectively solved, and the method is suitable for various application scenes such as industrial automation and service robots.
Owner:TSINGHUA UNIVERSITY

Weld defect identification method based on dense connection convolutional network model

The invention discloses a weld defect identification method based on a dense connection convolutional network model, and the method specifically comprises the following steps: S1, constructing a dense connection convolutional network model, and embedding a coordinate attention module behind a transition layer of a convolutional network; s2, data acquisition and processing: acquiring an RGB image of the welding seam through an industrial camera, constructing a data set of the image, and performing image enhancement and standardization processing; s3, performing hyper-parameter optimization, and performing global optimization on the constructed model by adopting a Bayesian optimization algorithm; s4, performing model training and verification, and training a dense connection convolutional network model by using the optimized hyper-parameter combination; and S5, defect identification: inputting a to-be-detected welding seam image into the trained dense connection convolutional network model, and outputting a defect category and a positioning result. According to the method, the transition layer of the convolutional network is embedded into the coordinate attention module, so that the convolutional network model more accurately positions the welding seam position, and the detail features of the welding seam are extracted.
Owner:SHANGHAI DONGXIN SOFTWARE ENG CO LTD +2