Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3910 results about "Vision based" patented technology

Semi-supervised medical image segmentation method and system based on visual language model

SOLUTION: A semi-supervised medical image segmentation method based on a visual language model includes the steps of: obtaining a medical image; inputting an unlabeled image and a text description into a visual language model, and obtaining a text-guided mask based on obtained dense image embedding and text embedding; inputting a labeled image into a student model, and calculating supervised loss by using obtained labeled image prediction; respectively inputting the unlabeled image into the student model and a teacher model to obtain unlabeled image prediction and a pseudo label, merging the text-guided mask with the pseudo label, and calculating semi-supervised loss by using the merged pseudo label and unlabeled image prediction; and performing medical image segmentation by using a trained student model on the basis of the supervised loss and the semi-supervised loss.EFFECT: A target segmentation region can be accurately identified by using advantages of text descriptions.SELECTED DRAWING: Figure 1
Owner:SHANDONG UNIV

Collaborative robot end control method based on vision and related equipment thereof

The invention relates to the technical field of equipment control, and provides a vision-based collaborative robot end control method and related equipment thereof. Multi-modal dynamic anti-interference filtering and semantic fusion are carried out on original visual data to obtain three-dimensional semantic point cloud data, and coordinate system alignment is carried out on the three-dimensional semantic point cloud data according to a robot-based coordinate system to obtain a target pose parameter set; performing dynamic path planning on the target pose parameter set to obtain a joint space motion track, and performing inverse kinematics solution and motor instruction synthesis on the joint space motion track according to robot D-H model parameters to obtain a motor control instruction, and performing closed-loop feedback optimization on the motor control instruction according to the state data of the end effector to obtain a dynamic correction control instruction. According to the method, through full-link optimization from sensing, planning to execution, the operation precision and safety of the collaborative robot in an uncertain environment are remarkably improved.
Owner:DONGGUAN BESON ROBOTIC TECH CO LTD

Mechanical arm dynamic deviation correction method and system based on visual driving and medium

The invention discloses a mechanical arm dynamic deviation correction method and system based on visual driving and a medium, and relates to the technical field of mechanical arm control. The method comprises the steps that when the tail end of a mechanical arm enters a preset machining space, an integrated 3D visual sensor is triggered to collect 3D point cloud of a workpiece to be machined; after pose recognition is carried out on the point cloud, the offset is recognized according to the teaching pose and the actual pose, and the initial offset is output; calling the multi-dimensional perception data, performing fusion correction, and outputting a correction offset; performing interference correction through an offset compensation model, and outputting a target offset; parameter adjustment and optimization are carried out according to the target offset, and a joint angle adjustment instruction is output; and performing correction closed-loop feedback according to the updated pose data. The technical problems of precision errors and low efficiency caused by deviation in the operation process of the mechanical arm are solved, and the technical effects that through dynamic deviation correction and multi-sensor data fusion, the operation precision and efficiency of the mechanical arm are improved, and stable operation in a complex environment is ensured are achieved.
Owner:ZHUHAI DEXIN ZHONGCHUANG INTELLIGENT TECHNOLOGY CO LTD

Uncoupling robot control system and method based on multi-source visual fusion

The embodiment of the invention provides an unhooking robot control method based on multi-source visual fusion, which is applied to the technical field of robot control and comprises the following steps: acquiring an RGB image, a depth image, an infrared image and IMU data through a multi-source sensing system mounted at the tail end of a robot; carrying out feature fusion identification by adopting a double-branch neural network, and outputting the boundary contour of the lifting hook and the three-dimensional coordinates of the optimal grabbing point; the visual coordinates are unified to a robot base coordinate system through a registration correction mechanism; a Transform prediction model is constructed based on the visual and inertial signals, and future pose changes of the lifting hook are estimated; a feedforward control track is generated to counteract swing of the lifting hook, and track correction is carried out in combination with visual servo feedback; and a joint instruction is generated through path planning and inverse kinematics solution, and the mechanical arm is driven to complete precise unhooking operation. According to the method, the recognition precision, the anti-interference capability and the operation success rate of unhooking operation in complex illumination and dynamic environments are effectively improved.
Owner:ANHUI HUADIAN SUZHOU POWER GENERATION

Vision-based traditional Chinese medicinal material defect detection method

The invention relates to the technical field of traditional Chinese medicinal material defect detection, in particular to a traditional Chinese medicinal material defect detection method based on vision, which comprises the following steps: regularly acquiring a time sequence image of a traditional Chinese medicinal material sample, acquiring a multi-view image, carrying out pixel alignment on the time sequence image, and carrying out structured organization on the multi-view image according to a shooting direction to associate camera parameters; and generating a registration time sequence image sequence and a multi-view image set. According to the method, the time sequence images of the traditional Chinese medicine samples are collected regularly, pixel alignment is carried out, image offset caused by environment illumination fluctuation and equipment jitter is eliminated, and time-space consistency of dynamic variable quantity calculation is ensured. Structured organization is carried out on multi-view-angle images according to shooting directions, camera parameters are associated, a geometric constraint relation between view angles is established, and the problem that three-dimensional reconstruction precision is insufficient due to view angle isolation in a traditional method is solved.
Owner:CANGNAN COUNTY QIUSHI TRADITIONAL CHINESE MEDICINE INNOVATION RES INST

Industrial part alignment method and system based on visual analysis and storage medium

The invention relates to the technical field of image processing, and discloses an industrial part alignment method and system based on visual analysis and a storage medium. The method comprises the steps that a three-view camera collects an industrial part image, and preprocessing is carried out through gradient magnitude local contrast enhancement to obtain an enhanced image; performing hierarchical feature extraction to identify edge contours and key control points to form a multi-dimensional feature set; and establishing a dynamic reference coordinate system based on the feature set to obtain a part space attitude matrix. And the attitude deviation is compensated through Z-axis offset and rotation coupling error analysis. Posture adjustment is decomposed into a plurality of sub-stages, an alignment track is optimized by adopting a variable speed planning strategy, and accurate alignment of the parts is achieved. The problems that multi-view visual information fusion is insufficient, a special recognition algorithm for geometrical characteristics of the industrial parts is lacked, and Z-axis offset and rotation coupling error compensation is inaccurate in the posture adjustment process are solved, and the precision and stability of alignment of the industrial parts are improved.
Owner:BEIJING TIANYUAN 3D TECH CO LTD

Dynamic Latent Space Adaptation Based on Spatiotemporal Kernal Context for Multiscale Rendering

A system for dynamic latent space adaptation using spatiotemporal kernel context for multiscale rendering with hierarchical and Lorentzian autoencoders. The Spatiotemporal Kernel Estimator (SKE) analyzes media through motion field, temporal recurrence, frequency band, and scene semantics analyzers to generate adaptive kernel parameters encoding content-specific importance distributions. The system dynamically adapts latent manifold geometry by modifying metric tensor properties according to kernel context, enabling content-aware compression that allocates representational capacity based on visual significance. A multiscale cache implements kernel-adaptive retention policies prioritizing important regions. An adaptive renderer provides intelligent level-of-detail selection based on zoom level and kernel-estimated importance, optimizing processing allocation. The self-optimizing architecture continuously refines kernel context and geometric adaptation based on user interaction and performance feedback, achieving superior compression ratios and perceptual quality. Applications include bandwidth-efficient video streaming, virtual reality, scientific visualization, and cognitive video analytics requiring intelligent context-aware visual processing.
Owner:ATOMBEAM TECH INC

Smart park safety management method and system based on AI visual identification

The invention relates to the technical field of safety monitoring, in particular to a smart park safety management method and system based on AI visual identification, and the method comprises the following steps: obtaining a video stream captured by a monitoring camera of a smart park, extracting the movement track data of a person through AI, and carrying out the statistics of the stay duration and access times of a target object in each grid; and constructing retention heat distribution of the grid region to obtain a personnel trajectory retention heat value. According to the invention, through the data processing of the monitoring video stream and the combination of the gridding analysis of the personnel motion trail, the accurate capture of the abnormal cruise behavior is enhanced, and through the calculation of the change rate of the access times between the adjacent grids, the statistics of the round-trip condition of the target among the plurality of areas and the combination of the passing authority data, the abnormal screening is carried out. The method improves the recognition capability of abnormal movement modes, achieves the dynamic adjustment of regional safety risk levels, improves the response capability of park security, enables the recognition capability of abnormal behaviors to be enhanced, and reduces the false alarm rate.
Owner:SHAOGUAN XINGZHITIANXIA NETWORK TECH CO LTD

Intelligent welding forming method and system for steel heating radiator for green building

The invention discloses an intelligent welding forming method and system for a green building steel heating radiator, and the method comprises the following steps: carrying out the surface defect recognition of a steel heating radiator base material based on an AI visual inspection system, recognizing a qualified base material, and automatically matching the type of a welding material from a material database according to the material and thickness parameters of the qualified base material. Welding parameters are intelligently matched through AI visual inspection and a neural network model, a laser and friction stir hybrid welding process is combined, traditional manual operation is replaced, the welding efficiency and precision are improved, and the problems of uneven welding seams and the like are solved; welding data are analyzed in real time through a multi-mode AI model, parameters are dynamically adjusted, intelligent defect recognition and repair welding are achieved in cooperation with 3D visual inspection, and the quality stability is improved through whole-process monitoring; smoke dust is treated through an environment-friendly process, acid pickling is replaced with mechanical rust removal, efficient recycling of materials is achieved through waste recycling, a green manufacturing system is constructed, and the sustainable development requirement of green buildings is met.
Owner:SICHUAN AOFEIER TECHNOLOGY CO LTD

Unmanned aerial vehicle navigation reasoning method and system based on visual perception and large language model

The invention discloses an unmanned aerial vehicle navigation reasoning method and system based on visual perception and a large language model, and belongs to the field of unmanned aerial vehicle navigation and artificial intelligence, and the method comprises the steps: collecting an image of a target area through a camera carried by an unmanned aerial vehicle; pixel-level perception is carried out on the image through a pre-trained visual perception model, and visual information is converted into structured data; converting the structured data into key statements which can be processed by a large language model by adopting retrieval enhancement generation and combining with a predefined knowledge base; adopting a prompt project and thinking chain technology, and outputting a navigation instruction by the large language model according to the key statement and the current task; when a target of an unknown type is encountered, an active learning process is started, and real-time updating of an edge end knowledge base is realized through human intervention; according to the method, the problems of understanding and decision making of the unmanned aerial vehicle on a dynamic scene in a complex environment are effectively solved, data-driven intelligent support is provided for autonomous flight of the unmanned aerial vehicle, and the method has the advantages of being high in real-time performance, high in environmental adaptability and good in decision making reliability.
Owner:AVIC JINCHENG UNMANNED SYST CO LTD

Fabric defect intelligent detection method and system based on AI visual identification

The invention relates to the technical field of fabric detection, and discloses a fabric defect intelligent detection method and system based on AI visual identification. According to the method, motion blur is quantized through motion state data, optical blur caused by fabric motion is eliminated through deconvolution solution, so that motion interference in the fabric transmission process is processed in a targeted mode, self-adaptive balance of the deblurring capacity and the feature retention capacity is achieved, and then based on the optical interference principle, the deblurring capacity and the feature retention capacity are improved. Through a dynamic calibration system combining hardware-level real-time compensation and multi-dimensional optical parameter calibration, dynamic optical parameter calibration of primary correction data is realized, then fabric defect characterization data is extracted to accurately obtain defect features, and finally, a detection-production line control closed loop is constructed through a quality quantitative index and a comprehensive risk value, so that fabric defect detection is realized. The fabric defect detection precision can be improved, so that the problem of high defect missing detection and false detection rate caused by optical data distortion due to movement and environment interference in a traditional method is effectively solved.
Owner:HANGZHOU HANGSIYUE TEXTILE TECH CO LTD

Belt tearing detection method, device and equipment based on visual identification

The invention relates to the technical field of visual identification, and discloses a belt tearing detection method, device and equipment based on visual identification, and the method comprises the steps: carrying out the dual-channel image collection and preprocessing of the surface of a conveying belt, and obtaining a multi-channel preprocessing image; extracting a thermal difference feature and a texture structure feature of the multi-channel preprocessed image through a double-flow feature network, and inputting the thermal difference feature and the texture structure feature into a quaternion material deformation analysis model for deformation gradient tensor and invariant parameter calculation to obtain a tear feature description vector; performing hierarchical progressive identification and three-dimensional reconstruction analysis to obtain target tearing feature data; and risk assessment is carried out based on the target tear feature data to obtain a tear grade classification result, the interference of ambient temperature drift and non-uniform illumination is effectively eliminated, the calculation efficiency and accuracy of tear detection are improved, and false alarms and missing alarms are reduced.
Owner:宁夏京能宁东发电有限责任公司

Rail transit environment foreign matter intrusion detection method and system based on visual model

The invention belongs to the technical field of rail transit, and discloses a rail transit environment foreign matter intrusion detection method and system based on a visual model. The method comprises the following steps: firstly, extracting a suspicious region sub-graph in an original image through preprocessing, and respectively extracting a coarse-grained global feature vector and a fine-grained local feature vector by using an image encoder; meanwhile, text prompt information matched with the image content is generated, and a corresponding text feature vector is extracted; the multi-level similarity of the image feature vector and the text feature vector is calculated, and the respective weight fusion is combined to obtain a comprehensive abnormal score; and finally, dynamically modeling abnormal score distribution based on a Gaussian mixture model to realize foreign matter invasion judgment. According to the method, image and text multi-modal information is fused, the detection precision of foreign matters in a complex scene is improved through multi-scale feature association, the dynamic environment adaptability is enhanced by adopting an adaptive threshold strategy, and the problems that a traditional method is high in false detection rate and insufficient in generalization ability in a complex background of a railway are effectively solved.
Owner:CHINA RAILWAY DESIGN GRP CO LTD

Joint manipulator automatic calibration method and device based on visual system

The invention relates to the technical field of visual automatic calibration, and discloses a joint manipulator automatic calibration method and device based on a visual system. The method comprises the following steps: carrying out multi-view image acquisition on a joint manipulator of the five-axis robot to obtain original image data; performing occlusion region extraction on the original image data to obtain a joint occlusion region set; on the basis of the joint occlusion region set, disordered image splicing and mark point feature extraction are carried out on the original image data, and joint feature image data are obtained; constructing an adaptive calibration equation based on the joint feature image data, and solving a target transformation relation between a global visual coordinate system and a local visual coordinate system; multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot are executed according to the target transformation relation, and joint error correction parameters are generated, high-precision joint error correction parameter calculation is achieved, and the positioning precision and the movement precision of a joint manipulator of the five-axis robot are greatly improved.
Owner:深圳市远望工业自动化设备有限公司

Visual language model illusion suppression method based on adaptive dynamic attention intervention

The invention discloses a visual language model illusion suppression method based on self-adaptive dynamic attention intervention. The method is used for reducing the problem that a visual language model generates wrong associated information. The method comprises the steps that text input, visual input and historical response are acquired, and an attention accumulation vector is initialized; in the layer-by-layer calculation process of the language model, dynamically adjusting a non-normalized attention matrix, enhancing the weight of a visual sensitive attention head, and performing visual Token pruning in a deep network to optimize cross-modal information interaction; and finally, generating an output Token based on the adjusted attention mechanism, and carrying out loop iteration until a complete response is generated. According to the method, a method of combining text deviation correction through self-adaptive attention head modification and visual attention convergence-based Token pruning is adopted, so that the performance of the model in a multi-modal task is remarkably improved, and the illusion phenomenon caused by a language modal dominant reasoning process is effectively relieved.
Owner:ZHEJIANG UNIV OF TECH

Deep forgery detection method based on visual language model

The invention discloses a deep forgery detection method based on a visual language model, and relates to the field of image forensics. The deep forgery detection method based on the visual language model aims to combine multi-source information to improve the discrimination capability of the model on a real image and a generated image. The method comprises the following steps: firstly, extracting image features through an image encoder of a pre-trained CLIP model; meanwhile, a frequency domain enhanced counterfeit perception adapter is embedded in the image encoder to mine potential anomalies of counterfeit images in the image domain and the frequency domain. Secondly, a manual feature extraction module is provided, discriminative low-dimensional features are extracted from the four aspects of the edge, the texture, the frequency and the symmetry of the image, and the discriminative low-dimensional features are used as auxiliary information input in the forgery detection process, so that the robustness and the interpretability of the model are improved; meanwhile, the text cue words are converted into feature vectors through a text encoder of a pre-training CLIP model; and finally, the model predicts a forgery score by calculating the cosine similarity between the image features and the text features so as to realize the discrimination of the authenticity of the image. According to the method, the problem that the detection capability of the model on the cross-dataset is insufficient is effectively improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Pipe gallery pipeline key connection point leakage monitoring system and method based on AI vision

The invention relates to the field of computer vision and industrial safety monitoring, and discloses a pipe gallery pipeline key connection point leakage monitoring system and method based on AI vision, and the system comprises a data collection module, an intelligent analysis module, a 3D modeling module, a space-time verification module and an alarm prediction module. Real-time detection and three-dimensional accurate positioning of a leakage area are realized in combination with neural network optimization driven by physical simulation parameters; further through optical flow tracking and fluid mechanics verification, false alarm events of non-physical rules are screened out; and based on Bayesian decision and diffusion equation prediction, generating a leakage risk heat map and triggering graded alarm. The method solves the technical problems of low leakage detection precision, high false alarm rate and inaccurate positioning in a complex pipe gallery scene, has the advantages of high environmental adaptability, high anti-interference capability and prospective risk pre-judgment, and is suitable for intelligent safety monitoring of underground pipe galleries and petrochemical engineering pipelines.
Owner:天津东方泰瑞科技有限公司 +2

Packaging defect eliminating method based on visual inspection

The invention discloses a package defect elimination method based on visual inspection, and relates to the technical field of image processing and computer vision, and the method comprises the following steps: setting a polarized light source with a fixed angle on a package production line, and uniformly illuminating a target package area to enable the incident light polarization direction and the package surface reflection direction to form a deviation angle; an industrial camera provided with a polarization filter is used for carrying out image collection on a packaging target area, and the polarization direction of the filter is perpendicular to the polarization direction of a light source so as to restrain reflection interference to the maximum extent. Reflection interference is eliminated through polarized light, and the image quality is improved; by combining a neural network and feature fusion scoring, the defect recognition accuracy and adaptability are improved; a closed-loop elimination control and data tracing mechanism is introduced, the elimination precision is guaranteed, quality responsibility investigation is supported, and the intelligence and reliability of the packaging defect detection and elimination system are integrally improved.
Owner:SUZHOU GILGAME SMART TECH CO LTD

System and method for extracting three-dimensional gluing contour of shoe sole based on visual single-line laser

The invention relates to the technical field of computer vision and industrial automation, in particular to a shoe sole three-dimensional gluing contour extraction system and method based on vision single-line laser, and aims to solve the problems that virtual calibration target spots cannot be accurately generated based on shoe sole geometry, the positions and sizes of the target spots are difficult to determine by combining curvature extreme values and principal component analysis in the prior art, and the production cost is low. The problem that a double-branch deep learning model cannot be adopted to fuse feature prediction transformation, and the re-projection error is increased is solved; a virtual calibration target spot is automatically generated based on sole geometry through a feature fusion calibration module, a grid is generated through point cloud processing and Poisson reconstruction, the position and size of the target spot are determined by combining a curvature extreme value and principal component analysis, a corresponding relation is established by utilizing two-dimensional and three-dimensional feature matching, initial alignment is realized through ICP and re-projection error optimization, and the target spot position and size are determined. A double-branch deep learning model is adopted to be fused with feature prediction transformation, iterative optimization is carried out through space consistency errors, and re-projection errors are reduced.
Owner:ANHUI UNIV

Control method of cable for charging unmanned ship based on visual identification

The invention provides an unmanned ship charging cable control method based on visual identification, and relates to the technical field of data processing, and the method comprises the steps: collecting continuous image frames of an unmanned ship charging area, identifying the spatial displacement and inclination angle change of an unmanned ship, analyzing the attitude offset of the unmanned ship caused by sea waves, calculating a water surface disturbance factor, and calculating the water surface disturbance factor. The method comprises the following steps: identifying the position of a charging interface, obtaining an initial positioning coordinate, carrying out prediction analysis on a disturbance trend in a preset time window, predicting a position change range of the charging interface, obtaining prediction coordinate data, planning a butt joint path of a cable and the charging interface, obtaining a compensation path sequence, and generating a cable propulsion instruction. And collecting an area image of the charging interface to obtain path tracking image data, judging whether the butt joint of the charging interface is successful or not, and if not, dynamically correcting the compensation path sequence. According to the invention, the cable is controlled through visual identification to charge the unmanned ship.
Owner:TIMES TIANHAI TECHNOLOGY CO LTD

Robot hierarchical reinforcement learning variable impedance control method based on vision and touch

The invention discloses a robot hierarchical reinforcement learning variable impedance control method based on vision and touch, and belongs to the field of robot control. According to the method, an upper-layer planning system and a lower-layer variable impedance control system are included, the upper-layer planning system fuses information from a visual sensor and a tactile sensor, feature coding is conducted on a visual image through a variational self-coding model, tactile feedback and mechanical arm state information are combined, and an action strategy adapting to the current environment and corresponding impedance parameters are formulated. And the lower-layer variable impedance control system realizes variable impedance control of the robot based on a depth kinematics and dynamics model, and allows the robot to dynamically adjust the motion and contact force according to the sensed information when interacting with the environment, thereby improving the adaptability to the environment change and the task execution efficiency. According to the method, the success rate and execution efficiency of the robot for executing the complex assembly task in the unstructured environment are remarkably improved, and high robustness is shown.
Owner:UNIV OF SCI & TECH OF CHINA

Robot intelligent welding method and system

The invention relates to the technical field of welding automation and robot visual perception, and discloses a robot intelligent welding method and system.The method comprises the steps that a three-dimensional point cloud image is constructed through a laser triangulation method; filtering and enhancing the point cloud image, and extracting a weld point cloud set by adopting a U-Net semantic segmentation model; carrying out space path identification and attitude regression based on the set to obtain a track element of a position-attitude pair; fitting a path by adopting a cubic B-spline curve and interpolating to generate a continuous welding track; and in the welding process, visual feedback is combined, and proportional differential control logic is introduced for closed-loop track error compensation. Compared with the technical problem that in the prior art, weld joint path recognition and accurate tracking cannot be stably achieved under complex structures such as curved surfaces and variable cross sections, a perception control closed loop based on visual servo is constructed, high-precision automatic welding of the complex weld joint structures is achieved, and the track fitting precision is improved.
Owner:XUZHOU MINGJIE METAL TECHNOLOGY CO LTD

Roof leakage point intelligent positioning method and system based on machine vision

The invention provides a roof leakage point intelligent positioning method and system based on machine vision, and relates to the technical field of constructional engineering detection.The method comprises the steps that 1, roof image data and surface texture, temperature field and environment illumination parameters are collected through a visible light and infrared thermal imaging two-channel sensing device, a multi-dimensional data set is constructed, and a multi-dimensional data set is established; generating visual data through dynamic range correction based on the multi-dimensional data set; step 2, based on visual data, performing detection area division on the acquired image, extracting a local temperature difference gradient, a surface texture abrupt change rate and a liquid flow track parameter, and establishing a multi-index characteristic matrix; according to the method, the features of the roof image are extracted through machine vision, intelligent positioning, risk grade division and dynamic evolution tracking of roof leakage points are realized, and the leakage positioning accuracy is improved.
Owner:THE 12TH CONSTR GRP OF SHAANXI CONSTR ENG CO LTD

Industrial robot multi-station intelligent sorting and stacking method based on visual guidance

The invention provides an industrial robot multi-station intelligent sorting and stacking method based on visual guidance, which relates to the technical field of robot control, and comprises the following steps: determining a target station by acquiring multi-station real-time state information, calculating an optimal grabbing time window by using a graph neural network state predictor, and obtaining a target grabbing time window; and the placement strategy is dynamically adjusted based on the visual image and the force sensing data in the article stacking process. According to the multi-station intelligent stacking system, the sorting efficiency in a multi-station scene can be improved, the grabbing precision of dynamic objects is enhanced, and stable and reliable intelligent stacking operation is achieved.
Owner:BEIJING CYBERROBOT TECH CO LTD

Vision-language interaction-based remote sensing image open vocabulary segmentation method and system

The invention discloses a remote sensing image open vocabulary segmentation method and system based on vision-language interaction. According to the method, the advantages of a multi-modal large language model and a semantic segmentation network based on a visual basic model are cooperatively utilized, language-pixel two-way mapping is taken as a core, five types of marks of images, texts, categories, objects and segmentation are introduced as carriers of cross-modal information, and three types of cross-modal fine-grained information interaction modules are taken as bridges, so that the cross-modal information interaction is realized. Bidirectional mapping and alignment of fine-grained information of the multi-modal large language model and the semantic segmentation network are realized, the open vocabulary segmentation capability of the semantic segmentation model is improved, and the method can adapt to different remote sensing scenes and category definitions. The method has the following advantages: the method has high performance, strong generalization ability and good expansibility, can provide any category of semantic segmentation maps for unlabeled target domain images based on instructions, and has high application value in the aspects of urban planning, map making, disaster response and the like.
Owner:WUHAN UNIV

Indoor robot navigation method based on multi-modal feature fusion

The invention relates to an indoor robot navigation method based on multi-modal feature fusion. The method comprises the following steps: constructing a semantic map based on visual observation environment information; the method comprises the following steps: acquiring an RGB image of an indoor scene object, converting the RGB image into point cloud data, and preprocessing the RGB image and the point cloud data; image multilayer semantic features and point cloud features in the RGB image and the point cloud data are extracted respectively, and initial fusion is carried out; performing weighted fusion on the multi-layer semantic features and the point cloud features of the image by adopting space-channel-cross-modal multi-attention dynamic cooperation; and predicting a long-term target in a map space from top to bottom based on the fused feature map and the semantic map, and performing navigation path planning based on the current position and the long-term target. Under low-cost hardware configuration, the robustness of environment perception, the real-time performance of decision response and the usability of system integration in a dynamic complex environment are comprehensively improved, and the method is particularly suitable for application scenes such as indoor service robots.
Owner:SOUTHWEST JIAOTONG UNIV

Device based on vision and 2D laser fusion positioning

The invention relates to the technical field of positioning, in particular to a device based on vision and 2D laser fusion positioning, and the device comprises a horizontal distance measurement module which is horizontally installed on a robot body and is used for collecting contour distance measurement data of a motion plane; the vertical vision module is installed on the robot body in a top view mode and used for collecting visual feature data of a vertical space; the calculation processing module is configured to construct a laser contour map and a visual feature map which are aligned in a coordinate system based on the contour distance measurement data and the visual feature data, determine an initial pose of the robot according to a given position or an optimal position of self-global matching of the laser contour map and the visual feature map, generate a predicted pose in combination with a motion measurement value, and send the predicted pose to the robot. And iteratively executing dual-mode observation filtering updating of laser contour matching optimization and visual reprojection matching optimization by taking the predicted pose as a center. According to the invention, through deep integration of a heterogeneous sensor architecture and a computing platform, an industrial scene-oriented all-weather positioning solution is constructed.
Owner:HANGZHOU LANXIN TECH CO LTD

PCB defect detection method based on visual converter combined with conditional diffusion

The invention belongs to the technical field of computer vision and deep learning, and particularly relates to a PCB defect detection method based on combination of a visual converter and conditional diffusion. Comprising the following steps: constructing an unlabeled PCB image data set and carrying out data preprocessing and enhancement to obtain a preprocessed image; executing a self-supervised pre-training task on the preprocessed image to obtain a feature extraction network; based on a conditional diffusion model, generating a synthetic defect PCB image and a label thereof by using the features output by the feature extraction network and the defect type control vector; mixing the synthetic defect image with a small number of real defect images to construct a training set; performing training adjustment on the defect detection model by adopting the training set to obtain a trained defect detection model; performing PCB defect detection by using the trained defect detection model; according to the method, the robustness and the cross-domain generalization ability are remarkably improved, the missed detection risk is reduced, and the rapid and stable quality control requirement of the production line is met.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Semi-supervised image semantic segmentation method and system based on visual basic model

The invention provides a semi-supervised image semantic segmentation method and system based on a visual basic model, and the method comprises the steps: constructing a multi-task model which comprises a visual basic model and a depth estimation basic model, and the visual basic model is connected with a task solution head, an adapter parameter efficient fine tuning module and a multi-modal cross fusion module; the task solution head comprises a semantic segmentation head and a depth estimation head; extracting semantic hierarchy features and a depth feature map of the RGB image, performing cross attention fusion on the semantic hierarchy features and the depth feature map, and inputting obtained fusion features into a semantic segmentation head and a depth estimation head respectively; semi-supervised learning is adopted to train a multi-task model, only parameters in the adapter parameter efficient fine tuning module and the multi-modal cross fusion module are trained, and a multi-task loss function is adopted. The image semantic segmentation model obtained through training can improve semantic segmentation performance, reduce training cost and is suitable for different tasks.
Owner:SHANGHAI JIAOTONG UNIV

Robot navigation method and system based on visual identification

The invention provides a robot navigation method and system based on visual identification, and relates to the technical field of computer vision, firstly, a continuous visual information set of the surrounding environment of a robot is obtained, the continuous visual information set comprises environment visual images of different directions of the advancing direction of the robot, and the object appearance features and the spatial arrangement relation are recorded; then establishing association mapping between the continuous visual information set and a preset navigation area to obtain a double-association mapping result; generating a semantic anchor point navigation path based on a double-correlation mapping result, wherein the semantic anchor point navigation path comprises a semantic anchor point sequence from the current position to the target position; real-time visual semantic features in the navigation process are obtained through a real-time visual collection module, and a semantic anchor navigation path is dynamically updated; finally, a final navigation execution path is output according to the updated path, the robot is driven to complete navigation operation, and the navigation capability and the intelligent level of the robot in a complex dynamic environment are improved.
Owner:CHENGDU AEROSPACE KAITE ELECTROMECHANICAL TECH CO LTD