Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

214 results about "Multimodality image fusion" patented technology

Road crack detection method and system based on fused image

The invention relates to the technical field of road crack detection, in particular to a road crack detection method and system based on a fused image. The method comprises the following steps: acquiring road multi-source monitoring data including a visible light image, infrared thermal imaging data and laser radar point cloud data, and performing multi-modal image fusion and road three-dimensional point cloud reconstruction to generate a fused road image and road three-dimensional modeling data; performing crack curvature analysis based on the fused road image to generate crack curvature data; performing reflection crack contour recognition and positioning on the fused road image through the crack curvature data to generate reflection crack initial positioning data; obtaining road base material data; and performing reflection crack stress field reconstruction on the road area according to the reflection crack initial positioning data to obtain a reflection crack stress field. According to the invention, through multi-modal fusion, curvature identification, stress field modeling and crack channel analysis, the accuracy and strain of road reflection crack detection are improved.
Owner:BINHAI BAY BRANCH OF DONGGUAN CITY URBAN MANAGEMENT & COMPREHENSIVE LAW ENFORCEMENT BUREAU

Visual positioning method and system for precise connector assembly

The invention relates to the technical field of visual positioning, and provides a visual positioning method and system for precise connector assembly. Performing multi-mode HDR fusion and distortion correction on the original connector image to obtain a multi-layer fusion image matrix; and performing feature region coarse positioning in combination with a composite convolutional neural network to obtain a connector coordinate set, performing geometric feature fine positioning according to the connector coordinate set and the multilayer fusion image matrix to obtain a key point coordinate set, and performing hand-eye calibration and dynamic compensation on the key point coordinate set to generate a pose instruction. And performing vision-force control hybrid assembly processing on the pose instruction according to the sensor feedback data to obtain assembly completion state data. Through multi-modal image fusion, deep learning coarse positioning, precise geometric registration and dynamic cooperation of vision and force control, the speed, precision and stability of precise connector assembly are improved.
Owner:DONGGUAN HAIHONG INTELLIGENT TECH CO LTD

Multi-modal image fusion method based on modal self-adaption and modal interaction compensation

The invention provides a multi-modal image fusion method based on modal self-adaption and modal interaction compensation, and the method comprises the following steps: S1, obtaining a multi-modal image fusion data set, and obtaining a training data set through preprocessing; S2, analyzing the modal difference characteristics of infrared and visible light images, and evaluating the correlation characteristics of image pairs in different scenes; s3, capturing a cross-modal feature dependency relationship through a self-attention mechanism; s4, a differential feature extraction strategy is adopted, model parameters are optimized through iterative training, and multi-modal image fusion is completed; s5, a modal interaction compensation module is additionally arranged, unit dynamic balance common features and modal exclusive features are fused, feature complementation is achieved in channel and space dimensions, parameters of the modal interaction compensation module are optimized, the model is made to learn the optimal fusion weight of the multi-modal features in a self-adaptive mode, and multi-modal fusion image generation optimization is achieved through the model; according to the invention, multi-modal image fusion can be accurately and effectively carried out.
Owner:FUZHOU UNIV

Power transmission line foreign matter detection method and system based on multi-modal image fusion

The invention discloses a power transmission line foreign matter detection method and system based on multi-modal image fusion, and relates to the technical field of intelligent operation and maintenance and state monitoring of a power system, a lightweight Ev-Mama architecture is introduced into a backbone network part of YOLOv13, the model keeps relatively low calculation complexity, and meanwhile, the power transmission line foreign matter detection efficiency is improved. And the modeling capability of the method on the long-range dependency relationship and the global semantic information is obviously enhanced. Besides, by using the CDIDF module, the EVCS module and the MHSAA module, on the basis of increasing a small amount of calculation, the scale sensing ability, the space structure modeling ability and the context understanding ability of the model are effectively improved, and the performance bottleneck of a traditional YOLO series network in the aspects of processing small targets, shielding targets and cross-scale information fusion is effectively relieved.
Owner:KUNMING UNIVERSITY

High-voltage equipment defect positioning system and method based on multi-modal image fusion

The invention relates to the technical field of electrical equipment defect monitoring, in particular to a high-voltage equipment defect positioning system and method based on multi-modal image fusion, and the system comprises an ultraviolet triggering module, a modal synchronization module, a space positioning module, a structure recognition module and a fusion diagnosis module. According to the method, automatic calibration of a suspected defect area can be realized by performing quantitative threshold judgment on a corona discharge signal in an ultraviolet image, the consistency of three-mode data in space and time dimensions is ensured, and the accuracy of the defect area is improved by comparing three types of coordinates of a lead hot spot, boundary geometry and a spot center in the image and performing registration offset correction. The image feature alignment precision is improved, the insulator chain light spot distribution path is extracted, the continuous distribution length is calculated and corresponds to the pollution level standard, automatic identification of the pollution level is achieved, a multi-modal image fusion feature block is constructed, and a risk sorting label is given. And the defect positioning accuracy and the comprehensive evaluation capability on the defect property and the influence degree are effectively improved.
Owner:SHANGHAI ZIHONG OPTOELECTRONICS TECH CO LTD

Poultry behavior abnormity real-time monitoring system based on multi-modal image fusion

The invention discloses a poultry behavior abnormity real-time monitoring system based on multi-modal image fusion, particularly relates to the technical field of intelligent breeding behavior recognition, and is used for solving the problem of poor behavior monitoring accuracy under feather shielding. The method comprises the following steps: firstly, through combined perception of a visible light image and an infrared image, extracting a claw track interruption point and an anus temperature gradient direction, and realizing analysis of a motion state of a sheltered area; then, in combination with the heat conduction delay characteristic and the group movement direction, the flexion and extension angle of the covered leg joint is inverted, and a complete gait sequence is generated; thirdly, multi-source features such as gaits, temperature differences and body postures are fused, and a dynamic deviation model of the individuals relative to the mass center of the group is constructed; and finally, generating a stress behavior threshold curve according to the ground temperature and the ammonia gas concentration, outputting an abnormal behavior type and confidence, and realizing intelligent distinguishing of mechanical obstacles and adaptive behaviors.
Owner:JIANGSU INST OF POULTRY SCI

Photovoltaic module surface defect online detection method based on multi-modal image fusion

The invention discloses a photovoltaic module surface defect online detection method based on multi-modal image fusion, and relates to the technical field of photovoltaic module detection.The method comprises the steps that a visible light line-scan digital camera, a short-wave infrared camera, an EL line-scan digital camera and a laser speckle imaging module are installed to collect real-time data, an optical image is subjected to dynamic window median filtering, and the surface defect of the photovoltaic module is obtained; the method comprises the following steps: carrying out time moving average combined dynamic threshold segmentation on an infrared image and an EL image, carrying out image-guided outlier filtering on 3D point cloud data, judging a pixel defect type based on gray-temperature correlation and calculating a fusion weight, respectively inputting the optical image, the infrared image, the EL image and 3D point cloud into GhostNet, lightweight ViT, CNN and PointNet-Lite to extract features, and carrying out image fusion on the 3D point cloud data. And after global average pooling and splicing, inputting into a channel attention module to generate a modal weight, and after weighted fusion, obtaining a defect type through a full connection layer. According to the invention, the accuracy of defect detection is effectively improved.
Owner:CHINALAND SOLAR ENERGY +1

Ship multi-modal image fusion identification method based on graph neural structure alignment

The invention particularly relates to a ship multi-modal image fusion recognition method based on graph neural structure alignment, and the method comprises the following steps: 1, obtaining a ship multi-modal image, and constructing a backbone neural network to extract the features of the multi-modal image; 2, constructing graph nodes of the graph neural network, and generating an adjacent edge relationship; 3, for graph structures constructed in different modes, adopting a two-level graph attention mechanism to complete structure alignment; step 4, utilizing an optimal transmission mechanism to realize structure alignment between the infrared and visible light modal diagrams; step 5, feature re-injection is carried out to fuse space coordinates and global information, and the positioning and expression ability of node features is improved; and step 6, training the constructed ship multi-modal image fusion recognition network by adopting local feature alignment loss, graph-level semantic consistency loss and classification supervision loss. According to the method, the problem of alignment errors caused by inconsistency of infrared and visible light modal images is solved, the structure and semantic information of the infrared and optical images are fully fused, the accuracy and robustness of cross-modal target recognition are effectively improved, and the method is suitable for complex scenes such as multi-modal ship recognition.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Deep learning-based multimodal image fusion method for soft tissue photoacoustic / ultrasound imaging

The invention discloses a deep learning-based multimodal image fusion method for soft tissue photoacoustic / ultrasound imaging. Steps: an ultrasound-photoacoustic imaging device acquires photoacoustic and ultrasound images of human soft tissue and performs size normalization processing; an input spatial transformation module converts the images to the YCbCr space; an input pre-convolution module modifies the number of data channels; an input multi-scale feature extraction module extracts salient features from the source images; an input filter prediction module derives multi-scale filters; and an input filter fusion and adaptive enhancement module combines the input source images to obtain the final fused result. The invention has superior fusion performance compared to several traditional fusion methods and deep learning-based fusion methods, and more importantly, it exhibits excellent real-time performance. Furthermore, various modes of photoacoustic / ultrasound fusion extension experiments have verified the effectiveness of the method proposed in the invention.
Owner:HARBIN INST OF TECH +1

Visible light-thermal infrared image semantic segmentation method and system driven by plug-and-play prompt

The invention relates to the technical field of multi-modal image fusion perception and scene understanding, in particular to a plug-and-play prompt-driven visible light-thermal infrared image semantic segmentation method and system, and the method comprises the steps: respectively extracting the feature representation of an input visible light image and a thermal infrared image through a dual-branch LoRA fine-tuning image encoder; converting a segmentation mask generated by the existing visible light-thermal infrared image semantic segmentation model into unified prompt information, including bounding box prompt or point prompt; using a prompt encoder to encode prompt information into prompt embedding; the prompt-based mask decoder fuses the image feature representation and the prompt embedded feature through a spatial channel cross attention mechanism, and generates a final segmentation mask through a classification head; in the training stage, parameterized random disturbance is added to the prompt of the true value mask, and error distribution of the prediction prompt is simulated. According to the method, high-precision semantic segmentation adaptive to various segmentation models can be realized without retraining, and the segmentation robustness in a complex scene is remarkably improved.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Lung tumor CT image 3D segmentation method and system based on multi-modal image fusion

PendingCN120976547AImage enhancementImage analysis3d segmentationTissue invasion
The invention relates to the technical field of medical image processing, in particular to a lung tumor CT image 3D segmentation method and system based on multi-modal image fusion. The method comprises the following steps: acquiring a lung tumor CT image; determining a first texture feature based on the lung tumor CT image; identifying a lung tumor boundary by using the first texture feature; detecting a nodule protrusion area in the boundary of the lung tumor; obtaining tissue infiltration data from the nodule protrusion area; determining a second texture feature according to the tissue infiltration data; determining a tumor heterogeneity feature according to the first texture feature and the second texture feature; evaluating the potential malignancy degree by utilizing tumor heterogeneity characteristics; and dividing a tumor risk area of the lung tumor CT image based on the potential malignancy degree. According to the invention, accurate heterogeneity identification and risk region division of the lung tumor CT image are realized based on a medical image processing technology, and the accuracy of lung tumor 3D segmentation is improved.
Owner:THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL

Power transmission line forest fire early warning method based on multi-modal image fusion

The invention discloses a power transmission line forest fire early warning method based on multi-modal image fusion, and aims to solve the problems that when an existing forest fire early warning method is in butt joint with the actual operation condition of a power transmission line, multi-source information complementation is difficult to consider at the same time, false detection and missing detection are often caused by environmental influence, and the early warning timeliness sometimes cannot meet the real-time requirement of power grid operation. According to the method, a power transmission line forest fire early warning model based on multi-modal image fusion is constructed, and the model comprises preprocessing, multi-scale feature extraction, fire point detection and positioning, risk level definition and early warning level division. And training the model according to the visible light image, the infrared image and the radar image at each moment in the process from the beginning to the end of the mountain fire combustion. The invention belongs to the technical field of power transmission line forest fire early warning.
Owner:BAISHAN POWER SUPPLY COMPANY OF STATE GRID JILIN ELECTRONICS POWER COMPANY

Differential feature guided spatial channel multi-modal image fusion method and system

The invention discloses a difference feature guided spatial channel multi-modal image fusion method and system, and the method specifically comprises the steps: 1, carrying out the fusion of a visible light multi-modal image and an infrared light multi-modal image at a plurality of layers, and carrying out the complementation of the missing information between two modals; obtaining image features of the visible light multi-modal image and image features of the infrared light multi-modal image; step 2, performing interaction and fusion on the two image features in the step 1 on a channel level to obtain features corresponding to the fused visible light multi-modal image and features corresponding to the fused infrared light multi-modal image; 3, performing spatial fusion on the features obtained in the step 2 to obtain fused spatial features; according to the invention, interactive fusion can be realized between the visible light mode and the infrared light mode, so that the quality of the fused image is improved. According to the method, the disadvantage of a single-mode image in a downstream task can be effectively solved, and complementary information of two modes can be better utilized.
Owner:SOUTHEAST UNIV

Multi-modal image fusion model construction method

PendingCN121147033AImage enhancementBiological modelsPattern recognitionInteractive modeling
The invention discloses a multi-modal image fusion model construction method, and particularly relates to the technical field of image fusion model construction. Interactive modeling of a modal contribution imbalance coefficient and a confidence coefficient estimation anomaly coefficient is introduced in a multi-modal image fusion process; a causal association between confidence prediction fluctuation and modal weight extreme is converted into a quantifiable cross-modal complementarity degeneration index, and the cross-modal complementarity degeneration index is used as a core adjustment mechanism to dynamically optimize a fusion weight and training regularization, so that the problem that a single mode is excessively amplified or mistakenly weakened is effectively inhibited on a fusion strategy level; according to the method, the structure retention capability of the model in edge details and texture regions is remarkably improved, and artifacts and information loss caused by modal complementarity degradation are avoided.
Owner:CHICHAO NETWORK TECHNOLOGY (WUXI) CO LTD

Robot automatic grabbing path planning method based on visual identification

The invention discloses a robot automatic grabbing path planning method based on visual identification, and relates to the technical field of intelligent grabbing. The method comprises the following steps: acquiring multi-view visual data of a target scene, and generating a scene three-dimensional compact reconstruction model through a multi-modal image fusion algorithm; performing target detection and feature extraction on the model, and screening an optimal capture point in combination with a visual attention mechanism; constructing a dynamic environment obstacle probability map, updating an obstacle state through time sequence visual tracking, and quantifying an interference weight; an initial grabbing path is planned based on an improved fast expansion random tree algorithm, path smoothness constraints and robot joint movement limit parameters are introduced, and path nodes are optimized through a Bezier curve; visual servo feedback and path deviation prediction are fused, path parameters are corrected in real time, and a continuous movement track is generated. The method effectively adapts to the dynamic environment, gives consideration to path safety, smoothness and mechanical arm motion characteristics, and remarkably improves the grabbing success rate and operation reliability.
Owner:TIANJIN UNIV OF SCI & TECH

Orthopedic disease diagnosis auxiliary method and system based on multi-modal image fusion

The invention relates to the technical field of orthopedic image processing, and particularly discloses an orthopedic disease diagnosis assisting method and system based on multi-modal image fusion, and the method comprises the steps: S1, collecting skeleton image data and soft tissue image data through imaging equipment; s2, performing cross-modal feature interaction evaluation based on the skeleton image data and the soft tissue image data after preprocessing and feature extraction; s3, performing feature dimensionality reduction on the image data to be subjected to dimensionality reduction, evaluating and optimizing a feature dimensionality reduction result, and marking the image data to be subjected to dimensionality reduction after feature dimensionality reduction as low-dimensional image data; and S4, performing disease category probability prediction based on the low-dimensional image data, and generating a disease category probability distribution report. The problems of excessive redundant information, poor feature interaction effect and low disease category probability prediction accuracy in the existing multi-modal image fusion technology are solved.
Owner:AFFILIATED HOSPITAL OF GUILIN MEDICAL UNIV

Underground structure leakage intelligent detection system based on multi-modal image fusion and deep learning recognition

The invention discloses an underground structure leakage intelligent detection system based on multi-modal image fusion and deep learning recognition. The system comprises an image acquisition and preprocessing module, a significance guide image fusion module, a bimodal target detection module and a feature fusion and output module. Compared with the prior art, the method has the following advantages: saliency guidance, channel attention and multi-modal joint training are combined, an information closed loop is constructed by a triple mechanism, and the image fusion quality is improved; bimodal parallel recognition and ANN fusion judgment are adopted to adapt to the image degradation condition in a complex environment, and the recognition accuracy of a weak signal area is improved; the system can be deployed in various underground structure scenes such as subways, tunnels, underground garages and pipe galleries, and is compatible with various hardware terminals; a lightweight feature extraction and rapid fusion module is provided, and the real-time processing requirement of edge computing nodes is met; a closed-loop detection system of image acquisition, fusion enhancement, depth identification and intelligent output is formed, the overall efficiency is high, and the false detection rate is low.
Owner:HARBIN INST OF TECH

Cross-modal-based VMama medical image fusion method and system combining packet ACmix convolution and selective clustering

The invention relates to the technical field of medical image and artificial intelligence crossing, and discloses a cross-modal-based VMama medical image fusion method and system combining grouped ACmix convolution and selective clustering, and the system comprises an input module, a preprocessing module, a feature extraction module, a multi-scale fusion module, and an output reconstruction module. According to the scheme, cross-modal attention is introduced into a visual state space model for the first time, a new multi-modal image fusion framework LMACV is obtained, the framework performs linkage optimization on ACmix and VMamba structures in a cross-modal medical image for the first time, feature reconstruction efficiency is enhanced through a selective clustering mechanism, and MSE and PSNR are remarkably superior to existing methods such as MPCT, FATFusion and MATR. By fusing a convolutional network and state space modeling, the network can effectively capture local texture details, and meanwhile, long-distance semantic association is reserved; besides, linear state updating and cross-modal dynamic alignment of attention guidance are effectively realized, and compared with the latest MPCT algorithm, the fusion speed is improved by 37.5%.
Owner:THE SECOND AFFILIATED HOSPITAL OF CHONGQING MEDICAL UNIV

Infrared and visible light image fusion method based on text semantic consistency guidance

The invention provides an infrared and visible light image fusion method based on text semantic consistency guidance, which relates to the technical field of multi-modal image fusion, and comprises the following steps: respectively carrying out fine-grained text semantic generation on infrared and visible light images and mapping the infrared and visible light images to a unified embedding space; bidirectional compensation and enhancement are carried out on text semantics through a cross-modal attention mechanism, and unified text semantic priori is constructed; a structure-intensity decoupling double-branch encoder is adopted, and structure texture features of visible light and intensity significant features of infrared light are extracted respectively; under the prior guidance of text semantics, the bimodal visual features are aligned in a shared semantic space through explicit semantic consistency constraint and implicit semantic distribution consistency constraint; and finally, taking the text semantic priori as a global modulation signal, and carrying out adaptive weighted fusion and decoding on the aligned features to generate a fused image. The problems that an existing method is insufficient in semantic modeling and poor in fusion result consistency are effectively solved.
Owner:XIAMEN UNIV OF TECH

Multi-modal image fusion method based on hierarchical semantic richness

The invention relates to a multi-modal image fusion method based on hierarchical semantic richness, and belongs to the technical field of multi-modal image fusion. According to the method, global information exchange is balanced by adopting multi-scale feature aggregation and redistribution, meanwhile, fusion and segmentation tasks are dynamically bridged, a progressive semantic dense injection strategy is introduced, and global semantics are injected into highly consistent infrared features through dense connection, so that the global semantic fusion is realized. And then the semantic-infrared mixed features are transmitted to visible light features. Secondly, two types of feature fusion modules are introduced, one is to use a cross-modal attention mechanism to carry out more comprehensive feature fusion, the other is to use semantic features as third input to enhance semantic representation of image fusion, and feature fusion is realized in a complex scene by dynamically balancing global semantic consistency and fine-grained local detail representation. According to the method, layering and enriching of semantic information are achieved through strategies such as semantic collection, distribution and injection, and the fusion visual effect and downstream perception performance are enhanced.
Owner:MINJIANG UNIVERSITY

Memory-enhanced deep unfolding multimodal image fusion method with enhanced downstream tasks

The application discloses a memory reinforcement deep unfolding multi-modal image fusion method with downstream task enhancement, which comprises the following steps: collecting infrared images and visible light images, and dividing training set and test set; establishing an optimization target and solving, to obtain an iterative formula; using a neural network to replace a proximal operator in the iterative formula, to obtain a neural network structure; sending the training set into the neural network to obtain a fusion picture, and calculating a total loss according to the fusion picture, the infrared image and the visible light image; updating parameters of the neural network according to the total loss to obtain an updated neural network; inputting the infrared image and the visible light image of the test set into the updated neural network to obtain a fusion image. The application makes the fused image have characteristics easy to be distinguished by a downstream task network, and can realize the best performance on a data set, while performance and interpretability are taken into account, and the application has rationality and applicability.
Owner:XI AN JIAOTONG UNIV

Neurosurgery stereotactic operation positioning system based on multi-modal image fusion

ActiveCN120694747AImage analysisGeometric image transformationStereotactic surgeryNeurosurgery
The invention relates to the technical field of stereotactic surgery, and discloses a neurosurgery stereotactic surgery positioning system based on multi-modal image fusion, which comprises a multi-modal image acquisition module, an intelligent registration fusion module, a three-dimensional modeling module, a surgery navigation engine, an augmented reality interface and a dynamic calibration module, a cross-modal elastic registration module is arranged, when multi-modal image fusion is carried out, the spatial matching degree of a functional image and a structural image is detected in real time through a nonlinear deformation compensation algorithm, and when registration deviation exists between white matter fiber bundles and tumor boundaries, an elastic deformation compensation field is dynamically generated by a system, so that the positioning accuracy of a neural functional area is improved; and by arranging the dynamic drift correction module, the brain tissue displacement change is perceived in real time based on laser surface scanning during deep target navigation, precise correction of the pose of the surgical instrument is realized, and the positioning reliability of the deep target is ensured.
Owner:ZHEJIANG RUICHUANG PRECISION MEDICAL TECH CO LTD

Infrared and visible light image fusion method based on multi-semantic deep collaboration

The invention provides an infrared and visible light image fusion method based on multi-semantic deep collaboration. The method mainly solves the problem that an existing method cannot fully integrate text modals and global consistency between image fusion and downstream tasks. Comprising the following steps: 1) constructing a dual-task parallel network structure, and efficiently establishing a deep correlation between image fusion and a downstream segmentation task; 2) designing a multi-semantic deep collaboration module, realizing effective fusion of multi-modal information by deeply integrating text features, pixel-level features and segmented semantic features, and meeting semantic requirements of downstream tasks; 3) guiding an image fusion and segmentation task by using deeper and fine-grained semantic information in a text mode, and enhancing semantic consistency between a fusion result and a downstream task; and 4) inputting the obtained multiple semantic features into a fusion decoder to generate a final image fusion result. The semantic comprehension and visual perception capabilities of the model can be effectively enhanced, and the multi-modal image fusion performance is improved.
Owner:XIDIAN UNIV

Battery cell appearance detection method and device based on multi-modal image fusion

The invention provides a battery cell appearance detection method and device based on multi-modal image fusion, and the method comprises the steps: obtaining a single-channel gray-scale map and a single-channel height map as input, carrying out the validity verification, generating a target fusion image through multi-modal fusion, and constructing a training set to train a deep learning model, and finally, accurate detection and result output of the appearance defects of the battery cell are realized. According to the invention, through combination of multi-modal image fusion and deep learning, the cell appearance detection precision and efficiency are significantly improved, and efficient and accurate intelligent quality control is realized.
Owner:GUANGDONG LYRIC ROBOT INTELLIGENT AUTOMATION CO LTD

Underground infrared and visible light image fusion method

The invention provides an underground infrared and visible light image fusion method, and particularly relates to the technical field of underground environment multi-mode image fusion, and the method comprises the steps: inputting a visible light image into a visible light image encoder, dynamically compensating a low-illumination region through an illumination feature enhancement module, and generating an illumination enhancement feature map; an infrared image is input into the infrared image encoder, and an infrared feature map is generated through the feature alignment module. The two features are input into a double-branch feature extraction module, visible light global semantic features are extracted through global branches, infrared local detail features are extracted through local branches, and a fusion feature graph is generated through splicing. And inputting the fused feature map into a decoder, and generating a fused image through channel compression and multi-layer up-sampling reconstruction. A multi-task loss function is adopted to jointly train an encoder and a decoder, feature consistency loss, illumination smoothness loss and edge retention loss are included, and the target recognition precision in a complex scene is improved.
Owner:XIAN UNIV OF SCI & TECH

Visual detection-based earphone appearance shell appearance defect detection method

The invention relates to an earphone appearance shell defect detection method based on visual detection, and belongs to the technical field of computer vision and artificial intelligence. The method comprises the following steps: acquiring illumination environment parameters of an earphone shell through a self-adaptive illumination adjustment module, adjusting light supplementing intensity and angle, and generating a preliminary imaging data sequence; performing multi-modal image fusion on the data, extracting texture, edge and morphological features of the earphone shell, and generating a multi-modal feature image; inputting the image into a deep learning detection model, and carrying out defect classification and positioning; according to the defect detection result and a preset grading standard, evaluating the defect grade and severity, and generating a defect grading report and a dynamic decision parameter; and finally, carrying out optimization and online updating on the deep learning model through a model training module. According to the method, the precision and efficiency of earphone shell defect detection are effectively improved, various defects can be automatically identified and classified, the severity of the defects can be evaluated in real time, a detection system is continuously optimized through adaptive learning, the method adapts to a complex production environment, and the quality control level in the production process is improved.
Owner:HUIZHOU TIANQI ELECTRONIC TECHNOLOGY CO LTD

Chip surface defect detection method based on multi-modal image fusion

The invention relates to the technical field of image processing, in particular to a chip surface defect detection method based on multi-modal image fusion. The method comprises the following steps: firstly, acquiring three-dimensional surface geometric data and dynamic thermal response data of a chip; performing physical model correction on the dynamic thermal response data based on the three-dimensional surface geometric data to generate a thermal conductivity surface graph; then, guiding a scanning acoustic microscope to carry out sub-surface detection on the chip based on an abnormal region identified by the thermal conductivity surface graph so as to obtain acoustic image data, fusing the three-dimensional surface geometric data, the thermal conductivity surface graph and the acoustic image data in a three-dimensional coordinate system, and constructing a surface-sub-surface integrated three-dimensional model of the chip; then, analyzing the integrated three-dimensional model by adopting a physical information neural network embedded with physical constraints, and outputting a defect type identification result; the physical accuracy and the detection rate of chip surface defect detection can be improved.
Owner:GUANGDONG CANEN PRECISION MASCH CO LTD

Skin lesion area reconstruction system based on multi-modal image fusion

The invention relates to the technical field of skin lesion area reconstruction, in particular to a skin lesion area reconstruction system based on multi-modal image fusion. The system comprises a deep multi-scale feature extraction module, a skin thermal anomaly feature extraction module, a hemodynamic feature extraction module, a depth and boundary feature extraction module, a skin surface feature extraction module, a fluorescence intensity distribution and morphological feature extraction module, a first feature fusion module, a second feature fusion module and a lesion area reconstruction module. According to the method, appearance, heat, blood flow, structure and surface features are extracted from the multi-modal data, then the multi-modal image data are integrated, finally, comprehensive representation and reconstruction of skin lesion appearance, function and structure information are achieved according to the synergistic effect of the multi-modal information, and the accuracy and clinical reference value of lesion area reconstruction are effectively improved.
Owner:HANGZHOU THIRD PEOPLES HOSPITAL (HANGZHOU HUIMIN HOSPITAL HANGZHOU THIRD AFFILIATED HOSPITAL OF ZHEJIANG UNIV OF TRADITIONAL CHINESE MEDICINE)

Multi-modal image fusion method based on dynamic pseudo supervision and semantic guidance

The invention provides a multi-modal image fusion method based on dynamic pseudo supervision and semantic guidance, belongs to the technical field of image fusion, and is used for realizing high-quality fusion of infrared and visible light images. The method comprises the following steps: firstly, taking a fusion quality index as guidance, adaptively generating and iteratively updating a pseudo-supervision image, and realizing dynamic alignment of a model training target and a fusion evaluation index; secondly, decomposing the characteristics of the input image into high-frequency and low-frequency components, and respectively carrying out multilayer convolution and channel attention enhancement; thirdly, cross-modal semantic embedding features of the infrared image, the visible light image and the pseudo-supervision image are utilized to calculate semantic similarity, affine modulation parameters are generated to perform semantic alignment and weight adjustment on the features, and the consistency of multi-modal feature fusion and target saliency are enhanced; according to the method, the feature alignment of fusion optimization and semantic guidance with consistent evaluation can be realized under the condition of no manual annotation, and the detail definition, the structural integrity and the visual perception quality of a fusion result are effectively improved.
Owner:DALIAN UNIV +1

Single photon and visible light image fusion method and system based on multi-scale Markov random field model

The invention provides a single photon and visible light image fusion method and system based on a multi-scale Markov random field model, and belongs to the technical field of multi-modal image fusion. In the long-distance target ranging and imaging process, a multi-sensor fusion method is used, the high resolution advantage of a visible light intensity image is utilized, the problems that a single-photon laser radar is small in point cloud density and low in image resolution are effectively solved, and the sensing capacity of the single-photon laser radar for scene target information is effectively enhanced; by using the fusion method based on the multi-scale Markov random field model, the structural consistency and the anti-interference capability are enhanced, global structural deviation or local overfitting under a single scale is effectively avoided, and the method is suitable for image fusion under multi-modal and complex scenes.
Owner:BEIJING CHANGCHENG INST OF METROLOGY & MEASUREMENT AVIATION IND CORP OF CHINA