Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

143 results about "Multimodality image fusion" patented technology

Power transmission line foreign matter detection method and system based on multi-modal image fusion

The invention discloses a power transmission line foreign matter detection method and system based on multi-modal image fusion, and relates to the technical field of intelligent operation and maintenance and state monitoring of a power system, a lightweight Ev-Mama architecture is introduced into a backbone network part of YOLOv13, the model keeps relatively low calculation complexity, and meanwhile, the power transmission line foreign matter detection efficiency is improved. And the modeling capability of the method on the long-range dependency relationship and the global semantic information is obviously enhanced. Besides, by using the CDIDF module, the EVCS module and the MHSAA module, on the basis of increasing a small amount of calculation, the scale sensing ability, the space structure modeling ability and the context understanding ability of the model are effectively improved, and the performance bottleneck of a traditional YOLO series network in the aspects of processing small targets, shielding targets and cross-scale information fusion is effectively relieved.
Owner:KUNMING UNIVERSITY

Poultry behavior abnormity real-time monitoring system based on multi-modal image fusion

The invention discloses a poultry behavior abnormity real-time monitoring system based on multi-modal image fusion, particularly relates to the technical field of intelligent breeding behavior recognition, and is used for solving the problem of poor behavior monitoring accuracy under feather shielding. The method comprises the following steps: firstly, through combined perception of a visible light image and an infrared image, extracting a claw track interruption point and an anus temperature gradient direction, and realizing analysis of a motion state of a sheltered area; then, in combination with the heat conduction delay characteristic and the group movement direction, the flexion and extension angle of the covered leg joint is inverted, and a complete gait sequence is generated; thirdly, multi-source features such as gaits, temperature differences and body postures are fused, and a dynamic deviation model of the individuals relative to the mass center of the group is constructed; and finally, generating a stress behavior threshold curve according to the ground temperature and the ammonia gas concentration, outputting an abnormal behavior type and confidence, and realizing intelligent distinguishing of mechanical obstacles and adaptive behaviors.
Owner:JIANGSU INST OF POULTRY SCI

Deep learning-based multimodal image fusion method for soft tissue photoacoustic / ultrasound imaging

The invention discloses a deep learning-based multimodal image fusion method for soft tissue photoacoustic / ultrasound imaging. Steps: an ultrasound-photoacoustic imaging device acquires photoacoustic and ultrasound images of human soft tissue and performs size normalization processing; an input spatial transformation module converts the images to the YCbCr space; an input pre-convolution module modifies the number of data channels; an input multi-scale feature extraction module extracts salient features from the source images; an input filter prediction module derives multi-scale filters; and an input filter fusion and adaptive enhancement module combines the input source images to obtain the final fused result. The invention has superior fusion performance compared to several traditional fusion methods and deep learning-based fusion methods, and more importantly, it exhibits excellent real-time performance. Furthermore, various modes of photoacoustic / ultrasound fusion extension experiments have verified the effectiveness of the method proposed in the invention.
Owner:HARBIN INST OF TECH +1

Differential feature guided spatial channel multi-modal image fusion method and system

The invention discloses a difference feature guided spatial channel multi-modal image fusion method and system, and the method specifically comprises the steps: 1, carrying out the fusion of a visible light multi-modal image and an infrared light multi-modal image at a plurality of layers, and carrying out the complementation of the missing information between two modals; obtaining image features of the visible light multi-modal image and image features of the infrared light multi-modal image; step 2, performing interaction and fusion on the two image features in the step 1 on a channel level to obtain features corresponding to the fused visible light multi-modal image and features corresponding to the fused infrared light multi-modal image; 3, performing spatial fusion on the features obtained in the step 2 to obtain fused spatial features; according to the invention, interactive fusion can be realized between the visible light mode and the infrared light mode, so that the quality of the fused image is improved. According to the method, the disadvantage of a single-mode image in a downstream task can be effectively solved, and complementary information of two modes can be better utilized.
Owner:SOUTHEAST UNIV

Multi-modal image fusion model construction method

PendingCN121147033AImage enhancementBiological modelsPattern recognitionInteractive modeling
The invention discloses a multi-modal image fusion model construction method, and particularly relates to the technical field of image fusion model construction. Interactive modeling of a modal contribution imbalance coefficient and a confidence coefficient estimation anomaly coefficient is introduced in a multi-modal image fusion process; a causal association between confidence prediction fluctuation and modal weight extreme is converted into a quantifiable cross-modal complementarity degeneration index, and the cross-modal complementarity degeneration index is used as a core adjustment mechanism to dynamically optimize a fusion weight and training regularization, so that the problem that a single mode is excessively amplified or mistakenly weakened is effectively inhibited on a fusion strategy level; according to the method, the structure retention capability of the model in edge details and texture regions is remarkably improved, and artifacts and information loss caused by modal complementarity degradation are avoided.
Owner:CHICHAO NETWORK TECHNOLOGY (WUXI) CO LTD

Robot automatic grabbing path planning method based on visual identification

The invention discloses a robot automatic grabbing path planning method based on visual identification, and relates to the technical field of intelligent grabbing. The method comprises the following steps: acquiring multi-view visual data of a target scene, and generating a scene three-dimensional compact reconstruction model through a multi-modal image fusion algorithm; performing target detection and feature extraction on the model, and screening an optimal capture point in combination with a visual attention mechanism; constructing a dynamic environment obstacle probability map, updating an obstacle state through time sequence visual tracking, and quantifying an interference weight; an initial grabbing path is planned based on an improved fast expansion random tree algorithm, path smoothness constraints and robot joint movement limit parameters are introduced, and path nodes are optimized through a Bezier curve; visual servo feedback and path deviation prediction are fused, path parameters are corrected in real time, and a continuous movement track is generated. The method effectively adapts to the dynamic environment, gives consideration to path safety, smoothness and mechanical arm motion characteristics, and remarkably improves the grabbing success rate and operation reliability.
Owner:TIANJIN UNIV OF SCI & TECH

Cross-modal-based VMama medical image fusion method and system combining packet ACmix convolution and selective clustering

The invention relates to the technical field of medical image and artificial intelligence crossing, and discloses a cross-modal-based VMama medical image fusion method and system combining grouped ACmix convolution and selective clustering, and the system comprises an input module, a preprocessing module, a feature extraction module, a multi-scale fusion module, and an output reconstruction module. According to the scheme, cross-modal attention is introduced into a visual state space model for the first time, a new multi-modal image fusion framework LMACV is obtained, the framework performs linkage optimization on ACmix and VMamba structures in a cross-modal medical image for the first time, feature reconstruction efficiency is enhanced through a selective clustering mechanism, and MSE and PSNR are remarkably superior to existing methods such as MPCT, FATFusion and MATR. By fusing a convolutional network and state space modeling, the network can effectively capture local texture details, and meanwhile, long-distance semantic association is reserved; besides, linear state updating and cross-modal dynamic alignment of attention guidance are effectively realized, and compared with the latest MPCT algorithm, the fusion speed is improved by 37.5%.
Owner:THE SECOND AFFILIATED HOSPITAL OF CHONGQING MEDICAL UNIV

Infrared and visible light image fusion method based on text semantic consistency guidance

The invention provides an infrared and visible light image fusion method based on text semantic consistency guidance, which relates to the technical field of multi-modal image fusion, and comprises the following steps: respectively carrying out fine-grained text semantic generation on infrared and visible light images and mapping the infrared and visible light images to a unified embedding space; bidirectional compensation and enhancement are carried out on text semantics through a cross-modal attention mechanism, and unified text semantic priori is constructed; a structure-intensity decoupling double-branch encoder is adopted, and structure texture features of visible light and intensity significant features of infrared light are extracted respectively; under the prior guidance of text semantics, the bimodal visual features are aligned in a shared semantic space through explicit semantic consistency constraint and implicit semantic distribution consistency constraint; and finally, taking the text semantic priori as a global modulation signal, and carrying out adaptive weighted fusion and decoding on the aligned features to generate a fused image. The problems that an existing method is insufficient in semantic modeling and poor in fusion result consistency are effectively solved.
Owner:XIAMEN UNIV OF TECH

Infrared and visible light image fusion method based on multi-semantic deep collaboration

The invention provides an infrared and visible light image fusion method based on multi-semantic deep collaboration. The method mainly solves the problem that an existing method cannot fully integrate text modals and global consistency between image fusion and downstream tasks. Comprising the following steps: 1) constructing a dual-task parallel network structure, and efficiently establishing a deep correlation between image fusion and a downstream segmentation task; 2) designing a multi-semantic deep collaboration module, realizing effective fusion of multi-modal information by deeply integrating text features, pixel-level features and segmented semantic features, and meeting semantic requirements of downstream tasks; 3) guiding an image fusion and segmentation task by using deeper and fine-grained semantic information in a text mode, and enhancing semantic consistency between a fusion result and a downstream task; and 4) inputting the obtained multiple semantic features into a fusion decoder to generate a final image fusion result. The semantic comprehension and visual perception capabilities of the model can be effectively enhanced, and the multi-modal image fusion performance is improved.
Owner:XIDIAN UNIV

Skin lesion area reconstruction system based on multi-modal image fusion

The invention relates to the technical field of skin lesion area reconstruction, in particular to a skin lesion area reconstruction system based on multi-modal image fusion. The system comprises a deep multi-scale feature extraction module, a skin thermal anomaly feature extraction module, a hemodynamic feature extraction module, a depth and boundary feature extraction module, a skin surface feature extraction module, a fluorescence intensity distribution and morphological feature extraction module, a first feature fusion module, a second feature fusion module and a lesion area reconstruction module. According to the method, appearance, heat, blood flow, structure and surface features are extracted from the multi-modal data, then the multi-modal image data are integrated, finally, comprehensive representation and reconstruction of skin lesion appearance, function and structure information are achieved according to the synergistic effect of the multi-modal information, and the accuracy and clinical reference value of lesion area reconstruction are effectively improved.
Owner:HANGZHOU THIRD PEOPLES HOSPITAL (HANGZHOU HUIMIN HOSPITAL HANGZHOU THIRD AFFILIATED HOSPITAL OF ZHEJIANG UNIV OF TRADITIONAL CHINESE MEDICINE)

Multi-modal image fusion method based on dynamic pseudo supervision and semantic guidance

The invention provides a multi-modal image fusion method based on dynamic pseudo supervision and semantic guidance, belongs to the technical field of image fusion, and is used for realizing high-quality fusion of infrared and visible light images. The method comprises the following steps: firstly, taking a fusion quality index as guidance, adaptively generating and iteratively updating a pseudo-supervision image, and realizing dynamic alignment of a model training target and a fusion evaluation index; secondly, decomposing the characteristics of the input image into high-frequency and low-frequency components, and respectively carrying out multilayer convolution and channel attention enhancement; thirdly, cross-modal semantic embedding features of the infrared image, the visible light image and the pseudo-supervision image are utilized to calculate semantic similarity, affine modulation parameters are generated to perform semantic alignment and weight adjustment on the features, and the consistency of multi-modal feature fusion and target saliency are enhanced; according to the method, the feature alignment of fusion optimization and semantic guidance with consistent evaluation can be realized under the condition of no manual annotation, and the detail definition, the structural integrity and the visual perception quality of a fusion result are effectively improved.
Owner:DALIAN UNIV +1

Single photon and visible light image fusion method and system based on multi-scale Markov random field model

The invention provides a single photon and visible light image fusion method and system based on a multi-scale Markov random field model, and belongs to the technical field of multi-modal image fusion. In the long-distance target ranging and imaging process, a multi-sensor fusion method is used, the high resolution advantage of a visible light intensity image is utilized, the problems that a single-photon laser radar is small in point cloud density and low in image resolution are effectively solved, and the sensing capacity of the single-photon laser radar for scene target information is effectively enhanced; by using the fusion method based on the multi-scale Markov random field model, the structural consistency and the anti-interference capability are enhanced, global structural deviation or local overfitting under a single scale is effectively avoided, and the method is suitable for image fusion under multi-modal and complex scenes.
Owner:BEIJING CHANGCHENG INST OF METROLOGY & MEASUREMENT AVIATION IND CORP OF CHINA

Defect detection method and system based on insulator multi-modal image fusion

The invention discloses a defect detection method and system based on insulator multi-modal image fusion, and relates to the technical field of electrical equipment defect detection, multi-modal image fusion is performed based on an attention mechanism improved RFN-Nest image fusion model, fusion of infrared image temperature anomaly features and visual image structure detail features is enhanced, and the defect detection accuracy is improved. The information entropy, mutual information and other indexes of the generated fusion image are remarkably superior to those of an original model and a traditional fusion method, high-quality data support is provided for a detection task, then defect detection is carried out based on an attention mechanism improved YOLOv8 target detection model, the defect feature discrimination capability and the anti-interference capability are improved, and the detection efficiency is improved. The problems of missing report and false report of insulator defects in a complex scene are effectively solved, the performance is remarkably improved compared with a traditional independent link design scheme, and the method can be directly applied to an actual electric power inspection scene.
Owner:GUANGYUAN POWER SUPPLY COMPANY OF STATE GRID SICHUAN ELECTRIC POWER

Substation multi-modal image fusion registration method, system, medium and equipment

The invention discloses a substation multi-modal image fusion registration method and system, a medium and equipment, and belongs to the technical field of power system multi-modal image analysis, and the method comprises the steps: obtaining a visible light image and an infrared image of a substation power equipment region, carrying out the edge detection, and extracting feature points; performing initial matching through a nearest neighbor and secondary neighbor distance ratio constraint to obtain a first matching point pair; further screening out a second matching point pair according to the matching point pair connecting line slope; solving an affine transformation matrix by adopting a least square method based on the screened point pairs; and finally, performing spatial transformation on the infrared image by using the matrix, and generating a registered infrared image which is spatially aligned with the visible light image by combining bilinear interpolation resampling. By implementing the method, the problems of high mismatching rate of feature points and insufficient registration precision caused by large difference of multi-band image imaging principles under the complex background of a transformer substation in the prior art can be solved.
Owner:GUANGZHOU XINDILI ENERGY TECHNOLOGY CO LTD

Multi-modal image fusion method based on deep coding and decoding axis interactive attention network

The invention relates to the technical field of image processing, in particular to a multi-modal image fusion method based on a deep coding and decoding axis interactive attention network, which comprises the following steps of: S1, acquiring data of an infrared image and a visible light image, normalizing the infrared image and the visible light image, and then inputting a model for feature extraction; the feature extraction method comprises an encoder, a fusion strategy and a decoder. S2, in a feature extraction stage of an encoder, multiple times of extraction is performed on two paths of infrared and visible light images through a convolutional neural network and transformer, so that local information of two modal images is captured, and encoding representation is obtained; s3, performing multi-time layered fusion on the features coded by the encoder by a fusion strategy; s4, reconstructing and decoding the fused image by a decoder; the feature representation capability is higher, the utilization rate of original information is higher, image detail mining is more sufficient, and the fusion effect is better.
Owner:SHAANXI SILK ROAD DIGITAL INTELLIGENT NAVIGATION TECHNOLOGY CO LTD

Forest fire rapid identification and accurate fire point positioning method based on air-ground cooperation

The invention discloses a forest fire rapid identification and accurate fire point positioning method based on air-ground cooperation. Comprising the following steps: a forest fire panoramic monitoring and positioning method based on a double-axis holder, a wildfire target intelligent identification algorithm based on multi-modal image fusion analysis, and a low-power-consumption air-ground cooperative communication network oriented to forest fire monitoring. The invention relates to an unmanned aerial vehicle-based forest wildfire target maneuvering inspection and positioning method and an air-ground fusion wildfire target accurate positioning method. According to the invention, 360-degree panoramic image acquisition is realized by using the double-shaft holder, and meanwhile, a fire source position self-adaptive adjustment mechanism is provided; multi-modal image fusion analysis is adopted in the image processing layer to resist interference, blind areas are compensated, and the fire source recognition accuracy is improved; a low-power-consumption air-ground cooperative communication network is constructed, so that the cost is reduced, and flight is continued; and the monitoring limitation is broken through by means of flexible inspection of the unmanned aerial vehicle, a closed-loop system of air-ground cooperation-information complementation-intelligent verification is formed, and the monitoring stability in a complex scene is remarkably enhanced.
Owner:ANHUI UNIV

Multi-modal image fusion method and system, medium and program product

The invention relates to the technical field of computer vision, in particular to a multi-modal image fusion method and system, a medium and a program product, and the method comprises the steps: firstly obtaining a bimodal image data set with consistent time and space, the bimodal image data set comprising a visible light image data set and an infrared image data set; then constructing an image fusion network by using the bimodal image data set, and constraining the image fusion network by using source image consistency loss, infrared target structure loss and visible light texture loss; performing target detection on the fused image output by the image fusion network by using a target detection network; carrying out the cooperative training of the target detection network and the image fusion network through the target detection loss, and obtaining the trained image fusion network and target detection network; and finally, low-altitude weak and small target detection is carried out by using the trained image fusion network and the target detection network. According to the invention, the detection precision and efficiency of the low-altitude weak and small target are obviously improved.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Generative multi-modal image fusion detection method based on state space model

The invention discloses a generative multi-modal image fusion detection method based on a state space model, and belongs to the technical field of multi-modal image processing and target detection.The method comprises the steps that a network model comprising a generator and a discriminator is adopted, the generator is composed of a feature flow, a fusion flow and a reconstruction flow, extracting low-level features of the source image by using a convolution module, a Mangbar module and a guide type Mangbar module in the feature flow; performing shallow layer and deep layer fusion by using a cross-modal interaction fusion module based on Mangban in the fusion stream, and guiding deep layer feature extraction by using a shallow layer fusion result; and generating a fusion image through up-sampling in the reconstruction stream, and integrating a target detection module to carry out end-to-end detection. Wherein the Mangbar module utilizes the linear complexity characteristic of a state space model to realize global perception modeling of image features. According to the method, the quality and information richness of multi-modal image fusion are effectively improved, and the accuracy and robustness of target detection in a complex scene are enhanced.
Owner:YANTAI UNIV

Iron tower screw state detection method and system based on multi-modal image fusion

The invention discloses an iron tower screw state detection method and system based on multi-modal image fusion, and belongs to the technical field of power system inspection, and the method comprises the steps: obtaining a multi-modal component image of a target iron tower; performing image preprocessing and feature extraction on the multi-modal component image to obtain a multi-scale feature map; constructing and training an iron tower screw identification model based on the multi-scale feature map; performing multi-stage detection and state evaluation on the obtained component image based on an iron tower screw identification model; and performing physical coordinate positioning on the missing screws in combination with multi-source information and a detection result, and establishing a cloud-edge collaborative closed-loop feedback mechanism to update and optimize the model. According to the method, the screws can be accurately recognized in a complex environment, multi-dimensional state monitoring of existence-deficiency-looseness of the screws is achieved, the management refinement level and the preventive maintenance capability are remarkably improved, federal learning is adopted to guarantee data privacy and enhance model generalization, and meanwhile, closed-loop feedback is adopted to enable the model to be continuously optimized.
Owner:CHINA TOWER CO LTD

Valve well gas leakage detection system based on image region segmentation

The invention relates to the technical field of image processing, and particularly discloses a valve well gas leakage detection system based on image region segmentation. The system comprises an image acquisition module, a preprocessing and enhancement module, a multi-scale feature extraction module, a semantic segmentation network module, a false leakage suppression module and a decision output module, through multi-modal image fusion, semantic segmentation and texture consistency verification, accurate identification and positioning of a leakage area are realized, false leakage interference is suppressed, and the detection reliability is improved.
Owner:GONGZUN INSTR (ZHEJIANG) CO LTD

Multi-modal image fusion and intelligent evaluation method for stem cell differentiation process

PendingCN121214127AMathematical modelsBiostatisticsEpigenetic ProfileCell Differentiation process
The invention belongs to the technical field of biological analysis and evaluation, and particularly relates to a stem cell differentiation process multi-modal image fusion and intelligent evaluation method, which comprises a stem cell differentiation process multi-modal image fusion method: on the basis of constructing a multi-modal near-physiological mechanics microenvironment, carrying out image analysis by adopting the multi-modal image fusion method; an intelligent evaluation method for the stem cell differentiation process: developing an intelligent evaluation method based on deep learning driving on the basis of constructing a multi-modal near-physiological mechanical microenvironment, and revealing a mechanical-epigenetic regulation mechanism; the synergistic application of the two solves the core observation and regulation problems in stem cell differentiation research, and provides an industrialized intelligent platform for regenerative medicine. According to the method, the data integrity can be improved, the resolution limitation is broken through, the experiment process is accelerated, accurate mechanism analysis can be realized, automatic decision making is promoted, and the research and development cost is reduced.
Owner:EAST CHINA JIAOTONG UNIVERSITY

Arthroscopic surgery real-time navigation method and system based on multi-modal image fusion

The invention relates to the technical field of medical image processing and surgical navigation, in particular to an arthroscopic surgery real-time navigation method and system based on multi-modal image fusion, and the method comprises the following steps: S1, collecting preoperative CT (Computed Tomography) or MRI (Magnetic Resonance Imaging) data of a patient to generate a three-dimensional model, and collecting an intra-operative arthroscopic real-time video to provide input data for subsequent registration; and S2, performing image preprocessing, including denoising and enhancement, on the three-dimensional model and the arthroscope video, and extracting key anatomical features and lesion boundaries to obtain a processing result which can be used for registration. According to the method, the preoperative three-dimensional model and the intra-operative arthroscope real-time video are fused, and multi-strategy registration, closed-loop correction and dynamic updating are combined, so that continuous registration and real-time navigation of multi-modal data are realized, and the problems that the traditional arthroscope surgical navigation mostly depends on a single-modal image, and the navigation accuracy is low are solved. And the problem of insufficient intraoperative registration precision caused by lack of multi-source information fusion is solved.
Owner:JIANGQIAO HOSPITAL JIADING DISTRICT SHANGHAI

Vehicle identity accurate identification method and system based on multi-modal image fusion

The invention discloses a vehicle identity accurate recognition method and system based on multi-modal image fusion, and particularly relates to the field of vehicle identity recognizing.The method comprises the steps that images are collected through visible light and infrared sensors which are synchronously calibrated, and preprocessing is conducted through adaptive filtering, CLAHE enhancement and distortion correction; a feature pyramid network and a channel attention mechanism are utilized to carry out adaptive fusion of multi-modal features, an improved residual convolutional network is combined to extract key features of a vehicle contour, a license plate, a vehicle logo and a vehicle window, a mixed matching algorithm is adopted, a cosine similarity and a weighted Euclidean distance are synthesized to calculate a feature matching degree, and the feature matching degree is calculated. Retrieval is carried out through a local-cloud distributed database; and when the fusion similarity exceeds a preset threshold value, outputting identity information such as a vehicle license plate, a brand model, a registration year and a color.
Owner:GUANGDONG FUDA INTELLIGENT CO LTD

Real-time positioning and classification recognition ore sorting system and method based on deep learning

The invention discloses an ore sorting system and method for real-time positioning and classification recognition based on deep learning, and relates to the technical field of AI vision, and the system comprises the following modules: an ore image multi-source fusion module which is used for collecting and preprocessing multi-modal ore images and fusing the multi-modal ore images to form a multi-modal image set; and the end-to-end deep learning reasoning module is used for combining an improved pyramid network structure and a local attention mechanism, extracting and fusing multi-scale feature information, performing reasoning calculation to output a recognition result, and obtaining a position coordinate, a category result and a judgment confidence coefficient of the ore. According to the method, through multi-modal image fusion and end-to-end deep learning reasoning, the position, category and contour information of the ore can be accurately recognized, the problem that the recognition rate of complex ore features, especially hard-to-process ores such as pock ores and linear ores, of a traditional method is low is effectively solved, the sorting precision is remarkably improved, and the sorting efficiency is improved. And mistaken discarding of concentrates and mistaken selection of tailings are reduced.
Owner:HUNAN JINSHI SORTING INTELLIGENT TECH CO LTD

Micro-nano visual lithography defect self-adaptive repairing method based on multi-modal image fusion

The invention discloses a micro-nano visual lithography defect self-adaptive repairing method based on multi-modal image fusion, which comprises the following steps: acquiring a multi-modal image of a wafer surface, inputting the multi-modal image into a cascade-feedback type double-branch defect analysis network, fusing feature maps of two extraction networks by adopting a decision-making layer self-adaptive fusion mechanism based on confidence coefficient weighting, and obtaining a micro-nano visual lithography defect self-adaptive repairing result. Generating a fusion defect probability graph; performing defect detection and classification based on the fused defect probability graph, and identifying the type, position and size parameters of the photoetching defect so as to adaptively determine repair parameters; repairing the defect area according to the repairing parameters; and after repairing is completed, the multi-modal image of the defect area is collected again, the similarity of the images before and after repairing is calculated, when the similarity is smaller than a preset threshold value, repairing parameters are determined again till the similarity is larger than or equal to the preset threshold value, and a final repairing image is output. The defect detection precision can be improved, and the repair effect is improved while the repair efficiency is improved.
Owner:GUANGDONG LUON MICRO-NANO VISION TECHNOLOGY CO LTD

Multi-modal image processing method based on frequency perception and interactive fusion

The invention discloses a multi-modal image processing method based on frequency perception and interactive fusion. The problem of image fusion under a low-light condition can be solved in a decoupling type double-branch parallel processing mode. The method comprises the following steps: importing a data set and performing characteristic decomposition, receiving visible light and infrared images, and decoupling each modal image into low-rank and sparse characteristics; carrying out interactive content fusion, carrying out cross-modal interaction on high and low frequency features through bilateral information exchange, carrying out enhancement, and reconstructing an intermediate fusion image; performing parallel global illumination estimation, inputting the visible light image into an illumination estimation network based on frequency domain processing, and generating an ideal illumination image; and performing physical illumination reconstruction and output, and performing brightness improvement on the intermediate fusion image to obtain a fusion image. According to the method, specialized processing of frequency perception and a cross-modal deep interaction mechanism are combined, and the quality and robustness of multi-modal image fusion under complex conditions are remarkably improved by decoupling content fusion and an illumination estimation task.
Owner:HOHAI UNIV

Multimodal image fusion method based on hesitant fuzzy variable granularity dictionary learning

The invention relates to the technical field of image fusion, in particular to a hesitant fuzzy variable granularity dictionary learning-based multi-modal image fusion method, which comprises the following steps of: firstly, adaptively selecting division granularity according to image quality, and partitioning a source image into blocks; then extracting image block features and calculating hesitant fuzzy membership degrees of the image block features so as to quantitatively represent uncertainty information in the image; obtaining a joint over-complete dictionary and a sparse coefficient through dictionary learning, and fusing the hesitant fuzzy entropy and a granularity coefficient to construct an adaptive weight; and finally, fusing the sparse coefficient by using the weight and reconstructing a fused image. The problems of image fuzzy processing, structure multi-scale expression and insufficient adaptive feature extraction capability are effectively solved.
Owner:SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING

A multi-modal feature fusion based harsh environment image enhancement and restoration method

ActiveCN122115236AImage enhancementNeural learning methodsPattern recognitionLeast squares optimization
The application provides a kind of multi-modal feature fusion-based harsh environment image enhancement and restoration method, belongs to image enhancement processing field;The method first obtains multi-modal image and calculates each modal degradation degree graph;Based on the degradation degree and gradient amplitude, a fusion weight is constructed, and the local dominant direction and coherence are obtained by structure tensor analysis;According to the coherence, the pixels are partitioned, and the adaptive processing of each modal gradient is carried out by using hard projection, isotropic fusion or soft projection strategy respectively;Based on the fusion weight, the processed gradient is weighted and fused, and the fusion gradient field is obtained combined with the enhancement factor g (x) ;Finally, an energy function containing pixel fidelity term and gradient fidelity term is constructed, and the enhanced fusion image is reconstructed by weighted least squares optimization;The application suppresses the contribution of blurred modal through degradation perception mechanism, and solves the gradient direction conflict through structure adaptive projection, which significantly improves the clarity and structure fidelity of multi-modal image fusion in harsh environment.
Owner:CHENGDU FANCHEN TECH CO LTD

Multimodal image fusion method and imaging system

The present application relates to the image processing method of optical remote sensing imaging detection, specifically relates to a multi-modal image fusion method and imaging system, in order to solve the existing technical image capture of target scene depends on time-sharing type imaging system or multi-device acquisition, resulting in different modal image registration difficulty, low registration accuracy, image edge details are inconsistent, the deficiency of poor fusion effect, the multi-modal image fusion method of the present application includes the original image and characteristic spectrum of target scene, multi-modal image preprocessing and multi-modal image fusion steps, based on the spectral radiation curve of target scene, spectral image and polarization image realizes multi-modal image fusion.Simultaneously, provide a kind of multi-modal image imaging system, including sub-aperture array compound eye unit, for obtaining the original image of target scene, realize the spectral image and polarization image of target scene are acquired simultaneously.
Owner:XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI

Multi-modal fusion progressive pre-training and self-adaptive fine tuning method based on visual large model

The invention provides a multi-modal fusion progressive pre-training and self-adaptive fine tuning method based on a visual large model, and relates to the technical field of deep learning and image fusion, and the method comprises the steps: firstly obtaining images of different modals, carrying out the preprocessing of the images, and obtaining a preprocessed multi-modal image; inputting the multi-modal image into the constructed multi-modal image fusion model, constructing a self-supervision loss function, and training the multi-modal image fusion model by using the multi-modal image based on the self-supervision loss function; and constructing a total task loss function based on the multi-modal image fusion downstream task, performing fine adjustment on the multi-modal image fusion model considering the multi-modal image fusion downstream task, and finally performing fusion on the multi-modal image based on the trained visual large model. According to the method, multi-modal optimization is realized through progressive pre-training and self-adaptive fine tuning, and the model performance is improved.
Owner:GUANGDONG UNIV OF TECH