Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

3603 results about "Imaging Feature" patented technology

Unmanned aerial vehicle image-based small object detection method for target areas

The present invention relates to the technical field of deep learning and computer vision. Disclosed is an unmanned aerial vehicle image-based small object detection method for target areas. The present invention crops images of obvious small objects in certain target areas, and annotates the small objects of different categories to form a raw training and testing dataset, so as to ensure the accuracy of data required in the early stage of the algorithm and further ensure the scientificity of the algorithm; uses the computing capability of an improved YOLOv7 detection model to collect image features of different degrees in the dataset, the improved YOLOv7 detection model using YOLOv7 as a basic model and adding to a neck network an MS-CET module, which is constituted by an improved self-attention mechanism and convolution module SPPCSP, and a BHC-FB module, which is constituted by bidirectional mixed convolution modules NConv and RPConv connected in parallel; and finally fuses different feature layers as a final judgment basis of an unmanned aerial vehicle for small object detection in the target areas, to further check the accuracy of the algorithm and criteria for dataset selection, thereby improving recognition accuracy.
Owner:CHONGQING UNIV OF TECH

Brain tumor multi-modal large model construction method and device, equipment and storage medium

The invention discloses a brain tumor multi-mode large model construction method, device and equipment and a storage medium, and is applied to the technical field of brain tumor imagines.The method comprises the steps that pixel-concept level alignment is conducted on a multi-mode MRI image and a pathological text; constructing a multi-modal feature fusion network for fusing image features and text features by adopting an attention mechanism of pathology perception and combining medical semantic information; training the multi-modal feature fusion network to generate an analysis report and a segmentation result; according to the technical scheme of multi-task cooperation, cross-modal pathological semantic accurate alignment, pathological knowledge graph injection and lightweight and continuous optimization parallelization, full-process coverage of brain tumor accurate segmentation, analysis report generation and prognosis prediction is achieved, the problems that a traditional model lacks pathological semantic support and is insufficient in clinical adaptability are solved, and the clinical adaptability of the traditional model is improved. And the deployment feasibility and the dynamic optimization capability are also considered.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Small target identification method and system for multi-modal fusion image in complex environment

The invention discloses a small target recognition method and system for a multi-modal fusion image in a complex environment, and belongs to the technical field of computer vision and image recognition, and the method comprises the steps: obtaining a visible light image, an infrared image and environment sensor data; image registration is carried out on visible light and infrared images, and a multi-scale image feature pyramid is constructed. And respectively extracting visible light and infrared image features to obtain visible light and infrared imaging feature data. And performing multi-modal data fusion on the visible light and infrared imaging feature data based on a cross-modal attention mechanism, and adaptively adjusting a fusion weight based on environmental sensor data to generate fusion features. And performing space-time enhancement processing on the fusion feature to obtain an enhanced fusion feature. And performing target tracking detection on the small target, and outputting position and category information of the small target. According to the method, the small target recognition capability in a severe environment is remarkably improved, and high precision and robustness can still be kept in a foggy, low-visibility and dark scene.
Owner:CHINA TOWER CO LTD +1

Multi-mode fruit sugar degree nondestructive testing method, device, system and medium

The invention discloses a multi-mode fruit sugar degree nondestructive testing method, device and system and a medium, and the method comprises the steps: preprocessing an input fruit near infrared spectrum image and a fruit visible light image to obtain a one-dimensional near infrared spectrum input image tensor and a two-dimensional visible light input image tensor; a pre-trained multi-modal deep learning fusion model is utilized to obtain the predicted fruit sugar degree, and the multi-modal deep learning fusion model is pre-trained to establish an input near infrared spectrum input image tensor and a visible light input image tensor. The method comprises a spectral feature extraction network, an image feature extraction network, a multi-modal feature fusion network and a regression prediction network. The objective of the invention is to solve the problem of limited prediction precision caused by the fact that single spectral information is susceptible to noise, illumination conditions, peel thickness, water content and other factors, and improve the accuracy of fruit sugar degree nondestructive testing.
Owner:HUNAN UNIV

Multimodal sentiment analysis method based on diffusion model and self-paced learning

The invention provides a multi-modal sentiment analysis method based on a diffusion model and self-paced learning. The method comprises the following steps: firstly, dividing a data set into a missing image modal data set and a complete modal data set according to image modal integrity; thirdly, constructing a feature alignment diffusion model, and performing image generation; training the diffusion model by adopting a self-paced learning strategy and a missing image data set; and based on the trained diffusion model, guiding a reverse process through text features to generate feature representation of the missing image. And carrying out weighted fusion on the generated image features and text features by using an attention mechanism, and dynamically adjusting contribution weights of all modalities to generate a complete multi-modal feature representation. And finally, integrating a missing modal completion result and the complete modal features to form a unified multi-modal representation, inputting the unified multi-modal representation into a multi-modal sentiment classification module, and outputting a sentiment classification result. According to the method, the problem of multi-modal sentiment analysis under random missing of image modals is effectively solved, and the generation quality and semantic consistency are improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Depth-guided three-dimensional Gaussian reconstruction method and system suitable for sparse view angle image

The invention discloses a depth-guided three-dimensional Gaussian reconstruction method and system suitable for a sparse view angle image, and belongs to the technical field of computer vision and three-dimensional reconstruction, and the method comprises the steps: carrying out the wavelet transformation super-resolution processing of a sparse multi-view angle image; outputting depth prior through a pre-trained monocular depth model, and extracting multi-view image features; constructing cross-view depth candidates by adopting a planar scanning stereo method, and generating initial depth distribution through feature similarity calculation; a self-attention-cross attention structure and deformable sampling are adopted to realize coarse-to-fine depth matching optimization; using an improved UNet network to fuse multi-scale features, and optimizing a depth estimation result; predicting three-dimensional Gaussian primitive parameters; and constructing a Gaussian field to generate a new visual angle image. According to the method, the problems of low accuracy, integrity and efficiency of existing sparse view angle three-dimensional reconstruction are solved. According to the invention, the precision and the detail fidelity of the depth map are improved, and high-fidelity three-dimensional reconstruction under the sparse visual angle condition is realized.
Owner:YUNNAN UNIV

Coal gangue recognition method based on visible-near-infrared spectrum and image multi-modal information fusion

The present invention belongs to the technical field of coal gangue recognition and sorting, and in particular, relates to a coal gangue recognition method based on visible-near-infrared spectrum and image multi-modal information fusion. The method includes: S1 collecting spectral information and image information about a sample to be recognized; S2 preprocessing the spectral information and the image information respectively; S3 extracting spectral features from a spectral data set by using a spectral feature extraction neural network model; and extracting image features from an image data set by using an image feature extraction neural network model; S4 inputting the spectral features and the image features obtained from feature extraction into a two-stream fusion network; S5 inputting extracted spectral features, extracted image features and a comprehensive feature into a spectral branch classifier, an image branch classifier and a fusion branch classifier, respectively; S6 calculating importance weights of a spectral branch, an image branch and an image-spectrum fusion branch; and S7 subjecting the importance weights and corresponding confidence to multiply-accumulate operation to obtain a score matrix of coal gangue, and using the score matrix to achieve coal gangue recognition.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Unmanned aerial vehicle imaging control method and device based on AI and flight control parameter fusion

The invention relates to an unmanned aerial vehicle imaging control method and device based on AI and flight control parameter fusion, and the method comprises the steps: carrying out the scene feature extraction of historical imaging data, obtaining image features under multiple scales, constructing an AI scene adaptation strategy according to the image features, solving an illumination intensity prediction model, and obtaining an optimal imaging parameter set of a corresponding camera; and acquiring original image data shot by the camera under the optimal imaging parameter set, optimizing the original image data based on an AI scene adaptation strategy, predicting the flight state of the unmanned aerial vehicle based on the flight control parameter to generate an imaging preset parameter, and using the imaging preset parameter to control shooting of the unmanned aerial vehicle. The AI scene adaptation strategy is constructed by using AI in combination with historical imaging data, the problem of poor scene adaptability is solved, in addition, the illumination intensity prediction model is constructed by combining flight control parameters, the illumination intensity prediction model is solved, the optimal imaging parameter set of the corresponding camera is obtained, and the defect of weak illumination adaptability is overcome.
Owner:HANGZHOU DUNJIA TECH CO LTD

Neural spline fields for image feature separation

Methods and systems are described for analyzing images. One or more machine learning models may be trained based on a plurality of images. The one or more machine learning models may comprise a model representing a feature in a scene. The one or more machine learning models may be trained to map input image coordinates to vectors of spline control points. Images may be reconstructed removing the feature from the scene.
Owner:THE TRUSTEES OF PRINCETON UNIV

Enameled plate surface defect detection method based on machine vision

The invention discloses an enamel plate surface defect detection method based on machine vision, and relates to the technical field of enamel plate surface defect detection. The method comprises the following steps: acquiring an original image of an enamel plate, and fusing the original image through a point-by-point optimal gray fusion method to obtain a real-time reflection suppression image; a defect-free reflection suppression image of the enamel plate is collected, matching analysis is conducted on the defect-free reflection suppression image and the real-time reflection suppression image through an improved PaDiM algorithm, a difference texture area is obtained, and coordinates of the difference texture area are recorded; in the difference texture area, two types of image features are extracted through a scale attention feature extraction method and fused, a real-time and standard fusion matrix is obtained, and a defect detection result is obtained through similarity matching; and converting the coordinates of the difference region to obtain a defect positioning result, calculating a defect comprehensive score based on a defect detection result, and grading to complete the surface defect detection of the enamel plate.
Owner:HUBEI SANXING TECH CO LTD

Multi-mode panoramic segmentation method for multi-view space-time alignment and implicit feature interaction

The invention belongs to the technical field of laser radar-camera panoramic segmentation, and particularly relates to a multi-mode panoramic segmentation method for multi-view space-time alignment and implicit feature interaction, and the method is executed by a multi-mode panoramic segmentation network, and comprises the steps: S1, obtaining laser radar point cloud data and multi-view camera image data in the same scene; s2, performing double-branch feature coding on the laser radar point cloud data; performing multi-scale image feature extraction on the camera image data; s3, generating false point cloud features with geometric perception capability; s4, generating semantic pixel features; s5, performing implicit fusion on the pseudo point cloud features generated in the S3 and the semantic pixel features generated in the S4 to obtain cross-modal fusion features; and S6, based on the cross-modal fusion features obtained in the S5, generating a unified panoramic segmentation result containing semantic tags and instance IDs. The method can effectively improve the robustness and precision of multi-mode panoramic segmentation in a complex urban environment.
Owner:CHONGQING UNIV OF TECH

Slope stability grade identification method and system based on multi-modal deep learning

The invention belongs to the technical field of geological disaster risk assessment, and particularly relates to a slope stability grade identification method and system based on multi-modal deep learning. Comprising the following steps: acquiring a remote sensing image and corresponding parameter data, and respectively preprocessing the remote sensing image and the corresponding parameter data to obtain a high-dimensional image feature vector and a parameter feature vector; performing dynamic fusion of multi-modal information on the high-dimensional image feature vector and the parameter feature vector through a modal-level gating and fine-grained weighting mechanism to obtain a fusion feature; and inputting the fusion features into a preset classifier to obtain probability distribution of slope stability levels, and outputting the class with the maximum probability as a stability level identification result. According to the method, through combined modeling of the image modality and the parameter modality and introduction of an improved gating attention mechanism in a feature fusion stage, a reliable explanatory basis can be provided while high-precision prediction is ensured, so that the engineering applicability and the popularization value of the model are enhanced.
Owner:SHAANXI PROVINCIAL GEOLOGICAL ENVIRONMENT MONITORING STATION +1

Bone tumor fine-grained classification model training and classification method and device

The invention provides a bone tumor fine-grained classification model training and classification method and device, and the method comprises the steps: constructing a multi-modal positive sample containing global / local positive lateral X-ray images, lesion attributes and patient information, and removing a false negative part in combination with text semantic similarity to construct a high-quality negative sample; global / local image features are extracted through double image encoders, lesion attribute keywords are converted into'entity-translation-existence 'triples based on a medical knowledge base, and basic information of a patient and global / local semantic features of lesion attributes are extracted through a text encoder; infoNCE contrast loss is constructed for global images and global semantics based on contrast learning, global image-text feature alignment and local image-text feature alignment are realized in combination with a local mutual information loss and classification loss training model calculated based on a DV variational formula, medical term semantics are deeply combined, the training stability is improved, and the training efficiency is improved. The accuracy and robustness of bone tumor subtype classification are remarkably improved, and reliable support is provided for clinical precise diagnosis.
Owner:BEIHANG UNIV

Oral cavity image recognition method and system based on deep learning, and storage medium

The invention provides a deep learning-based oral cavity image recognition method, a storage medium and a deep learning-based oral cavity image recognition system. The method comprises the steps of deploying a federated learning framework and collecting a multi-modal oral cavity image data set; extracting local features to obtain image features, and generating a modal adaptive weight map; a multi-head self-attention mechanism is used for fusing the cross-modal features to generate a fused feature map, and deconvolution up-sampling is carried out to form high-resolution multi-modal feature representation. A tooth segmentation mask is generated based on this representation, and an initial diagnostic report is generated. And aggregating the attention weight of each client through an encryption protocol, and generating interpretable decision support data. And finally, generating a structured clinical report by using a natural language. According to the method, the Grad-CAM thermodynamic diagram is combined with the encrypted and aggregated attention weight, so that the privacy security is guaranteed, the model interpretability is enhanced, the clinical credibility and the diagnosis decision efficiency are improved, and the problems of insufficient diagnosis precision of complex lesions and insufficient utilization of multi-modal information in the prior art are solved.
Owner:CHONGQING THREE GORGES MEDICAL COLLEGE +1

Super-lens image restoration method based on fuzzy prior and semantic segmentation

The invention discloses a super-lens image restoration method based on fuzzy prior and semantic segmentation. The method comprises the following steps: eliminating an image visual angle displacement error by adopting a feature point automatic registration algorithm; the method comprises the following steps: designing a parallel multi-scale convolution adapter on the basis of an encoder optimized by a pre-trained visual Transform model, and generating encoding features adaptive to the degradation characteristics of a super lens; splicing the degraded image and the semantic segmentation mask into a joint tensor along a channel dimension, and obtaining a prior feature based on a fuzzy prior extraction network; performing spatial feature enhancement on the semantic segmentation mask based on a semantic segmentation graph convolutional network to obtain graph features aligned with the coding features; splicing the coding features, the prior features and the image features along channel dimensions to generate fusion features, and finally outputting a preliminary recovery image; and optimizing the whole model by adopting a staged training strategy, and finally outputting a high-quality recovery graph. According to the method, the physical rationality and detail fidelity of the recovery quality are improved.
Owner:江苏优众微纳半导体科技有限公司

Fastener defect automatic identification method based on deep learning

The invention discloses a fastener defect automatic identification method based on deep learning, and the method comprises the following steps: S1, collecting fastener image data, and generating a standardized image sample; s2, executing edge feature enhancement and multi-scale filtering operation to generate an image feature map; s3, inputting into an improved TransNeXt model, and outputting a defect stage label, an identification confidence coefficient and a feature coding vector; s4, screening the position of the suspicious region, executing a local region decoding operation, and generating a defect response thermodynamic diagram; s5, extracting a bounding box, morphological characteristics and response intensity distribution of the maximum response region, and constructing a defect candidate region set; s6, non-defect areas are screened out, and a final defect recognition result is obtained; and S7, constructing a defect distribution diagram and a statistical information table, and outputting a defect type, a position coordinate, an area estimation value and an identification confidence level. According to the invention, high-precision automatic identification and structured information extraction of fastener defects are realized.
Owner:WUXI ZHIGULIAN TECHNOLOGY CO LTD

Shearing behavior dynamic correction method and system based on wear state recognition

The invention relates to the technical field of image recognition, in particular to a shearing behavior dynamic correction method and system based on wear state recognition. Acquiring a digital image sequence of the cutting edge area; constructing an image feature separation network, and performing parallel feature extraction on the preprocessed digital image sequence; identifying a pixel-level high-frequency texture discontinuous region and an edge gradient direction field by utilizing continuity characteristics of a cutting edge surface periodic texture mode to obtain a target defect probability graph; identifying a projection shadow area and a low-frequency illumination halation of the edge of the bulge by using a backlighting imaging model to obtain an interference artifact probability graph; establishing spatial mutual exclusion constraints of the target defect probability graph and the interference artifact probability graph in a pixel space, and generating a defect binary mask; performing multi-dimensional texture feature calculation on an area corresponding to the defect binary mask, and constructing a surface state feature vector; according to the invention, based on the surface state feature vector, the defect mode category is discriminated, and the corresponding shearing correction parameter is generated.
Owner:SUZHOU LILAI IRON & STEEL CO LTD

Screw fastening quality evaluation method and system based on image features

The invention belongs to the technical field of image processing, and particularly relates to a screw fastening quality evaluation method and system based on image features, and the method comprises the steps: collecting a screw fastening image, and obtaining a complex response diagram through a Log-Gabor filter; obtaining a phase congruency diagram according to each complex response diagram; acquiring gradient directions of pixel points in the phase consistency graph, and constructing an accumulator graph according to the gradient directions and offset points in a preset radius range; according to salient points in the accumulator graph and local signal-to-noise ratios of the salient points, screw positioning points are determined through two-dimensional quadratic polynomial function fitting; the torque value of the screw is obtained according to the screw positioning point driving torque measuring device, and whether the screw is fastened or not is judged. According to the method, the screw is positioned by using phase consistency in the frequency domain, so that the interference of specular reflection and extreme shadow on positioning is overcome, the accuracy of screw positioning is improved, and screw fastening quality evaluation is assisted.
Owner:SCHNEIDER SHAANXI BAOGUANG ELECTRICAL APP CO LTD +1

Image data management method and system for high-precision size visual inspection

The invention discloses an image data management method and system for high-precision size visual inspection, and relates to the technical field of image data, the image data management method for high-precision size visual inspection comprises the following steps: S1, collecting a data set, and carrying out region segmentation processing; s2, correcting and enhancing the data set, and extracting contour feature parameters; s3, mapping association is carried out, and screening classification is carried out; s4, parameters are optimized and adjusted, and a standardized size detection feature set is generated; s5, constructing a size detection model to obtain a target size quantification detection result; and S6, performing accuracy verification, and establishing an image feature size parameter association database. According to the invention, through synchronous acquisition of a multi-view image data set and a physical size calibration data set, and in combination with calibration reference of a geometric quantity standard appliance, the problems of incomplete information and large physical size mapping deviation of a traditional single-view image are solved.
Owner:CHANGCHUN AUTOMOBILE IND INST

Remote sensing image target statistical method and system fusing large language model and visual cue driving

The invention provides a remote sensing image target statistical method and system fusing a large language model and visual prompt driving. The method comprises the following steps: acquiring a remote sensing instance segmentation image to be processed and a visual prompt thereof; inputting a to-be-processed remote sensing instance segmentation image and a visual prompt thereof into the trained remote sensing image target statistical model, and outputting a remote sensing image target statistical result; the training comprises the following steps: introducing a large language model and visual cue into an encoder architecture of a GrondingDINO model to obtain a remote sensing image target statistical model; inputting a remote sensing instance segmented image and the visual cue thereof into an encoder, and outputting an image feature, a visual cue feature and a text feature; the feature intensifier carries out fusion processing on the output of the encoder; a language-guided query selection module calculates cross-modal query according to the fusion processing result, and a cross-modal decoder obtains a target statistical result of the image based on the fusion processing result and the cross-modal query; and training by using the training data and outputting the trained model.
Owner:WUHAN UNIV

Target detection method and device in wide-area complex scene, equipment and storage medium

The invention relates to a target detection method and device in a wide-area complex scene, equipment and a storage medium, and the method comprises the steps: carrying out the multi-level feature extraction of a wide-area complex scene image through employing a backbone network, and obtaining a low-level feature, a middle-level feature and a high-level feature; encoding the high-level features by adopting an EDFPT module to obtain encoded features; inputting the low-level features, the middle-level features and the coding features into a feature fusion module to obtain fusion features; screening a fixed number of image features from the fusion features by adopting an IoU-perceived query selection strategy to obtain an initial query vector; and processing the initial query vector by adopting a decoder with an auxiliary prediction head to obtain a wide-area complex scene image target detection result. The method has higher feature sensitivity to dense small targets and special-shaped targets in a wide-area complex scene image, the detection precision is effectively improved while the light weight of the model is ensured, and the omission ratio is reduced compared with RT-DETR.
Owner:NANCHANG UNIV +1

Vacuum cup surface defect detection method based on image features

The invention discloses a vacuum cup surface defect detection method based on image features, and relates to the technical field of image processing and intelligent detection. According to the vacuum cup surface defect detection method based on the image features, multi-angle image acquisition and preprocessing are performed, structural features such as texture, shape and reflection of the surface of a vacuum cup are extracted, a texture disturbance index, a contour offset index and a reflection focusing index are constructed, and the indexes are compared with a preset defect judgment interval, so that defect identification of images at various angles is realized; by constructing a multi-angle image feature extraction and index analysis mechanism, feature information such as texture, contour and reflection of the surface of the vacuum cup is accurately extracted, various abnormal degrees are quantified, automatic identification and marking are realized in combination with the preset defect judgment interval, the detection coverage range and the judgment accuracy can be effectively improved, and the detection efficiency is improved. The method enhances the recognition capability of complex defects, and is suitable for appearance quality inspection scenes with high requirements.
Owner:POLYBELL(GUANGZHOU) LTD

Construction and use method of potential diffusion model for SAR image super-resolution

The invention provides a construction and use method of a potential diffusion model for SAR image super-resolution. The construction and use method comprises the steps of obtaining an original SAR image and inputting the original SAR image into a real degradation model to generate a degraded SAR image; inputting the degraded SAR image into an automatic encoder to generate a structure-enhanced submerged space feature map; and inputting the structure-enhanced latent space feature map and the degraded SAR image into a potential diffusion model to generate an SAR super-resolution image. The method has the beneficial effects that a two-stage training strategy is adopted, different optimization targets are focused in stages, the training efficiency is improved, and meanwhile, the learning ability of the model to SAR image features is enhanced; sAR imaging key degradation factors are comprehensively covered, so that a generated low-resolution sample is closer to a real scene, high-quality data support is provided for model training, and model learning is prevented from being separated from an actual degradation rule; sAR specific interference such as speckle noise is effectively simulated, and the anti-noise training effect of the model is enhanced.
Owner:NANKAI UNIV

Medical image automatic identification system based on neural network

The invention discloses a medical image automatic identification system based on a neural network, and relates to the technical field of medical image identification. The method is used for solving the problem that early recognition of neurodegenerative diseases is difficult due to medical image and genome data splitting and poor model interpretability in the prior art. The method comprises the following steps: firstly, extracting multi-scale features of a brain structure through a three-dimensional convolutional neural network and a self-attention mechanism, calculating a multi-gene risk score based on a risk site, and encoding the score into a feature vector; secondly, using a cross attention mechanism to take gene features as query vectors, fusing the gene features with image features, and generating brain structure anomaly features under gene regulation; then, gradient weighting class activation mapping is applied to generate a visual thermodynamic diagram, and gene-image association weight weighting is combined to construct a brain region risk distribution diagram; and finally, a high-risk brain region space coordinate set is extracted through threshold segmentation, and an accurate quantification basis is provided for early recognition.
Owner:MEIZHICOMSCOPE TECHNOLOGY (WENZHOU) CO LTD

Device for evaluating consciousness level and storage medium

The invention discloses a device for awareness level evaluation and a storage medium. The apparatus comprises: a processor; the device realizes the following operations: collecting clinical information of a person to be assessed and a task state electroencephalogram signal under a target stimulation normal form; extracting frequency domain characteristics of a specific frequency band and spatial-temporal characteristics of a target event related potential based on the task state electroencephalogram signal, and combining the frequency domain characteristics and the spatial-temporal characteristics into a corresponding electroencephalogram topographic map; inputting the corresponding electroencephalogram topographic map into a multi-modal large language model, and performing image feature extraction by using an image encoder to obtain electroencephalogram features; inputting clinical information into the multi-modal large language model, and performing text feature extraction by using a text encoder to obtain text features; and performing cross-modal attention calculation fusion on the electroencephalogram features and the text features by using a cross-modal fusion module to realize consciousness evaluation so as to output a consciousness evaluation result. By means of the scheme, the consciousness level of the patient can be automatically and accurately evaluated.
Owner:UNION STRONG (BEIJING) TECH CO LTD

Remote sensing scene classification method for small sample multi-modal prototype learning

The invention belongs to the computer vision technology, and particularly relates to a small sample multi-modal prototype learning-oriented remote sensing scene classification method, which comprises the following steps of: acquiring RGB (Red, Green and Blue) images with category labels and text prompts of the RGB images as a support set; establishing a text prototype, an RGB prototype and a hyperspectral prototype of each category according to the support set; and extracting to-be-classified query set image features by using a pre-trained CLIP image encoder, calculating cosine similarities between the query set image features and the text prototype, the RGB prototype and the hyperspectral prototype of each category of the support set, taking the cosine similarities as input of a multi-layer perceptron, and obtaining the category of the to-be-classified RGB image through classification of the multi-layer perceptron. High-precision and high-robustness remote sensing scene classification is realized under the small sample condition, only prototype and similarity calculation is needed in the reasoning stage, and deployment and expansion are easy.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Image feature matching method and system based on self-supervised learning

The invention relates to the technical field of computer vision and deep learning, in particular to an image feature matching method and system based on self-supervised learning, and provides an image feature matching method and system based on self-supervised learning. Features are extracted by means of an online network model and a momentum encoder, and image features are matched by adopting a multi-order sequence comparison algorithm after the model is trained through a mixed comparison loss function. Multiple innovations of asymmetric data enhancement, a mixed loss function, an attention mechanism and a multi-order sequence comparison algorithm are fused, optimal balance of matching accuracy, training efficiency, robustness and calculation overhead is achieved, and the method is suitable for scenes of intelligent traffic fee evasion detection, security monitoring, medical image analysis, industrial quality inspection and the like.
Owner:SHENZHEN UNIV

Hub temperature anomaly detection method and system based on image feature fusion

The invention provides a hub temperature anomaly detection method and system based on image feature fusion, and relates to the technical field of image processing, and the method comprises the steps: obtaining the related data of a hub, the data comprises the surface temperature distribution image, the internal temperature data and the vibration spectrum data of a to-be-detected hub, and the braking data and the charging and discharging state data of a vehicle system; constructing a working condition simulation environment based on the braking data and the charging and discharging state data, and loading the working condition simulation environment to a pre-constructed first digital twinborn model to obtain a second digital twinborn model; based on the internal temperature data, the vibration spectrum data and the material thermal parameters of the to-be-detected hub, generating temperature distribution information by using a second digital twinborn model; and feature fusion processing is carried out on the surface temperature distribution image and the temperature distribution information through the feature fusion model to generate the target feature information of the to-be-detected hub, so that the temperature detection result of the to-be-detected hub is further generated, and the comprehensiveness and accuracy of hub temperature anomaly detection are improved.
Owner:SHANXI JIAOKE INFORMATION SYST ENG CO LTD +1

Multi-mode credible dialogue type retrieval enhancement generation system for medical diagnosis

The invention discloses a medical diagnosis-oriented multi-modal credible dialogue type retrieval enhancement generation system, and relates to the field of artificial intelligence, in the system, a multi-modal image analysis module generates a structured image feature data packet based on a multi-modal large model according to an original medical image; the knowledge database construction module is used for constructing a multi-center collaborative medical knowledge database; the image knowledge retrieval module is used for obtaining image feature associated knowledge; the clinical knowledge retrieval module is used for obtaining a candidate knowledge set; a multi-dimensional weighting reordering module obtains a knowledge list; the medical inquiry generation module is used for generating structured preliminary diagnosis based on the user questions, the image abstract and the knowledge list; the credibility verification module is used for verifying the structured preliminary diagnosis to obtain a final diagnosis report; according to the method, the medical knowledge retrieval precision can be improved, the image information fusion capability is enhanced, and high-credibility and traceable intelligent diagnosis and treatment auxiliary service oriented to doctor-patient scenes is realized.
Owner:PEKING UNION MEDICAL COLLEGE HOSPITAL

Array image demosaicing method based on dynamic convolution and adaptive coding

The invention provides an array image demosaicing method based on dynamic convolution and adaptive coding, which relates to the technical field of image processing, and comprises the following steps: acquiring single-channel original image data and a corresponding color filtering array arrangement type identifier; converting the arrangement type identifier into a multi-dimensional physical feature vector, and inputting the multi-dimensional physical feature vector into a feature processor of a neural network model to generate a weighting coefficient vector; carrying out weighted combination on the plurality of special arrangement transformation matrixes through a weighting coefficient vector to obtain a transformation component, adding the transformation component and a basic convolution kernel parameter to obtain a dynamic convolution kernel parameter, and carrying out directional modulation on a specific spatial position of the dynamic convolution kernel parameter based on a direction weight component in a multi-dimensional physical feature vector; and performing convolution operation on the original image data by using the dynamic convolution kernel parameters, extracting multi-scale features, reconstructing image features, and outputting multi-channel color image data, so that the method can be adaptive to different color filter array types, and the demosaicing precision and generalization capability are improved.
Owner:BEIJING HAOMO TECH CO LTD