Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

7353 results about "Imaging Feature" patented technology

Multi-modal medical image data intelligent processing system

The invention discloses a multi-modal medical image data intelligent processing system, relates to the field of medical image analysis, and is applied to multi-modal medical image whole-process analysis of CT, MRI, PET, ultrasound and the like. According to the system, different modal image features are extracted and fused through a cross-modal manifold fusion network; a semantic guidance dynamic registration engine optimizes registration parameters to ensure that the registration error is less than or equal to 1.5 mm; the multi-task collaborative diagnosis network realizes multiple tasks such as disease classification; the clinical knowledge embedding and interpretable module generates a structured report and is in butt joint with an HIS system. Meanwhile, the model is optimized through a federated learning architecture, the adaptability of newly added data is improved by more than or equal to 20%, and intelligent processing and analysis of multi-modal medical images are realized.
Owner:SHANDONG JUNKANGLIN MEDICAL TECHNOLOGY CO LTD

Magnetic core intelligent cutting parameter self-adaptive optimization system based on multi-mode sensing

The invention provides a magnetic core intelligent cutting parameter self-adaptive optimization system based on multi-mode perception, and relates to the technical field of data processing.The method comprises the steps that a multi-mode sensor module is integrated on magnetic core cutting equipment, and the module comprises a force sensor, a visual sensor and a temperature sensor; the acquisition units are respectively used for acquiring cutting force dynamic signals, cutting track image sequences and cutter temperature time sequence data in real time; magnetic core surface texture features and three-dimensional contour data are captured through a visual sensor, and an initial cutting parameter set is generated in combination with a magnetic core material type recognition result, associated parameters in a historical process database and preset process constraint conditions; and first workpiece trial cutting is executed based on the initial cutting parameter set, multi-modal data fusion collection is synchronously started, cutting force frequency domain feature vectors, a tool temperature change rate curve and cutting surface defect image features are obtained, and multi-modal data are obtained. According to the invention, multi-objective collaborative optimization of processing efficiency and energy consumption is realized.
Owner:BEIJING CRYSTAL MAGNETIC TECH CO LTD

Weldment welding seam automatic detection method and device based on machine vision

The invention discloses a weldment welding seam automatic detection method and device based on machine vision, and relates to the technical field of machine vision intelligent detection. The weldment welding seam automatic detection method and device based on machine vision comprises the steps that S1, surface images and forming feature data of a weldment are collected and preprocessed to construct a standardized image feature data set; s2, the boundary clearness of the weld joint is evaluated by combining the edge strength and the contour coherence, and the main contour extraction range is dynamically adjusted; s3, analyzing abnormal focusing characteristics of the candidate area, and adjusting a defect labeling range and a detection priority; and S4, integrating the boundary definition and the abnormal focusing features, analyzing the structure abnormality, and dynamically controlling and verifying a resource allocation strategy. The problems that in the weldment detection process, obvious light reflection and texture blurring phenomena exist in a heat affected area at a weld joint, a traditional image enhancement and edge extraction algorithm is difficult to stably recognize microdefects, and the credibility of a detection result is reduced are solved.
Owner:WUXI TIENENG PRECISION MASCH CO LTD

Unmanned aerial vehicle target detection method based on frequency-space joint attention and dynamic fusion

The invention relates to the technical field of computer vision detection, in particular to an unmanned aerial vehicle target detection method based on frequency-space joint attention and dynamic fusion, and the method comprises the steps: obtaining an unmanned aerial vehicle image data set, carrying out the preprocessing, and dividing a training set and a test set; constructing a target detection model, inputting the training set into the target detection model to extract image features, sequentially performing frequency domain detail enhancement, spatial domain salient region extraction and multi-scale feature adaptive fusion based on the image features, and establishing a feature sequence; screening the feature sequence to obtain an initial target query, and finishing target classification and positioning on the initial target query through a decoder; training a target detection model by using the training set, and inputting the test set into the trained target detection model to generate a detection result; on the premise that the real-time reasoning advantage of RT-DETR is kept as much as possible, the problems that in an unmanned aerial vehicle scene, a target is prone to missing detection, the scale change is large, the background is complex, and the target is fuzzy are effectively solved, and the detection precision is improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Cutting workpiece defect detection method and system based on image feature feedback

The invention discloses a cut workpiece defect detection method and system based on image feature feedback, and relates to the technical field of image processing.The method comprises the steps that cut workpiece technological characteristics are obtained, a preset defect type library is constructed, a hardware system is built, and parameters are initialized; synchronously acquiring a multi-view original image, and storing and associating annotation information; de-noising the original image, enhancing the contrast, and extracting a region of interest ROI; extracting texture, shape, edge and gray features from the ROI, and screening through a Relief-F algorithm to obtain an optimal feature subset; inputting into an SVM (Support Vector Machine) model for reasoning, and screening to obtain an effective defect detection result; and calculating an evaluation index and generating a feedback signal, and performing iterative optimization after adjusting parameters. The system comprises an acquisition module, a master control module, a data processing module and a display module. Through the precise design and closed-loop feedback of the whole process, the precision, efficiency and long-term adaptability of defect detection of the complex cutting workpiece are improved, and the industrial quality management and control requirements are met.
Owner:苏州艾克夫电子有限公司

Unmanned aerial vehicle image-based small object detection method for target areas

The present invention relates to the technical field of deep learning and computer vision. Disclosed is an unmanned aerial vehicle image-based small object detection method for target areas. The present invention crops images of obvious small objects in certain target areas, and annotates the small objects of different categories to form a raw training and testing dataset, so as to ensure the accuracy of data required in the early stage of the algorithm and further ensure the scientificity of the algorithm; uses the computing capability of an improved YOLOv7 detection model to collect image features of different degrees in the dataset, the improved YOLOv7 detection model using YOLOv7 as a basic model and adding to a neck network an MS-CET module, which is constituted by an improved self-attention mechanism and convolution module SPPCSP, and a BHC-FB module, which is constituted by bidirectional mixed convolution modules NConv and RPConv connected in parallel; and finally fuses different feature layers as a final judgment basis of an unmanned aerial vehicle for small object detection in the target areas, to further check the accuracy of the algorithm and criteria for dataset selection, thereby improving recognition accuracy.
Owner:CHONGQING UNIV OF TECH

Wind turbine generator data analysis and fault diagnosis method and system based on big data and artificial intelligence

The invention discloses a wind turbine generator data analysis and fault diagnosis method and system based on big data and artificial intelligence. According to the method, a blade image, a vibration signal, audio data and operation parameters are synchronously acquired through an unmanned aerial vehicle multi-mode sensor and a ground monitoring system, and a multi-source heterogeneous data set is constructed; after the data is classified and preprocessed, image features, vibration time-frequency domain features and operation parameter key value pairs are extracted respectively; dimensionality reduction is carried out by using an auto-encoder, feature-level space-time alignment is realized through an improved DTW algorithm, and a multi-dimensional fault feature matrix is generated; a hierarchical diagnosis model including a GRU auto-encoder, an MLP network and an attention mechanism CNN is constructed, and training is carried out by taking minimization of sub-model deviation as an optimization target; and finally, fusing multi-source features to realize fault classification, and generating a visual diagnosis report. According to the method, efficient fusion and accurate diagnosis of multi-source heterogeneous data are realized, and the accuracy and the real-time performance of fault detection of the wind turbine generator are remarkably improved.
Owner:NAT ENERGY GRP DONGTAI OFFSHORE WIND POWER CO LTD

Medical image segmentation method and system based on guiding information and multi-dimensional attention mechanism

The invention discloses a medical image segmentation method and system based on guidance information and a multi-dimensional attention mechanism. The method comprises the following steps: collecting an original dermatoscope image for preprocessing; constructing a segmentation model, wherein the segmentation model comprises a double-path image encoder, a guide information encoder and a mask decoder; the two-way image encoder is used for extracting local detail features and global context semantic information in the image; the guide information encoder is used for converting a coarse segmentation mask predicted by the last round of network into guide feature information; the mask decoder fuses the image feature information and the guide feature information, gradually restores and refines the coarse-grained feature map, and finally outputs an accurate lesion segmentation mask; constructing a loss function, and training the segmentation model by using the preprocessed data; and inputting a to-be-segmented original dermatoscope image into the trained segmentation model, and outputting a lesion region segmentation mask of the image. According to the method, the segmentation precision and the model generalization ability can be improved, and the multi-scale lesion processing ability is enhanced.
Owner:ZHEJIANG UNIV +1

Warehousing checking method based on multi-mode sensing technology, robot and warehousing system

The invention discloses a storage checking method based on a multi-modal sensing technology, a robot and a storage system, and belongs to the technical field of storage management and intelligent sensing fusion. Multi-modal data such as a visual image, space depth, radio frequency sensing and infrared temperature are collected, and an image feature vector, a three-dimensional point cloud model, a radio frequency response matrix and a temperature map are constructed; generating a fusion recognition vector through a multi-channel fusion network based on an attention mechanism, and dynamically adjusting a modal weight; constructing an article space distribution map, and marking a perception missing region; automatically complementing low-confidence region data based on a priority scheduling algorithm; performing joint verification on the original fusion result and the completion result to form a final inventory list; if the confidence coefficient of a certain article is lower than an early warning threshold continuously for multiple times, triggering an abnormal alarm and generating a traceable sensing sequence; the method is suitable for a high-precision inventory task in a complex storage scene, and has the advantages of high recognition robustness, intelligent completion mechanism, traceable abnormity and the like.
Owner:DIGITAL WHALE (SHANDONG) ENERGY TECH CO LTD

Cross-modal interaction image restoration method fusing text semantic guidance and visual structure prior

The invention discloses a cross-modal interactive image restoration method fusing text semantic guidance and visual structure priori, which comprises the following steps of: firstly, acquiring natural language description input by a user and an image to be restored, and generating a semantic segmentation map of the image through a semantic segmentation model; encoding the text and image semantics by using a pre-trained cross-modal encoding model to obtain text and semantic features; guiding a semantic alignment attention module through Prompt to realize deep fusion of multi-modal semantic features and image space features; structural enhancement and regulation of image features are realized by constructing a text guide weight graph, performing element-level modulation on the text guide weight graph and the optimized semantic segmentation graph, constructing a cross-modal structure semantic feature graph and generating a structural modulation factor; a four-stage image restoration network is adopted, and a high-quality restoration image conforming to semantic guidance and structure prior is generated step by step. According to the method, the semantic consistency, the structural integrity and the visual reality sense of an image restoration result are improved.
Owner:NANJING UNIV OF POSTS & TELECOMM

Deep forgery detection method based on visual language model

The invention discloses a deep forgery detection method based on a visual language model, and relates to the field of image forensics. The deep forgery detection method based on the visual language model aims to combine multi-source information to improve the discrimination capability of the model on a real image and a generated image. The method comprises the following steps: firstly, extracting image features through an image encoder of a pre-trained CLIP model; meanwhile, a frequency domain enhanced counterfeit perception adapter is embedded in the image encoder to mine potential anomalies of counterfeit images in the image domain and the frequency domain. Secondly, a manual feature extraction module is provided, discriminative low-dimensional features are extracted from the four aspects of the edge, the texture, the frequency and the symmetry of the image, and the discriminative low-dimensional features are used as auxiliary information input in the forgery detection process, so that the robustness and the interpretability of the model are improved; meanwhile, the text cue words are converted into feature vectors through a text encoder of a pre-training CLIP model; and finally, the model predicts a forgery score by calculating the cosine similarity between the image features and the text features so as to realize the discrimination of the authenticity of the image. According to the method, the problem that the detection capability of the model on the cross-dataset is insufficient is effectively improved.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Brain tumor multi-modal large model construction method and device, equipment and storage medium

The invention discloses a brain tumor multi-mode large model construction method, device and equipment and a storage medium, and is applied to the technical field of brain tumor imagines.The method comprises the steps that pixel-concept level alignment is conducted on a multi-mode MRI image and a pathological text; constructing a multi-modal feature fusion network for fusing image features and text features by adopting an attention mechanism of pathology perception and combining medical semantic information; training the multi-modal feature fusion network to generate an analysis report and a segmentation result; according to the technical scheme of multi-task cooperation, cross-modal pathological semantic accurate alignment, pathological knowledge graph injection and lightweight and continuous optimization parallelization, full-process coverage of brain tumor accurate segmentation, analysis report generation and prognosis prediction is achieved, the problems that a traditional model lacks pathological semantic support and is insufficient in clinical adaptability are solved, and the clinical adaptability of the traditional model is improved. And the deployment feasibility and the dynamic optimization capability are also considered.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Systems and methods for automatic medical report generation

The decision process of a first machine learning (ML) model may be explained based on a second ML model implemented on an apparatus. The apparatus may obtain a prediction about an image made based on the first ML model. The apparatus may further determine visual concepts associated with the image that may have been used by the first ML model to make the prediction, and determine respective contributions of the visual concepts to the prediction made by the first ML model. The apparatus may then generate, based on the second ML model, a textual description that explains the respective contributions of the visual concepts to the prediction made by the first ML model. The second ML model may determine respective image features associated with the visual concepts, map the determined image features to corresponding text features, and generate the textual description based at least on the text features.
Owner:SHANGHAI UNITED IMAGING INTELLIGENCE CO LTD

Small-sample defect identification method based on cross-modal text semantic driving

The invention discloses a few-sample defect identification method based on cross-modal text semantic driving, and belongs to the technical field of image processing. Aiming at the problem of insufficient generalization of a detection model caused by scarcity of abnormal samples and dynamic evolution of defect types in an industrial quality inspection scene, an unknown defect type can be accurately identified only by a small amount of normal data by establishing a dynamic feature recombination mechanism and an adaptive discrimination boundary; according to the method, a simulation sample similar to a real defect in form is generated on a normal sample through a matching relation between text description and image features; when a defect type which is not seen is encountered, a comparison standard of image textures can be automatically adjusted according to text semantics, subtle differences between a normal area and an abnormal area can be accurately distinguished, dependence on real defect data is not needed, and the sample defect identification precision is further improved.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Image annotation method and system applied to brain MRI (Magnetic Resonance Imaging) image segmentation

The embodiment of the invention discloses an image annotation method and system applied to brain MRI image segmentation, and the method comprises the steps: obtaining a brain MRI image data set of a target object, and the brain MRI image data set comprises original image sequences of a plurality of scanning levels; performing multi-modal feature fusion processing on the original image sequence to generate an enhanced image feature set; calling a multi-layer cascade segmentation network to perform hierarchical feature extraction on the enhanced image feature set to obtain a multi-scale anatomical structure feature map; and performing region boundary optimization processing based on the multi-scale anatomical structure feature map, and generating a marked brain structure segmentation image. Therefore, the boundary of each structure of the brain can be accurately defined, the segmented image is more accurate and clearer, and the image segmentation and marking of the brain MRI image can be accurately and clearer realized.
Owner:SHENZHEN NUCLEAR MAP MEDICAL TECHNOLOGY CO LTD

Intelligent fermentation process regulation and control method and system based on multi-modal perception

The invention relates to the technical field of data processing. The fermentation process intelligent regulation and control method and system based on multi-modal perception are provided, and the method comprises the following steps: carrying out image feature extraction processing on microorganism image data to generate a morphological feature vector, and carrying out metabolic feature dimension reduction processing on metabonomics data to generate a metabolic feature matrix; performing time sequence alignment processing to generate a fusion feature matrix, and performing abnormal marking processing on the metabonomics data to generate abnormal marking data; carrying out correlation intensity calculation processing on the morphological change of the microorganisms and the concentration fluctuation of the metabolites to generate a dynamic correlation intensity curve; constructing a cross-dimensional anomaly recognition model and a multi-modal collaborative prediction model, and generating a regulation and control parameter suggested value; the parameters of the multi-modal collaborative prediction model are updated through a feedback learning mechanism, the feature fusion weight of the fusion feature matrix is optimized, the accuracy of anomaly detection and regulation decision is improved, and the risk of stability fluctuation in the fermentation process is reduced.
Owner:HEBEI YIJIAEN INTELLIGENT TECH CO LTD

Image processing method applied to printed matter surface color difference detection

The invention discloses an image processing method applied to printed matter surface color difference detection, which relates to the technical field of image processing, and comprises the following steps: acquiring a digital image of a printed matter to be detected, performing illumination non-uniformity correction, and converting the corrected digital image into a CIELAB color space; performing region segmentation on the digital image based on double constraint conditions of color gradient and texture boundary to generate a detection region graph; feature parameters are extracted based on the detection area graph, and a feature data set is generated; establishing a dynamic reference model by utilizing process parameters and material characteristics of the printed matter; and calculating the distance between the feature data set and the dynamic reference model, identifying color difference regions and generating color difference scores, and screening and grading the color difference regions according to the color difference scores. According to the method, a multi-layer image pyramid and homomorphic filtering combined illumination correction technology is adopted, and an adaptive weight fusion mechanism based on image features is introduced, so that the illumination nonuniformity is effectively eliminated while the definition of printing details is kept.
Owner:GUANG ZHOU BEIDE PACKAGING & PRINTING CO LTD

Improved YOLOv5 protective equipment detection method combining channel selection and attention mechanism

The invention discloses an improved YOLOv5 protective equipment detection method combining channel selection and an attention mechanism, and the method comprises the steps: replacing an original spatial pyramid pooling fusion module SPPF of a YOLOv5 backbone network part with a channel selection multi-scale fusion module CSMF, and enabling the channel selection multi-scale fusion module CSMF to be based on a feature adaptive weighted fusion strategy of a multi-channel selection mechanism. According to context information and semantic distribution of input image features, fusion weights of feature channels under different receptive fields are dynamically adjusted, so that more discriminative multi-scale feature integration is realized. According to the invention, the adaptive capacity of the model to multi-scale targets is improved through the CSMF module, and the problem of inaccurate recognition of small targets such as gloves by a traditional model is solved; after the CBAM is adopted to enhance the MBconv module, the model can more effectively distinguish a key area and a background area in an image, and the robustness of target shielding and background interference in a complex scene is enhanced.
Owner:XIAN UNIV OF TECH

Concrete structure apparent defect identification method and device based on inspection robot

The invention discloses a concrete structure apparent defect recognition method and device based on an inspection robot, and the method comprises the steps: collecting an apparent image of a target concrete structure in real time through the inspection robot in response to an apparent defect recognition instruction of the target concrete structure; obtaining a defect identification model based on YOLOv5; inputting the processed apparent image into a defect identification model, carrying out global image feature extraction and context image feature extraction on the apparent image through each Transform module in the backbone network, and carrying out merging processing on the global image features and the context image features through an SPPF module; performing key feature enhancement on the merged image features through a channel attention layer and a space attention layer of each CBAM module in the neck network, performing feature fusion on the enhanced merged image features to obtain fused image features corresponding to each CBAM module, and performing defect identification on each fused image feature through the head network to obtain a fused image feature corresponding to each CBAM module; and obtaining a defect identification result output by each prediction head module.
Owner:BCEG CIVIL ENGINEERING CO LTD

Environment detection method and system based on multi-modal data fusion and deep learning

The invention provides an environment detection method and system based on a sample target detection model. The method comprises the following steps: synchronously acquiring an environment image, a video stream and physical parameters by using a multi-mode sensor; decomposing the data into image features and environmental parameter components through a dual-time sequence control signal, and realizing space-time alignment by adopting a linear phase filter; constructing a foreground region template based on the depth information, and generating target recognition feature representation containing an abnormal blurred target; adversarial training is carried out on the lightweight target detection network in combination with a transfer learning strategy, the network integrates convolutional features and a Transform attention mechanism, and the weight is dynamically adjusted through environmental parameters; fusing a target result and sensor data in real-time detection, and inputting a decision tree model for risk grading; and after the early warning is triggered, reconstructing a false detection sample through an online learning mechanism and iteratively optimizing the model. The system correspondingly comprises a multi-modal data acquisition module, a data enhancement and annotation module, a model training module, a real-time detection and fusion module and an early warning and optimization module. According to the invention, through multi-source data fusion, dynamic data enhancement and an adaptive compensation mechanism, the small target detection precision, the environmental adaptability and the real-time early warning capability are significantly improved.
Owner:SHANDONG HUANFA INSPECTION & TESTING CO LTD

Small target identification method and system for multi-modal fusion image in complex environment

The invention discloses a small target recognition method and system for a multi-modal fusion image in a complex environment, and belongs to the technical field of computer vision and image recognition, and the method comprises the steps: obtaining a visible light image, an infrared image and environment sensor data; image registration is carried out on visible light and infrared images, and a multi-scale image feature pyramid is constructed. And respectively extracting visible light and infrared image features to obtain visible light and infrared imaging feature data. And performing multi-modal data fusion on the visible light and infrared imaging feature data based on a cross-modal attention mechanism, and adaptively adjusting a fusion weight based on environmental sensor data to generate fusion features. And performing space-time enhancement processing on the fusion feature to obtain an enhanced fusion feature. And performing target tracking detection on the small target, and outputting position and category information of the small target. According to the method, the small target recognition capability in a severe environment is remarkably improved, and high precision and robustness can still be kept in a foggy, low-visibility and dark scene.
Owner:CHINA TOWER CO LTD +1

Extreme sea condition parameter identification system based on deep learning

The invention discloses an extreme sea condition parameter identification system based on deep learning, and relates to the technical field of ship navigation auxiliary equipment, in particular to a self-adaptive sea condition identification device which is used for acquiring image data and inertial measurement data of a current sea condition; the wave field visual depth estimation module is used for extracting visible light image features and infrared image features of a wave area from image data of the current sea condition, fusing the extracted visible light image features and infrared image features, using an encoder-decoder architecture and fusing an energy function to obtain a pixel-level wave height field, and outputting the pixel-level wave height field. The three-dimensional reconstruction of the wave surface is realized; the multi-modal data fusion module uses a filter dynamic model and a cost function to eliminate space-time asynchronous errors between inertial measurement data and visual perception data, performs multi-modal data fusion, and outputs wave field real-time parameterization information. According to the invention, the sea condition parameter real-time high-precision identification capability of the autonomous unmanned ship or the offshore carrying platform can be improved.
Owner:WUHAN UNIV OF TECH

Image relighting using machine learning

A method, apparatus, non-transitory computer readable medium, and system for image generation includes obtaining an input image and an input prompt, where the input image depicts an object and the input prompt describes a lighting condition for the object, generating relighted image features based on the input image and the input prompt, where the relighted image features represent the object with the lighting condition, and generating a synthetic image based on the relighted image features, where the synthetic image depicts the object with the lighting condition.
Owner:ADOBE INC

Multi-feature fusion diagnosis system and method for L1-L4 lumbar vertebra segments

The invention provides an L1-L4 lumbar vertebra segment-oriented multi-feature fusion diagnosis system and method, and the system comprises an image preprocessing module which is used for receiving a lumbar vertebra CT image sequence of a patient; a centrum anatomy partition module; the multi-dimensional image feature extraction module is used for extracting four types of quantitative features from each sub-region; the clinical multi-modal data coding module is used for independently acquiring and processing three types of clinical data: a multi-modal graph attention fusion network; and the segment-level diagnosis output module outputs diagnosis results of three levels. Through a parallel processing architecture and an optimized feature extraction algorithm, the whole diagnosis process only needs 45 seconds from data input to report generation, time is saved compared with manual film reading, and the consistency of diagnosis results is remarkably improved.
Owner:NANJING WANGSHI INTELLIGENT TECHNOLOGY CO LTD

Thyroid tumor diagnosis method and system based on ultrasonic and cytological image conjoint analysis

The invention discloses a thyroid tumor diagnosis method and system based on ultrasonic and cytological image conjoint analysis. The method comprises the following steps: converting a thyroid B ultrasonic image of a patient to be diagnosed into an image feature vector FUS; performing structured feature extraction on the thyroid cytological image of the patient to be diagnosed to obtain a structured feature vector FCYTO; the image feature vector FUS and the structured feature vector FCYTO are converted into a fusion Token sequence; and inputting the Token into a multi-modal feature fusion and prediction network based on a Transform architecture, and carrying out feature fusion and classification prediction so as to obtain the probability that the thyroid tumor of the patient to be diagnosed is malignant. The thyroid tumor diagnosis based on multi-modal fusion is carried out on the basis of ultrasonic and cytological images, so that the diagnosis accuracy is improved.
Owner:金凤实验室

Semi-supervised target detection method for visible light-infrared multi-mode fusion scene

The invention provides a semi-supervised target detection method for a visible light-infrared multi-mode fusion scene. The method comprises the following steps: constructing a semi-supervised visible light-infrared multi-modal image data set based on an LLVIP data set; on the basis of a YOLOv11 model architecture, constructing a target detection model oriented to multi-modal image feature fusion, and training the target detection model by using a semi-supervised visible light-infrared multi-modal image data set in a deep learning end-to-end mode to obtain a trained target detection model; and inputting a to-be-detected multi-modal image into the trained target detection model, and outputting a target detection result of the to-be-detected multi-modal image by the trained target detection model. The method is based on a semi-supervised learning normal form, so that the precision and robustness of target detection in a multi-modal fusion scene are improved, and the requirements for high efficiency and reliability of target recognition in practical application scenes such as intelligent traffic and intelligent security and protection are met.
Owner:BEIJING JIAOTONG UNIV

Target detection method based on YOLO model, electronic equipment and storage medium

The invention discloses a target detection method based on a YOLO model, electronic equipment and a storage medium, and relates to the technical field of target detection. Comprising the following steps: inputting an aerial image of an unmanned aerial vehicle into a trained target YOLO model; the target YOLO model comprises a backbone network, a neck network and a head network, and a feature extraction module in the backbone network performs multi-scale feature extraction by adopting a double-branch architecture attention mechanism; performing multi-scale feature extraction on the aerial image by adopting a dual-branch architecture attention mechanism through a feature extraction module in the backbone network, and constructing to obtain a plurality of layers of first comprehensive image features of the aerial image; performing feature fusion on the first comprehensive image features of different levels through a neck network to obtain second comprehensive image features of multiple levels; and inputting the multiple levels of second comprehensive image features into a head network to obtain a target detection result of the aerial image. According to the invention, the accuracy of small target detection can be improved.
Owner:HUNAN UNIV OF TECH

Video multi-mode sentiment analysis method and device based on multi-layer perceptron fusion

The invention discloses a video multi-mode sentiment analysis method and device based on multi-layer perceptron fusion, and relates to the technical field of sentiment analysis, and the method comprises the steps: S1, extracting text features, image features and audio features in a video; extracting time sequence information in the image features and the audio features to obtain time sequence image features and time sequence audio features; s2, constructing a video multi-modal sentiment analysis model comprising a multi-modal feature capture module, a multi-layer perceptron fusion module and a sentiment classifier, and constructing a loss function according to modal similarity and modal heterogeneity between modals; s3, training the model; and S4, inputting the text features, the time sequence image features and the time sequence audio features into the trained model to obtain emotion polarity probability distribution. According to the method, similarity loss and heterogeneity loss are constructed, and sequence, channel and modal dimension fusion is carried out by using a multi-layer perceptron, so that the calculation complexity and memory consumption are reduced, and the integrity and discrimination capability of the multi-modal emotion features are improved.
Owner:HUAQIAO UNIVERSITY

Three-dimensional shielded target tracking method based on multi-modal space-time interaction

The invention discloses a three-dimensional shielding target tracking method based on multi-modal space-time interaction, and relates to the technical field of target tracking. The method comprises the following steps: acquiring a point cloud and an image and preprocessing to obtain global fusion features; obtaining an initial detection frame and region-of-interest features through region proposal network processing; projecting the non-empty voxel point cloud to the image features, and reconstructing shielded target features; convolution and neural network processing are utilized to obtain a refined detection frame; screening legal detection frames through distance calculation and legality judgment; the bipartite graph and the self-adaptive channel graph are adopted for convolution, and appearance correlation scores are calculated; and matching the detection frame and the trajectory based on a Hungary algorithm to realize whole-course tracking. The target identification accuracy and robustness are improved, the shielding problem is solved, the accuracy of the detection frame is ensured, the correlation accuracy is improved by using the bipartite graph and the adaptive convolution, the nodes are matched in combination with the geometric cost matrix, and whole-course tracking and error calibration are realized.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Optical remote sensing image feature landslide information extraction method

The invention relates to the technical field of earthquakes, in particular to an optical remote sensing image feature landslide information extraction method. A double-time-phase NDVI time sequence analysis strategy is adopted, and landslide area recognition is achieved by constructing a vegetation index difference chart before and after an earthquake. A landslide mass area before an earthquake presents a high NDVI value due to complete vegetation coverage, and a vegetation index of the area in an image after a disaster is significantly attenuated due to surface disturbance, so that an obvious change response is formed. If the normalized vegetation index does not change, a non-landslide area can be determined, and the greater the change of the normalized vegetation index is, the greater the possibility of landslide occurrence is, and a landslide preselection area is determined; then Otsu threshold segmentation is carried out by combining cloud layer features and water body features, cloud layer and water body change parts are effectively eliminated, object-oriented geometric shape rule fine recognition is carried out on a preselected area, a DEM is adopted to calculate a slope value to constrain low-lying or flat earth surface changes, and efficient and accurate landslide information extraction is achieved.
Owner:SEISMOLOGICAL BUREAU OF GANSU PROVINCE CHINA EARTHQUAKE ADMINISTRATION