Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

36 results about "Modal method" patented technology

Tea withering intelligent control system and method based on deep learning and multi-modal fusion

The invention discloses an intelligent tea withering control system based on multi-modal feature fusion and time sequence prediction. The system is composed of a multi-modal feature extraction module, a time sequence modeling module, a transfer learning module and an intelligent decision control module, a Transform attention mechanism is adopted to construct a cross-modal fusion framework, RGB images, hyperspectrum and time sequence information are fused, and accurate recognition and trend prediction of the withering state are achieved. According to the system, an adaptive attention fusion network is designed, optimal fusion of images, spectrums and grade information is realized through dynamic distribution of modal weights, and the recognition accuracy and stability are remarkably improved. The classification accuracy of 180 verification samples reaches 93.33%, and is improved by 15.93%-34.83% compared with that of a single-mode method. And cross-environment self-adaption is realized through fusion transfer learning, and the performance is improved by 6.7%-14.3%. The intelligent decision engine optimizes temperature and humidity parameters based on a multivariable coupling control theory, the control precision reaches + / -0.5 DEG C and + / -2.0%, and the response time is 2.5 seconds.
Owner:JIANGSU OCEAN UNIV +1

Ship propulsion shafting fault diagnosis method based on multi-modal attention fusion

The invention discloses a ship propulsion shafting fault diagnosis method based on multi-modal attention fusion, and relates to the technical field of ship fault diagnosis. According to the method, multi-modal data such as vibration parameters, lubricating oil parameters and cooling water parameters of key parts of a ship propulsion shafting are synchronously collected and preprocessed; a modal specific feature extraction strategy is adopted, a multi-modal attention fusion mechanism including intra-modal self-attention, inter-modal cross attention and adaptive dynamic weight distribution is designed, and heterogeneous modal features are effectively integrated. Compared with a single-mode method and a traditional feature splicing method, the method has the advantage that the diagnosis accuracy is remarkably improved. The method has high robustness, and can still maintain high diagnosis accuracy even under the condition that part of sensors fail.
Owner:CHINA SHIP SCIENTIFIC RESEARCH CENTER

Classroom teacher teaching performance description method, model and system based on multi-modal data fusion and storage medium

The invention relates to the technical field of data processing, in particular to a classroom teacher teaching performance description method, model and system based on multi-modal data fusion and a storage medium. By introducing a multi-modal data fusion strategy, cooperative processing of classroom teacher visual information, audio information and text information is realized, and the accuracy and time sequence continuity of teacher target perception are significantly improved. Compared with a traditional single-mode method, the method not only can extract the posture, expression and action characteristics of the teacher from the visual mode, but also can recognize the voice emotion and the side language signal from the audio mode, and achieves the full-dimensional description of the teaching behavior of the teacher in combination with the text semantics. The objective of the invention is to solve the problem of how to perform multi-dimensional evaluation on teaching performance of a classroom teacher based on data of multiple modalities.
Owner:YUNNAN NORMAL UNIV

Image fuzzy detection method based on fusion of frequency domain analysis and deep learning

PendingCN120807454AImage enhancementImage analysisOptical flowModal method
The invention provides an image fuzzy detection method based on frequency domain analysis and deep learning fusion. The method comprises the following steps: S1, frequency domain feature extraction and quantification; s2, spatial domain feature extraction and modeling; s3, carrying out multi-modal feature fusion; s4, joint optimization and post-treatment are carried out; s5, outputting and verifying; through complementarity design of frequency domain and deep learning, complex fuzzy detection requirements of static images and video streams are covered, high efficiency and reliability are verified in industrial quality inspection, video conferences and other scenes, energy attenuation characteristics caused by global blur are accurately captured through frequency domain analysis, motion blur and out-of-focus blur are effectively distinguished, and the method is suitable for large-scale popularization and application. According to the method, local texture degradation of deep learning network modeling, dynamic track abnormity analysis of an optical flow network and complex scenes covering static images and video streams are realized through a bidirectional feature fusion mechanism, the mAP of mixed fuzzy detection is effectively improved compared with a single-mode method, and dynamic fuzzy and static out-of-focus fuzzy are effectively distinguished.
Owner:YIREN (SHANGHAI) TECH CO LTD

Methods and systems for real time video driven human 3-d posture estimation

PendingUS20250336236A1Image enhancementImage analysisEngineeringModal method
The disclosure relates generally to methods and systems for real time video driven human 3-dimensional (3-D) posture estimation during physical activities. Conventional techniques do not exploit temporal information, they do not give smooth transition of postures over time. Furthermore, the techniques that exploit the temporal information suffer from higher time requirements due to two state computations. The present disclosure solves the technical problems in the art with the methods and systems for real time video driven human 3-D posture estimation during physical activities. The present invention discloses a smart-phone camera based automatic posture monitoring system designed with an auto-encoder based architecture. The disclosed auto-encoder based cross-modal method uses monocular video (2-D image sequences) from a single low-end mobile device (for example, smart-phone camera) for estimating human 3-D posture in real time (˜5 fps) with high accuracy (less than 1 cm error per joint location).
Owner:TATA CONSULTANCY SERVICES LTD

Multi-modal document analysis method and device, electronic equipment and storage medium

The invention relates to the technical field of multi-modal document analysis, and discloses a multi-modal document analysis method and device, electronic equipment and a storage medium. The method comprises the steps that picture elements in a document are automatically recognized and extracted through a third-party library module, placeholders are generated in combination with picture positions, and the placeholders are stored in the third-party library module; and then semantic understanding and text generation are carried out by using a multi-modal large model, so that the visual content is converted into a structurable information text. The method has the advantages that by introducing the multi-modal large language model, unified analysis of multi-modal information such as texts, pictures, tables and flow charts is achieved, layout semantics and logic structures of documents are reserved, fusion analysis of image-text content is achieved, the integrity and the intelligent level of document analysis are remarkably improved, and the method is suitable for large-scale popularization and application. Compared with a traditional single-mode method, the method has the advantages that unified understanding and structured expression can be carried out on multi-mode content, and richer and more accurate structured results can be output.
Owner:GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)

A cross-modal method for visual recognition in large-scale point cloud maps

The present application relates to a kind of cross-modal methods for visual identification in large-scale point cloud map, comprising: based on the RGB image data and point cloud data collected in global map, construct data set;The data set is converted cross-modal;Using converted data set, the preset network model is trained, and cross-modal positioning model is obtained;Wherein, the preset network model includes: multi-scale feature encoder, cascaded cross attention module and projection converter;RGB image data not included in the data set is input into the cross-modal positioning model, and the position of sensor in global map is obtained.The present application can significantly improve the image-to-point cloud cross-modal location recognition accuracy and stability in unknown indoor and outdoor environment, and simultaneously due to its lightweight design, increase the feasibility of the cross-modal location recognition device in practical application deployment.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A simulation data generation method of metal material multi-modal fusion

The application belongs to the field of material data, and discloses a simulation data generation method of metal material multi-modal fusion. The method comprises the following steps: collecting data, extracting features to obtain a feature vector, fusing the feature vector through a multi-modal method to obtain a metal material multi-modal feature training data set, constructing a metal material simulation data diffusion model and a diffusion model loss function meeting physical condition constraints, training the diffusion model, and outputting high-quality simulation metal material data meeting the basic physical law of the material. The simulation data generation method of metal material multi-modal fusion provided by the application introduces physical constraints, so that the generated simulation data meets experimental data from a statistical perspective and meets the basic physical law. The method can effectively generate simulation data of metal materials under the condition of limited data, significantly reduces the experimental and calculation costs, and simultaneously fuses numerical and image modal data, which is more in line with the organization-property evolution relationship of the material.
Owner:HUAZHONG UNIV OF SCI & TECH

Multimodal transport method and system suitable for salt mist environment

The invention discloses a multimodal transport method and system suitable for a salt mist environment, and the method comprises the steps: building a multimodal transport salt mist adaptive transport capacity grid model with city / hub nodes as vertexes and connected road sections as edges; configuring a salt mist adaptation device at the node, and marking the node protection capability, the availability time window and the supply capability; meanwhile, a model hypothesis is set, and cargo transportation integrity, facility availability, transportation rules and a parameter acquisition mode are determined; defining a multi-dimensional decision variable, a parameter and a state quantity; incorporating the salt mist environment influence into cost composition, and constructing a salt mist correlation total cost model; constructing a mixed integer programming model containing path connection and flow conservation, protection capability constraint, energy / endurance constraint and corrosion / exposure constraint by taking the minimum total cost as a target; and designing an improved hybrid immune-particle swarm dual optimization algorithm, and carrying out optimization solution, optimal path output, mode selection and protection parameter configuration through hybrid coding, security repair operators and a memory mechanism.
Owner:HUBEI UNIV OF TECH

Fatigue gait pattern recognition device and method based on multi-modal sensors

The present application relates to a fatigue gait pattern recognition device and method based on a multi-modal sensor, multi-modal sensor data is collected based on a multi-modal sensor, gait segmentation is performed after preprocessing, data samples after gait segmentation are obtained, including multi-modal sensor data and corresponding data features; an improved feature fusion model is constructed, and the model is trained to be stable with data samples; multi-modal sensor data is collected, and the preprocessed data is input into the trained model, and the fatigue gait pattern recognition result is output; the device includes two groups of multi-modal sensors for collecting multi-modal sensor data, and the synchronous transceiver device sends the corresponding two groups of multi-modal sensor data after matching; the controller acquires the multi-modal sensor data and recognizes the fatigue gait pattern. The present application realizes high-precision classification of fatigue gait pattern, and the accuracy is improved by 15%-20% compared with traditional single-mode method; based on dynamic threshold classification and multi-dimensional feature analysis, user-specific rehabilitation suggestions are generated, and walking ability and safety are improved.
Owner:ZHEJIANG UNIV OF TECH

Vision-point cloud multi-modal fusion-based scene identification method and system

The invention discloses a scene identification method and system based on vision-point cloud multi-modal fusion, and the method comprises the steps: collecting a visual image and laser radar point cloud data in a target scene, carrying out the preprocessing, and extracting corresponding visual image features and point cloud features; generating a global feature, a visual modal local feature and a point cloud modal local feature in the target scene; performing global retrieval in a preset scene database by utilizing the global features to obtain a plurality of candidate positions and corresponding global matching scores; performing local matching and geometric verification on the candidate positions based on the visual modal local features and the point cloud modal local features to obtain a local matching score of each candidate position; and reordering the candidate positions according to the global matching score and the local matching score, and outputting an optimal scene recognition result. The method solves the problems of poor feature alignment, insufficient modal fusion, weak complex scene adaptability and data redundancy in the existing multi-modal method.
Owner:BEIHANG UNIV

Multi-modal false comment identification method and system

The invention discloses a multi-modal false comment recognition method and system. The method comprises the steps that text data and behavior data are preprocessed to generate word embedding vectors and behavior feature vectors; inputting the word embedding vectors into a plurality of multi-scale space weighted convolution modules in parallel to generate multi-scale semantic features; inputting the multi-scale semantic features into a multi-scale context aggregation module to generate a semantic feature fusion vector; inputting the semantic feature fusion vector into a BiLSTM-Attention module to generate a semantic enhancement vector; and after the behavior feature vector and the semantic enhancement vector are spliced, inputting into a graph Laplacian layer for classification, and outputting a false comment identification result. False comment identification is carried out through a multi-modal method, the defects of a traditional single-modal method are overcome, and the identification accuracy and generalization ability are improved.
Owner:SHENZHEN UNIV

Spectrum library optimization method and system based on sensitivity analysis

The invention discloses a spectrum library optimization method and system based on sensitivity analysis, and belongs to the field of semiconductor optical measurement, and the optimization method comprises the steps: obtaining a simulation spectrum; carrying out sensitivity analysis on preset library building parameters based on the simulated spectrum, and obtaining the sensitivity of each library building parameter; obtaining a parameter value range of each library building parameter, and obtaining a division weight of each library building parameter based on the sensitivity of each library building parameter and the parameter value range of each library building parameter; calculating the grid division number of each library building parameter based on the division weight of each library building parameter; grid points are obtained according to the grid division number of each library building parameter, each grid point corresponds to a set of geometric parameters, and a parameter network is obtained; and geometric parameters corresponding to each grid point in the parameter network are simulated through a Fourier modal method, and an optimized spectrum library is generated. The spectrum library sample obtained through the optimization method is balanced in distribution and small in deviation, and the precision and stability of the neural network model are improved.
Owner:SHENZHEN ANGSTROM EXCELLENCE TECH CO LTD

Dual-mode fusion wavelet enhancement upper six pieces of hyperspectral classification method and system

This invention discloses a dual-modal fusion wavelet-enhanced six-image hyperspectral classification method and system. The method includes: acquiring spectral and RGB image data of the six images; extracting features from the spectral and RGB image data respectively; enhancing the feature-extracted data using a learnable wavelet enhancement module; flattening and stitching the enhanced data; fusing the stitched feature sequences to obtain a fused feature sequence; embedding location information and category labels into the fused feature sequence, and feeding it into a Transformer encoder for global context modeling and classification. This invention's dual-modal fusion wavelet-enhanced six-image hyperspectral classification method and system achieves the fusion and utilization of hyperspectral imaging, near-infrared spectroscopy, and digital image features through feature-level and data-level fusion, fully utilizing the spatial-spectral information of multiple data sources, leveraging the advantages of various data, and effectively overcoming the limitations of single-modal methods.
Owner:CHINA TOBACCO HENAN IND CO LTD

Multi-modal method for classifying thyroid nodule based on ultrasound and infrared thermal images

ActiveUS12482248B2Image enhancementImage analysisRadiologyModal method
The present disclosure provides a multi-modal method for classifying a thyroid nodule based on ultrasound (US) and infrared thermal (IRT) images. Based on ultrasound and infrared thermal images and in combination with a multi-modal learning method, the present disclosure provides an adaptive multi-modal hybrid (AmmH) model which is composed of three parts: an intra-modal hybrid encoder (HIME), an adaptive cross-modal encoder (ACME), and a multilayer perceptron (MLP) head. The HIME is capable of modeling a global feature while extracting a local feature. The ACME is capable of customizing personalized modality-weights according to different cases and performing information interaction and fusion of inter-modal features. The MLP head classifies a fused feature obtained. The method enables the AmmH model to automatically classify a thyroid nodule of a subject based on ultrasound and infrared thermal images of the subject, providing a doctor with an objective and accurate classification result to assist diagnosis.
Owner:WUHAN UNIV

Multi-modal method for interacting with 3D models

PendingUS20260253327A1SimulationClassical mechanics
The present disclosure concerns a methodology that allows a user to “orbit” around a model on a specific axis of rotation and view an orthographic floor plan of the model. A user may view and “walk through” the model while staying at a specific height above the ground with smooth transitions between orbiting, floor plan, and walking modes.
Owner:COSTAR REALTY INFORMATION INC

Multi-modal network few-sample image classification method based on visual text prompt

The invention relates to the technical field of image classification, in particular to a visual text prompt-based multi-modal network few-sample image classification method, which comprises the following steps of: acquiring an image, a learnable prompt and a manual prompt; obtaining a frequency domain prompt fusion feature map; obtaining an original image feature vector; inputting the original image and the frequency domain visual fusion feature map into an image encoder to obtain a map frequency domain prompt fusion feature vector; respectively inputting the learnable prompt and the manual prompt into a text encoder to obtain a learnable prompt feature vector and a manual prompt feature vector; calculating a similarity score between the original image feature vector and the manual prompt feature vector; calculating a similarity score of the graph frequency domain prompt fusion feature vector and the learnable prompt feature vector; and constructing the regularization of the minimum divergence to the consistency among the vision, the text prompt and the distribution. According to the method, the problem of coordination and integration of different modal information and consistency of cross-modal semantics in an existing multi-modal method is solved.
Owner:CHANGZHOU UNIV

Expressway scene-oriented interpretable hierarchical reasoning multi-modal method and system

The invention belongs to the technical field of electric digital data processing, and particularly relates to an expressway scene-oriented interpretable hierarchical reasoning multi-modal method and system, and the method comprises the steps: obtaining multi-modal data, and carrying out the preprocessing of the multi-modal data; performing feature coding and semantic fusion on the preprocessed multi-modal data to form uniform features; performing hierarchical recursive collaborative reasoning on the unified features, including performing global semantic understanding and long-term reasoning through a high-level part to obtain a final semantic state; data updating, fine-grained calculation and supplementary inference are carried out through the low-layer part; a natural language answer is generated through an encoder based on the output of the high-level part and the low-level part under the condition of the final semantic state, and the contribution proportion of the natural language answer is calculated by calculating the interpretation distribution of the high-level part and the low-level part; after interpretability analysis is carried out on the natural language answers, causal activation mapping calculation is carried out on each generated result, and a visual interpretation graph is obtained.
Owner:SHANDONG HI SPEED GRP CO LTD +1

Method for monitoring three-dimensional large deformation of beam structure based on combination of modal method and geometric precision beam

PendingCN121168119AGeometric CADMeasurement devicesMixed beamAlgorithm
The invention relates to a modal method and geometric precision beam mixed three-dimensional large deformation monitoring method for a beam structure, belongs to the technical field of structure health monitoring, solves the problems of high generalized strain calculation complexity of the structure and too high requirement on the mounting position of a sensor in the prior art, and comprises the following steps: S1, initializing calculation; mathematical representation is carried out on the geometric configuration of the to-be-monitored beam structure, and a global coordinate system and a local coordinate system are established; s2, obtaining generalized strain of a beam center reference line of the to-be-monitored beam type structure based on a modal superposition method; s3, calculating a spatial position vector of a beam center reference line; s4, calculating a spatial position vector of any point on the to-be-monitored beam structure after deformation; and S5, outputting the obtained spatial position vector of the to-be-monitored beam structure after deformation as a monitoring result, and providing the monitoring result to a structure instability risk early warning program.
Owner:BEIHANG UNIV

Industrial video anomaly identification method and device, electronic equipment and storage medium

The invention relates to an industrial video anomaly recognition method and device, electronic equipment and a storage medium, and the method comprises the steps: inputting a multi-view video, sensor data and equipment operation data into a visual language model, obtaining a depth feature representation, carrying out the multi-modal depth fusion of the depth feature representation through an industrial equipment knowledge graph, and obtaining an industrial equipment anomaly recognition result. A multi-modal fusion sequence is obtained; performing global context information enhancement on the features of the multi-modal fusion sequence through a space-time diagram to obtain final feature representation; performing video anomaly detection on the final feature representation, and outputting a video anomaly detection result; multi-view videos, sensor data and equipment operation data are integrated, multi-modal feature extraction is realized by using a visual language model, deep fusion is performed in combination with an industrial equipment knowledge graph, global context information is enhanced by means of a space-time diagram, the problem of incomplete feature representation of a traditional single-modal method is effectively solved, and the method is suitable for large-scale popularization and application. The method has the advantage that the industrial video anomaly detection accuracy and the system robustness are improved.
Owner:WUHAN SURVEYING GEOTECHN RES INST OF MCC +1

Frequency Domain Modal Method for Determining Stability of Vehicle-to-Grid Oscillations in New Energy Vehicles Applicable to Converter Impedance Measurement

PendingCN122371132AConvertersNew energy
This invention discloses a frequency-domain modal method for determining the stability of vehicle-to-grid oscillations in new energy systems, applicable to converter impedance measurement. Specifically, it involves: acquiring the topology and component parameters of an electrified railway traction power supply system; performing frequency scanning on the single-input single-output frequency-domain impedance obtained through measurement or mathematical modeling in the frequency band to be analyzed, forming the system frequency-domain node admittance matrix at each frequency; performing eigenvalue decomposition to obtain the amplitude-frequency curve of the modal admittance or modal impedance; determining the oscillation frequency based on the extreme values ​​of the curve; determining the system stability based on the real and imaginary parts of the modal admittance or impedance at the oscillation frequency; and calculating the participation factor of each node in the unstable oscillation mode. This invention is applicable to low-frequency oscillation stability analysis of traction power supply systems with arbitrary topologies that do not contain three-phase converters. The SISO frequency-domain impedance of the single-phase voltage source converter interface power supply and locomotive load can be directly measured through impedance scanning without the need for dq-domain conversion.
Owner:SOUTHWEST JIAOTONG UNIV

Pipeline management method based on modal displacement and electronic equipment

The invention relates to the technical field of engineering machinery, and provides a pipeline management method based on modal displacement and electronic equipment, a multi-body dynamics-fluid pressure coupling simulation model is constructed based on a hydraulic excavator working device pipeline system, and target position parameters of a lug are obtained through a free modal method and a modal displacement algorithm. The multi-mode superposition effect principle is applied to layout optimization of the engineering machinery pipeline lug, accurate optimization of the position of the pipeline lug is achieved by modeling and analyzing the coupling vibration problem of a multi-order mode, and the problem of stress concentration caused by resonance is effectively avoided. Furthermore, actually-measured acquisition data are acquired through load spectrum acquisition, dynamic response analysis is performed to execute dynamic monitoring and early warning, and a dynamically-optimized and continuously-monitored closed-loop system is established through comparative analysis of simulation and actually-measured data, so that the simulation precision is improved, dynamic adjustment and early warning of actual working conditions are realized, and the working efficiency is improved. And the applicability and the reliability of the system are obviously improved.
Owner:LIUZHOU LIUGONG EXCAVATORS CO LTD +2

Multimodal data fusion counterfeit content identification method and device

The invention discloses a multi-modal data fusion counterfeited content identification method and device, the device comprises an identification module and an indicating lamp installed at the front side of the identification module, the right side of the identification module is provided with a response module, and the lower side of the identification module is provided with a processing module. A main processor is arranged in the outer wall of the upper end of the processing module, an acquisition module is arranged on the right side of the recognition module, and an acquisition module group is arranged on the outer wall of the front end of the acquisition module so as to acquire corresponding data information, and spatial forgery traces, such as facial fusion boundary and inconsistent illumination, of images are extracted through independent modal encoders; if the inter-frame motion of the video is abnormal, such as the head posture is not coherent and the voiceprint of the audio is deviated from the lip language, the detection accuracy is improved by using cross-modal complementary information. In an FF + + data set test, the AUC index of a complex forged sample is improved by 12-18% compared with a single-mode method.
Owner:郑竹青

Non-contact human respiration rate measurement method based on video and frequency modulated continuous wave radar information fusion

The application discloses a kind of non-contact human respiratory rate measurement methods based on video and frequency modulation continuous wave radar information fusion, comprising:1 respectively processes video and radar bimodal data to obtain pixel motion trajectory and distance angle chart time sequence;2 video measurement part uses spectral subtraction and principal component analysis technique to carry out pretreatment, radar measurement part uses static clutter removal and average filtering technique to carry out pretreatment, two kinds of modal all use empirical mode decomposition technique to carry out single modal measurement;3 feature level fusion, after bimodal signal using multivariate singular spectrum analysis extraction shared respiratory signal is pretreated;4 decision level fusion, according to the signal-to-noise ratio weighted frequency domain estimation value summation of feature level fusion result and single modal measurement result obtains final measurement result.The application can measure respiration under a variety of spontaneous motion scene, measurement result is compared with single modal method and has significant improvement, can effectively expand the application range of video and radar fusion measurement.
Owner:HEFEI UNIV OF TECH

Urban air-ground radio map redrawing method based on multi-modal information fusion

This invention presents a multimodal information fusion-based urban air-to-ground radio map redrawing method. By integrating prior knowledge from partially incomplete maps with a small amount of RSS measurements, a more comprehensive and robust information representation can be obtained. Through modal fusion of multi-class occlusion wireless channel models and multi-class virtual obstacle map models, wireless signal features and geographic information features are used together to redraw the radio map, better capturing their correlation and complementarity. The method of this invention is experimentally validated using an open-source RSS measurement dataset and compared with single-modal methods to evaluate its beneficial effects. The evaluation results demonstrate the advantages of this invention in wireless signal reception strength estimation and virtual obstacle map redrawing tasks, and show its effectiveness in reducing the amount of data required for the task and improving redrawing accuracy.
Owner:SHENZHEN UNIV

Near-infrared spectrum feature selection method and device, electronic equipment and storage medium

PendingCN122451412AAlgorithmFt ir spectra
The application relates to a near-infrared spectrum feature selection method and device, electronic equipment and a storage medium, wherein the method comprises: purifying an original spectrum to obtain spectrum data; performing parallel calculation of multiple heterogeneous strategies based on the spectrum data to generate an importance curve meeting a preset diversification condition; extracting a consistency index based on the importance curve, determining a basic weight according to the consistency index, and correcting the basic weight to generate a final importance curve for selecting near-infrared spectrum features according to the corrected basic weight. Thus, the problems in the related art that the single modal method has inherent bias, is sensitive to data disturbance and cannot balance resolution and smoothness, leading to incomplete feature selection, poor model stability, performance fluctuation under limited sample size or data disturbance, and lack of systematic multi-modal fusion strategy, so that the feature selection accuracy cannot be guaranteed while the stability and interpretability of the results are considered.
Owner:GUODIAN ENVIRONMENTAL PROTECTION RES INST CO LTD +3

Artificial intelligence industrial visual inspection method, system and equipment

The invention belongs to the technical field of artificial intelligence, and provides an artificial intelligence industrial visual detection method, system and equipment, and the detection method comprises the steps: image collection and multi-modal data acquisition, defect sample generation, dynamic preprocessing and noise suppression, multi-modal feature fusion, and dynamic balance defect detection. According to the method, 2D HDR images, 3D point cloud and polarized light data of a product are synchronously collected, reflective interference is inhibited, microdefect contrast is enhanced, multi-modal data fusion is matched with cross-scale analysis, inter-modal feature alignment and intra-modal multi-scale defect capture are achieved, the limitation of a traditional single-modal method is solved, the defect omission ratio and the defect false detection rate are reduced, and the product quality is improved. Diversified samples are generated based on real defect data, physical simulation is combined to verify defect rationality, rare defect training data volume is increased, data distribution of different production lines is aligned through adversarial training, cross-line detection accuracy fluctuates, synthetic data cooperation domain self-adaption is achieved, and the detection rate of rare defects in a small sample scene is increased.
Owner:INNER MONGOLIA FINANCE AND ECONOMICS UNIVERSITY +1

Space-time-frequency decoupled and frequency-domain enhanced cross-modal video adversarial noise generation method

The application discloses a kind of spatio-temporal frequency decoupling and frequency domain enhancement cross-modal video confrontation disturbance generation method, belong to computer vision technical field, including steps: obtaining the video frame sequence containing clean image, initialize each frame disturbance;According to clean image and disturbance, generate confrontation frame and generate enhanced confrontation frame;The intermediate features of clean image and enhanced confrontation frame are extracted respectively;Constitute spatial high-frequency component and time low-frequency component;Constitute joint loss;With video frame sequence to minimize joint loss update disturbance to generate the optimal disturbance of each frame;With optimal disturbance generation optimal confrontation frame constructs confrontation video.The application is through the multi-component feature constraint of spatio-temporal frequency decoupling and optimization stage frequency domain-scale joint enhancement, effectively overcome the deficiency of existing frame-by-frame cross-modal method in migration stability and representation destruction comprehensiveness, can better meet the actual demand in video model security and robustness evaluation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Alzheimer's disease early analysis method based on multi-modal features and neural network

The invention discloses an Alzheimer's disease early analysis method based on multi-modal features and a neural network, and belongs to the technical field of neural engineering and artificial intelligence crossing. Comprising the following steps: step 1, acquiring multi-modal data; 2, data preprocessing is carried out; and step 3, feature extraction and feature fusion are carried out. 4, two parallel training routes are adopted in the modeling stage. And 5, carrying out multi-modal model fusion. And 6, outputting a result. According to the method, joint modeling can be carried out on EEG, MRI and gene multi-modal data, and Alzheimer disease risk prediction can be rapidly completed. Compared with a traditional single-mode method, the scheme has remarkable advantages in the aspects of accuracy and robustness, the misjudgment rate can be reduced, and a tool with higher application value is provided for clinical early diagnosis.
Owner:BEIJING UNIV OF TECH

A multi-granularity medical text information guided 3D multi-modal fusion method

The present application relates to the technical field of data processing, in particular to a 3D multi-modal fusion method guided by multi-granularity medical text information, comprising: extracting key description information in 3D images and image reports as input data, encoding to obtain a vector corresponding to the input data; performing image feature extraction and aggregation based on KD at a fine-grained level to obtain slice-level features; multiplying the slice importance score obtained based on coarse-grained FT with the slice-level features to obtain final overall 3D image features; fusing the overall 3D image features, residual data stream features and complete text vectors to obtain multi-modal features for classification; and inputting the multi-modal features into a classification head to obtain model prediction output. The present application makes up for the lack of text utilization in 3D medical multi-modal methods, solves the problems of difficulty in accurately identifying fine-grained features and difficulty in effectively solving coarse-grained redundancy in 3D medical images, and is not limited to specific diseases in method design.
Owner:BEIJING INST OF TECH