Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

19 results about "Cross modality" patented technology

Cross-modality translation is the process of converting from the affective, sensory, or evaluative perceptions of pain to a graded number, word, line, or color scale (e. However, it should be noted that many cross-modality equivalence classes have been demonstrated.

Rolling bearing vibration signal multi-mode fault classification method

The invention discloses a rolling bearing vibration signal multi-mode fault classification method. The method comprises the following steps: collecting a vibration time sequence signal of a rolling bearing; the vibration time sequence signals are input into a time sequence branch network and a space branch network in parallel, and the time sequence branch network extracts time sequence dependence characteristics of the signals through a one-dimensional convolutional neural network and a bidirectional gating circulation unit; the spatial branch network converts the vibration time sequence signal into a Markov transform field image, and extracts spatial structure features of the image by using a two-dimensional convolutional neural network and a window Transform-based visual network; performing bidirectional interaction and weighted fusion on the time sequence features and the spatial features through a cross-modal attention mechanism to obtain fusion features; and inputting the fusion features into a classifier, and outputting a fault classification result of the rolling bearing. According to the method, the problems that a traditional single-mode fault diagnosis model is insufficient in adaptability to complex working conditions, multi-mode feature fusion is insufficient, and the generalization ability is weak due to model structure redundancy are solved.
Owner:HARBIN INST OF TECH

Electroencephalogram-myoelectricity collaborative limb movement intention decoding method and system based on symmetric cross-modal attention network

The invention provides an electroencephalogram-myoelectricity collaborative limb movement intention decoding method and system based on a symmetric cross-modal attention network, and belongs to the technical field of neural rehabilitation engineering and movement intention decoding. Comprising the following steps: acquiring an electroencephalogram signal and an electromyographic signal to be decoded; inputting the electroencephalogram signals into an electroencephalogram channel specificity feature extraction network, and extracting multi-scale space-time oscillation features; the electromyographic signals are input into an electromyographic channel specificity feature extraction network, and dynamic time sequence mode features are extracted; inputting the electroencephalogram features and the myoelectricity features into a symmetric cross-modal attention module, carrying out bidirectional feature interaction and calibration, and generating fusion enhancement features; the symmetric cross-modal attention module realizes mutual enhancement and alignment between the electroencephalogram signals and the electromyographic signals by taking own features as queries and taking another modal feature as a key and a value through a cross attention mechanism; and inputting the fusion enhancement feature into a classifier, and decoding to obtain a corresponding motion intention category.
Owner:NINGXIA UNIVERSITY

Multi-source threat detection method based on hybrid expert model

According to the multi-source threat detection method based on the hybrid expert model, real-time collection and structured processing of network flow, system logs and user behavior data are achieved through a multi-mode intelligent collection engine, and high-quality multi-source input is provided for upper-layer analysis; the double-branch feature extractor carries out deep analysis on the network flow time sequence mode and the log semantic context to generate fine-grained feature vectors; the hybrid expert reasoning framework is based on expert models in three fields of a dynamic routing gating network, intelligent scheduling network behaviors, log semantics and user portraits, combines space-time alignment features through a cross-modal attention mechanism, and constructs an interpretable attack evidence chain in combination with a causal reasoning engine. Finally, a full-link closed loop from multi-modal data acquisition, feature collaborative extraction and intelligent threat reasoning is realized, and while millisecond-level real-time response is ensured, the complex internal threat detection accuracy is obviously improved.
Owner:THE QUARTERMASTER RES INST OF THE GENERAL LOGISTICS DEPT OF THE CPLA

Medical image fusion method based on visual state space and gated attention mechanism

PendingCN121502717AImage enhancementImage analysisCross modalityFeature extraction
The invention relates to the technical field of multi-modal medical image fusion, provides a medical image fusion method based on a visual state space and a gated attention mechanism, and effectively captures cross-modal local details and a global dependency relationship through collaborative combination of a convolutional neural network and a state space model. Specifically, a multi-scale feature extraction module is introduced to extract hierarchical features of different scales of each mode, and the module further introduces an optimized channel-space attention module to enhance cross-scale feature expression and inter-channel interaction. In order to model long-distance spatial dependence and enhance cross-modal feature interaction, a multi-scale visual state space module is developed based on SSM, and the module can realize global context aggregation while keeping fine-grained structure details. And finally, integrating the information of the two modes together through a specific fusion strategy, and generating a single image with rich details.
Owner:CHANGCHUN UNIV

System for positioning and answering surgical visual questions

The invention discloses a system for positioning and answering surgical visual questions. The system is characterized in that an expert surgical hybrid model composed of a large language visual model and a small language visual model is constructed; in a surgical hybrid model, respectively extracting image features and text features by using the large language visual model, then performing cross-modal feature extraction and fusion on the image features and the text features to output features, and then converting the output features into text embedding; image features extracted by the large-language visual model and text embedding of the large-language visual model are obtained through the small-language visual model, then cross-modal feature extraction and fusion are conducted on the image features and the text embedding to output fusion features, and finally positioning and answering of surgical visual problems are conducted through the fusion features. According to the method, the reasoning ability, the accuracy and the visual consistency in the operation visual question-answering task can be enhanced, and the excellent question positioning and answering ability is achieved.
Owner:ZHEJIANG ACAD OF TRADITIONAL CHINESE MEDICINE

Automated moodboard augmentation via cross-modal generative association making

A method for automated moodboard augmentation via cross-modal generative association making is described. The method includes specifying, by a user, a region to augment in their digital workspace, including at least one selected image. The method also includes inferring a representative text, label, or description for the at least one selected image. The method further includes creating a basis for concept blending based on the representative text, label, or description inferred for the at least one selected image. The method also includes generating images in response to an adjustable slider, as adjusted by the user, to adjust how much the generated images should resemble directly adjacent images, including the at least one selected image.
Owner:TOYOTA JIDOSHA KK

A road scene target detection method, system, device and medium based on a Transformer and cross-modal

PendingCN122435554APattern recognitionCross modality
A road scene target detection method, system, device and medium based on a Transformer and cross-modal, the method comprising: constructing a road scene target detection network model based on a Transformer and cross-modal; using an RGB and infrared dual-modal road scene target detection dataset for model training, and using the trained model to realize road scene target detection; the system, device and medium are used to realize the method; the application constructs three key modules of cross-modal feature fusion, direction structure enhancement and scale perception attention, systematically improves the feature representation capability of the model from three aspects of modal complementation, structure expression and scale modeling, and the three modules are mutually coordinated, so that the overall network can still maintain stable and accurate perception performance in a complex environment, and a technical solution with high robustness and generalization capability is provided for multi-modal target detection and scene understanding.
Owner:XIAN UNIV OF POSTS & TELECOMM

A target detection method, system, device and medium based on cross-modal fusion and guided attention mechanism

A target detection method, system, device and medium based on cross-modal fusion and guided attention mechanism, the target detection method comprising: acquiring a visible-infrared image paired dataset, processing and dividing the dataset to obtain a training set, a validation set and a test set; constructing a multi-modal target detection network; setting network training parameters; training and optimizing the multi-modal target detection network using the training set, outputting a training weight file after training is completed, and verifying the weight file using the validation set to select the weight file with the highest precision as the optimal weight file; loading the test set and the optimal weight file into the multi-modal target detection network to detect targets in the test set and obtain a target detection result; the system, device and medium are used to carry and implement the method; the application has lower false detection and error detection, and improves the robustness and accuracy of target detection.
Owner:XIDIAN UNIV

An intelligent auxiliary method and system based on multi-modal deep learning

The application discloses an intelligent auxiliary method and system based on multi-modal deep learning, relates to medical image processing and data processing technology, and comprises the following steps: text features are extracted based on illness description text data, and image features of medical image data are extracted; the extracted text features and image features are respectively input into a pre-trained LSTM, so that the LSTM outputs context illness text features and local lesion features; the output context illness text features and local lesion features are weighted and fused by using an attention fusion mechanism; the weighted and fused features are encoded by using an encoder, so that text-related lesion features are output by the encoder; and the text-related lesion features and the local lesion features are decoded by using a decoder, so that a predicted lesion condition is output based on the decoder. Through fusion of text and image features in cross modalities, the application improves the prediction accuracy of the model.
Owner:CHONGQING MEDICAL UNIV SHAOXING KEQIAO MEDICAL LAB TECH RES CENT

Cognitive state recognition method based on multi-modal feature fusion

PendingCN120873813ASensorsDiagnostic recording/measuringCross modalityLabeled data
The invention provides a cognitive state recognition method based on multi-modal feature fusion, which comprises the following steps of: firstly, capturing a time evolution rule of a cross mode based on a multi-head attention mechanism, and adjusting contribution degree of cross-modal features through dynamic weight to extract common features related to a cognitive state; secondly, designing a double-model collaborative verification strategy, screening out high-quality pseudo-label data of a target domain, and performing self-training on the high-quality pseudo-label data to avoid dependence of a domain adaptation method on an auxiliary module; quantitative and qualitative comprehensive experiment comparison shows that the recognition precision of the cognitive state can be remarkably improved, and the designed pseudo-label optimization mechanism can be migrated to related cognitive state recognition tasks on the premise that the complexity of the model is not increased.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method, device, equipment, medium and program product for training neural network model

ActiveCN115115049BNeural learning methodsPattern recognitionCross modality
The application provides a neural network model training method and device, equipment, medium and program product; wherein, the method comprises: obtaining the characteristics of multiple modalities of an image; performing fusion processing based on the image characteristics and title text characteristics to obtain first fusion characteristics of cross modalities; calling a first neural network model based on the characteristics of multiple modalities to perform multiple single-modality prediction tasks to obtain corresponding single-modality prediction results and determine corresponding single-modality loss; calling the first neural network model based on the first fusion characteristics to perform multiple cross-modality prediction tasks to obtain corresponding cross-modality prediction results and determine corresponding cross-modality loss; and performing back propagation based on the single-modality loss and the cross-modality loss to update the parameters of the first neural network model. Through the application, one model can be used to perform multiple different modality tasks, thereby improving modeling efficiency.
Owner:TENCENT TECH WUHAN

An abnormal video detection method, system and terminal based on artificial intelligence

The present invention discloses an abnormal video detection method, system and terminal based on artificial intelligence, which relates to the field of electronic information technology. The detection method includes: converting the original video into a frame sequence to obtain a high-dimensional feature vector of each frame image; extracting the text features of the original video to obtain a text feature vector; using the late fusion method and the cross-modal attention mechanism to perform feature fusion on the high-dimensional feature vector and the text feature vector to obtain a multimodal feature sequence; using a multi-layer Transformer encoder to perform feature processing on the multimodal feature sequence to obtain a processed feature sequence; post-processing the processed feature sequence to obtain an abnormal information value; configuring an abnormal information threshold, when the abnormal information value is greater than the abnormal information threshold, it is judged as an abnormal video; when the abnormal information value is less than or equal to the abnormal information threshold, it is judged as a non-abnormal video. The present invention can make full use of multimodal information, optimize feature fusion results, and improve detection performance.
Owner:BEIJING INST OF TECH

Multi-mode signal processing method for audio emotion recognizer

The invention provides a multi-modal signal processing method for an audio emotion recognizer, and belongs to the technical field of artificial intelligence, natural language processing and large models.The method comprises the steps that audio features and text features are extracted through a long-short-term memory network and a pre-trained BERT model respectively and divided into a modal invariant part and a modal specific part; the core of the method is that a modal binding mechanism is designed: for modal invariant features, the consistency information of the modal invariant features is enhanced by calculating a cross modal weight matrix; for modal specific features, a cross attention mechanism is adopted to capture a dynamic complementary relationship, after the bound features are combined with classification and position embedding, the features are sent to a Transform encoder for deep interactive learning, and finally, a classifier outputs an emotion recognition result. According to the method, optimization is carried out through a composite loss function including task loss, similarity loss and difference loss, accurate distinguishing and deep fusion of multi-modal features are achieved, and high efficiency and robustness are achieved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

A consistent cross-modal hashing retrieval method and system

The application relates to the technical field of data retrieval, and discloses a cross-modal hash retrieval method and system with consistency, which comprises the following steps: S1. acquiring heterogeneous data; S2. acquiring hash codes of the heterogeneous data, and performing hash code learning on the heterogeneous data; S3. performing hash function learning according to the obtained optimal hash code; and S4. mapping the heterogeneous data to the same low-rank Hamming space through the hash function after the function learning is completed, using an exclusive OR operation to the similarity of the heterogeneous data and a retrieval set, returning a result with high similarity, and completing cross-modal retrieval of the heterogeneous data to be retrieved. The application solves the problems of insufficient retrieval precision and complicated optimization process in the prior art, and has the characteristics of being capable of overcoming the heterogeneity of cross modalities.
Owner:GUANGDONG UNIV OF TECH

A medical image multi-modal calculation method and system based on cross-modal generation

This application discloses a medical image multimodal computing method and system based on cross-modal generation, applicable to the field of medical imaging technology. The method includes: acquiring brain image data and preprocessing the brain image data; training a dual-path Siamese neural network based on the preprocessed brain image data; wherein the dual-path Siamese neural network includes a cross-modal generation network and a brain image diagnosis network; inputting the first intermediate layer feature map of the acquired cross-modal generation network and the second intermediate layer feature map of the brain image diagnosis network into a cross-modal feature fusion module for feature fusion to obtain a diagnostic result; acquiring brain image data of a target user, inputting the target user's brain image data into the trained dual-path Siamese neural network and the cross-modal feature fusion module, and outputting an evaluation result.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Incomplete shape symmetry prediction method and system based on multi-modal feature fusion

The application relates to a multi-modal feature fusion-based incomplete shape symmetry prediction method and system, which comprises the following steps: inputting an RGB image and a depth image of an incomplete shape into a channel perception network for fusion; cross-modal features extracted by the channel perception network are fused with color and depth modal features; the cross-modal features are input into a self-reconstruction network of a three-dimensional variational autoencoder for reconstruction to obtain a complete geometric shape, which provides global geometric features for a subsequent symmetry prediction network; and the cross-modal features and the global geometric features are input into the symmetry prediction network for symmetry prediction to obtain predicted symmetry of the incomplete shape. The method can reconstruct the incomplete shape into a complete geometric shape, overcome the defect that the incomplete shape lacks global geometric features, and realize symmetry prediction of the incomplete shape.
Owner:NAT UNIV OF DEFENSE TECH

Network intrusion detection method and system based on multi-modal deep learning

The application belongs to the technical field of Internet security, and discloses a network intrusion detection method and system based on multi-modal deep learning, which comprises collecting and modal attribution of multi-source data from different channels, constructing a time-continuous multi-track dynamic perception track set, integrating a modal attention distribution mechanism, dynamically distributing modal parameters according to the modal energy compound of each type of modal under different perception tracks, and outputting an initial index tensor after multi-track modal fusion; receiving the initial index tensor, establishing a cross-modal dependency relationship by constructing a bidirectional response tensor between the modal and the track; constructing a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship, adopting a cross-modal residual connection mechanism, preserving low-order modal coupling features through a parallel residual path, and outputting a modal linkage node; the recognition ability of the detection on complex intrusion behaviors is improved, and more comprehensive and accurate detection on complex network attack behaviors is realized.
Owner:聊城大学东昌学院 +1

An intelligent interactive intention recognition method and system based on multi-modal data

The application discloses a multi-modal data-based intelligent interaction intention recognition method and system, constructs a multi-smart ring body recognition architecture based on a user, and the multi-smart ring body recognition architecture comprises a plurality of concentric rings with different diameters; multi-modal data of the user is collected, the multi-modal data is preprocessed, and the multi-modal data is respectively mapped to the circumferential surfaces of different concentric rings according to different types of the multi-modal data; static features and dynamic features of the multi-modal data located on the concentric rings are respectively calculated, the static features and the dynamic features are fused to obtain fused features; the fused features are input into an intention classification model to recognize the user intention; the multi-modal data can be mapped to different concentric rings to solve the space-time dislocation problem between cross modalities, and the concentric rings can rotate to dynamically adjust the relative positions of the modal data, so that the static features and the dynamic features can be accurately calculated, and the fused features can be used for accurate recognition of the user intention.
Owner:DEDE JIE (FOSHAN) TECHNOLOGY CO LTD

Low-altitude air route resource quantitative evaluation method and system and storage medium

PendingCN122155088AAccurately model complexityAccurately model dependenciesForecastingBiological modelsCross modalitySimulation
The application relates to the technical field of low-altitude airspace intelligent management, in particular to a low-altitude air route resource quantitative evaluation method and system and a storage medium, which comprises the following steps: inputting standardized multi-modal data into a multi-modal deep fusion evaluation model, extracting deep features of each mode through a mode-specific encoder, realizing information fusion between modes through a cross-modal attention fusion module, and extracting space-time dependent features through a space-time feature extraction network; calculating a plurality of basic evaluation indexes of low-altitude air route resources based on the space-time dependent features, dynamically fusing the plurality of basic evaluation indexes through an evaluation output layer using a weighting mechanism, and generating a comprehensive score of low-altitude air route resources. Through the innovative mode-specific encoder and cross-modal attention fusion module, deep features of each mode are extracted, the complex interaction and dependency relationship between modes are more accurately modeled, and the evaluation is more comprehensive and accurate.
Owner:CHINA AERO GEOPHYSICAL SURVEY & REMOTE SENSING CENT FOR LAND & RESOURCES