Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

29 results about "Cross modality" patented technology

Cross-modality translation is the process of converting from the affective, sensory, or evaluative perceptions of pain to a graded number, word, line, or color scale (e. However, it should be noted that many cross-modality equivalence classes have been demonstrated.

Digital intelligent customer service system based on multiple modes

The invention discloses a multi-modal-based digital intelligent customer service system. Data compensation is carried out on initial multi-modal problem data by adopting a dynamic adaptive compensation mechanism through a multi-modal input module. And carrying out dynamic weight distribution on the target multi-modal problem data by adopting a cross-modal attention mechanism through a multi-modal comparison learning module. And performing intention recognition on the target multi-modal semantic vector through an intention recognition module according to the historical interaction data and a preset dynamic knowledge graph. And inputting the user intention data into the target reinforcement learning reply model through a reply generation module to construct reply data. And performing potential problem prediction by adopting the target multi-modal semantic vector, the historical interaction sequence corresponding to the user and the time data through a potential problem prediction module. And by utilizing the dynamic updating capability of the knowledge graph, the context sensitive identification of the user problem is realized, and the identification accuracy is improved. And the potential problem prediction module turns from passive response to active service, so that the interaction efficiency and the user satisfaction are remarkably improved.
Owner:GUANGZHOU JIAYIN COMMUNICATION CO LTD

Rolling bearing vibration signal multi-mode fault classification method

The invention discloses a rolling bearing vibration signal multi-mode fault classification method. The method comprises the following steps: collecting a vibration time sequence signal of a rolling bearing; the vibration time sequence signals are input into a time sequence branch network and a space branch network in parallel, and the time sequence branch network extracts time sequence dependence characteristics of the signals through a one-dimensional convolutional neural network and a bidirectional gating circulation unit; the spatial branch network converts the vibration time sequence signal into a Markov transform field image, and extracts spatial structure features of the image by using a two-dimensional convolutional neural network and a window Transform-based visual network; performing bidirectional interaction and weighted fusion on the time sequence features and the spatial features through a cross-modal attention mechanism to obtain fusion features; and inputting the fusion features into a classifier, and outputting a fault classification result of the rolling bearing. According to the method, the problems that a traditional single-mode fault diagnosis model is insufficient in adaptability to complex working conditions, multi-mode feature fusion is insufficient, and the generalization ability is weak due to model structure redundancy are solved.
Owner:HARBIN INST OF TECH

VR interactive control management system and method

The invention discloses a VR interactive control management system and method, and relates to the technical field of virtual reality. The method comprises the steps of obtaining multi-source interaction data of a user in a virtual reality environment, performing intra-modal representation conversion and fusion, and constructing a unified multi-modal input tensor; semantic intention representation is extracted based on the cross-modal perception structure; generating a control instruction vector in combination with the historical state information and the current semantic intention; mapping the control instruction vector into an equipment control signal set conforming to various VR terminal interface specifications; after the equipment executes the control instruction, multi-dimensional feedback information is collected, the structure of the multi-dimensional feedback information is reconstructed, feedback representation capable of flowing back to the sensing module is generated, and closed-loop interaction between sensing and control is achieved. Through unifying a multi-source interaction data structure, a dynamically adaptive control instruction vector and a high-precision cross-modal semantic representation mechanism are constructed, and a closed-loop interaction process with consistent sensing and control structures, flexible response and semantic alignment in a virtual reality system is realized.
Owner:HANGZHOU KAILIN CULTURE TECHNOLOGY CO LTD +1

Cross-modal image and text corpus association analysis system

The invention provides a cross-modal image and text corpus association analysis system, and relates to the field of image and text analysis. Comprising a data collection and preprocessing module, a feature extraction module, a cross-modal association learning module and a model training and optimization module. The data collection and preprocessing module collects image and text data from multiple sources and preprocesses the image and text data; the feature extraction module extracts image and text features by using CNN and NLP models; the cross-modal association learning module enhances the semantic consistency of image and text features through feature alignment, weighted summation, an attention mechanism and a cross-modal interaction unit; the model training and optimizing module adopts an unsupervised learning method to train a model and uses an optimization algorithm to adjust parameters; according to the system, through an innovative cross-modal association learning mechanism and an advanced deep learning model, the accuracy and reliability of cross-modal image and text corpus association analysis are effectively improved.
Owner:CHINA SOUTHERN POWER GRID ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Electroencephalogram-myoelectricity collaborative limb movement intention decoding method and system based on symmetric cross-modal attention network

The invention provides an electroencephalogram-myoelectricity collaborative limb movement intention decoding method and system based on a symmetric cross-modal attention network, and belongs to the technical field of neural rehabilitation engineering and movement intention decoding. Comprising the following steps: acquiring an electroencephalogram signal and an electromyographic signal to be decoded; inputting the electroencephalogram signals into an electroencephalogram channel specificity feature extraction network, and extracting multi-scale space-time oscillation features; the electromyographic signals are input into an electromyographic channel specificity feature extraction network, and dynamic time sequence mode features are extracted; inputting the electroencephalogram features and the myoelectricity features into a symmetric cross-modal attention module, carrying out bidirectional feature interaction and calibration, and generating fusion enhancement features; the symmetric cross-modal attention module realizes mutual enhancement and alignment between the electroencephalogram signals and the electromyographic signals by taking own features as queries and taking another modal feature as a key and a value through a cross attention mechanism; and inputting the fusion enhancement feature into a classifier, and decoding to obtain a corresponding motion intention category.
Owner:NINGXIA UNIVERSITY

Network intrusion detection method and system based on multi-modal deep learning

The invention belongs to the technical field of Internet security, and discloses a multi-modal deep learning-based network intrusion detection method and system, which comprises the steps of carrying out collection and modal attribution on multi-source data from different channels, constructing a time-continuous multi-track dynamic sensing track set, integrating a modal attention distribution mechanism, and carrying out multi-modal deep learning. According to the modal energy sensing complexes of each type of modals under different sensing orbits, dynamically distributing modal parameters, and outputting an initial index tensor after multi-orbit modal fusion; receiving an initial index tensor, and establishing a cross-modal dependency relationship by constructing a bidirectional response tensor between a modal and an orbit; constructing a semantic trigger matrix based on semantic tags in a cross-modal dependency relationship, adopting a cross-modal residual connection mechanism, retaining low-order modal coupling features through a parallel residual path, and outputting modal linkage nodes; according to the invention, the identification capability of detection on complex intrusion behaviors is improved, and more comprehensive and accurate detection on complex network attack behaviors is realized.
Owner:聊城大学东昌学院 +1

Multi-source threat detection method based on hybrid expert model

According to the multi-source threat detection method based on the hybrid expert model, real-time collection and structured processing of network flow, system logs and user behavior data are achieved through a multi-mode intelligent collection engine, and high-quality multi-source input is provided for upper-layer analysis; the double-branch feature extractor carries out deep analysis on the network flow time sequence mode and the log semantic context to generate fine-grained feature vectors; the hybrid expert reasoning framework is based on expert models in three fields of a dynamic routing gating network, intelligent scheduling network behaviors, log semantics and user portraits, combines space-time alignment features through a cross-modal attention mechanism, and constructs an interpretable attack evidence chain in combination with a causal reasoning engine. Finally, a full-link closed loop from multi-modal data acquisition, feature collaborative extraction and intelligent threat reasoning is realized, and while millisecond-level real-time response is ensured, the complex internal threat detection accuracy is obviously improved.
Owner:THE QUARTERMASTER RES INST OF THE GENERAL LOGISTICS DEPT OF THE CPLA

Medical image fusion method based on visual state space and gated attention mechanism

PendingCN121502717AImage enhancementImage analysisCross modalityFeature extraction
The invention relates to the technical field of multi-modal medical image fusion, provides a medical image fusion method based on a visual state space and a gated attention mechanism, and effectively captures cross-modal local details and a global dependency relationship through collaborative combination of a convolutional neural network and a state space model. Specifically, a multi-scale feature extraction module is introduced to extract hierarchical features of different scales of each mode, and the module further introduces an optimized channel-space attention module to enhance cross-scale feature expression and inter-channel interaction. In order to model long-distance spatial dependence and enhance cross-modal feature interaction, a multi-scale visual state space module is developed based on SSM, and the module can realize global context aggregation while keeping fine-grained structure details. And finally, integrating the information of the two modes together through a specific fusion strategy, and generating a single image with rich details.
Owner:CHANGCHUN UNIV

System for positioning and answering surgical visual questions

The invention discloses a system for positioning and answering surgical visual questions. The system is characterized in that an expert surgical hybrid model composed of a large language visual model and a small language visual model is constructed; in a surgical hybrid model, respectively extracting image features and text features by using the large language visual model, then performing cross-modal feature extraction and fusion on the image features and the text features to output features, and then converting the output features into text embedding; image features extracted by the large-language visual model and text embedding of the large-language visual model are obtained through the small-language visual model, then cross-modal feature extraction and fusion are conducted on the image features and the text embedding to output fusion features, and finally positioning and answering of surgical visual problems are conducted through the fusion features. According to the method, the reasoning ability, the accuracy and the visual consistency in the operation visual question-answering task can be enhanced, and the excellent question positioning and answering ability is achieved.
Owner:ZHEJIANG ACAD OF TRADITIONAL CHINESE MEDICINE

A Method and System for Generating Cross-Modal Video Adversarial Examples

ActiveCN115496966BCharacter and pattern recognitionCross modalityFeature vector
The present invention discloses a method and system for generating cross-modal video adversarial samples, relating to the technical field of deep learning, including: obtaining clean video samples and converting them into a series of picture frames; extracting features from each picture frame to obtain corresponding feature vectors; determining key frames in the series of picture frames according to the feature vectors; dividing each key frame to obtain a series of image patch pictures and calculating the gradient score of each image patch picture; selecting the image patch picture with the largest gradient score as the local picture; adding perturbations to the local picture to obtain an adversarial frame and calculating the similarity between the local picture and the adversarial frame; updating the perturbations until the similarity reaches the minimum value, and taking the corresponding adversarial frame as the picture adversarial sample; using the picture adversarial sample to replace the corresponding key frame to obtain the video adversarial sample. The method and system for generating video adversarial samples in the present invention have high generation efficiency and strong concealment, and improve the cross-modal transferability of adversarial samples.
Owner:广州市省信软件有限公司

Automated moodboard augmentation via cross-modal generative association making

A method for automated moodboard augmentation via cross-modal generative association making is described. The method includes specifying, by a user, a region to augment in their digital workspace, including at least one selected image. The method also includes inferring a representative text, label, or description for the at least one selected image. The method further includes creating a basis for concept blending based on the representative text, label, or description inferred for the at least one selected image. The method also includes generating images in response to an adjustable slider, as adjusted by the user, to adjust how much the generated images should resemble directly adjacent images, including the at least one selected image.
Owner:TOYOTA JIDOSHA KK

A road scene target detection method, system, device and medium based on a Transformer and cross-modal

PendingCN122435554APattern recognitionCross modality
A road scene target detection method, system, device and medium based on a Transformer and cross-modal, the method comprising: constructing a road scene target detection network model based on a Transformer and cross-modal; using an RGB and infrared dual-modal road scene target detection dataset for model training, and using the trained model to realize road scene target detection; the system, device and medium are used to realize the method; the application constructs three key modules of cross-modal feature fusion, direction structure enhancement and scale perception attention, systematically improves the feature representation capability of the model from three aspects of modal complementation, structure expression and scale modeling, and the three modules are mutually coordinated, so that the overall network can still maintain stable and accurate perception performance in a complex environment, and a technical solution with high robustness and generalization capability is provided for multi-modal target detection and scene understanding.
Owner:XIAN UNIV OF POSTS & TELECOMM

A target detection method, system, device and medium based on cross-modal fusion and guided attention mechanism

A target detection method, system, device and medium based on cross-modal fusion and guided attention mechanism, the target detection method comprising: acquiring a visible-infrared image paired dataset, processing and dividing the dataset to obtain a training set, a validation set and a test set; constructing a multi-modal target detection network; setting network training parameters; training and optimizing the multi-modal target detection network using the training set, outputting a training weight file after training is completed, and verifying the weight file using the validation set to select the weight file with the highest precision as the optimal weight file; loading the test set and the optimal weight file into the multi-modal target detection network to detect targets in the test set and obtain a target detection result; the system, device and medium are used to carry and implement the method; the application has lower false detection and error detection, and improves the robustness and accuracy of target detection.
Owner:XIDIAN UNIV

An intelligent auxiliary method and system based on multi-modal deep learning

The application discloses an intelligent auxiliary method and system based on multi-modal deep learning, relates to medical image processing and data processing technology, and comprises the following steps: text features are extracted based on illness description text data, and image features of medical image data are extracted; the extracted text features and image features are respectively input into a pre-trained LSTM, so that the LSTM outputs context illness text features and local lesion features; the output context illness text features and local lesion features are weighted and fused by using an attention fusion mechanism; the weighted and fused features are encoded by using an encoder, so that text-related lesion features are output by the encoder; and the text-related lesion features and the local lesion features are decoded by using a decoder, so that a predicted lesion condition is output based on the decoder. Through fusion of text and image features in cross modalities, the application improves the prediction accuracy of the model.
Owner:CHONGQING MEDICAL UNIV SHAOXING KEQIAO MEDICAL LAB TECH RES CENT

Cognitive state recognition method based on multi-modal feature fusion

PendingCN120873813ASensorsDiagnostic recording/measuringCross modalityLabeled data
The invention provides a cognitive state recognition method based on multi-modal feature fusion, which comprises the following steps of: firstly, capturing a time evolution rule of a cross mode based on a multi-head attention mechanism, and adjusting contribution degree of cross-modal features through dynamic weight to extract common features related to a cognitive state; secondly, designing a double-model collaborative verification strategy, screening out high-quality pseudo-label data of a target domain, and performing self-training on the high-quality pseudo-label data to avoid dependence of a domain adaptation method on an auxiliary module; quantitative and qualitative comprehensive experiment comparison shows that the recognition precision of the cognitive state can be remarkably improved, and the designed pseudo-label optimization mechanism can be migrated to related cognitive state recognition tasks on the premise that the complexity of the model is not increased.
Owner:NANJING UNIV OF POSTS & TELECOMM

Method, device, equipment, medium and program product for training neural network model

ActiveCN115115049BNeural learning methodsPattern recognitionCross modality
The application provides a neural network model training method and device, equipment, medium and program product; wherein, the method comprises: obtaining the characteristics of multiple modalities of an image; performing fusion processing based on the image characteristics and title text characteristics to obtain first fusion characteristics of cross modalities; calling a first neural network model based on the characteristics of multiple modalities to perform multiple single-modality prediction tasks to obtain corresponding single-modality prediction results and determine corresponding single-modality loss; calling the first neural network model based on the first fusion characteristics to perform multiple cross-modality prediction tasks to obtain corresponding cross-modality prediction results and determine corresponding cross-modality loss; and performing back propagation based on the single-modality loss and the cross-modality loss to update the parameters of the first neural network model. Through the application, one model can be used to perform multiple different modality tasks, thereby improving modeling efficiency.
Owner:TENCENT TECH WUHAN

An abnormal video detection method, system and terminal based on artificial intelligence

The present invention discloses an abnormal video detection method, system and terminal based on artificial intelligence, which relates to the field of electronic information technology. The detection method includes: converting the original video into a frame sequence to obtain a high-dimensional feature vector of each frame image; extracting the text features of the original video to obtain a text feature vector; using the late fusion method and the cross-modal attention mechanism to perform feature fusion on the high-dimensional feature vector and the text feature vector to obtain a multimodal feature sequence; using a multi-layer Transformer encoder to perform feature processing on the multimodal feature sequence to obtain a processed feature sequence; post-processing the processed feature sequence to obtain an abnormal information value; configuring an abnormal information threshold, when the abnormal information value is greater than the abnormal information threshold, it is judged as an abnormal video; when the abnormal information value is less than or equal to the abnormal information threshold, it is judged as a non-abnormal video. The present invention can make full use of multimodal information, optimize feature fusion results, and improve detection performance.
Owner:BEIJING INST OF TECH

A Time-Varying Cross-Mode Identification Method Based on Nonlinear Frequency Modulation Modal Decomposition

ActiveCN115455349BComplex mathematical operationsCross modalityAlgorithm
The present invention belongs to the technical field of civil engineering structural health monitoring data analysis, and relates to a time-varying cross-modal identification method based on non-linear frequency modulation modal decomposition. First, the time-frequency distribution of the collected acceleration response is calculated by using the short-time Fourier transform, and the ridge line of the time-frequency distribution is extracted; secondly, the time modal coefficients at different frequencies are calculated by using the time-frequency distributions at the positions of each sensor; then, the time modal correlation coefficients at the time-frequency points of the ridge line are calculated, and the initial center frequency of the cross-modal is determined by combining the ridge line frequencies; finally, the instantaneous frequency of the cross-modal is extracted by the non-linear frequency modulation modal decomposition method. The present invention uses the time modal correlation coefficient to determine the initial center frequency, and can accurately identify the modal parameters of each order of the time-varying structure even in the case of the existence of cross-modal in the time-varying structure.
Owner:HEBEI UNIV OF TECH

Multi-mode signal processing method for audio emotion recognizer

The invention provides a multi-modal signal processing method for an audio emotion recognizer, and belongs to the technical field of artificial intelligence, natural language processing and large models.The method comprises the steps that audio features and text features are extracted through a long-short-term memory network and a pre-trained BERT model respectively and divided into a modal invariant part and a modal specific part; the core of the method is that a modal binding mechanism is designed: for modal invariant features, the consistency information of the modal invariant features is enhanced by calculating a cross modal weight matrix; for modal specific features, a cross attention mechanism is adopted to capture a dynamic complementary relationship, after the bound features are combined with classification and position embedding, the features are sent to a Transform encoder for deep interactive learning, and finally, a classifier outputs an emotion recognition result. According to the method, optimization is carried out through a composite loss function including task loss, similarity loss and difference loss, accurate distinguishing and deep fusion of multi-modal features are achieved, and high efficiency and robustness are achieved.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Method and system for distributed learning and adaptation in autonomous driving vehicles

The present teaching relates to system, method, medium for in-situ perception in an autonomous driving vehicle. A plurality of types of sensor data acquired continuously by a plurality of types of sensors deployed on the vehicle are first received, where the plurality of types of sensor data provide information about surrounding of the vehicle. Based on at least one model, one or more items are tracked from a first of the plurality of types of sensor data acquired by one or more of a first type of the plurality of types of sensors, wherein the one or more items appear in the surrounding of the vehicle. At least some of the one or more items are then automatically labeled on-the-fly via either cross modality validation or cross temporal validation of the one or more items and are used to locally adapt, on-the-fly, the at least one model in the vehicle.
Owner:PLUSAI INC

A consistent cross-modal hashing retrieval method and system

The application relates to the technical field of data retrieval, and discloses a cross-modal hash retrieval method and system with consistency, which comprises the following steps: S1. acquiring heterogeneous data; S2. acquiring hash codes of the heterogeneous data, and performing hash code learning on the heterogeneous data; S3. performing hash function learning according to the obtained optimal hash code; and S4. mapping the heterogeneous data to the same low-rank Hamming space through the hash function after the function learning is completed, using an exclusive OR operation to the similarity of the heterogeneous data and a retrieval set, returning a result with high similarity, and completing cross-modal retrieval of the heterogeneous data to be retrieved. The application solves the problems of insufficient retrieval precision and complicated optimization process in the prior art, and has the characteristics of being capable of overcoming the heterogeneity of cross modalities.
Owner:GUANGDONG UNIV OF TECH

HUD-based driver behavior-oriented prediction system and abnormity early warning system

The invention discloses an HUD-based driver-oriented behavior prediction system and an abnormal early warning system, and belongs to the technical field of intelligent driving safety, and the system comprises a data collection and preprocessing module, a feature extraction module, a cross-modal attention mechanism module and a behavior prediction module. The data acquisition and preprocessing module is used for acquiring multi-modal data of a driver in real time and preprocessing the multi-modal data to obtain preprocessed multi-modal data; the feature extraction module performs feature extraction based on the preprocessed multi-modal data; the cross-modal attention mechanism module is used for calculating an attention weight between modals and carrying out dynamic weighted fusion on features of each modal data in each time window to obtain cross-modal features; the behavior prediction module predicts a behavior of the driver based on the cross-modal features. According to the system, multi-modal information fusion and accurate time sequence modeling are realized, and the behavior prediction accuracy is improved.
Owner:NANJING BOTUO VISION TECH CO LTD

VR interactive control management system and method

The present invention discloses a VR interactive control management system and method, which relates to the field of virtual reality technology. The method includes: obtaining multi-source interaction data of users in a virtual reality environment, and performing intra-modal representation conversion and fusion to construct a unified multi-modal input tensor; extracting semantic intent representation based on a cross-modal perception structure; generating a control instruction vector by combining historical state information with current semantic intent; mapping the control instruction vector into a set of device control signals that conform to multiple VR terminal interface specifications; after the device executes the control instruction, collecting multi-dimensional feedback information and reconstructing its structure to generate a feedback representation that can flow back to the perception module, thereby realizing a closed-loop interaction between perception and control. By unifying the multi-source interaction data structure, constructing a dynamically adaptive control instruction vector, and a high-precision cross-modal semantic representation mechanism, a closed-loop interaction process with consistent structure, flexible response, and semantic alignment of perception and control in a virtual reality system is realized.
Owner:HANGZHOU KAILIN CULTURE TECHNOLOGY CO LTD +1

A medical image multi-modal calculation method and system based on cross-modal generation

This application discloses a medical image multimodal computing method and system based on cross-modal generation, applicable to the field of medical imaging technology. The method includes: acquiring brain image data and preprocessing the brain image data; training a dual-path Siamese neural network based on the preprocessed brain image data; wherein the dual-path Siamese neural network includes a cross-modal generation network and a brain image diagnosis network; inputting the first intermediate layer feature map of the acquired cross-modal generation network and the second intermediate layer feature map of the brain image diagnosis network into a cross-modal feature fusion module for feature fusion to obtain a diagnostic result; acquiring brain image data of a target user, inputting the target user's brain image data into the trained dual-path Siamese neural network and the cross-modal feature fusion module, and outputting an evaluation result.
Owner:SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Incomplete shape symmetry prediction method and system based on multi-modal feature fusion

The application relates to a multi-modal feature fusion-based incomplete shape symmetry prediction method and system, which comprises the following steps: inputting an RGB image and a depth image of an incomplete shape into a channel perception network for fusion; cross-modal features extracted by the channel perception network are fused with color and depth modal features; the cross-modal features are input into a self-reconstruction network of a three-dimensional variational autoencoder for reconstruction to obtain a complete geometric shape, which provides global geometric features for a subsequent symmetry prediction network; and the cross-modal features and the global geometric features are input into the symmetry prediction network for symmetry prediction to obtain predicted symmetry of the incomplete shape. The method can reconstruct the incomplete shape into a complete geometric shape, overcome the defect that the incomplete shape lacks global geometric features, and realize symmetry prediction of the incomplete shape.
Owner:NAT UNIV OF DEFENSE TECH

Network intrusion detection method and system based on multi-modal deep learning

The application belongs to the technical field of Internet security, and discloses a network intrusion detection method and system based on multi-modal deep learning, which comprises collecting and modal attribution of multi-source data from different channels, constructing a time-continuous multi-track dynamic perception track set, integrating a modal attention distribution mechanism, dynamically distributing modal parameters according to the modal energy compound of each type of modal under different perception tracks, and outputting an initial index tensor after multi-track modal fusion; receiving the initial index tensor, establishing a cross-modal dependency relationship by constructing a bidirectional response tensor between the modal and the track; constructing a semantic trigger matrix based on the semantic labels in the cross-modal dependency relationship, adopting a cross-modal residual connection mechanism, preserving low-order modal coupling features through a parallel residual path, and outputting a modal linkage node; the recognition ability of the detection on complex intrusion behaviors is improved, and more comprehensive and accurate detection on complex network attack behaviors is realized.
Owner:聊城大学东昌学院 +1

An intelligent interactive intention recognition method and system based on multi-modal data

The application discloses a multi-modal data-based intelligent interaction intention recognition method and system, constructs a multi-smart ring body recognition architecture based on a user, and the multi-smart ring body recognition architecture comprises a plurality of concentric rings with different diameters; multi-modal data of the user is collected, the multi-modal data is preprocessed, and the multi-modal data is respectively mapped to the circumferential surfaces of different concentric rings according to different types of the multi-modal data; static features and dynamic features of the multi-modal data located on the concentric rings are respectively calculated, the static features and the dynamic features are fused to obtain fused features; the fused features are input into an intention classification model to recognize the user intention; the multi-modal data can be mapped to different concentric rings to solve the space-time dislocation problem between cross modalities, and the concentric rings can rotate to dynamically adjust the relative positions of the modal data, so that the static features and the dynamic features can be accurately calculated, and the fused features can be used for accurate recognition of the user intention.
Owner:DEDE JIE (FOSHAN) TECHNOLOGY CO LTD

Low-altitude air route resource quantitative evaluation method and system and storage medium

PendingCN122155088AAccurately model complexityAccurately model dependenciesForecastingBiological modelsCross modalitySimulation
The application relates to the technical field of low-altitude airspace intelligent management, in particular to a low-altitude air route resource quantitative evaluation method and system and a storage medium, which comprises the following steps: inputting standardized multi-modal data into a multi-modal deep fusion evaluation model, extracting deep features of each mode through a mode-specific encoder, realizing information fusion between modes through a cross-modal attention fusion module, and extracting space-time dependent features through a space-time feature extraction network; calculating a plurality of basic evaluation indexes of low-altitude air route resources based on the space-time dependent features, dynamically fusing the plurality of basic evaluation indexes through an evaluation output layer using a weighting mechanism, and generating a comprehensive score of low-altitude air route resources. Through the innovative mode-specific encoder and cross-modal attention fusion module, deep features of each mode are extracted, the complex interaction and dependency relationship between modes are more accurately modeled, and the evaluation is more comprehensive and accurate.
Owner:CHINA AERO GEOPHYSICAL SURVEY & REMOTE SENSING CENT FOR LAND & RESOURCES

A semantically rich dialogue generation method integrating visual context

This paper discloses a method for generating semantically rich dialogues that integrates visual context. Challenging audiovisual scene perception datasets are collected for model training. The overall model, based on the Transformer, designs and implements a multi-step cross-modal attention mechanism, capturing heterogeneous semantic associations between different modalities in spatiotemporal dimensions in a fine-grained manner. Multimodal feature representations are then jointly constructed into a spatiotemporal graph structure, which is then used for cross-modal learning and reasoning using a graph convolutional network. Finally, the method decodes and generates rich and accurate dialogue responses that are consistent with the current context. By integrating multimodal data and conducting cross-modal interactions, the present invention captures multi-angle, fine-grained, progressive feature interactions and semantic associations between modalities, achieving visual-language cross-modal semantic alignment, improving the model's semantic understanding and reasoning capabilities, and ultimately generating informative and high-quality responses.
Owner:NORTHWESTERN POLYTECHNICAL UNIV