Multimodal neural learning network model, method, device and medium

Through dual-branch image path encoding and soft label contrast loss, the problem of insufficient alignment accuracy in the EEG-image multimodal neural learning network model is solved, and adaptive adjustment of the visual modality and enhanced robustness in high-noise environments are achieved.

CN120429833BActive Publication Date: 2025-09-12SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510926274.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-12
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

The existing EEG-image multimodal neural learning network model has deficiencies in image modality alignment accuracy and robustness, lacks a dynamic adjustment mechanism, and does not introduce prompt optimization strategies and soft label loss function design, resulting in low model processing accuracy.

Method used

A dual-branch image path encoding is adopted, and the image modal structure is adaptively adjusted through the cross-attention mechanism and dynamic filter generation network. Combined with the soft label contrast loss, visual adaptive cue learning is introduced to enhance the visual modality alignment capability.

Benefits of technology

The alignment accuracy and robustness of EEG signals and images are improved, and the semantic alignment and generalization capabilities of the model in high-noise environments are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429833B_ABST
    Figure CN120429833B_ABST
Patent Text Reader

Abstract

The present application proposes a multimodal neural learning network model, method, device and medium. After encoding the EEG signal and the original image, the EEG features and image features are obtained, and then the similarity between the two is calculated to complete the comparative learning of the EEG signal and the original image signal. The encoding of the original image includes obtaining the first primary feature of the embedding representation level based on the original image, obtaining the second primary feature of the embedding representation level based on the filtered image, fusion of the primary features and reasoning. In this way, the embedding representation level primary features of the original image and the filtered image are respectively extracted through a dual-branch image path, and the embedding representation level fusion is performed using an attention gating mechanism, thereby realizing adaptive adjustment of the image modal structure, overcoming the problem of the single path of traditional image encoding, and effectively enhancing the alignment capability of the visual modality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a multimodal neural learning network model, method, device and medium. Background Art

[0002] In recent years, with the deepening of brain-computer interface (BCI) research, decoding and reconstructing human visual perception content based on electroencephalogram (EEG) signals has gradually become an important direction in the field of neural decoding. Many works have attempted to align EEG signals with pre-trained visual models to achieve tasks such as image classification, retrieval or reconstruction, such as the Contrastive Language-Image Pre-Training (CLIP) model. Among them, the first representative method proposed an EEG encoder that integrates channel attention and spatiotemporal convolution, and combined it with CLIP embedding for alignment. The second representative method further introduced a semantic decoupling mechanism and geometric consistency loss to improve cross-modal semantic consistency. However, these methods still have the following technical defects:

[0003] First, existing methods often rely on a "single-path + static alignment" approach, comparing the representations extracted by the EEG encoder with pre-trained visual features. This approach fails to dynamically adjust the representation of the image modality itself based on the actual modality alignment. For example, despite constructing a joint semantic space and introducing a decoupling module, the image encoding path itself remains unadapted, and image features serve only as fixed reference vectors. This limits alignment accuracy when EEG signals exhibit significant ambiguity.

[0004] Second, existing methods generally ignore the plasticity of the internal structure of image patch embedding representations (patch tokens) during visual encoding. Neither the direct mapping approach (the first representative approach) nor the semantically relevant subspace approach (the second representative approach) addresses how to dynamically adjust visual patch token features to better serve the semantic alignment of EEG modalities. The lack of a more fine-grained token-level tuning mechanism limits the ability of the image modality to combat interference from physiological modalities.

[0005] Thirdly, prompt tuning strategies have not yet been incorporated into existing EEG image alignment frameworks. Although previous work has demonstrated the effectiveness of prompts in image-language alignment, no method has attempted to introduce prompt embedding representations to guide the fusion of visual features and EEG signals in cross-modal EEG-image scenarios. This also makes it difficult for models to obtain interpretable prompt signals in zero-shot scenarios lacking class label supervision.

[0006] Finally, in terms of loss function design, traditional methods often use a hard contrast loss function, which only considers the label relationship between positive and negative samples and ignores the fact that the EEG modality itself may have ambiguous semantic responses across multiple images. This overly hard contrast strategy can lead to misjudgment of implicit connections between negative samples, thereby distancing images that should have a certain degree of semantic connection and affecting the overall feature space structure.

[0007] In summary, the existing EEG-image oriented multimodal neural learning network model has low model processing accuracy due to the above technical defects. Summary of the Invention

[0008] This application proposes a multimodal neural learning network model, method, device and medium, which can solve one of the problems existing in the background technology.

[0009] To achieve the above objectives, this application adopts the following technical solutions:

[0010] In a first aspect, a multimodal neural learning network model is provided, the model comprising:

[0011] An EEG signal encoding module is used to encode EEG signals and obtain EEG features;

[0012] An image encoding module, used to encode the original image to obtain image features; and

[0013] A similarity calculation module is used to calculate the similarity between the EEG features and the image features to obtain the task results.

[0014] Wherein, the image encoding module includes:

[0015] A first image encoding submodule is used to extract features from the original image to obtain a first primary feature at an embedding representation level;

[0016] a second image encoding submodule, configured to filter the original image to obtain a filtered image, and perform feature extraction on the filtered image to obtain a second primary feature at an embedding representation level;

[0017] a fusion module, configured to fuse the first primary feature with the second primary feature based on a cross-attention mechanism at the embedding representation level to obtain a secondary feature; and

[0018] An inference module is used to obtain the image feature based on the inference result of the secondary feature.

[0019] Based on the above technical solution, after encoding the EEG signal and the original image, the EEG features and image features are obtained, and then the similarity between the two is calculated to complete the comparative learning of the EEG signal and the original image signal. The encoding of the original image includes obtaining the first primary feature of the embedding representation level based on the original image, obtaining the second primary feature of the embedding representation level based on the filtered image, fusion of the primary features and reasoning. In this way, the embedding representation level primary features of the original image and the filtered image are respectively extracted through the dual-branch image path, and the embedding representation level fusion is performed using the attention gating mechanism, thereby realizing adaptive adjustment of the image modal structure, overcoming the problem of the single path of traditional image encoding, and effectively enhancing the alignment capability of the visual modality.

[0020] In a possible design of the first aspect, the EEG signal encoding module includes:

[0021] A time series perturbation layer is used to assign a learnable weight to each channel and each time point of the preprocessed EEG signal to obtain an intermediate representation with enhanced time series amplitude, and to superimpose a learnable bias matrix on the intermediate representation to obtain an enhanced representation; and

[0022] The EEG signal encoding layer is used to flatten the enhanced representation to obtain flattened EEG data, linearly project the flattened EEG data to obtain an initial feature representation, and perform layer normalization on the initial feature representation to obtain the EEG feature.

[0023] Based on the above technical solution, by applying learnable weights and biases to the EEG channel-time series data, dynamic re-addition of noise interference is achieved, which enhances the model's perception of key spatiotemporal segments of EEG signals and improves the model's robustness in extracting useful semantic representations in high-noise environments, while maintaining extremely low model complexity.

[0024] In a possible design manner of the first aspect, the second image encoding submodule includes:

[0025] A filter parameter generation network, configured to convert the original image into a dynamic filter parameter vector, wherein the dynamic filter parameter vector is defined by a batch size, a filter kernel size, and a number of channels;

[0026] a filter application network, configured to perform dynamic convolution on the original image using the dynamic filter parameter vector to obtain the filtered image; and

[0027] The image block embedding layer is used to extract features from the filtered image to obtain the second primary features.

[0028] In a possible design manner of the first aspect, the filter parameter generation network includes: 3 layers of lightweight convolution layers, a batch normalization layer, an activation layer, an adaptive average pooling layer and 2 layers of fully connected layers arranged in sequence.

[0029] Based on the above technical solution, the parameters of the dynamic filter are generated based on the image content, which further realizes the adaptive adjustment of the image modal structure. The image features can be locally adjusted according to the semantic degree of the EEG signal, overcoming the problem of static invariance of the traditional image encoding path and effectively enhancing the alignment capability of the visual modality.

[0030] In a possible design manner of the first aspect, the fusion module includes:

[0031] a cross-attention layer, configured to perform cross-attention calculation on the first primary feature and the second primary feature to obtain a cross-modal attention weight matrix; and

[0032] The feature vector layer is used to use the attention weight matrix to perform weighted fusion on the first primary feature and the second primary feature to obtain the secondary feature.

[0033] In a possible design manner of the first aspect, the image encoding module further includes:

[0034] A visual adaptive cue learning module is used to add the cue embedding representation learned from the EEG signal to the secondary feature.

[0035] Based on the above technical solution, prompt learning is introduced in the EEG-image alignment scenario to simulate the human neural attention mechanism, enabling the model to automatically generate interpretable prompts based on EEG signals without relying on explicit text guidance, thereby improving the flexibility and adaptability of modal fusion.

[0036] In a possible design manner of the first aspect, the model adopts a soft label contrast loss, where the soft label contrast loss includes a hard label contrast loss term and a contrast loss term based on softening of intra-modal distribution similarity.

[0037] Based on the above technical solution, soft label contrast loss is adopted to alleviate the problem of overly extreme division of positive and negative samples in hard contrast learning. It is especially suitable for fuzzy response scenarios in EEG modalities, enabling the model to more realistically simulate the continuous mapping between physiological signals and semantic space, thereby improving the semantic alignment accuracy and generalization ability.

[0038] In a second aspect, a data processing method based on a multimodal neural learning network model is provided, the processing method comprising:

[0039] Encode the EEG signal to obtain EEG features;

[0040] Encode the original image to obtain image features; and

[0041] Calculate the similarity between the EEG feature and the image feature to obtain the task result,

[0042] The encoding of the original image includes:

[0043] Performing feature extraction on the original image to obtain a first primary feature of an embedding representation level;

[0044] Filtering the original image to obtain a filtered image, and performing feature extraction on the filtered image to obtain a second primary feature of an embedding representation level;

[0045] fusing the first primary feature with the second primary feature based on a cross-attention mechanism at the embedding representation level to obtain a secondary feature; and

[0046] The image features are obtained based on the inference results of the secondary features.

[0047] In a third aspect, an electronic device is provided, comprising: a processor, and a memory coupled to the processor, the memory being used to store a computer program; and the processor being used to execute the computer program stored in the memory, so that the electronic device performs the processing method described in the second aspect.

[0048] In a fourth aspect, a computer-readable storage medium is provided, comprising a computer program or instructions, which, when executed on a computer, causes the computer to execute the processing method described in the second aspect.

[0049] In a fifth aspect, a computer program product is provided, comprising: a computer program or instructions, which, when the computer program or instructions are run on a computer, causes the computer to execute the processing method described in the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0051] Figure 1 This is an architecture diagram of the multimodal neural learning network model provided in Example 1 of the present application;

[0052] Figure 2This is an architectural diagram of a multimodal neural cue learning framework for EEG-image alignment provided in Example 2 of the present application;

[0053] Figure 3 This is a statistical comparison of the Top-1 and Top-5 accuracy rates of the within-subject predictions provided in Example 2 of this application;

[0054] Figure 4 This is a statistical comparison of the Top-1 and Top-5 accuracy rates across subjects provided in Example 2 of this application;

[0055] Figure 5 This is a schematic diagram of similarity prediction results for five sample pairs provided in Example 2 of the present application;

[0056] Figure 6 This is a schematic diagram of similarity prediction results for four sample pairs provided in Example 2 of the present application;

[0057] Figure 7 This is a schematic diagram of similarity prediction results for the three sample pairs provided in Example 2 of the present application;

[0058] Figure 8 This is a schematic diagram of the similarity prediction results of the two sample pairs of subjects provided in Example 2 of the present application;

[0059] Figure 9 This is a schematic diagram of the similarity prediction results of the sample pair of subject 1 provided in Example 2 of the present application;

[0060] Figure 10 This is a schematic diagram of similarity prediction results for 10 sample pairs of subjects provided in Example 2 of the present application;

[0061] Figure 11 This is a schematic diagram of similarity prediction results for 9 sample pairs of subjects provided in Example 2 of the present application;

[0062] Figure 12 This is a schematic diagram of similarity prediction results for eight sample pairs provided in Example 2 of the present application;

[0063] Figure 13 This is a schematic diagram of similarity prediction results for 7 sample pairs provided in Example 2 of the present application;

[0064] Figure 14 This is a schematic diagram of the similarity prediction results of 6 sample pairs of subjects provided in Example 2 of the present application. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0066] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0068] Example 1

[0069] like Figure 1 As shown, this embodiment provides a multimodal neural learning network model 100, and the model 100 includes:

[0070] The EEG signal encoding module 101 is used to encode the EEG signal to obtain EEG features;

[0071] An image encoding module 102 is used to encode the original image to obtain image features; and

[0072] The similarity calculation module 103 is used to calculate the similarity between the EEG feature and the image feature to obtain the task result.

[0073] The image encoding module 102 includes:

[0074] The first image encoding submodule 201 is used to extract features from the original image to obtain first primary features at the embedding representation level;

[0075] The second image encoding submodule 202 is configured to filter the original image to obtain a filtered image, and perform feature extraction on the filtered image to obtain a second primary feature at an embedding representation level;

[0076] A fusion module 203 is configured to fuse the first primary feature with the second primary feature based on a cross-attention mechanism at the embedding representation level to obtain a secondary feature; and

[0077] The inference module 204 is configured to obtain the image features based on the inference results of the secondary features.

[0078] Specifically, multimodality means that the input of this model is data of at least two modalities. In this embodiment, it mainly involves EEG signals and images.

[0079] EEG, also known as electroencephalography (EEG), is a noninvasive diagnostic technique that uses electrodes to record the electrical activity of neurons in the cerebral cortex. It is primarily used to assess brain function and diagnose epilepsy, sleep disorders, and brain diseases. It captures spontaneous electrical signals from the brain and generates analyzable waveforms. It offers high temporal resolution, is radiation-free, and is easy to use.

[0080] An image is a vivid, lifelike description or portrait of an objective object. It is the most commonly used information carrier in human social activities. Alternatively, an image is a representation of an objective object, containing relevant information about the depicted object. It is people's primary source of information.

[0081] It is understandable that before encoding, the original data usually needs to be preprocessed accordingly to make the preprocessed data suitable for subsequent encoding, such as filtering, completion, time alignment, standardization and other processing.

[0082] By calculating the similarity between the encoded EEG features and image features, comparative learning between EEG and images can be achieved.

[0083] The embedding representation is the smallest working unit of the model during operation and can be represented by the corresponding token or patch.

[0084] Cross-Attention is a computational mechanism that enhances features through cross-sequence or cross-modal information interaction. Its core principle is to focus on and integrate key information by dynamically calculating the correlation weights between different input sources. It is widely used in natural language processing, multimodal learning, and computer vision.

[0085] Inference is the process of using extracted features to make predictions or decisions. In small amounts of data learning, inference models can make accurate predictions with the support of a small number of features.

[0086] Based on the above technical solution, after encoding the EEG signal and the original image, the EEG features and image features are obtained, and then the similarity between the two is calculated to complete the comparative learning of the EEG signal and the original image signal. The encoding of the original image includes obtaining the first primary feature of the embedding representation level based on the original image, obtaining the second primary feature of the embedding representation level based on the filtered image, fusion of the primary features and reasoning. In this way, the embedding representation level primary features of the original image and the filtered image are respectively extracted through the dual-branch image path, and the embedding representation level fusion is performed using the attention gating mechanism, thereby realizing adaptive adjustment of the image modal structure, overcoming the problem of the single path of traditional image encoding, and effectively enhancing the alignment capability of the visual modality.

[0087] In one possible design, the EEG signal encoding module includes:

[0088] A time series perturbation layer is used to assign a learnable weight to each channel and each time point of the preprocessed EEG signal to obtain an intermediate representation with enhanced time series amplitude, and to superimpose a learnable bias matrix on the intermediate representation to obtain an enhanced representation; and

[0089] The EEG signal encoding layer is used to flatten the enhanced representation to obtain flattened EEG data, linearly project the flattened EEG data to obtain an initial feature representation, and perform layer normalization on the initial feature representation to obtain the EEG feature.

[0090] Specifically, the EEG signal tensor can be expressed as [batch_size, channels, timesteps], where batch_size represents the number of samples selected for one training, channels represents the number of channels, and timesteps represents the time step.

[0091] The temporal perturbation layer may include: a learnable weight parameter layer and a learnable bias parameter layer.

[0092] The learnable weight parameter layer can multiply the EEG signal by an independent learnable scalar weight at each channel and each time point to obtain an intermediate representation with enhanced temporal amplitude.

[0093] The learnable bias parameter layer can add a learnable bias matrix of the same size [channels, timesteps] to the above intermediate representation, thereby further performing fine-grained perturbation adjustment on the EEG signal to enhance the model's adaptability to background noise or weak EEG fluctuations.

[0094] The EEG signal encoding layer may include: a flattening layer, a fully connected layer, and a layer normalization layer.

[0095] The flattening layer can receive the enhanced representation after perturbation enhancement, which is also expressed as [batch_size, channels, timesteps]. The enhanced representation is flattened to obtain flattened EEG data. The dimension of the flattened EEG data is [channels×timesteps].

[0096] The fully connected layer can perform linear projection on the flattened EEG data, mapping it to the specified feature dimension embedding space to obtain the initial feature representation.

[0097] The layer normalization layer can perform layer normalization on the initial feature representation to obtain EEG features, thereby improving the stability of model training and the consistency of feature distribution, and enhancing the expression ability of cross-modal alignment.

[0098] Based on the above technical solution, by applying learnable weights and biases to the EEG channel-time series data, dynamic re-addition of noise interference is achieved, which enhances the model's perception of key spatiotemporal segments of EEG signals and improves the model's robustness in extracting useful semantic representations in high-noise environments, while maintaining extremely low model complexity.

[0099] In one possible design, the second image encoding submodule includes:

[0100] A filter parameter generation network, configured to convert the original image into a dynamic filter parameter vector, wherein the dynamic filter parameter vector is defined by a batch size, a filter kernel size, and a number of channels;

[0101] a filter application network, configured to perform dynamic convolution on the original image using the dynamic filter parameter vector to obtain the filtered image; and

[0102] The image block embedding layer is used to extract features from the filtered image to obtain the second primary features.

[0103] In one possible design, the filter parameter generation network includes: 3 lightweight convolution layers, a batch normalization layer, an activation layer, 1 adaptive average pooling layer and 2 fully connected layers arranged in sequence.

[0104] Specifically, the setting of lightweight convolutional layers can significantly reduce the amount of computation and model parameters while maintaining accuracy.

[0105] Batch normalization is a deep learning optimization technique designed to address the problem of internal covariate shift in deep network training. It accelerates training and improves model generalization by stabilizing the input distribution of each layer. Its core principles include normalization and restoring expressive power through linear transformation.

[0106] The activation layer can use the ReLU activation function.

[0107] After the original image passes through three layers of lightweight convolutional layers, batch normalization layers, and activation layers in sequence, a downsampled intermediate feature map is obtained. The intermediate feature map is processed by an adaptive average pooling layer to obtain batch features. The batch features are mapped to a dynamic filter parameter vector through two layers of fully connected layers. The dynamic filter parameter vector can be expressed as [B, H×W×C], where B represents the batch size, H×W represents the filter kernel size, and C represents the number of channels.

[0108] The filter application network receives a dynamic filter parameter vector and applies it to the original image, implementing dynamic convolution.

[0109] First, the filter application network receives as input the original image and a dynamic filter parameter vector. The original image is a tensor of shape (B, C, H, W), and the dynamic filter parameter vector is the dynamic filter weights of shape (B, C × K_h × K_w), where K_h is the height of the dynamic filter kernel and K_w is the width of the dynamic filter kernel. The filter tensor is then reshaped to (B, C × K_h × K_w, 1, 1) to accommodate the subsequent point-wise weighting operations. Subsequently, to extract local region information, the image is converted into a locally expanded representation at each location, resulting in a local perception tensor of shape (B, K_h × K_w, H_out, W_out), where H_out is the height of each locally expanded block of the image and W_out is the height of each locally expanded block of the image. Then, for each channel i, the corresponding single-channel image (B, 1, H, W) is extracted and locally expanded. Simultaneously, the corresponding channel weight parameters (B, K_h × K_w, 1, 1) are extracted from the dynamic filter. The two results are element-wise multiplied and summed over the channel dimension to obtain the filtered output of the channel (B, 1, H_out, W_out). The above process is repeated for all channels, and finally the outputs of all channels are spliced ​​into the complete filtering result (B, C, H_out, W_out), realizing the dynamic perception filtering process based on image content and obtaining the final image coding features after dynamic filtering.

[0110] Based on the above technical solution, the parameters of the dynamic filter are generated based on the image content, which further realizes the adaptive adjustment of the image modal structure. Under the supervision of the cross-modal alignment loss function with the EEG signal, the filter parameter generation network can dynamically adjust the filter parameters of the image and apply them to the original image to obtain more efficient image coding features. The obtained image coding features are aligned with the EEG signal again to obtain a new round of supervision loss, and dynamic adjustment is performed based on the feedback of the supervision loss to perform a positive optimization cycle, thereby overcoming the problem of static invariance of the traditional image coding path and effectively enhancing the alignment capability of the visual modality.

[0111] In one possible design, the fusion module includes:

[0112] a cross-attention layer, configured to perform cross-attention calculation on the block-level embedding features of the first primary feature and the second primary feature to obtain a cross-modal attention weight matrix; and

[0113] The feature vector layer is used to use the attention weight matrix to perform weighted fusion on the block-level embedding features of the first primary feature and the second primary feature to obtain the secondary feature.

[0114] In one possible design, the image encoding module further includes:

[0115] A visually adaptive cue learning module is used to add cue embedding representations to the secondary features.

[0116] Specifically, a new hint embedding representation is introduced and directly concatenated before the secondary features. The concatenated features are then fed into the inference module. This hint embedding representation is randomly initialized at the beginning of model training and is subsequently iteratively updated based on the cross-modal alignment loss function. The hint embedding representation is essentially a trainable parameter that is updated as the neural network trains, based on feedback from the loss function.

[0117] Specifically, the hint embedding representation is used to guide the model to better learn the global features of the image.

[0118] Based on the above technical solution, prompt learning is introduced in the EEG-image alignment scenario to simulate the human neural attention mechanism, enabling the model to automatically generate interpretable prompts based on EEG signals without relying on explicit text guidance, thereby improving the flexibility and adaptability of modal fusion.

[0119] In a possible design manner of the first aspect, the model adopts a soft label contrast loss, where the soft label contrast loss includes a hard label contrast loss term and a contrast loss term based on softening of intra-modal distribution similarity.

[0120] Specifically, for hard label contrast loss, define B as the number of samples in a batch during training (batch_size), , The i-th EEG signal sample and the j-th image sample obtained by the EEG signal encoder and image encoder in a batch are defined as their similarity Calculated as:

[0121]

[0122] The hard label contrast loss term is then defined as:

[0123]

[0124] Now define , are all the output feature matrices of a batch obtained by the EEG signal encoding module and the image encoding module, and are their transposed matrices respectively, are hyperparameters, are fixed values, and the soft labels Defined as: .

[0125] definition , , then the contrast loss term based on the softening of intra-modal distribution similarity is defined as:

[0126]

[0127] in, represents the KL divergence. The final overall soft label contrast loss is defined as:

[0128]

[0129] Based on the above technical solution, soft label contrast loss is adopted to alleviate the problem of overly extreme division of positive and negative samples in hard contrast learning. It is especially suitable for fuzzy response scenarios in EEG modalities, enabling the model to more realistically simulate the continuous mapping between physiological signals and semantic space, thereby improving the semantic alignment accuracy and generalization ability.

[0130] This embodiment also provides a data processing method based on a multimodal neural learning network model, the processing method comprising:

[0131] Encode the EEG signal to obtain EEG features;

[0132] Encode the original image to obtain image features; and

[0133] Calculate the similarity between the EEG feature and the image feature to obtain the task result,

[0134] The encoding of the original image includes:

[0135] Performing feature extraction on the original image to obtain a first primary feature of an embedding representation level;

[0136] Filtering the original image to obtain a filtered image, and performing feature extraction on the filtered image to obtain a second primary feature of an embedding representation level;

[0137] fusing the first primary feature with the second primary feature based on a cross-attention mechanism at the embedding representation level to obtain a secondary feature; and

[0138] The image features are obtained based on the inference results of the secondary features.

[0139] The content of the above processing method is similar to that of the model and will not be repeated here.

[0140] Example 2

[0141] like Figure 2As shown, this embodiment provides a multimodal neural cue learning framework (NeuralCLIP) for EEG image alignment. The framework is based on the pre-trained CLIP-ViT-L / 14 model. By constructing four key modules, namely EEG signal perturbation enhancement coding, two-stream dynamic self-cue image coding, distributed-aware embedding representation (token) level fusion module, and visual adaptive cue learning mechanism, it optimizes the visual modality representation from EEG signals for modality alignment and completes the 200-category (200-way) zero-shot retrieval task from EEG signals to images.

[0142] The CLIP-ViT-L / 14 model is a variant of the CLIP model that adopts the Vision-Transformer (ViT) architecture and has powerful image coding capabilities.

[0143] The specific technical implementation of the above framework is as follows:

[0144] 1. EEG signal perturbation enhancement coding module:

[0145] This module is divided into two parts: one is the EEG signal timing perturbation layer, and the other is the EEG signal encoding layer.

[0146] The EEG signal timing perturbation layer receives the EEG signal tensor from the preprocessing stage, whose shape is [batch_size, channels, timesteps], and performs linear perturbation on the data of each channel at each time point to improve the model's robustness to signal changes. The processed output maintains the same shape as the original input and is used for feature extraction by the subsequent EEG encoder.

[0147] The EEG signal timing perturbation layer includes the following parts:

[0148] The learnable weight parameter layer is used to multiply the input EEG signal tensor by an independent learnable scalar weight at each channel and each time point to obtain an intermediate representation with enhanced temporal amplitude.

[0149] The learnable bias parameter layer is used to add a learnable bias matrix of the same size [channels, timesteps] to the above intermediate representation, so as to further perform fine-grained perturbation adjustment on the EEG signal and enhance the model's adaptability to background noise or weak EEG fluctuations.

[0150] The overall strategy implemented by the EEG signal timing perturbation layer is a point-by-point learnable linear transformation in the channel-time dimension, which acts as a lightweight perturbator and can automatically learn the optimal perturbation method for each channel-time point during training.

[0151] The function of the EEG signal encoding layer is to receive the perturbation-enhanced EEG signal data, whose shape is [batch_size, channels, timesteps], flatten the input EEG signal data, and map it to a unified embedding space to obtain the EEG representation vector [batch-size, feature_dim] for alignment with image features, where feature_dim represents the feature dimension.

[0152] The EEG signal encoding layer includes the following parts:

[0153] The flattening layer is used to flatten the EEG signal data after perturbation enhancement to obtain flattened EEG data of [channels×timesteps] dimensions.

[0154] The fully connected layer (Linear) is used to linearly project the flattened EEG data of the input [channels × timesteps] dimension and map it to the specified feature_dim embedding space (e.g., 768 dimensions) to obtain the initial EEG feature representation.

[0155] The layer normalization layer (LayerNorm) is used to normalize the initial EEG feature representation, improve the stability of model training and the consistency of feature distribution, thereby enhancing the expressive ability of cross-modal alignment.

[0156] The EEG signal encoding layer is responsible for extracting discriminative representations from raw neural signals. The output EEG embedding can be directly used for comparative learning with image modalities in a shared space. Its specific structure is shown in Table 1:

[0157] Table 1 Structure of EEG signal encoding layer

[0158]

[0159] Among them, Reshape means flattening, represents the perturbation output of the timing perturbation layer after the EEG signal encoding module, R represents a set of real numbers, T represents the time step of the EEG signal sample, d represents the dimension, and Normalize over d represents the normalization operation on the features of d dimensions.

[0160] 2. Dual-stream dynamic self-prompting image encoding module:

[0161] The function of this module is to receive the input original image x, x∈[batch_size,C,H,W], and extract the patch level (patch_level) representation of the original image and the dynamic filtered image in a dual-stream manner.

[0162] After receiving the original image x, one branch directly passes the original image through a layer of image block embedding layer with frozen weights to obtain the original image block-level embedding (patch embedding) vector, and the other branch passes it through the dynamic filtering layer to obtain the dynamically filtered image features, and then passes it through the image block embedding layer shared with the first branch to obtain the filtered image block-level embedding vector.

[0163] The dual-stream dynamic self-hinting image encoding module consists of the following parts:

[0164] The image content-based filter generation network first receives the original image x, which is first subjected to three layers of lightweight convolution (Conv2d), batch normalization (BatchNorm2d), and ReLU activation to obtain a downsampled intermediate feature map. This intermediate feature map is then converted into a batch feature [B, hidden_dim] through adaptive average pooling (AdaptiveAvgPool2d(1,1)), where hidden_dim represents the hidden layer dimension. The batch feature is then mapped to a dynamic filter parameter vector [B, H × W × C] through two fully connected layers. The purpose of this network is to generate a set of dynamic filter parameters specific to the original image.

[0165] The dynamic image filter application network receives the dynamic filter parameters generated by the filter generation network and applies them to the original image to achieve dynamic convolution. The processing process is as follows: the original image x and the dynamic filter parameters are first structured into an independent H×W convolution kernel for each channel. Then, a local image block is expanded through a two-dimensional convolution layer. Each channel is then multiplied by the corresponding learnable filter kernel weight and summed to obtain the filtered image, with a final shape of [B, C, H, W].

[0166] The image block embedding layer passes the original image input and the filtered image features through the image block embedding layer (a convolutional layer with shared and frozen weights) to obtain feature representations at two embedding levels: the original image block-level embedding and the filtered image block-level embedding.

[0167] 3. Fusion Module

[0168] The function of this module is to achieve dynamic fusion of the original image block-level embedding and the filtered image block-level embedding based on the cross-attention mechanism.

[0169] The fusion module consists of the following parts:

[0170] The crisscross attention layer performs crisscross attention on the input original image block-level embedding and the filtered image block-level embedding, generating a crisscross attention matrix. In the crisscross attention mechanism, one sequence (the query sequence) focuses on the other sequence (the key sequence and the value sequence) through an attention mechanism, allowing the model to capture the relationship between the two sequences.

[0171] The feature vector layer is used to use the obtained cross-attention matrix to perform weighted fusion of the original image block-level embedding and the filtered image block-level embedding to obtain the fused embedding feature representation.

[0172] 4. Visual Adaptive Prompt Learning Module:

[0173] The role of this module is to learn a set of visual cue vectors to guide the image encoder part (CLIP-VIT block) in the CLIP model to better learn the global features of the image.

[0174] The visual adaptive cue learning module includes the following parts:

[0175] The visual cue layer is used to provide prompts for the fused embedded feature representation obtained by the fusion module. By introducing visual cue embedding and category embedding and concatenating them with the fused embedded feature representation, an embedded feature vector with prompt function is obtained.

[0176] The inference layer is used to embed feature vectors with prompt functions based on further encoding and splicing. This part uses the pre-trained open source CLIP-VIT model, freezes the weights of the model, and is only responsible for inferring the encoding to obtain the global features of the image after inference.

[0177] The image feature projection layer is used to map the global image features obtained by the image encoder module to the same dimension as the EEG signal sample features. It is a fully connected layer to facilitate the subsequent calculation of the loss function.

[0178] This embodiment also provides a testing method based on the aforementioned multimodal neural cue learning framework for EEG image alignment (NeuralCLIP), which specifically includes:

[0179] Step 1: Perform dynamic filtering on the input original image through the dynamic filter layer to obtain the image features after dynamic filtering.

[0180] Step 2: The image features of the original image and the dynamically filtered image features obtained in step 1 are embedded through a shared convolutional layer (the first layer of CLIP-ViT, with frozen weight parameters) to obtain embedded representation (token) vectors.

[0181] Step 3: The embedded representations of the original image and the dynamically filtered image obtained in step 2 are processed by a distribution-aware embedded representation level fusion module to obtain fused embedded representation features.

[0182] Step 4: The embedded representation features obtained in step 3, the classification block embedding representation (cls token), the prompt embedding representation (prompt token) generated by the visual adaptive prompt learning module, and the position embedding block (position embedding token) are sent to the image encoder part (CLIP-VIT block) for inference. Finally, the classification block embedding representation result of the last layer is taken as the encoding feature of the image.

[0183] Step 5: Input the input EEG signal features directly into the EEG encoder for projection to obtain EEG coding features. This step is actually completed synchronously and in parallel with step 1.

[0184] Step 6: Calculate the similarity between all EEG encoding features and all image features, and calculate the TOP-1 and TOP-5 accuracy rates for the entire retrieval recall pool. Specifically, for all test samples in the recall pool (which consists of N EEG signal samples and corresponding N image samples, where each EEG signal sample has exactly one image sample as its positive sample), calculate the similarity between each EEG signal sample and all N image samples, obtain the similarity value with the N image samples, and sort the similarities from high to low. If the image with the EEG signal sample as the positive sample ranks first in similarity, it is a TOP-1 accuracy hit; if it ranks in the top 5, it is a TOP-5 accuracy hit. The final TOP-1 and TOP-5 accuracy rates for the retrieval test are calculated as the quotient of the number of hits M and N: M / N.

[0185] The above model achieves state-of-the-art performance on the 200-way zero-shot retrieval task. Figure 3 and Figure 4 As shown in Figure 2. On the THINGS-EEG2 dataset, we outperform the state-of-the-art model (ATM) by 25.0% and 20.6% in the within-subject TOP-1 and TOP-5 accuracy criteria, respectively, and by 5.2% and 6.6% in the cross-subject TOP-1 and TOP-5 accuracy criteria, respectively. The model calculates the cross-modal sample similarity matrix as shown in Figure 2. Figures 5 to 14As shown in the figure, the vertical axis corresponds to EEG signal samples, and the horizontal axis corresponds to image samples. These similarity matrices show the predicted similarity between EEG signals and image samples. Brighter similarity matrices (colors closer to yellow) indicate higher predicted similarity. The diagonal line represents the predicted similarity value for positive sample pairs. That is, the brightness of the colors on the diagonal line is clearly distinguishable from the brightness of the other parts. This indicates that the model is better at distinguishing positive and negative samples, and is more likely to retrieve truly matching samples, indicating better retrieval performance.

[0186] An embodiment of the present application also provides an electronic device, comprising: a processor, and a memory coupled to the processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the method described in any one of the above embodiments.

[0187] The electronic device may be a computing device such as a desktop computer, a notebook computer, a PDA, a cloud server, etc. The electronic device may include, but is not limited to, a processor and a memory.

[0188] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the electronic device, connecting various parts of the entire device using various interfaces and lines.

[0189] The memory may be used to store the computer program, and the processor implements various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.

[0190] The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function, and the like; the data storage area may store data generated based on the use of the mobile phone. Furthermore, the memory may include high-speed random access memory (RAM) and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0191] The embodiment of the present application also provides a storage medium, which is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium.

[0192] An embodiment of the present application further provides a computer program product, including: a computer program or instructions, which, when executed on a computer, causes the computer to execute any of the above-mentioned possible implementation methods.

[0193] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications are also considered to be within the scope of protection of the present application.

Claims

1. A multimodal neural learning network model, characterized in that: The model includes: An EEG signal encoding module is used to encode EEG signals and obtain EEG features; An image encoding module, used to encode the original image to obtain image features; and A similarity calculation module is used to calculate the similarity between the EEG features and the image features to obtain the task results. Wherein, the image encoding module includes: A first image encoding submodule is used to extract features from the original image to obtain a first primary feature at an embedding representation level; a second image encoding submodule, configured to filter the original image to obtain a filtered image, and perform feature extraction on the filtered image to obtain a second primary feature at an embedding representation level; a fusion module, configured to fuse the first primary feature with the second primary feature based on a cross-attention mechanism at the embedding representation level to obtain a secondary feature; and An inference module, configured to obtain the image features based on the inference results of the secondary features, The second image encoding submodule includes: A filter parameter generation network, configured to convert the original image into a dynamic filter parameter vector, wherein the dynamic filter parameter vector is defined by a batch size, a filter kernel size, and a number of channels; a filter application network, configured to perform dynamic convolution on the original image using the dynamic filter parameter vector to obtain the filtered image; and The image block embedding layer is used to extract features from the filtered image to obtain the second primary features.

2. The model according to claim 1, wherein: The EEG signal encoding module includes: A time series perturbation layer is used to assign a learnable weight to each channel and each time point of the preprocessed EEG signal to obtain an intermediate representation with enhanced time series amplitude, and to superimpose a learnable bias matrix on the intermediate representation to obtain an enhanced representation; and The EEG signal encoding layer is used to flatten the enhanced representation to obtain flattened EEG data, linearly project the flattened EEG data to obtain an initial feature representation, and perform layer normalization on the initial feature representation to obtain the EEG feature.

3. The model according to claim 1, wherein: The filter parameter generation network includes: 3 lightweight convolution layers, a batch normalization layer, an activation layer, an adaptive average pooling layer and 2 fully connected layers arranged in sequence.

4. The model according to claim 1, wherein The fusion module includes: a cross-attention layer, configured to perform cross-attention calculation on the first primary feature and the second primary feature to obtain a cross-modal attention weight matrix; and The feature vector layer is used to use the attention weight matrix to perform weighted fusion on the first primary feature and the second primary feature to obtain the secondary feature.

5. The model according to claim 1, wherein: The image encoding module also includes: A visual adaptive cue learning module is configured to add cue embedding representations to the secondary features, wherein the cue embedding representations are iteratively updated according to a loss function.

6. The model according to claim 1, wherein: The model adopts a soft label contrast loss, which includes a hard label contrast loss term and a contrast loss term based on softening of intra-modal distribution similarity.

7. A data processing method based on a multimodal neural learning network model, characterized in that: The processing method comprises: Encode the EEG signal to obtain EEG features; Encode the original image to obtain image features; and Calculate the similarity between the EEG feature and the image feature to obtain the task result, The encoding of the original image includes: Performing feature extraction on the original image to obtain a first primary feature of an embedding representation level; Filtering the original image to obtain a filtered image, and performing feature extraction on the filtered image to obtain a second primary feature of the embedding representation level; fusing the first primary feature with the second primary feature based on a cross-attention mechanism at the embedding representation level to obtain a secondary feature; and Based on the inference result of the secondary feature, the image feature is obtained, Filtering the original image to obtain a filtered image, and performing feature extraction on the filtered image to obtain a second primary feature of the embedding representation level, specifically comprising: A filter parameter generation network, configured to convert the original image into a dynamic filter parameter vector, wherein the dynamic filter parameter vector is defined by a batch size, a filter kernel size, and a number of channels; a filter application network, configured to perform dynamic convolution on the original image using the dynamic filter parameter vector to obtain the filtered image; and The image block embedding layer is used to extract features from the filtered image to obtain the second primary features.

8. An electronic device, characterized in that: The electronic device includes: a processor, and a memory coupled to the processor, The memory is used to store computer programs; and The processor is configured to execute the computer program stored in the memory, so that the electronic device executes the processing method according to claim 7.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a computer program or instructions. When the computer program or instructions are executed on a computer, the computer is caused to execute the processing method according to claim 7 .

Citation Information

Patent Citations

  • Infrared weak and small target tracking method based on semi-supervised twin network

    CN114299111A

  • Model training method and device, electronic equipment, storage medium and program product

    CN117216546A