Cross-modal Decision Confidence Estimation Method and System Based on Generative Adversarial Learning

Through the generation and adversarial learning method, using eye movement signals to generate EEG features, the complexity and high cost of EEG signal acquisition are solved, and accurate decision-making confidence estimation is achieved in daily scenarios.

CN116439720BActive Publication Date: 2025-07-04SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310416057.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2025-07-04
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

In the prior art, the acquisition process of EEG signals is complex and costly, which affects the practical application of decision-making confidence estimation, especially the wearing of EEG caps and the use of conductivity gels, resulting in a decrease in comfort and a poor signal quality.

Method used

Generative adversarial learning method is adopted, by collecting the subject's EEG signals and eye movement signals, extracting features and training the EEG generative model, so that the model can generate EEG features from eye movement signals, and using a multimodal classifier for decision-making confidence estimation, reducing dependence on EEG signals.

Benefits of technology

It improves the accuracy and convenience of decision-making confidence estimation, solves the complexity and cost of EEG signal acquisition, and enables decision-making confidence estimation to be applied in daily scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116439720B_ABST
    Figure CN116439720B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a cross-modal decision confidence estimation method and system based on generative adversarial learning. The method includes: collecting electroencephalogram (EEG) signals and eye movement signals of a subject during the decision-making process; extracting EEG features and eye movement features from the EEG signals and eye movement signals; performing a first training on an EEG generation model to extract decision confidence features through the EEG features and eye movement features, and performing a second training on the relationship between eye movement and EEG through generative adversarial learning from the decision confidence features; inputting the obtained real eye movement signals of the subject into the trained EEG generation model to obtain predicted EEG signals, and performing decision confidence estimation on the real eye movement signals and the predicted EEG signals based on a multi-modal classifier. The embodiment of the present invention utilizes generative adversarial learning to improve the decision confidence estimation ability based on eye movement signals, solves the problem of high complexity and cost in the process of collecting EEG signals, and ensures accurate decision confidence estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal brain-computer interfaces, and in particular, to a cross-modal decision confidence estimation method and system based on generative adversarial learning. Background Art

[0002] Decision confidence refers to the degree of confidence of an individual in the optimality or correctness of their decision when making a judgment or decision. Decision confidence is a common psychological phenomenon in people's daily lives and is one of the important factors affecting decision-making results.

[0003] In the prior art, physiological signals such as eye movement and electroencephalogram can be used to estimate the confidence level of an individual during the decision-making process. Although the above-mentioned multimodal methods using eye movement and electroencephalogram are more reliable than unimodal methods, the multimodal methods mean that it is necessary to spend more costs to collect multimodal data. Especially, the acquisition of physiological signals requires the use of contact devices, which greatly limits its application in real scenarios. In particular, the process of collecting electroencephalogram signals is very complex.

[0004] In the process of implementing the present invention, the inventors found that there are at least the following problems in the related art:

[0005] In the experiment, the subject needs to correctly wear an electroencephalogram cap and inject conductive gel to ensure that the electrodes are in the corresponding positions. This is a relatively complicated and time-consuming task. Secondly, the wearing of the electroencephalogram cap will affect the comfort of the subject during the experiment to a certain extent, and as time goes by, the electroencephalogram gel will slowly dry out, affecting the quality of the collected electroencephalogram signals. Therefore, the use of electroencephalogram caps is restricted in practical applications, and it is difficult to obtain electroencephalogram signals for judging decision confidence. Summary of the Invention

[0006] In order to at least solve the problem that it is difficult to obtain electroencephalogram signals for judging decision confidence in the prior art.

[0007] In a first aspect, an embodiment of the present invention provides a cross-modal decision confidence estimation method based on generative adversarial learning, including:

[0008] Collecting electroencephalogram signals and eye movement signals of a subject during the decision-making process;

[0009] Extracting electroencephalogram features and eye movement features from the electroencephalogram signals and the eye movement signals;

[0010] Performing a first training on the electroencephalogram generation model to extract decision confidence features through the electroencephalogram features and the eye movement features, and performing a second training on the generative adversarial learning from the decision confidence features to determine the relationship between eye movement and electroencephalogram, so that the trained electroencephalogram generation model can generate corresponding electroencephalogram features from the input eye movement features;

[0011] Input the obtained true eye movement signal of the subject into the trained electroencephalogram (EEG) generation model to obtain a predicted EEG signal, and perform decision confidence estimation on the true eye movement signal and the predicted EEG signal based on a multimodal classifier.

[0012] In a second aspect, an embodiment of the present invention provides a cross-modal decision confidence estimation system based on generative adversarial learning, including:

[0013] A signal acquisition program module for acquiring the EEG signal and the eye movement signal of the subject during the decision-making process;

[0014] A feature extraction program module for extracting EEG features and eye movement features from the EEG signal and the eye movement signal;

[0015] A generative model training program module for performing a first training on the EEG generation model to extract decision confidence features through the EEG features and the eye movement features, and a second training for generative adversarial learning to determine the relationship between eye movement and EEG from the decision confidence features, so that the trained EEG generation model can generate corresponding EEG features from the input eye movement features;

[0016] An estimation program module for inputting the obtained true eye movement signal of the subject into the trained EEG generation model to obtain a predicted EEG signal, and performing decision confidence estimation on the true eye movement signal and the predicted EEG signal based on a multimodal classifier.

[0017] In a third aspect, an electronic device is provided, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the cross-modal decision confidence estimation method based on generative adversarial learning according to any embodiment of the present invention.

[0018] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the cross-modal decision confidence estimation method based on generative adversarial learning according to any embodiment of the present invention are implemented.

[0019] The beneficial effects of the embodiments of the present invention are as follows: By using generative adversarial learning, the decision confidence estimation ability based on eye movement signals is improved, the problem of complex and high-cost EEG signal acquisition process is solved, and decision confidence estimation can be applied to daily scenarios, and accurate decision confidence estimation can be ensured even when only eye movement signals are required. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 is a flowchart of a cross-modal decision confidence estimation method based on generative adversarial learning provided by an embodiment of the present invention;

[0022] Figure 2 is a framework diagram of a cross-modal decision confidence estimation method based on generative adversarial learning provided by an embodiment of the present invention;

[0023] Figure 3 is a test data diagram of a cross-modal decision confidence estimation method based on generative adversarial learning provided by an embodiment of the present invention;

[0024] Figure 4 is a test schematic diagram of a cross-modal decision confidence estimation method based on generative adversarial learning provided by an embodiment of the present invention;

[0025] Figure 5 is a structural schematic diagram of a cross-modal decision confidence estimation system based on generative adversarial learning provided by an embodiment of the present invention;

[0026] Figure 6 is a structural schematic diagram of an embodiment of an electronic device for cross-modal decision confidence estimation based on generative adversarial learning provided by an embodiment of the present invention. Detailed implementation manners

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0028] As Figure 1 shown is a flowchart of a cross-modal decision confidence estimation method based on generative adversarial learning provided by an embodiment of the present invention, including the following steps:

[0029] S11: Collect the electroencephalogram (EEG) signals and eye movement signals of the subject during the decision-making process;

[0030] S12: Extract EEG features and eye movement features from the EEG signals and the eye movement signals;

[0031] S13: Perform the first training of extracting decision confidence features on the EEG generation model through the EEG features and the eye movement features, and perform the second training of generating adversarial learning to determine the relationship between eye movement and EEG from the decision confidence features, so that the trained EEG generation model can generate corresponding EEG features from the input eye movement features;

[0032] S14: Input the obtained real eye movement signal of the subject into the trained EEG generation model to obtain a predicted EEG signal, and perform decision confidence estimation on the real eye movement signal and the predicted EEG signal based on a multi-modal classifier.

[0033] In this embodiment, considering that EEG signals are difficult to collect and are accompanied by scene limitations, while other physiological signals, such as eye movement signals, are relatively easier to collect, and only need to wear an eye tracker to complete the real-time collection of eye movement signals. Therefore, the research goal of this method is to establish a reliable and robust model, so that the model can perform confidence decision estimation in the absence of the EEG modality and also achieve relatively satisfactory performance.

[0034] For step S11, in the initial stage of this method, a certain amount of EEG signals and eye movement signals are still required to test the subjects who need decision confidence estimation, so as to obtain EEG signals and eye movement signals.

[0035] As an implementation manner, the collection of the EEG signal and the eye movement signal of the subject during the decision-making process includes:

[0036] Collect the EEG signal and the eye movement signal of the subject during the decision-making process based on an EEG acquisition device and an eye movement acquisition device.

[0037] In this embodiment, the EEG acquisition device can adopt an ESI NeuroScan wet electrode EEG cap in an experimental environment, a medical environment or other environmental scenarios, and the eye movement signal can be collected using a Tobii Pro X3-120 eye tracker. Specifically, when performing decision confidence estimation on the subject, various tests are displayed to the subject through a display screen. The subject wearing an ESI NeuroScan wet electrode EEG cap and a Tobii Pro X3-120 eye tracker will give feedback on various tests. At this time, the EEG acquisition device and the eye movement acquisition device collect the EEG signal and the eye movement signal of the subject during the decision-making process.

[0038] As an implementation manner, after collecting the EEG signal and the eye movement signal of the subject during the decision-making process, the method further includes preprocessing the EEG signal and the eye movement signal, including:

[0039] Perform baseline correction on the EEG signals, remove the 50 Hz AC power noise in the EEG signals after baseline correction, and remove the low-frequency signals and high-frequency invalid signals in the EEG signals based on a 1-75 HZ band-pass filter;

[0040] Remove the signals affected by light in the eye movement signals based on the principal component analysis method.

[0041] In this embodiment, perform processing such as baseline correction, artifact removal, and filtering on the EEG signals. Specifically, the 50 Hz AC power noise in the EEG signals can be removed, and a 1-75 Hz band-pass filter is used to remove the low-frequency and high-frequency invalid signals in the EEG signals. For the eye movement signals, the principal component analysis method can be used to remove the part affected by light in the eye movement signals. Through preprocessing, the accuracy of feature correlation in subsequent steps can be further improved.

[0042] For step S12, in order to find the internal relationship between EEG and eye movement features, it is necessary to extract EEG features and eye movement features from the EEG signals and the eye movement signals in advance.

[0043] Specifically, extracting EEG features and eye movement features from the EEG signals and the eye movement signals includes:

[0044] Perform short-time Fourier transform on the EEG signals based on a fixed-length Hanning window, divide the EEG signals after short-time Fourier transform into multiple frequency bands, and determine the differential entropy features of the EEG in the multiple frequency bands.

[0045] Determine the mean values, variances, differential entropy features of eye movement in multiple frequency bands, and eye movement features in the X-axis and Y-axis directions through the eye movement signals, where the multiple frequency bands of the differential entropy features of eye movement include: 0-0.2 Hz, 0.2-0.4 Hz, 0.4-0.6 Hz, 0.6-0.8 Hz, and the eye movement features include: fixation time, blink frequency.

[0046] For step S13, in this method, the training of the EEG generation model can be divided into two parts: in the first training stage, perform high-level feature extraction and generate high-level EEG (electroencephalography) features. Specifically, use a DAE (deep auto encoder) to learn the high-level features of each modality for identifying the decision confidence level, and then in the second training stage, train GANs (generative adversarial networks) to generate features of the EEG modality according to eye movement. In the test stage, only the eye movement signals are required, and the corresponding EEG features are generated from them.

[0047] As an implementation, the above training includes:

[0048] In the first training, based on the deep autoencoder, determine the reconstruction losses of the EEG features and the eye movement features, generate the eye movement high-order features and the EEG high-order features representing the decision confidence through the reconstruction losses, and determine the eye movement high-order features and the EEG high-order features as the decision confidence features;

[0049] In the second training, input the eye movement high-order features into the EEG generation model to obtain the predicted EEG high-order features, discriminate the cross-entropy loss between the EEG high-order features and the predicted EEG high-order features, and perform generative adversarial learning on the EEG generation model based on the cross-entropy loss.

[0050] In this implementation, the first training stage: Feature extraction:

[0051] Suppose The data representing the eye movement signal, The data representing the EEG signal. N represents the quantity, and d1 and d2 are the dimensions of the initial EEG and eye movement features. E eye 、D eye Represent the encoder and decoder of eye movement, E eeg 、D eeg Represent the encoder and decoder of the EEG signal, u eye 、v eye 、u eeg 、v eeg Represent their respective parameters. The output of the encoder can be expressed as:

[0052] O eye =E eye (X eye ;u eye )·O eeg =E eeg (X eeg ;u eeg )

[0053] Correspondingly, the output of the decoder is:

[0054]

[0055] The reconstruction loss of the DAE can be expressed as the reconstruction losses loss rec (L RC ) of the eye movement and the EEG features:

[0056]

[0057] Then, this method selects the aggregation-based fusion as the multi-modal fusion strategy, which directly combines Oeye and O eeg are concatenated (replaced with O1 and O2, and after concatenation it is (O1, O2)). Calculate the cross - entropy loss loss between the predicted emotion category and the true category Y cls :

[0058] loss cls = CrossEntropy(E s (O1, O2), Y)

[0059] The optimization of the entire model is achieved by minimizing the sum of the reconstruction loss and the classification loss:

[0060] loss = λ rec loss rec + λ cls loss cls

[0061] where λ rec and λ cls are the balance parameters between the losses.

[0062] Second training stage: EEG high - level feature generation. After the first training stage, based on the deep auto - encoder, the high - level features of eye movement O eye and EEG O eeg are obtained respectively. Then, O eye is used as the input of GANs to guide the generator to generate the corresponding EEG features.

[0063]

[0064] where θ represents the parameters of the generative adversarial network G.

[0065] The discriminator D is a binary classifier used to distinguish between real modality pairs and predicted modality pairs. Give the real multimodal data (O eye , O eeg ) a label of 1, and give the predicted multimodal (O eye , ) a label of 0. Minimize the cross - entropy loss loss D to train the discriminator (similarly, replace O eye , O eeg with O1, O2):

[0066]

[0067] For the generator, its goal is to make the discriminator unable to distinguish the generated EEG features from the real EEG features, and at the same time be closest to the corresponding real EEG features. Therefore, it is optimized by minimizing loss G :

[0068]

[0069] Among them, λ g and λ mse respectively represent the balance parameters between the two.

[0070] In addition, this method also adopts a content loss function to encourage to be close to O eeg . This can be achieved by minimizing the Euclidean distance between them, resulting in the mean squared error (MSE) loss L-MSE being defined as:

[0071]

[0072] Among them, L-MSE encourages learning the detailed information for completing the EEG modality. Therefore, the overall loss function of the generative adversarial network G can be described as:

[0073]

[0074] The training structures of the above two stages are as shown in the upper half of Figure 2 . Through the training of the above method, the trained EEG generation model can generate corresponding EEG features from the input eye movement features.

[0075] For step S14, after the training of steps S11 - S13, the EEG generation model can simulate the corresponding EEG signal according to the eye movement signal of the subject. The reason for such processing is that the estimation of the subject's decision-making confidence is not a one-time thing and needs to be continuously followed up. However, it is impossible to let the user keep estimating in the laboratory or medical scenario all the time, which is also a consumption of money and time for the subject. Therefore, let the subject use high-precision EEG and eye movement acquisition devices in the initial laboratory or medical scenario to train an EEG generation model exclusive to the subject.

[0076] Such as Figure 2As shown in the lower part of [], cross-modal was carried out in the subsequent decision confidence estimation test stage, that is, electroencephalogram (EEG) signals were not collected, and only eye movement signals were used for prediction. Although this method has a certain loss in recognition rate compared with the multi-modal method, the cross-modal method reduces the dependence on EEG signals, making this work easier to be popularized in practical applications. For example, the subject can use a portable eye tracker to collect real eye movement signals in a home environment scenario or other scenarios (due to different eye movement acquisition devices, there is a slight difference in the accuracy of the eye movement signals collected compared with those in step S11). The collected real eye movement signals are input into the trained EEG generation model to generate predicted EEG signals. At this time, only the eye movement of the subject needs to be collected to very conveniently realize the estimation of decision confidence, and the accuracy of the estimation is also guaranteed.

[0077] It can be seen from this implementation that by using generative adversarial learning, the ability to estimate decision confidence based on eye movement signals is improved, the problem of complex and high-cost EEG signal acquisition process is solved, and decision confidence estimation can be applied to daily scenarios, and accurate decision confidence estimation can also be ensured when only eye movement signals are required.

[0078] The method was experimentally demonstrated. The method was experimented on the SEED-VPDC dataset. This dataset is a multi-modal dataset, including EEG signals and eye movements, and is used to measure five-level decision confidence. The experiment consisted of 135 trials, each trial containing an image corresponding to a decision. The stimulus materials included three types selected from the Caltech 101 dataset. 14 subjects participated in the experiment, and eye movements and electroencephalogram signals were recorded simultaneously throughout the experiment. The eye movement data of one subject was incomplete, and the eye movement and electroencephalogram signals of 13 subjects were complete.

[0079] For eye movement signals, 22 features were extracted by a Tobii Pro X3-120 screen eye tracker, including pupil diameter, fixation duration, blink duration, and saccade duration. According to the international 10-20 system, EEG signals were recorded by a 62-channel active AgCl electrode cap with an ESI Neuroscan system at a sampling rate of 1000 Hz. For data preprocessing, a band-pass filter between 0.3 and 50 Hz was applied to each channel to filter noise, and the LDS (linear dynamic system) method was used to smooth the features. In non-overlapping 1-second time windows, DE (Differential entropy) features were extracted from 5 frequency bands (i.e., δ: 1-3 Hz, θ: 4-7 Hz, α: 8-13 Hz, β: 14-30 Hz, and γ: 31-50 Hz) of each sample, which has been proven to have the best performance in decision confidence classification.

[0080] This method adopts a five-fold cross-validation method and a subject (i.e., the participant) related classification setting. Two classifiers, SVM (support vector machine) and DNNS (deep neural network with shortcut connections) are selected as baselines, and the radial basis function kernel is used to search for C in SVM from the parameter space of 2 [-5:10] The parameter space of. The DNNS method adopts four hidden layers and one output layer, and the size of the hidden layer is searched from 16 to 256, and the learning rate is set to 0.001. These two classifiers are respectively tested and trained on eye movement and multimodal data. To further verify the performance gap between the cross-modal model and the multimodal model of this method, an aggregation-based fusion is selected as the multimodal fusion strategy, which connects the high-level eye movement and EEG features extracted from the DAE. Finally, to verify the performance of the cross-modal method of this method, the model of DAL (deep adversarial learning) is tested on the SEED-VPDC dataset, which directly generates EEG information from the main eye movement features.

[0081] As Figure 3 shown are the experimental results including the accuracy and F1 score of different methods. The average accuracy and standard deviation of the model of this method are compared with other methods.

[0082] For electroencephalogram and eye movement, from Figure 3 it is found that the DNNS method is significantly better than the SVM method in any modality, which proves the superiority of the neural network. In addition, whether using the SVM method or the DNNS method, the classification ability of the EEG signal for decision confidence is stronger than that of eye movement, which indicates that in the decision confidence recognition task, the EEG information is more reliable than eye movement. From Figure 4 the confusion matrix shown, it can be seen that for low decision confidence levels (1 and 2), eye movement has a relatively high recognition rate, while the EEG signal has a stronger ability to distinguish extreme confidence levels (1 and 5). This indicates that there are complementary representations for measuring decision confidence between eye movement and EEG signals.

[0083] For the cross-modal and unimodal of this method, the cross-modal method proposed by this method is superior to the DNNS method trained and tested only on eye movement signals, with the accuracy rate increased by about 5.43% and the F1 score increased by 4.13%. In addition, it can be seen from the confusion matrix that the recognition rate of this method has a relatively large increase at all levels, especially at the extreme confidence level, with a maximum increase of 11.55% at the first level and 8.82% at the fifth level. Even when using eye movement as the only input, the electroencephalogram knowledge with stronger ability to recognize extreme confidence levels can be learned through this method.

[0084] For the cross-modal and multimodal of this method, this method is as competitive as the DNNS multimodal method, but there is still a performance gap compared with the DAE multimodal method. It can be seen from the confusion matrix that this method mainly performs poorly in recognizing extreme confidence levels. This can be explained by the fact that the generated EEG information cannot completely replace the real EEG signal. However, the moderate decrease in accuracy is considered acceptable compared with the multimodal method because this method only tests the eye movement signals, thus reducing the dependence on EEG signals, which makes the decision confidence measurement more applicable and feasible.

[0085] For the comparison of the cross-modal of this method with the existing cross-modal, compared with the DAL method, the DAL method generates EEG information from primary eye movement signals without going through the high-level feature extraction process. This method has an increase in accuracy of about 5.80% and an increase in F1 score of 5.16%. The DAE method obtains better performance than the multimodal DNNS method, which indicates that the high-level features extracted from the first training stage are more beneficial for decision confidence classification. More importantly, it is difficult to generate primary high-dimensional EEG features from low-dimensional eye movement features.

[0086] Generally speaking, this method proposes a cross-modal method based on generative adversarial learning for the decision confidence measurement task. In this method, the internal relationship between eye movement and EEG features in the high-level feature space can be learned during the training stage. The experimental results on the SEED-VPDC dataset show that this method is superior to the unimodal method trained and tested only on eye movement signals. This indicates that in the absence of EEG, EEG features can be generated from eye movement features, which supplements the information of the EEG modality to a certain extent.

[0087] As Figure 5 shown is a schematic structural diagram of a cross-modal decision confidence estimation system based on generative adversarial learning provided by an embodiment of the present invention. The system can execute the cross-modal decision confidence estimation method based on generative adversarial learning described in any of the above embodiments and is configured in a terminal.

[0088] A cross-modal decision confidence estimation system 10 based on generative adversarial learning provided by this embodiment includes: a signal acquisition program module 11, a feature extraction program module 12, a generative model training program module 13, and an estimation program module 14.

[0089] Among them, the signal acquisition program module 11 is used to collect the electroencephalogram (EEG) signal and the eye movement signal of the subject during the decision-making process; the feature extraction program module 12 is used to extract EEG features and eye movement features from the EEG signal and the eye movement signal; the generative model training program module 13 is used to perform a first training on the EEG generative model to extract decision confidence features through the EEG features and the eye movement features, and a second training to determine the relationship between eye movement and EEG through generative adversarial learning from the decision confidence features, so that the trained EEG generative model can generate corresponding EEG features from the input eye movement features; the estimation program module 14 is used to input the obtained real eye movement signal of the subject into the trained EEG generative model to obtain a predicted EEG signal, and perform decision confidence estimation on the real eye movement signal and the predicted EEG signal based on a multi-modal classifier.

[0090] An embodiment of the present invention also provides a non-volatile computer storage medium, and the computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the cross-modal decision confidence estimation method based on generative adversarial learning in any of the above method embodiments;

[0091] As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as:

[0092] Collect the electroencephalogram (EEG) signal and the eye movement signal of the subject during the decision-making process;

[0093] Extract EEG features and eye movement features from the EEG signal and the eye movement signal;

[0094] Perform a first training on the EEG generative model to extract decision confidence features through the EEG features and the eye movement features, and a second training to determine the relationship between eye movement and EEG through generative adversarial learning from the decision confidence features, so that the trained EEG generative model can generate corresponding EEG features from the input eye movement features;

[0095] Input the obtained real eye movement signal of the subject into the trained EEG generative model to obtain a predicted EEG signal, and perform decision confidence estimation on the real eye movement signal and the predicted EEG signal based on a multi-modal classifier.

[0096] As a non-volatile computer-readable storage medium, it can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present invention. One or more program instructions are stored in the non-volatile computer-readable storage medium and, when executed by a processor, perform the cross-modal decision confidence estimation method based on generative adversarial learning in any of the above method embodiments.

[0097] Figure 6 FIG. is a schematic hardware structure diagram of an electronic device for the cross-modal decision confidence estimation method based on generative adversarial learning provided in another embodiment of the present application, as Figure 6 shown, the device includes:

[0098] One or more processors 610 and a memory 620, Figure 6 Taking one processor 610 as an example. The device for the cross-modal decision confidence estimation method based on generative adversarial learning may further include: an input device 630 and an output device 640.

[0099] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected through a bus or other means, Figure 6 Taking connection through a bus as an example.

[0100] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the cross-modal decision confidence estimation method in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, implements the cross-modal decision confidence estimation method based on generative adversarial learning in the above method embodiments.

[0101] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data, etc. In addition, the memory 620 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories may be connected to the mobile device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0102] The input device 630 can receive input digital or character information. The output device 640 may include a display device such as a display screen.

[0103] The one or more modules are stored in the memory 620 and, when executed by the one or more processors 610, perform the cross-modal decision confidence estimation method based on generative adversarial learning in any of the above method embodiments.

[0104] The above product can execute the method provided in the embodiments of the present application, and has functional modules and beneficial effects corresponding to the execution of the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.

[0105] The non-volatile computer-readable storage medium may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created according to the use of the device, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0106] An embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the cross-modal decision confidence estimation method based on generative adversarial learning in any embodiment of the present invention.

[0107] The electronic devices in the embodiments of the present application exist in various forms, including but not limited to:

[0108] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0109] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, such as tablet computers.

[0110] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable vehicle navigation devices.

[0111] (4) Other electronic devices with data processing functions.

[0112] In this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising" and "including" not only include those elements, but also other elements not explicitly listed, or elements inherent to such a process, method, article, or device. Without further limitation, elements defined by the statement "comprising..." do not exclude the existence of additional identical elements in the process, method, article, or device that includes the said elements.

[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cross-modal decision confidence estimation method based on generative adversarial learning, comprising: Collecting electroencephalogram (EEG) signals and eye movement signals of a subject during the decision-making process; Extracting EEG features and eye movement features from the EEG signals and the eye movement signals; Performing a first training on an EEG generation model to extract decision confidence features through the EEG features and the eye movement features, and performing a second training on the relationship between eye movement and EEG through generative adversarial learning from the decision confidence features, so that the trained EEG generation model can generate corresponding EEG features from the input eye movement features. Wherein, in the first training, the reconstruction losses of the EEG features and the eye movement features are determined based on a deep autoencoder, and the eye movement high-order features and EEG high-order features for representing decision confidence are generated through the reconstruction losses, and the eye movement high-order features and EEG high-order features are determined as decision confidence features. In the second training, the eye movement high-order features are input into the EEG generation model to obtain predicted EEG high-order features, the cross-entropy loss between the EEG high-order features and the predicted EEG high-order features is discriminated, and generative adversarial learning is performed on the EEG generation model based on the cross-entropy loss; Inputting the obtained real eye movement signals of the subject into the trained EEG generation model to obtain predicted EEG signals, and performing decision confidence estimation on the real eye movement signals and the predicted EEG signals based on a multi-modal classifier.

2. The cross-modal decision confidence estimation method according to claim 1, wherein, The collecting of the EEG signals and the eye movement signals of the subject during the decision-making process includes: Collecting the EEG signals and the eye movement signals of the subject during the decision-making process based on an EEG acquisition device and an eye movement acquisition device.

3. The cross-modal decision confidence estimation method according to claim 1, wherein, After collecting the EEG signals and the eye movement signals of the subject during the decision-making process, the method further includes preprocessing the EEG signals and the eye movement signals, including: Performing baseline correction on the EEG signals, removing the 50 Hz alternating current power noise in the EEG signals after baseline correction, and removing the low-frequency signals and high-frequency invalid signals in the EEG signals based on a 1-75 HZ band-pass filter; Removing the signals affected by light in the eye movement signals based on the principal component analysis method.

4. The cross-modal decision confidence estimation method according to claim 1, wherein, The extracting of the EEG features and the eye movement features from the EEG signals and the eye movement signals includes: Performing short-time Fourier transform on the EEG signals based on a fixed-length Hanning window, dividing the EEG signals after short-time Fourier transform into multiple frequency bands, and determining the differential entropy features of the EEG in the multiple frequency bands.

5. The cross-modal decision confidence estimation method according to claim 4, wherein, The method further includes: Determining the mean, variance, differential entropy features of the eye movement in multiple frequency bands, and eye movement features in the X-axis and Y-axis directions of the pupil through the eye movement signals, wherein the multiple frequency bands of the differential entropy features of the eye movement include: 0-0.2 Hz, 0.2-0.4 Hz, 0.4-0.6 Hz, 0.6-0.8 Hz, and the eye movement features include: fixation time, blink frequency.

6. A cross-modal decision confidence estimation system based on generative adversarial learning, comprising: A signal acquisition program module for collecting the EEG signals and the eye movement signals of the subject during the decision-making process; A feature extraction program module for extracting electroencephalogram (EEG) features and eye movement features from the EEG signals and the eye movement signals; A generation model training program module for performing a first training on an EEG generation model to extract decision confidence features through the EEG features and the eye movement features, and performing a second training on the decision confidence features to generate an adversarial learning for determining the relationship between eye movement and EEG, so that the trained EEG generation model can generate corresponding EEG features from the input eye movement features. In the first training, based on a deep autoencoder, the reconstruction losses of the EEG features and the eye movement features are determined, and eye movement high-order features and EEG high-order features for representing decision confidence are generated through the reconstruction losses, and the eye movement high-order features and the EEG high-order features are determined as decision confidence features. In the second training, the eye movement high-order features are input into the EEG generation model to obtain predicted EEG high-order features, the cross-entropy loss between the EEG high-order features and the predicted EEG high-order features is discriminated, and the EEG generation model is subjected to generative adversarial learning based on the cross-entropy loss; An estimation program module for inputting the obtained real eye movement signals of the subject into the trained EEG generation model to obtain predicted EEG signals, and estimating the decision confidence of the real eye movement signals and the predicted EEG signals based on a multi-modal classifier.

7. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the cross-modal decision confidence estimation method according to any one of claims 1-5.

8. A storage medium, on which a computer program is stored, characterized in that, When the program is executed by a processor, it implements the steps of the cross-modal decision confidence estimation method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Electroencephalogram based behavior decision prediction system

    CN106175757A

  • Electroencephalogram emotion migration model training method and system and electroencephalogram emotion recognition method and device

    CN112690793A