Emotion classification method based on functional near infrared spectrum signals

By constructing the fNIRS-CNN model, combining the delayed hemodynamic response of functional near-infrared spectral signals and the activation characteristics of different brain regions, the feature redundancy and calculation cost problems of the emotional classification method of functional near-infrared spectral signals in the prior art are solved, and higher accuracy and better feature expression are achieved.

CN119939353APending Publication Date: 2025-05-06ZHONGBEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082562.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, the emotional classification method of functional near-infrared spectral signals mainly relies on traditional machine learning methods or multimodal deep learning methods, and fails to fully utilize the characteristics of functional near-infrared spectral signals, resulting in feature redundancy, insufficient key feature expression or unnecessary computational costs.

Method used

A emotion classification method based on functional near-infrared spectral signals is proposed. By constructing the fNIRS-CNN model, combining the delayed hemodynamic response of functional near-infrared spectral signals and the activation characteristics of different brain regions, the architecture of the deep learning model is designed to achieve accurate emotion classification.

Benefits of technology

Through the fNIRS-CNN model, the feature redundancy and calculation cost problems in the emotional classification of functional near-infrared spectral signals can be effectively solved, achieving higher accuracy and better feature expression, and improving the accuracy of emotional classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939353A_ABST
    Figure CN119939353A_ABST
Patent Text Reader

Abstract

The invention discloses an emotion classification method based on functional near infrared spectrum signals. The method comprises the following steps: S110, collecting signal data of functional near infrared spectrums of different emotions; s120, signal preprocessing is carried out; s130, a functional near infrared spectrum signal emotion classification model of the fNIRS-CNN is constructed; s140, the constructed fNIRS-CNN is subjected to training; and S150, giving a classification result according to the functional near infrared spectrum signal emotion classification method. According to the fNIRS-CNN model, accurate emotion classification can be carried out only through functional near infrared spectrum signal data, and the architecture of the deep learning model is designed by combining two characteristics of functional near infrared spectrum signals, delayed hemodynamic response and activation modes of different brain regions. A deep learning algorithm model designed by combining the two characteristics can effectively solve the problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0003] The present invention relates to the field of artificial intelligence and machine learning technology, and in particular to a method for classifying emotions based on functional near-infrared spectroscopy signals, providing strong technical support for the practical application of functional near-infrared spectroscopy signals. Background Art

[0005] The role of emotion classification and its application in functional near-infrared spectroscopy signals is an important research direction in the field of affective computing and biomedicine in recent years. Emotion classification aims to identify the emotional state of an individual by analyzing different signals (such as speech, facial expressions, behavioral patterns, physiological signals, etc.). The role of near-infrared emotion classification and its application in functional near-infrared spectroscopy signals is an important research direction in the field of affective computing and biomedicine in recent years. Emotion classification aims to identify the emotional state of an individual by analyzing different signals (such as speech, facial expressions, behavioral patterns, physiological signals, etc.). The application of near-infrared spectroscopy signals in emotion classification, especially for the detection and identification of emotional reactions, is gradually gaining high attention. Functional near-infrared spectroscopy technology is a non-invasive brain function detection method that does not require electrodes to be implanted directly into the brain and is relatively safe and comfortable to use. It can provide brain activity data with high temporal resolution and is suitable for real-time monitoring of emotional changes.

[0006] In terms of recognition methods, Soleymani M, Pantic M, Pun T. and Xu Y, Hübener I, Seipp A, in previous studies, mainly used traditional machine learning methods [1, 2] to fuse various features and use classifiers for classification. The early feature fusion methods were more of a simple splicing or normalization, and the late fusion (decision layer fusion) was to make the final classification result by integrating the results of each modality's separate classifier. Commonly used classifiers are generally SVM, K-Nearest Neighbor (KNN), linear discriminant analysis, random forest, etc. Although traditional machine learning methods have achieved certain results, the differences and complementarities between different modalities have not been fully utilized, which may lead to feature redundancy, insufficient expression of key features, or unnecessary computational costs [3].

[0007] With the rapid development of deep learning in the fields of image and speech, deep learning methods based on multimodal physiological signal emotion recognition have emerged on the basis of machine learning. The earliest research on multimodal physiological signal emotion recognition using deep learning was in 2016. Liu et al. [4] used a bimodal deep autoencoder (BDAE) to distinguish different emotions. This is a classic joint representation fusion model that projects different modalities into a joint space and extracts the shared representation of the modalities to distinguish different emotions. It achieved an accuracy of 91.01% on the SEED dataset EEG and eye movement signals, and greatly improved the recognition effect of the four binary classification tasks on the DEAP dataset EEG and peripheral physiological signals.

[0008] For the existing functional near-infrared spectroscopy signal emotion classification and recognition problem, there are only traditional machine learning algorithms and multimodal deep learning algorithms based on functional near-infrared spectroscopy signals and other signals (EEG), such as Figure 1 As shown in the figure, there are few deep learning methods based only on functional near-infrared spectroscopy signals.

[0009] Traditional machine learning methods for emotion classification of functional near infrared spectroscopy signals were proposed in papers [1], [2], [3]:

[0010] Soleymani M, Pantic M, Pun T. Multimodal emotion recognition inresponse to videos[J]. IEEE transactions on affective computing, 2011, 3(2):211-223.

[0011] Xu Y, Hübener I, Seipp A, et al. From the lab to the real-world: Aninvestigation on the influence of human movement on Emotion Recognition using physiological signals[C]. In: 2017 IEEE International Conference on PervasiveComputing and Communications Workshops (PerCom Workshops). IEEE, 2017. 345-350.

[0012] Zhang Y, Cheng C, Zhang Y. Multimodal emotion recognition using ahierarchical fusion convolutional neural network[J]. IEEE access, 2021, 9:7943-7951

[0013] The paper [4] proposed a new machine learning method for multimodal emotion classification of functional near infrared spectroscopy signals and other signals:

[0014] [4]Liu W, Zheng W, Lu B. Emotion recognition using multimodal deeplearning[C]. In: Neural Information Processing: 23rd InternationalConference, ICONIP 2016, Kyoto, Japan, October 16-21, 2016, Proceedings, PartII 23. Springer, 2016. 521-529. Summary of the invention

[0016] The technical problem to be solved by the present invention is to provide an emotion classification method based on functional near infrared spectroscopy signals in view of the deficiencies of the prior art.

[0017] The technical solution of the present invention is as follows:

[0018] An emotion classification method based on functional near infrared spectroscopy signals, characterized by comprising steps S110-S150:

[0019] S110: Collect signal data of functional near infrared spectra of different emotions, assign data labels using one-hot encoding according to different emotions, and convert category labels into numerical forms that can be processed by computers;

[0020] S120: signal preprocessing; constructing training data set and test data set;

[0021] S130: Constructing the functional near infrared spectroscopy signal emotion classification model of fNIRS-CNN; the network structure includes input end, DHR Model module, Global Model module and output end;

[0022] S140: training the constructed fNIRS-CNN;

[0023] S150: According to the functional near infrared spectroscopy signal emotion classification method, a classification result is given.

[0024] The emotion classification method, the data preprocessing method in step S120, comprises steps S210-S230:

[0025] S210 Modified Beer-Lambert law converts optical density The concentration changes of two hemoglobins HbO and HbR are converted into near-infrared light absorption;

[0026] The S220 uses a bandpass filter with a passband of 0.01-0.1 Hz for filtering;

[0027] The S230 baseline correction solves the problem of baseline drift by subtracting the average value of a reference interval from the NIR spectral signal.

[0028] The emotion classification method, step S210: the modified Beer-Lambert law is specifically implemented as follows:

[0029] The optical density The concentration changes of two hemoglobins HbO and HbR generated by near-infrared light absorption are converted into the modified Beer-Lambert law as follows:

[0030] in, and Wavelength The extinction coefficients of HbO and HbR at is the differential path length factor and l is the distance between the source and the detector.

[0031] The emotion classification method, step S220: using a bandpass filter with a passband of 0.01-0.1 Hz, allowing signals in the range of 0.01 Hz to 0.1 Hz to pass through, while suppressing frequency components below 0.01 Hz and above 0.1 Hz.

[0032] The emotion classification method, step S230: baseline correction solves the baseline drift problem by subtracting the average value of a reference interval from the near-infrared spectral signal, eliminating signal drift caused by instrument noise, environmental changes, sample changes, etc., to ensure that the analysis results are more accurate. Select a flat interval that does not contain any meaningful absorption peak area, has small signal changes, and has no absorption characteristics to represent the baseline. Within the selected reference interval, calculate the average or median of the signal intensity in the interval as an estimate of the baseline drift. Remove the baseline drift from the entire spectral signal and achieve correction by subtracting the average value of the reference interval from the original spectrum. The formula for baseline correction is as follows:

[0033]

[0034] in, is the original spectral signal, is the mean value of the reference interval, is the corrected spectral signal after removing the baseline drift.

[0035] In the emotion classification method, the construction of the fNIRS-CNN model in step S130 includes:

[0036] 310 At the input end, the signal data after data preprocessing is used as the input of the network;

[0037] 320 In the DHR Model module, there are five convolution blocks conv, which are used to extract the characteristics of delayed hemodynamic response of functional near-infrared spectroscopy signal data; each convolution block consists of a 2D convolution layer, a batch normalization layer and a Sigmoid activation layer; the specific calculation process includes: the input data passes through the first convolution (Conv) block, the second Conv block, the third Conv block, the fourth Conv block, and the fifth Conv block in turn, and the signal data in each test channel is subjected to dimensionality reduction feature extraction, and finally the extracted feature matrix is ​​obtained; each convolution operation is that the input data is a two-dimensional matrix X, the convolution kernel is W, is the position in the output data The value of is the input image at position The value of is the convolution kernel at position The weight of is the size of the convolution kernel; then the formula for the convolution operation is:

[0038]

[0039] 330 The Global Model module includes a depth-wise separable convolution consisting of two convolution blocks; the first convolution (DWConv) block uses depth-wise convolution to extract the correlation features of functional near-infrared spectral signals in different brain regions; the second convolution (PWConv) block uses point-wise convolution to increase the dimension of the feature matrix after depth-wise convolution to obtain more features and reduce the number of parameters of the convolution block; the output data of the DHR Model module is used as the input of the Global Model module, and the output of this module is obtained by convolution operations of DWConv and PWConv in sequence;

[0040] 340 Flatten the feature matrix output by each module in one dimension and input it to the fully connected layer for the final classification task; in the neural network, y is the output vector, and its dimension is usually the same as the number of neurons in this layer; W is the weight matrix, with a size of m×n, where m is the output dimension and n is the input dimension; x is the input vector, with a size of n×1, corresponding to the output of the previous layer; is the bias vector, with a size of m×1, and each output neuron has a corresponding bias term. The formula of the linear layer can be expressed as:

[0041]

[0042] In the emotion classification method, the training set data constructed in step S120 is used to train the model in step S140; the specific steps are as follows: first, the training set data is input into the network to calculate the result of this round of iterative calculation, the result is compared with the label assigned by the corresponding one-hot encoding, and the loss value is calculated using the cross entropy loss function after smoothing the label; secondly, the gradient is calculated by the stochastic gradient descent optimizer and the weights in the network are updated by back propagation; then the above process is iterated until the error requirements are met to obtain the network training model; finally, the generalization ability of the model is verified using the test data set. The loss function in the above steps is calculated as follows:

[0043] Label smoothing is often used to prevent DNNs from overconfidence by weighted averaging of hard targets and uniform distribution over labels. The network predicts each class label k∈{1,...,k}: where zi is the logits (the i-th element of the model output output vector).

[0044]

[0045] Where yk = 1 represents the basic truth value, yk = 0 represents the remaining truth values. The cross entropy loss function is defined as:

[0046]

[0047] Label smoothing is defined as, where ε is the smoothing parameter, which is set to 0.1 by default, and uk = 1 / K for uniform distribution:

[0048]

[0049] Finally, the cross entropy loss function with label smoothing is written as:

[0050] .

[0051] The present invention proposes an fNIRS-CNN model that can accurately classify emotions with only functional near infrared spectroscopy signal data. Figure 2 As shown in the figure. This method combines two characteristics of functional near-infrared spectroscopy signals, delayed hemodynamic response and activation mode of different brain regions, to design the architecture of the deep learning model. The deep learning algorithm model designed by combining these two characteristics can effectively solve the above problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 : An algorithm based on multimodal deep learning of functional near-infrared spectroscopy signals and other signals (EEG);

[0054] Figure 2 : The fNIRS-CNN model can accurately classify emotions only with functional near infrared spectroscopy signal data;

[0055] Figure 3 : Overall flow chart;

[0056] Figure 4 : Flowchart of data preprocessing method;

[0057] Figure 5 :fNIRS-CNN network structure diagram;

[0058] Figure 6 : Convolutional block structure diagram;

[0059] Figure 7 : Depth-wise convolution structure diagram;

[0060] Figure 8 : Point-by-point convolution structure diagram; DETAILED DESCRIPTION

[0062] The present invention is described in detail below in conjunction with specific embodiments.

[0063] The present invention provides an emotion classification method based on functional near infrared spectroscopy signals. The overall process is as follows: Figure 3 As shown. It includes steps S110-S150, which are as follows

[0064] S110: Collect signal data of functional near infrared spectra of different emotions, assign data labels using one-hot encoding according to different emotions, and convert category labels into numerical forms that can be processed by computers;

[0065] S120: Signal preprocessing: Modified Beer-Lambert law, filtering and baseline correction. Construct training and test datasets;

[0066] S130: Construct a functional near-infrared spectroscopy signal emotion classification model based on fNIRS-CNN.

[0067] S140: Train the constructed fNIRS-CNN.

[0068] S150: According to the functional near infrared spectroscopy signal emotion classification method, a classification result is given.

[0069] The data preprocessing method in step S120 is as follows: Figure 4 As shown, it includes steps S210-S230. Specifically, S210 modifies the Beer-Lambert law to convert the optical density The S220 uses a bandpass filter with a passband of 0.01-0.1 Hz for filtering; the S230 baseline correction solves the baseline drift problem by subtracting the average value of a reference interval from the near-infrared spectral signal.

[0070] S210: Functional near-infrared spectroscopy signals may be interfered by motion artifacts during the acquisition process. For example, the subject's head movement or other physiological noise (such as heartbeat, breathing, etc.) may affect the propagation and absorption of light. These artifacts may cause false changes in the signal. The absorption coefficient of hemoglobin also changes with the change of wavelength. The propagation path of light in the tissue is not only affected by absorption, but also experiences scattering. According to the classic Beer-Lambert law, the absorbance is proportional to the concentration of the substance, the optical path length and the absorption coefficient of the substance. However, in the actual application of functional near-infrared spectroscopy signals, since the tissue is a complex multi-layer structure and the propagation and scattering of light are also affected to a certain extent, the classic Beer-Lambert law needs to be corrected. The specific implementation process of the corrected Beer-Lambert law is as follows:

[0071] The optical density The concentration changes of two hemoglobins HbO and HbR generated by near-infrared light absorption are converted into the modified Beer-Lambert law as follows:

[0072]

[0073] in, and Wavelength The extinction coefficients of HbO and HbR at is the differential path length factor and l is the distance between the source and the detector.

[0074] S220: A bandpass filter with a passband of 0.01-0.1 Hz is used, which allows signals in the range of 0.01 Hz to 0.1 Hz to pass through, while suppressing frequency components below 0.01 Hz and above 0.1 Hz.

[0075] S230: Baseline correction solves the baseline drift problem by subtracting the average value of a reference interval from the near-infrared spectral signal, eliminating signal drift caused by instrument noise, environmental changes, sample changes, etc., to ensure more accurate analysis results. Select a flat interval that does not contain any meaningful absorption peak area, has small signal changes, and has no absorption characteristics to represent the baseline. Within the selected reference interval, calculate the average or median of the signal intensity in the interval as an estimate of the baseline drift. Remove the baseline drift from the entire spectral signal and achieve correction by subtracting the average value of the reference interval from the original spectrum. The formula for baseline correction is as follows:

[0076]

[0077] in, is the original spectral signal, is the mean value of the reference interval, is the corrected spectral signal after removing the baseline drift.

[0078] In step S130, the fNIRS-CNN model is designed to combine the two characteristics of the functional near infrared spectroscopy signal. The network structure is as follows: Figure 5 As shown in Figure 1, it includes input terminal, DHR Model module, Global Model module and output terminal. The details are as follows:

[0079] 310 At the input end, the signal data after data preprocessing is used as the input of the network.

[0080] 320 In the DHR Model module, five convolution blocks conv are included to extract the characteristics of the delayed hemodynamic response of the functional near infrared spectroscopy signal data. The convolution block structure is as follows Figure 6As shown in the figure, each convolution block consists of a 2D convolution layer, a batch normalization layer, and a Sigmoid activation layer. The specific calculation process includes: the input data passes through the first convolution (Conv) block, the second Conv block, the third Conv block, the fourth Conv block, and the fifth Conv block in turn, and the signal data in each test channel is reduced in dimension and feature extracted, and finally the extracted feature matrix is ​​obtained. Each convolution operation is based on the input data of the two-dimensional matrix X, the convolution kernel is W, is the position in the output data The value of is the input image at position The value of is the convolution kernel at position The weight of is the size of the convolution kernel. Then the formula for the convolution operation is:

[0081]

[0082] 330 The structure in the Global Model module includes a depth-wise separable convolution consisting of two convolution blocks. The first convolution (DWConv) block uses depth-wise convolution, such as Figure 7 As shown in , it is used to extract the correlation features of functional near-infrared spectral signals of different brain regions. The second convolution (PWConv) block uses point-by-point convolution, as shown in Figure 8 As shown in the figure, it is used to increase the dimension of the feature matrix after depth-by-depth convolution to obtain more features and reduce the number of parameters of the convolution block; the output data of the DHR Model module is used as the input of the Global Model module, and is sequentially convolved by DWConv and PWConv to obtain the output of this module.

[0083] Compared with standard convolution, depthwise separable convolution significantly reduces computational cost. is the kernel size, is the number of input channels, is the number of output channels, is the feature map size. The computational cost of standard convolution is defined as:

[0084]

[0085] The computational cost of depth-wise separable convolution is:

[0086]

[0087] 340 Flatten the feature matrix output by the previous module in one dimension and input it to the fully connected layer for the final classification task. In a neural network, y is the output vector, and its dimension is usually the same as the number of neurons in this layer; W is the weight matrix, with a size of m×n, where m is the output dimension and n is the input dimension; x is the input vector, with a size of n×1, corresponding to the output of the previous layer; is the bias vector, with a size of m×1, and each output neuron has a corresponding bias term. The formula of the linear layer can be expressed as:

[0088]

[0089] In step S140, the model is trained using the training set data constructed in step S120. The specific steps are as follows: first, the training set data is input into the network to calculate the result of this round of iterative calculation, the result is compared with the corresponding one-hot encoded label, and the loss value is calculated using the cross entropy loss function after smoothing the label; secondly, the gradient is calculated by the stochastic gradient descent optimizer and the weights in the network are updated by back propagation; then the above process is iterated until the error requirements are met to obtain the network training model; finally, the generalization ability of the model is verified using the test data set. The loss function in the above steps is calculated as follows:

[0090] Label smoothing is often used to prevent DNNs from overconfidence by weighted averaging of hard targets and uniform distribution over labels. The network predicts each class label k∈{1,...,k}: where zi is the logits (the i-th element of the model output output vector).

[0091]

[0092] Where yk = 1 represents the basic truth value, yk = 0 represents the remaining truth values. The cross entropy loss function is defined as:

[0093]

[0094] Label smoothing is defined as, where ε is the smoothing parameter, which is set to 0.1 by default, and uk = 1 / K for uniform distribution:

[0095]

[0096] Finally, the cross entropy loss function with label smoothing is written as:

[0097]

[0098] The following describes in detail the implementation process of the present invention in the Open Access fNIRS Dataset for Classification of Unilateral Finger- and Foot-Tapping published by Sujin Bak 2019 to study the fNIRS technology in action classification and the self-built functional near-infrared spectroscopy signal dataset of positive and negative emotions.

[0099] Example 1: Open Access fNIRS Dataset for Classification of Unilateral Finger- and Foot-Tapping Functional Near Infrared Spectroscopy Signal Classification

[0100] The specific steps are as follows:

[0101] The raw data in the Open Access fNIRS Dataset for Classification of Unilateral Finger-and Foot-Tapping dataset was divided into training set and test set data, and the data was further preprocessed. The modified Beer-Lambert law was applied to the data, and the formula is as follows:

[0102]

[0103] The data are then band-pass filtered and baseline corrected. The training data set and test data set are constructed with the preprocessed data in a ratio of 80% and 20%, respectively. The preprocessed data are intercepted using a sliding window, and the input signal data size is 40×40.

[0104] Taking the signal data intercepted by a sliding window as an example, a convolutional neural network model based on functional near-infrared spectroscopy signal classification is constructed and trained. The specific steps are as follows:

[0105] First, construct the HDR Model module. The HDR Model module is used to extract the characteristics of the delayed hemodynamic response of the functional near-infrared spectral signal data. Each convolution block consists of a 2D convolution layer, a batch normalization layer, and a Sigmoid activation layer. The specific calculation process includes: the input image passes through the first convolution (Conv) block, the second Conv block, the third Conv block, the fourth Conv block, and the fifth Conv block in sequence, and the signal data in each test channel is subjected to dimensionality reduction feature extraction, and finally the extracted feature matrix is ​​obtained.

[0106] Secondly, the Global Model module is constructed. The correlation features extracted from different brain regions in the Backbone network are fused to enhance the feature expression capability. The first convolution (DWConv) block uses depth-wise convolution to extract the correlation features of functional near-infrared spectral signals in different brain regions. The second convolution (PWConv) block uses point-wise convolution to increase the dimension of the feature matrix after depth-wise convolution to obtain more features and reduce the number of parameters of the convolution block.

[0107] Again, the feature matrix output by the previous module is flattened in one dimension and input into the fully connected layer for the final classification task.

[0108] Finally, the model is trained using the training data set. The specific steps are as follows: the signal data is input into the network to calculate the result of this round of iterative calculation, the result is compared with the label value, and the loss value is calculated using the loss function. The loss function uses the cross entropy loss after smoothing the label. At the same time, the performance of the current model is evaluated using the test data set, and the next step of parameter optimization is performed based on the evaluation results. Next, the gradient is calculated using the stochastic gradient descent optimizer and the weights in the network are updated through back propagation. The above process is iterated until the error requirements are met to obtain a network training model. The experimental results in this embodiment are shown in Table 1.

[0109] Table 1 Experimental results of fNIRS-CNN on the Open Access fNIRS Dataset for Classification of Unilateral Finger- and Foot-Tapping dataset

[0110] Accuracy Precision Recall LSTM 57.49 59.13 57.49 fNIRS-CNN 68.53 69.75 68.88

[0111] Compared with the traditional time series deep learning model LSTM, fNIRS-CNN can better adapt to functional near-infrared spectral signal data, and the accuracy, precision and recall rate in the classification results have been greatly improved.

[0112] Example 2: Self-built functional near-infrared spectroscopy signal dataset of positive and negative emotions

[0113] The data set of this embodiment is obtained by manual testing. The environment has a greater impact on data collection and there are more factors of human influence, so the preprocessing requirements for data are more stringent.

[0114] The original data in the self-built functional near-infrared spectroscopy signal data set of positive and negative emotions are divided into training set and test set data, and the data is further preprocessed. The modified Beer-Lambert law is applied to the data, and the formula is as follows:

[0115]

[0116] The data were then band-pass filtered and baseline corrected. The training data set and the test data set were constructed with the preprocessed data in a ratio of 80% and 20%, respectively. The preprocessed data were intercepted by window interception after the stimulation point, and the input signal data size was 44×660.

[0117] In this embodiment, the specific steps of constructing and training the fNIRS-CNN emotion classification model are the same as the model training steps in Example 1, and the training process is iterated until the error requirement is met to obtain a network training model. The experimental results of the fNIRS-CNN model in this embodiment are shown in Table 2.

[0118] Table 2 Experimental results of fNIRS-CNN on self-built dataset

[0119] Accuracy Precision Recall LSTM 61.20 59.8 61.5 fNIRS-CNN 72.40 71.8 62.02

[0120] Compared with the traditional time series deep learning model LSTM, fNIRS-CNN can better adapt to functional near-infrared spectral signal data, and the accuracy, precision and recall rate in the classification results have been significantly improved.

[0121] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.

Claims

1. A method for emotion classification based on functional near infrared spectroscopy signals, characterized in that: The method comprises steps S110-S150: S110: Collect signal data of functional near infrared spectra of different emotions, assign data labels using one-hot encoding according to different emotions, and convert category labels into numerical forms that can be processed by computers; S120: signal preprocessing; constructing training data set and test data set; S130: Constructing the functional near infrared spectroscopy signal emotion classification model of fNIRS-CNN; the network structure includes input end, DHR Model module, Global Model module and output end; S140: training the constructed fNIRS-CNN; S150: According to the functional near infrared spectroscopy signal emotion classification method, a classification result is given.

2. The emotion classification method according to claim 1, characterized in that: The data preprocessing method in step S120 includes steps S210-S230: S210 Modified Beer-Lambert law converts optical density The concentration changes of two hemoglobins HbO and HbR are converted into near-infrared light absorption; The S220 uses a bandpass filter with a passband of 0.01-0.1 Hz for filtering; The S230 baseline correction solves the problem of baseline drift by subtracting the average value of a reference interval from the NIR spectral signal.

3. The emotion classification method according to claim 1, characterized in that: The specific implementation process of step S210: the modified Beer-Lambert law is as follows: The optical density The concentration changes of two hemoglobins HbO and HbR generated by near-infrared light absorption are converted into the modified Beer-Lambert law as follows: ;in, and Wavelength The extinction coefficients of HbO and HbR at is the differential path length factor and l is the distance between the source and the detector.

4. The emotion classification method according to claim 1, characterized in that: The step S220: using a bandpass filter with a passband of 0.01-0.1 Hz, allowing signals in the range of 0.01 Hz to 0.1 Hz to pass through, while suppressing frequency components below 0.01 Hz and above 0.1 Hz.

5. The emotion classification method according to claim 1, characterized in that: The step S230: baseline correction solves the problem of baseline drift by subtracting the average value of a reference interval from the near-infrared spectrum signal, eliminating signal drift caused by instrument noise, environmental changes, sample changes, etc., to ensure more accurate analysis results; selects a flat interval representing the baseline that does not contain any meaningful absorption peak area, has small signal changes, and has no absorption characteristics; within the selected reference interval, calculates the average or median value of the signal intensity in the interval as an estimated value of the baseline drift; removes the baseline drift from the entire spectral signal, and corrects it by subtracting the average value of the reference interval from the original spectrum; the formula for baseline correction is as follows: ;in, is the original spectral signal, is the mean value of the reference interval, is the corrected spectral signal after removing the baseline drift.

6. The emotion classification method according to claim 1, characterized in that: The construction of the fNIRS-CNN model in step S130 includes: 310 At the input end, the signal data after data preprocessing is used as the input of the network; 320 In the DHR Model module, there are five convolution blocks conv, which are used to extract the characteristics of delayed hemodynamic response of functional near-infrared spectroscopy signal data; each convolution block consists of a 2D convolution layer, a batch normalization layer and a Sigmoid activation layer; the specific calculation process includes: the input data passes through the first convolution (Conv) block, the second Conv block, the third Conv block, the fourth Conv block, and the fifth Conv block in turn, and the signal data in each test channel is subjected to dimensionality reduction feature extraction, and finally the extracted feature matrix is ​​obtained; each convolution operation is that the input data is a two-dimensional matrix X, the convolution kernel is W, is the position in the output data The value of is the input image at position The value of is the convolution kernel at position The weight of is the size of the convolution kernel; then the formula for the convolution operation is: ; The Global Model module includes a depth-wise separable convolution consisting of two convolution blocks; the first convolution (DWConv) block uses depth-wise convolution to extract the correlation features of functional near-infrared spectral signals in different brain regions; the second convolution (PWConv) block uses point-wise convolution to increase the dimension of the feature matrix after depth-wise convolution to obtain more features and reduce the number of parameters of the convolution block; the output data of the DHR Model module is used as the input of the Global Model module, and is sequentially convolved by DWConv and PWConv to obtain the output of this module; 340 Flatten the feature matrix output by each module in one dimension and input it to the fully connected layer for the final classification task; In the neural network, y is the output vector, and its dimension is usually the same as the number of neurons in this layer; W is the weight matrix, with a size of m×n, where m is the output dimension and n is the input dimension; x is the input vector, with a size of n×1, corresponding to the output of the previous layer; is the bias vector, with a size of m×1, and each output neuron has a corresponding bias term; The formula of the linear layer can be expressed as: 。 7. The emotion classification method according to claim 1, characterized in that: In step S140, the model is trained using the training set data constructed in step S120; the specific steps are as follows: first, the training set data is input into the network to calculate the result of this round of iterative calculation, the result is compared with the label assigned by the corresponding one-hot encoding, and the loss value is calculated using the cross entropy loss function after smoothing the label; secondly, the gradient is calculated by the stochastic gradient descent optimizer and the weights in the network are updated by back propagation; then the above process is iterated until the error requirement is met to obtain the network training model; finally, the generalization ability of the model is verified using the test data set; the loss function in the above steps is calculated as follows: Label smoothing is often used to prevent DNN overconfidence by weighted averaging of hard targets and uniform distribution over labels; the network predicts each class label k∈{1,…,k}: where zi is the i-th element of the logits model output vector; ; where yk = 1 represents the basic truth value, and yk = 0 represents the remaining truth values. The cross entropy loss function is defined as: ; Label smoothing is defined as, where ε is the smoothing parameter, which is set to 0.1 by default, and uk = 1 / K is a uniform distribution: ; Finally, the cross entropy loss function with label smoothing is written as: 。