Method for identifying sentiment of old people based on ViT network and recurrence plot

By using the Vision Transformer network and recursive graph in the emotional recognition of the elderly, using the electrocardiogram signal for emotion recognition, the problem of unsatisfactory recognition of the recognition rate in the prior art is solved, and high-accuracy emotional recognition of the elderly is achieved.

CN120180067APending Publication Date: 2025-06-20NORTHEAST FORESTRY UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311744713.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and quickly identify the emotional state of the elderly, especially because the EEG signal is weak and susceptible to noise interference, resulting in the unsatisfactory accuracy of emotional recognition.

Method used

The EKS emotion recognition method based on Vision Transformer network and recursive graph is adopted, and the variational modal decomposition is optimized through the dung beetle algorithm, the ECG signal is converted into two-dimensional image data, and the Vision Transformer network is input for identification.

Benefits of technology

The high accuracy of emotional recognition in elderly people is achieved, with a test set accuracy of 97.20%, which is better than traditional machine learning models and other deep learning methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180067A_ABST
    Figure CN120180067A_ABST
Patent Text Reader

Abstract

The invention discloses a method for recognizing sentiment of old people based on a ViT network and a recurrence plot, and belongs to the technical field of sentiment recognition. The method specifically comprises the following steps: acquiring original electrocardiosignals of old people under stimulation of different emotion videos; the collected original electrocardiosignals are preprocessed; variational mode decomposition is optimized by adopting a dung beetle optimization algorithm, and the preprocessed electrocardiosignals are decomposed; performing arc tangent standardization on the decomposed electrocardiosignal data sample; performing coding processing on the standardized signal sample by using a recurrence plot method, and converting the signal sample into two-dimensional image data; labeling the two-dimensional image data, and dividing the two-dimensional image data into a training set, a verification set and a test set; according to the method, transfer learning is adopted, a pre-training weight of a Vision Transform on an ImageNet-22k data set is loaded, two-dimensional image data is trained, the minimum loss of a verification set is taken as a benchmark, the training weight at the moment is obtained and stored in a target model, model verification and comparison are carried out through a test set, and a result shows that a Vision Transform network is combined with a recurrence plot, so that the accuracy of the model verification is improved. The accuracy rate of emotion recognition through the electrocardiosignals reaches up to 97.20%, and compared with an electroencephalogram recognition method and a traditional machine recognition method, the electrocardiosignal recognition method based on the Vision Transform network and the recurrence plot has the obvious advantages that signal collection is convenient and fast, and recognition is accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of emotion recognition, and particularly relates to a method for recognizing the emotional state of electrocardiogram signals of the elderly based on the Vision Transformer network and recursive graphs. Background Art

[0002] Emotion recognition is a means of recognizing a person's emotional state by analyzing physiological signals, facial expressions, and other information, and it has a wide range of applications in the fields of human-computer interaction, intelligent education, intelligent medicine, etc.

[0003] In China, the proportion of the elderly population in the total population has been increasing year by year. Research shows that when the elderly are immersed in a negative psychological state caused by emotions, the adrenaline and corticosteroids secreted by their bodies will be more than those of the elderly in a non-negative emotional state. This will cause the physical functions of the elderly in a negative emotional state to degenerate faster and the aging speed to increase. When the elderly are in a very excited state, their blood flow rate will accelerate, thus increasing the risk of cardiovascular diseases. It is obvious that the risk of illness caused by the imbalance of the emotional state of the elderly is a common problem that thousands of families have to face today. However, emotional problems are often difficult to detect, and there is little research on emotion recognition for the elderly population. Therefore, in the face of the social problem of aging, how to accurately and quickly recognize the emotional state of the elderly has become an urgent problem to be solved.

[0004] Currently, most research uses electroencephalogram signals for emotion recognition. However, considering factors such as the complex process of wearing electroencephalogram devices and the weak electroencephalogram signals being easily interfered by noise, the present invention proposes a method for recognition using electrocardiogram signals.

[0005] Traditional electrocardiogram signal emotion recognition methods mainly rely on machine learning algorithms to extract features from electrocardiogram signal data and then input them into machine learning models for recognition, such as Support Vector Machine (SVM), K Nearest Neighbors (KNN), etc. However, machine learning algorithms are difficult to effectively reflect the high-dimensional and non-linear relationships of electrocardiogram signal data. Therefore, the accuracy of emotion recognition is not ideal.

[0006] Deep learning proposed by HINTON et al. has powerful representation capabilities due to its large and flexible model structure, enabling it to process large-scale electrocardiogram (ECG) signal data. The deep distributed representation of deep learning gives it good generalization ability, allowing it to capture the inherent characteristics that complex ECG signal data may share in different environments. Commonly used deep learning methods include convolutional neural networks (CNNs), auto-encoders (AEs), Transformers, etc. Among them, vision models based on Transformers have been widely applied in other computer vision tasks such as object detection, image segmentation, image generation, and image annotation. Therefore, the present invention proposes a method for recognizing the emotions of the elderly using a Vision Transformer network combined with a recurrence plot, aiming to provide new ideas for the emotion recognition of the elderly. Summary of the Invention

[0007] The invention designs an ECG signal emotion recognition method based on a Vision Transformer network and a recurrence plot for the elderly population. Relying on the powerful representation capabilities and good generalization ability of the Transformer, it can process large-scale ECG signal data and capture the inherent characteristics that complex ECG signal data may share in different environments. The variational mode decomposition is optimized using the dung beetle algorithm, and then the decomposed signals are converted into two-dimensional image data using the recurrence plot image conversion method, and then input into the Vision Transformer network to achieve the purpose of quickly and accurately recognizing the emotions of the elderly.

[0008] By comparing the emotion recognition accuracy rates of ECG signals on different models, the present invention gives relatively ideal suggestions on which model to use for ECG signal emotion recognition of the elderly, aiming to provide ideas for the emotion recognition of the elderly. The technical solution of the present invention is divided into the following four stages:

[0009] (1) The experimental instruments include the ErgoLAB PPG wireless pulse sensor and the ErgoLAB human-machine-environment synchronization platform V3.0. The above instruments are all provided by Beijing Jinfa Technology Co., Ltd.

[0010] Based on the method of video-induced emotions, the original ECG signals are collected using a pulse sensor respectively. The collected ECG signals are input into ErgoLAB to eliminate noise and redundant signals;

[0011] (2) The dung beetle optimization algorithm is used to optimize the number of modes \(k\) and the penalty factor \(\alpha\) of variational mode decomposition. The value range of the number of modes \(k\) is set in \([3, 12]\), and the value range of the penalty factor \(\alpha\) is set in \([500, 3000]\). The fitness function is selected. After 18 iterations of the electrocardiogram signal, the fitness values converge to the minimum, indicating that the optimal state has been reached. The number of modes decomposed by the optimized electrocardiogram signal is 8 respectively.

[0012] Save the electrocardiogram signal data after decomposition. The length of each sample is 1920, and the sampling step is 20. Each sample is normalized by the arctangent as shown in Equation (1).

[0013]

[0014] In the formula: \(z_i\) is the normalized data, and \(\theta\) is the adjustment parameter;

[0015] (3) The standardized signal samples are encoded by a recurrence plot to generate two-dimensional image data. 100 samples of each video clip under each type of electrocardiogram signal are taken, and the two-dimensional image data is labeled according to the ratio of 6:2:2, and divided into a training set, a validation set, and a test set;

[0016] (4) To accelerate the network training speed, transfer learning is adopted, and the pre-trained weight file trained on the ImageNet-22k dataset is loaded for training; the number of iterations is set to 30; the cross-entropy loss function commonly used for multi-classification is selected as the loss function; to improve the efficiency of memory alignment and matrix multiplication, the training batch is set to 8; the initial learning rate is set to 0.001, and the learning rate decay is carried out with the cosine annealing strategy, which can not only make the loss in the early training process drop rapidly, but also ensure the accurate optimization of the loss function in the later stage; the gradient descent optimization algorithm, which is relatively mainstream, is selected as the optimization method; the number of modes \(k\) is taken as 5, and the penalty factor \(\alpha\) is taken as 2000; the network that uses the dung beetle optimization algorithm to optimize variational mode decomposition is denoted as DVVision Transformer; the two-dimensional image data is input into the Vision Transformer network for training, and the training weight with the lowest validation set loss during the iteration process is saved. The training weight is loaded and the two-dimensional image data of the test set is input. The results show that the accuracy of the test set of the electrocardiogram signal reaches 97.20% respectively, which can meet the requirements of elderly emotion recognition.

[0017] Beneficial effects

[0018] (1) The present invention uses the dung beetle algorithm to optimize variational mode decomposition, which can effectively obtain feature information. Especially compared with the validation set under Base, the accuracy is improved by more than 5%.

[0019] (2) The present invention uses a recurrence plot to transform signal data into a two-dimensional image, extracts rich non-linear information in the original signal, and can achieve a high accuracy by utilizing the characteristics of Transformer being sensitive to image data.

[0020] (3) The recognition accuracy of the method proposed by the present invention for electrocardiogram signals is 97.20%.

[0021] (4) In terms of the recognition accuracy of electrocardiogram signals, the method proposed by the present invention has at least increased by 9.4%, 11.13%, and 12.61% respectively compared with SVM, NB, and KNN; compared with ResNet34, EfficientNet-B0, and VGG16, it has at least increased by 1.14%, 0.54%, and 3.34% respectively. Description of the Drawings

[0022] Figure 1 Flow chart of the present invention

[0023] Figure 2 Wearing effect diagram of PPG wireless pulse sensor

[0024] Figure 3 Schematic diagram of electrocardiogram signal after transformation by recurrence plot

[0025] Figure 4 Vision Transformer network architecture diagram

[0026] Figure 5 Position encoding process diagram

[0027] Figure 6 Overall architecture diagram of Encoder Block

[0028] Figure 7 Schematic diagram of multi-head attention mechanism

[0029] Figure 8 Confusion matrix of electrocardiogram signal test

[0030] Best Mode for Carrying Out the Invention

[0031] The present invention will be further described in detail below with reference to the drawings and examples.

[0032] A method for recognizing the emotions of the elderly based on the ViT network and recurrence plot provided by the present invention includes:

[0033] S1. Obtain the original electrocardiogram signals of the elderly population under different emotional video stimuli through experiments, and preprocess the collected original electrocardiogram signals;

[0034] S2. Optimize variational mode decomposition using the dung beetle optimization algorithm, decompose the preprocessed electrocardiogram (ECG) signal, and perform arctangent normalization on the decomposed ECG signal data samples.

[0035] S3. Encode the standardized signal samples using the recurrence plot method, convert them into two-dimensional image data, label the two-dimensional image data, and divide it into a training set, a validation set, and a test set.

[0036] S4. Adopt transfer learning, load the pre-trained weights of Vision Transformer on the ImageNet-22k dataset, train the two-dimensional image data, take the lowest loss of the validation set as the benchmark, obtain the training weights at this time, save them to the target model, and perform model validation and comparison through the test set.

[0037] In step S1 of this embodiment, the original ECG signals of the elderly population under different emotional video stimuli are obtained through experiments; the preprocessing of the collected original ECG signals includes:

[0038] The experimental instruments selected in the present invention include an ErgoLAB PPG wireless pulse sensor for collecting ECG signals and an ErgoLAB human-machine environment synchronization platform V3.0. The above instruments are all provided by Beijing Jinfa Technology Co., Ltd.

[0039] The experiment uses film clips of various themes to induce the emotional states of the subjects. By referring to the production processes of the Deap dataset (Database for Emotion Analysis Using Physiological Signals) and the Seed dataset (SJTUEmotion EEG Dataset), six appropriate film clips are selected. The emotions expected to be induced in the subjects include but are not limited to happiness, sadness, fear, disgust, etc. When delimiting the labels, the present invention uses the dimensional theory for emotion classification. In research, valence and arousal are commonly used to quantify human emotions. Valence is used to reflect the positive and negative degrees of emotions; arousal reflects the excitement degree of a person in a certain state. According to the emotion dimension theory, the dimension scoring range is 1-9, and the scoring result higher than 5 is classified into the high-level group, otherwise it is the low-level group. The six film clips respectively induce four emotional states of high arousal and high valence, high arousal and low valence, low arousal and high valence, and low arousal and low valence, and label them with 0, 1, 2, and 3 respectively. Based on the method of inducing emotions by video, a pulse sensor is used to collect ECG signals. The collected ECG signals are input into ErgoLAB to eliminate noise and redundant signals for subsequent emotion classification.

[0040] In step S2 of this embodiment, the variational mode decomposition is optimized by using the dung beetle optimization algorithm, and the preprocessed electrocardiogram signal is decomposed. The arctangent normalization of the decomposed electrocardiogram signal data samples includes:

[0041] Variational Mode Decomposition is a signal decomposition method based on variational Bayesian theory. It can decompose complex multi-component signals into a series of intrinsic mode functions with narrow bandwidths. The advantages of variational mode decomposition are that it can effectively decompose non-stationary signals, is robust to noise, can automatically determine the number of decompositions, and can control the sparsity and smoothness of the decomposition results.

[0042] In order to make the collected electrocardiogram signal better reflect its internal characteristics, the present invention uses variational mode decomposition to decompose the collected electrocardiogram signal data. In the variational mode decomposition algorithm, the mode number k and the penalty factor α are two relatively important hyperparameters. If the value of k is too small, the original signal feature information will be lost. Conversely, it will lead to frequency aliasing; α determines the bandwidth of each intrinsic mode function after variational mode decomposition. Therefore, the dung beetle optimization algorithm (Dung Beetle Optimizer) is used to optimize the mode number k and the penalty factor α of the variational mode decomposition. The value range of the mode number k is set in [3, 12], and the value range of the penalty factor α is set in [500, 3000]. The fitness function is selected as H(x)=-∑[p(x i )log2(p(x i ))]. After the electrocardiogram signal iterates 18 times, the fitness value converges to the minimum, indicating that the optimal state has been reached at this time. The decomposed mode numbers after optimization are 8 respectively. Save the decomposed electrocardiogram signal data. Each sample length is 1920, and the sampling step is 20. The arctangent normalization is performed on each sample as shown in Equation (1).

[0043]

[0044] In the formula: zi is the normalized data, and θ is the adjustment parameter.

[0045] In step S3 of this embodiment, the normalized signal samples are encoded by the recurrence plot method, converted into two-dimensional image data, and labels are assigned to the two-dimensional image data. It is divided into a training set, a validation set, and a test set, including:

[0046] The recurrence plot is a commonly used processing method in the field of nonlinear signal research and has been widely applied in the fields of medicine and mechanical flaw detection. Using the idea of phase space reconstruction, the recurrence plot reconstructs time series data into a high-dimensional phase space and then presents it in the form of a two-dimensional image, which can reveal the internal information of the time series data and provide relevant prior knowledge of predictability and similarity for analyzing the non-stationarity, chaos, and periodicity of the time series data.

[0047] For the given time series data X = [x1, x2, … x i , … x n , perform phase space reconstruction:

[0048] X i = {x i , x i+τ ,..., x i+(m-1)τ} (2)

[0049] where: i = 1, 2, …, N - (m - 1) i , and τ is the delay time.

[0050] Calculate the distance between two points after phase space reconstruction:

[0051] d ij = ||X i - X j || (3)

[0052] where: X i and X j are the points after reconstructing the original signal, and ||·|| is the norm operation.

[0053] Subtract the selected threshold ε from d ij , and input the result into the Heaviside function to obtain the value of each point on the recurrence matrix. The specific formula is:

[0054]

[0055] where: R ij can take two values, which are 1 and 0 respectively. When the value is 1, it indicates that the distance between two points in the reconstructed phase space is less than the threshold, indicating that the two points have a recurrence relationship; when the value is 0, it indicates that the distance between two points in the reconstructed phase space is greater than the threshold, indicating that the two points do not have a recurrence relationship. Therefore, after the original time series data is processed by recurrence, it will be converted into a recurrence plot composed of points of different colors.

[0056] The standardized signal samples are encoded using a recurrence plot to generate two-dimensional image data. 100 samples are taken from each film segment under each type of electrocardiogram signal, and the two-dimensional image data is labeled according to a ratio of 6:2:2 and divided into a training set, a validation set, and a test set. The schematic diagram of the electrocardiogram signal after conversion by the recurrence plot is as shown in Figure 2 shown.

[0057] In step S4 of this embodiment, transfer learning is adopted. The pre-trained weights of Vision Transformer on the ImageNet-22k dataset are loaded to train the two-dimensional image data. Based on the lowest loss of the validation set, the training weights at this time are obtained and saved to the target model. Model validation and comparison through the test set include:

[0058] Transformer is different from previous RNN and CNN networks. It does not have any recurrence and convolution structures in its architecture but is completely based on the self-attention mechanism. And it has the most advanced research results in the fields of natural language processing, computer vision, and multi-modal. It is mainly composed of an encoder and a decoder. The encoder is similar to a convolutional layer and is used to extract the features of the input data, while the decoder is responsible for converting the extracted features into output results. In the Vision Transformer network, only the encoder part is adopted, and position encoding is performed before the encoder, and the decoder part is replaced by a fully connected layer. The architecture of the Vision Transformer network is as shown in Figure 3 shown.

[0059] For a standard Transformer network, the input data must be a sequence of vectors (tokens), that is, a two-dimensional matrix [num_token, token_dim]. Taking Vision Transformer-B / 16 as an example, the length of each token is 768. For image data, its size is [H, W, C]. Therefore, it is necessary to convert the data into a two-dimensional matrix through the Embedding layer. First, the input image [H×W] is divided into patches of size 16×16 to obtain HW / 256 patches, and then each patch is mapped to a token through a linear mapping to obtain a token with a length of 768, thus obtaining a two-dimensional matrix of size [HW / 256, 768]. Finally, it is concatenated with a trainable class token of size [1, 768] to obtain a two-dimensional matrix of size [HW / 256 + 1, 768] as the input data. The position encoding process is as shown in Figure 4 shown.

[0060] The Transformer Encoder extracts features by repeatedly stacking the Encoder Block L (repeated 12 times in VisionTransformer - B / 16) times, as Figure 5 shown. The input data is first normalized through the Layer Norm layer, then passed through the Multi - Head Attention layer so that the model can focus on information in different positions and different sub - spaces and allocate appropriate weights. Then, through the DropPath layer, some redundant information is discarded to prevent overfitting. The output obtained at this time is shortcut - connected to the original input as the input of the next Layer Norm. The result of the re - normalized data is passed through the MLP Block to enhance the model's expressive ability. The MLP Block layer consists of two linear layers, a Dropout layer, and a GELU activation function. Then, after passing through the DropPath again, the obtained result is shortcut - connected to the input again. This is the overall architecture of one Encoder Block.

[0061] The Multi - Head Attention layer, as the key architecture of Vision Transformer, is similar to the self - attention mechanism. The self - attention mechanism includes the following steps:

[0062] For sequence data with an output length of i, there are corresponding i input nodes x1, x2,..., x i , and then through f(x), the input is mapped to a1, a2,..., a i , and then respectively dot - product with three trainable transformation matrices W q , W k , W v to obtain the corresponding q i , k i , v i As shown in Equation (5):

[0063] q i = a i W q , k i = a i W k , v i = a i W v (5)

[0064] Among them, q i is the query vector, which matches the corresponding k i , and v i represents the information extracted from a i , qi The matching with k i is to calculate the correlation between the two. The greater the correlation, the greater the weight of the corresponding v i . In the Vision Transformer network, the correlation weight is calculated in the form of scaled dot product, as shown in Equation (6):

[0065]

[0066] where d k represents the vector length. The input matrix size is appropriately adjusted by scaling to avoid too small gradient after subsequent normalization, which affects network training. Then, weight(q t , k t ) is normalized using the softmax function. In summary, the self-attention mechanism can be represented by matrix multiplication, as shown in Equation (7):

[0067]

[0068] Multi-Head Attention is similar to Self-Attention. a i is respectively passed through W q , W k , W v to obtain the corresponding q i , k i , v i . According to the number of heads h used, q i , k i , v i are all divided into h parts. For each head, the method in Equation (7) of Self-Attention is used. The h heads are concatenated by concat, and matrix multiplication is performed with the trainable W O matrix to obtain the final output result. The calculation process is as shown in Equation (8):

[0069]

[0070] The shape of the token output after passing through the Transformer Encoder remains unchanged. The class token at this time is extracted, and then the final classification result is obtained through the MLP Head. The MLP Head consists of Linear + tanh activation function + Linear.

[0071] In order to verify the effectiveness of the dung beetle optimization algorithm and variational mode decomposition, the present invention uses the training set and verification set of the ECG signal as input samples and uses ablation experiments for analysis. In order to speed up the network training speed, transfer learning is used to load the pre-trained weight file trained on the ImageNet-22k dataset for training; the number of iterations is set to 30; the loss function selects the cross entropy loss function commonly used in multi-classification; in order to improve the efficiency of memory alignment and matrix multiplication, the training batch needs to be set to the power of 2, and the training batch in this study is set to 8; a larger learning rate will make the training process difficult to converge, on the contrary, a smaller learning rate will cause the network to fall into a local optimum, therefore, the initial learning rate of the present invention is set to 0.001, and the learning rate is decayed by the cosine annealing strategy, which can not only make the loss of the early training process drop rapidly, but also ensure the accurate optimization of the later loss function; the optimization method selects the more mainstream gradient descent optimization algorithm. The network using recursive graph and Vision Transformer is taken as the base network, and the network using only recursive graph and Vision Transformer is denoted as Base; the network using variational mode decomposition is denoted as VBase, where the number of modes k is 5 and the penalty factor α is 2000; the network using the dung beetle optimization algorithm to optimize variational mode decomposition is denoted as DVBase. The specific model parameter settings are shown in Table 1.

[0072] Table 1 Model parameter settings

[0073]

[0074] As the number of iterations increases, the loss of ECG signals shows a downward trend and converges within 30 epochs. During training, when the input data is ECG signals, the loss of the training set and validation set under DVBase is the lowest compared with the loss of the training set and validation set under VBase and Base, and the loss at the 30th epoch is only 0.042 and 0.003 respectively; when the input data is ECG signals, except for DVBase, the loss of VBase and Base increases and fluctuates to varying degrees, and the loss of the training set and validation set under Base reaches more than 0.3. This is because the skin of the elderly is loose and the extracted original signal has more interference, so the network training effect is more dependent on the extraction of feature signals. The training effect of the DVBase network is better than that of VBase and Base, and the superiority of the dung beetle optimization algorithm and variational mode decomposition is obvious.

[0075] To verify the superiority of the Vision Transformer network in recognizing the emotions of the elderly, the network weight parameters at the lowest validation set loss in 30 epochs of the electrocardiogram signal were saved, and the test set data was imported to evaluate the training effect. The accuracy rate of the electrocardiogram signal was as high as 97.20%, and the confusion matrix was as shown in Figure 2 . It shows that after the electrocardiogram signal is optimized by the dung beetle algorithm for variational mode decomposition, the decomposed signal is converted by the recurrence plot and input into the Vision Transformer network, and this method is effective.

[0076] The image modality conversion method has a great influence on the recognition accuracy. After comparison, the recognition accuracy of the signal data converted by the recurrence plot is better than the other three image modality conversion methods, and the emotion recognition accuracy of the electrocardiogram signal reaches 97.20%; while the recognition accuracy of the signal data converted by the Markov transition field is much lower than this, indicating that this method is not suitable for the emotion recognition of the elderly.

[0077] The emotion recognition model based on Vision Transformer adopted in the present invention has a higher accuracy rate than the traditional machine learning model. Analyzing the reasons, the traditional machine learning model cannot well reflect the non-linear characteristics between electrocardiogram signal data, while deep learning can achieve end-to-end mapping, which helps to solve non-linear problems. Considering that the Transformer architecture has the characteristic of self-attention mechanism, the relative position information between elements in the input sequence can be retained through position encoding, so it has advantages in processing time series data compared with other CNN architectures, which is also one of the reasons for selecting the Vision Transformer network in this study.

Claims

1. An electrocardiogram signal emotion recognition method for the elderly, based on the Vision Transformer network and recursive graph, characterized in that Including: S1. Obtain the original electrocardiogram signals of the elderly population under different emotional video stimuli through experiments, and preprocess the collected original electrocardiogram signals; S2. Optimize the variational mode decomposition using the dung beetle optimization algorithm, decompose the preprocessed electrocardiogram signals, and perform arctangent normalization on the decomposed electrocardiogram signal data samples; S3. Encode and process the standardized signal samples using the recurrence plot method, convert them into two-dimensional image data, label the two-dimensional image data, and divide it into a training set, a validation set, and a test set; S4. Adopt transfer learning, load the pre-trained weights of Vision Transformer on the ImageNet-22k dataset, train the two-dimensional image data, take the lowest loss of the validation set as the benchmark, obtain the training weights at this time, and save them to the target model, and perform model verification and comparison through the test set.

2. The emotion recognition method based on the Vision Transformer network and recursive graph according to claim 1, characterized in that Obtain the original electrocardiogram signals of the elderly population under different emotional video stimuli, and preprocess the collected original electrocardiogram signals: The experimental instruments include the ErgoLAB PPG wireless pulse sensor and the ErgoLAB human-machine-environment synchronization platform V3.

0. The above instruments are all provided by Beijing Jinfa Technology Co., Ltd. Based on the method of video-induced emotion, use the pulse sensor to collect the original electrocardiogram signals, and input the collected electrocardiogram signals into ErgoLAB to eliminate noise and redundant signals.

3. The emotion recognition method based on the Vision Transformer network and recursive graph according to claim 1, characterized in that Optimize the variational mode decomposition using the dung beetle optimization algorithm, and decompose the preprocessed electrocardiogram signals; perform arctangent normalization on the decomposed electrocardiogram signal data samples: The dung beetle optimization algorithm is used to optimize the number of modes \(k\) and the penalty factor \(\alpha\) of variational mode decomposition. The value range of the number of modes \(k\) is set in \([3, 12]\), and the value range of the penalty factor \(\alpha\) is set in \([500, 3000]\). The fitness function is selected as \(H(x)=-\sum[p(x i )\log_2(p(x i ))]\). After 18 iterations of the electrocardiogram signal, the fitness value converges to the minimum, indicating that the optimal state has been reached. The number of modes decomposed by the optimized electrocardiogram signal is 3, 4, and 8 respectively. The electrocardiogram signal data after decomposition are saved respectively. Each sample length is 1920, and the sampling step is 20. Arc tangent normalization is performed on each sample as shown in Equation (1): In the formula: zi is the standardized data, and θ is the adjustment parameter.

4. The emotion recognition method based on the Vision Transformer network and recursive graph according to claim 1, characterized in that Encode and process the standardized signal samples using the recurrence plot method to convert them into two-dimensional image data; label the two-dimensional image data and divide it into a training set, a validation set, and a test set: Encode and process the standardized signal samples using the recurrence plot to generate two-dimensional image data. Take 100 samples of each film segment under each electrocardiogram signal, and label the two-dimensional image data according to the ratio of 6:2:2, and divide it into a training set, a validation set, and a test set.

5. The emotion recognition method based on the Vision Transformer network and recursive graph according to claim 1, characterized in that Adopt transfer learning, load the pre-trained weights of Vision Transformer on the ImageNet-22k dataset, train the two-dimensional image data, take the lowest loss of the validation set as the benchmark, obtain the training weights at this time, and save them to the target model, and perform model verification and comparison through the test set: To accelerate the network training speed, transfer learning is adopted, and the pre-trained weight file trained on the ImageNet-22k dataset is loaded for training; the number of iterations is set to 30; the cross-entropy loss function commonly used for multi-classification is selected as the loss function; to improve the efficiency of memory alignment and matrix multiplication, the training batch is set to 8; the initial learning rate is set to 0.001, and the learning rate decay is carried out with the cosine annealing strategy, which can not only make the loss in the early training process drop rapidly, but also ensure the accurate optimization of the loss function in the later stage; the optimization method selects the relatively mainstream gradient descent optimization algorithm; the number of modes k is taken as 5, and the penalty factor α is taken as 2000; the dung beetle optimization algorithm is used to optimize the network of variational mode decomposition, denoted as DVVision Transformer; the two-dimensional image data is input into the Vision Transformer network for training, the training weights of the iteration with the lowest validation set loss during the iteration process are saved, the training weights are loaded and the two-dimensional image data of the test set is input. The results show that the test set accuracy of the electrocardiogram signal reaches 97.20%, which can meet the requirements of elderly emotion recognition.