Emotional feature recognition method and system based on electrocardiosignals
By integrating the time-domain, frequency-domain, and nonlinear features of electrocardiogram signals and combining them with a deep learning model, the problem of distinguishing between anxiety and depression in traditional methods has been solved, thereby improving the accuracy and sensitivity of emotion recognition and supporting mental health monitoring in universities.
Patent Information
- Application Number
- CN202511581294.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-10-31
AI Technical Summary
In the monitoring of mental health in colleges and universities, existing technologies and traditional methods are unable to effectively distinguish between anxiety and depression with similar physiological manifestations, leading to misjudgments, failure to accurately capture non-linear characteristics, and affecting the counseling strategies of mental health teachers.
By fusing the time-domain, frequency-domain, and nonlinear features of electrocardiogram signals, a structured emotion feature matrix is constructed. Then, the attention mechanism and recurrent neural network in the deep learning model are used to adaptively enhance key features and generate emotion recognition results.
It achieves comprehensive coverage and dynamic optimization of the emotional characteristics of electrocardiogram signals, improves the accuracy and sensitivity of emotion recognition, reduces noise interference, and provides more reliable emotion monitoring support.
Smart Images

Figure CN121528518A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of signal processing technology, and in particular to a method and system for identifying emotion features based on electrocardiogram signals. Background Technology
[0002] In college mental health monitoring, students wear portable electrocardiogram (ECG) devices to collect signals in real time to identify emotions such as calmness, depression, and anxiety, assisting mental health teachers in timely intervention. Emotional changes will simultaneously trigger changes in the time domain, frequency domain, and nonlinear characteristics of ECG signals. However, traditional techniques mostly focus on time domain features such as heart rate and RR interval standard deviation, or only obtain low-frequency and high-frequency features through power spectrum analysis. They do not pay enough attention to nonlinear features that reflect the complex dynamic laws of signals. Furthermore, single-dimensional features are difficult to distinguish between different emotions with similar physiological manifestations, which can easily lead to feature confusion.
[0003] Anxiety caused by exam pressure and depression caused by interpersonal relationship problems in students can both manifest as changes in time and frequency domain characteristics, such as increased heart rate and decreased HF component proportion. Traditional methods that rely solely on these characteristics are prone to misdiagnosis. In fact, depression enhances the regularity of electrocardiogram signals and significantly reduces sample entropy, while anxiety leads to increased signal chaos and increased sample entropy. Traditional methods, because they do not extract these nonlinear features, cannot accurately capture the differences between the two, which may lead school counselors to misdiagnose depression as anxiety and adopt guidance strategies that are not appropriate for the students. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method and system for emotion feature recognition based on electrocardiogram (ECG) signals, improve the completeness of the extraction of emotion-related features from ECG signals, and provide more reliable support for emotion monitoring and recognition applications.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, an emotion feature recognition method based on electrocardiogram (ECG) signals, the method comprising: Step 1: Acquire raw electrocardiogram (ECG) signals and preprocess them to obtain preprocessed ECG signals; Step 2: Based on the preprocessed electrocardiogram signal, extract time-domain features, frequency-domain features, and nonlinear features related to emotional state to form an initial feature set; Step 3: Fuse the multiple features in the initial feature set to construct a structured emotion feature matrix that integrates spatial distribution and multi-dimensional feature information; Step 4: Perform feature structure analysis on the structured emotion feature matrix. By defining a reference feature vector and calculating the distribution relationship between multiple associated feature vectors and the reference feature vector, a feature association region is constructed. Within the feature association region, two feature reference points with temporal relationship are selected respectively, and a feature adjustment path is formed based on the temporal connection relationship between the two feature reference points. A feature association metric is generated based on the feature adjustment path, and the feature association metric is used to adjust the structured emotion feature matrix to obtain the adjusted structured emotion feature matrix. Step 5: Input the adjusted structured emotion feature matrix into the deep learning model. The attention mechanism in the deep learning model adaptively enhances the key features related to emotion recognition in the structured emotion feature matrix, generating an enhanced emotion feature representation. Step 6: Based on the enhanced emotion feature representation, the time dependence of the electrocardiogram signal is obtained through the recurrent neural network in the deep learning model to generate the final emotion recognition result.
[0006] Secondly, an emotion feature recognition system based on electrocardiogram signals includes: The acquisition and processing module is used to acquire raw electrocardiogram (ECG) signals and preprocess them to obtain preprocessed ECG signals. The feature extraction module is used to extract time-domain features, frequency-domain features, and nonlinear features related to emotional state based on preprocessed electrocardiogram signals, forming an initial feature set. The feature construction module is used to fuse multiple types of features in the initial feature set to construct a structured sentiment feature matrix that integrates spatial distribution and multi-dimensional feature information; The feature adjustment module is used to perform feature structure analysis on the structured emotion feature matrix. It constructs feature association regions by defining reference feature vectors and calculating the distribution relationship between multiple associated feature vectors and the reference feature vectors. Within the feature association regions, two feature reference points with temporal relationships are selected, and feature adjustment paths are formed based on the temporal connection relationship between the two feature reference points. Based on the feature adjustment paths, feature correlation metrics are generated, and the feature correlation metrics are used to adjust the structured emotion feature matrix to obtain the adjusted structured emotion feature matrix. The feature enhancement module is used to input the adjusted structured emotion feature matrix into the deep learning model, and adaptively enhance the key features related to emotion recognition in the structured emotion feature matrix through the attention mechanism in the deep learning model, thereby generating an enhanced emotion feature representation. The result generation module is used to generate the final emotion recognition result by obtaining the time dependence of electrocardiogram signals through a recurrent neural network in a deep learning model based on the enhanced emotion feature representation.
[0007] Thirdly, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0008] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0009] The above-described solution of the present invention has at least the following beneficial effects: By integrating time-domain, frequency-domain, and nonlinear features, the model comprehensively covers the dimensions of physiological signal changes related to emotions, providing a complete feature foundation for accurate recognition. Simultaneously, through feature association region construction, temporal feature adjustment path generation, and correlation metric adjustment, the model achieves dynamic feature optimization, adaptively selecting effective features and suppressing noise interference, making the feature matrix more aligned with recognition needs. Furthermore, the combination of attention mechanisms and recurrent neural networks in the deep learning model not only strengthens the focus on key emotional features but also accurately captures the temporal dependence of electrocardiogram signals, enhancing the deep learning model's sensitivity to emotional changes.
[0010] ECG signals can be conveniently collected through wearable devices such as smart bracelets and ECG patches. Combined with their optimization of signal processing and deep learning model inference processes, they have the characteristics of low latency and high adaptability, and can be widely used in diverse scenarios such as mental health monitoring, intelligent cockpit emotion intervention, human-computer interaction, and educational emotion feedback, with broad prospects for industrialization. Attached Figure Description
[0011] Figure 1 This is a flowchart illustrating the emotion feature recognition method based on electrocardiogram signals provided in an embodiment of the present invention.
[0012] Figure 2 This is a schematic diagram of an emotion feature recognition system based on electrocardiogram signals provided in an embodiment of the present invention.
[0013] Figure 3 This is a schematic diagram of the feature structure analysis of the structured emotion feature matrix provided by the embodiments of the present invention. It is constructed by defining a reference feature vector and calculating the distribution relationship between multiple associated feature vectors and the reference feature vector to build a feature association region. Detailed Implementation
[0014] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0015] like Figure 1 As shown, embodiments of the present invention propose a method for emotion feature recognition based on electrocardiogram signals, the method comprising the following steps: Step 1: Acquire raw electrocardiogram (ECG) signals and preprocess them to obtain preprocessed ECG signals; Step 2: Based on the preprocessed electrocardiogram signal, extract time-domain features, frequency-domain features, and nonlinear features related to emotional state to form an initial feature set; Step 3: Fuse the multiple features in the initial feature set to construct a structured emotion feature matrix that integrates spatial distribution and multi-dimensional feature information; Step 4: Perform feature structure analysis on the structured emotion feature matrix. By defining a reference feature vector and calculating the distribution relationship between multiple associated feature vectors and the reference feature vector, a feature association region is constructed. Within the feature association region, two feature reference points with temporal relationship are selected respectively, and a feature adjustment path is formed based on the temporal connection relationship between the two feature reference points. A feature association metric is generated based on the feature adjustment path, and the feature association metric is used to adjust the structured emotion feature matrix to obtain the adjusted structured emotion feature matrix. Step 5: Input the adjusted structured emotion feature matrix into the deep learning model. The attention mechanism in the deep learning model adaptively enhances the key features related to emotion recognition in the structured emotion feature matrix, generating an enhanced emotion feature representation. Step 6: Based on the enhanced emotion feature representation, the time dependence of the electrocardiogram signal is obtained through the recurrent neural network in the deep learning model to generate the final emotion recognition result.
[0016] In this embodiment of the invention, by fusing three types of features—time domain, frequency domain, and nonlinearity—the dimensions of physiological signal changes related to emotions are comprehensively covered, providing a complete feature foundation for accurate recognition. Simultaneously, through the construction of feature association regions, the generation of time-series feature adjustment paths, and the adjustment of correlation metrics, dynamic optimization of features is achieved, adaptively selecting effective features and suppressing noise interference, making the feature matrix more aligned with recognition needs. Furthermore, the combination of the attention mechanism and recurrent neural network in the deep learning model not only strengthens the focus on key emotional features but also accurately captures the time dependence of electrocardiogram signals, improving the sensitivity of the deep learning model to emotional changes.
[0017] ECG signals can be conveniently collected through wearable devices such as smart bracelets and ECG patches. Combined with their optimization of signal processing and deep learning model inference processes, they have the characteristics of low latency and high adaptability, and can be widely used in diverse scenarios such as mental health monitoring, intelligent cockpit emotion intervention, human-computer interaction, and educational emotion feedback, with broad prospects for industrialization.
[0018] In a preferred embodiment of the present invention, step 1 above, which involves acquiring the raw electrocardiogram (ECG) signal and preprocessing the raw ECG signal to obtain a preprocessed ECG signal, may include: In this embodiment of the invention, a portable ECG acquisition device suitable for university students is selected. The device is equipped with disposable ECG electrodes. Before acquisition, the skin of the student's chest or the corresponding acquisition site on the limb is cleaned to remove sweat, oil, and other impurities, ensuring good contact between the electrodes and the skin. Then, the electrodes are attached according to the device specifications, typically selecting chest leads, such as V1-V6, or limb leads, such as the wrist or ankle. The electrodes are connected to the device host, ensuring a secure and tight connection. The acquisition device is then turned on, and sampling parameters that conform to ECG signal acquisition standards are set. Typically, the sampling rate is set to 250Hz-500Hz to ensure complete signal acquisition. To capture the waveform details of the electrocardiogram (ECG) signal, the student remains seated or in a normal active state while the device continuously collects ECG signals. The raw signals are stored in real time in digital form in the device's memory or transmitted wirelessly to a supporting terminal. The storage format uses a common medical signal format. Strong electromagnetic environments are avoided during the acquisition process to reduce the impact of external interference on the raw signals. The stored raw ECG signal data is then exported from the acquisition device or terminal and imported into the signal processing system. The data is first preliminarily sorted to remove invalid signal segments generated during the start and end of the acquisition phase due to device startup and wearing adjustments, while retaining the valid raw signal data from the stable acquisition phase.
[0019] Raw electrocardiogram (ECG) signals often contain low-frequency baseline drift caused by breathing and slight body movement. Adaptive filtering or wavelet transform is used to separate and filter out these slowly changing interferences, restoring a stable baseline and highlighting key waveforms such as the QRS complex, P wave, and T wave. For high-frequency noise such as electromyographic interference and power supply interference introduced during acquisition, low-pass or notch filtering is used to remove high-frequency noise above 20Hz and specific frequency power supply interference, preserving the effective frequency components of the ECG signal. When students exhibit large limb movements, motion artifacts are generated in the signal. These artifacts are then processed through signal... Amplitude detection and waveform morphology analysis identify artifact signal segments with abnormal amplitude and distorted waveforms. Short-term artifacts are repaired by interpolation with adjacent normal signals, while long-term severe artifacts are directly removed. The noise-removed signal undergoes amplitude standardization to uniformly adjust the signal amplitude to a fixed range, such as between -1 and 1, eliminating signal amplitude differences caused by physiological differences and varying electrode contact levels among different students. Subsequently, the signal is segmented and regularized according to a fixed time window or according to the heartbeat cycle to ensure that each segment has a consistent length and a complete waveform, ultimately resulting in a stable, clean, and standardized preprocessed ECG signal.
[0020] In a preferred embodiment of the present invention, step 2 above, which involves extracting time-domain features, frequency-domain features, and nonlinear features related to emotional state based on preprocessed electrocardiogram signals to form an initial feature set, may include: In this embodiment of the invention, step 220 involves extracting time-domain features to characterize the statistical properties of heart rate fluctuations based on the preprocessed electrocardiogram (ECG) signal, forming a subset of time-domain features. Specifically, this includes: first, performing QRS complex identification on the preprocessed stationary ECG signal, focusing on capturing the R-wave peak value, as the R-wave is the most amplitude-maximum and easily identifiable feature in the ECG signal. By identifying consecutive R-wave positions, the time interval between every two adjacent R-waves is determined; this interval is the normal sinus beat interval, or NN interval for short. Next, based on the obtained continuous NN interval sequence, various statistical properties characterizing heart rate fluctuations are calculated. The temporal characteristics of the heart rate are analyzed. Heart rate is calculated by dividing 60 by the duration of a single NN interval to obtain the instantaneous heart rate at that moment. The average heart rate is then calculated by averaging all instantaneous heart rates. This characteristic directly reflects the overall heart rate level. The standard deviation of the SDNN (Standard Deviation of Heart Rates) is calculated by first calculating the average of all NN intervals, then subtracting this average from each NN interval to obtain the difference between each NN interval and the average. Each difference is squared, and the sum of all squared results is obtained. This sum of squared differences is then divided by the total number of NN intervals minus one, and finally, the square root of the result is taken. This is known as SDNN, a feature that reflects the overall degree of heart rate fluctuation. RMSSD, or root mean square of the difference between adjacent NN intervals, is calculated by first calculating the difference between every two adjacent NN intervals, subtracting the previous NN interval from the next, resulting in a series of adjacent differences. Each adjacent difference is squared, and the sum of all squared results is obtained. The sum of squared adjacent differences is then divided by the total number of adjacent differences, and the square root of the result is the RMSSD. This feature primarily reflects the regulatory effect of the vagus nerve on heart rate. NN50 and pNN50 are also calculated. NN50 refers to the average square root of adjacent NN intervals. The number of NN interval differences with an absolute value greater than 50 milliseconds is calculated by first calculating the absolute value of the difference between each adjacent NN interval, and then counting the number of values greater than 50 milliseconds, which is NN50. pNN50 is the ratio of NN50 to the total number of adjacent NN intervals. The result of dividing the number of NN50 by the total number of adjacent NN intervals is pNN50. These two features are also used to reflect short-term heart rate fluctuations. The average heart rate, SDNN, RMSSD, NN50, pNN50 and other features calculated above are compiled and summarized to form a time-domain feature subset used to characterize the statistical properties of heart rate fluctuations.
[0021] Step 221: Based on the preprocessed ECG signal, perform power spectrum analysis on the heartbeat cycle signal on which the time-domain feature subset is based, and extract frequency-domain features to characterize the state of autonomic nervous activity, forming a frequency-domain feature subset; specifically, this includes: First, obtaining the NN interval sequence used to calculate the time-domain features in step 220. This sequence is the heartbeat cycle signal. Since the NN interval sequence is a time series with unequal intervals, it needs to be interpolated first. A linear interpolation method is used, and a fixed sampling frequency is selected, usually 100 Hz, to interpolate the unequal NN intervals. The time series is converted into an equally spaced time series to ensure the accuracy of power spectrum analysis. Then, power spectrum analysis is performed on the interpolated equally spaced NN interval series. First, Fourier transform is performed on the series to convert the time domain signal into a frequency domain signal, and the power spectral density distribution of the signal is obtained. According to the standard of heart rate variability frequency domain analysis, the power spectrum is divided into different frequency bands, including low frequency band, high frequency band and very low frequency band. The frequency range of the low frequency band is 0.04 Hz to 0.15 Hz, and the frequency range of the high frequency band is 0.15 Hz to 0.4 Hz.
[0022] Then, the power values of each frequency band are calculated. For the low-frequency band, the power spectral density values corresponding to all frequency points within the band are statistically analyzed, and these values are integrated and summed to obtain the low-frequency power, abbreviated as LF. For the high-frequency band, the power spectral density values corresponding to all frequency points within the band are statistically analyzed, and the high-frequency power, abbreviated as HF, is obtained after integration and summation. The total power is calculated by adding the power of the low-frequency band, the power of the high-frequency band, and the power of other frequency bands to obtain the total power of the signal. The LF / HF ratio is calculated by dividing the low-frequency power by the high-frequency power. The calculated features such as LF, HF, total power, and LF / HF ratio are compiled and summarized to form a frequency domain feature subset used to characterize the state of autonomic nervous activity. Among them, LF mainly reflects the joint action of the sympathetic and vagus nerves, HF mainly reflects the activity state of the vagus nerve, and the LF / HF ratio reflects the balance between the sympathetic and vagus nerves.
[0023] Step 222: Based on the preprocessed ECG signal, the signal complexity, which forms the basis for time-domain and frequency-domain feature subset analysis, is quantified, and nonlinear features used to characterize the nonlinear dynamics of the ECG signal are extracted to form a nonlinear feature subset. Specifically, this includes: First, determining the signal source for extracting nonlinear features. This can be the preprocessed original ECG signal waveform or the NN interval sequence obtained in step 220. Both types of signals reflect the dynamic changes of the ECG signal and are the basis for time-domain and frequency-domain feature analysis. Their complexity changes are closely related to emotional state. Next, for the selected signal, various nonlinear features used to quantify signal complexity are calculated, and sample entropy is calculated. First, the embedding dimension is set, usually 2, and the similarity tolerance is usually 0.2 times the signal standard deviation. The signal sequence is reconstructed into multiple vectors according to the embedding dimension, and then the relationship between each vector and all other vectors is calculated. Similarity is determined when the absolute value of the difference between corresponding elements of vectors is less than the similarity tolerance. The number of similar vectors for each vector is counted, and the similarity probability of each vector is calculated. The natural logarithm of the similarity probabilities of all vectors is taken, and the average is calculated. The embedding dimension is then changed, and the above operation is repeated. Finally, the difference between the two averages is used as the sample entropy. This feature reflects the irregularity and complexity of the signal; the more complex the signal, the larger the sample entropy value. Approximate entropy is calculated by setting the same embedding dimension and similarity tolerance as the sample entropy. The signal sequence is reconstructed into vectors, and the ratio of the number of similar vectors (including itself) of each vector to the total number of vectors is calculated, i.e., the similarity probability. The natural logarithm of all similar probabilities is taken, and the average is calculated. The embedding dimension is changed, and the operation is repeated. The difference between the two averages is the approximate entropy. This feature is similar to sample entropy and is used to characterize the complexity of the signal, but it is less sensitive to the amount of data.
[0024] The fractal dimension is calculated using the box-counting method. First, a series of boxes of different sizes are determined, and these boxes are used to sequentially cover the ECG signal waveform curve. The minimum number of boxes required to cover the curve each time is counted. Then, a double logarithmic coordinate graph is plotted with the logarithm of the box side length as the x-axis and the logarithm of the number of boxes as the y-axis. The slope of the graph is obtained through linear fitting, and this slope is the box-counting dimension. The larger the fractal dimension, the more complex the ECG signal waveform and the more obvious the nonlinear characteristics. The calculated sample entropy, approximate entropy, fractal dimension, and other features are compiled and summarized to form a subset of nonlinear features used to characterize the nonlinear dynamics of the ECG signal.
[0025] Step 223 involves integrating the time-domain feature subset, frequency-domain feature subset, and nonlinear feature subset to form an initial feature set. Specifically, this includes: First, reviewing the time-domain, frequency-domain, and nonlinear feature subsets to clarify the specific feature names and corresponding calculation results within each subset, ensuring the source and representational meaning of each feature are clearly traceable. For example, SDNN in the time-domain feature subset corresponds to the overall degree of heart rate fluctuation, HF in the frequency-domain feature subset corresponds to vagal nerve activity, and sample entropy in the nonlinear feature subset corresponds to signal complexity. Then, according to the feature type order (time-domain features first, then frequency-domain features, and finally nonlinear features), or according to the importance of features in representing emotional states, all features in the three feature subsets are sequentially arranged and combined. During the arrangement process, it is ensured that the position of each feature is fixed, without repetition or omission, forming a complete feature set framework. The specific numerical values of each feature are then filled into the set framework, making the set contain all features representing the statistical characteristics of heart rate fluctuations, the state of autonomic nervous activity, and the nonlinear dynamics of the signal, ultimately forming an initial feature set that comprehensively reflects the correlation between electrocardiogram signals and emotional states.
[0026] By extracting and integrating three types of features, the initial feature set can comprehensively cover the multidimensional changes in electrocardiogram signals caused by emotional changes. The time-domain features capture the statistical regularity of heart rate fluctuations, the frequency-domain features reflect the activity balance of the autonomic nervous system, and the nonlinear features characterize the complex dynamic characteristics of the signal. This provides a comprehensive and sufficient feature basis for accurately distinguishing different emotions with similar physiological manifestations, thereby improving the accuracy of emotion recognition.
[0027] In a preferred embodiment of the present invention, step 3 above, which involves fusing multiple types of features in the initial feature set to construct a structured emotion feature matrix that integrates spatial distribution and multi-dimensional feature information, may include: In this embodiment of the invention, step 330 involves extracting feature values from the time-domain feature subset, frequency-domain feature subset, and nonlinear feature subset in the initial feature set to obtain the numerical sets of each feature subset. Specifically, this includes: first, clarifying the specific composition of the three feature subsets in the initial feature set; the time-domain feature subset includes features such as average heart rate, SDNN, RMSSD, NN50, and pNN50; the frequency-domain feature subset includes features such as LF, HF, total power, and the LF / HF ratio; and the nonlinear feature subset includes features such as sample entropy, approximate entropy, and fractal dimension. Next, feature value extraction is performed on the time-domain feature subset, extracting each feature pair within the subset one by one. The specific values are all derived from the calculation results of the preprocessed ECG signals in step 220. For example, the average heart rate, SDNN, RMSSD, NN50, and pNN50 values of a certain sample are extracted and arranged in a fixed order of time-domain features, such as average heart rate, then SDNN, then RMSSD, NN50, and finally pNN50, to form the time-domain feature value sequence of the sample. The above operation is repeated for all collected student emotion monitoring samples to obtain the time-domain feature value sequence of each sample. The time-domain feature value sequences of all samples are summarized to form a set of values for the time-domain feature subset.
[0028] Then, following the same logic, the numerical set of the frequency domain feature subset is extracted. Specific values for LF, HF, total power, and LF / HF ratio within each sample's frequency domain feature subset are extracted. These values come from the power spectrum analysis results of step 221. The frequency domain feature values of each sample are arranged in a fixed order, such as first LF, then HF, then total power, and finally the LF / HF ratio, forming a sequence of frequency domain feature values for each sample. This sequence is then summarized for all samples to obtain the numerical set of the frequency domain feature subset. Finally, the numerical set of the nonlinear feature subset is extracted. Specific values for sample entropy, approximate entropy, and fractal dimension within each sample's nonlinear feature subset are extracted. These values come from the signal complexity quantification results of step 222. The nonlinear feature values of each sample are arranged in a fixed order, such as first sample entropy, then approximate entropy, and finally fractal dimension, forming a sequence of nonlinear feature values for each sample. This sequence is then summarized for all samples to obtain the numerical set of the nonlinear feature subset.
[0029] Step 331: Based on the numerical set, normalize the time-domain feature subset, frequency-domain feature subset, and nonlinear feature subset respectively to obtain standardized time-domain feature vectors, frequency-domain feature vectors, and nonlinear feature vectors. Specifically, this includes: First, determining the normalization method, selecting the min-max normalization method, which can uniformly map feature values to the range of 0 to 1, eliminating the influence of differences in dimensions and numerical ranges of different features, and adapting to different features in the electrocardiogram signals of college students, such as the numerical difference between heart rate and entropy values; Normalizing the numerical set of the time-domain feature subset, taking the average heart rate feature in the time-domain feature subset as an example, first traversing the average heart rate of all samples. Find the maximum and minimum values among them. For the average heart rate value of each sample, subtract the minimum average heart rate from the average heart rate value, and divide the difference by the difference between the maximum and minimum average heart rates to calculate the normalized average heart rate value of the sample. Following the same calculation method, normalize all sample values of each feature in the time-domain feature subset, such as SDNN, RMSSD, NN50, and pNN50, to obtain the normalized value sequence of the time-domain feature of each sample. Arrange the normalized value sequence of the time-domain feature of each sample in feature order to form the standardized time-domain feature vector of the sample. The standardized time-domain feature vectors of all samples together constitute the time-domain feature vector set.
[0030] Next, the numerical set of the frequency domain feature subset is normalized. Taking the LF feature as an example, the maximum and minimum values of the LF values of all samples are first found. The minimum LF value of each sample is subtracted from the LF value, and the difference is divided by the difference between the maximum and minimum LF values to obtain the normalized LF value of that sample. In this way, the frequency domain features such as HF, total power, and LF / HF ratio are normalized in turn to form a standardized frequency domain feature vector for each sample. The sets of frequency domain feature vectors are then collected. Finally, the numerical set of the nonlinear feature subset is normalized. Taking the sample entropy feature as an example, the maximum and minimum values of the sample entropy values of all samples are found. The minimum sample entropy value of each sample is subtracted from the sample entropy value, and the difference is divided by the difference between the maximum and minimum sample entropy values to obtain the normalized sample entropy value of that sample. The approximate entropy and fractal dimension are processed in the same way to form a standardized nonlinear feature vector for each sample. The sets of nonlinear feature vectors are then collected.
[0031] Step 332 involves concatenating the time-domain feature vector, frequency-domain feature vector, and nonlinear feature vector in sequence according to the feature category dimension to form a multi-dimensional fused feature vector. Specifically, this includes: First, determining a fixed order for vector concatenation. To ensure clear and traceable feature categories, the time-domain feature vector, frequency-domain feature vector, and nonlinear feature vector are concatenated in that order. This order aligns with the logical order of feature extraction and normalization, meeting the needs of multi-dimensional feature correlation analysis in emotion recognition. Taking a single student emotion monitoring sample as an example, the standardized time-domain feature vector of the sample is first obtained. This vector contains normalized values of time-domain features such as average heart rate and SDNN, arranged in feature order to form a continuous numerical sequence. Next, the standardized frequency-domain feature vector of the sample is obtained, containing normalized values of frequency-domain features such as LF and HF, also arranged in feature order to form a numerical sequence. Finally, the standardized nonlinear feature vector of the sample is obtained, containing normalized values of nonlinear features such as sample entropy and approximate entropy, arranged in feature order to form a numerical sequence.
[0032] The three numerical sequences are concatenated in a predetermined order: first, the numerical sequence of the time-domain feature vector is preserved intact; then, the numerical sequence of the frequency-domain feature vector is appended to its end; finally, the numerical sequence of the nonlinear feature vector is appended to the end of the frequency-domain feature vector numerical sequence. The three sequences are then combined into a long numerical sequence that encompasses the standardized information of the sample's time-domain, frequency-domain, and nonlinear features, forming the sample's multi-dimensional fusion feature vector. This concatenation operation is repeated for all student emotion monitoring samples to obtain the multi-dimensional fusion feature vector corresponding to each sample.
[0033] Step 333: Based on the multi-dimensional fusion feature vectors of multiple samples, arrange and combine them according to the sample dimensions to construct a multi-row, multi-column structured emotion feature matrix with rows representing feature samples and columns representing feature types. Specifically, this includes: First, determining the row and column feature definitions of the matrix. To adapt to the dual correlation analysis of samples and feature types in the subsequent feature structure analysis, the rows of the matrix correspond to emotion monitoring samples, and each row vector represents a multi-dimensional fusion feature of a sample; the columns of the matrix correspond to feature types, and each column vector represents the standardized value of all samples on a specific feature. Next, determining the number of rows and columns of the matrix. The total number of all student emotion monitoring samples is counted, and this number is the number of rows in the matrix. The total number of features contained in a single multi-dimensional fusion feature vector is counted, which is the sum of the number of time-domain features, the number of frequency-domain features, and the number of nonlinear features. This total number is the number of columns in the matrix. For example, if there are 5 time-domain features, 4 frequency-domain features, and 3 nonlinear features, then a single fusion vector contains 12 features, and the number of columns in the matrix is 12; if there are 100 monitoring samples, then the number of rows in the matrix is 100.
[0034] Then, the vectors are arranged and combined. The multi-dimensional fusion feature vector of the first sample is selected and used as the first row of the matrix. Each feature value in the vector is filled into the corresponding column of the first row in sequence. That is, the first value of the vector is filled into the first column of the first row, the second value is filled into the second column of the first row, and so on until all columns of the first row are filled. Next, the multi-dimensional fusion feature vector of the second sample is selected and used as the second row of the matrix. The feature values are filled into each column of the second row in the same way. According to the sample collection order or numbering order, the multi-dimensional fusion feature vectors of all samples are used as the row vectors of the matrix and filled into the corresponding columns row by row. Finally, a two-dimensional matrix with multiple rows and columns is formed. This matrix is the preliminary structured emotion feature matrix. The row direction fully presents the feature information of all emotion monitoring samples, and the column direction clearly distinguishes different types of features, realizing the structured association between samples and features.
[0035] Step 334 involves enhancing the spatial distribution features of the structured emotion feature matrix. This is achieved by calculating the statistical relationships between feature columns and supplementing the spatial distribution descriptors that describe feature correlations, ultimately forming a structured emotion feature matrix that integrates spatial distribution and multi-dimensional feature information. Specifically, this includes: First, determining the calculation type of the statistical relationship between feature columns, using the Pearson correlation coefficient as the statistical relationship indicator. This indicator accurately quantifies the linear correlation between two feature columns, aligning with the needs of analyzing the correlation between different features in electrocardiogram signals, such as HF and RMSSD, caused by emotional changes. It can effectively capture the differences in correlation between features under emotions such as anxiety and depression. For the initially constructed structured emotion feature matrix, any two different feature columns are selected, such as the column representing RMSSD and the column representing HF, to calculate the correlation coefficient. The calculation process is as follows: First, calculate the average of all sample values for these two feature columns respectively. For each sample, subtract the mean of the first feature column from the sample's value in the first feature column to obtain the first difference; simultaneously, subtract the mean of the second feature column from the sample's value in the second feature column to obtain the second difference; multiply the first and second differences for each sample to obtain the product of differences for that sample; sum the products of differences for all samples to obtain the sum of difference products; divide the sum of difference products by the sample size minus one to obtain the covariance of the two feature columns; calculate the standard deviation of all sample values for each feature column separately, i.e., first calculate the difference between each sample value and the mean of that column, square the difference, sum the results, divide by the sample size minus one, and finally take the square root of the result to obtain the standard deviation of each; divide the covariance of the two feature columns by the product of the standard deviations of the first and second feature columns to obtain the Pearson correlation coefficient between the two feature columns.
[0036] Following the above method, the Pearson correlation coefficient between each pair of different feature columns in the matrix is calculated sequentially. All correlation coefficients are then arranged in the order of the feature columns to form a square correlation coefficient matrix. Each element in this matrix represents the degree of correlation between the corresponding two feature columns, which is the spatial distribution descriptor describing the feature correlation. Finally, this correlation coefficient matrix is used as supplementary information and integrated with the preliminary structured sentiment feature matrix. The integration method is to concatenate the correlation coefficient matrix to the right or below the preliminary matrix, so that the final matrix contains both the multi-dimensional feature values of each sample and the correlation information between all features, forming a structured sentiment feature matrix that integrates spatial distribution and multi-dimensional feature information.
[0037] By constructing a structured emotion feature matrix, we have achieved deep integration of three types of features: time domain, frequency domain, and nonlinearity. At the same time, we have integrated the spatial distribution correlation information between features, reducing the defects of single-dimensional features or lack of feature correlation information. The spatial distribution descriptors in the matrix can accurately capture these differences, effectively avoiding misjudgment of emotions caused by incomplete feature information or missing correlations. This provides reliable support for psychological teachers to accurately distinguish students' emotion types and formulate appropriate guidance strategies.
[0038] In a preferred embodiment of the present invention, step 4 above involves performing feature structure analysis on the structured emotion feature matrix. This is achieved by defining a reference feature vector and calculating the distribution relationship between multiple associated feature vectors and the reference feature vector to construct a feature association region. Within each feature association region, two feature reference points with a temporal relationship are selected, and a feature adjustment path is formed based on the temporal connection relationship between the two feature reference points. A feature correlation metric is generated based on the feature adjustment path, and this metric is used to adjust the structured emotion feature matrix, resulting in an adjusted structured emotion feature matrix. This adjustment may include: In this embodiment of the invention, step 440 involves selecting feature samples with typical emotional representations from the structured emotion feature matrix, defining them as reference feature vectors. Specifically, in the structured emotion feature matrix, each row corresponds to a student's electrocardiogram signal feature sample at a certain moment, each column corresponds to a type of feature, and each sample is clearly labeled with an emotion category. These labels are consistent with the emotion types focused on in university mental health monitoring, including calm, depressed, and anxious. When selecting reference feature vectors, for each target emotion category to be identified, samples that best reflect the core physiological characteristics of the emotion within that category are selected from the matrix. Taking calm emotion as an example, the selection criteria are that the central rate of the time-domain feature is stable within the normal range, such as 60-80 beats / minute, and the standard deviation of the RR interval (SDNN) is relatively low. For example, in the frequency domain features, the proportion of high-frequency (HF) components is high, and the ratio of low-frequency to high-frequency (LF / HF) is small. In the nonlinear features, the sample entropy is at a moderate level. Taking anxiety as an example, the screening criteria are a significant increase in heart rate, such as exceeding 90 beats / minute, and a decrease in SDNN. In the frequency domain features, the proportion of HF components decreases, the LF / HF ratio increases, and the sample entropy in the nonlinear features increases significantly. Taking depression as an example, the screening criteria are a slight increase in heart rate, such as 70-85 beats / minute, and a decrease in SDNN. In the frequency domain features, the proportion of HF components decreases, the LF / HF ratio increases slightly, and the sample entropy in the nonlinear features decreases significantly. The samples that meet the corresponding emotion screening criteria and have the most representative feature performance are determined as the reference feature vector for that emotion category. All feature structure analyses use this vector as the core benchmark.
[0039] Step 441: Based on the reference feature vector, calculate the Euclidean distance between all other feature samples in the structured emotion feature matrix and the reference feature vector in the feature space to obtain the distance distribution sequence. Specifically, this includes: after determining the reference feature vector, for each feature sample in the structured emotion feature matrix other than the reference vector, calculate the Euclidean distance between them in the feature space. The feature space is composed of three types of features: time domain, frequency domain, and nonlinear features. Each feature sample and the reference feature vector have a one-to-one corresponding feature value in each dimension of the three types of features. When calculating the Euclidean distance between a single sample and the reference vector, start from the time domain feature dimension, take the heart rate value of the sample, subtract the heart rate value of the reference feature vector from the heart rate value to obtain the difference in the heart rate dimension, and then square the difference. Next, take the SDNN value of the sample, subtract the SDNN value of the reference vector from it to obtain the difference in the SDNN dimension and square it. Following the same method, the squared differences between the sample and the reference vector are calculated sequentially across all time-domain feature dimensions. After completing the time-domain feature dimension calculations, the process shifts to the frequency-domain feature dimensions, calculating the squared differences between the sample and the reference vector in each frequency-domain feature dimension, such as LF value, HF value, and LF / HF ratio. Subsequently, the nonlinear feature dimensions are calculated, with the squared differences between the sample and the reference vector calculated in each nonlinear feature dimension, such as sample entropy and Lempel-Ziv complexity. After the squared differences across all feature dimensions have been calculated, all the squared difference results are summed to obtain a total value. Finally, the square root of this total value is taken, and the result is the Euclidean distance between the current feature sample and the reference feature vector. Following the complete calculation process described above, all other feature samples in the matrix are processed one by one, and each obtained Euclidean distance is arranged sequentially according to the calculation order to form a distance distribution sequence.
[0040] Step 442 involves statistical analysis of the distance distribution sequence to determine the dynamic distance threshold used to define the feature association range. Specifically, this includes: after obtaining the distance distribution sequence, a comprehensive statistical analysis of all distance values within the sequence is performed to calculate the core statistical indicators of the sequence. The first step is to calculate the arithmetic mean of the sequence by summing all Euclidean distance values in the sequence, then dividing the sum by the number of distance values (i.e., the number of feature samples involved in the calculation) to obtain the average. The second step is to calculate the standard deviation of the sequence by subtracting the average from each distance value to obtain the deviation value of each distance from the average, then squaring each deviation value to obtain the squared deviation value. All deviations are then squared. The variance is obtained by summing the squared deviations and dividing this sum by the number of distance values. The standard deviation is obtained by taking the square root of the variance. The dynamic distance threshold is determined based on the calculated mean and standard deviation. Specifically, the mean is added to twice the standard deviation. The dynamic distance threshold is determined by the actual statistical characteristics of the current distance distribution sequence, rather than a fixed value. It can be flexibly adjusted according to the ECG characteristics of different emotion categories and different student groups to ensure that the selected associated samples have a real strong correlation with the reference feature vector, avoiding the problem of overly broad or narrow selection of associated samples caused by a fixed threshold.
[0041] Step 443: Based on the dynamic distance threshold, select a subset of feature samples from the structured emotion feature matrix that are closest to the reference feature vector, defining it as the associated feature vector set. Specifically, this includes: using the dynamic distance threshold determined in step 442 as the selection criterion, comparing the Euclidean distance between each of the remaining feature samples in the structured emotion feature matrix and the reference feature vector with the threshold. If the Euclidean distance of a feature sample is less than or equal to the dynamic distance threshold, it indicates that the sample is close to the reference feature vector in the feature space, and its time-domain, frequency-domain, and nonlinear feature combinations are similar to the emotion feature pattern represented by the reference vector. High similarity indicates a strong correlation between the student's emotional state and the target emotional category, such as anxiety or depression. If the Euclidean distance of a feature sample is greater than the dynamic distance threshold, it indicates that the feature pattern of the sample differs significantly from that of the reference feature vector, and the corresponding emotional state is less correlated with the target emotion. All feature samples whose Euclidean distance is less than or equal to the dynamic distance threshold are selected and integrated into a feature sample subset. This subset is formally defined as the associated feature vector set. Each sample in this set carries feature information that is highly correlated with the target emotion, providing a data foundation for accurately analyzing the changing patterns of emotional characteristics.
[0042] Step 444: Construct a feature association region centered on the position of the reference feature vector in the feature space and with a dynamic distance threshold as the spatial radius. Specifically, the feature space is a multi-dimensional space composed of time-domain, frequency-domain, and nonlinear features. The position of the reference feature vector in this space is determined by the eigenvalues of all its feature dimensions. Each feature dimension corresponds to a coordinate component in the space. When constructing the feature association region, the coordinate position of the reference feature vector in the feature space is first set as the origin of the region. Then, the dynamic distance threshold calculated in step 442 is used as the spatial radius of the region. A closed spatial region is delineated around the origin in the feature space. The boundary of this region is formed by connecting the coordinate positions of all feature samples whose Euclidean distance to the reference feature vector is exactly equal to the dynamic distance threshold in the feature space. The region completely contains the reference feature vector and all samples in the set of associated feature vectors selected in step 443. The region's exterior contains feature samples with weaker correlation to the target emotion. By constructing this feature association region, target emotion-related samples scattered throughout the feature space can be concentrated into a specific range, effectively eliminating the influence of irrelevant features and interfering samples, allowing feature analysis to focus on effective information.
[0043] Step 445: Within the feature association region, based on the timestamp information of the feature samples, select two adjacent feature samples in chronological order as the first feature reference point and the second feature reference point, respectively. Specifically, each feature sample in the structured emotion feature matrix corresponds to a specific acquisition time of the student's electrocardiogram signal. Therefore, each sample is accompanied by a unique timestamp. The timestamp accurately records the acquisition time of the sample and can be directly used to determine the temporal relationship between samples. Within the feature association region, the timestamp information of all samples is first collected. Then, according to the acquisition time from earliest to latest, all feature samples within the region are sorted to form a chronologically arranged sample sequence. From this sequence... Two samples that are adjacent in time sequence are selected, namely the sample corresponding to the previous collection time and the sample corresponding to the immediately following collection time. After clarifying the temporal relationship between the two samples, the sample with the earlier collection time is determined as the first feature reference point, and the sample adjacent to it with the later collection time is determined as the second feature reference point. Adjacent temporal samples are selected as reference points because the changes in ECG signal characteristics at adjacent times can most realistically and delicately reflect the dynamic evolution of students' emotions over time. For example, the changes in ECG characteristics of students from mild anxiety to moderate anxiety in a short period of time will be reflected in adjacent samples. This provides an accurate temporal basis for constructing a feature adjustment path that conforms to the laws of emotional change.
[0044] Step 446: Based on the coordinate positions of the first and second feature reference points in the feature space, calculate the spatial vector relationship between the two feature reference points. Specifically, the coordinate positions of the first and second feature reference points in the feature space are composed of the eigenvalues of each dimension of the time-domain, frequency-domain, and nonlinear features they contain. Each feature dimension's value corresponds to a component of the coordinates, and the number of coordinate components is consistent with the total number of feature dimensions. When calculating the spatial vector relationship, the same operational logic is performed for each feature dimension: first, obtain the eigenvalue of the second feature reference point in that dimension; then, obtain the eigenvalue of the first feature reference point in the same dimension; and finally, subtract the eigenvalue of the first reference point from the eigenvalue of the second reference point to obtain the spatial vector relationship. The difference in eigenvalues between two reference points in a feature dimension is calculated as follows: For example, in the sample entropy dimension of nonlinear features, the difference in sample entropy is obtained by subtracting the sample entropy value of the first reference point from the sample entropy value of the second reference point; similarly, in the HF dimension of frequency domain features, the difference in HF is obtained by subtracting the HF value of the first reference point from the HF value of the second reference point. All the eigenvalue differences corresponding to all feature dimensions are arranged sequentially according to the time domain, frequency domain, and nonlinear feature categories to form a complete vector. This vector is the spatial vector between the two feature reference points. Each component of this vector accurately reflects the specific changes in the corresponding feature dimension between two adjacent time-series reference points, fully presenting the positional differences and trends of the two reference points in the feature space.
[0045] Step 447: Based on the spatial vector relationship, determine the direction vector from the first feature reference point to the second feature reference point, and establish a linear connection between the two feature reference points. Specifically, this includes: determining the direction vector based on the spatial vector calculated in step 446. The vector components of this direction vector are exactly the same as those of the spatial vector. Its core function is to clarify the direction of feature change, pointing from the first feature reference point to the second feature reference point. This intuitively reflects the direction of change of each feature dimension over time, from the time of collection at the first reference point to the time of collection at the second reference point. For example, if the vector component of the sample entropy dimension is positive, it indicates that the sample entropy is increasing from the first reference point to the second reference point, which corresponds to the student's emotion possibly developing towards anxiety. After determining the direction vector, a linear connection between the two feature reference points is established. That is, it is assumed that in the feature space, the feature change from the first feature reference point to the second feature reference point proceeds along a straight line trajectory. The emotional feature state at any time between the two reference points can be represented by the corresponding point on this straight line. The establishment of this linear connection is based on the continuous pattern of emotional changes at adjacent times, providing a trajectory basis that conforms to the logic of emotional evolution for constructing feature adjustment paths.
[0046] Step 448: Based on the linear connection relationship, construct a straight path extending from the first feature reference point to the second feature reference point to form a feature adjustment path. Specifically, this includes: using the linear connection relationship between the first and second feature reference points established in step 447 as the trajectory basis, construct a straight path in the feature space. The starting point of this path is the coordinate position of the first feature reference point in the feature space, and the ending point is the coordinate position of the second feature reference point. The path extends completely along the linear connection direction between the two. Each point on the path corresponds to a complete set of feature values. The components of each dimension of these feature values start from the corresponding component of the first feature reference point and gradually and evenly transition to the corresponding component of the second feature reference point along the trend indicated by the direction vector. For example, if the sample entropy value of the first reference point is 0.8 and the sample entropy value of the second reference point is 1.2, then the sample entropy value of the points on the path between the two will gradually increase from 0.8 to 1.2, corresponding to the gradual increase of the student's anxiety. This straight path completely simulates the continuous change process of features between two adjacent time-series reference points and is defined as the feature adjustment path. Based on the feature change pattern on this path, key information for emotion recognition will be extracted.
[0047] Step 449: Based on the spatial relationship between two feature reference points on the feature adjustment path, calculate the magnitude of change of each feature dimension along the path direction to form a feature correlation measurement parameter. Specifically, this includes: for the first and second feature reference points corresponding to the feature adjustment path, analyze the magnitude of change of each feature dimension along the path direction to reflect the sensitivity of that dimension to emotional changes. When calculating the magnitude of change of a single feature dimension, first obtain the feature value of the first feature reference point in that dimension, then obtain the feature value of the second feature reference point in the same dimension. Subtract the feature value of the first reference point from the feature value of the second reference point to obtain the absolute change of the feature value of that dimension. Then take the absolute value of this absolute change; the result is the magnitude of change of that feature dimension along the path direction. For example, in nonlinear features... For the sample entropy dimension of the feature, if the sample entropy of the first baseline point is 0.8 and the sample entropy of the second baseline point is 1.2, the difference between the two is 0.4, and the absolute value of the difference is 0.4. In the heart rate dimension of the time domain feature, if the heart rate of the first baseline point is 85 beats / minute and the heart rate of the second baseline point is 90 beats / minute, the difference is 5, and the change amplitude is 5. Following the above method, the change amplitudes of all time domain feature dimensions, frequency domain feature dimensions, and nonlinear feature dimensions are calculated in sequence. The change amplitudes of all feature dimensions are arranged in the order of time domain-frequency domain-nonlinearity to form a complete set of parameters. This set is the feature correlation measurement parameter. The value of each parameter directly reflects the degree of fluctuation of the corresponding feature dimension in the process of emotional change in adjacent time series. The larger the value, the more sensitive the dimension is to emotional change.
[0048] Step 450 involves normalizing the feature correlation measurement parameters to obtain an assessment of the relative contribution of each feature dimension to emotion recognition. Specifically, this includes: first, summing the feature correlation measurement parameters obtained in step 449, adding the magnitudes of change for all feature dimensions to obtain the total magnitude of change for all feature dimensions. For example, if the total magnitude of change for each time-domain feature dimension is 10, the total magnitude of change for each frequency-domain feature dimension is 8, and the total magnitude of change for each nonlinear feature dimension is 12, then the total magnitude of change for all feature dimensions is 30. Subsequently, for each feature dimension's magnitude of change, this magnitude is used... Dividing by the sum of the changes in all feature dimensions yields a ratio. For example, if the change in a certain nonlinear feature (sample entropy) is 0.4 and the sum is 30, then the ratio for that dimension is 0.4 / 30≈0.013. This ratio is the result after normalization of that feature dimension. Since it reflects the proportion of the change in a single feature dimension to the total change in all feature dimensions, the higher the proportion, the more significant the role of that feature dimension in the evolution of emotion over time, and the greater its impact on emotion state recognition. Therefore, this ratio is directly used as the evaluation result of the relative contribution of each feature dimension in emotion recognition.
[0049] Step 451: Based on the relative contribution evaluation results, establish the proportional relationship for feature adjustment, and based on this proportional relationship, differentiate the feature dimensions in the original structured emotion feature matrix. Specifically, this includes: establishing the proportional relationship for feature adjustment based on the relative contribution evaluation results of each feature dimension; a feature dimension with a higher relative contribution means it plays a more crucial role in emotion recognition, and its corresponding adjustment coefficient is larger, thus highlighting its feature differences through a larger adjustment; a feature dimension with a lower relative contribution has a smaller impact on emotion recognition, and its corresponding adjustment coefficient is smaller, requiring only slight adjustments to avoid feature distortion. For example, for nonlinear features with a high relative contribution (such as sample entropy, contributing 15%), the adjustment coefficient is set to 1.2; for frequency domain features with a medium relative contribution (such as HF components, contributing 1%), the adjustment coefficient is set to 1.2. For the first feature, the adjustment coefficient is set to 1.1. For the second feature, the adjustment coefficient is set to 1.05. After determining the adjustment coefficient for each dimension, for each feature sample in the original structured emotion feature matrix, the feature value of that dimension in the sample is adjusted according to the adjustment coefficient corresponding to each feature dimension. Specifically, the feature value of the sample in that dimension is multiplied by the corresponding adjustment coefficient to obtain the adjusted feature value. For example, if the sample entropy value of a sample is 1.0, multiplying it by the adjustment coefficient 1.2 will result in an adjusted sample entropy value of 1.2. If the HF value of a sample is 200, multiplying it by the adjustment coefficient 1.1 will result in an adjusted HF value of 220. In this way, the differential adjustment of each feature dimension is achieved, amplifying the differences of key features and weakening the interference of secondary features.
[0050] Step 452 involves standardizing and integrating the differentiated feature matrix to form a structured emotion feature matrix with adjusted feature distribution. Specifically, the feature matrix after differentiation adjustment exhibits significant differences in the range of feature values across different feature dimensions. For example, the heart rate value for time-domain features may range from 60 to 120, and the sample entropy value for non-linear features may range from 0.5 to 2.0. These scale differences can interfere with the judgment of feature importance; therefore, standardization and integration are necessary. Standardization is performed dimension by dimension. For each feature dimension, firstly, the feature values of all feature samples under that dimension are collected after adjustment, and the maximum and minimum values of that dimension are calculated. Then, the feature value of each feature sample under that dimension is standardized by subtracting the minimum value of that dimension from the feature value of that sample, resulting in a difference. The value is then divided by the difference between the maximum and minimum values of that dimension to obtain the standardized feature value. For example, the maximum value of a certain non-linear feature dimension after adjustment is 2.0, and the minimum value is 0.5. The adjusted value of this dimension for a certain sample is 1.2. The standardization operation is (1.2-0.5) / (2.0-0.5)=0.7 / 1.5≈0.47, that is, the standardized value of this sample in this dimension is 0.47. Through this operation, the feature values of each feature dimension are uniformly mapped to the range of 0 to 1 to ensure that the scale of each feature dimension is consistent. After all feature dimensions have been standardized, the standardized feature values are rearranged according to the sample dimension (rows) and feature category dimension (columns) of the original structured emotion feature matrix to form a new matrix. This matrix is the structured emotion feature matrix after feature distribution adjustment.
[0051] By selecting reference feature vectors of typical emotions, combining dynamic distance thresholds to filter related samples and construct feature association regions, we can accurately focus on feature samples related to the target emotion, effectively eliminating the influence of physiological noise features and interference samples unrelated to emotion. Based on time-series benchmarks, we can construct feature adjustment paths, calculate the magnitude of change in each dimension and convert it into relative contribution, accurately identify the feature dimensions that play a key role in distinguishing emotions, improve the distinguishability between different emotions, and ensure that psychological teachers can accurately identify students' true emotional states.
[0052] In a preferred embodiment of the present invention, step 5 above, which involves inputting the adjusted structured emotion feature matrix into a deep learning model and adaptively enhancing key features related to emotion recognition in the structured emotion feature matrix through the attention mechanism in the deep learning model to generate an enhanced emotion feature representation, may include: In this embodiment of the invention, step 550 involves inputting the adjusted structured emotion feature matrix into the feature analysis layer of the deep learning model, and obtaining the distribution of correlation strength between features through feature interaction calculation. Specifically, the feature analysis layer consists of two fully connected network layers. The number of neurons in the first fully connected network is exactly the same as the number of feature dimensions in the structured emotion feature matrix, and the number of neurons in the second fully connected network is set to twice the number of feature dimensions. This setting is to more fully capture the interaction relationship between different feature dimensions. First, the adjusted structured emotion feature matrix is converted into a sequence of feature vectors. Each feature vector corresponds to a row of sample data in the matrix, containing the values of all feature dimensions, such as time-domain features (e.g., heart rate and RR interval standard deviation), frequency-domain features (e.g., low-frequency and high-frequency components), and nonlinear features (e.g., sample entropy). These feature vector sequences are input into the first fully connected network, which performs a linear transformation on each feature vector. Specifically, the value of each feature dimension is multiplied by the weight corresponding to that dimension, and then all the multiplication results are summed, plus a fixed bias. The value is calculated to obtain an intermediate vector with the same dimension as the original feature vector. Then, all intermediate vectors are input into the second fully connected network. Similarly, a high-dimensional interaction vector is obtained through linear transformation. The linear transformation here is similar to that of the first layer. The value of each intermediate vector is multiplied by the corresponding weight, and then the results are added together with the bias value. However, the magnitude of the weight and bias value is different from that of the first layer. Then, the product of the corresponding values of any two feature dimensions in the high-dimensional interaction vector is calculated. This product is used as the initial correlation strength between the two feature dimensions. For example, the product of the corresponding values of the nonlinear feature sample entropy and the frequency domain feature high-frequency component in the high-dimensional interaction vector is calculated to obtain the initial correlation strength between the two. Then, the average value of the initial correlation strength of each feature dimension with all other feature dimensions is calculated. This average value is the overall correlation strength between the feature dimension and other features. The final distribution result can clearly reflect the mutual influence between various features, such as the correlation strength between the time domain feature heart rate and the frequency domain feature low-frequency component, and the correlation strength between the nonlinear feature sample entropy and the time domain feature RR interval standard deviation, etc.
[0053] Step 551: Based on the distribution of correlation strength between features, calculate the impact of each feature dimension in the structured emotion feature matrix on the emotion classification task. Specifically, this includes: first, collecting a large amount of historical emotion recognition data, which contains samples of different emotions such as calmness, depression, and anxiety, along with their corresponding correct classification results. In this historical data, the frequency of each feature dimension in correct classification is calculated. For example, in all samples correctly identified as anxiety, the proportion of times the sample entropy value is in the high range is the frequency of sample entropy in correct anxiety classification. Similarly, in samples correctly identified as depression, the proportion of times the sample entropy value is in the low range is the frequency of sample entropy in correct depression classification. The average of these two frequencies is then taken as the overall correct classification frequency of sample entropy. For other features... For dimensions like heart rate, the frequency of their correct classification across various emotions is statistically analyzed and averaged. Then, the maximum value is extracted from the distribution of correlation strength between each feature dimension and other feature dimensions. For example, the maximum value among the correlation strengths of sample entropy and all other features is used. This maximum value is multiplied by the frequency of correct classification for that feature dimension to obtain the first median value. Next, the magnitude of the change in the value of that feature dimension across all samples is calculated by subtracting the minimum value from the maximum value of that feature dimension across all samples. The result is added to the first median value. The sum is the assessment of the influence of that feature dimension on the emotion classification task. In this process, because sample entropy has a higher frequency of correct classification when distinguishing between anxiety and depression, and its correlation strength with other features is also stronger, its influence assessment will be higher than that of time-domain features such as heart rate, which perform similarly in the two emotions.
[0054] Step 552 involves normalizing the impact assessment to obtain the attention focus parameters for each feature dimension. Specifically, this involves summing the impact assessment values of all feature dimensions, then dividing the impact assessment value of each feature dimension by this sum. The result is the attention focus parameter for that feature dimension. This transformation ensures that the sum of the attention focus parameters for all feature dimensions is 1, intuitively reflecting the relative importance of each feature dimension within the overall feature set. For example, if the total impact assessment value of all feature dimensions is 300, and the impact assessment value of sample entropy is 60, then the attention focus parameter for sample entropy is 60 divided by 300. Similarly, if the impact assessment value of heart rate is 30, its attention focus parameter is 30 divided by 300. These parameters clearly demonstrate that sample entropy is more important than heart rate in emotion classification.
[0055] Step 553: Adjust the importance of the corresponding feature dimensions in the structured emotion feature matrix based on the attention focus parameter. Specifically, this includes: first, determining a base adjustment coefficient, which is obtained through multiple historical data validations. During the validation process, different coefficient values are tried to observe the effect of the adjusted features on emotion recognition. Finally, a coefficient value is selected that moderately enhances the key features without affecting the effectiveness of other features. For each feature dimension, each value of that dimension in the structured emotion feature matrix is multiplied by (1 plus the product of the attention focus parameter of that feature dimension and the base adjustment coefficient) to obtain the adjusted feature value. For example, if the attention focus parameter of the sample entropy is high, multiplying its value by (1 plus the product of the high parameter and the base adjustment coefficient) will result in a significant enhancement. Conversely, for feature dimensions with low attention focus parameters, multiplying their values by (1 plus the product of the low parameter and the base adjustment coefficient) will result in a smaller enhancement, thus highlighting the importance of high-parameter features. This makes key features that distinguish between anxiety and depression, such as sample entropy, more significant.
[0056] Step 554 integrates the importance-adjusted feature dimensions to generate an enhanced emotional feature representation. Specifically, this involves arranging all values of each adjusted feature dimension in rows according to the original order of time-domain, frequency-domain, and nonlinear features, forming a new feature matrix. This arrangement maintains the logical relationship between feature categories. During the arrangement process, each feature dimension's value is checked against a reasonable physiological range. For example, the normal range for heart rate is typically 60 to 100 beats per minute. If a value exceeds this range, it is compared to the average value of that feature dimension. If it is higher than the average, it is adjusted to the average plus twice the standard deviation of that feature dimension; if it is lower than the average, it is adjusted to the average minus twice the standard deviation of that feature dimension. This adjustment ensures that all values conform to the actual physiological signal patterns. The resulting enhanced emotional feature representation retains the multi-dimensional information from the original features, including time-domain and frequency-domain nonlinearity, while highlighting key features like sample entropy that effectively distinguish between anxiety and depression through importance adjustment. This provides a more accurate feature foundation for analyzing the time dependence of electrocardiogram signals using recurrent neural networks.
[0057] By using feature interaction calculation and impact assessment, the important role of nonlinear features such as sample entropy in distinguishing between anxiety and depression can be accurately captured. With the help of attention focusing parameters for targeted adjustment, the confusion between the two types of emotions can be reduced. The enhanced emotional feature representation is more in line with the physiological signal differences corresponding to different emotions, improves the accuracy of emotion recognition, avoids psychological teachers adopting guidance strategies that do not meet the needs of students due to misjudgment, and better assists universities in carrying out mental health monitoring work.
[0058] In a preferred embodiment of the present invention, step 6 above, which involves obtaining the time dependence of the electrocardiogram signal through a recurrent neural network in a deep learning model based on the enhanced emotion feature representation, to generate the final emotion recognition result, may include: In this embodiment of the invention, step 660 involves inputting the enhanced emotional feature representation into the time series processing layer of a recurrent neural network in chronological order. By processing the feature representations at each time point in sequence, the evolutionary pattern of features over time is extracted. Specifically, the time series processing layer adopts a long short-term memory network structure, which is designed specifically to capture the dynamic changes of time series data such as electrocardiogram signals. Its core consists of multiple long short-term memory units, and the number of units matches the number of feature dimensions in the enhanced emotional feature representation. For example, if the enhanced emotional feature representation contains 30 feature dimensions, then the number of long short-term memory units in the time series processing layer is set to 30, and each unit corresponds to processing the temporal changes of one feature dimension.
[0059] First, the enhanced emotional feature representation is divided into time series segments. Based on the ECG signal acquisition interval, the continuous feature data is divided into fixed-length time windows, each corresponding to a time series segment. For example, if the ECG signal is acquired at a frequency of 250 samples per second, and each 10-second time window is a segment, then each time window contains feature representations at 10 time points. Each time point's feature representation contains values across 30 feature dimensions, covering key information such as time-domain features (heart rate), frequency-domain features (high-frequency components), and non-linear features (sample entropy). The divided time... The sequence fragments are input into the time series processing layer in chronological order. The feature representations of each time point are sequentially entered into the corresponding long short-term memory units. For each long short-term memory unit, when processing the feature value at the current time point, it first receives the processing result of the previous time point for that feature dimension. The feature value at the current time point is then fused with the processing result of the previous time point. Specifically, the feature value at the current time point is multiplied by the weight value corresponding to the unit, and the processing result of the previous time point is multiplied by another set of weight values. The two multiplication results are added together, and then the bias value of the unit is added to obtain the fused intermediate result.
[0060] By sequentially processing the feature representations at all time points within each time window, the numerical change trend of each feature dimension at different time points is recorded, thereby extracting the feature evolution law that changes over time. For example, in the time series of anxiety, the value of the non-linear feature of sample entropy will gradually increase over time, while the value of the heart rate feature will show a continuous fluctuating upward evolution law; while in the time series of depression, the value of sample entropy will gradually decrease over time, while the value of the heart rate feature will show a slow upward evolution law followed by a stable evolution law.
[0061] Step 661: Based on the feature evolution law, the long-term feature dependencies across time periods are obtained through the state transfer process of the memory unit in the recurrent neural network, forming a feature representation including temporal information. Specifically, the memory unit of the recurrent neural network, namely the Long Short-Term Memory (LSTM), contains three core structures: a forget gate, an input gate, and an output gate. Each structure corresponds to an independent set of weights and biases, used to accurately filter and transfer the emotional feature information of the electrocardiogram signal in a temporal manner, adapting to the capture of the temporal evolution law of features such as sample entropy and heart rate under emotions such as anxiety and depression. First, the forget gate is processed. The forget gate receives the feature representation at the current time point and the hidden state of the memory unit at the previous time point. The feature representation at the current time point is multiplied by the first set of weights corresponding to the forget gate, and the hidden state at the previous time point is multiplied by the second set of weights corresponding to the forget gate. The two multiplication results are added together, plus the bias value of the forget gate, to obtain the intermediate calculation result of the forget gate. Then, the intermediate calculation result is input into the sigmoid activation function to obtain the output value of the forget gate. This output value is used to determine the... The sigmoid activation function determines which information in the previous memory unit needs to be retained and which needs to be forgotten. The specific calculation process involves: first, using the natural constant as the base, negating the intermediate calculation result of the forgetting gate to calculate the exponential decay value; then, dividing 1 by 1 and adding the exponential decay value, the final quotient is the output value of the sigmoid function. This output value always lies between 0 and 1. The closer the value is to 1, the higher the proportion of information retained; the closer the value is to 0, the higher the proportion of information forgotten. For example, regarding the characteristic evolution of continuously increasing sample entropy in anxiety, when the intermediate calculation result is a large positive number, the exponential decay value approaches 0, and the sigmoid output value is close to 1. In this case, a large amount of the trend information of increasing sample entropy in the previous moment will be retained. Conversely, for occasional interference information unrelated to emotional evolution, such as a sudden increase in heart rate, the intermediate calculation result will be a small negative number, and the sigmoid output value will be close to 0. This type of interference information will be effectively forgotten, ensuring that the memory unit focuses on the core emotional feature temporal information.
[0062] Next, the input gate is processed. The input gate also receives the feature representation at the current time point and the hidden state at the previous time point. First, the feature representation at the current time point is multiplied by the first set of weights corresponding to the input gate, and the hidden state at the previous time point is multiplied by the second set of weights corresponding to the input gate. These two multiplication results are added together, along with the input gate's bias value, to obtain the first intermediate result of the input gate. This result is input into the sigmoid activation function to obtain the update coefficient of the input gate. The calculation process of the sigmoid activation function is exactly the same as in the forgetting gate. Similarly, the first intermediate result is converted to a value between 0 and 1. This value serves as the update coefficient, controlling the proportion of new information entering the memory unit. When the value is close to 1, new feature information at the current time point will be prioritized for inclusion in the memory unit; when the value is close to 0, the contribution of the current new information to the emotional time-series pattern is small, and the proportion entering the memory unit will be significantly reduced. Second, the feature representation at the current time point is multiplied by another set of weights, and the hidden state at the previous time point is multiplied by yet another set of weights. The weights are multiplied, and the results are added together, along with the corresponding bias value, to obtain the second intermediate result of the input gate. This result is then input into the tanh activation function to obtain candidate update information. The specific calculation process of the tanh activation function is as follows: First, using the natural constant as the base, the exponent values of the second intermediate result and the exponent values of the opposite of the second intermediate result are calculated respectively. Then, the difference between the exponent value and the exponential decay value is subtracted, and the result is divided by the sum of the exponent value and the exponential decay value. The final quotient is the output value of the tanh function. This output value is always between -1 and 1, which can accurately reflect both the positive and negative changes of feature values, perfectly adapting to the opposite evolution trend of core features under different emotions, and providing comprehensive and accurate new feature information for the memory unit. Finally, the update coefficient is multiplied with the candidate update information to obtain the new information that needs to be updated in the memory unit. The update coefficient will proportionally distribute different parts of the candidate update information, strengthening new information that matches the emotional time sequence pattern and suppressing irrelevant information.
[0063] Next, the cell state of the memory unit is updated by multiplying the output value of the forget gate by the cell state of the previous time step to obtain the historical feature information that needs to be retained. This historical information is then added to the new information obtained from the input gate to obtain the cell state of the memory unit at the current time step, completing the cell state transmission and update. At this point, the cell state contains both the filtered historical emotional feature time sequence information and the regulated current new feature information, completely recording the emotional feature evolution trajectory from the start of the time window to the current time step. Finally, the output gate is processed by multiplying the feature representation of the current time step by the first set of weight values corresponding to the output gate, and multiplying the hidden state of the previous time step by the second set of weight values corresponding to the output gate. These two multiplication results are added together, along with the bias value of the output gate, to obtain the intermediate result of the output gate. This result is then input into the sigmoid activation function to obtain the output coefficients. Here, sigmoid is used to obtain the output coefficients. The calculation process of the id activation function remains the same as described above, outputting a value between 0 and 1 to control the proportion of information transferred from the current cell state to the hidden state. Simultaneously, the current cell state is input into the tanh activation function to obtain the cell state output information. The calculation process of the tanh activation function is completely consistent with that of the input gate, converting the cell state value to a range between -1 and 1. This conversion effectively compresses extreme values that may occur in the cell state due to feature enhancement, avoiding numerical anomalies from interfering with subsequent calculations, while fully preserving the temporal trend of the emotional features contained in the cell state. Finally, the output coefficient is multiplied by the cell state output information to obtain the hidden state of the memory unit at the current moment. This hidden state is a comprehensive representation that integrates the current time-point feature information and historical feature information, focusing on core emotional features while fully carrying temporal correlations.
[0064] Through the aforementioned memory unit state transfer process, all time points within each time window are processed sequentially, ensuring that the hidden state of each time point contains feature information from the start of that time window to the current time. This allows for the acquisition of long-term feature dependencies across time periods. For example, when processing anxiety data containing multiple consecutive time windows, it can continuously capture the long-term dependency of sample entropy increasing from slightly to significantly within multiple time windows, while clearly recording the correlation of heart rate fluctuations and increases across different time windows. When processing depressed data, it can capture the long-term temporal correlation of continuously decreasing sample entropy and a slow increase in heart rate followed by stabilization. By arranging the hidden states of all time points in chronological order, a feature representation containing complete temporal information is ultimately formed, laying the foundation for deep integration of emotional features and accurate emotion identification.
[0065] Step 662: The feature representation is input into the feature fusion layer. A multi-layer nonlinear transformation network is used to deeply integrate the temporal features and extract high-level emotional features with discriminative power. Specifically, the feature fusion layer consists of three fully connected networks, each containing several neurons. The number of neurons decreases sequentially from input to output to achieve gradual compression and deep integration of the temporal features. The number of neurons in the first fully connected network is set to twice the dimension of the feature representation containing temporal information; the number of neurons in the second fully connected network is set to half that of the first layer; and the number of neurons in the third fully connected network is set to half that of the second layer. This setting effectively reduces the feature dimension and improves subsequent processing efficiency while retaining key feature information. First, the feature representation containing temporal information is input into the first fully connected network. This feature representation is a sequence of hidden states arranged in chronological order. This sequence is first expanded into a one-dimensional vector row-wise and used as the input vector of the first fully connected network. For each value in the input vector, it is multiplied by the weight value of the corresponding neuron in the first fully connected network. All multiplication results are added together, along with the bias of the corresponding neuron. The intermediate output of the first fully connected network is obtained by taking the values and inputting them into an activation function. This intermediate output is then transformed using a nonlinear transformation to enhance the expressive power of the features, resulting in the output vector of the first fully connected network. The output vector of the first fully connected network is then input into the second fully connected network, and the above calculation process is repeated. Each value in the output vector is multiplied by the weight value of the corresponding neuron in the second layer. All multiplications are summed, a bias value is added, and the result is then input into the activation function for a nonlinear transformation, yielding the output vector of the second fully connected network. Next, the output vector of the second fully connected network is input into the third fully connected network, and the same process of multiplying and adding the values with the weights and the bias value is performed. This is then input into the activation function for a nonlinear transformation, yielding the output vector of the third fully connected network. This output vector represents the deeply integrated high-level emotion feature, which incorporates the temporal characteristics and multi-dimensional features of the electrocardiogram signal. It can highlight the key differences between different emotions. For example, for anxiety and depression, the high-level emotion feature clearly reflects the differences in the temporal change trend of sample entropy and the temporal differences in the correlation between heart rate and high-frequency components, providing discriminative information.
[0066] Step 663: Based on high-level emotion features, the feature space is mapped to the emotion category space through a feature dimension transformation layer to obtain preliminary discrimination scores for each emotion. Specifically, the feature dimension transformation layer is a fully connected network with the number of neurons matching the number of emotion categories to be identified. For example, if three emotions—calm, depressed, and anxious—need to be identified, the fully connected network would have three neurons, each corresponding to one emotion category and responsible for calculating the preliminary discrimination score for that category. First, the high-level emotion features are used as the input vector to the feature dimension transformation layer. The dimension of this input vector matches the dimension of the output vector from the feature fusion layer. For each value in the input vector, it is multiplied by the weight value corresponding to each neuron in the feature dimension transformation layer. For example, the first value in the input vector is multiplied by the weight value of the first neuron. One weight value is multiplied, the second value of the input vector is multiplied by the second weight value of the first neuron, and so on, adding all the multiplication results related to the neuron, plus the bias value corresponding to the neuron, to obtain the intermediate calculation result of the neuron. This result is the preliminary discrimination score of the corresponding emotion category. Following the above method, the intermediate calculation results of the three neurons are calculated separately to obtain the preliminary discrimination scores of calm emotion, depressed emotion, and anxious emotion. The magnitude of these scores reflects the probability of the current ECG signal sample corresponding to each emotion. For example, if a sample has a high preliminary discrimination score for anxious emotion, it means that the sample is more likely to belong to anxious emotion. However, the score at this time has not been normalized and does not have probabilistic meaning. It is only used as the basis for probability transformation.
[0067] Step 664: Input the preliminary discrimination scores into the probability transformation layer to generate the final emotion category probability distribution as the emotion recognition result. Specifically, the core function of the probability transformation layer is to convert the preliminary discrimination scores output by the feature dimension transformation layer into probabilistically meaningful values, making the discrimination results for various emotions more intuitive and interpretable. First, it receives all the preliminary discrimination scores output by the feature dimension transformation layer, assuming they are scores for calm, depressed, and anxious emotions. It then calculates the exponential values corresponding to these preliminary discrimination scores. Specifically, using the natural constant as the base, it performs exponential operations on each preliminary discrimination score to obtain the exponential value corresponding to each score. Then, it sums the exponential values of all emotion categories to obtain the exponential sum. Finally, it divides the exponential value of each emotion category by... The probability value for each emotion category is obtained by summing the exponents. For example, dividing the exponent value of calm emotion by the sum of exponents yields the probability value of calm emotion; dividing the exponent value of depressed emotion by the sum of exponents yields the probability value of depressed emotion; and dividing the exponent value of anxiety emotion by the sum of exponents yields the probability value of anxiety emotion. The sum of these probability values is 1, forming the final probability distribution of emotion categories. The emotion category with the highest probability value is the emotion recognition result corresponding to the current ECG signal sample. For example, if the probability value of anxiety emotion is the highest, it means that the emotion corresponding to the sample is anxiety. At the same time, the probability values of all emotion categories are retained as the complete output of the emotion recognition result, which helps psychological teachers understand the probability of the sample corresponding to each type of emotion and provides a more comprehensive reference for psychological intervention.
[0068] By fully capturing the temporal dependence of emotion-related features in electrocardiogram signals, the state transfer between the time series processing layer and the memory unit can completely preserve the long-term evolutionary patterns of features under different emotions. In particular, it can clearly capture the temporal differences of key features such as sample entropy in anxiety and depression, avoiding misjudgment caused by similar features at a single time point. The deep integration of the feature fusion layer organically combines temporal information with multi-dimensional features, and the extracted high-level emotional features are more discriminative and can accurately reflect the essential differences between different emotions.
[0069] like Figure 2 As shown, embodiments of the present invention also provide an emotion feature recognition system based on electrocardiogram signals, including: The acquisition and processing module is used to acquire raw electrocardiogram (ECG) signals and preprocess them to obtain preprocessed ECG signals. The feature extraction module is used to extract time-domain features, frequency-domain features, and nonlinear features related to emotional state based on preprocessed electrocardiogram signals, forming an initial feature set. The feature construction module is used to fuse multiple types of features in the initial feature set to construct a structured sentiment feature matrix that integrates spatial distribution and multi-dimensional feature information; The feature adjustment module is used to perform feature structure analysis on the structured emotion feature matrix. It constructs feature association regions by defining reference feature vectors and calculating the distribution relationship between multiple associated feature vectors and the reference feature vectors. Within the feature association regions, two feature reference points with temporal relationships are selected, and feature adjustment paths are formed based on the temporal connection relationship between the two feature reference points. Based on the feature adjustment paths, feature correlation metrics are generated, and the feature correlation metrics are used to adjust the structured emotion feature matrix to obtain the adjusted structured emotion feature matrix. The feature enhancement module is used to input the adjusted structured emotion feature matrix into the deep learning model, and adaptively enhance the key features related to emotion recognition in the structured emotion feature matrix through the attention mechanism in the deep learning model, thereby generating an enhanced emotion feature representation. The result generation module is used to generate the final emotion recognition result by obtaining the time dependence of electrocardiogram signals through a recurrent neural network in a deep learning model based on the enhanced emotion feature representation.
[0070] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0071] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0072] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0073] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for emotion feature recognition based on electrocardiogram signals, characterized in that, The method includes: Step 1: Acquire raw electrocardiogram (ECG) signals and preprocess them to obtain preprocessed ECG signals; Step 2: Based on the preprocessed electrocardiogram signal, extract time-domain features, frequency-domain features, and nonlinear features related to emotional state to form an initial feature set; Step 3: Fuse the multiple features in the initial feature set to construct a structured emotion feature matrix that integrates spatial distribution and multi-dimensional feature information; Step 4: Perform feature structure analysis on the structured emotion feature matrix. By defining a reference feature vector and calculating the distribution relationship between multiple associated feature vectors and the reference feature vector, a feature association region is constructed. Within the feature association region, two feature reference points with temporal relationship are selected respectively, and a feature adjustment path is formed based on the temporal connection relationship between the two feature reference points. A feature association metric is generated based on the feature adjustment path, and the feature association metric is used to adjust the structured emotion feature matrix to obtain the adjusted structured emotion feature matrix. Step 5: Input the adjusted structured emotion feature matrix into the deep learning model. The attention mechanism in the deep learning model adaptively enhances the key features related to emotion recognition in the structured emotion feature matrix, generating an enhanced emotion feature representation. Step 6: Based on the enhanced emotion feature representation, the time dependence of the electrocardiogram signal is obtained through the recurrent neural network in the deep learning model to generate the final emotion recognition result.
2. The emotion feature recognition method based on electrocardiogram signals according to claim 1, characterized in that, Based on preprocessed electrocardiogram signals, time-domain features, frequency-domain features, and nonlinear features related to emotional states are extracted to form an initial feature set, including: Based on the preprocessed electrocardiogram signal, time-domain features used to characterize the statistical properties of heart rate fluctuations are extracted to form a subset of time-domain features; Based on the preprocessed electrocardiogram signal, power spectrum analysis is performed on the heart beat cycle signal on which the time-domain feature subset is based to extract frequency domain features used to characterize the state of autonomic nervous activity, thus forming a frequency domain feature subset; Based on preprocessed ECG signals, the signal complexity of the time-domain feature subset and frequency-domain feature subset analysis basis is quantified, and nonlinear features used to characterize the nonlinear dynamics of ECG signals are extracted to form a nonlinear feature subset. The time-domain feature subset, frequency-domain feature subset, and nonlinear feature subset are integrated to form the initial feature set.
3. The emotion feature recognition method based on electrocardiogram signals according to claim 2, characterized in that, The multiple features in the initial feature set are fused to construct a structured sentiment feature matrix that integrates spatial distribution and multi-dimensional feature information, including: Feature values are extracted from the time-domain feature subset, frequency-domain feature subset, and nonlinear feature subset in the initial feature set to obtain the numerical set of each feature subset; Based on the numerical set, the time-domain feature subset, frequency-domain feature subset, and nonlinear feature subset are normalized respectively to obtain the standardized time-domain feature vector, frequency-domain feature vector, and nonlinear feature vector. The time-domain feature vector, frequency-domain feature vector, and nonlinear feature vector are concatenated in order according to the feature category dimension to form a multi-dimensional fused feature vector; Based on the multi-dimensional fusion feature vectors of multiple samples, the sample dimensions are arranged and combined to construct a multi-row, multi-column structured emotion feature matrix that represents feature samples in the row direction and feature types in the column direction. Spatial distribution feature enhancement is performed on the structured emotion feature matrix. By calculating the statistical relationship between feature columns, spatial distribution descriptors describing feature correlation are added, ultimately forming a structured emotion feature matrix that integrates spatial distribution and multi-dimensional feature information.
4. The emotion feature recognition method based on electrocardiogram signals according to claim 3, characterized in that, Feature structure analysis is performed on the structured sentiment feature matrix. This involves defining a reference feature vector and calculating the distribution relationship between multiple associated feature vectors and the reference feature vector to construct feature association regions, including: Feature samples with typical emotional representations are selected from the structured emotion feature matrix and defined as reference feature vectors; Based on the reference feature vector, calculate the Euclidean distance between all other feature samples in the structured emotion feature matrix and the reference feature vector in the feature space to obtain the distance distribution sequence; Statistical analysis is performed on the distance distribution sequence to determine the dynamic distance threshold used to divide the feature association range; Based on the dynamic distance threshold, a subset of feature samples that are closest to the reference feature vector are selected from the structured emotion feature matrix and defined as the associated feature vector set. The feature association region is constructed with the position of the reference feature vector in the feature space as the center and the dynamic distance threshold as the spatial radius.
5. The emotion feature recognition method based on electrocardiogram signals according to claim 4, characterized in that, Within the feature association region, two feature reference points with a temporal relationship are selected respectively, and a feature adjustment path is formed based on the temporal connection relationship between the two feature reference points, including: Within the feature association region, based on the timestamp information of the feature samples, two adjacent feature samples in time sequence are selected as the first feature reference point and the second feature reference point, respectively. Based on the coordinate positions of the first and second feature reference points in the feature space, calculate the spatial vector relationship between the two feature reference points; Based on the spatial vector relationship, determine the direction vector from the first feature reference point to the second feature reference point, and establish a linear connection relationship between the two feature reference points; Based on the linear connection relationship, a straight path is constructed extending from the first feature reference point to the second feature reference point, forming a feature adjustment path.
6. The emotion feature recognition method based on electrocardiogram signals according to claim 5, characterized in that, Based on the feature adjustment path, a feature correlation metric is generated, and the feature correlation metric is used to adjust the structured sentiment feature matrix, resulting in an adjusted structured sentiment feature matrix, including: Based on the spatial relationship between two feature reference points on the feature adjustment path, the change magnitude of each feature dimension along the path direction is calculated to form feature correlation measurement parameters; The feature correlation measurement parameters are normalized to obtain the relative contribution assessment of each feature dimension in emotion recognition. Based on the relative contribution assessment results, a proportional relationship for feature adjustment is established, and based on the proportional relationship, the feature dimensions in the original structured emotion feature matrix are adjusted differentially. The differentiated feature matrix is standardized and integrated to form a structured emotion feature matrix after feature distribution adjustment.
7. The emotion feature recognition method based on electrocardiogram signals according to claim 6, characterized in that, The adjusted structured emotion feature matrix is input into a deep learning model. The attention mechanism within the deep learning model adaptively enhances key emotion-recognition-related features in the structured emotion feature matrix, generating an enhanced emotion feature representation, including: The adjusted structured emotion feature matrix is input into the feature analysis layer of the deep learning model, and the distribution of the correlation strength between features is obtained through feature interaction calculation. Based on the distribution of correlation strength between features, the influence of each feature dimension in the structured emotion feature matrix on the emotion classification task is evaluated. The impact assessment is normalized to obtain the attention focus parameters for each feature dimension; Based on the attention focus parameters, the importance of the corresponding feature dimensions in the structured emotion feature matrix is adjusted; The importance-adjusted feature dimensions are integrated to generate an enhanced emotional feature representation.
8. The emotion feature recognition method based on electrocardiogram signals according to claim 7, characterized in that, Based on the enhanced emotion feature representation, the time dependence of electrocardiogram signals is obtained through a recurrent neural network in a deep learning model to generate the final emotion recognition result, including: The enhanced emotional feature representations are input into the time series processing layer of the recurrent neural network in chronological order. By processing the feature representations at each time point in turn, the evolutionary pattern of features over time is extracted. Based on the feature evolution law, the long-term feature dependencies across time periods are obtained through the state transfer process of the memory unit of the recurrent neural network, forming a feature representation that includes time-series information. The feature representation is input into the feature fusion layer, and the temporal features are deeply integrated through a multi-layer nonlinear transformation network to extract high-level emotion features with discriminative power. Based on high-level emotion features, the feature space is mapped to the emotion category space through a feature dimension transformation layer to obtain preliminary discrimination scores for various emotions; The initial discrimination score is input into the probability transformation layer to generate the final emotion category probability distribution, which serves as the emotion recognition result.
9. An emotion feature recognition system based on electrocardiogram signals, wherein the system implements the method as described in any one of claims 1 to 8, characterized in that, include: The acquisition and processing module is used to acquire raw electrocardiogram (ECG) signals and preprocess them to obtain preprocessed ECG signals. The feature extraction module is used to extract time-domain features, frequency-domain features, and nonlinear features related to emotional state based on preprocessed electrocardiogram signals, forming an initial feature set. The feature construction module is used to fuse multiple types of features in the initial feature set to construct a structured sentiment feature matrix that integrates spatial distribution and multi-dimensional feature information; The feature adjustment module is used to perform feature structure analysis on the structured emotion feature matrix. It constructs feature association regions by defining a reference feature vector and calculating the distribution relationship between multiple associated feature vectors and the reference feature vector. Within the feature association region, two feature reference points with temporal relationship are selected respectively, and feature adjustment paths are formed based on the temporal connection relationship between the two feature reference points; The feature correlation metric is generated based on the feature adjustment path, and the feature correlation metric is used to adjust the structured emotion feature matrix to obtain the adjusted structured emotion feature matrix. The feature enhancement module is used to input the adjusted structured emotion feature matrix into the deep learning model, and adaptively enhance the key features related to emotion recognition in the structured emotion feature matrix through the attention mechanism in the deep learning model, so as to generate an enhanced emotion feature representation. The result generation module is used to generate the final emotion recognition result by obtaining the time dependence of electrocardiogram signals through a recurrent neural network in a deep learning model based on the enhanced emotion feature representation.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Multi-modal signal fusion emotion recognition method based on attention mechanism
CN120477781A
Multi-modal emotion recognition method for depression tendency of teenagers
CN120570609A
Method for realizing a multi-channel convolutional recurrent neural network EEG emotion recognition model using transfer learning
US20230039900A1