Emotion state quantification analysis system based on physiological parameter fusion
Patent Information
- Application Number
- CN202610986522.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的目的是提供基于生理参数融合的情绪状态量化分析系统,解决信号模态单一、情绪表征片面的问题,并通过RBM无监督融合摆脱对大量情绪标签的依赖,利用海量无标注数据训练,降低标注成本,有效缓解小样本场景下模型过拟合的技术难题
[0014]本发明的有益效果为:全模态覆盖,表征全面:融合 8 类信号,覆盖中枢、自主神经与行为维度,有效解决单一模态表征片面、可靠性低的问题;
Smart Images

Figure CN122805267A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a quantitative analysis system for emotional states based on the fusion of physiological parameters. Background Technology
[0002] The objective and accurate quantification of emotional states is a core technological requirement in fields such as human-computer interaction, mental health monitoring, intelligent healthcare, and vehicle safety. Traditional emotion recognition relies heavily on single behavioral signals such as facial expressions and voice, which are easily affected by faking and environmental interference, resulting in high subjectivity and low reliability. With the development of physiological signal acquisition technology, physiological signals such as electroencephalography (EEG), electrocardiography (ECG), and conductance of skin (GSR) have become the mainstream carriers for emotion quantification research because they can directly reflect the activity of the central and autonomic nervous systems, possessing advantages such as high objectivity, difficulty in faking, and high responsiveness.
[0003] Currently, emotion analysis technologies based on physiological signals are mainly divided into two categories: one is single-modal physiological signal analysis, which uses only single signals such as EEG and ECG to extract features and achieves emotion recognition through traditional machine learning (SVM, LDA) or simple neural networks; the other is multimodal fusion analysis, which focuses on 2-3 modalities such as EEG + speech or EEG + facial expressions and uses shallow fusion methods such as decision-level voting and simple splicing.
[0004] While existing technologies have achieved preliminary emotion recognition, they suffer from significant shortcomings in areas such as signal coverage dimensions, feature fusion effectiveness, quantification accuracy, generalization ability, and engineering practicality, making it difficult to meet the demands of high-precision, real-time, and scalable emotion quantification scenarios. Therefore, this invention proposes an emotion state quantification analysis system based on physiological parameter fusion, specifically addressing the deficiencies of existing technologies. Summary of the Invention
[0005] The purpose of this invention is to provide an emotion state quantification analysis system based on physiological parameter fusion, which solves the problems of single signal modality and one-sided emotion representation. It also eliminates the dependence on a large number of emotion labels by using RBM unsupervised fusion, utilizes massive unlabeled data for training, reduces labeling costs, and effectively alleviates the technical problem of model overfitting in small sample scenarios.
[0006] To achieve the above objectives, the present invention provides the following solution: A quantitative analysis system for emotion states based on physiological parameter fusion includes: The multimodal signal acquisition module is used to acquire multimodal signals. The data preprocessing and feature extraction module is used to perform data preprocessing and feature extraction on the multimodal signal to obtain several single-modal emotion feature vectors; The multimodal feature fusion module is used to concatenate the single-modal emotion feature vectors into an original feature matrix and perform global standardization processing, then perform unsupervised iterative training on the Restricted Boltzmann Machine (RBM). The multimodal concatenated feature matrix to be predicted is input into the trained RBM, and the emotion representation fusion feature is output. The emotion quantification analysis module is used to input the fusion features of the emotion representation into the emotion quantification model and output the valence, arousal and discrete emotion category of the emotion. The emotion quantification model is constructed by particle swarm optimization support vector machine or lightweight regression network.
[0007] Optionally, the multimodal signals include electroencephalogram (EEG), electrocardiogram (ECG), electrodermal conductance, respiration, body temperature, facial expression, speech, and eye movement signals.
[0008] Optionally, data preprocessing and feature extraction of the multimodal signals include: Each type of signal is cleaned, denoised, and normalized separately. Extract α / β / θ / δ / γ band power, power ratio, asymmetry, approximate entropy, sample entropy, fractal dimension, and LZC complexity from preprocessed EEG signals; Extract the mean, standard deviation, RMSSD, pNN20, and pNN50 of heart rate variability from the preprocessed ECG signals; The mean, variance, rise time, number of peaks, and frequency characteristics of the preprocessed electrodermal signals were extracted. The frequency, amplitude, standard deviation of respiratory interval, and rate of change of respiratory depth were extracted from the preprocessed respiratory signals. The mean, rate of change, and short-term fluctuations of the preprocessed body temperature signal were extracted. The distance between key points, AU activation intensity, and probability of expression category were extracted from the preprocessed facial expressions. The time-domain energy, zero-crossing rate, fundamental frequency, frequency-domain Mel-frequency cepstral coefficients, nonlinear fractal dimension, and Hurst exponent are extracted from the preprocessed speech signal. The fixation percentage, saccade amplitude, blink frequency, pupil mean and variance were extracted from the preprocessed eye movement signals.
[0009] Optionally, concatenating the single-modal sentiment feature vectors into the original feature matrix and performing global standardization includes: The single-modal feature vectors after dimensionality reduction by principal component analysis are horizontally concatenated along the sample dimension to construct the original feature matrix. The global mean of all samples in the same dimension is subtracted from each feature, and then divided by the global standard deviation of all samples in the same dimension to obtain the standardized feature matrix.
[0010] Optionally, unsupervised iterative training of the Restricted Boltzmann Machine (RBM) includes: Based on the standardized feature matrix, the Restricted Boltzmann Machine (RBM) is trained using the contrastive divergence algorithm in an unsupervised iterative manner, including forward encoding of real samples in positive phase, sampling and reconstructing samples in negative phase, gradient calculation and parameter updating, until the network converges.
[0011] Optionally, the forward encoding of the positive-phase true samples and the reconstructed samples from the negative-phase samples include: Input the standardized feature matrix, calculate the hidden layer activation probability, and perform binary sampling on the hidden layer to obtain the true hidden layer probability; Based on the true hidden layer probability, the mean of the visible layer is reconstructed in reverse. The hidden layer is then updated after sampling the mean of the visible layer to obtain the reconstructed hidden layer probability.
[0012] Optionally, the gradient calculation and parameter update include: The difference between the true hidden layer probability and the reconstructed hidden layer probability is calculated dimension by dimension, and the weight matrix, visible layer bias, hidden layer bias and adaptive noise standard deviation are updated.
[0013] Optionally, the trained Restricted Boltzmann Machine (RBM) outputs emotion representation fusion features including: The multimodal splicing feature matrix to be predicted is standardized. The standardized input features are divided by the noise standard deviation in each dimension, and matrix multiplication is performed with the trained weight matrix. The trained hidden layer bias vector is superimposed, and the hidden layer activation probability matrix is output by the sigmoid activation function. The hidden layer activation probability matrix is used as the emotion representation fusion feature.
[0014] The beneficial effects of this invention are: full modality coverage and comprehensive representation: it integrates 8 types of signals, covering the central nervous system, autonomic nervous system and behavioral dimensions, and effectively solves the problems of one-sided and low reliability of single modality representation; Deep unsupervised fusion with high information utilization: The RBM feature layer unsupervised fusion is adopted to explore the intrinsic correlation across modalities, remove redundant information, and improve fusion efficiency and generalization ability; Reduced labeling dependence and adaptation to small samples: RBM does not require sentiment labels and can be trained using massive amounts of unlabeled data, alleviating the problem of overfitting in small samples and reducing the cost of engineering implementation; Continuous quantitative output with higher precision: Constructing a two-dimensional quantitative system of valence and arousal to achieve a breakthrough in the quantitative analysis of emotions from qualitative classification; Modular design with strong scalability: Each module is independently pluggable, making it easy to replace algorithms, adapt to multiple devices and deployment scenarios, and highly practical. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a framework diagram of the emotional state quantification analysis system based on physiological parameter fusion according to an embodiment of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 As shown, this embodiment proposes a quantitative analysis system for emotional states based on the fusion of physiological parameters, including: The multimodal signal acquisition module is used to acquire multimodal signals. The data preprocessing and feature extraction module is used to perform data preprocessing and feature extraction on the multimodal signal to obtain several single-modal emotion feature vectors; The multimodal feature fusion module is used to concatenate the single-modal emotion feature vectors into an original feature matrix and perform global standardization processing, then perform unsupervised iterative training on the Restricted Boltzmann Machine (RBM). The multimodal concatenated feature matrix to be predicted is input into the trained RBM, and the emotion representation fusion feature is output. The emotion quantification analysis module is used to input the fusion features of the emotion representation into the emotion quantification model and output the valence, arousal and discrete emotion category of the emotion. The emotion quantification model is constructed by particle swarm optimization support vector machine or lightweight regression network. Results output, storage and visualization module: Displays sentiment quantification results in real time, saves raw data, feature data and analysis reports, and supports data backtracking and visualization.
[0020] Furthermore, the multimodal signals include electroencephalogram (EEG), electrocardiogram (ECG), electroskin conductance, respiration, body temperature, facial expression, speech, and eye movement signals.
[0021] Specifically, the data collection (after expansion) synchronously collects the following eight types of signals: 1. Electroencephalography (EEG): Reflects central nervous system activity and is used for core emotional characteristics. 2. Electrocardiography (ECG): Heart rate and heart rate variability (HRV). 3. Gestational skin conductance (GSR): Skin conductance level and skin conductance response. 4. Respiration (RESP): Respiratory rate, respiratory amplitude, and respiratory variability. 5. Body temperature (TEMP): Body surface temperature and its rate of change. 6. Facial expression: Key point coordinates, autonomic aggregator (AU) action units, and expression intensity. 7. Speech signal: Fundamental frequency, energy, speech rate, Mel features, and nonlinear features. 8. Eye movement signal: Fixation duration, saccade speed, blink frequency, and pupil diameter.
[0022] Furthermore, the data preprocessing and feature extraction of the multimodal signals include: Each type of signal is cleaned, denoised, and normalized separately. Extract α / β / θ / δ / γ band power, power ratio, asymmetry, approximate entropy, sample entropy, fractal dimension, and LZC complexity from preprocessed EEG signals; Extract the mean, standard deviation, RMSSD, pNN20, and pNN50 of heart rate variability from the preprocessed ECG signals; The mean, variance, rise time, number of peaks, and frequency characteristics of the preprocessed electrodermal signals were extracted. The frequency, amplitude, standard deviation of respiratory interval, and rate of change of respiratory depth were extracted from the preprocessed respiratory signals. The mean, rate of change, and short-term fluctuations of the preprocessed body temperature signal were extracted. The distance between key points, AU activation intensity, and probability of expression category were extracted from the preprocessed facial expressions. The time-domain energy, zero-crossing rate, fundamental frequency, frequency-domain Mel-frequency cepstral coefficients, nonlinear fractal dimension, and Hurst exponent are extracted from the preprocessed speech signal. The fixation percentage, saccade amplitude, blink frequency, pupil mean and variance were extracted from the preprocessed eye movement signals.
[0023] Furthermore, concatenating the single-modal sentiment feature vectors into the original feature matrix and performing global standardization includes: The single-modal feature vectors after dimensionality reduction by principal component analysis are horizontally concatenated along the sample dimension to construct the original feature matrix. The global mean of all samples in the same dimension is subtracted from each feature, and then divided by the global standard deviation of all samples in the same dimension to obtain the standardized feature matrix.
[0024] Specifically, the energy function of a restricted Boltzmann machine is: ; in, v For visible layer units,h For hidden layer units, D This represents the number of visible layer cells. K This represents the number of hidden layer units. As weight, For visible layer bias, For hidden layer bias. For the first i The standard deviation of Gaussian noise for each visible cell.
[0025] Hidden unit activation probability: ; Visible layer cell reconstruction distribution: ; in, For the sigmoid function, It follows a Gaussian distribution.
[0026] Furthermore, unsupervised iterative training of Restricted Boltzmann Machines (RBMs) includes: Based on the standardized feature matrix, the Restricted Boltzmann Machine (RBM) is trained using the contrastive divergence algorithm in an unsupervised iterative manner, including forward encoding of real samples in positive phase, sampling and reconstructing samples in negative phase, gradient calculation and parameter updating, until the network converges.
[0027] Furthermore, the forward encoding of the positive-phase true samples and the reconstructed samples from the negative-phase samples include: Input the standardized feature matrix, calculate the hidden layer activation probability, and perform binary sampling on the hidden layer to obtain the true hidden layer probability; Based on the true hidden layer probability, the mean of the visible layer is reconstructed in reverse. The hidden layer is then updated after sampling the mean of the visible layer to obtain the reconstructed hidden layer probability.
[0028] Furthermore, the gradient calculation and parameter update include: The difference between the true hidden layer probability and the reconstructed hidden layer probability is calculated dimension by dimension, and the weight matrix, visible layer bias, hidden layer bias and adaptive noise standard deviation are updated.
[0029] Furthermore, the trained Restricted Boltzmann Machine (RBM) outputs emotion representation fusion features including: The multimodal splicing feature matrix to be predicted is standardized. The standardized input features are divided by the noise standard deviation in each dimension, and matrix multiplication is performed with the trained weight matrix. The trained hidden layer bias vector is superimposed, and the hidden layer activation probability matrix is output by the sigmoid activation function. The hidden layer activation probability matrix is used as the emotion representation fusion feature.
[0030] Specifically, the training process includes the following steps: Phase 1: Data Preparation and Initialization Step 1: The low-dimensional features after dimensionality reduction by 8-way single-modal PCA are concatenated and standardized for preprocessing. At the same time, all learnable parameters of the RBM network are initialized to generate a dataset and initial network weights that conform to the Gaussian-Bernoulli RBM input specification. This is a necessary pre-step for iterative training. Without the output of this stage, the second-stage loop training cannot be carried out.
[0031] The multimodal original feature matrix is constructed by horizontally concatenating the feature vectors of EEG, ECG, GSR, RESP, TEMP, facial expression, speech, and eye tracking after dimensionality reduction by single-modal PCA in the sample dimension to generate a global multimodal feature matrix. The number of rows in the matrix is equal to the total number of samples N, and the number of columns is fixed at 75, corresponding to the total feature dimension after concatenation of the 8 modalities.
[0032] Constructing the original feature matrix of the multimodal mode: ; N is the number of samples.
[0033] Step 2: Global Standardization Process (Adapting to the Gaussian Unit Distribution of the Visible Layer) Standardize the 75-dimensional feature matrix after splicing dimension by dimension: Subtract the global mean of all samples in that dimension from each feature dimension, and then divide by the global standard deviation of all samples in that dimension. After processing, all feature dimensions satisfy the zero mean and unit variance distribution, matching the distribution assumption of the Gaussian input units in the visible layer of the RBM.
[0034] Global normalization (for visible layer Gaussian cells): ; Ensure that each dimension has zero mean and unit variance.
[0035] Step 3: RBM network parameter initialization: Define network dimensions: number of visible layer units D=75 (corresponding to the dimension of concatenated features), number of hidden layer units K=40 (dimension of final fused feature output); Weight matrix W: Generate a 75-row, 40-column random normal distribution matrix, and reduce the values to 0.01 times for small weight initialization to avoid gradient explosion caused by excessively large initial weights; Visible layer bias vector a: generates a one-dimensional vector of length 75 with all values being 0; Hidden layer bias vector b: generates a one-dimensional vector of length 40 with all values being 0; Gaussian noise standard deviation vector sigma: Generates a one-dimensional vector of length 75, with all values initialized to 1, used to constrain the variance of the Gaussian distribution in the visible layer.
[0036] Phase Two: Contrast Divergence (CD-k) Training: Using the standardized batch samples from Phase 1 and the initialized network parameters as input, 500 iterations are performed. In each iteration, two symmetric computation branches are split: positive phase (encoding of real samples) and negative phase (decoding of sample reconstruction). The gradient is calculated using the error between real samples and reconstructed samples, and all network parameters are updated in reverse to continuously reduce the multimodal feature reconstruction error until the network converges. This phase is the core link for RBM to learn the joint distribution among 8 modalities and mine modal complementary information. The specific training parameters and calculation methods provided in the revised draft are retained below.
[0037] The learning rate is set to 0.001, the batch size for a single training session is 32, the total number of iterations is 500, and the number of Gibbs sampling steps is k=1 (CD-1 is a one-step comparison of divergence to simplify sampling, balancing training speed and fitting accuracy).
[0038] Batch sample reading divides the standardized global dataset into batches of fixed size 32 samples, and feeds each batch into the network to complete one round of gradient update.
[0039] Positive Phase: Forward Encoding of Real Samples (Forward Propagation, Extracting Real Feature Associations) ① Input the current batch standardized sample matrix v0, with the number of rows equal to the batch size of 32 and the number of columns of 75; ② Calculate the activation probability of hidden layer units: First, divide the input sample dimension by the noise standard deviation sigma of the corresponding dimension, then perform matrix multiplication with the weight matrix W, superimpose the hidden layer bias vector b, and finally input the whole into the sigmoid activation function to obtain a 32-row, 40-column hidden layer activation probability matrix h0_prob; ③ Binary Sampling of Hidden Layers: Generate a random matrix of the same size (0~1), compare the hidden layer activation probability matrix with the random matrix element by element, mark the positions with probabilities greater than the random values as 1, and the rest as 0, to obtain the binary discrete hidden layer state matrix h0.
[0040] Negative Phase: Gibbs Sampling Reconstructs Samples (Reverse Decoding, Generating Reconstructed Pseudo-Samples) ① Reverse Reconstructs the Visible Layer Mean Based on the Discrete Hidden Layer h0 Obtained by the Positive Phase: The hidden layer state matrix is transposed and multiplied by the weight matrix W, and the visible layer bias vector a is superimposed. The result is multiplied by the corresponding noise standard deviation sigma in each dimension to obtain the Gaussian distribution mean matrix v1_mean of the reconstructed sample; ② Gaussian Sampling Generates Reconstructed Sample v1: Based on the mean matrix v1_mean, Gaussian random noise of the same dimension and standard deviation sigma is superimposed to obtain a 32-row, 75-column reconstructed sample matrix v1; ③ Secondary Calculation of Hidden Layer Probability Corresponding to the Reconstructed Sample: The reconstructed sample v1 is divided by the noise standard deviation sigma and multiplied by the weight matrix W, and the hidden layer bias b is superimposed. After sigmoid activation, the hidden layer probability matrix h1_prob corresponding to the reconstructed sample is obtained.
[0041] Gradient Calculation and Parameter Update (Error Backward Correction) ① Weight Gradient dW: The correlation matrix between the real sample and the real hidden layer probability is subtracted from the correlation matrix between the reconstructed sample and the reconstructed hidden layer probability. The result is divided by the batch sample size of 32 to obtain the weight update gradient. ② Visible Layer Bias Gradient da: The difference between the real sample v0 and the reconstructed sample v1 is calculated dimension by dimension, and then the average value is taken over the batch dimension. ③ Hidden Layer Bias Gradient db: The difference between the real hidden layer probability h0_prob and the reconstructed hidden layer probability h1_prob is calculated dimension by dimension, and then the average value is taken over the batch dimension. ④ Parameter Iterative Update: All parameters are equal to the original parameters plus the product of the learning rate and the corresponding gradient. The weight matrix W, visible layer bias a, and hidden layer bias b are updated synchronously. ⑤ Adaptive Noise Standard Deviation Update: The square of the error between the real sample v0 and the reconstructed mean v1_mean is calculated dimension by dimension. The average value is taken over the batch dimension and the square root is taken to update the sigma vector. The noise is adaptively matched to the features of each dimension for reconstruction. After each iteration of the reconstruction error monitoring, all standardized samples are reconstructed using the current network parameters. The global average of the squared element-wise errors between the original standardized samples and the reconstructed samples is calculated as the reconstruction loss index, which is used to judge the degree of network convergence.
[0042] Phase 3: Feature Extraction After all iterations are completed in Phase 2 and the network parameters converge, the system exits the training loop, constructs a fixed inference function, inputs any 8 modal standardized features, and directly outputs a 40-dimensional unified unsupervised emotion fusion representation. The output features can be fed into the backend PSO-SVM or regression network to complete the valence, arousal, and discrete emotion quantification. This is the final business output of the entire RBM fusion module.
[0043] 1. Input data is standardized to the multimodal spliced feature matrix to be predicted. The global mean and global standard deviation saved in the first training stage are reused to complete the standardization according to the same rules, ensuring that the data distribution is consistent with the training set. 2. The hidden layer fusion feature calculation divides the standardized input features by the noise standard deviation sigma in each dimension, performs matrix multiplication with the weight matrix W completed in stage two training, and superimposes the trained hidden layer bias vector b. The whole input is the sigmoid activation function, and the output is the hidden layer activation probability matrix. 3. The continuous fusion feature output does not perform hidden layer binary sampling. Instead, the continuous probability value output by sigmoid is directly used as the final fusion feature. The number of rows in the matrix is equal to the number of input samples, and the number of columns is fixed at 40 dimensions. This 40-dimensional vector is the unified emotion representation that captures the joint distribution of the 8 modalities of central nervous system, autonomic nervous system, and peripheral behavior. It is then input into the backend emotion quantification model to complete the prediction.
[0044] After training, the fusion features are extracted.
[0045] Example 1: Applying a physiological parameter-based emotional state quantification analysis system to the real-time quantification analysis of emotional states in healthy adults includes: 1. Target audience: One healthy adult male aged 25 years was selected, with no history of mental illness or sleep disorders, and no strenuous exercise or alcohol / caffeine intake in the 24 hours prior to the experiment.
[0046] 2. Experimental Environment and Equipment: Environment: A quiet, low-light, soundproof laboratory with a temperature of 24℃ and humidity of 50%; Data acquisition equipment: EEG: 32-channel dry electrode EEG headband, sampling rate 250Hz; ECG: Chest patch ECG sensor, sampling rate 1000Hz; Conductivity for skin: Finger clip-on GSR sensor, sampling rate 100Hz; Respiration: Abdominal breathing band, sampling rate 50Hz; Body temperature: Ear temperature sensor, sampling rate 1Hz; Facial expressions: High-definition camera (30fps); Voice: Noise-canceling microphone (16kHz); Eye movement: Desktop eye tracker, sampling rate 120Hz.
[0047] 3. Experimental Procedure: Step 1: Data Collection Subjects sat quietly for 5 minutes (baseline), and then watched videos that evoked four types of emotions in turn: calm, joy, sadness, and anger (2 minutes for each type). Eight types of signals were collected synchronously throughout the process, for a total of 40 minutes of data collection.
[0048] Step 2: Preprocessing and Feature Extraction Process each type of signal separately: EEG: Removes electrooculography artifacts, filters 0.5–45Hz, extracts 20-dimensional features such as α / β / θ / δ / γ power and approximate entropy, and reduces the dimensionality to 8 dimensions using PCA; ECG: Extract 15-dimensional features such as HRV mean, RMSSD, and pNN50, and reduce the dimensionality to 6 dimensions using PCA; GSR: Extracts 10-dimensional features such as mean, variance, and number of peaks, and reduces the dimensionality to 5 dimensions using PCA; RESP: Extracts 8-dimensional features such as frequency, amplitude, and standard deviation of interval, and reduces the dimensionality to 4 dimensions using PCA; TEMP: Extracts 3-dimensional features such as mean, rate of change, and short-term fluctuations, and reduces the dimensionality to 2-dimensional using PCA; Facial expressions: 68 key points and AU activation intensity, etc., were extracted, and PCA was used to reduce the dimensionality to 6 dimensions; Speech: Extract 18-dimensional features including MFCC, fundamental frequency, and energy, and reduce the dimensionality to 7 dimensions using PCA; Eye movement: 10-dimensional features such as fixation percentage and pupil diameter were extracted, and PCA was used to reduce the dimensionality to 5 dimensions.
[0049] The final result is an 8-way low-dimensional single-modal feature vector.
[0050] Step 3: RBM Unsupervised Feature Fusion: The 8-way features are concatenated into an original feature matrix (number of samples × 43), which is then input into the RBM model: RBM structure: 43 nodes in the visible layer and 20 nodes in the hidden layer; Training: The contrastive divergence (CD-k) algorithm was used for 50 unsupervised training rounds; Output: 20-dimensional unified emotion fusion feature.
[0051] Step 4: Sentiment Quantification Prediction: Input the fused features into the PSO-SVM model: Training: Fine-tuning the model using labeled emotional video data; Example of output: Calm video: Valence = 0.12, Arousal = 0.25 → Category: Calm; Pleasant Video: Valence = 0.78, Arousal = 0.62 → Category: Pleasant; Sadness Video: Valence = -0.65, Arousal = 0.38 → Category: Sadness; Angry video: Valence = -0.72, Arousal = 0.85 → Category: Tension / Anger.
[0052] Step 5: Output and Storage of Results The system displays the valence-wakefulness two-dimensional curve in real time, and simultaneously saves the original data, feature data and quantitative reports, supporting subsequent retrospective analysis.
[0053] 4. Implementation Results: In this case, the system achieved an average recognition accuracy of 92.3% for the four types of emotions, with a valence / arousal quantification error of less than 0.08, which is 7.5% higher than the traditional two-modal fusion method, verifying the effectiveness and reliability of the present invention.
[0054] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A quantitative analysis system for emotional states based on the fusion of physiological parameters, characterized in that, include: The multimodal signal acquisition module is used to acquire multimodal signals. The data preprocessing and feature extraction module is used to perform data preprocessing and feature extraction on the multimodal signal to obtain several single-modal emotion feature vectors; The multimodal feature fusion module is used to concatenate the single-modal emotion feature vectors into an original feature matrix and perform global standardization processing, then perform unsupervised iterative training on the Restricted Boltzmann Machine (RBM). The multimodal concatenated feature matrix to be predicted is input into the trained RBM, and the emotion representation fusion feature is output. The emotion quantification analysis module is used to input the fusion features of the emotion representation into the emotion quantification model and output the valence, arousal and discrete emotion category of the emotion. The emotion quantification model is constructed by particle swarm optimization support vector machine or lightweight regression network.
2. The emotional state quantitative analysis system based on physiological parameter fusion according to claim 1, characterized in that, The multimodal signals include electroencephalogram (EEG), electrocardiogram (ECG), skin conductance, respiration, body temperature, facial expression, speech, and eye movement signals.
3. The emotional state quantitative analysis system based on physiological parameter fusion according to claim 2, characterized in that, Data preprocessing and feature extraction of the multimodal signals include: Each type of signal is cleaned, denoised, and normalized separately. Extract α / β / θ / δ / γ band power, power ratio, asymmetry, approximate entropy, sample entropy, fractal dimension, and LZC complexity from preprocessed EEG signals; Extract the mean, standard deviation, RMSSD, pNN20, and pNN50 of heart rate variability from the preprocessed ECG signals; The mean, variance, rise time, number of peaks, and frequency characteristics of the preprocessed electrodermal signals were extracted. The frequency, amplitude, standard deviation of respiratory interval, and rate of change of respiratory depth were extracted from the preprocessed respiratory signals. The mean, rate of change, and short-term fluctuations of the preprocessed body temperature signal were extracted. The distance between key points, AU activation intensity, and probability of expression category were extracted from the preprocessed facial expressions. The time-domain energy, zero-crossing rate, fundamental frequency, frequency-domain Mel-frequency cepstral coefficients, nonlinear fractal dimension, and Hurst exponent are extracted from the preprocessed speech signal. The fixation percentage, saccade amplitude, blink frequency, pupil mean and variance were extracted from the preprocessed eye movement signals.
4. The emotional state quantitative analysis system based on physiological parameter fusion according to claim 1, characterized in that, The process of concatenating the single-modal sentiment feature vectors into an original feature matrix and then performing global standardization includes: The single-modal feature vectors after dimensionality reduction by principal component analysis are horizontally concatenated along the sample dimension to construct the original feature matrix. The global mean of all samples in the same dimension is subtracted from each feature, and then divided by the global standard deviation of all samples in the same dimension to obtain the standardized feature matrix.
5. The emotional state quantitative analysis system based on physiological parameter fusion according to claim 4, characterized in that, Unsupervised iterative training of Restricted Boltzmann Machines (RBMs) includes: Based on the standardized feature matrix, the Restricted Boltzmann Machine (RBM) is trained using the contrastive divergence algorithm in an unsupervised iterative manner, including forward encoding of real samples in positive phase, sampling and reconstructing samples in negative phase, gradient calculation and parameter updating, until the network converges.
6. The emotional state quantitative analysis system based on physiological parameter fusion according to claim 5, characterized in that, The forward encoding of the positive-phase true samples and the reconstructed samples from the negative-phase sampling include: Input the standardized feature matrix, calculate the hidden layer activation probability, and perform binary sampling on the hidden layer to obtain the true hidden layer probability; Based on the true hidden layer probability, the mean of the visible layer is reconstructed in reverse. The hidden layer is then updated after sampling the mean of the visible layer to obtain the reconstructed hidden layer probability.
7. The emotional state quantitative analysis system based on physiological parameter fusion according to claim 6, characterized in that, The gradient calculation and parameter update include: The difference between the true hidden layer probability and the reconstructed hidden layer probability is calculated dimension by dimension, and the weight matrix, visible layer bias, hidden layer bias and adaptive noise standard deviation are updated.
8. The emotional state quantitative analysis system based on physiological parameter fusion according to claim 1, characterized in that, The trained Restricted Boltzmann Machine (RBM) outputs the following emotion representation fusion features: The multimodal splicing feature matrix to be predicted is standardized. The standardized input features are divided by the noise standard deviation in each dimension, and matrix multiplication is performed with the trained weight matrix. The trained hidden layer bias vector is superimposed, and the hidden layer activation probability matrix is output by the sigmoid activation function. The hidden layer activation probability matrix is used as the emotion representation fusion feature.