VR motion sickness state detection method based on multi-modal physiological parameter fusion
Through multimodal physiological parameter fusion and machine learning algorithms, the problem of insufficient analysis of single indicators for VR motion sickness detection is solved, and accurate detection and user experience optimization of motion sickness are achieved.
Patent Information
- Application Number
- CN202510917881.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing VR motion sickness detection methods have problems such as incomplete analysis of single indicators, easy to overfit, poor generalization ability, and simple and rough evaluation methods.
Multimodal physiological parameter fusion method is adopted, and the synchronous acquisition and processing of electrocardiogram and electrocardiogram signals is combined with MRMR algorithm for feature screening, and classification and regression models are constructed using multiple machine learning algorithms, and VR device parameters are dynamically adjusted to reduce motion sickness symptoms.
Accurate detection of VR motion sickness status is achieved, the generalization ability and robustness of the model are improved, and the severity of motion sickness can be reflected in real time and optimized the user experience.
Smart Images

Figure CN120392030A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent health and biofeedback, and specifically to a method for detecting VR motion sickness state based on multi-modal physiological parameter fusion. Background Art
[0002] In recent years, virtual reality (VR) technology has been greatly popularized and widely applied in various fields, including medical training, education, entertainment, and public safety. Although VR technology can already provide users with excellent immersive experiences, side effects that affect the user experience still exist. In particular, a series of discomforts called "motion sickness" such as eye fatigue, dizziness, nausea, and disorientation that users experience during VR experiences can cause great physical discomfort to users. Therefore, detecting the user's motion sickness state efficiently and accurately is crucial for improving the user experience. Currently, the detection of the user's VR motion sickness state mainly includes subjective measurement methods and objective measurement methods. The subjective measurement method is mainly based on the Simulator Sickness Questionnaire (SSQ) and the VR Motion Sickness Questionnaire (VRSQ). This measurement method is easy to implement and helps to collect a wide range of data, but it may be affected by individual differences and environmental factors, and it cannot perform real-time measurement, resulting in deviations in the results. The objective measurement method is mainly based on physiological characteristic signals of users such as electroencephalogram (EEG), electrocardiogram (ECG), eye movement, and electrodermal activity (EDA). Among them, the method for detecting the user's motion sickness state based on EEG signals, although having advantages such as high time resolution, the introduction of an EEG signal acquisition system will affect the user's VR operation, and it is also more vulnerable to interference from other noises, affecting the accuracy of the data. Heart rate changes are closely related to emotions and physical reactions and can effectively reflect the user's physiological state. Among them, the method for detecting the user's motion sickness state based on ECG signals has the characteristic of strong adaptability and can be used in various VR environments. EDA signals can effectively reflect the user's emotions and physiological reactions, especially in a state of tension or anxiety, where the changes are rapid and obvious. Among them, the method for detecting the user's motion sickness state based on EDA signals has the characteristics of strong sensitivity and strong adaptability. In addition, ECG and EDA signal acquisition devices are lightweight and have strong anti-interference ability, making them more suitable for long-term VR scenarios. Currently, most studies only use a single objective physiological index to analyze the motion sickness state, but a single or limited physiological index often cannot fully and effectively reflect the motion sickness symptoms, and lacks universality and cannot comprehensively explain the problem. Some studies have begun to apply algorithms such as machine learning to fuse multi-dimensional indicators for the detection of VR motion sickness. This is different from the traditional single-indicator evaluation method. Machine learning algorithms can build intelligent classifiers, effectively fuse multi-dimensional data, and more accurately identify the motion sickness state, avoiding misjudgment caused by the limitations of single indicators. However, these methods still have certain limitations. First, the algorithms used in existing studies are few, and the constructed models may have overfitting problems and poor generalization ability. These problems can be reduced by integrating multiple models, enhancing the robustness and generalization ability of the models. Second, the current evaluation methods for motion sickness detection mainly focus on constructing binary classification models (for example, only distinguishing between high and low levels of motion sickness). This evaluation method is relatively simple. However, the change in the degree of motion sickness is actually a continuous process. Therefore, using a multi-classification model or performing regression prediction to evaluate motion sickness can provide more accurate and meaningful results; For example, a motion sickness detection method with the application number CN110613429A. This invention introduces a motion sickness detection method that uses electroencephalogram (EEG) signals to evaluate motion sickness. However, in a VR environment, electrical signals generated by scalp muscle activities (such as blinking and facial expressions) may be confused with EEG signals, affecting the accuracy of the results and interfering with the user's comfort and freedom of movement. Compared with EEG signals, the collection of electrocardiogram (ECG) and galvanic skin response (GSR) signals is more convenient and reliable, and is more sensitive to changes in motion sickness, making it more suitable for long-term monitoring while wearing. In addition, this invention can only identify the occurrence or non-occurrence of motion sickness and cannot distinguish different degrees of motion sickness states (such as mild, moderate, and severe). In fact, motion sickness is a continuous change process. Using a three-classification method or performing regression prediction of motion sickness can more comprehensively explain the severity of motion sickness, and predicting the severity is more practical than simply judging the occurrence or non-occurrence; As for a method for detecting visual-induced motion sickness in naked-eye 3D displays with the application number CN108836322A, this invention introduces a method for detecting visual-induced motion sickness in naked-eye 3D displays based on EEG. This invention only utilizes the correlation between features and response variables for feature screening. Although this method is simple and easy to understand, there are problems such as redundant features not being effectively eliminated, the inability to capture non-linear relationships, and poor adaptability to noise and high-dimensional data. In contrast, using the MRMR algorithm to process the data can extract the key features that are most relevant and non-redundant to the VR motion sickness state from multi-modal high-dimensional data. In addition, this invention only relies on EEG signals as the data source to predict the motion sickness state, but motion sickness is a complex process involving the coordinated response of multiple systems. A single information dimension cannot explain the problem. Although the response of EEG signals to motion sickness is significant, the signals themselves have a large amount of noise and are greatly affected by the environment, individual characteristics, and equipment quality. In contrast, electrodermal signals can reflect sympathetic nerve activity and are relatively stable, and electrocardiogram signals can characterize the dynamic regulation of the autonomic nervous system and have a large signal intensity and are easy to collect. The two are fused and complementary, which can provide a more comprehensive characterization and higher stability of the motion sickness response. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for detecting VR motion sickness state based on the fusion of multi-modal physiological parameters to solve the problems of incomplete single-index analysis, easy overfitting and poor generalization ability of the algorithm, and simple and rough evaluation method in the above-mentioned background technology.
[0004] To achieve the above purpose, the present invention provides the following technical solution: A method for detecting VR motion sickness state based on the fusion of multi-modal physiological parameters, and this detection method includes the following steps: S1. Data collection is carried out through the data collection module: The electrocardiogram signal is collected in the form of 3-lead electrode points fixed around the heart, and the sampling frequency is 1024 Hz. The electrodermal signal is collected in the form of 2-lead electrode points placed at the fingertips, and the sampling frequency is 64 Hz. During the VR operation process, the electrocardiogram and electrodermal signals are respectively collected from the operator's heart signal and finger skin surface signal, and the signals are sent to the collection device through the Bluetooth sensor device. The two source signals are synchronously collected during the normal operation of the operator, ensuring the consistency of the data in the time dimension; S2. Process the electrocardiogram signals collected by the data acquisition module through the signal processing module: First, filter out the electromyogram interference through a Butterworth filter, and set the band-stop frequency range to 49 - 51 Hz. Secondly, use a FIR (Finite Impulse Response) digital filter to process the power frequency interference. Finally, correct the baseline drift through an IIR (Infinite Impulse Response). By preprocessing the electrocardiogram signals, the Heart Rate Variability (HRV) signals in the ECG signals are extracted. Through time-domain analysis and non-linear analysis of the HRV signals, 7 corresponding time-domain features and 6 non-linear features are obtained respectively. The total number of extracted electrocardiogram signal features is 13; S3. Process the skin conductance signals collected by the data acquisition module through the signal processing module: First, eliminate the power frequency noise through a low-pass filter. Secondly, use the wavelet transform method to reduce the noise such as motion artifacts and baseline drift. After the above preprocessing, the Skin Conductance Value (SC) and Skin Conductance Level (SCL) in the skin conductance signals are extracted. The skin conductance signal features including the EDA time-domain features SC, SCL, the maximum value EMAX of the skin conductance value, the minimum value EMIN, the standard deviation ESTD, the variance EVAR, and the range EPD are also used as signal features. The total number of skin conductance signal features is 7; S4. Perform Z-score standardization on the physiological feature data obtained by the signal processing module through the feature processing module to enhance data stability and unify the feature scale. Secondly, use the sliding window function to expand the standardized physiological feature data to obtain a large amount of data for training the model. Expand the 20 indicators (13 electrocardiogram features, 7 skin conductance features), expand from 480 rows of data in time series to 30,000 rows. Finally, use the MRMR (Minimum-RedundancyMaximum-Relevancy) algorithm to rank the importance scores of the signal features and perform feature screening. Finally, after sorting the 20 features of ECG + EDA, 18 features are selected to reduce the risk of overfitting, improve the generalization ability of the model, thereby improving the calculation efficiency and enhancing the robustness; S5. Perform fusion training through the model training module: A total of 18 pieces of physiological signal feature data obtained by the feature processing module (12 electrocardiogram + 6 galvanic skin responses) are used as input features X to identify, classify, and perform regression prediction on the total score of the VRSQ questionnaire. In the classification of the motion sickness state, the total score of the VRSQ questionnaire is divided into three gradient intervals. 0 - 33 points, 0 - 66 points, and 67 - 100 points are respectively marked as low, medium, and high motion sickness levels. The above three different markings are used as the response variables of the motion sickness state classification model. In the regression prediction of the motion sickness state, the total score of the VRSQ questionnaire is used as the response variable of the motion sickness state regression model. The specific approach is to first randomly select 85% of the data in the dataset for training and 15% of the data for testing. Secondly, the ten-fold cross-validation method is adopted. The training set is divided into ten subsets. Each time, one subset is used as the validation set, and the remaining nine subsets are used as the training set. This process is repeated ten times, with the validation set being changed each time. Finally, the results of the ten times are averaged. Finally, a motion sickness state classification model and a motion sickness state regression model are constructed to perform fusion training on the selected physiological feature data. S6. Based on the classification (low, medium, high) and regression scores (0 - 100 points) output by the model training module through the motion sickness improvement module, dynamically output suggestions for adjusting VR device parameters, and select the optimal combination from 16 preset virtual scenarios (4 image qualities × 4 visual areas) to reduce the user's motion sickness symptoms. If the user is in a high motion sickness state (>70 points), the module outputs a combination suggestion of selecting low image quality (720P / 480P) + small field of view (90° / 120°). If the user is in a medium motion sickness state (30 - 70 points), the module outputs a combination suggestion of using medium image quality (1080P) + medium field of view (120° / 150°). If the user is in a low motion sickness state (30 - 70 points), the module outputs a suggestion to maintain the current scenario or switch to high image quality (4K) + large field of view (180°).
[0005] Preferably, the time domain features in S2 include: Inter-Beat Interval (IBI), that is, the time interval between two consecutive heartbeats; Standard Deviation of NN intervals (SDNN); Standard Deviation of Successive Differences (SDSD); Root Mean Square of the Mean of the Sum of Squares of Differences between Adjacent R-R Intervals (RMSSD); Percentage of NN20 intervals; Percentage of NN50 intervals, which is the proportion of adjacent R-R intervals with a difference exceeding 50 ms.
[0006] Preferably, in S2, a Poincare scatter plot is drawn using the HRV signal and elliptical fitting is performed. The extracted non-linear features include: S is the area of the fitted ellipse; SD1 is the semi-minor axis, representing the instantaneous heart rate to heart rate variability or short-term variability; SD2 is the semi-major axis; CSI is the ratio of SD1 to SD2.
[0007] Preferably, the calculation formula for IBI is: ; where N is the total number of normal heart beats, is the i-th R-R interval; The calculation formula for SDNN is: ; where N is the total number of normal heart beats, is the i-th R-R interval; The calculation formula for SDSD is: ; where N is the total number of R-R intervals, is the change in the R-R interval of the i-th heartbeat interval, is the average value of ; The calculation formula for RMSSD is: <SHAPE \*> ; where N is the total number of R-R intervals, is the length (ms) of two adjacent heartbeat cycles; The calculation formulas for pNN20 and pNN50 are: ; ; where NN is the total number of R-R intervals, NN20 is the number of R-R differences higher than 20 ms in the time recorded by the signal, and NN50 is the number of R-R differences higher than 50 ms in the time recorded by the signal.
[0008] Preferably, the calculation formula for SD1 is: ; where N is the total number of R-R intervals, is the length (ms) of two adjacent heartbeat cycles; The calculation formula of SD2 is as follows: ; Where N is the total number of R-R intervals, is average value; The calculation formula of CSI is as follows: .
[0009] Preferably, the discrete wavelet transform formula in S3 is as follows: ; ;[[ID=2S]] Where and are the approximation coefficient and the detail coefficient obtained by decomposition respectively, is the input signal, and are the low-pass filter and the high-pass filter respectively; ; ; Where and are the approximation part and the detail part obtained by reconstruction respectively.
[0010] Preferably, the motion sickness state classification model in S5 is constructed by using decision tree, adaptive boosting, neural network, support vector machine and k-nearest neighbor algorithm.
[0011] Preferably, the motion sickness state regression model in S5 is constructed by using decision tree, adaptive boosting, neural network, Gaussian process regression and kernel approximation algorithm.
[0012] Compared with the prior art, the beneficial effects of the present invention are: the VR motion sickness state detection method based on multi-modal physiological parameter fusion: 1. Select skin conductance and electrocardiogram data with high correlation with motion sickness. Electrocardiogram and skin conductance signals respectively reflect the activities of the autonomic nervous system and the sympathetic nerve, capture the physiological responses of motion sickness from different dimensions, and through multi-modal fusion analysis, more accurately extract the physiological state information during the occurrence of motion sickness, so as to more comprehensively reflect the motion sickness state of the user; 2. Under 16 virtual reality environments, 5 classification algorithms and regression algorithms are respectively used for modeling and verification. Through the comparative analysis of the multi-scenario test results, the best-performing adaptive boosting algorithm is finally selected to construct the classification model and the regression model. This strategy effectively reduces the risk of overfitting by verifying the performance of multiple algorithms in different scenarios, and at the same time significantly improves the generalization ability of the model in practical applications; 3. Considering that motion sickness is a continuously changing process, this method uses a three-class motion sickness detection system (low, medium, high), which can provide higher accuracy and reliability. By introducing a regression prediction model, it can accurately output the current motion sickness state score of the user, ranging from 0 to 100 points. The classification model quickly judges the motion sickness level, and the regression model accurately quantifies the severity, meeting the requirements of different application scenarios; 4. This system uses traditional machine learning algorithms (such as decision trees, SVM) instead of deep learning to optimize the computational complexity. Compared with the deep learning model, the training time of the machine learning model is shortened by more than 50%, which is suitable for deployment on embedded devices and has a lower real-time detection delay; 5. Through the electrocardiogram data, galvanic skin response data, and subjective evaluation data of the subjects collected by the data acquisition module, the signal processing module extracts and processes the above electrocardiogram, galvanic skin response, and subjective evaluation data. The feature screening module expands and screens the extracted data to reduce the dimension to improve the generalization performance of the model. The screened feature indicators are input into the model training module, and the corresponding VR motion sickness state is used as the data label of the integrated model for model training. Through the classification and regression results of the model, the motion sickness level of the current operator is evaluated and described. By combining the three types of data of electrocardiogram, galvanic skin response, and subjective evaluation, a VR motion sickness state detection process with high sensitivity and high reliability is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0015] Please refer to Figure 1 , the present invention provides a technical solution: a VR motion sickness state detection method based on multi-modal physiological parameter fusion. The detection method includes the following steps: S1. Data collection is performed through the data collection module: The electrocardiogram (ECG) signal is collected by fixing 3-lead electrode points around the heart, with a sampling frequency of 1024 Hz. The electrodermal activity (EDA) signal is collected by placing 2-lead electrode points at the fingertips, with a sampling frequency of 64 Hz. During the VR operation, the ECG and EDA signals are respectively collected from the operator's heart signal and the finger skin surface signal, and the signals are sent to the collection device through the Bluetooth sensor device. The two source signals are synchronously collected during the normal operation of the operator, ensuring the consistency of the data in the time dimension; S2. The ECG signal collected by the data collection module is processed through the signal processing module: First, the myoelectric interference is filtered out by the Butterworth filter, and the stopband frequency range is set to 49 - 51 Hz. Secondly, the power frequency interference is processed by the FIR (Finite Impulse Response) digital filter. Finally, the baseline drift is corrected by the IIR (Infinite Impulse Response). By preprocessing the ECG signal, the heart rate variability (HRV) signal in the ECG signal is extracted. Through the time-domain analysis and non-linear analysis of the HRV signal, 7 corresponding time-domain features and 6 non-linear features are obtained respectively. The total number of ECG signal features extracted is 13; The time-domain features in S2 include: Inter-Beat Interval (IBI), that is, the time interval between two consecutive heartbeats; Standard Deviation of NN intervals (SDNN); Standard Deviation of Successive Differences (SDSD); Root Mean Square of the Mean of the Sum of Squares of Differences between Adjacent R-R Intervals (RMSSD); Percentage of NN20 intervals, that is, the proportion of adjacent R-R intervals with a difference exceeding 20 ms; Percentage of NN50 intervals, that is, the proportion of adjacent R-R intervals with a difference exceeding 50 ms; In S2, the Poincare scatter plot is drawn using the HRV signal and elliptical fitting is performed. The non-linear features extracted include: S is the area of the fitted ellipse; SD1 is the semi-minor axis, representing the instantaneous heart rate to heart rate variability or short-term variability; SD2 is the semi-major axis; CSI is the ratio of SD1 to SD2; The calculation formula for IBI is as follows: ; where N is the total number of normal heartbeats, is the i-th R-R interval; The calculation formula for SDNN is as follows: ; where N is the total number of normal heartbeats, is the i-th R-R interval; The calculation formula for SDSD is as follows: ; where N is the total number of R-R intervals, is the change in the heartbeat interval of the i-th R-R interval, is the average value of; The calculation formula for RMSSD is as follows: ; where N is the total number of R-R intervals, is the length (ms) of two adjacent heartbeat cycles; The calculation formulas for pNN20 and pNN50 are as follows: ; ; where NN is the total number of R-R intervals, NN20 is the number of R-R differences higher than 20 ms in the time recorded by the signal, and NN50 is the number of R-R differences higher than 50 ms in the time recorded by the signal; The calculation formula for SD1 is as follows: ; where N is the total number of R-R intervals, is the length (ms) of two adjacent heartbeat cycles; The calculation formula for SD2 is as follows: ; where N is the total number of R-R intervals, is the average value of; The calculation formula for CSI is as follows: .
[0016] S3. Process the electrodermal signals collected by the data acquisition module through the signal processing module: First, eliminate power frequency noise through a low-pass filter. Second, use the wavelet transform method to reduce noise such as motion artifacts and baseline drift. After the above preprocessing, extract the skin conductance value (SC) and skin conductance level (SCL) in the electrodermal signal. The electrodermal signal features including the EDA time-domain features SC, SCL, the maximum value EMAX, the minimum value EMIN, the standard deviation ESTD, the variance EVAR, and the range EPD of the skin conductance value are also used as signal features. There are a total of 7 electrodermal signal features; The discrete wavelet transform formula in S3 is as follows: ; ; where and are the approximate coefficient and the detail coefficient obtained by decomposition respectively, is the input signal, and are the low-pass filter and the high-pass filter respectively; ; ; where and are the approximate part and the detail part obtained by reconstruction respectively.
[0017] The physiological indicators are shown in the following table: Full name of the indicator Abbreviation in English Meaning of the indicator Standard deviation of the R-R interval of sinus beats SDNN Measure the overall variability of heart rate Heart rate variability RMSSD Reflects the degree of heart rate variability in a short period of time Standard deviation of the difference between adjacent RR intervals SDSD Measure the variability between adjacent heartbeat intervals Proportion of adjacent RR intervals with a difference exceeding 20ms pNN20 Reflect short-term heart rate variability Proportion of adjacent RR intervals with a difference exceeding 50ms pNN50 Measure the proportion of adjacent heartbeat intervals with a difference exceeding 50ms Median absolute deviation of heart rate HR_MAD Reflects the variability of heart rate Area of the ellipse drawn by Poincare isoclines S Represents the overall heart rate variability Semi-minor axis of the Poincare scatter plot ellipse SD1 Instantaneous difference between two adjacent R-R intervals Semi-major axis of the Poincare scatter plot ellipse SD2 Continuity or long-term difference between two adjacent groups of R-R intervals Ratio of SD1 and SD2 CSI Indicates the relative influence of sympathetic and parasympathetic nerves on heart rate variability Respiratory rate BR Measure the rate of respiration Skin conductance value SC Characterizes the skin conductance change caused by sensory stimulation Skin conductance level SCL Describes the overall effect of sensory stimulation on skin conductance
[0018] S4. Perform Z-score standardization on the physiological feature data obtained by the signal processing module through the feature processing module to enhance data stability and unify the feature scale. Second, use the sliding window function to expand the standardized physiological feature data to obtain a large amount of data for training the model. Expand the 20 indicators (13 electrocardiogram features and 7 electrodermal features), expanding from 480 rows of data to 30,000 rows in time series. Finally, use the MRMR (Minimum-Redundancy Maximum-Relevancy) algorithm to sort the importance scores of the signal features and perform feature screening. Finally, sort the 20 features of ECG+EDA and select 18 features to reduce the risk of overfitting, improve the generalization ability of the model, thereby improving the calculation efficiency and enhancing the robustness; Z-score normalization (also known as standard deviation normalization) is a commonly used data normalization method that transforms data into a distribution with a mean of 0 and a standard deviation of 1. This method is applicable to converting data with different dimensions or ranges into a comparable standard scale.
[0019] ; The sliding window function is a technique for data processing. Its principle is to define a fixed-size window that slides over the data sequence and processes the data within each window in turn. During the processing, the window moves one data point at a time, and each movement generates a new window, thereby achieving data expansion and processing.
[0020] The principle of the MRMR algorithm is to find a set of features in the original feature set that has the maximum relevance (Max-Relevance) to the final output result, but the minimum redundancy among the features (Min-Redundancy). It first calculates the correlation between each feature and the response variable; then considers the redundancy between the features; finally, by comprehensively considering the correlation and redundancy, it ranks all the features according to their importance scores and selects the ones with high scores as input features;
[0021] S5. Conduct fusion training through the model training module: A total of 18 physiological signal feature data obtained by the feature processing module (12 electrocardiograms + 6 galvanic skin responses) are used as input features X to identify, classify, and perform regression prediction on the total score of the VRSQ questionnaire: In the classification of the motion sickness state, the total score of the VRSQ questionnaire is divided into three gradient intervals, and 0 - 33 points, 0 - 66 points, and 67 - 100 points are respectively marked as low, medium, and high motion sickness degrees. The above three different markings are used as the response variables of the motion sickness state classification model. In the regression prediction of the motion sickness state, the total score of the VRSQ questionnaire is used as the response variable of the motion sickness state regression model. The specific approach is to first randomly select 85% of the data in the dataset for training and 15% of the data for testing. Secondly, the ten-fold cross-validation method is adopted. The training set is divided into ten subsets. Each time, one subset is used as the validation set, and the remaining nine subsets are used as the training set. This process is repeated ten times, with the validation set changed each time. Finally, the average of the ten results is taken. Finally, a motion sickness state classification model and a motion sickness state regression model are constructed to perform fusion training on the screened physiological feature data; In S5, the motion sickness state classification model is constructed using decision tree, adaptive boosting, neural network, support vector machine, and k-nearest neighbor algorithms; In S5, the motion sickness state regression model is constructed using decision tree, adaptive boosting, neural network, Gaussian process regression, and kernel approximation algorithms; Decision tree classification and regression model The decision tree algorithm model is a supervised learning algorithm that mimics the human decision-making process. It visualizes data features and decision rules through a tree-like diagram structure, performs feature selection and decision division recursively, and is characterized by being easy to understand and interpret.
[0022] The prediction of the decision tree classification model is based on the majority class of the leaf nodes, while the prediction of the decision tree regression model is based on the average value of the leaf nodes. The construction process of the decision tree model is as follows:
[0023] (1) When constructing a classification model, select Information Gain as the criterion for splitting the data set; when constructing a regression model, select Mean Squared Error (MSE) as the splitting criterion for the data set. The formula for Information Gain is as follows:
[0024] ; where S is the data set, c is the number of classes, represents the proportion of samples of class i in the data set S.
[0025] ; where S is the data set, A is the attribute, are all possible values of attribute A, represents the number of samples in the data set S under the condition that the value of attribute A is v.
[0026] (2) Select the best feature according to the splitting criterion, and recursively divide the data set until the stopping condition is met.
[0027] (3) Construct leaf nodes, and the leaf nodes represent the final predicted values.
[0028] (4) Prune to remove unimportant branches in the tree to improve the generalization ability of the model and avoid the risk of overfitting.
[0029] Adaptive Boosting Classification and Regression Model Adaptive Boosting (AdaBoost) is an ensemble learning method that gradually reduces the error in classification or regression by weighted combination of multiple weak learners (such as decision trees), and has strong generalization ability and high robustness. The construction process of the Adaptive Boosting model is as follows:
[0030] (1) Select a decision tree as the weak learner and assign initial weights to each training sample.
[0031] (2) In the construction of the classification model, train the weak classifier in each round of iteration and calculate the error rate of the weak classifier, and set the weight of the current weak classifier according to the error rate. In the construction of the regression model, set the weight of the current weak regressor based on the prediction error (residual) of the weak regressor in each round of iteration. The calculation formula is as follows:
[0032] ; in is the weight of sample i, is the true label of the sample, is the predicted label of the current classifier, is an indicator function that equals 1 if the classification is wrong and 0 otherwise.
[0033] ; If the error rate of the classifier is high, its weight Will be small, on the contrary, when the error rate is low It will be bigger.
[0034] (3) Increase the weights of samples with large misclassification and prediction errors and perform normalization.
[0035] (4) Repeat steps (2) and (3) until the error converges to a sufficiently small value.
[0036] (5) All weak learners are weighted together to form a strong learner. The final prediction result of the model is: ; ; Where T is the number of weak splitters, For each weak classifier corresponding weight, Predict labels for each weak classifier, is the output of the final classifier. If the value is positive, the prediction is class 1; if it is negative, the prediction is class -1; is the predicted value of the final regressor.
[0037] Neural Network Classification and Regression Models A neural network model is a computational model that simulates the neuronal structure of the human brain. It extracts features and recognizes patterns through multi-layered neuronal connections. It possesses powerful automatic learning and nonlinear mapping capabilities, making it suitable for handling complex pattern recognition and regression problems. The goal of a classification model is to minimize classification error, while the goal of a regression model is to minimize the difference between predicted and actual values, enabling the neural network to accurately predict continuous values. The process of building a neural network model is as follows:
[0038] (1)Design the input layer, hidden layer, and output layer of the neural network. The number of nodes in the input layer is equal to the number of input features; the activation function of the hidden layer (introducing non-linearity to enable the neural network to learn complex patterns) selects the ReLU (Rectified Linear Unit) function; the activation function of the output layer of the classification model selects the softmax function (for multi-classification problems) to output the probability of each class; the output layer of the regression model does not set an activation function and directly outputs a continuous numerical value.
[0039] (2)For the classification model, select the categorical cross-entropy loss function to measure the gap between the predicted class and the true class. For the regression model, select the mean squared error (MSE) loss function to measure the gap between the predicted value and the true value. The optimization algorithm selects the gradient descent algorithm.
[0040] (3)Train the model through forward propagation, calculating the loss, and backpropagation. Forward propagation includes the calculations of each layer of the neural network to obtain the prediction results; backpropagation updates the network weights through the gradient descent algorithm to minimize the loss function.
[0041] (4)Repeat step (3) until the model loss is small enough.
[0042] Support Vector Machine Classification Model The support vector machine (SVM) is a supervised learning algorithm that maximizes the margin between classes by finding the optimal hyperplane. It is characterized by its ability to effectively handle high-dimensional data and has good generalization ability, especially performing well when the dataset is small or the noise is low. The support vector machine classification model focuses on maximizing the margin. The model construction process is as follows:
[0043] (1)Select the Gaussian radial basis kernel (RBF kernel) function to map multi-modal feature data such as electrocardiogram and galvanic skin response to a high-dimensional space for linear separability.
[0044] (2)Introduce a soft margin and select the optimal penalty parameter (C) through grid search.
[0045] (3)Solve the convex optimization problem and construct the optimal hyperplane to complete the model construction. The convex optimization problem is as follows:
[0046] ; ; where is the sample feature vector, is the sample label, is the hyperplane parameter vector.
[0047] K-Nearest Neighbor Classification Model The K-Nearest Neighbor Classification Model (K-NN) is an instance-based learning algorithm. By calculating the distances between the samples to be classified and the samples in the training set, it selects the nearest K neighbors to vote or vote with weights to determine the class. The construction process of the K-NN model is as follows:
[0048] (1) Standardize the pre-training data to ensure that the scales of different features are consistent.
[0049] (2) Select the Euclidean distance metric and select the optimal k value through cross-validation. If the k value is too small, it may be affected by noise and lead to overfitting. If the k value is too large, it may lead to underfitting.
[0050] (3) The KNN algorithm completes the classification task through a voting mechanism, that is, it selects the K nearest samples, and then predicts the class of the new sample according to the label that appears most frequently among these K samples.
[0051] Gaussian Process Regression Model Gaussian Process Regression (GPR) is a non-parametric Bayesian regression method, which is widely used in dealing with regression problems, especially suitable for small-sample, high-dimensional, and complex regression tasks. Its basic idea is to establish the relationship between data points through a Gaussian process and can flexibly fit complex non-linear functions. The construction process of the GPR model is as follows:
[0052] (1) Define a Gaussian process, and use the squared exponential, Matern5 / 2, exponential, and quadratic rational as kernel functions to construct different GPR models respectively, and select the model with the best performance. The Gaussian process formula is as follows:
[0053] ; where is the mean function, usually assumed to be 0, is the kernel function (also called the covariance function).
[0054] (2) Calculate the covariance matrix of the training data and the covariance matrix of the training data and the test data.
[0055] (3) Calculate the predicted value of the Gaussian process regression based on the covariance of the training data points and the test data points (4) Optimize the hyperparameters of the kernel function in the Gaussian process regression by maximizing the marginal likelihood estimation (MLE, Maximum Likelihood Estimation). The formula is as follows:
[0056] ; Kernel approximation regression model The kernel approximation regression model is a non - linear regression method. It maps data to a high - dimensional space through a kernel function and performs linear regression in this space. It has the characteristics of flexibility, non - parametric, and local weighting, and can adapt to the local characteristics of data and capture the dependence relationships between variables. The construction process of the kernel approximation regression model is as follows:
[0057] (1) Select the Gaussian radial basis kernel (RBF kernel) function to map the input data to a high - dimensional feature space.
[0058] (2) Construct the objective function and obtain the optimal parameters by taking the derivative of the objective function . The formula of the objective function is as follows:
[0059] ; where is an N×N kernel matrix, , that is, the kernel function value between each pair of data points and . is the parameter (weight) of the model
[0060] (3) Predict the test points according to the optimal parameters obtained in (2) and calculate the kernel matrix.
[0061] (4) Select a suitable regularization parameter through cross - validation and optimize the length scale of the Gaussian kernel by maximizing the marginal likelihood estimation .
[0062] S6. Based on the classification (low, medium, high) and regression scores (0 - 100 points) output by the model training module, the motion sickness improvement module dynamically outputs suggestions for adjusting VR device parameters, selects the optimal combination from 16 preset virtual scenarios (4 picture qualities × 4 visual areas) to reduce the user's motion sickness symptoms. If the user is in a high motion sickness state (>70 points), the module outputs a suggestion to select a combination of low picture quality (720P / 480P) + small field of view (90° / 120°); if the user is in a medium motion sickness state (30 - 70 points), the module outputs a suggestion to adopt a combination of medium picture quality (1080P) + medium field of view (120° / 150°); if the user is in a low motion sickness state (30 - 70 points), the module outputs a suggestion to maintain the current scenario or switch to a combination of high picture quality (4K) + large field of view (180°) Example - Evaluate the accuracy and effectiveness through the verification module The purpose of this module is to verify the accuracy and effectiveness of the VR motion sickness state detection method based on multimodal physiological parameter fusion and evaluate the stability and reliability of this method under different conditions.
[0063] First, in terms of verifying the experimental data collection, 32 subjects were recruited to participate in the collection of this verification experiment data, including 16 males and 16 females, with an average age of 22.9 years. Among them, 42% had never used VR, 48% had used it very rarely, and 10% belonged to other categories. Second, the experimental conditions included hardware devices, virtual reality scenarios, and virtual reality tasks. The hardware devices included 1 head-mounted VR all-in-one device (Pico Neo 3) and 1 high-performance computer; the virtual reality scenarios were designed and built using Unity3D software; 16 virtual reality scenarios were designed, with 4 levels of virtual picture quality (4K, 1080P, 720P, 480P) × 4 levels of visible areas (180°, 150°, 120°, 90°) to simulate virtual reality scenarios under different conditions. Based on the above 9 scenario combinations, two tasks were set: Task 1 was a collection task, and Task 2 was a tracking task. The collection task required collecting as many distributed balls as possible in the built maze within a given time, and the tracking task required moving along the trajectory generated by the balls within a given time. Finally, the subjective questionnaire data, task performance data, and physiological signal data (galvanic skin response, electrocardiogram) of the subjects were collected. The total score of the VRSQ questionnaire among them was classified and recognized and regression predicted. In the classification and recognition of the motion sickness state, the three gradient intervals of the total score of the VRSQ questionnaire, 0 - 33 points, 34 - 66 points, and 67 - 100 points, were respectively marked as low, medium, and high motion sickness states. In the regression prediction of the motion sickness state, the total score of the VRSQ questionnaire was used as the response variable of the motion sickness state regression model. Through the comprehensive processing of the data collection module, signal processing module, feature screening module, and model training module, this method outputs the classification results (low, medium, high) of the subjects' motion sickness states and the regression prediction scores (ranging from 0 to 100 points).
[0064] The verification results are shown in the following table. This method has high accuracy and stability and can effectively detect the motion sickness state in different VR environments.
[0065] Classification and recognition results of the motion sickness state:
[0066] Regression prediction results of the motion sickness state:
[0067] In the above specific embodiments, the object, technical solution and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and do not limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A method for detecting VR motion sickness state based on multi-modal physiological parameter fusion, characterized in that: The detection method includes the following steps: S1. Data acquisition is performed through a data acquisition module: The electrocardiogram (ECG) signal is collected by fixing 3-lead electrode points around the heart, with a sampling frequency of 1024 Hz. The electrodermal (EDA) signal is collected by placing 2-lead electrode points at the fingertips, with a sampling frequency of 64 Hz. During the VR operation, the ECG and EDA signals are respectively collected from the operator's heart signal and finger skin surface signal, and the signals are sent to the collection device through a Bluetooth sensor device. The two source signals are synchronously collected during the normal operation of the operator, ensuring the consistency of data in the time dimension. S2. The ECG signal collected by the data acquisition module is processed through a signal processing module: First, the electromyogram interference is filtered out by a Butterworth filter, and the band-stop frequency range is set to 49 - 51 Hz. Secondly, the power frequency interference is processed by a FIR (Finite Impulse Response) digital filter. Finally, the baseline drift is corrected by an IIR (Infinite Impulse Response). Through the preprocessing of the ECG signal, the heart rate variability (HRV) signal in the ECG signal is extracted. By performing time-domain analysis and non-linear analysis on the HRV signal, 7 corresponding time-domain features and 6 non-linear features are obtained respectively, and a total of 13 ECG signal features are extracted. S3. The EDA signal collected by the data acquisition module is processed through a signal processing module: First, the power frequency noise is eliminated by a low-pass filter. Secondly, the wavelet transform method is used to denoise the noise such as motion artifacts and baseline drift. After the above preprocessing, the skin conductance value (SC) and skin conductance level (SCL) in the EDA signal are extracted. The EDA signal features including the EDA time-domain features SC, SCL, the maximum value EMAX of the skin conductance value, the minimum value EMIN, the standard deviation ESTD, the variance EVAR, and the range EPD are also used as signal features, and a total of 7 EDA signal features are obtained. S4. The physiological feature data obtained by the signal processing module is standardized by Z-score through a feature processing module to enhance data stability and unify the feature scale. Secondly, the sliding window function is used to expand the data of the standardized physiological feature data to obtain a large amount of data for training the model. The 20 indicators (13 ECG features and 7 EDA features) are expanded from 480 rows of data to 30000 rows in time series. Finally, the MRMR (Minimum-Redundancy Maximum-Relevancy) algorithm is used to rank the importance scores of the signal features and perform feature screening. Finally, 18 features are selected after ranking the 20 features of ECG + EDA to reduce the risk of overfitting, improve the generalization ability of the model, thereby improving the calculation efficiency and enhancing the robustness. S5. Fusion training is performed through a model training module : A total of 18 physiological signal feature data obtained by the feature processing module (12 electrocardiograms + 6 skin conductance) are used as input feature X to identify, classify, and perform regression prediction on the total score of the VRSQ questionnaire: In the classification of the dizziness state, the total score of the VRSQ questionnaire is divided into three gradient intervals. 0 - 33 points, 0 - 66 points, and 67 - 100 points are respectively marked as low, medium, and high dizziness levels. The above three different markings are used as the response variables of the dizziness state classification model. In the regression prediction of the dizziness state, the total score of the VRSQ questionnaire is used as the response variable of the dizziness state regression model. The specific approach is to first randomly select 85% of the data in the dataset for training and 15% of the data for testing. Secondly, the ten-fold cross-validation method is adopted. The training set is divided into ten subsets. Each time, one subset is used as the validation set, and the remaining nine subsets are used as the training set. This process is repeated ten times, with the validation set being changed each time. Finally, the average of the ten results is taken. Finally, a motion sickness state classification model and a motion sickness state regression model are constructed to perform fusion training on the selected physiological feature data; S6. Based on the classification (low, medium, high) and regression scores (0 - 100 points) output by the model training module, the motion sickness improvement module dynamically outputs suggestions for adjusting VR device parameters, and selects the optimal combination from 16 preset virtual scenarios (4 image qualities × 4 viewing areas) to reduce the user's motion sickness symptoms. If the user is in a high motion sickness state (>70 points), the module outputs a combination suggestion of selecting low image quality (720P / 480P) + small viewing field (90° / 120°). If the user is in a medium motion sickness state (30 - 70 points), the module outputs a combination suggestion of using medium image quality (1080P) + medium viewing field (120° / 150°). If the user is in a low motion sickness state (30 - 70 points), the module outputs a suggestion to maintain the current scenario or switch to high image quality (4K) + large viewing field (180°).
2. The method for detecting VR motion sickness state based on multimodal physiological parameter fusion according to claim 1, wherein: The time domain features in S2 include: Inter-Beat Interval (IBI), that is, the time interval between two consecutive heartbeats; Standard Deviation of NN intervals (SDNN); Standard Deviation of Successive Differences (SDSD); Root Mean Square of the Mean of the Squared Differences between Adjacent R-R Intervals (RMSSD); Percentage of NN20 intervals, which is the proportion of adjacent R-R intervals with a difference exceeding 20ms; Percentage of NN50 intervals, which is the proportion of adjacent R-R intervals with a difference exceeding 50ms.
3. A method for detecting VR motion sickness state based on multimodal physiological parameter fusion according to claim 1, characterized in that: In S2, the Poincare scatter plot is drawn using the HRV signal and elliptical fitting is performed. The non-linear features extracted include: S is the area of the fitted ellipse; SD1 is the semi-minor axis, representing the instantaneous heart rate to heart rate variability or short-term variability; SD2 is the semi-major axis; CSI is the ratio of SD1 to SD2.
4. A method for detecting VR motion sickness state based on multi-modal physiological parameter fusion according to claim 2, characterized in that: The calculation formula of the IBI is as follows: ; where N is the total number of normal heartbeats, is the i-th R-R interval; The calculation formula of the SDNN is as follows: ; where N is the total number of normal heartbeats, is the i-th R-R interval; The calculation formula of the SDSD is as follows: ; where N is the total number of R-R intervals, is the change in the heartbeat interval of the i-th R-R interval, is the average value of; The calculation formula of the RMSSD is as follows: ; where N is the total number of R-R intervals, is the length (ms) of two adjacent cardiac cycles; The calculation formulas of the pNN20 and pNN50 are as follows: ; ; Where NN is the total number of R-R intervals, NN20 is the number of R-R differences higher than 20 ms in the time recorded by the signal, and NN50 is the number of R-R differences higher than 50 ms in the time recorded by the signal.
5. The method for detecting VR motion sickness state based on multi-modal physiological parameter fusion according to claim 3, wherein: The calculation formula of the SD1 is as follows: ; where N is the total number of R-R intervals, is the length (ms) of two adjacent cardiac cycles; The calculation formula of the SD2 is as follows: ; where N is the total number of R-R intervals, is average value; The calculation formula of the CSI is as follows: 。 6. The method for detecting VR motion sickness state based on multi-modal physiological parameter fusion according to claim 1, wherein: The discrete wavelet transform formula in the S3 is as follows: ; ; Among them and are the approximate coefficient and the detail coefficient obtained by decomposition respectively, is the input signal, and are the low-pass filter and the high-pass filter respectively; ; ; wherein and are respectively the approximated part and the detailed part obtained by reconstruction.
7. A method for detecting VR motion sickness state based on multi-modal physiological parameter fusion according to claim 1, characterized in that: The motion sickness state classification model in the S5 is constructed using decision trees, adaptive boosting, neural networks, support vector machines, and k-nearest neighbor algorithms.
8. A method for detecting VR motion sickness state based on multi-modal physiological parameter fusion according to claim 1, characterized in that: The motion sickness state regression model in the S5 is constructed using decision trees, adaptive boosting, neural networks, Gaussian process regression, and kernel approximation algorithms.
Citation Information
Patent Citations
Wearable motion load detection device and method based on dynamic ECG
CN109998522A
Motion sickness detection system for autonomous vehicle
CN115768671A
VR (Virtual Reality) vertigo objective assessment method and assessment system fused with multi-dimensional electroencephalogram characteristics
CN116942179A
Scene demonstration evaluation system and method for public management of virtual experiment cases
CN118035651A
Passenger carsickness degree evaluation method and system
CN119066545A
Cited By
Cardiovascular disease recognition system and device based on electrocardio fusion and medium
CN120570618A