A VR disease assessment method based on a bimodal network
By constructing a VR disease evaluation method based on a dual-modal network, using 3D-ResNet and 3D-CBAM attention mechanisms to imitate the two paths of the human visual system, the problem of data acquisition difficulties in the existing VR disease evaluation methods is solved, and high-precision VR disease evaluation is achieved.
Patent Information
- Application Number
- CN202210896558.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-07-27
AI Technical Summary
The existing VR disease evaluation methods require the participation of a large number of subjects, and the physiological signal-based methods rely on special equipment and there are many interference factors in the signal acquisition process, which leads to difficulty in data extraction and affects the development of VR technology.
A VR disease evaluation method based on a dual-modal network is constructed, and the appearance flow network and motion flow network are designed respectively. The 3D-ResNet and 3D-CBAM attention mechanism are used to realize VR disease evaluation by fusing the output of two subnets, imitating the two paths of the human visual system.
The objective estimation of VR disease is achieved, the prediction accuracy of VR disease evaluation network is improved, and the problem of difficulty in data acquisition in existing methods is solved.
Smart Images

Figure CN115147004B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a VR disease evaluation method based on a bimodal network, belonging to the field of virtual reality. Background Art
[0002] Virtual Reality (VR) integrates various technologies such as computer graphics, multimedia, and simulation to simulate a 360° virtual environment, bringing users a realistic and immersive viewing and interaction experience. In recent years, the VR technology and industry have developed rapidly. Especially in 2021, the popularity of the "metaverse" concept has once again promoted the rapid development of VR technology. However, discomfort symptoms such as dizziness, nausea, and vomiting are extremely likely to occur during VR experiences. In severe cases, there may even be arrhythmia, collapse, etc., resulting in problems such as limited VR product experience duration and poor experience effects, seriously affecting the development of VR technology and industry, and being one of the bottleneck problems affecting the development of VR technology.
[0003] This discomfort during VR experiences is similar to that of carsickness and seasickness, belonging to the category of motion sickness and simulator sickness, and is called virtual reality sickness, i.e., VR sickness. Before experiencing VR products, accurately evaluating the disease-inducing degree of the visual content of VR products is an important research topic in this field. Currently, VR disease evaluation mainly includes subjective evaluation methods and objective evaluation methods. The subjective evaluation method mainly evaluates by having the subjects fill out questionnaires after VR experiences. The objective evaluation method takes the subjective evaluation results as a benchmark and estimates the disease by obtaining the physiological signals (such as skin conductance, brain waves, heart rate, etc.) of the subjects during VR experiences. The above two types of methods require implementing VR experience experiments and a large number of subjects; especially the method based on physiological signal measurement requires relying on special equipment, and there are many interference factors during the signal acquisition process, making it difficult to extract useful data. Research shows that the occurrence of VR sickness is closely related to the vestibular system and the visual system. During VR experiences, the mismatch between the motion perceived by the visual system and the vestibular system is one of the fundamental reasons for discomfort. Therefore, scholars have carried out research on objective evaluation methods of VR sickness based on 360° VR videos. Extract features such as motion in the video and use machine learning and other methods to estimate the disease-inducing degree. In recent years, with the development of deep learning, deep convolutional autoencoders, recurrent neural networks, etc. have been applied to the field of VR disease evaluation. Summary of the Invention
[0004] The object of the present invention is to provide a VR disease evaluation method based on a bimodal network. The human visual system includes two pathways, one is the ventral stream responsible for object recognition, and the other is the dorsal stream responsible for motion recognition. In view of this, the present invention designs an appearance stream network and a motion stream network respectively, mimicking the two pathways of the human visual system; the appearance stream network takes continuous RGB video frames as input to extract video frame information; the motion stream network takes the corresponding continuous optical flow image sequence of the RGB video as input to extract motion information; the final VR disease evaluation is achieved by fusing the results of the two sub-networks.
[0005] The technical solution of the present invention is realized as follows: A VR disease evaluation method based on a bimodal network, characterized in that: First, on the basis of the 2D-ResNet50 model, all convolutional kernels are extended into 3D convolutional kernels, and appearance stream and motion stream sub-networks are constructed based on 3D-ResNet; Then, 3D convolution is introduced into the 2D-CBAM attention mechanism, and attention modules are added to the two sub-networks respectively to strengthen the features in channels and space; Finally, a weighted average backend fusion method is used to achieve VR disease evaluation; the specific steps are as follows:
[0006] Step 1: Construct an appearance stream sub-network for VR disease evaluation, including the following sub-steps:
[0007] Step 101: Reduce the size of each frame image of the 360° VR color video and crop it to a fixed size of 112×112; and decompose the video into multiple consecutive frame groups, and each group includes consecutive L frame images;
[0008] Step 102: Divide it into four categories according to the subjective VR disease score value of the VR video: comfortable, slightly uncomfortable, moderately uncomfortable, and severely uncomfortable, and use the classification result as the true value for model training;
[0009] Step 103: On the basis of the traditional 2D-ResNet50, keep the network structure unchanged, and extend all convolutional kernels into 3D convolutional kernels, so that the convolutional kernels can perform sliding convolutional operations on the image frame sequence in the time dimension; the input sizes of the Conv1, Conv2_x, Conv3_x, Conv4_x, and Conv5_x layers of the improved appearance stream sub-network are respectively: L×112×112, L×56×56, The network structure configurations are respectively: 7×7×7, 64, stride 2, After extracting features layer by layer, connect to a fully connected layer, and use a Softmax classifier and a cross-entropy loss function for classification;
[0010] Step 104: Expand the traditional 2D-CBAM attention mechanism into a 3D structure, and consider the change of depth parameters when extracting channel features and spatial features each time. For the input feature map F 3D , calculate the channel attention feature map M CA and the spatial attention feature map M SA in the order of channel - space respectively:
[0011]
[0012]
[0013] where, σ represents the Sigmoid activation function, ω0, ω1 are weights, and ω0 ∈ R C / r×C , C is the number of channels, ω1 ∈ R C ×C / r , and for and are shared; and are two feature descriptors obtained by solving the feature map F 3D through max - pooling and average - pooling operations respectively, The parameter r is usually taken as 16; NLP is a hidden layer in the CBAM attention model, that is, a shared multi - layer perceptron; f 7×7×7 represents a 3D convolutional layer of 7×7×7; and are two feature descriptors obtained by solving F' 3D through max - pooling and average - pooling operations respectively,
[0014] The output feature map of the attention module is F" 3D :
[0015]
[0016] Add 3D - CBAM modules after the 3D convolution of the Conv1 layer in the appearance flow sub - network and after each residual block of the four layers such as Conv2_x respectively;
[0017] Step 105: Use the color video continuous frame group images as the input of the appearance flow sub - network, each image is 3 channels × 112 pixels × 112 pixels, and conduct training.
[0018] Step 2: Build a motion flow sub - network for VR disease evaluation, including the following sub - steps:
[0019] Step 201: Use the FlowNet2.0 algorithm to obtain the optical flow map sequence of the 360° VR color video; repeat Step 101 to obtain multiple consecutive optical flow map sequence groups for each video; and use the classification result of Step 102 as the ground truth for model training.
[0020] Step 202: Repeat Steps 103 and 104 to construct the motion flow sub-network, using the images in the consecutive optical flow map sequence group as the input of the network, with each image being 2 channels × 112 pixels × 112 pixels, and perform training.
[0021] Step 3: Adopt the method of backend fusion to perform weighted average fusion on the outputs of the appearance flow sub-network and the motion flow sub-network to obtain the final VR disease evaluation result Average_f:
[0022]
[0023] where the weight w i ≥ 0, and T = 2, P(x i ) represents the output results of the two sub-networks.
[0024] The positive effect of the present invention is to achieve an objective estimation of VR disease, using the RGB frames and the corresponding optical flow map sequences of the 360° VR video as inputs respectively, constructing the appearance flow and motion flow sub-networks based on the 3D-ResNet architecture; in order to make the network better reflect and imitate the human perception system, integrating the 3D-CBAM channel attention and spatial attention mechanisms; and obtaining the VR disease estimation result through the method of weighted average fusion; improving the performance of the VR disease evaluation network and effectively enhancing the prediction accuracy. Description of the Drawings
[0025] Figure 1 Sub-network structure diagram of 3D-ResNet50 based on the 3D-CBAM attention mechanism.
[0026] Figure 2 Overall structure diagram of the VR disease evaluation bimodal network. Detailed Embodiment
[0027] The following further describes the present invention in conjunction with the drawings and embodiments: As Figure 1-2As shown, a VR disease evaluation method based on a bimodal network is characterized in that: First, on the basis of the 2D-ResNet50 model, all convolutional kernels are extended into 3D convolutional kernels, and appearance flow and motion flow sub-networks are constructed based on 3D-ResNet; Then, 3D convolution is introduced into the 2D-CBAM attention mechanism, and attention modules are added to the two sub-networks respectively to strengthen the features in channels and space; Finally, a weighted average backend fusion method is used to implement VR disease evaluation. In the embodiment, the VR stereo video public dataset provided by Padmanaban et al. is used. This dataset consists of 19 stereo VR videos with a duration of 60 seconds each, and the VR disease scores of each video are given. The specific steps are as follows:
[0028] Step 1: Construct an appearance flow sub-network for VR disease evaluation, including the following sub-steps:
[0029] Step 101: Preprocess each video in the dataset, shrink each frame image of the VR color video, and crop it to a fixed size of 112×112; Use data augmentation techniques such as image translation, rotation, and flipping to expand the experimental samples; And decompose the video into multiple consecutive frame groups, and each group includes consecutive L frame images;
[0030] Step 102: Divide the videos into four categories according to the subjective VR disease score values of the videos: comfortable, slightly uncomfortable, moderately uncomfortable, and severely uncomfortable, and use the classification results as the true values for model training;
[0031] Step 103: On the basis of the traditional 2D-ResNet50, keep the network structure unchanged, and expand all convolutional kernels into 3D convolutional kernels, so that the convolutional kernels can perform sliding convolutional operations on the image frame sequence in the time dimension; The input sizes of the Conv1, Conv2_x, Conv3_x, Conv4_x, and Conv5_x layers of the improved appearance flow sub-network are: L×112×112, L×56×56, The network structure configurations are: 7×7×7, 64, stride 2, After extracting features layer by layer, connect to a fully connected layer, and use a Softmax classifier and a cross-entropy loss function for classification;
[0032] Step 104: Expand the traditional 2D-CBAM attention mechanism into a 3D structure, and consider the change of depth parameters when extracting channel features and spatial features each time. For the input feature map F 3D , calculate the channel attention feature map M CA and the spatial attention feature map M SA in the order of channel-space:
[0033]
[0034]
[0035] Among them, σ represents the Sigmoid activation function, ω0 and ω1 are weights, and ω0 ∈ R C / r×C , C is the number of channels, and ω1 ∈ R C ×C / r , and for and are shared; and are the two feature descriptors calculated by the feature map F 3D through max-pooling and average-pooling operations respectively, The parameter r usually takes 16; MLP is a hidden layer in the CBAM attention model, that is, a shared multi-layer perceptron; f 7×7×7 represents a 3D convolutional layer of 7×7×7; and are for F' 3D The two feature descriptors calculated by max-pooling and average-pooling operations respectively,
[0036] The output feature map of the attention module is F" 3D :
[0037]
[0038] Add 3D-CBAM modules after the 3D convolution of the Conv1 layer in the appearance flow sub-network and after each residual block of the four layers such as Conv2_x respectively;
[0039] Step 105: Use the color video continuous frame group images as the input of the appearance flow sub-network. Each image is 3 channels × 112 pixels × 112 pixels, and perform model training.
[0040] Step 2: The present invention constructs a motion flow sub-network for VR disease evaluation, which specifically includes the following sub-steps:
[0041] Step 201: Use the FlowNet2.0 algorithm to obtain the optical flow map sequence of each VR color video in the dataset; repeat Step 101 to obtain multiple continuous optical flow map sequence groups of each video; and use the classification result of Step 102 as the ground truth for model training;
[0042] Step 202: Repeat Steps 103 and 104 to construct a motion flow sub-network. Use the images in the continuous optical flow map sequence group as the input of the network. Each image is 2 channels × 112 pixels × 112 pixels, and perform model training to obtain the motion flow sub-network.
[0043] Step 3: In an approach of backend fusion, the outputs of the appearance flow sub-network and the motion flow sub-network are fused by weighted average to obtain the final VR disease assessment result Average_f. The entire dual-modal network structure is as shown in Figure 2 the figure;
[0044]
[0045] where the weight w i ≥ 0, and T = 2, P(x i ) represents the output results of the two sub-networks;
[0046] The two sub-networks independently perform the VR experience comfort classification task. The model training parameters are shown in Table 1:
[0047]
[0048] During testing, the final result is obtained after fusing the prediction scores of the two sub-networks, as shown in Table 2:
[0049]
[0050] Performance comparison between the dual-modal network of the present invention and other methods.
Claims
1. A VR disease assessment method based on a bimodal network, characterized in that: First, based on the 2D-ResNet50 model, all convolutional kernels are extended into 3D convolutional kernels, and appearance flow and motion flow sub-networks are constructed based on 3D-ResNet. Then, 3D convolution is introduced into the 2D-CBAM attention mechanism, and attention modules are added to the two sub-networks respectively to strengthen the features in channels and space. Finally, a weighted average backend fusion method is used to achieve VR disease assessment. The specific steps are as follows: Step 1: Construct an appearance flow sub-network for VR disease assessment, including the following sub-steps: Step 101: Resize each frame image of the 360° VR color video and crop it to a fixed size of 112×112; and decompose the video into multiple consecutive frame groups, each group including consecutive L frame images; Step 102: Divide it into four categories according to the subjective VR disease score value of the VR video: comfortable, slightly uncomfortable, moderately uncomfortable, and severely uncomfortable, and use the classification result as the ground truth for model training; Step 103: Based on the traditional 2D-ResNet50, while keeping the network structure unchanged, expand all convolutional kernels into 3D convolutional kernels so that the convolutional kernels can perform sliding convolutional operations on the time-dimensional image frame sequence; the input sizes of the Conv1, Conv2_x, Conv3_x, Conv4_x, and Conv5_x layers of the improved appearance flow sub-network are: L×112×112, L×56×56, The network structure configurations are respectively: 7×7×7, 64, stride 2, After extracting features layer by layer, connect to a fully connected layer, and use the Softmax classifier and cross-entropy loss function for classification; Step 104: Expand the traditional 2D-CBAM attention mechanism into a 3D structure, and consider the change of depth parameters when extracting channel features and spatial features each time. For the input feature map F 3D , calculate the channel attention feature map M CA and the spatial attention feature map M SA respectively in the order of channel - space: Among them, σ represents the Sigmoid activation function, ω0, ω1 are weights, and ω0 ∈ R C / r×C , C is the number of channels, ω1 ∈ R C×C / r , and for and are shared; and are two feature descriptors calculated by the average pooling and max pooling operations on the feature map F 3D respectively; the parameter r takes 16; MLP is a hidden layer in the CBAM attention model, that is, a shared multi-layer perceptron; f represents a 3D convolutional layer of 7×7×7; 7×7×7 and are two feature descriptors calculated by the average pooling and max pooling operations on F' 3D respectively; The output feature map of the attention module is F" 3D : Add 3D-CBAM modules after the 3D convolution of the Conv1 layer and after each residual block of the four layers of Conv2_x, Conv3_x, Conv4_x, and Conv5_x in the appearance flow sub-network respectively; Step 105: Use the consecutive frame group images of the color video as the input of the appearance flow sub-network, each image being 3 channels × 112 pixels × 112 pixels, and conduct training; Step 2: Construct a motion flow sub-network for VR disease assessment, including the following sub-steps: Step 201: Use the FlowNet2.0 algorithm to obtain the optical flow map sequence of the 360° VR color video; repeat Step 101 to obtain multiple consecutive optical flow map sequence groups for each video; and use the classification result of Step 102 as the ground truth for model training; Step 202: Repeat Steps 103 and 104 to construct the motion flow sub-network, use the images in the consecutive optical flow map sequence group as the input of the network, each image being 2 channels × 112 pixels × 112 pixels, and conduct training; Step 3: Adopt a backend fusion method to perform weighted average fusion on the outputs of the appearance flow sub-network and the motion flow sub-network to obtain the final VR disease assessment result Average_f: Among them, the weight w i ≥ 0, and T = 2, P(x i ) represents the output results of two sub-networks.
Citation Information
Patent Citations
Sign language translation implementation method and device
CN110532912A
Lens movement recognition method based on attention mechanism 3D residual network
CN112016434A
Micro-expression recognition method and system based on optical flow and RGB modal contrast learning
CN113139479A