Motion redirection method based on Resnet-152

By using a deep residual network model based on ResNet-152, user emotions and scene factors are dynamically identified, and the redirection threshold is adjusted in real time. This solves the limitations and insufficient user experience of virtual space roaming in existing technologies, and improves the efficiency and immersion of users' virtual space roaming.

CN121785468APending Publication Date: 2026-04-03ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing redirected walking methods are difficult to implement in real-life virtual space roaming, and the influence of emotions on visual and tactile perception is not fully utilized, resulting in limitations in the use of threshold ranges.

Method used

A deep residual network model based on ResNet-152 is used to dynamically identify users' emotional states. Combined with personalized influence factors and decay factors, the redirection threshold gain is adjusted in real time to adapt to different users and complex environments.

Benefits of technology

It improves the efficiency and immersion of users roaming in virtual space, reduces the discomfort of redirected walking, and enhances users' ability to rotate and translate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121785468A_ABST
    Figure CN121785468A_ABST
Patent Text Reader

Abstract

The Resnet-152-based motion redirection method provided by the invention comprises the following steps: carrying out a reference emotion influence measurement experiment on a user in a virtual space, and counting and fitting threshold ranges of the user in different emotion states; comparing and integrating the actually influenced difference value data of the user with preset standard group difference value data, and calculating a personalized influence factor reflecting the sensitive degree of the user; defining a deep network structure, constructing a deep residual network model based on a ResNet-152 algorithm, and performing training classification on emotion representatives of scene factors to obtain a scene emotion classification model; in the virtual roaming process of the user, dynamic key frame extraction is carried out on a viewport scene, and the ResNet-152 scene emotion classification model is utilized to carry out classification identification on key frames; if the recognition results of continuous multiple frames are consistent, determining a current scene type factor, counting the cumulative change times of the scene type in the previous virtual roaming, determining a recession factor according to the cumulative times, and then entering the next step; if the identification result is not consistent, returning to the previous step to continue extraction and identification; and dynamically integrating the personalized influence factor and the decline factor, calculating an optimal threshold gain at the current moment, and performing real-time deployment on control parameters of the redirection walking method by using the gain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of motion redirection, emotion recognition, and virtual reality technologies, and specifically to a motion redirection method based on ResNet-152. Background Technology

[0002] ROOM-SCALE virtual reality allows users to wear head-mounted displays and freely explore virtual environments through their own physical movements (such as walking), but its fundamental limitation is the need for a large physical space. Redirected walking is a promising method to overcome this limitation. The core idea of ​​redirected walking is to manipulate users' walking and rotating behaviors in ways that are imperceptible to them, thereby creating the feeling of an infinitely vast virtual space within a limited physical space.

[0003] This technology primarily leverages the characteristic that visual perception can, to some extent, correct for tactile sensations when there is a discrepancy between bodily and visual perception. A common method for redirecting walking is to apply gains to walking distance, curvature, or rotation angle in virtual reality and explore the range of gains that the user cannot perceive (i.e., the detection threshold: DT). However, current redirection methods still have certain requirements regarding the layout and size of the physical space, making it difficult to achieve complete redirected walking in everyday life.

[0004] Currently, mainstream research assumes that visual perception plays a decisive role in situations where there is inconsistency between visual and somatosensory perception. Their main research focus is exploring how auditory, tactile, and olfactory senses can influence the dominance of visual perception, thereby enhancing or weakening the effectiveness of redirected walking. One effective method is to provide users with scene-related, physically fixed auditory cues, thus reducing the user's attention to redirection. This method has successfully influenced the threshold range of redirected walking based on auditory and visual cues. However, current research rarely explores the impact of emotions when people perceive inconsistencies between visual and somatosensory perception. Previous studies have demonstrated that emotions play a significant role in user decision-making and affect various parts of the brain, such as the frontal lobe, mid-vena cava, and cerebellum. These perceptual abilities directly influence rotation and translation in normal individuals. Therefore, it is reasonable to believe that emotions affect the threshold range in redirected walking. Furthermore, virtual environments often contain implicit cues, and different environments evoke different emotions in users. Therefore, using the same threshold for virtual roaming tasks in different scenarios has certain limitations. Summary of the Invention

[0005] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a motion redirection method based on Resnet-152 (Residual Network-152 layers).

[0006] Experiments have demonstrated that users' subjective perception of redirection varies under different emotional states, and the range of thresholds they can perceive also differs. When users freely roam in a virtual space, this invention dynamically identifies the user's potential emotional state based on scene factors and dynamically assigns an appropriate threshold to the user according to the corresponding emotional state, applying the threshold to the current redirection algorithm. The impact of scene factors often varies for different users; therefore, this invention sets personalized influence factors to individually adjust the threshold range. In complex environments, the influence of scene factors on users gradually weakens with fluctuations; therefore, this invention records the user's influence changes and sets a decay factor to gradually reduce the impact of scene factors on threshold changes over time, providing users with the most suitable threshold gain during translation and rotation.

[0007] This is a process of generating the optimal redirection threshold based on the user's location and surrounding environment factors.

[0008] This invention is a motion redirection method based on ResNet-152, comprising the following steps: Step 1: Conduct a baseline emotion impact measurement experiment on users in a virtual space, and statistically analyze and fit the threshold range of users under different emotional states; Step 2: Based on the user fitting results in Step 1, compare and integrate the difference data of the actual impact on users with the preset standard group difference data to calculate the personalized impact factor that reflects the user's sensitivity. Step 3: Define the deep network structure, construct a deep residual network model based on the ResNet-152 algorithm, train and classify the emotion representations of scene factors to obtain a scene emotion classification model; Step 4: During the user's virtual roaming, dynamic keyframes are extracted from the viewport scene, and the keyframes are classified and identified using the ResNet-152 scene emotion classification model. Step 5: Determine the consistency of the classification results: If the recognition results of multiple consecutive frames are consistent, determine the current scene type factor, and count the cumulative number of changes of this scene type in the previous virtual roaming. Determine the decay factor based on the cumulative number of changes, and then proceed to step 6; If the recognition results are inconsistent, return to step 4 to continue extraction and recognition. Step 6: Dynamically integrate the personalized influencing factor and the decay factor, calculate the optimal threshold gain at the current moment, and use this gain to deploy the control parameters of the redirection walk method in real time.

[0009] Furthermore, the specific implementation of steps 1 and 2 includes: Step 1.1: Record the raw data of the user's perception threshold under different emotional arousal states by setting virtual guided stimuli; and perform curve fitting on the raw data to remove outliers and establish an accurate correspondence model between the user's emotional intensity and the perception threshold range. Step 1.2: Extract the user's threshold feature value based on the correspondence model, and compare the difference with the group average feature value in the standard database; Step 1.3: Based on the difference ratio obtained from the comparison, the personalized influence factor is generated in a quantitative manner to characterize the user's sensitivity shift relative to the standard group.

[0010] Furthermore, the process of constructing and training the ResNet-152 network structure in step 3 includes: Step 3.1: Build the main architecture of the ResNet-152 network, which includes an input convolutional layer, a four-stage cascaded residual extraction layer, a global average pooling layer, and a fully connected classification layer from the input end. Step 3.2: In the cascaded residual extraction layers of the four stages, bottleneck residual units with stacked numbers of 3, 8, 36 and 3 are configured in sequence to form a feature extraction network with a total depth of 152 layers. Step 3.3: Construct the internal structure of the bottleneck residual unit, which sequentially includes a first 1×1 convolutional layer for dimensionality reduction, a 3×3 convolutional layer for spatial feature extraction, and a second 1×1 convolutional layer for dimensionality recovery, and set an identity mapping skip connection across these three convolutional layers. Step 3.4: Input the scene image dataset with emotion labels into the built network, and use the stochastic gradient descent algorithm to iteratively update the network weights until the loss function converges.

[0011] Furthermore, step 5, concerning the consistency determination and the determination of the degradation factor, includes: Step 5.1: Set the consistency judgment window to the most recent N frame images. The recognition result is considered consistent only when all keyframes in this window are identified as the same scene category by the model. Step 5.2: Establish a scene history database and update the cumulative frequency of occurrence of the currently determined scene type factors in this roaming in real time; Step 5.3: Calculate the decay factor based on the principle of diminishing marginal utility. As the cumulative frequency of occurrence increases, the value of the decay factor is reduced non-linearly to simulate the user's adaptive desensitization to repetitive scenarios.

[0012] Furthermore, the integration and deployment process in step 6 includes: Step 6.1: Using the preset basic redirection threshold as the base, the user baseline threshold is obtained by statically adjusting the base using the personalized influence factor obtained in Step 2. Step 6.2: Using the decay factor obtained in Step 5 as a time-varying adjustment coefficient, dynamically reduce the user baseline threshold and calculate the optimal threshold gain at the current moment. Step 6.3: Apply the optimal threshold gain mapping to the translation gain, rotation gain, or curvature gain parameters of the redirection walk algorithm, and execute the new control strategy in the next frame rendering.

[0013] This invention proposes a redirection threshold adjustment method based on Resnet-152. This method not only significantly improves the user's sense of presence, immersion, and embodiment, but also increases the user's turning and rotating abilities to a certain extent by adjusting the threshold range, thereby reducing the occurrence of resets.

[0014] This invention proposes an innovative ResNet-152-based redirection threshold adjustment method to address the problem of users using a single threshold gain in different virtual scenarios. Firstly, this invention introduces the concepts of personalized influencing factors and decay factors to select the most suitable threshold gain for the user. Secondly, this invention uses ResNet-152 to dynamically detect scene factors, monitoring changes in scene factors in real time and adjusting the user's threshold gain accordingly.

[0015] The advantages of this invention are: it can not only effectively improve the user experience, but also improve the ability to respond promptly to different scenarios, quickly and accurately adjust the threshold range, and improve the user's roaming efficiency. Attached Figure Description

[0016] Figure 1 This is a flowchart for determining the consistency of scenario factors in this invention.

[0017] Figure 2 This is a structural diagram of the Resnet-152 network of the present invention; Figure 3 This is a flowchart of the method of the present invention. Detailed Implementation

[0018] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0019] The motion redirection method based on ResNet-152 includes the following steps: (1) Conduct a simple experiment to measure the impact of emotions on users in virtual space, and statistically analyze and fit the threshold range under different emotions.

[0020] In the rotation and translation experiment, this invention first requires participants to undergo emotional induction, and then, wearing an HMD, rotate (move) their bodies at a certain angle under a certain rotation (translation) gain operation. Next, participants are asked to express their perception of the amount of rotation (movement) in the VR using the 2AFC method. The data is then fitted using a psychometric function (1) to obtain... , The value is determined to find a curve that fits its own characteristics, and at the same time, the present invention estimates the corresponding translation gain range for different emotional states based on the curve. (by a 25% threshold) Lower limit and 75% threshold (Upper limit composition) and rotational gain range (by a 25% threshold) Lower limit and 75% threshold (Upper limit composition).

[0021] (1)

[0022] (2) Based on the user fitting results, the difference in actual impact on the user is integrated with the difference in standard data to calculate the personalized impact factor.

[0023] After obtaining the translation gain range and rotation gain range under different emotions, in order to avoid instability caused by insufficient data, this invention cannot directly apply them as usable gain ranges, but can only use them as reference samples. The personalized influencing factor is determined by integrating the relationship between the measured data and the standard data. The specific determination formulas are as shown in (2) and (3), using the translation gain personalized influencing factor of positive emotions. For example, firstly, we statistically measure the positive emotions under simple conditions. The difference was calculated using the previously fitted upper and lower limits. Similarly, the normal emotional state under simple measurement was calculated. Standard measurements of positive emotions and the normal emotional state measured by standard By comparing the threshold difference under simple measurement with the threshold difference under standard conditions, the degree to which a user will be affected by scene factors during this virtual roaming can be determined, and the results are more individualized.

[0024] (2) (3) (3) Define the network structure and use the Resnet-152 algorithm to train and classify the emotion representatives of scene factors.

[0025] The ResNet-152 architecture is as follows: starting from the input layer, it sequentially connects to an input convolutional layer, four stages of residual extraction layers, a global average pooling layer, and a fully connected classification layer. Each input convolutional layer is followed by a max pooling layer; the four stages of residual extraction layers are each composed of stacked bottleneck residual units, with the first residual unit in each stage being downsampled upon input; the global average pooling layer is followed by a fully connected classification layer, and finally, the classification result is output through a Softmax activation function. These four stages of residual extraction layers are designated as Stage 1, Stage 2, Stage 3, and Stage 4.

[0026] The input convolutional layer and max pooling layer are used for preliminary feature extraction and dimensionality reduction of the input image; the four-stage residual extraction layer is used to extract deep and abstract feature information of the input image step by step, and solves the gradient vanishing problem in deep network training through residual connections; the global average pooling layer is used to compress the spatial information of each feature channel into a single value, significantly reducing the number of model parameters and preventing overfitting; the fully connected classification layer is used to map the extracted high-dimensional features to the category space of the samples, and the output of this layer is a one-dimensional vector containing the probability of scene emotion category; the softmax activation function is used to transform the output of the fully connected layer into a probability distribution to achieve scene emotion classification.

[0027] Furthermore, based on the data size of the input image of the ResNet-152 network (e.g., 224x224x3), the kernel size, stride, number of kernels, and number of stacked residual units for each layer are designed.

[0028] The scene dynamic recognition method based on ResNet-152 proposed in this embodiment classifies emotional attributes by extracting deep visual features of scene factors in images. Therefore, for the input virtual roaming scene image, the size of the convolutional kernels and stride in the ResNet-152 network is designed according to the image resolution and feature extraction requirements. In this embodiment, residual neural networks possess powerful feature learning capabilities, especially their deep structure, which can capture extremely abstract and complex visual patterns. The residual unit is the core of ResNet, mainly used to maintain effective gradient propagation while increasing network depth.

[0029] like Figure 2 As shown, the established ResNet-152 network includes: one initial convolutional layer (Conv1), one max pooling layer (MaxPool), four residual stages (Stages 1-4), one global average pooling layer (Global Avg Pool), one fully connected layer (FC), and one softmax activation layer. Among them: Stage 1 contains 3 stacked bottleneck residual units; Stage 2 contains 8 stacked bottleneck residual units; Stage 3 contains 36 stacked bottleneck residual units; Stage 4 contains 3 stacked bottleneck residual units.

[0030] Each bottleneck residual unit contains a series of sequentially connected 1x1 convolutional layers (dimensionality reduction), 3x3 convolutional layers (feature extraction), and 1x1 convolutional layers (dimensionality increase), as well as identity mapping skip connections across these three layers. After all convolutional layers (except fully connected layers), there are batch normalization (BN) layers and ReLU activation functions.

[0031] Through this deep residual structure, the network can effectively extract complex emotion-related features from scene images. Simultaneously, the use of global average pooling layers replaces multiple traditional fully connected layers, significantly reducing model parameters and improving training speed and generalization ability. By combining the input image size of the ResNet-152 network and appropriately designing the number of convolutional kernels (256, 512, 1024, 2048 respectively) and stride at each stage, different types of scene emotion features can be effectively extracted and distinguished. Therefore, based on the differences in the extracted deep features, the scene emotion category to which the input keyframe belongs can be identified.

[0032] In this embodiment, the specific parameters for designing the ResNet-152 network are shown in Table 1: Table 1 Resnet-152 Network Parameters

[0033] Finally, the scene image dataset with emotion labels is input into the constructed ResNet-152 network for training to obtain the trained ResNet-152 network; and the trained ResNet-152 network is validated using the validation set to obtain the final ResNet-152 network. (4) The key frames of the virtual roaming scene are dynamically extracted and classified by the Resnet-152 algorithm.

[0034] During virtual roaming, the scene information captured by a user's gaze is dynamically changing, and consequently, the user's emotional state also changes dynamically. To address this real-time and dynamic behavior, this system employs methods such as... Figure 1The process architecture of this invention is as follows: during the user's virtual roaming, the scene information captured by the HMD device is used to segment the scene in real time, the key frames of the segmented scene are extracted, and each frame is handed over to the ResNet-152 model for real-time judgment to obtain the type score sequence of each frame. After normalizing the multi-frame sequence, this invention will obtain the total type probability sequence.

[0035] (5) If multiple consecutive frames are consistent, determine the scene type factor, count the number of scene factor changes that occurred in the previous virtual roaming, determine the decay factor, and proceed to step (6); otherwise, return to step (4).

[0036] If the ResNet-152 model determines that the scene has high consistency, i.e., the score distribution is concentrated and the score of a single class is higher than 0.9, then the number of changes in the scene for the corresponding class is taken as the score. Simultaneously, it increments, and the decay factor of this invention is obtained based on the number of changes. The formula is as follows (4) (Here, we also take the translation gain of positive emotions as an example). If the Resnet-152 model determines that the scene does not have high consistency, that is, the score distribution is scattered and there is no single category score higher than 0.9, then return to step 4 to perform frame extraction judgment.

[0037] (4)

[0038] (6) Dynamically integrate the user's influence factors and decay factors, determine the most suitable threshold gain, and deploy the redirection method.

[0039] The recession factors generated above and personalized influencing factors By combining this with the standard threshold range, we can obtain the lower limit of the threshold that the user will actually use. and threshold upper limit The specific generation formulas are as follows (5) (6) (Here, we also take the translation gain of positive emotions as an example). Finally, the actual threshold is deployed to the redirection method to guide users.

[0040] (5) (6) The embodiments described in this specification are merely examples of implementations of the inventive concept. The scope of protection of this invention should not be considered as limited to the specific forms stated in the embodiments. The scope of protection of this invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.

Claims

1. A motion redirection method based on ResNet-152, comprising the following steps: Step 1: Conduct a baseline emotion impact measurement experiment on users in a virtual space, and statistically analyze and fit the threshold range of users under different emotional states; Step 2: Based on the user fitting results in Step 1, compare and integrate the difference data of the actual impact on users with the preset standard group difference data to calculate the personalized impact factor that reflects the user's sensitivity. Step 3: Define the deep network structure, construct a deep residual network model based on the ResNet-152 algorithm, train and classify the emotion representations of scene factors to obtain a scene emotion classification model; Step 4: During the user's virtual roaming, dynamic keyframes are extracted from the viewport scene, and the keyframes are classified and identified using the ResNet-152 scene emotion classification model. Step 5: Determine the consistency of the classification results: If the recognition results of multiple consecutive frames are consistent, determine the current scene type factor, and count the cumulative number of changes of this scene type in the previous virtual roaming. Determine the decay factor based on the cumulative number of changes, and then proceed to step 6; If the recognition results are inconsistent, return to step 4 to continue extraction and recognition. Step 6: Dynamically integrate the personalized influencing factor and the decay factor, calculate the optimal threshold gain at the current moment, and use this gain to deploy the control parameters of the redirection walk method in real time.

2. The motion redirection method based on ResNet-152 as described in claim 1, characterized in that, The specific implementation of steps 1 and 2 includes: Step 1.1: Record the raw data of the user's perception threshold under different emotional arousal states by setting virtual guided stimuli; and perform curve fitting on the raw data to remove outliers and establish an accurate correspondence model between the user's emotional intensity and the perception threshold range. Step 1.2: Extract the user's threshold feature value based on the correspondence model, and compare the difference with the group average feature value in the standard database; Step 1.3: Based on the difference ratio obtained from the comparison, the personalized influence factor is generated in a quantitative manner to characterize the user's sensitivity shift relative to the standard group.

3. The motion redirection method based on ResNet-152 as described in claim 1, characterized in that, Step 3, the process of constructing and training the ResNet-152 network, includes: Step 3.1: Build the main architecture of the ResNet-152 network, which includes an input convolutional layer, a four-stage cascaded residual extraction layer, a global average pooling layer, and a fully connected classification layer from the input end. Step 3.2: In the cascaded residual extraction layers of the four stages, bottleneck residual units with stacked numbers of 3, 8, 36 and 3 are configured in sequence to form a feature extraction network with a total depth of 152 layers. Step 3.3: Construct the internal structure of the bottleneck residual unit, which sequentially includes a first 1×1 convolutional layer for dimensionality reduction, a 3×3 convolutional layer for spatial feature extraction, and a second 1×1 convolutional layer for dimensionality recovery, and set an identity mapping skip connection across these three convolutional layers. Step 3.4: Input the scene image dataset with emotion labels into the built network, and use the stochastic gradient descent algorithm to iteratively update the network weights until the loss function converges.

4. The motion redirection method based on ResNet-152 as described in claim 1, characterized in that, Step 5, concerning the consistency assessment and the determination of the decline factor, includes: Step 5.1: Set the consistency judgment window to the most recent N frame images. The recognition result is considered consistent only when all keyframes in this window are identified as the same scene category by the model. Step 5.2: Establish a scene history database and update the cumulative frequency of occurrence of the currently determined scene type factors in this roaming in real time; Step 5.3: Calculate the decay factor based on the principle of diminishing marginal utility. As the cumulative frequency of occurrence increases, the value of the decay factor is reduced non-linearly to simulate the user's adaptive desensitization to repetitive scenarios.

5. The motion redirection method based on ResNet-152 as described in claim 1, characterized in that, The integration and deployment process in step 6 includes: Step 6.1: Using the preset basic redirection threshold as the base, the user baseline threshold is obtained by statically adjusting the base using the personalized influence factor obtained in Step 2. Step 6.2: Using the decay factor obtained in Step 5 as a time-varying adjustment coefficient, dynamically reduce the user baseline threshold and calculate the optimal threshold gain at the current moment. Step 6.3: Apply the optimal threshold gain mapping to the translation gain, rotation gain, or curvature gain parameters of the redirection walk algorithm, and execute the new control strategy in the next frame rendering.