Redirection walking method and system based on deep reinforcement learning, terminal and storage medium

Through deep reinforcement learning, the feasible entropy of the real space is calculated, and the walking paths in virtual reality are dynamically adjusted, which solves the problem of high collision between users and obstacles, and improves the smoothness and immersion of virtual reality roaming.

CN120259595APending Publication Date: 2025-07-04SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510335787.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing redirection walking algorithm has a high chance of collision with obstacles in virtual reality, and lacks flexibility and ability to adapt to complex environments, resulting in poor roaming experience.

Method used

Using a method based on deep reinforcement learning, the user's location and virtual space state are calculated in real time, and the walking path is dynamically adjusted. The user's walking path is optimized to avoid collisions by calculating the spatial feasible entropy of the real space.

Benefits of technology

Effectively reduce the probability of collision between users and physical boundaries, improve immersion experience and roaming fluency, adapt to complex physical space layout, and reduce VR vertigo.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259595A_ABST
    Figure CN120259595A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of virtual reality redirection, and particularly provides a redirection walking method and system based on deep reinforcement learning, a terminal and a storage medium, and the method comprises the steps: obtaining the physical positions of a user in a real space and a virtual space in real time; dividing a walkable area of a user in a real space, and calculating a space feasible entropy in the real space; calculating a real-virtual environment state; based on the real-virtual environment state and the space feasible entropy, reward functions of deep reinforcement learning are calculated, and the reward functions comprise a position reward function, a space alignment reward function and a reset penalty function; judging whether the walking state of the user needs to be reset or not, and if not, using a position reward function and a space alignment reward function; and if so, calculating the optimal reset direction, and guiding the walking path of the user to the optimal safe area. According to the invention, the collision probability between the user and the obstacle is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of virtual reality redirection, and particularly relates to a redirected walking method, system, terminal and storage medium based on deep reinforcement learning. Background Art

[0002] In virtual reality, when a user experiences a virtual world through a virtual device, if they can walk realistically, it will significantly enhance the immersion and realism. This experience of real walking can make virtual roaming more natural, and users can perceive the surrounding environment and scene changes, thereby improving the quality of the virtual reality experience. Currently, various interaction methods have been developed in virtual reality technology. Among them, using real walking as an input method to control the character in the virtual environment has become a research hotspot. This technology enables users to roam naturally in the virtual space.

[0003] In virtual reality, the real walking behavior of the user needs to be mapped to the movement in the virtual world. Traditional walking mapping methods use "equivalent mapping" to convert the user's position, corresponding the position of the user in the physical world with the position in the virtual space one by one. However, this method has problems, especially when the sizes of the virtual space and the real space do not match. For example, when the area of the real space is inconsistent with the area of the virtual world, it becomes very difficult to accurately map the user's walking behavior to the virtual environment. This results in the spatial mapping often not being very accurate, affecting the user experience.

[0004] Currently, the main solution is to adopt the Redirected Walking (RDW) technology, which actively adjusts the user's real movement trend without the user's awareness by adding subtle offsets between the user's virtual and real movements, so as to achieve roaming in a small real space within a large virtual scene. Limited by the user's sensitive perception ability, the movement deviation adopted by redirected walking can only adjust the user's movement to a certain extent, resulting in the user often colliding with the boundary of the real space. The common solution is to use a reset strategy for explicit intervention to force the user to return to the physical space, but this also leads to the interruption of the roaming experience. To reduce collisions, a redirected walking algorithm is proposed, which integrates multiple gains through various strategies to improve the movement control ability of redirected walking. Although the redirected walking algorithm can optimize the redirection strategy of the roaming process, there are still certain defects, which are as follows: First, the traditional redirected walking algorithm mainly adopts a fixed guidance mode. This method can achieve the docking of virtual and reality to a certain extent, but it lacks flexibility and cannot effectively adapt to the changes of the physical space or dynamic adjustment. For example, when the environment or physical space where the user is located changes, this fixed guidance method may not be able to adjust in time, resulting in poor performance of the algorithm in complex and changing space layouts. Second, the guidance strategy is optimized by the characteristics of the physical space or virtual space. For example, adjustments are made by analyzing the geometric shape of the space, the distribution of obstacles, etc. However, in the face of complex environments or changing space layouts, there is a lack of sufficient detailed representation, which may limit the user's walking experience and prevent the best obstacle avoidance and walking path planning in various environments. Third, when the user approaches the boundary of the physical space, the traditional redirected walking algorithm usually adopts a reset strategy to forcibly adjust the user's walking direction or position to ensure that the user does not encounter obstacles. However, this reset strategy relies on a fixed mode or simple direction adjustment and does not consider the complex physical space layout. When the real space layout is relatively complex, these strategies may cause the user to collide with obstacles again soon after being reset, interrupting the user's walking process and affecting the user experience. Summary of the Invention

[0005] Aiming at the problem that the probability of collision between the user and obstacles in roaming using the redirected walking algorithm in the prior art is high, the present invention provides a redirected walking method, system, terminal and storage medium based on deep reinforcement learning to solve the above technical problems.

[0006] In a first aspect, the present invention provides a redirected walking method based on deep reinforcement learning, including:

[0007] Obtain the physical position of the user in the real space and the physical position in the virtual space in real time;

[0008] Divide the walkable area of the user in the real space based on the user's physical position in the real space, and calculate the spatial feasibility entropy in the real space based on the walkable area;

[0009] Calculate the real-virtual environment state based on the user's physical position in the real space and the user's physical position in the virtual space;

[0010] Calculate the reward function of deep reinforcement learning based on the real-virtual environment state and the spatial feasibility entropy, where the reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function;

[0011] Determine whether the walking state of the user needs to be reset. If it does not need to be reset, use the position reward function and the spatial alignment reward function; if it needs to be reset, use the reset penalty function to calculate the optimal reset direction and guide the user's walking path to the optimal safe area within the walkable area.

[0012] Further, dividing the walkable area of the user in the real space based on the user's physical position in the real space includes:

[0013] Define the real space S P , define the user's physical position x in the real space P , where x P ∈S P ;

[0014] Construct a visibility polygon VP(x P ), where the polygon includes all points visible from x P ;

[0015] Establish polar coordinates and model the visibility polygon VP(x P ) in the polar coordinates to obtain the walkable area.

[0016] Further, calculating the spatial feasibility entropy in the real space based on the walkable area includes:

[0017] Discretize the walkable area and calculate the free walking entropy e free , spatial blockage entropy e obs and distance penalty function f dis ;

[0018] Based on the free walking entropy e free , spatial blockage entropy e obs and distance penalty function f dis calculate the spatial feasibility entropy E sw of the user's physical position.

[0019] Further, the feasible walking area is discretized, and the free walking entropy e of the discretized feasible walking area is calculated free , the space blocking entropy e obs and the distance penalty function f dis , and based on the free walking entropy e free , the space blocking entropy e obs and the distance penalty function f dis , the space feasible entropy E of the user's physical position is calculated sw , including:

[0020] Taking as the step size, sample the visibility polygon VP(x P ), and discretely divide the visibility polygon VP(x P ) into n consecutive visible triangles {VT1(x P ), VT2(x P ), …, VT n (x P ));

[0021] Model the walking freedom under x P , and define the probability p i for the user to transition to a specific VT P (x i ) as:

[0022]

[0023] where A(·) represents the polygon area function;

[0024] Based on , obtain the free walking entropy e P at x free :

[0025]

[0026] where e free represents the uniformity of the motion probability distribution from one position to the surrounding area;

[0027] Model the space blockage under x P to quantify the blockage degree of each sampled VT i (x P ) and the change of the blockage degree in different sampling directions within VP(x P ), and the space blocking entropy e P at x obs is defined as:

[0028]

[0029] where eobs Indicates the relative degree of obstacle for the user's movement in various directions from x P for safety comparison of different physical locations;

[0030] Use the distance penalty function f dis to adjust the entropy of x P The distance penalty function f dis is defined as follows:

[0031]

[0032] where d(·) represents the Euclidean distance between x P and the nearest obstacle, and represent the walkable area and the blocked area respectively, and λ represents the scaling factor, with λ = 0.1 set;

[0033] Based on the free - walking entropy e free 、spatial - blockage entropy e obs and the distance penalty function f dis calculate the spatial - feasibility entropy E sw of the user's physical location:

[0034] E sw (x P ) = e free (x P )e obs (x P )f dis (x P )(5).

[0035] Furthermore, based on the user's physical location in the real space and the virtual space, calculate the real - virtual environment state, including:

[0036] Define as the two - dimensional position of the user in the virtual space S V at time t, as the two - dimensional forward direction in the virtual space S V , define as the two - dimensional position of the user in the real space S P at time t, as the two - dimensional forward direction in the real space S V , define the motion state as

[0037] At and uniformly emit k rays to detect the distance between the user and the surrounding environment, and use the real - virtual projection function to project and and Align, and compare the virtual ray and the real ray at each emission angle, where the emission angle

[0038] Calculate the safety of the user at a certain θ i :

[0039]

[0040] where w s represents the relative distance measure, represents the distance from the two-dimensional position in the real space along the direction θ i to the obstacle, represents the distance from the two-dimensional position in the virtual space along the direction θ i to the obstacle, and when it means walking safely;

[0041] Calculate the real-virtual environment state ε t :

[0042]

[0043] where is the diagonal length of S P .

[0044] Furthermore, based on the real-virtual environment state and the spatial feasibility entropy, calculate the reward function of the deep reinforcement learning, and the reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function, including:

[0045] Calculate the reward function of the deep reinforcement learning:

[0046]

[0047] where R P represents the position reward function, R a represents the spatial alignment reward function, R c represents the reset penalty function; the position reward function R P is:

[0048]

[0049] where is the walking speed of the user in the real space, v min is the minimum movement threshold;

[0050] The calculation steps of the spatial alignment reward function R a are:

[0051] Define the distance weight w dis :

[0052]

[0053] Based on the distance weight w dis Calculate the alignment factor W align :

[0054]

[0055] Based on the alignment factor W align , calculate the spatial alignment reward function R a :

[0056]

[0057] wherein, represents the rotation speed of the user;

[0058] The reset penalty function R c is:

[0059]

[0060] wherein, v max is the maximum walking speed.

[0061] Furthermore, use the reset penalty function to calculate the optimal reset direction, and guide the user's walking path to the optimal safe area within the walkable area, including:

[0062]

[0063] In a second aspect, the present invention provides a redirected walking method system based on deep reinforcement learning, including:

[0064] A position acquisition module, configured to acquire the physical position of the user in the real space and the physical position in the virtual space in real time;

[0065] A spatial feasibility entropy module, configured to divide the walkable area of the user in the real space based on the physical position of the user in the real space, and calculate the spatial feasibility entropy in the real space based on the walkable area;

[0066] A state calculation module, configured to calculate the real-virtual environment state based on the physical position of the user in the real space and the physical position in the virtual space;

[0067] A reward function calculation module, configured to calculate the reward function of deep reinforcement learning based on the real-virtual environment state and the spatial feasibility entropy, and the reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function;

[0068] A reset calculation module, configured to determine whether the walking state of the user needs to be reset. If no reset is required, a position reward function and a spatial alignment reward function are used; if a reset is required, a reset penalty function is used to calculate an optimal reset direction, and the user's walking path is guided to an optimal safe area within the walkable area.

[0069] In a third aspect, a terminal is provided, including:

[0070] A processor and a memory, where

[0071] the memory is used to store a computer program,

[0072] the processor is configured to call and run the computer program from the memory, so that the terminal executes the method of the above terminal.

[0073] In a fourth aspect, a computer storage medium is provided. Instructions are stored in the computer-readable storage medium, and when it runs on a computer, the computer is made to execute the methods described in the above aspects.

[0074] The beneficial effects of the present invention are as follows. The redirection walking method, system, terminal and storage medium based on deep reinforcement learning provided by the present invention effectively reduce the collision probability between the user and the physical boundary by calculating the spatial feasible entropy in the real space, and while aligning the virtual and the physical, dynamically adjust the user's walking path, so that the user always faces a more suitable area for walking, and can adjust the walking speed, direction and path without the user's awareness, making the walking smoother and more natural, reducing VR dizziness, improving the immersion experience, and guiding the user to a more open direction when approaching the physical boundary through the optimal reset direction, thereby reducing secondary collisions, enhancing the fluency and freedom of VR roaming, and thus being able to adapt to complex physical space layouts. In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very broad application prospect. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0076] Figure 1 is a schematic flowchart of the method of an embodiment of the present invention.

[0077] Figure 2 is a schematic block diagram of the system of an embodiment of the present invention.

[0078] Figure 3 is a schematic structural diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners

[0079] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0080] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.

[0081] The method for redirected walking based on deep reinforcement learning provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the system for the method for redirected walking based on deep reinforcement learning runs in the computer device.

[0082] Figure 1 is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1 The execution subject may be a system for the method for redirected walking based on deep reinforcement learning. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0083] For the convenience of understanding the present invention, the following further describes the method for redirected walking based on deep reinforcement learning provided by the present invention with reference to the principle of the method for redirected walking based on deep reinforcement learning of the present invention and the process of managing pluggable module materials in the embodiments.

[0084] Specifically, as Figure 1 shown, the method for redirected walking based on deep reinforcement learning includes:

[0085] S1. Real-time obtain the physical position of the user in the real space and the physical position of the user in the virtual space.

[0086] S2. Divide the walkable area of the user in the real space based on the physical position of the user in the real space, and calculate the spatial walkable entropy in the real space based on the walkable area.

[0087] Define the real space S P , define the physical position x of the user in the real space P , where x P ∈S P ; construct the visibility polygon VP(xP ), the polygon includes all visible points from x P ; establish polar coordinates, and model the visibility polygon VP(x P ) in polar coordinates to obtain the walkable area.

[0088] Discretize the walkable area and calculate the free walk entropy e free , spatial block entropy e obs and distance penalty function f dis ; calculate the spatial feasible entropy E free of the user's physical location based on the free walk entropy e obs , spatial block entropy e dis and distance penalty function f sw .

[0089] At step size, sample the visibility polygon VP(x P ), and discretely divide the visibility polygon VP(x P ) into n consecutive visible triangles {VT1(x P ), VT2(x P ), …, VT n (x P )};

[0090] Model the walking freedom at x P , and define the probability p i that the user transitions to a specific VT P (x i ) as:

[0091]

[0092] where A(·) represents the polygon area function;

[0093] Based on obtain the free walk entropy e P at x free :

[0094]

[0095] where e free represents the uniformity of the motion probability distribution from one position to the surrounding area;

[0096] Model the spatial block at x P to quantify the block degree of each sampled VT i (x P ) and the change of the block degree in different sampling directions within VP(x P ), xP The spatial blockage entropy e at obs is defined as:

[0097]

[0098] where e obs represents the relative obstacle degree for the user to move in various directions from x P and is used for safe comparison of different physical positions;

[0099] The distance penalty function f dis is used to adjust the entropy of x P . The distance penalty function f dis is defined as follows:

[0100]

[0101] where d(·) represents the Euclidean distance between x P and the nearest obstacle, and represent the walkable area and the blocked area respectively, and λ represents the scaling factor, with λ set to 0.1;

[0102] Based on the free - walking entropy e free , the spatial blockage entropy e obs and the distance penalty function f dis , calculate the spatial feasible entropy E of the user's physical position sw :

[0103] E sw (x P ) = e free (x P )e obs (x P )f dis (x P )(5).

[0104] S3. Calculate the real - virtual environment state based on the user's physical position in the real space and the physical position in the virtual space.

[0105] Define as the two - dimensional position of the user in the virtual space S V at time t, as the two - dimensional forward direction in the virtual space S V . Define as the two - dimensional position of the user in the real space S P at time t, as the two - dimensional forward direction in the real space S V . Define the motion state as

[0106] at and uniformly emit k-rays to detect the distance between the user and the surrounding environment, and use the real-virtual projection function to with and align, compare the virtual rays and the real rays at each emission angle, where the emission angle

[0107] calculate the safety of the user at a certain θ i :

[0108]

[0109] where w s represents the relative distance measure, represents the distance from the two-dimensional position in the real space along the direction θ i to the obstacle, represents the distance from the two-dimensional position in the virtual space along the direction θ i to the obstacle, and when it means safe to walk;

[0110] calculate the real-virtual environment state ε t :

[0111]

[0112] where is the diagonal length of S P .

[0113] S4. Calculate the reward function of deep reinforcement learning based on the real-virtual environment state and the spatial feasibility entropy. The reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function.

[0114] Calculate the reward function of deep reinforcement learning:

[0115]

[0116] where R P represents the position reward function, R a represents the spatial alignment reward function, R c represents the reset penalty function;

[0117] The position reward function R P is:

[0118]

[0119] where is the walking speed of the user in the real space, v min is the minimum movement threshold;

[0120] The spatial alignment reward function R a is calculated as follows:

[0121] Define the distance weight w dis :

[0122]

[0123] Based on the distance weight w dis calculate the alignment factor W align :

[0124]

[0125] Based on the alignment factor W align , calculate the spatial alignment reward function R a :

[0126]

[0127] where, represents the rotational speed of the user;

[0128] The reset penalty function R c is:

[0129]

[0130] where, v max is the maximum walking speed.

[0131] S5. Determine whether the walking state of the user needs to be reset. If it does not need to be reset, use the position reward function and the spatial alignment reward function. If it needs to be reset, use the reset penalty function to calculate the optimal reset direction and guide the user's walking path to the optimal safe area within the walkable area.

[0132] Using the reset penalty function to calculate the optimal reset direction and guide the user's walking path to the optimal safe area within the walkable area, including:

[0133]

[0134] Specifically, this application defines the walkable area of the user's current physical position and discretizes it. Here, for the physical space s P , define the user position x P ∈ S P , we determine the walkable area by constructing the visibility polygon VP(x P ), and this polygon contains from xP All visible points. We model VP(x P ) in polar coordinates and sample VP(x s ) at a step size of θ P = π / 90, discretely dividing it into n consecutive visible triangles {VT1(x P ), VT2(x P ), …, VT n (x P )}. In addition, we use an occupancy grid map to describe the degree of blockage at different positions in S P .

[0135] This application calculates the free walking entropy e free . We model the freedom of movement at x P . First, define the probability p i that the user transitions to a specific VT P (x i ) as:

[0136]

[0137] where A(·) represents the polygon area function. According to Equation (1), the free walking entropy e P at x free is defined as:

[0138]

[0139] The free walking entropy value calculated using Equation (2) is higher in open areas than in closed areas. However, e free only represents the uniformity of the probability distribution of movement from one position to the surrounding area. In this case, positions close to obstacles and positions far from obstacles may produce similar entropy values. Therefore, we further impose additional constraints on VP(x P ) to describe the obstacle level. We model the spatial blockage entropy e obs to quantify the degree of blockage of each sampled VT i (x P ) and take into account the variation in the degree of blockage in different sampling directions within VP(x P ). The spatial blockage entropy e P at x obs is defined as:

[0140]

[0141] Through Equation (3), the degree to which the user can move from x PThe relative degree of obstacle to movement in all directions is used to safely compare different physical positions. However, for actual VR roaming, determining the actual safe navigation range is crucial. Therefore, we use a distance penalty function f dis to adjust the entropy of x P so that it decreases near obstacles and remains stable far from obstacles. The formula for f dis is defined as follows:

[0142]

[0143] where d(·) calculates the Euclidean distance between x P and the nearest obstacle, and represent the walkable area and the blocked area respectively, and λ represents the scaling factor. Here, we set λ = 0.1.

[0144] This application calculates the spatial walking entropy E sw . According to equations (2), (3), and (4), the spatial walking entropy E sw is defined as follows:

[0145] E sw (x P ) = e free (x P )e obs (x P )f dis (x P )(5)

[0146] Through equation (5), we can effectively evaluate the feasibility and safety of each physical position according to the spatial structure and layout.

[0147] This application proposes a DRL-based spatial walking ability perception redirected walking algorithm. The algorithm learns two key aspects: the physical SWE distribution and the alignment between virtual and physical user states. The training of our algorithm utilizes an improved PPO neural network. We enhance the prediction ability of the agent by adding a shared feature extraction head with a hidden layer and a GRU layer before the policy and value networks. Specifically, the agent is trained in a simulation environment using pre-generated virtual trajectories. By allowing the simulated user to explore various trajectories, our agent learns the optimal policy through trial and error with different gains under a specific S P To provide sufficient learning information, we extract various features of the user's virtual physical state to feed into the neural network.

[0148] This application models the environmental state. Define and is the user's two-dimensional position in the virtual space S at time t V and the physical space S P , and and represents the two-dimensional forward direction, and the motion state is expressed as We detect the distance between the user and the surrounding environment by uniformly emitting k rays from and . We apply the virtual physical projection function to align with and , allowing the comparison of virtual and physical rays at each emission angle θ i = 2πi / k (0 ≤ i < k). Obviously, when , the walking safety is guaranteed. The real-time feasible walking safety w of a specific θ i is defined as the relative distance measure, and its formula is as follows: s

[0149]

[0150] The environmental state ε t is defined as follows:

[0151]

[0152] where is the diagonal length of S P . Here, we set k = 18 to sample the surrounding environmental conditions as detailed as possible. In addition, we normalize all position and distance data of u t by dividing them by the respective diagonal lengths of the relevant spaces to ensure data consistency between u t and ε t . We use the state stack to form a complete observation o t by connecting 10 consecutive past states:

[0153] o t = [(u t-9 , ε t-9 ), (u t-8 , ε t-8 ), …, (u t , ε t )](8)

[0154] We set the action output by the policy network as the applied redirection gain to allow the agent to directly control the redirection, where and Adjust the translation, rotation, and curvature gain separately. We use the existing detection threshold to obtain the redirection gain: The symbol indicates the user's turning direction. In this application, we utilize action repetition to improve learning. Specifically, a is executed in the next 5 consecutive steps t , but the individual rewards for each time step are cumulative. Subsequently, the average value is used as the overall reward of a t and fed back to the agent until the next action at a t+5 is determined.

[0155] For the reward function R of DRL, we model it by calculating the position reward, spatial alignment reward, and reset penalty. Given u t and ε t , the single-step reward R is defined as:

[0156]

[0157] where R P represents the position reward, R a represents the spatial alignment reward, and R c represents the reset penalty.

[0158] For the position reward, since there are large magnitude differences in the E sw values at different positions, directly using them for reward calculation may cause severe oscillations. Therefore, we apply min-max normalization to the E sw values of each sampled physical position to standardize the distribution, so that all sampled E sw converge between 0 and 1, and thus obtain the normalized entropy distribution map E sw (S P ). We use the average entropy for R P calculation to ensure that the overall position reward distribution is centered around 0, and its formula is as follows:

[0159]

[0160] where is the physical walking speed, and v min is the minimum movement threshold. In this application, we set v min = 0.2 m / s.

[0161] This application models the spatial alignment reward. We uniformly emit k rays to sample the user's surrounding environment, and these rays jointly summarize all potential turning directions of the user Considering that the user tends to maintain the current walking direction, so θ closer to the user's current direction ishould have higher importance in spatial alignment. In this case, we model the direction weight w dir as follows:

[0162]

[0163] Based on this, we further estimate the user's actual walking ability in each sampling direction. Using Equation (6), the relative walking safety of each θ i can be determined, thus giving the derivation of the overall spatial orientation. However, Equation (6) is not sufficient to describe the level of navigation safety, because even if is relatively large, but under conditions, it will produce a very small value. To solve this problem, we compare with the minimum safety distance l s to determine the absolute safety margin. In addition, when is large enough to meet the necessary safety requirements, the increase in the safety level should gradually decrease. Therefore, we combine the relative and absolute safety to define the distance weight w dis as:

[0164]

[0165] Here, set l s = 2m. According to Equation (12), if the given θ i has a high absolute safety level then a higher safety weight can be derived; conversely, its safety weight is mainly determined by the relative safety level, which is consistent with the safety requirements in RDW. Using Equations (11) and (12), we define the alignment factor W i for each θ align as:

[0166]

[0167] Furthermore, the spatial alignment reward R a is finally defined as:

[0168]

[0169] where, represents the user's rotational speed. Equation (13) is used to attenuate the reward in Equation (14). As the level of consistency increases, the negative reward generated by Equation (14) decreases. In addition, we use as the base term to ensure that the magnitudes of R a and R P are comparable.

[0170] When a reset occurs, we apply the reset penalty function R cWe will R c The definition is as follows:

[0171]

[0172] where v max As the maximum walking speed, we set v max =2m / s, It represents the maximum intersection distance between the light rays emitted by all sampling points in SP and the surrounding physical obstacles. Formula (15) increases the reset penalty in open physical space and appropriately reduces the reset penalty in crowded physical space.

[0173] This application models the reset optimization. Given the user's current walkable area Potential reset direction is from arrive Any point x in a Here, x a The degree of openness is determined by its absolute safety area C(x a ) to estimate, that is, x a The area of ​​the maximum inscribed circle with C(x2)=C(x3)>(x1). Because C(x2)=C(x3)>(x1), C(x2) and C(x3) can provide higher walking safety than C(x1). Therefore, resetting to x2 or x3 is a better choice. When the absolute safety areas are of equal size, the feasibility outside the absolute safety areas at different center positions is different. For example, C(x2) is located near multiple physical boundaries, which means that the feasibility outside C(x2) is low and the user may collide with the boundary again after leaving C(x2). In contrast, the area around C(x3) is more open, reducing the possibility of the user colliding again after leaving C(x3). Therefore, the absolute safety area with a higher total SWE provides a stronger expectation of uninterrupted walking and is more suitable as a turning target for resetting. In this case, the optimal reset direction V opt Defined as:

[0174]

[0175] By utilizing V opt ,We redirect users to the most feasible and safe direction within the current feasible area, thereby enhancing the continuous walking distance after the reset.

[0176] In some embodiments, the redirection walking method system based on deep reinforcement learning may include multiple functional modules composed of computer program segments. The computer program of each program segment in the redirection walking method system based on deep reinforcement learning may be stored in the memory of a computer device and executed by at least one processor to perform (see Figure 1Description) Function of the redirected walking method based on deep reinforcement learning.

[0177] In this embodiment, the system of the redirected walking method based on deep reinforcement learning can be divided into multiple functional modules according to the functions it performs, such as Figure 2 shown. The functional modules of the system 200 may include: a position acquisition module 210, a spatial feasibility entropy module 220, a state calculation module 230, a reward function module 240, and a reset calculation module 250. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0178] The position acquisition module is used to acquire the physical position of the user in the real space and the physical position in the virtual space in real time;

[0179] The spatial feasibility entropy module is used to divide the feasible walking area of the user in the real space based on the physical position of the user in the real space, and calculate the spatial feasibility entropy in the real space based on the feasible walking area;

[0180] The state calculation module is used to calculate the real-virtual environment state based on the physical position of the user in the real space and the physical position in the virtual space;

[0181] The reward function calculation module is used to calculate the reward function of deep reinforcement learning based on the real-virtual environment state and the spatial feasibility entropy, and the reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function;

[0182] The reset calculation module is used to determine whether the walking state of the user needs to be reset. If it does not need to be reset, the position reward function and the spatial alignment reward function are used; if it needs to be reset, the reset penalty function is used to calculate the optimal reset direction, and the user's walking path is guided to the optimal safe area within the feasible walking area.

[0183] Figure 3 FIG. 24 is a schematic structural diagram of a terminal 300 provided in an embodiment of the present invention, and the terminal 300 can be used to execute the redirected walking method based on deep reinforcement learning provided in the embodiment of the present invention.

[0184] Among them, the terminal 300 may include: a processor 310, a memory 320, and a communication unit 330. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation to the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine some components, or different component arrangements.

[0185] Among them, the memory 320 can be used to store the execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. When the execution instructions in the memory 320 are executed by the processor 310, the terminal 300 can execute some or all of the steps in the above method embodiments.

[0186] The processor 310 is the control center of the storage terminal, connecting various parts of the entire electronic terminal through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 320, and by calling the data stored in the memory, it executes various functions of the electronic terminal and / or processes data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 310 can include only a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single arithmetic core or can include multiple arithmetic cores.

[0187] The communication unit 330 is used to establish a communication channel so that the storage terminal can communicate with other terminals. It receives user data sent by other terminals or sends user data to other terminals.

[0188] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the various embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.

[0189] Therefore, the present invention calculates the spatial feasible entropy in the real space, thereby effectively reducing the collision probability between the user and the physical boundary. And while achieving virtual-physical alignment, it dynamically adjusts the user's walking path so that it always faces a more suitable walking area, and can adjust the walking speed, direction, and path without the user's awareness, making the walking smoother and more natural, reducing VR dizziness, and improving the immersive experience. Moreover, when the user approaches the physical boundary, the optimal reset direction guides the user towards a more open direction, thereby reducing secondary collisions and enhancing the fluency and freedom of VR roaming, so as to adapt to complex physical space layouts. The technical effects achievable by this embodiment can be referred to the descriptions in the above text and will not be elaborated here.

[0190] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., which can store program codes, and includes several instructions for causing a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.

[0191] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the terminal embodiments, since they are basically similar to the method embodiments, the descriptions are relatively simple, and the relevant parts can refer to the descriptions in the method embodiments.

[0192] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the systems or modules can be in electrical, mechanical, or other forms.

[0193] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or may be distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0194] In addition, in each embodiment of the present invention, each functional module may be integrated in a processing module, may exist separately as individual physical modules, or two or more modules may be integrated in one module.

[0195] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions should all be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A redirected walking method based on deep reinforcement learning, characterized in that including: Obtain the physical location of the user in the real space and the physical location in the virtual space in real time; Divide the walkable area of the user in the real space based on the physical location of the user in the real space, and calculate the spatial feasibility entropy in the real space based on the walkable area; Calculate the real-virtual environment state based on the physical location of the user in the real space and the physical location in the virtual space; Calculate the reward function of deep reinforcement learning based on the real-virtual environment state and the spatial feasibility entropy, and the reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function; Determine whether the walking state of the user needs to be reset. If it does not need to be reset, use the position reward function and the spatial alignment reward function; If it needs to be reset, use the reset penalty function to calculate the optimal reset direction and guide the user's walking path to the optimal safe area within the walkable area.

2. The method according to claim 1, characterized in that, Dividing the walkable area of the user in the real space based on the physical location of the user in the real space includes: Define the real space S P , define the user's physical position x in the real space P , where, x P ∈S P ; Construct a visibility polygon VP(x P ), the polygon including all points visible from x P ; Establish polar coordinates and model the visibility polygon VP(x P ) in polar coordinates to obtain the walkable area.

3. The method according to claim 1, characterized in that, Calculating the spatial feasibility entropy in the real space based on the walkable area includes: Discretize the walkable area and calculate the free walk entropy \(e\) of the discretized walkable area free 、space blocking entropy \(e\) obs and distance penalty function \(f\) dis ; Based on the free walking entropy e free , the space blocking entropy e obs and the distance penalty function f dis calculate the spatial feasible entropy E of the user's physical location sw .

4. The method according to claim 3, wherein Discretize the walkable area and calculate the free walk entropy e of the discretized walkable area free , the space block entropy e obs and the distance penalty function f dis , and calculate the space feasible entropy E of the user's physical location based on the free walk entropy e free , the space block entropy e obs and the distance penalty function f dis , including: sw ​ With step size, sample the visibility polygon VP(x P ), and discretely divide the visibility polygon VP(x P ) into n consecutive visible triangles {VT1(x P ), VT2(x P ), …, VT n (x P )}; Model the freedom of movement for x P Define the probability p that the user transitions to a specific VT i (x P ) as i being: where A(·) represents the polygon area function; Based on obtain x P the free walking entropy e at free : Among them, e free represents the uniformity of the motion probability distribution from one position to the surrounding area; Modeling the spatial blockage for x P to quantify each sampled VT i (x P ) and the change in the blockage degree in different sampling directions within VP(x P ). The spatial blockage entropy e P at x obs is defined as: Among them, e obs represents the relative obstacle degree of the user's movement from x P in all directions, and is used for safe comparison of different physical positions; Use the distance penalty function f dis , to adjust the entropy of x P . The distance penalty function f dis is defined as follows: where d(·) represents the Euclidean distance between x P and the nearest obstacle, and represent the walkable area and the blocked area respectively, λ represents the scaling factor, and λ is set to 0.1; Based on the free walking entropy e free , the space blocking entropy e obs and the distance penalty function f dis calculate the spatial feasible entropy E of the user's physical location sw : E sw (x P ) = e free (x P ) e obs (x P ) f dis (x P ) (5).

5. The method according to claim 1, characterized in that Calculating the real-virtual environment state based on the physical location of the user in the real space and the physical location in the virtual space includes: Definition is the two-dimensional position of the user in the virtual space S V at time t, is the two-dimensional forward direction in the virtual space S V at time t. Define as the two-dimensional position of the user in the real space S P at time t, is the two-dimensional forward direction in the real space S V at time t. Define the motion state as At and uniformly emit k-rays to detect the distance between the user and the surrounding environment, and use the real-virtual projection function to and and align, and compare the virtual rays and the real rays at each emission angle, where the emission angle Calculate the security of the user at a certain θ i : where w s represents a relative distance measure, represents the distance from a two-dimensional position in the real space along the direction θ i to the obstacle, represents the distance from a two-dimensional position in the virtual space along the direction θ i to the obstacle, and when it indicates safe walking; Calculate the real-virtual environment state ε t : Among them, is the length of the diagonal of S P .

6. The method according to claim 1, wherein Calculating the reward function of deep reinforcement learning based on the real-virtual environment state and the spatial feasibility entropy, and the reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function, includes: Calculating the reward function of deep reinforcement learning: Among them, R P represents the position reward function, R a represents the spatial alignment reward function, R c represents the reset penalty function; The position reward function R P is as follows: Among them, is the walking speed of the user in the real space, v min is the minimum movement threshold; The calculation steps of the spatial alignment reward function R a are as follows: Define the distance weight w dis : Based on the distance weight w dis Calculate the alignment factor W align : Based on the alignment factor W align , calculate the spatial alignment reward function R a : Among them, represents the rotation speed of the user; The reset penalty function R c is as follows: where v max is the maximum walking speed.

7. The method according to claim 6, characterized in that Using the reset penalty function to calculate the optimal reset direction and guiding the user's walking path to the optimal safe area within the walkable area includes:

8. A redirected walking method and system based on deep reinforcement learning, characterized in that including: A position acquisition module for obtaining the physical location of the user in the real space and the physical location in the virtual space in real time; A spatial feasibility entropy module for dividing the walkable area of the user in the real space based on the physical location of the user in the real space and calculating the spatial feasibility entropy in the real space based on the walkable area; A state calculation module for calculating the real-virtual environment state based on the physical location of the user in the real space and the physical location in the virtual space; A reward function calculation module for calculating the reward function of deep reinforcement learning based on the real-virtual environment state and the spatial feasibility entropy, and the reward function includes a position reward function, a spatial alignment reward function, and a reset penalty function; A reset calculation module for determining whether the walking state of the user needs to be reset. If it does not need to be reset, use the position reward function and the spatial alignment reward function; If it needs to be reset, use the reset penalty function to calculate the optimal reset direction and guide the user's walking path to the optimal safe area within the walkable area.

9. A terminal, characterized in that, including: A memory for storing the program of the redirection walking method based on deep reinforcement learning; A processor for implementing the steps of the redirection walking method based on deep reinforcement learning as described in any one of claims 1-7 when executing the program of the redirection walking method based on deep reinforcement learning.

10. A computer-readable storage medium storing a computer program, characterized in that, The readable storage medium stores a program for the redirected walking method based on deep reinforcement learning. When the program for the redirected walking method based on deep reinforcement learning is executed by a processor, the steps of the redirected walking method based on deep reinforcement learning as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Predictive redirection control method and system for complex virtual scene

    CN120491655A

  • Method and system for generating virtual road map of predictive redirection controller

    CN120510340A

  • VR redirection walking method and system based on deep reinforcement learning

    CN122223281A