Light tracking positioning method and system for VR large-space immersive touring
By combining deep learning models and ray tracing technology, the ray tracing method of convolutional neural networks and recurrent neural networks is used to solve the problem of traditional positioning methods being computationally large in complex scenarios and being sensitive to scene changes, achieving high-precision and real-time positioning effects, and is suitable for applications such as VR large-space immersive tours.
Patent Information
- Application Number
- CN202510416293.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing positioning methods are computationally large in complex scenarios and are sensitive to scene changes, and cannot meet the requirements of high precision and real-time, especially in application scenarios such as large space VR tours, museum VR tours, scenic spot VR tours, and inspection robot navigation.
Combining deep learning models and ray tracing technology, the network structure combined with convolutional neural network and recurrent neural network is used to reduce the computational complexity and improve the robustness of scene parameter changes through the prediction of the interaction between rays and scenes. The Monte Carlo method is used to generate CSI data and an adaptive dynamic adjustment mechanism is introduced.
It significantly improves positioning accuracy and generalization capabilities of the model, enhances the adaptability and real-time nature of the system in a dynamic environment, and provides high-precision and high-efficiency positioning solutions for application scenarios such as VR large-space immersive tours.
Smart Images

Figure CN120353336A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and computer vision, and specifically includes a ray tracing positioning method and system for VR large-space immersive tours. Background Art
[0002] High-precision positioning technology plays a crucial role in many fields, especially in large-space VR tours, museum VR tours, scenic area VR tours, and inspection robot navigation. These application scenarios require the system to be able to accurately determine the position of the user or device in real time in order to provide corresponding services or perform specific tasks. For example, in large-space VR tours and museum VR tours, accurate positioning not only helps prevent collisions between participants and avoid collisions with obstacles, but also provides personalized interactive experiences and service content based on the specific position and line-of-sight direction of the participants. In scenic area VR tours, accurate positioning can bring a more realistic immersive travel experience to tourists, making them feel as if they are on the spot even when they are in a different place. Inspection robots rely on high-precision positioning to ensure that they can navigate autonomously in complex environments and effectively complete various detection and maintenance tasks.
[0003] However, existing positioning methods often fail to meet the high-precision requirements of the above applications. Although traditional ray tracing methods can simulate the wireless signal propagation process in the real world to achieve relatively accurate positioning, they face challenges such as huge computational complexity and sensitivity to changes in scene parameters when dealing with complex scenarios. This not only limits their real-time performance but also reduces their applicability in dynamic environments.
[0004] To address these problems, this method proposes a ray tracing positioning method with efficient parameter fine-tuning: by combining a deep learning network to approximate the traditional ray tracing process. The core of this method is to use a neural network model to predict the interaction between rays and the scene, thereby reducing the computational complexity and improving the robustness to changes in scene parameters. Specifically, a large amount of scene data needs to be collected first, including CSI (Channel State Information) under different conditions and corresponding scene parameters. Then, these data are used to train a deep learning model so that it can learn the relationship between complex scene parameters and CSI. Compared with traditional physics-based methods, this data-driven approach reduces the dependence on physical models and also enables the model to better adapt to changes in the scene.
[0005] The introduction of deep learning models has brought several significant advantages. First, it greatly reduces the computational complexity, making real-time positioning possible, which is particularly important for application scenarios that require quick responses. Second, since the model is trained with a large amount of actual data, it has a higher tolerance for minor changes in scene parameters, enhancing the stability and reliability of the system. In addition, this method can also capture some non-linear relationships that are difficult to describe by traditional physical models, further improving the accuracy of positioning. Summary of the Invention
[0006] To solve the technical problems in the above background, the present invention provides a ray tracing positioning method for VR large-space immersive tours, and the steps include:
[0007] Collect scene parameters in a VR large space and generate training data;
[0008] Process the training data to obtain processed data;
[0009] Input the processed data into the constructed ray tracing positioning model and output the ray tracing positioning result.
[0010] Preferably, the scene parameters include scene geometric information and electromagnetic material properties; the scene geometric information includes: object position and shape; wherein, the object position in the scene represents the position of the object in the scene in three-dimensional coordinates, and the shape in the scene describes the shape of the object using triangular meshes or other geometric representation methods; the electromagnetic material properties include: the relative permittivity, magnetic permeability, and conductivity of the object.
[0011] Preferably, the method for generating the training data includes: using a physical simulation tool to generate ray tracing data, including the electromagnetic wave propagation paths, reflections, and refraction information in different scenes, where each sample contains the scene parameters and the corresponding ray tracing results;
[0012] At the same time, expand the scene parameters through scaling, rotation, and translation operations; and normalize the scene parameters and ray tracing results to the range of 0 to 1.
[0013] Preferably, the constructed ray tracing positioning model adopts a network structure combining a convolutional neural network and a recurrent neural network, including a preliminary extraction layer, a convolutional neural network layer, a recurrent network layer, and an output head;
[0014] Among them, the preliminary extraction layer includes a convolutional layer for preliminary feature extraction of data;
[0015] The convolutional neural network layer consists of several convolutional blocks, a multi-layer perceptron, and an attention module; among them, each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation function, the convolutional kernel size is 3×3, and the number of channels gradually increases, which is used to extract spatial features at different levels; a residual connection is added between the convolutional layers to accelerate the convergence of the model;
[0016] The recurrent neural network layer is a bidirectional long short-term memory network added after the convolutional neural network layer, which is used to capture the time series features in the ray tracing process; among them, the bidirectional long short-term memory network is used to learn the information of the ray from both positive and negative directions to enhance the context understanding ability of the model;
[0017] The output head is used to integrate the features of the output of the recurrent neural network layer, uses a dropout layer to prevent overfitting, and outputs the ray tracing result through a fully connected layer; the ray tracing result includes the propagation path, reflection times, and propagation time of the ray.
[0018] Preferably, the attention module is used to receive the features of the middle layer of the model as input and output the corresponding weight value ζ i :
[0019]
[0020] Among them, e i represents the attention score, e i = MLP(f i ), where f i represents the feature vector of the middle layer of the model, and MLP represents the multi-layer perceptron; N is the number of samples.
[0021] Preferably, the steps of outputting the optical tracking positioning result include: converting the geometric and electromagnetic calculation results of ray tracing into CSI data that can be used for positioning, estimating the path integral by the Monte Carlo method, and combining multi-antenna and multi-subcarrier technologies to generate a CSI matrix; the Monte Carlo method formula is as follows:
[0022]
[0023] Among them, represents the CSI approximately obtained by the Monte Carlo method under the given scenario model I and frequency f; H(I,f,w i ) represents the electromagnetic field response of the path w i at the frequency; N w represents the number of sampled paths; w i represents the i-th sampled path;
[0024] CSI is a complex-valued matrix, which represents the signal propagation characteristics of different antennas for different subcarriers:
[0025]
[0026] Among them, H represents the final complex-valued CSI matrix; represents the frequency of the j-th subcarrier; f c represents the center frequency; Δf represents the subcarrier spacing; N t and N r respectively represent the number of transmit and receive antennas; N s represents the number of subcarriers; represents the n t th transmit antenna and the n r th receive antenna's CSI.
[0027] Preferably, the total loss function of the optical tracking positioning model includes:
[0028]
[0029] Among them, λ1, λ2, and λ3 represent weight coefficients; represents the mean square error loss function; represents the cross-entropy loss function after combining the attention mechanism; represents the geometric consistency loss function;
[0030] The positioning loss function is used to measure the difference between the simulated CSI and the actual CSI:
[0031]
[0032] Among them, γ1 and γ2 represent weight coefficients; MSE amp represents the CSI amplitude error loss function; MSE delay represents the CSI delay error function, and R(x, y) represents the regularization term.
[0033] Preferably, an adaptive dynamic adjustment mechanism is introduced to dynamically adjust the model parameters according to real-time data, so that the optical tracking positioning model can quickly adapt to environmental changes, including:
[0034]
[0035] θ represents the model parameters.
[0036] The present invention also provides an optical tracking positioning system for VR large-space immersive tours. The system is used to implement the above method and includes: a collection module, a processing module, and an output module;
[0037] The collection module is used to collect the scene parameters under VR large space and generate training data;
[0038] The processing module is used to process the training data to obtain processed data;
[0039] The output module is used to input the processed data into the constructed ray tracing positioning model and output the ray tracing positioning result.
[0040] Preferably, the constructed ray tracing positioning model adopts a network structure combining a convolutional neural network and a recurrent neural network, including a preliminary extraction layer, a convolutional neural network layer, a recurrent network layer, and an output head;
[0041] Among them, the preliminary extraction layer includes a convolutional layer for preliminary feature extraction of data;
[0042] The convolutional neural network layer is composed of several convolutional blocks, a multi-layer perceptron, and an attention module; among them, each convolutional block is composed of a convolutional layer, a batch normalization layer, and a ReLU activation function, the convolutional kernel size is 3×3, and the number of channels gradually increases for extracting spatial features at different levels; a residual connection is added between the convolutional layers to accelerate model convergence;
[0043] The recurrent network layer is a bidirectional long short-term memory network added after the convolutional neural network layer for capturing time series features in the ray tracing process; among them, the bidirectional long short-term memory network is used to learn light information from both positive and negative directions to enhance the model's context understanding ability;
[0044] The output head is used to integrate the features of the output of the recurrent network layer, uses a dropout layer to prevent overfitting, and outputs the ray tracing result through a fully connected layer; the ray tracing result includes the propagation path of the light, the number of reflections, and the propagation time.
[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0046] By combining a deep learning model with ray tracing technology, the present invention effectively solves the problems of large computational amount and sensitivity to scene changes of traditional positioning methods in complex scenes. This method uses a network structure combining a convolutional neural network and a recurrent neural network, which can efficiently extract spatial features and time series features, and focuses on key ray paths with the help of an attention mechanism, significantly improving the positioning accuracy and the generalization ability of the model. At the same time, by generating CSI data through the Monte Carlo method and introducing an adaptive dynamic adjustment mechanism, the adaptability and real-time performance of the system in a dynamic environment are further enhanced, providing a high-precision and high-efficiency positioning solution for application scenarios such as VR large-space immersive tours, and having important practical application value and broad market prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for use in the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0048] Figure 1 It is a schematic flowchart of the method of the embodiment of the present invention. Specific implementation manners
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0051] Embodiment 1
[0052] This embodiment provides a ray tracing positioning method for VR large-space immersive tours, and the steps include:
[0053] S1. Collect scene parameters in a VR large space and generate training data.
[0054] The scene parameters collected in this embodiment include scene geometric information and electromagnetic material characteristics in the form of a three-dimensional matrix. Among them, the scene geometric information includes parameters such as the position and shape of objects in the scene. "Object position" represents the position of an object in the scene using three-dimensional coordinates; "shape" describes the shape of an object using triangular meshes or other geometric representation methods; electromagnetic material characteristics include parameters such as the relative permittivity, permeability, and conductivity of an object, and these parameters will affect the reflection, refraction, etc. characteristics of electromagnetic waves on the surface of the object.
[0055] After that, a physical simulation tool (such as rendering engines like Lumverse 3D, Unreal Engine, Unity 3D, etc.) is used to generate a large amount of ray tracing data, including information such as the propagation path, reflection, and refraction of electromagnetic waves in different scenes. Each sample contains scene parameters (such as object position, electromagnetic properties) and corresponding ray tracing results (such as propagation time, path loss).
[0056] S2. Process the training data to obtain processed data;
[0057] Expand scene parameters through operations such as scaling, rotation, and translation to enhance the generalization ability of the model. For example, randomly change the positions and electromagnetic properties of objects to generate more training samples. Also, normalize the scene parameters and ray tracing results to the range of 0 to 1 for easy model learning.
[0058] Since the output of the subsequent model is the ray tracing result, which mainly includes information such as the propagation path of the ray and the number of reflections, other parameters such as propagation time and ray propagation direction can be calculated.
[0059] 1) Propagation path: Represents the sequence of objects and the path direction that the ray passes through in the scene. In a deep learning model, the propagation path can be encoded as a vector or matrix, where each element represents the probability or certainty of the ray passing through a certain object.
[0060] 2) Number of reflections: Represents the number of times the ray reflects during propagation. This is a discrete value and can be used as an output dimension of the model.
[0061] S3. Input the processed data into the constructed ray tracing positioning model and output the ray tracing positioning result.
[0062] This embodiment adopts a network structure that combines a convolutional neural network (CNN) and a recurrent neural network (RNN), which includes a preliminary extraction layer, a convolutional neural network layer, a recurrent network layer, and an output head.
[0063] (1) The "preliminary extraction layer" is a convolutional layer used for preliminary feature extraction of data.
[0064] (2) The "convolutional neural network layer" consists of multiple convolutional blocks, multi-layer perceptrons (MLPs), and attention modules. Each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation function. The convolutional kernel size is [specific size], and the number of channels gradually increases to extract spatial features at different levels. Residual connections are added between convolutional layers to alleviate the vanishing gradient problem in deep networks and accelerate model convergence.
[0065] Traditional ray tracing methods cannot effectively capture important ray paths in the scene, resulting in low positioning accuracy. For example, the reflection characteristics of a metal object may be stronger than those of a non-metal object. To solve this problem, this method introduces an attention mechanism in the deep learning model, which can dynamically assign different weights to different ray paths, enabling the model to focus on key paths and improve positioning accuracy. By calculating the correlation between the ray path and the target position, the weights of the path are dynamically adjusted, making the model pay more attention to the paths that have a greater impact on the positioning result and reducing the interference of noise paths. Specifically, this embodiment calculates the weights by adding an "attention module", which receives the features of the intermediate layer of the model as input and outputs the corresponding weight values:
[0066]
[0067] Among them, e i represents the attention score, and e i = MLP(f i ), where f i represents the feature vector of the intermediate layer of the model, and MLP represents the multi-layer perceptron, which is used to learn the non-linear relationship between the input features and the attention score; N is the number of samples.
[0068] (3) "Recurrent neural network layer": Add a bidirectional long short-term memory network (Bidirectional LSTM) after the "convolutional neural network layer" to capture the time series features in the ray tracing process. The bidirectional LSTM can learn the information of the ray from both forward and backward directions, enhancing the context understanding ability of the model.
[0069] (4) The "output head" integrates the features of the output of the "recurrent neural network layer", uses a dropout layer to prevent overfitting, and outputs the ray tracing results through a fully connected layer, including information such as the propagation path of the ray, the number of reflections, and the propagation time. In some classification task scenarios (such as judging the propagation path category of the ray or the object interaction type), a softmax layer is also connected after the fully connected layer (output layer) to convert the original output of the neural network into a probability distribution.
[0070] The training process of the above ray tracing positioning model is as follows:
[0071] (1) Data loading: Load the preprocessed data into the model and train it in batches.
[0072] (2) Forward propagation: Input the scene parameters into the CNN layer to extract spatial features; then transfer the features to the bidirectional LSTM layer to learn the time series features; finally, output the ray tracing results through the fully connected layer.
[0073] (3) Calculate the loss: Compare the predicted ray tracing results with the true values and calculate the loss function.
[0074] (4) Backward propagation: Update the model parameters according to the gradient of the loss function and use the Adam optimizer to update the weights.
[0075] Training evaluation: Evaluate the performance of the model on the validation set, record metrics such as the loss value and accuracy, and select the optimal model for saving.
[0076] The output of the ray tracing positioning result includes CSI (Channel State Information) generation, which is a key step in converting the geometric and electromagnetic calculation results of ray tracing into CSI data that can be used for positioning. By estimating the path integral using the Monte Carlo method and combining multi-antenna and multi-subcarrier technologies, the generated CSI matrix can accurately describe the propagation characteristics of wireless signals in complex environments, providing important support for high-precision positioning.
[0077] The CSI generated by ray tracing can accurately reflect the propagation characteristics of wireless signals in complex environments, providing a basis for high-precision positioning. CSI contains rich environmental information, such as path delay, angle information, etc., which can be used for environmental perception and scene modeling. During the ray tracing process, the system records the parameters of each ray propagation path, including the path delay that is, the propagation path parameters of the ray from the transmitter to the receiver; the angle information and that is, the angle of arrival and the angle of departure of the ray; the amplitude and phase of the electromagnetic field, that is, the change in the electromagnetic field calculated based on the interaction between the ray and the object. The above parameter expressions are as follows:
[0078]
[0079] where, Q i represents the intersection sequence on path i, and these intersections are the intersections of the signal with objects (such as walls, furniture, etc.) in the scene during propagation, used to describe the reflection, refraction, etc. paths of the signal; represents the transmission matrix of path i.
[0080] In theory, CSI can be calculated by integrating all possible propagation paths, but due to the lack of a closed solution for path integration, the Monte Carlo method is used for numerical estimation in practice, that is:
[0081]
[0082] where, represents the CSI approximately obtained by the Monte Carlo method under the given scene model I and frequency f; H(I,f,w i ) represents the electromagnetic field response of path w i at the frequency; N w represents the number of sampled paths; w i represents the i-th sampled path.
[0083] CSI is a complex-valued matrix, representing the signal propagation characteristics on different antennas and different subcarriers:
[0084]
[0085] where, H represents the final complex-valued CSI matrix; represents the frequency of the j-th subcarrier; f c represents the center frequency; Δf represents the subcarrier spacing; N t and N r represent the number of transmit and receive antennas respectively; N s represents the number of subcarriers; represents the n t -th transmit antenna and the n r -th receive antenna's CSI.
[0086] The above expression is as follows:
[0087]
[0088] where and represent the radiation pattern functions of the receive and transmit antennas respectively; represents the phase change caused by the path delay; f represents the electromagnetic frequency.
[0089] The total loss function of the optical ray tracing positioning model is as follows:
[0090]
[0091] where λ1, λ2, and λ3 represent weight coefficients; represents the mean squared error loss function; represents the cross-entropy loss function after incorporating the attention mechanism; represents the geometric consistency loss function.
[0092] Mean squared error: Used to measure the difference between the predicted ray tracing results (such as the number of reflections, propagation time) and the true values. For example, if the model predicts the number of reflections to be 2 while the true number of reflections is 3, then the MSE will calculate this difference:
[0093]
[0094] where y i represents the true value of the ray tracing result; represents the predicted value of the ray tracing result; N represents the number of samples.
[0095] Cross-entropy loss: Used to determine whether the predicted ray path is correct, treating it as a classification problem, that is, classifying the ray path to determine which objects the ray passes through in the scene and how it propagates:
[0096]
[0097] where zi The class label indicating the true ray path. For example, 0 can indicate that the ray does not pass through a certain object, and 1 indicates that the ray has passed through the object; The class probability of the ray path predicted by the model, which is a probability value ranging from 0 to 1, representing the confidence level that the model believes the ray belongs to this class.
[0098] In ray tracing, each possible ray path can be regarded as a class. For example, the ray can pass through object A, object B, or not pass through any object. By training a deep learning model, it can predict the probability of which class the ray belongs to. The cross-entropy loss function measures the accuracy of the model prediction by comparing the probability distribution of the true class and the probability distribution of the predicted class. When the class probability predicted by the model is exactly the same as the true class, the cross-entropy loss value is the smallest, otherwise the loss value increases.
[0099] The loss function combined with the attention mechanism:
[0100]
[0101] Among them, ζ i Represents the attention weight of the ray tracing of the i-th ray (described above), Represents the loss term corresponding to the i-th ray path.
[0102] The geometric consistency loss function: Ensure that the predicted ray path is geometrically reasonable and satisfies the physical laws of electromagnetic wave propagation:
[0103]
[0104] Among them, d i Represents the propagation direction of the true ray; Represents the predicted propagation direction.
[0105] The above optimization process of CSI includes:
[0106] (1) Input: Simulated CSI, that is, the channel state information generated by ray tracing; actual CSI, that is, the channel state information measured from the real environment.
[0107] (2) Loss function design: Use the mean square error (MSE) as the loss function to measure the difference between the simulated CSI and the actual CSI. The loss function combines the amplitude and delay information of the CSI and introduces a regularization term to prevent overfitting.
[0108] (3) Gradient descent optimization: Use the RMSprop optimizer for gradient descent optimization to adjust the scene parameters to minimize the loss function. Learning rate: Set to 0.03. Momentum: Set to 0.6. Number of iterations: Set to 100 iterations in the positioning task.
[0109] (4) Output: The optimized target position, i.e., the final positioning result.
[0110] The loss function is designed as follows:
[0111] The loss function provides a direction for the optimization process by measuring the difference between the simulated CSI and the actual CSI. The gradient descent optimization algorithm adjusts the target position according to the gradient information of the loss function to minimize the value of the loss function. By combining the amplitude and delay information of the CSI, the loss function can more comprehensively reflect the propagation characteristics of the wireless signal, thereby improving the positioning accuracy.
[0112] Composite loss function:
[0113]
[0114] where γ1 and γ2 represent weight coefficients; MSE amp represents the CSI amplitude error loss function; MSE delay represents the CSI delay error function, and R(x, y) represents the regularization term.
[0115] CSI amplitude error: The amplitude of the CSI reflects the signal strength, which is one of the important pieces of information in the positioning process. By minimizing the difference between the simulated CSI amplitude and the actual CSI amplitude, the positioning accuracy can be improved.
[0116]
[0117] where N tot represents the total number of elements in the CSI matrix; a i represents the actually measured CSI amplitude; represents the simulated CSI amplitude.
[0118] CSI delay error: The delay (or phase) of the CSI reflects the signal propagation time, which is another important piece of information in the positioning process. By minimizing the difference between the simulated CSI delay and the actual CSI delay, the positioning accuracy can be further improved.
[0119]
[0120] where represents the actually measured CSI delay; represents the simulated CSI delay.
[0121] Regularization term: To prevent the optimization process from falling into local minima and improve the generalization ability of the model, a regularization term is introduced. The regularization term is usually a simple quadratic term used to constrain the change of the target position.
[0122] R(x,y) = γ2(x 2 +y 2 )
[0123] Among them, x and y represent the coordinates of the target position, and γ2 represents the regularization coefficient, which is used to balance the weights of the regularization term and other error terms.
[0124] The indoor environment changes dynamically, such as personnel movement, object position changes, etc., which affect the adaptability of the model. In this embodiment, by introducing an adaptive dynamic adjustment mechanism, the model parameters are dynamically adjusted according to real-time data, so that the optical tracking positioning model can quickly adapt to environmental changes and maintain high positioning accuracy. By monitoring environmental changes, the online learning mode of the model is triggered, and the model is fine-tuned using a small batch of data to update the model parameters to adapt to new environmental conditions.
[0125] 1. Design of online learning mode: Real-time data collection and preprocessing, small batch data buffering, incremental model fine-tuning, model evaluation and switching.
[0126] 2. Real-time data collection and preprocessing: In a wireless indoor positioning system, the dynamic changes in the environment (such as personnel movement, object position changes, etc.) will have a significant impact on the results of ray tracing. In order to enable the model to quickly adapt to these changes, we need to design an online learning mode to collect and process data in real time.
[0127] (1) Data collection: Through sensors deployed in the environment (such as wireless access points, LiDAR, depth cameras, etc.), data such as channel state information (CSI), environmental geometry information, and object positions are collected in real time.
[0128] (2) Data preprocessing: The collected data is preprocessed, including operations such as filtering, normalization, and data augmentation, to improve data quality and the robustness of the model. For example, sliding window filtering is performed on CSI data to reduce the influence of noise; coordinate transformation and normalization are performed on geometric information to meet the requirements of model input.
[0129] 3. Small batch data buffering: In order to implement online learning, the real-time collected data is stored in the buffer in the form of small batches. The size of the small batch data can be adjusted according to computing resources and update frequency, usually between dozens and hundreds of samples.
[0130] (1) Buffer management: Design a circular buffer to store the recently collected small batch data. When the buffer is full, new data will replace the oldest data to ensure that the data in the buffer always reflects the latest environmental state.
[0131] (2) Data Sampling: Randomly sample a small batch of data from the buffer for fine-tuning the model. Uniform sampling or weighted sampling can be used during sampling. Weighted sampling can be adjusted according to the importance of the data and the uncertainty of the model.
[0132] 4. Incremental Model Fine-tuning: Use a small batch of data to perform incremental fine-tuning on the model and update the model parameters.
[0133] Gradient Calculation: Input the small batch of data into the model and calculate the gradient between the predicted value and the true value. The loss function can use the same total loss function as in offline training (a weighted sum of mean squared error, cross-entropy loss, and geometric consistency loss).
[0134]
[0135] θ represents the model parameters.
[0136]
[0137] (2) Parameter Update: Update the model parameters according to the calculated gradient using an optimization algorithm (such as Adam, RMSprop, etc.). To prevent the model from quickly deviating from the knowledge learned previously, a smaller learning rate can be used. Using a smaller learning rate can limit the update amplitude of the parameters, enabling the model to gradually adapt to the new data and avoid quickly deviating from the knowledge learned previously. This embodiment uses the Adam optimization algorithm:
[0138]
[0139] where α represents the learning rate; θ t represents the updated model parameters; ∈ represents a small constant; represents the first moment estimate, represents the second moment estimate; where β1 and β1 represent hyperparameters; represents the loss function gradient of the model parameters.
[0140] Model Evaluation and Switching: During online learning, periodically evaluate the model to ensure that its performance meets the requirements. Use a small portion of the reserved validation data to calculate the performance metrics of the model (such as recalculating the loss function value). When it is detected that the model performance deteriorates, trigger the offline training process and retrain the model using more data. After the offline training is completed, deploy the updated model to the online learning mode and continue with real-time updates. The overall process of the present invention is as Figure 1 shown.
[0141] Embodiment 2
[0142] This embodiment also provides a ray tracing positioning system for VR large-space immersive tours, including: a collection module, a processing module, and an output module. Among them, the collection module is used to collect scene parameters in a VR large space and generate training data; the processing module is used to process the training data to obtain processed data; the output module is used to input the processed data into the constructed ray tracing positioning model and output the ray tracing positioning result.
[0143] Next, in combination with this embodiment, it will be detailed how the present invention solves technical problems in actual work.
[0144] First, use the collection module to collect scene parameters in a VR large space and generate training data.
[0145] The scene parameters collected in this embodiment include scene geometric information in the form of a three-dimensional matrix and electromagnetic material properties. Among them, the scene geometric information includes parameters such as the position and shape of objects in the scene. "Object position" represents the position of an object in the scene using three-dimensional coordinates; "shape" describes the shape of an object using triangular meshes or other geometric representation methods; electromagnetic material properties include parameters such as the relative permittivity, magnetic permeability, and conductivity of an object, and these parameters will affect the reflection, refraction, etc. of electromagnetic waves on the surface of the object.
[0146] After that, use physical simulation tools (such as rendering engines like Lumverse 3D, Unreal Engine, Unity 3D, etc.) to generate a large amount of ray tracing data, including information such as the propagation path, reflection, and refraction of electromagnetic waves in different scenes. Each sample contains scene parameters (such as object position, electromagnetic properties) and corresponding ray tracing results (such as propagation time, path loss).
[0147] Then the processing module processes the training data to obtain processed data;
[0148] Expand the scene parameters through operations such as scaling, rotation, and translation to increase the generalization ability of the model. For example, randomly change the position and electromagnetic properties of objects to generate more training samples. And normalize the scene parameters and ray tracing results to the range of 0 to 1 for easy model learning.
[0149] Since the output of the subsequent model is the ray tracing result, mainly including information such as the propagation path of the light and the number of reflections, other parameters such as propagation time and light propagation direction can be calculated.
[0150] 1) Propagation path: Represents the sequence of objects and the path direction that the light passes through in the scene. In a deep learning model, the propagation path can be encoded as a vector or matrix, and each element represents the probability or certainty of the light passing through a certain object.
[0151] 2) Number of reflections: Represents the number of times light is reflected during propagation. This is a discrete value and can be an output dimension of the model.
[0152] The final output module inputs the processed data into the constructed ray tracing positioning model and outputs the ray tracing positioning result.
[0153] This embodiment adopts a network structure that combines a convolutional neural network (CNN) and a recurrent neural network (RNN), which includes a preliminary extraction layer, a convolutional neural network layer, a recurrent network layer, and an output head.
[0154] (1) The "preliminary extraction layer" is a convolutional layer used for preliminary feature extraction of data.
[0155] (2) The "convolutional neural network layer" consists of multiple convolutional blocks, a multi-layer perceptron (MLP), and an attention module. Each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation function. The convolutional kernel size is [value], and the number of channels gradually increases to extract spatial features at different levels. Residual connections are added between convolutional layers to alleviate the vanishing gradient problem in deep networks and accelerate model convergence.
[0156] Traditional ray tracing methods cannot effectively capture important light paths in the scene, resulting in low positioning accuracy. For example, a metal object may have stronger light reflection characteristics than a non-metal object. To solve this problem, this method introduces an attention mechanism in the deep learning model, which can dynamically assign different weights to different light paths, enabling the model to focus on key paths and improve positioning accuracy. By calculating the correlation between the light path and the target position, the weight of the path is dynamically adjusted, making the model pay more attention to the paths that have a greater impact on the positioning result and reducing the interference of noise paths. Specifically, this embodiment calculates the weight by adding an "attention module", which receives the features of the middle layer of the model as input and outputs the corresponding weight value:
[0157]
[0158] where, e i represents the attention score, e i = MLP(f i ), where f i represents the feature vector of the middle layer of the model, and MLP represents the multi-layer perceptron, which is used to learn the non-linear relationship between the input features and the attention score; N is the number of samples.
[0159] (3) "Recurrent Neural Network Layer": After the "Convolutional Neural Network Layer", a Bidirectional Long Short-Term Memory Network (Bidirectional LSTM) is added to capture the time series features during the ray tracing process. The Bidirectional LSTM can learn the information of light from both forward and backward directions, enhancing the model's context understanding ability.
[0160] (4) The "Output Head" integrates the features of the output of the "Recurrent Neural Network Layer", uses a dropout layer to prevent overfitting, and outputs the ray tracing results through a fully connected layer, including information such as the propagation path of the light, the number of reflections, and the propagation time. In some classification task scenarios (such as determining the category of the light propagation path or the type of object interaction), a softmax layer is also connected after the fully connected layer (output layer) to convert the original output of the neural network into a probability distribution.
[0161] The training process of the above ray tracing positioning model is as follows:
[0162] (1) Data Loading: Load the preprocessed data into the model and train it in batches.
[0163] (2) Forward Propagation: Input the scene parameters into the CNN layer to extract spatial features; then transfer the features to the Bidirectional LSTM layer to learn the time series features; finally, output the ray tracing results through the fully connected layer.
[0164] (3) Calculate Loss: Compare the predicted ray tracing results with the true values and calculate the loss function.
[0165] (4) Backward Propagation: Update the model parameters according to the gradient of the loss function and use the Adam optimizer for weight update.
[0166] Training Evaluation: Evaluate the performance of the model on the validation set, record metrics such as the loss value and accuracy, and select the optimal model for saving.
[0167] The output of the ray tracing positioning result includes CSI (Channel State Information) generation, which is a key step in converting the geometric and electromagnetic calculation results of ray tracing into CSI data that can be used for positioning. By estimating the path integral through the Monte Carlo method and combining multi-antenna and multi-subcarrier technologies, the generated CSI matrix can accurately describe the propagation characteristics of wireless signals in complex environments, providing important support for high-precision positioning.
[0168] The CSI generated by ray tracing can accurately reflect the propagation characteristics of wireless signals in complex environments, providing a basis for high-precision positioning. CSI contains rich environmental information, such as path delay, angle information, etc., and can be used for environmental perception and scene modeling. During the ray tracing process, the system will record the parameters of each ray propagation path, including the path delay i.e., the propagation path parameters of light from the transmitter to the receiver; angle information and i.e., the angle of arrival and the angle of departure of light; the amplitude and phase of the electromagnetic field, i.e., the change in the electromagnetic field calculated based on the interaction between light and objects. The above parameter expressions are as follows:
[0169]
[0170] where Q i represents the sequence of intersection points on path i, which are the intersection points of the signal with objects (such as walls, furniture, etc.) in the scene during propagation, and are used to describe the reflection, refraction, etc. paths of the signal; represents the transmission matrix of path i.
[0171] In theory, CSI can be calculated by integrating over all possible propagation paths. However, since there is no closed - form solution for path integration, the Monte Carlo method is used for numerical estimation in practice, i.e.:
[0172]
[0173] where represents the CSI approximately obtained by the Monte Carlo method under the given scene model I and frequency f; H(I,f,w i ) represents the electromagnetic field response of path w i at the frequency; N w represents the number of sampled paths; w i represents the i - th sampled path.
[0174] CSI is a complex - valued matrix, representing the signal propagation characteristics on different sub - carriers for different antenna pairs:
[0175]
[0176] where H represents the final complex - valued CSI matrix; represents the frequency of the j - th sub - carrier; f c represents the center frequency; Δf represents the sub - carrier spacing; N t and N r represent the number of transmit and receive antennas respectively; N s represents the number of sub - carriers; represents the CSI between the n t - th transmit antenna and the n r - th receive antenna.
[0177] The above expression is as follows:
[0178]
[0179] Among them, and respectively represent the radiation pattern functions of the receiving and transmitting antennas; represents the phase change caused by the path delay; f represents the electromagnetic frequency.
[0180] The total loss function of the light tracing positioning model is as follows:
[0181]
[0182] Among them, λ1, λ2, and λ3 represent the weight coefficients; represents the mean square error loss function; represents the cross-entropy loss function after combining the attention mechanism; represents the geometric consistency loss function.
[0183] Mean square error: It is used to measure the difference between the predicted ray tracing results (such as the number of reflections, propagation time) and the true values. For example, if the number of reflections predicted by the model is 2, while the true number of reflections is 3, then the MSE will calculate this difference:
[0184]
[0185] Among them, y i represents the true value of the ray tracing result; represents the predicted value of the ray tracing result; N represents the number of samples.
[0186] Cross-entropy loss: It is used to judge whether the predicted ray path is correct, and it is regarded as a classification problem, that is, the classification of the ray path, to judge which objects the ray passes through in the scene and how it propagates:
[0187]
[0188] Among them, Z i represents the class label of the true ray path. For example, 0 can represent that the ray does not pass through a certain object, and 1 represents that the ray passes through the object; represents the class probability of the ray path predicted by the model, which is a probability value ranging from 0 to 1, indicating the confidence level that the model believes the ray belongs to this class.
[0189] In ray tracing, each possible ray path can be regarded as a category. For example, a ray can pass through object A, object B, or not pass through any object. By training a deep learning model, it can predict the probability that a ray belongs to a certain category. The cross-entropy loss function measures the accuracy of the model's prediction by comparing the probability distribution of the true category and the probability distribution of the predicted category. When the category probability predicted by the model is exactly the same as the true category, the cross-entropy loss value is the smallest; otherwise, the loss value increases.
[0190] The loss function combined with the attention mechanism:
[0191]
[0192] where ζ i represents the attention weight of the ray tracing of the i-th ray (as described above), represents the loss term corresponding to the i-th ray path.
[0193] Geometric consistency loss function: Ensure that the predicted ray path is geometrically reasonable and satisfies the physical laws of electromagnetic wave propagation:
[0194]
[0195] where d i represents the propagation direction of the true ray; represents the predicted propagation direction.
[0196] The above optimization process of CSI includes:
[0197] (1) Input: Simulated CSI, that is, the channel state information generated by ray tracing; actual CSI, that is, the channel state information measured from the real environment.
[0198] (2) Loss function design: Use the mean square error (MSE) as the loss function to measure the difference between the simulated CSI and the actual CSI. The loss function combines the amplitude and delay information of the CSI, and at the same time introduces a regularization term to prevent overfitting.
[0199] (3) Gradient descent optimization: Use the RMSprop optimizer for gradient descent optimization to adjust the scene parameters to minimize the loss function. Learning rate: Set to 0.03. Momentum: Set to 0.6. Number of iterations: Set to 100 iterations in the positioning task.
[0200] (4) Output: The optimized target position, that is, the final positioning result.
[0201] The loss function is designed as follows:
[0202] The loss function provides a direction for the optimization process by measuring the difference between the simulated CSI and the actual CSI. The gradient descent optimization algorithm adjusts the target position according to the gradient information of the loss function to minimize the value of the loss function. By combining the amplitude and delay information of CSI, the loss function can more comprehensively reflect the propagation characteristics of wireless signals, thereby improving the positioning accuracy.
[0203] Composite loss function:
[0204]
[0205] where γ1 and γ2 represent weight coefficients; MSE amp represents the CSI amplitude error loss function; MSE delay represents the CSI delay error function, and R(x, y) represents the regularization term.
[0206] CSI amplitude error: The amplitude of CSI reflects the signal strength, which is one of the important pieces of information in the positioning process. By minimizing the difference between the simulated CSI amplitude and the actual CSI amplitude, the positioning accuracy can be improved.
[0207]
[0208] where N tot represents the total number of elements in the CSI matrix; a i represents the actually measured CSI amplitude; represents the simulated CSI amplitude.
[0209] CSI delay error: The delay (or phase) of CSI reflects the signal propagation time, which is another important piece of information in the positioning process. By minimizing the difference between the simulated CSI delay and the actual CSI delay, the positioning accuracy can be further improved.
[0210]
[0211] where, represents the actually measured CSI delay; represents the simulated CSI delay.
[0212] Regularization term: To prevent the optimization process from falling into local minima and improve the generalization ability of the model, a regularization term is introduced. The regularization term is usually a simple quadratic term used to constrain the change of the target position.
[0213] R(x, y) = γ2(x 2 + y 2 )
[0214] Among them, x and y represent the coordinates of the target position, and γ2 represents the regularization coefficient, which is used to balance the weights of the regularization term and other error terms.
[0215] The indoor environment changes dynamically, such as the movement of people and the change of object positions, which affects the adaptability of the model. In this embodiment, by introducing an adaptive dynamic adjustment mechanism, the model parameters are dynamically adjusted according to real-time data, so that the ray tracing positioning model can quickly adapt to environmental changes and maintain high positioning accuracy. By monitoring environmental changes, the online learning mode of the model is triggered, and the model is fine-tuned using a small batch of data to update the model parameters to adapt to the new environmental conditions.
[0216] 1. Online learning mode design: Real-time data collection and preprocessing, small batch data buffering, incremental model fine-tuning, model evaluation and switching.
[0217] 2. Real-time data collection and preprocessing: In a wireless indoor positioning system, the dynamic changes in the environment (such as the movement of people and the change of object positions) will have a significant impact on the results of ray tracing. In order to enable the model to quickly adapt to these changes, we need to design an online learning mode to collect and process data in real time.
[0218] (1) Data collection: Through sensors deployed in the environment (such as wireless access points, LiDAR, depth cameras, etc.), data such as channel state information (CSI), environmental geometry information, and object positions are collected in real time.
[0219] (2) Data preprocessing: The collected data is preprocessed, including operations such as filtering, normalization, and data augmentation, to improve the data quality and the robustness of the model. For example, sliding window filtering is performed on the CSI data to reduce the influence of noise; coordinate transformation and normalization are performed on the geometric information to meet the requirements of the model input.
[0220] 3. Small batch data buffering: In order to implement online learning, the data collected in real time is stored in the buffer in the form of small batches. The size of the small batch data can be adjusted according to the computing resources and update frequency, usually between dozens and hundreds of samples.
[0221] (1) Buffer management: Design a circular buffer to store the recently collected small batch data. When the buffer is full, the new data will replace the oldest data to ensure that the data in the buffer always reflects the latest environmental state.
[0222] (2) Data sampling: Randomly sample small batch data from the buffer for model fine-tuning. Uniform sampling or weighted sampling can be used during sampling, and weighted sampling can be adjusted according to the importance of the data and the uncertainty of the model.
[0223] 4. Incremental Model Fine-tuning: Use a small batch of data to perform incremental fine-tuning on the model and update the model parameters.
[0224] Gradient Calculation: Input the small batch of data into the model and calculate the gradient between the predicted value and the true value. The loss function can use the same total loss function as in offline training (the weighted sum of mean squared error, cross-entropy loss, and geometric consistency loss).
[0225]
[0226] θ represents the model parameters.
[0227]
[0228] (2) Parameter Update: According to the calculated gradient, use an optimization algorithm (such as Adam, RMSprop, etc.) to update the model parameters. To prevent the model from quickly deviating from the knowledge learned previously, a smaller learning rate can be used. Using a smaller learning rate can limit the update amplitude of the parameters, enabling the model to gradually adapt to the new data and avoid quickly deviating from the knowledge learned previously. This embodiment adopts the Adam optimization algorithm:
[0229]
[0230] where α represents the learning rate; θ t represents the updated model parameters; ∈ represents a small constant; represents the first moment estimate, represents the second moment estimate; where β1 and β1 represent hyperparameters; represents the loss function with respect to the gradient of the model parameters.
[0231] Model Evaluation and Switching: During the online learning process, periodically evaluate the model to ensure that its performance meets the requirements. Use a small portion of the reserved validation data to calculate the performance metrics of the model (such as recalculating the loss function value). When it is detected that the model performance degrades, trigger the offline training process and retrain the model with more data. After the offline training is completed, deploy the updated model into the online learning mode and continue with real-time updates.
[0232] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A ray tracing positioning method for VR large-space immersive tour, characterized in that the steps Including: Collecting scene parameters in a large VR space and generating training data; Processing the training data to obtain processed data; Inputting the processed data into the constructed ray tracing positioning model and outputting the ray tracing positioning result.
2. The light tracing positioning method for VR large-space immersive tour according to claim 1, wherein The scene parameters include scene geometric information and electromagnetic material properties; the scene geometric information includes: object position and shape; wherein, the object position in the scene represents the position of the object in the scene using three-dimensional coordinates, and the shape in the scene describes the shape of the object using triangular meshes or other geometric representation methods; the electromagnetic material properties include: relative permittivity, permeability, and conductivity of the object.
3. The ray tracing positioning method for VR large-space immersive tour according to claim 1, wherein, The method for generating the training data includes: using a physical simulation tool to generate ray tracing data, including electromagnetic wave propagation paths, reflection, and refraction information in different scenes, where each sample includes the scene parameters and the corresponding ray tracing results; Meanwhile, expanding the scene parameters through scaling, rotation, and translation operations; and normalizing the scene parameters and ray tracing results to the range of 0 to 1.
4. The light tracking and positioning method for VR large-space immersive tour according to claim 1, wherein, The constructed ray tracing positioning model adopts a network structure combining a convolutional neural network and a recurrent neural network, including a preliminary extraction layer, a convolutional neural network layer, a recurrent network layer, and an output head; Among them, the preliminary extraction layer includes a convolutional layer for preliminary feature extraction of data; The convolutional neural network layer consists of several convolutional blocks, a multi-layer perceptron, and an attention module; wherein, each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation function, with a convolutional kernel size of 3×3 and the number of channels gradually increasing, used to extract spatial features at different levels; adding residual connections between convolutional layers to accelerate model convergence; The recurrent network layer adds a bidirectional long short-term memory network after the convolutional neural network layer to capture the time series features in the ray tracing process; wherein, the bidirectional long short-term memory network is used to learn the information of light from both positive and negative directions to enhance the model's context understanding ability; The output head is used to integrate the features of the output of the recurrent network layer, uses a dropout layer to prevent overfitting, and outputs the ray tracing result through a fully connected layer; the ray tracing result includes the propagation path, reflection times, and propagation time of the light.
5. The light tracing positioning method for VR large-space immersive tour according to claim 4, characterized in that, The attention module is used to receive the features of the intermediate layer of the model as input and output the corresponding weight value ζ i : Among them, e i represents the attention score, and e i = MLP(fi), where f i represents the feature vector of the intermediate layer of the model, MLP represents the multi-layer perceptron; N is the number of samples.
6. The light tracking and positioning method for VR large-space immersive tour according to claim 1, characterized in that The steps for outputting the ray tracing positioning result include: converting the geometric and electromagnetic calculation results of ray tracing into CSI data available for positioning, estimating the path integral through the Monte Carlo method, and combining multi-antenna and multi-subcarrier technologies to generate a CSI matrix; the Monte Carlo method formula is as follows: Among them, represents the CSI approximately obtained by the Monte Carlo method under the given scenario model I and frequency f; H(I, f, w i ) represents the electromagnetic field response of path w i at the frequency; N w represents the number of sampled paths; w i represents the i-th sampled path; CSI is a complex-valued matrix representing the signal propagation characteristics of different antenna pairs on different subcarriers: Among them, H represents the final CSI complex value matrix; represents the frequency of the j-th subcarrier; f c represents the center frequency; Δf represents the subcarrier spacing; N t and N r respectively represent the numbers of transmit and receive antennas; N s represents the number of subcarriers; represents the n t -th transmit antenna and the CSI between the n r -th receive antenna.
7. The light tracking positioning method for VR large-space immersive tour according to claim 6, characterized in that, The total loss function of the ray tracing positioning model includes: Among them, λ1, λ2, and λ3 represent weight coefficients; represents the mean squared error loss function; represents the cross-entropy loss function after incorporating the attention mechanism; represents the geometric consistency loss function; Using a positioning loss function to measure the difference between the simulated CSI and the actual CSI: Among them, γ1 and γ2 represent weight coefficients; MSE amp represents the CSI amplitude error loss function; MSE delay represents the CSI delay error function, and R(x, y) represents the regularization term.
8. The light tracking and positioning method for VR large-space immersive tour according to claim 7, characterized in that, Introducing an adaptive dynamic adjustment mechanism to dynamically adjust the model parameters according to real-time data, enabling the ray tracing positioning model to quickly adapt to environmental changes, including: θ represents the model parameters.
9. A ray tracing positioning system for VR large-space immersive tours, the system being used to implement the method according to any one of claims 1-8, characterized in that, Including: A collection module, a processing module, and an output module; The collection module is used to collect scene parameters in a large VR space and generate training data; The processing module is used to process the training data to obtain processed data; The output module is used to input the processed data into the constructed ray tracing localization model and output a ray tracing localization result.
10. The ray tracing positioning system for VR large-space immersive tour according to claim 9, wherein, The constructed ray tracing localization model adopts a network structure combining a convolutional neural network and a recurrent neural network, and includes a preliminary extraction layer, a convolutional neural network layer, a recurrent network layer, and an output head; Among them, the preliminary extraction layer includes a convolutional layer for preliminary feature extraction of data; The convolutional neural network layer is composed of a number of convolutional blocks, a multi-layer perceptron, and an attention module; among them, each convolutional block is composed of a convolutional layer, a batch normalization layer, and a ReLU activation function, the convolutional kernel size is 3×3, and the number of channels gradually increases, which is used to extract spatial features at different levels; a residual connection is added between the convolutional layers to accelerate model convergence; The recurrent network layer is a bidirectional long short-term memory network added after the convolutional neural network layer, which is used to capture the time series features in the ray tracing process; among them, the bidirectional long short-term memory network is used to learn the information of light from both positive and negative directions to enhance the model's context understanding ability; The output head is used to integrate the features of the output of the recurrent network layer, uses a dropout layer to prevent overfitting, and outputs the ray tracing result through a fully connected layer; the ray tracing result includes the propagation path of the light, the number of reflections, and the propagation time.