Large-space VR multi-target high-dynamic real-time tracking method and system

Through the graph neural network method based on spatiotemporal structure, combined with data preprocessing and knowledge distillation technology, the positioning accuracy and response delay problems of multi-objective real-time tracking in large space VR are solved, and high-precision multi-objective tracking and personalized services are achieved, improving the security and interactivity of the VR experience.

CN120276591APending Publication Date: 2025-07-08北京渲光科技有限公司
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510331708.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Traditional positioning methods are difficult to meet the needs of multi-objective real-time tracking in large-space VR, especially when dealing with frequent changes in locations of a large number of participants, there are problems of insufficient positioning accuracy, response delay and data sparsity, which affect service quality.

Method used

The graph neural network (GNN) method based on spatiotemporal structure is adopted to build a real-time tracking model through data preprocessing, multi-task shared backbone feature extraction layer and knowledge distillation technology, including teacher model and student model, and positioning and behavior recognition are used to combine interpolation and time smoothing methods to solve data sparseness, real-time processing of the model is achieved.

Benefits of technology

It improves the accuracy of positioning and behavior recognition, ensures security and personalized services in the VR environment, improves users' immersive experience, is suitable for multi-person interaction in large space VR scenarios, and promotes the application of cultural tourism, education and training fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276591A_ABST
    Figure CN120276591A_ABST
Patent Text Reader

Abstract

The invention discloses a large-space VR (virtual reality) multi-target high-dynamic real-time tracking method and system. The method comprises the following steps: collecting behavior data of a plurality of targets in a VR large space; constructing a real-time tracking model based on the behavior data; and completing real-time tracking of a plurality of targets in the VR large space by using the real-time tracking model. According to the method, the problems of inaccurate positioning and slow response in a traditional method are solved, and the immersive experience of the user is improved through personalized service. According to the invention, the multi-user interaction VR experience even in a large-scale public place is safe and rich in educational significance, and the wide application and development of the VR technology in multiple fields such as cultural tourism, educational training and the like are promoted. Through the mode, not only can the safety of participants be better protected, but also unprecedented interactive learning experience can be brought to the participants, and cultural inheritance and exchange are further promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multi-object tracking in VR space, and specifically to a large-space VR multi-object high-dynamic real-time tracking method and system. Background Art

[0002] In large-space virtual reality (VR) applications, users experience an immersive tour by wearing professional VR devices. This technology brings unprecedented interaction and learning opportunities for users. However, due to the enclosed nature of the VR environment, participants cannot directly perceive the real world outside, which poses potential risks to their safety. For example, in a virtual museum or art exhibition, multiple visitors are moving around simultaneously, and they may accidentally bump into each other, obstacles, or walls. Therefore, ensuring the safety of each participant has become a crucial task.

[0003] To enhance the user experience, in addition to basic safety guarantees, highly customized services need to be provided according to the positions and behaviors of the participants. For example, by analyzing the facial orientation of the participants to estimate their gaze direction, and then determining what the participants are paying attention to, and providing detailed explanations or presenting more abundant visual materials accordingly. Such personalized services can not only enhance the user's immersion, but also deepen their understanding and memory of the exhibits. For example, when visiting a virtual museum, if the system can identify that a participant is focusing on a specific cultural relic, it can immediately provide comprehensive information such as the historical background and cultural value of the cultural relic, and even allow the visitors to observe the cultural relic from different angles through 3D models, greatly enriching the visiting experience.

[0004] However, implementing the above functions faces huge technical challenges. Traditional positioning methods are difficult to meet the requirements of multi-object real-time tracking in a high-dynamic environment. Especially when dealing with the frequent changes in the positions of a large number of participants, problems such as insufficient positioning accuracy and response delay often occur. In addition, due to the existence of the data sparsity problem, existing methods are difficult to accurately capture the behavioral details of each participant, thus affecting the service quality. Summary of the Invention

[0005] In response to the above background problems, the present invention proposes a new method based on a graph neural network (GNN) with a spatio-temporal structure. This method first solves the problem of data sparsity through data preprocessing, making subsequent feature extraction more accurate and effective. Then, a neural network architecture with a multi-task shared backbone feature extraction layer is adopted to improve the accuracy of positioning and behavior recognition. This architecture designs multiple task-specific output layers to reduce the overall model complexity while improving its flexibility and adaptability. Finally, knowledge distillation technology is used to "compress" the complex model into a lightweight version, which can not only ensure the performance of the model, but also significantly improve its real-time processing ability.

[0006] To achieve the above object, the present invention provides a large-space VR multi-target high-dynamic real-time tracking method, and the steps include:

[0007] Collect the behavior data of multiple targets in the large VR space;

[0008] Based on the behavior data, construct a real-time tracking model; the real-time tracking model includes: a teacher model and a student model, and both the teacher model and the student model are composed of an input stage, a spatio-temporal graph neural network, and an output stage; wherein, the teacher model contains several pairs of spatio-temporal graph neural networks, and the student model contains 1 pair of spatio-temporal graph neural networks;

[0009] Use the real-time tracking model to complete the real-time tracking of multiple targets in the large VR space.

[0010] Preferably, the real-time tracking model solves the problem of data sparsity through data preprocessing in the input stage;

[0011] Then, adopt a spatio-temporal graph neural network with a multi-task shared backbone feature extraction layer to improve the accuracy of positioning and behavior recognition;

[0012] Finally, use the knowledge distillation technology to compress its own complex model.

[0013] Preferably, the difference between the teacher model and the student model lies in the stacking layer number of the spatio-temporal graph neural network; the teacher model learns a higher-level feature representation through a multi-layer stacking method; the student model realizes model lightweighting through one layer of the spatio-temporal graph neural network to meet dynamic adaptability.

[0014] Preferably, the real-time tracking model adopts a combined method of interpolation and time smoothing to solve the problem of data sparsity; wherein, the interpolation adopts the spline interpolation method:

[0015] S(t0 = a i + b i (t - t i ) + c i (t - t i ) 2 + d i (t - t i ) 3

[0016] wherein, S(t) represents the signal value obtained by spline interpolation at the current time point t; a i , b i , c i , d i all represent the coefficients of the piecewise polynomial; t represents the current time point; ti represents the time position corresponding to the i-th data point;

[0017] The time smoothing adopts the Gaussian smoothing method:

[0018]

[0019] where y(t) represents the signal value after Gaussian smoothing at time point t; i0 represents the offset of time point t; k represents the size of the smoothing window; σ represents the standard deviation of the Gaussian kernel.

[0020] Preferably, the spatio-temporal graph neural network improves the accuracy of positioning and behavior recognition through spatio-temporal feature aggregation:

[0021]

[0022] where, represents the aggregated feature; N s (i1) and N t (i1) represent the spatial neighborhood and temporal neighborhood of node i1 respectively; x′ j,s represents the spatial node feature; x′ j,s represents the temporal node feature; Aggregate represents the aggregation operation.

[0023] Preferably, the output stage adopts a combination of three "fully connected" and "Softmax" to output three types of labels, namely the position label probability distribution, the state label probability distribution, and the behavior label probability distribution.

[0024] Preferably, the real-time tracking model adopts a joint loss function to balance the weights of different tasks:

[0025]

[0026] where β represents the balance coefficient; represents the original loss function of the student model; represents the distillation loss function.

[0027] The present invention also provides a large-space VR multi-object high-dynamic real-time tracking system, which is used to implement the above method, including: a collection module, a construction module, and a tracking module;

[0028] The collection module is used to collect the behavior data of multiple targets in the large space of VR;

[0029] The building block is used to construct a real-time tracking model based on the behavior data; the real-time tracking model includes: a teacher model and a student model, and both the teacher model and the student model are composed of an input stage, a spatio-temporal graph neural network, and an output stage; wherein, the teacher model contains several pairs of spatio-temporal graph neural networks, and the student model contains one pair of spatio-temporal graph neural networks;

[0030] The tracking module is used to utilize the real-time tracking model to complete the real-time tracking of multiple targets in a large VR space.

[0031] Preferably, the real-time tracking model solves the problem of data sparsity through data preprocessing in the input stage;

[0032] Then, a spatio-temporal graph neural network with a multi-task shared backbone feature extraction layer is adopted to improve the accuracy of positioning and behavior recognition;

[0033] Finally, knowledge distillation technology is used to compress its own complex model.

[0034] Preferably, the difference between the teacher model and the student model lies in the stacking layers of the spatio-temporal graph neural network; the teacher model learns higher-level feature representations through a multi-layer stacking method; the student model realizes model lightweighting through one layer of the spatio-temporal graph neural network to meet dynamic adaptability.

[0035] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0036] The present invention not only solves the problems of inaccurate positioning and slow response existing in the traditional method, but also improves the user's immersive experience through personalized services. It makes the multi-person interactive VR experience in large public places safe and educational, and promotes the wide application and development of VR technology in multiple fields such as cultural tourism and education and training. In this way, not only can the safety of participants be better protected, but also an unprecedented interactive learning experience can be brought to them, further promoting the inheritance and exchange of culture. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.

[0038] Figure 1 It is a schematic diagram of the model structure of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Embodiment 1

[0042] This embodiment provides a large-space VR multi-target high-dynamic real-time tracking method, which is characterized in that the steps include:

[0043] A large-space VR multi-target high-dynamic real-time tracking method, which is characterized in that the steps include:

[0044] S1. Collect the behavior data of multiple targets in the large VR space.

[0045] The input data of this embodiment is the Wi-Fi received signal strength indication. In addition, other solutions such as GPS, Bluetooth, RFID, and ultra-wideband can also be used. However, because Wi-Fi access points (APs) are widely available, no additional infrastructure and hardware need to be installed; by using the existing Wi-Fi network, no additional hardware devices are required, reducing the deployment cost; the Wi-Fi signal coverage is wide and can usually cover the entire building or multiple floors, suitable for large-scale indoor positioning; the Wi-Fi signal has strong penetration ability and can penetrate walls and obstacles to a certain extent, suitable for complex indoor environments; the change of Wi-Fi signal strength can reflect the relative position change between the device and the access point, suitable for dynamic environments; most modern devices (such as VR, smartphones, tablets, laptops) support Wi-Fi, with wide compatibility; by using the existing Wi-Fi network, data collection is simple, and only the RSSI value between the device and the access point needs to be recorded; the positioning accuracy through the RSSI value is usually within the range of 1-5 meters, suitable for most indoor positioning requirements; in summary, this embodiment uses Wi-Fi RSSI as the input data of the model.

[0046] S2. Build a real-time tracking model based on the behavior data.

[0047] In this embodiment, the real-time tracking model includes: a teacher model and a student model, both of which are composed of an input stage, a spatio-temporal graph neural network, and an output stage; among them, the teacher model contains N pairs (where N ≥ 3, and N = 5 is selected in this embodiment) of spatio-temporal graph neural networks, and the student model contains 1 pair of spatio-temporal graph neural networks, and its structure is as Figure 1 shown.

[0048] First, in the model input stage, data preprocessing is used to solve the data sparsity problem in the received data. The data sparsity problem usually manifests as: missing in the time dimension, that is, some time steps may have no data records; missing in the space dimension, that is, some nodes may have no data at some time steps; noisy data, that is, the recorded data may contain noise, resulting in signal strength fluctuations. These problems above will affect the performance of the model. Especially in a dynamic environment, data sparsity and noise may cause the model to be unable to accurately capture spatio-temporal relationships.

[0049] The data preprocessing steps include: serialization, selecting a time window, interpolation, time smoothing, and normalization.

[0050] The specific steps are as follows;

[0051] Serialization: It is necessary to convert the RSSI data into a time series format, that is, for each location, record its RSSI measurement values at different time points.

[0052] Selecting a time window: It is necessary to determine an appropriate time window size to capture the dynamic changes of the signal in the time dimension. The selection of the time window will affect the performance and computational complexity of the model.

[0053] Then, a combined method of interpolation and time smoothing is adopted, that is, first interpolate and then smooth. Specifically, first use data interpolation technology to fill in the missing values, and then use time smoothing technology to reduce data fluctuations and noise. Data interpolation is a method of estimating unknown data points through known data points, which can effectively fill in the missing parts of the data. Time smoothing technology reduces data fluctuations and noise and enhances data continuity by applying a smoothing operation on the time series. Among them, the interpolation algorithms can adopt linear interpolation, polynomial interpolation, and spline interpolation, etc., and the time smoothing algorithms can adopt moving average, exponentially weighted moving average, and Gaussian smoothing, etc.

[0054] In this embodiment, a combination of spline interpolation and Gaussian smoothing is adopted, which meets high-precision interpolation and is sensitive to noise. Specifically, spline interpolation: Use a piecewise polynomial function to fit the known data points, and estimate the missing values through the piecewise polynomial function, which is applicable to scenarios where the data changes relatively smoothly and high-precision interpolation is required:

[0055] S(t) = a i +b i (t - ti ) + c i (t - t i ) 2 + d i (t - t i ) 3

[0056] S(t) represents the signal value obtained by spline interpolation at the current time point t; a i , b i , c i , d i all represent the coefficients of the piecewise polynomial; t represents the current time point; t i represents the time position corresponding to the i-th data point.

[0057] Gaussian smoothing: At each time step, the Gaussian kernel function is used to perform weighted averaging on the surrounding data to smooth the data, which is suitable for scenarios where the data needs to be smoothed and is sensitive to noise:

[0058]

[0059] where y(t) represents the signal value after Gaussian smoothing at the time point t; i0 represents the offset of the time point t; k represents the size of the smoothing window; σ represents the standard deviation of the Gaussian kernel.

[0060] Finally, normalization can make the input data have a unified scale numerically, avoiding excessive numerical differences between different features; normalization can accelerate the convergence speed during the training process of the neural network, making the weight update process of the neural network smoother and more stable; normalization can eliminate the influence of part of the noise and outliers in the input data on the network, making the neural network more focused on learning the core features and patterns in the data.

[0061] Spatio-temporal graph neural network (ST-GNN) is an extended graph neural network architecture that models the dynamic characteristics of data through a spatio-temporal graph structure and is specifically used to process data with spatio-temporal dependencies. It not only considers the spatial relationship between nodes (such as the proximity of RSSI data), but also introduces the time dimension to capture the change of the signal over time. Its characteristics are:

[0062] (1) In the ordinary graph neural network, the graph structure is static, while in the ST-GNN, the graph structure changes over time.

[0063] (2) The features of the nodes include not only spatial information (such as RSSI value), but also time information (such as the change of signal strength over time).

[0064] (3) The ST-GNN will dynamically update the graph structure according to the time series to adapt to the changes in the environment.

[0065] The steps to construct a spatio-temporal graph neural network for constructing a spatio-temporal graph are as follows:

[0066] (1) Construction of the spatial graph: Similar to ordinary GNNs, kNN is used to construct the spatial graph, which represents the spatial relationship between nodes.

[0067] (2) Construction of the temporal graph: Introducing the time dimension, the measurement values of the same node at different time points are connected, that is, the nodes at adjacent time points are connected by edges, and the weight of the edge can represent the temporal correlation.

[0068] (3) In the spatial graph and the temporal graph, each measurement point is regarded as a node in the graph, and the feature of each node is its corresponding feature vector, which represents the Wi-Fi RSSI signal strength measured at that position. The edges between nodes are constructed by the kNN (k-Nearest Neighbors) algorithm, and the weight of the edge is determined by the similarity between different feature vectors, that is, the Euclidean distance between feature vectors.

[0069] (4) Representation of the spatio-temporal graph: The spatio-temporal graph can be represented as G=(V,E s ,E t ), where E s represents the spatial edge, and E t represents the temporal edge.

[0070] After that, convolution operations are performed on the spatio-temporal graph. Through spatial correlation modeling and temporal dynamics analysis, efficient feature extraction in complex environments is achieved; at the same time, combined with lightweight design and multi-task cooperation, the accuracy and robustness of multi-object tracking are significantly improved while ensuring real-time performance, so as to meet the stringent requirements of high dynamics and multi-interactions in large-space VR scenarios. Specifically as follows:

[0071] (1) Attention mechanism

[0072]

[0073] Among them, LeakyReLU(·) represents the non-linear activation function; the subscript τ represents space or time; N τ (j) represents the spatial neighborhood or temporal neighborhood of node j; W represents the feature transformation matrix, which is used to transform the input feature x k into a new feature representation; is the normalization operation, ensuring that the sum of the neighborhood weights of each node is 1.

[0074] (2) Spatio-temporal feature aggregation: In ST-GNN, the new feature of a node depends not only on its spatial domain but also on its temporal domain. Therefore, the spatio-temporal aggregation operation is:

[0075]

[0076] Among them, represents the aggregated features; N s (i1) and N t (i1) represent the spatial neighborhood and temporal neighborhood of node i1 respectively; x′ j,s represents the spatial node features; x′ j,t represents the temporal node features; Aggregate represents the aggregation operation.

[0077] (3) Spatiotemporal graph update: In each layer, the graph structure is dynamically updated according to the spatiotemporal features of the current node to capture the dynamic changes of the signal. Specifically, the attention mechanism is used to calculate the weight q i,j between node i and node j, and then whether this edge exists is dynamically adjusted according to the magnitude of the weight. If q i,j < ζ, this edge can be deleted; conversely, if q i,j ≥ ζ, an edge can be added, where ζ is a threshold.

[0078]

[0079] Among them, μ is a learnable attention vector; W is a feature transformation matrix; ||· represents the concatenation operation, which concatenates the feature vectors of node i and node j so that the attention mechanism can consider the features of both the central node and the neighborhood nodes at the same time; LeakyReLU is an activation function used to introduce non-linearity, which allows a part of negative values to pass through to avoid the problem of gradient vanishing.

[0080] The number of stacked layers of the "spatiotemporal graph neural network" is the biggest difference between the teacher model and the student model. In the teacher model, through N layers of stacking (where N ≥ 3, in this embodiment, N = 5 is taken to meet the requirements of multi-task feature extraction), the model can gradually learn higher-level feature representations. In the student model, only 1 layer is used to make the model lightweight and meet the real-time requirements.

[0081] In the output stage, a combination of three "fully connected" and "Softmax" is adopted to output three labels, namely "position label probability distribution", "status label probability distribution", and "behavior label probability distribution".

[0082] Among them, the position label is as follows:

[0083] 1) Objective: Predict the position of the user in the large-space indoor.

[0084] 2) Label design: The position label is a classification label indicating the area or position where the user is located. For example, if the large-space indoor area is divided into 50 different small areas, the position label is an integer, and the range is.

[0085] The status labels are as follows:

[0086] 1) Objective: To identify the user's current activity status (walking, stationary, and the face orientation when stationary).

[0087] 2) Label design: The activity status label is a multi-class label indicating the user's current activity status. For example, the label for "running" is 0, the label for "walking" is 1, and the labels for "stationary and facing east", "stationary and facing west", "stationary and facing south", "stationary and facing north" are 2, 3, 4, 5 respectively. The labels for "stationary and facing southeast", "stationary and facing northeast", "stationary and facing southwest", "stationary and facing northwest" are 6, 7, 8, 9 respectively.

[0088] The behavior labels are as follows:

[0089] 1) Objective: To identify the gestures performed by the user, such as pointing gestures, grasping gestures, waving gestures, rotating gestures, zooming gestures, swiping gestures, palm-outward gestures, bending the four fingers to hold the thumb, opening and closing a clenched fist, etc.

[0090] Pointing gesture: The user points a finger in a certain direction or at an object to express selection or indication.

[0091] Grasping gesture: The user simulates the action of grasping an object and "grasps" an item in the virtual world by closing the fingers.

[0092] Waving gesture: The user gently waves the arm to perform specific operations, such as switching interfaces, opening menus, or greeting.

[0093] Rotating gesture: The user rotates the wrist or arm to make the virtual object rotate accordingly. This gesture allows the user to perform more precise manipulation of virtual items.

[0094] Zooming gesture: The user adjusts the size of the virtual object by opening and closing the fingers. This gesture is very useful when viewing details or adjusting the layout.

[0095] Swiping gesture: The user swipes a finger on the virtual interface to move the cursor, scroll the page, or switch views. This gesture is widely used in browsing and navigation.

[0096] Palm-outward gesture: In some VR systems, the user can perform specific functions, such as recalibrating the perspective or returning to the main interface, by turning the palm outward.

[0097] Bending the four fingers to hold the thumb: This action may be used to drag content or perform other operations that require fine control.

[0098] Clench and then open the fist: In some VR devices, this action may be used to switch interfaces or perform other important operations.

[0099] The gesture label is a classification label indicating the type of gesture performed by the user. For example, if there are 9 different gestures, the gesture label can be an integer in the range [0, 8].

[0100] Finally, a joint loss function is designed for the model to balance the weights of different tasks:

[0101]

[0102] where β represents the balance coefficient; represents the original loss function of the student model; represents the distillation loss function.

[0103] Among them, the distillation loss function is expressed as follows:

[0104]

[0105] where N represents the number of samples; is the predicted probability distribution of the teacher model for each class; is the predicted probability distribution of the student model for each class; represents the KL divergence, which is used to measure the difference between the output distribution of the teacher model and the output distribution of the student model. By minimizing this difference, the student model can learn the "soft" labels of the teacher model and thus inherit the knowledge and generalization ability of the teacher model. The KL divergence is non - negative, and when the two distributions are exactly the same, the KL divergence is 0. Therefore, the goal of the student model is to minimize this value. and are the original output feature vectors of the teacher model and the student model respectively. T is the temperature parameter, which is used to adjust the "soft" degree of the output distribution of the teacher model. By increasing the temperature parameter, the output distribution of the teacher model can be made smoother, thus providing richer information.

[0106] S3. Use the real - time tracking model to complete the real - time tracking of multiple targets in a large VR space.

[0107] Deploy the constructed real - time tracking model in the VR device to achieve the real - time tracking of multiple targets in a large VR space.

[0108] Example Two

[0109] This embodiment also provides a large-space VR multi-target high-dynamic real-time tracking system, including: a collection module, a construction module, and a tracking module; the collection module is used to collect the behavior data of multiple targets in the large VR space; the construction module is used to construct a real-time tracking model based on the behavior data; the real-time tracking model includes: a teacher model and a student model, both the teacher model and the student model are composed of an input stage, a spatio-temporal graph neural network, and an output stage; among them, the teacher model contains several pairs of spatio-temporal graph neural networks, and the student model contains 1 pair of spatio-temporal graph neural networks; the tracking module is used to utilize the real-time tracking model to complete the real-time tracking of multiple targets in the large VR space.

[0110] Next, in combination with this embodiment, it will be detailed how the present invention solves the technical problems in actual work.

[0111] First, use the collection module to collect the behavior data of multiple targets in the large VR space. The collection module is the collection device of the VR device itself. The input data of this embodiment is the Wi-Fi received signal strength indication. In addition, other solutions such as GPS, Bluetooth, RFID, and ultra-wideband can also be used. However, because Wi-Fi access points (APs) are widely available, there is no need to install additional infrastructure and hardware; using the existing Wi-Fi network, there is no need for additional hardware devices, reducing the deployment cost; the Wi-Fi signal coverage is wide, usually covering the entire building or multiple floors, suitable for large-scale indoor positioning; the Wi-Fi signal has strong penetration ability and can penetrate walls and obstacles to a certain extent, suitable for complex indoor environments; the change of Wi-Fi signal strength can reflect the relative position change between the device and the access point, suitable for dynamic environments; most modern devices (such as VR, smartphones, tablets, laptops) support Wi-Fi, with wide compatibility; using the existing Wi-Fi network, data collection is simple, only need to record the RSSI value between the device and the access point; the positioning accuracy through the RSSI value is usually within the range of 1-5 meters, suitable for most indoor positioning requirements; in summary, this embodiment uses Wi-Fi RSSI as the input data of the model.

[0112] After that, the construction module deployed on the VR device constructs a real-time tracking model based on the behavior data. In this embodiment, the real-time tracking model includes: a teacher model and a student model, both the teacher model and the student model are composed of an input stage, a spatio-temporal graph neural network, and an output stage; among them, the teacher model contains N pairs (where N≥3, and N = 5 is selected in this embodiment) of spatio-temporal graph neural networks, and the student model contains 1 pair of spatio-temporal graph neural networks, and its structure is as Figure 1 shown.

[0113] First, in the model input stage, data preprocessing is used to address the data sparsity issue in the received data. The data sparsity issue typically manifests as follows: missing values in the time dimension, i.e., there may be no data records for some time steps; missing values in the spatial dimension, i.e., some nodes may have no data at some time steps; and noisy data, i.e., the recorded data may contain noise, resulting in fluctuations in signal strength. These issues can affect the performance of the model. Especially in a dynamic environment, data sparsity and noise may cause the model to fail to accurately capture spatio-temporal relationships.

[0114] The data preprocessing steps include: serialization, selecting a time window, interpolation, temporal smoothing, and normalization.

[0115] The specific steps are as follows:

[0116] Serialization: The RSSI data needs to be converted into a time series format, i.e., for each location, record its RSSI measurement values at different time points.

[0117] Selecting a time window: It is necessary to determine an appropriate time window size to capture the dynamic changes of the signal in the time dimension. The selection of the time window will affect the performance and computational complexity of the model.

[0118] Then, a combined method of interpolation and temporal smoothing is adopted, with interpolation first and then smoothing. That is, first, data interpolation techniques are used to fill in the missing values, and then temporal smoothing techniques are used to reduce data fluctuations and noise. Data interpolation is a method of estimating unknown data points through known data points and can effectively fill in the missing parts of the data. Temporal smoothing techniques reduce data fluctuations and noise and enhance data continuity by applying smoothing operations to the time series. Among them, interpolation algorithms can include linear interpolation, polynomial interpolation, and spline interpolation, etc., and temporal smoothing algorithms can include moving average, exponentially weighted moving average, and Gaussian smoothing, etc.

[0119] In this embodiment, a combination of spline interpolation and Gaussian smoothing is adopted, which meets the requirements of high-precision interpolation and is sensitive to noise. Specifically, spline interpolation: Use a piecewise polynomial function to fit the known data points and estimate the missing values through the piecewise polynomial function, which is suitable for scenarios where the data changes relatively smoothly and high-precision interpolation is required.

[0120] S(t) = a i + b i (t - t i ) + c i (t - t i ) 2 + d i (t - t i ) 3

[0121] Among them, S(t) represents the signal value obtained by spline interpolation at the current time point t; a i 、b i 、c i 、d i all represent the coefficients of the piecewise polynomial; t represents the current time point; t i represents the time position corresponding to the i-th data point.

[0122] Gaussian smoothing: At each time step, the Gaussian kernel function is used to perform weighted averaging on the surrounding data to smooth the data, which is applicable to scenarios where the data needs to be smoothed and is sensitive to noise:

[0123]

[0124] Among them, y(t) represents the signal value after Gaussian smoothing at the time point t; i0 represents the offset of the time point t; k represents the size of the smoothing window; σ represents the standard deviation of the Gaussian kernel.

[0125] Finally, normalization can make the input data have a unified scale numerically, avoiding excessive numerical differences between different features; normalization can accelerate the convergence speed during the training process of the neural network, making the weight update process of the neural network smoother and more stable; normalization can eliminate the influence of part of the noise and outliers in the input data on the network, making the neural network more focused on learning the core features and patterns in the data.

[0126] Spatio-Temporal Graph Neural Network (ST-GNN) is an extended graph neural network architecture that models the dynamic characteristics of data through a spatio-temporal graph structure and is specifically used to process data with spatio-temporal dependencies. It not only considers the spatial relationship between nodes (such as the proximity of RSSI data), but also introduces the time dimension to capture the changes of signals over time. Its characteristics are:

[0127] (1) In the ordinary graph neural network, the graph structure is static, while in ST-GNN, the graph structure changes over time.

[0128] (2) The features of nodes include not only spatial information (such as RSSI values) but also time information (such as the change of signal strength over time).

[0129] (3) ST-GNN dynamically updates the graph structure according to the time series to adapt to environmental changes.

[0130] The steps for the Spatio-Temporal Graph Neural Network to construct a spatio-temporal graph are as follows:

[0131] (1) Construction of the spatial graph: Similar to ordinary GNN, kNN is used to construct the spatial graph to represent the spatial relationship between nodes.

[0132] (2) Construction of the time graph: Introduce the time dimension and connect the measurement values of the same node at different time points, that is, connect the nodes at adjacent time points with edges, and the weight of the edge can represent the time correlation.

[0133] (3) In the spatial graph and the time graph, each measurement point is regarded as a node in the graph, and the feature of each node is its corresponding feature vector, which represents the Wi-Fi RSSI signal strength measured at that position. The edges between nodes are constructed by the kNN (k-Nearest Neighbors) algorithm, and the weight of the edge is determined by the similarity between different feature vectors, that is, the Euclidean distance between the feature vectors.

[0134] (4) Representation of the spatio-temporal graph: The spatio-temporal graph can be represented as G=(V, E s , W t ), where E s represents the spatial edge, and E t represents the time edge.

[0135] After that, perform convolution operations on the spatio-temporal graph. Through spatial correlation modeling and time dynamics analysis, achieve efficient feature extraction in complex environments; at the same time, combine lightweight design and multi-task collaboration to significantly improve the accuracy and robustness of multi-object tracking while ensuring real-time performance, so as to meet the stringent requirements of high dynamics and multi-interactions in large-space VR scenarios. Specifically as follows:

[0136] (1) Attention mechanism

[0137]

[0138] Among them, LeakReLU(·) represents the non-linear activation function; the subscript τ represents space or time; N τ (j) represents the spatial neighborhood or time neighborhood of node j; W represents the feature transformation matrix, which is used to transform the input feature x k into a new feature representation; is the normalization operation to ensure that the sum of the neighborhood weights of each node is 1.

[0139] (2) Spatio-temporal feature aggregation: In ST-GNN, the new feature of a node depends not only on its spatial domain but also on its time domain. Therefore, the spatio-temporal aggregation operation is:

[0140]

[0141] Among them, represents the aggregated feature; N s (i1) and N t (i1) represent the spatial neighborhood and time neighborhood of node i1 respectively; x′ j,sRepresents spatial node features; x′ j,t Represents temporal node features; Aggregate represents the aggregation operation.

[0142] (3) Spatiotemporal graph update: In each layer, the graph structure is dynamically updated according to the spatiotemporal features of the current nodes to capture the dynamic changes of the signals. Specifically, the attention mechanism is used to calculate the weight q i,j between node i and node j, and then whether this edge exists is dynamically adjusted according to the magnitude of the weight. If q i,j < ζ, this edge can be deleted; conversely, if q i,j ≥ ζ, an edge can be added, where ζ is a threshold.

[0143]

[0144] Among them, μ is a learnable attention vector; W is a feature transformation matrix; ||· represents the concatenation operation, which concatenates the feature vectors of node i and node j so that the attention mechanism can consider the features of the central node and the neighborhood nodes simultaneously; LeakyReLU is an activation function used to introduce non-linearity, which allows a part of negative values to pass through to avoid the problem of gradient disappearance.

[0145] The number of stacked layers of the "spatiotemporal graph neural network" is the biggest difference between the teacher model and the student model. In the teacher model, through N layers of stacking (where N≥3, in this embodiment, N = 5 is taken to meet the requirements of multi-task feature extraction), the model can gradually learn higher-level feature representations. In the student model, only 1 layer is used, so that the model is lightweight and meets the requirements of real-time performance.

[0146] In the output stage, a combination of three "fully connected" and "Softmax" is adopted to output three types of labels, namely "position label probability distribution", "state label probability distribution", and "behavior label probability distribution".

[0147] The position labels are as follows:

[0148] 1) Objective: Predict the position of the user in the large indoor space.

[0149] 2) Label design: The position label is a classification label indicating the area or position where the user is located. For example, if the large indoor area is divided into 50 different small areas, the position label is an integer, and the range is.

[0150] The state labels are as follows:

[0151] 1) Objective: Identify the current activity state of the user (walking, stationary, and the face orientation when stationary).

[0152] 2) Label Design: The activity status label is a multi-class label indicating the user's current activity status. For example, the label for "running" is 0, the label for "walking" is 1, and the labels for "stationary and facing east", "stationary and facing west", "stationary and facing south", "stationary and facing north" are 2, 3, 4, 5 respectively. The labels for "stationary and facing southeast", "stationary and facing northeast", "stationary and facing southwest", "stationary and facing northwest" are 6, 7, 8, 9 respectively.

[0153] The behavior labels are as follows:

[0154] 1) Objectives: Identify the gestures performed by the user, such as pointing gestures, grasping gestures, waving gestures, rotating gestures, scaling gestures, swiping gestures, palm-outward gestures, bending the four fingers to hold the thumb, clenching the fist and then opening it, etc.

[0155] Pointing Gesture: The user points a finger in a certain direction or at an object to express a selection or indication.

[0156] Grasping Gesture: The user simulates the action of grasping an object and "grasps" an item in the virtual world by closing the fingers.

[0157] Waving Gesture: The user gently waves the arm to perform a specific operation, such as switching interfaces, opening a menu, or greeting.

[0158] Rotating Gesture: The user rotates the wrist or arm to make the virtual object rotate accordingly. This gesture allows the user to perform more precise manipulation of the virtual item.

[0159] Scaling Gesture: The user adjusts the size of the virtual object by opening and closing the fingers. This gesture is very useful when viewing details or adjusting the layout.

[0160] Swiping Gesture: The user swipes a finger on the virtual interface to move the cursor, scroll the page, or switch views. This gesture is widely used in browsing and navigation.

[0161] Palm-Outward Gesture: In some VR systems, the user can perform a specific function, such as recalibrating the view or returning to the main interface, by turning the palm outward.

[0162] Bending the Four Fingers to Hold the Thumb: This action may be used to drag content or perform other operations that require fine control.

[0163] Clenching the Fist and then Opening It: In some VR devices, this action may be used to switch interfaces or perform other important operations.

[0164] The gesture label is a classification label indicating the type of gesture performed by the user. For example, if there are 9 different gestures, the gesture label can be an integer in the range [0, 8].

[0165] Finally, the combined loss function is designed for the model to balance the weights of different tasks:

[0166]

[0167] where β represents the balance coefficient; represents the original loss function of the student model; represents the distillation loss function.

[0168] Among them, the distillation loss function is expressed as follows:

[0169]

[0170] where N represents the number of samples; is the predicted probability distribution of the teacher model for each category; is the predicted probability distribution of the student model for each category; represents the KL divergence, which is used to measure the difference between the output distribution of the teacher model and the output distribution of the student model. By minimizing this difference, the student model can learn the "soft" labels of the teacher model, thereby inheriting the knowledge and generalization ability of the teacher model. The KL divergence is non - negative, and when the two distributions are exactly the same, the KL divergence is 0. Therefore, the goal of the student model is to minimize this value. and are the original output feature vectors of the teacher model and the student model respectively. T is the temperature parameter, which is used to adjust the "soft" degree of the output distribution of the teacher model. By increasing the temperature parameter, the output distribution of the teacher model can be made smoother, thereby providing richer information.

[0171] Finally, the real - time tracking function of the VR device is optimized by using the implementation tracking model.

[0172] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A large-space VR multi-target high-dynamic real-time tracking method, characterized in that the steps Including: Collecting the behavior data of multiple targets in a large VR space; Constructing a real-time tracking model based on the behavior data; The real-time tracking model includes: a teacher model and a student model, both of which are composed of an input stage, a spatio-temporal graph neural network, and an output stage; among them, the teacher model contains several pairs of spatio-temporal graph neural networks, and the student model contains 1 pair of spatio-temporal graph neural networks; Using the real-time tracking model to complete the real-time tracking of multiple targets in a large VR space.

2. The large-space VR multi-target high-dynamic real-time tracking method according to claim 1, wherein The real-time tracking model solves the problem of data sparsity through data preprocessing in the input stage; Then, a spatio-temporal graph neural network with a multi-task shared backbone feature extraction layer is adopted to improve the accuracy of positioning and behavior recognition; Finally, knowledge distillation technology is used to compress its own complex model.

3. The large-space VR multi-target high-dynamic real-time tracking method according to claim 1, characterized in that The difference between the teacher model and the student model lies in the stacking layers of the spatio-temporal graph neural network; the teacher model learns higher-level feature representations through a multi-layer stacking method; the student model realizes model lightweighting through one layer of the spatio-temporal graph neural network to meet dynamic adaptability.

4. The large-space VR multi-target high-dynamic real-time tracking method according to claim 2, characterized in that The real-time tracking model adopts a combined method of interpolation and temporal smoothing to solve the problem of data sparsity; among them, interpolation adopts the spline interpolation method: S(t) = a i + b i (t - t i ) + c i (t - t i ) 2 + d i (t - t i ) 3 Among them, S(t) represents the signal value obtained by spline interpolation at the current time point t; a i , b i , c i , d i all represent the coefficients of the piecewise polynomial; t represents the current time point; t i represents the time position corresponding to the i-th data point; Temporal smoothing adopts the Gaussian smoothing method: Among them, y(t) represents the signal value after Gaussian smoothing at time point t; i0 represents the offset of time point t; k represents the size of the smoothing window; σ represents the standard deviation of the Gaussian kernel.

5. The large-space VR multi-target high-dynamic real-time tracking method according to claim 2, characterized in that The spatio-temporal graph neural network improves the accuracy of positioning and behavior recognition through spatio-temporal feature aggregation: Among them, represents the aggregated feature; N s (i1) and N t (i1) respectively represent the spatial neighborhood and temporal neighborhood of node i1; x′ j,s represents the spatial node feature; x′ j,t represents the temporal node feature; Aggregate represents the aggregation operation.

6. The large-space VR multi-target high-dynamic real-time tracking method according to claim 1, characterized in that The output stage adopts a combination of three "fully connected" and "Softmax" to output three types of labels, namely the probability distribution of position labels, the probability distribution of state labels, and the probability distribution of behavior labels.

7. The large-space VR multi-target high-dynamic real-time tracking method according to claim 1, wherein The real-time tracking model adopts a joint loss function to balance the weights of different tasks: Among them, β represents the balance coefficient; represents the original loss function of the student model; represents the distillation loss function.

8. A large-space VR multi-target high-dynamic real-time tracking system, which is used to implement the method described in any one of claims 1-7, and is characterized in that, Including: A collection module, a construction module, and a tracking module; The collection module is used to collect the behavior data of multiple targets in a large VR space; The construction module is used to construct a real-time tracking model based on the behavior data; The real-time tracking model includes: a teacher model and a student model, both of which are composed of an input stage, a spatio-temporal graph neural network, and an output stage; among them, the teacher model contains several pairs of spatio-temporal graph neural networks, and the student model contains 1 pair of spatio-temporal graph neural networks; The tracking module is used to complete the real-time tracking of multiple targets in a large VR space by using the real-time tracking model.

9. The large-space VR multi-target high-dynamic real-time tracking system according to claim 8, characterized in that, The real-time tracking model solves the problem of data sparsity through data preprocessing in the input stage; Then, a spatio-temporal graph neural network with a multi-task shared backbone feature extraction layer is adopted to improve the accuracy of positioning and behavior recognition; Finally, knowledge distillation technology is used to compress its own complex model.

10. The large-space VR multi-target high-dynamic real-time tracking system according to claim 8, characterized in that, The difference between the teacher model and the student model lies in the stacking layers of the spatio-temporal graph neural network; the teacher model learns higher-level feature representations through a multi-layer stacking method; the student model realizes model lightweighting through one layer of the spatio-temporal graph neural network to meet dynamic adaptability.

Citation Information

Patent Citations

  • Speech enhancement method based on cross-layer similarity knowledge distillation

    CN114067819A

  • Human body behavior recognition method combining key nodes and enhanced data guidance

    CN115035600A

  • Large-space multi-person VR interactive experience system and method

    CN116679834A

  • Cross-modal image rotation target identification method and system based on knowledge distillation

    CN118781331A

  • Garbage classification putting behavior identification method and system based on AI algorithm

    CN119049133A