Indoor pedestrian positioning method based on ResNet and Transform networks

By combining ResNet and Transformer networks, the motion characteristics of IMU data are extracted and time series characteristics are captured, which solves the problem of difficulty in distinguishing cumulative error and speed in IMU positioning, and achieves higher positioning accuracy and robustness.

CN120084336APending Publication Date: 2025-06-03CHANGZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510180705.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing IMU-based indoor positioning method has cumulative error problems, which leads to a significant increase in positioning errors. It is difficult for a single neural network to distinguish the magnitude of speed, resulting in a degradation of dynamic performance and waste of resources.

Method used

The indoor pedestrian positioning method based on ResNet and Transformer networks is adopted to extract the motion characteristics of IMU data through ResNet, predict the instantaneous velocity, and input the velocity sequence into the Transformer model. The time series features are captured through the self-attention mechanism to optimize the trajectory of the motion direction.

Benefits of technology

It significantly improves the accuracy and robustness of indoor positioning, can better adapt to complex environments, reduce computing complexity and energy consumption, and improves the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120084336A_ABST
    Figure CN120084336A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of inertial navigation, in particular to an indoor pedestrian positioning method based on a ResNet and Transform network, which comprises the following steps: firstly, acquiring IMU (Inertial Measurement Unit) data of a smart phone, acquiring a real motion track of a pedestrian through equipment carrying a Tango technology, and preprocessing the acquired data; extracting motion features of the IMU data through a ResNet network, and predicting instantaneous speed based on residual mapping between learning input data; then, inputting the predicted instantaneous speed sequence into a Transform model, capturing time sequence features of IMU data through a self-attention mechanism, and combining with speed prediction to optimize a trajectory in a motion direction; and finally, evaluating the accuracy of navigation positioning according to the predicted trajectory and the real trajectory. According to the method, the deep learning technology and IMU data characteristics are fully combined, the accuracy and robustness of trajectory prediction are remarkably improved, and the method has high practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of inertial navigation, and in particular to an indoor pedestrian positioning method based on ResNet and Transformer networks. Background Art

[0002] Positioning technology has become an indispensable part of our daily lives, enabling various applications such as navigation, tracking, and environmental perception. Precise indoor positioning technology is particularly important because people spend a significant amount of time searching for destinations inside buildings, such as shopping malls, airports, and offices.

[0003] Due to the attenuation of satellite signals and signal blockage by walls and roofs, traditional Global Positioning System (GPS)-based positioning has limitations when used indoors. A promising approach to indoor positioning is Visual Simultaneous Localization and Mapping (SLAM) technology. Visual SLAM uses cameras to map the environment and estimate the position of the camera in the map, showing great potential in achieving precise indoor positioning. However, it is usually impractical to deploy unobstructed cameras in indoor environments, and the power consumption resulting from continuous use of cameras is not suitable for daily use.

[0004] Thus, the importance of the IMU (Inertial Measurement Unit) in indoor positioning is reflected. The IMU consists of sensors such as accelerometers and gyroscopes and is used to measure the movement and orientation of a device. IMUs are commonly found in consumer devices such as smartphones and provide basic indoor positioning capabilities. However, they suffer from the problem of cumulative error, that is, the measurement error will accumulate over time, resulting in significant positioning errors.

[0005] To address the limitations of IMU-based positioning, deep learning-based methods have emerged. By using neural networks to predict motion, the cumulative error generated by the IMU can be effectively reduced. These neural networks learn from large-scale datasets to predict and correct positioning errors and improve positioning accuracy. Combining deep learning with IMUs and other environmental sensors has become an effective way to solve the problem of cumulative error in IMU positioning. Existing single neural networks cannot distinguish the magnitude of speed, resulting in excessive focus on direction estimation when the speed is slow, leading to degraded dynamic performance, reduced estimation accuracy, slower convergence speed, poorer robustness, and resource waste. Therefore, it is necessary to use a combination of multiple models and optimize the structure and training method of the neural network. Summary of the Invention

[0006] The technical problem to be solved by the present invention is: to solve the problems existing in the prior art in the above-mentioned background art, and to provide an indoor pedestrian positioning method based on ResNet and Transformer networks.

[0007] The technical solution adopted by the present invention to solve its technical problems is: an indoor pedestrian positioning method based on ResNet and Transformer networks, including the following steps: S1. Collect the IMU data of the smartphone, and at the same time obtain the real movement trajectory of the pedestrian through the device equipped with Tango technology, and preprocess the collected data; S2. Extract the motion features of the IMU data through the ResNet network, and predict the instantaneous velocity based on learning the residual mapping between the input data; S3. Input the predicted instantaneous velocity sequence into the Transformer model, capture the time series features of the IMU data through the self-attention mechanism, and optimize the trajectory of the motion direction in combination with the velocity prediction; S4. Evaluate the accuracy of the navigation and positioning according to the predicted trajectory and the real trajectory.

[0008] Further, the collection of the IMU data of the smartphone in step S1 is specifically: collect the data of the three-axis acceleration and the three-axis angular velocity through the IMU sensor on the smartphone, and the collection frequency is 200Hz.

[0009] Furthermore, the device equipped with Tango technology uses a smartphone with the model number ASUS_A002A, which integrates the Tango technology developed by Google, and this Tango technology can obtain the real movement trajectory of people through visual recognition technology.

[0010] Furthermore, during the collection of the IMU data, hold the smartphone in a natural way to simulate the daily use scenario.

[0011] Furthermore, the preprocessing in step S1 is specifically: S11. Denoise the acceleration and angular velocity data through Kalman filtering to filter out high-frequency interference signals; S12. Realize the synchronization of the IMU data and the real trajectory in time through the Bluetooth module: S13. First perform coordinate transformation on the IMU data to unify it to the global inertial coordinate system to eliminate the influence of the device attitude change on the data. The calculation formula is as follows:

[0012]

[0013] In this formula, and represents the acceleration vector and angular velocity vector in the device coordinate system of the device equipped with Tango technology at the i moment, and represents the acceleration vector and angular velocity vector in the device coordinate system of the IMU device at the moment; represents the rotation matrix from the IMU device coordinate system to the device coordinate system of the device equipped with Tango technology at the i moment. The calculation formula of its quaternion representation is as follows:

[0014] wherein, represents the quaternion from the IMU device coordinate system to the device coordinate system of the device equipped with Tango technology at the i moment, and respectively represent the attitude quaternions of the IMU and the device equipped with Tango technology in their respective device coordinate systems at the initial moment; represents the rotation quaternion from the IMU device coordinate system to the device coordinate system of the device equipped with Tango technology at the initial moment, which is determined by fixing the initial relative attitude of the two devices.

[0015] Furthermore, the motion features of the IMU data are extracted through the ResNet network in the step S2, specifically: S21. Input the preprocessed IMU acceleration and angular velocity data into the ResNet network, and map the original data to a high-dimensional feature space through the input layer; S22. Construct multiple residual blocks in the residual network. Each residual block extracts local motion features in the IMU data through the combination of a convolutional layer, a batch normalization layer, and an activation function, and at the same time alleviates the problem of gradient disappearance through residual connections; S23. Use the cascading of multiple residual blocks to generate a global feature vector to capture the comprehensive motion features of the IMU data; S24. Perform dimensionality reduction processing on the feature vector through the global average pooling layer to reduce redundant information; S25. Through the fully connected network of the output layer, map the dimensionality-reduced feature vector to the predicted value of the instantaneous velocity.

[0016] Even further, based on the predicted value of the instantaneous velocity, adopt the mean square error loss function to optimize the model parameters to minimize the error between the predicted velocity and the true velocity. Its formula is:

[0017] In this formula, n represents the number of samples, is the predicted speed value of the model for the sample, while is the corresponding actual speed value.

[0018] Furthermore, the time series features of the IMU data captured in step S3 are specifically as follows: S31. Use the speed sequence predicted by the ResNet network as the input, and jointly input it into the encoder structure of the Transformer network together with the time series features of the IMU data; S32. Map the input data to high-dimensional feature vectors through the embedding layer, and add position encoding information to clarify the order of each data point in the time series; S33. Apply the multi-head self-attention mechanism in the encoder to calculate the correlation between the input data and capture the motion features of long-term dependencies in the time series. The calculation formula is:

[0019] where Q , K , V represent query, key, and value respectively, d k is the dimension of the key matrix; S34. Use the multi-layer encoder structure to further extract deep features, and map the feature vectors to the predicted values of the motion direction through the fully connected layer.

[0020] Even further, according to the magnitude of the speed sequence, dynamically adjust the data window for direction prediction. The specific method is as follows: When the speed magnitude exceeds the set threshold, discard some data points in the window to reduce the weight of historical data. The adaptive adjustment function is:

[0021] where, m is the number of data points in the window to be adjusted, v low and v high are the lower and upper limits of the speed respectively; optimize the direction prediction through the cosine similarity loss function. The formula is:

[0022] where, is the angle between the predicted direction and the true direction, is the magnitude of the predicted speed; Combine the optimized direction prediction and speed prediction results to generate a complete motion trajectory.

[0023] Further, the specific steps of step S4 are as follows: S41. Align the predicted trajectory generated from the velocity sequence and direction sequence with the real trajectory recorded using Tango technology to ensure consistency in the time dimension; S42. Calculate the position information of each trajectory point in the predicted trajectory using an integral algorithm, and calculate the absolute trajectory error and relative trajectory error between the predicted trajectory and the real trajectory respectively based on the evaluation index of trajectory error. ATE And the relative trajectory error RTE , where the formula for the absolute trajectory error ATE is:

[0024] Where is the i-th predicted position, is the corresponding real position, N is the total number of trajectory points; the formula for the relative trajectory error RTE is:

[0025] Where k is the length of the evaluation time window, and are the i + j-th predicted position and real position respectively; By statistically analyzing the absolute trajectory error ATE and the relative trajectory error RTE , evaluate the positioning accuracy of the model under different speeds and different environments.

[0026] Advantages of the present invention: The combination of ResNet and Transformer networks in the present invention can give full play to the advantages of ResNet in extracting local features and Transformer in capturing global information; The deep residual structure of ResNet can efficiently extract local details in images, while the self-attention mechanism of Transformer can process global information and further integrate the complex relationships between features, thus significantly improving the positioning accuracy; The global information processing ability of Transformer enables the model to better cope with complex indoor environmental changes. This combination method significantly improves the performance under multiple conditions, showing higher tracking accuracy and robustness to environmental factors; By fusing ResNet and Transformer, the model can effectively reduce the computational complexity and energy consumption while maintaining high performance; This method can better adapt to complex environments when dealing with indoor pedestrian positioning, and the model can more accurately identify and track targets; After pre-training, the model based on ResNet and Transformer can quickly adapt to new tasks or datasets through transfer learning, further enhancing the generalization ability and application scope of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The present invention will be further described below in conjunction with the drawings and embodiments.

[0028] Figure 1 is a flowchart of the indoor pedestrian positioning method based on the ResNet and Transformer networks of the present invention.

[0029] Figure 2 is a diagram for collecting experimental data of the indoor pedestrian positioning method based on the ResNet and Transformer networks.

[0030] Figure 3 is a diagram of the speed branch model of the indoor pedestrian positioning method based on the ResNet and Transformer networks of the present invention.

[0031] Figure 4 is a diagram of the direction branch model of the indoor pedestrian positioning method based on the ResNet and Transformer networks of the present invention.

[0032] Figure 5 is a diagram of the window dynamic adjustment of the indoor pedestrian positioning method based on the ResNet and Transformer networks of the present invention.

[0033] Figure 6 is a diagram comparing the errors of five models of the indoor pedestrian positioning method based on the ResNet and Transformer networks of the present invention on two datasets.

[0034] Figure 7 is a diagram comparing the trajectories of different models of the indoor pedestrian positioning method based on the ResNet and Transformer networks of the present invention on two datasets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The present invention will now be described in further detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, and therefore only showing the components related to the present invention.

[0036] As Figure 1As shown, the indoor pedestrian positioning method based on ResNet and Transformer networks has the following process: Step 1, collect the IMU data of the smartphone, and at the same time obtain the real movement trajectory of the pedestrian through the device equipped with Tango technology, and preprocess the collected data; Step 2, extract the movement features of the IMU data through the ResNet network, and predict the instantaneous speed based on learning the residual mapping between the input data; Step 3, input the predicted instantaneous speed sequence into the Transformer model, capture the time series features of the IMU data through the self-attention mechanism, and optimize the trajectory of the movement direction in combination with the speed prediction; Step 4, evaluate the accuracy of the navigation positioning according to the predicted trajectory and the real trajectory.

[0037] If the magnitude of the speed is predicted first and the magnitude of the speed is added to the loss function of the direction, it is found through experiments that it is better to use the transformer network for processing time series to predict the movement direction.

[0038] In Step 1, the IMU sensor on the smartphone collects triaxial acceleration and triaxial angular velocity data, with a collection frequency of 200 Hz. The collection methods include various movement methods such as holding and swinging the arm, and the collection locations are divided into two different environments: indoor and outdoor. During the implementation of this embodiment, a smartphone with the model ASUS_A002A was used as the main data collection device. This phone integrates the Tango technology developed by Google. Tango technology is an advanced augmented reality (AR) and computer vision technology that can use the depth sensor and camera in the device to capture the three-dimensional structure of the environment in real time. Through visual recognition technology, the Tango system can accurately track the position and movement trajectory of the device in space, generate high-precision movement labels as a reference for the real trajectory. It should be noted that the device for collecting IMU data is any smartphone, and the device equipped with Tango technology is a smartphone with the specified model ASUS_A002A.

[0039] As Figure 2As shown, during the data collection process, the device equipped with Tango technology (i.e., the smartphone with the model ASUS_A002A) is fixed on the chest of the collector to record the movement trajectory. During the collection, the experimenter holds the smartphone in a natural way to simulate daily usage scenarios, such as walking, turning, going up and down stairs, etc. The way of holding can capture the subtle dynamic changes generated by the device following the hand movement, thus comprehensively reflecting the movement pattern of the device in a complex environment. The collected IMU data, including three-axis acceleration and three-axis angular velocity information, is affected by various factors, such as device vibration, internal noise of the sensor, and external environmental interference, etc. These high-frequency noise signals may significantly reduce the quality of the data, thus affecting the accuracy and stability of the subsequent positioning model. To solve this problem, the Kalman filter algorithm is adopted in the data processing stage to denoise the IMU data.

[0040] The two devices are connected via Bluetooth, and the synchronization of the IMU data and the real trajectory in time is achieved by recording the relative start time of operation: the data is subjected to coordinate system transformation and unified to the global inertial coordinate system to eliminate the influence of the device's attitude change on the data. The calculation formula is as follows:

[0041]

[0042] In this formula, and represent the acceleration vector and angular velocity vector in the coordinate system of the device equipped with Tango technology at time i , and and represent the acceleration vector and angular velocity vector in the coordinate system of the IMU device at time ; represents the rotation matrix from the coordinate system of the IMU device to the coordinate system of the device equipped with Tango technology at time i . Its calculation formula in quaternion representation is as follows:

[0043] Among them, represents the quaternion from the coordinate system of the IMU device to the coordinate system of the device equipped with Tango technology at time i , and respectively represent the attitude quaternions of the IMU and the device equipped with Tango technology in their respective device coordinate systems at the initial moment; represents the rotation quaternion from the coordinate system of the IMU device to the coordinate system of the device equipped with Tango technology at the initial moment, which is determined by fixing the initial relative attitude of the two devices.

[0044] As shown Figure 3 in the figure, step 2 is specifically as follows: First, the preprocessed IMU acceleration and angular velocity data are segmented into several data blocks with a window size of 200. Each data block represents the historical IMU data at the time point of the velocity to be predicted. These data blocks are input into the Residual Network (ResNet) with a batch size of 32. Through the input layer, the six-dimensional original data is mapped to a 64-dimensional high-dimensional feature space; Second, multiple residual blocks are constructed in the ResNet. Each residual block extracts local motion features in the IMU data by combining convolutional layers, batch normalization layers, and activation functions, and at the same time alleviates the problem of gradient disappearance through residual connections; Then, multiple residual blocks are cascaded to generate a global feature vector to capture the comprehensive motion characteristics of the IMU data; Next, the feature vector is reduced in dimension through the max-pooling layer to reduce redundant information; Then, through the fully connected network in the output layer, the reduced feature vector is mapped to the predicted value of the instantaneous velocity; Finally, based on the predicted value of the instantaneous velocity, the mean square error loss function is used to optimize the model parameters to minimize the error between the predicted velocity and the true velocity. The formula is as follows:

[0045] In this formula, n represents the number of samples, is the predicted velocity value of the model for the th sample, while is the corresponding actual velocity value.

[0046] Step 3 is specifically as follows: The velocity sequence predicted by the Residual Network (ResNet) is used as the input and jointly input into the encoder structure of the Transformer network together with the time series features of the IMU data. As Figure 4 shown in the figure, the Transformer network maps the input data to a high-dimensional feature vector through the embedding layer and adds position encoding information to clarify the order of each data point in the time series. The multi-head self-attention mechanism is applied in the encoder to calculate the correlation between the input data and capture the long-term dependent motion features in the time series. The calculation formula is as follows:

[0047] where Q , K , V represent query, key, and value respectively, d k is the dimension of the key matrix; The multi-layer encoder structure is used to further extract deep features, and the feature vector is mapped to the predicted value of the motion direction through the fully connected layer. According to the size of the velocity sequence, the data window for direction prediction is dynamically adjusted; As Figure 5As shown in the figure, the specific method is as follows: when the speed magnitude exceeds the set threshold, some data points in the window are discarded to reduce the weight of historical data. The adaptive adjustment function is:

[0048] Among them, m is the number of data points in the window to be adjusted, v low and v high are the lower and upper limits of the speed respectively; the direction prediction is optimized through the cosine similarity loss function, and its formula is:

[0049] Among them, is the angle between the predicted direction and the true direction, is the magnitude of the predicted speed; Combining the optimized direction prediction and speed prediction results, a complete motion trajectory is generated.

[0050] Step 4 is specifically as follows: Align the predicted trajectory generated from the speed sequence and direction sequence with the true trajectory recorded using Tango technology to ensure consistency in the time dimension; Use the integral algorithm to calculate the position information of each trajectory point in the predicted trajectory. Based on the evaluation index of trajectory error, calculate the absolute trajectory error ATE and relative trajectory error RTE between the predicted trajectory and the true trajectory respectively. Among them, the calculation formula of the absolute trajectory error ATE is:

[0051] Among them, is the i-th predicted position, is the corresponding true position, N is the total number of trajectory points; the calculation formula of the relative trajectory error RTE is:

[0052] Among them, k is the length of the evaluation time window, and are the (i + j)-th predicted position and true position respectively; By statistically analyzing the absolute trajectory error ATE and relative trajectory error RTE evaluate the positioning accuracy of the model under different speeds and different environments. Optimize the network structure or training parameters of the model according to the evaluation results to further improve the positioning accuracy and robustness.

[0053] Such as Figure 6As shown, the error comparison of the five models on two datasets. Due to the influence of the receptive field limitation of the TCN network, it may not be able to fully capture all key information when processing long-sequence data, which results in a larger prediction error compared with other network structures. The other several models adjusted the structures and parameters of their respective networks, and basically achieved the most accurate prediction under the original data error; while the combination of the ResNet and Transformer networks in this embodiment, considering both the spatial and temporal information of the task, can provide a more comprehensive perspective and higher prediction accuracy, that is, the effect is slightly better than several existing improved models.

[0054] As Figure 7 shown, the trajectories of the five models for some data on two datasets. Due to the errors, noises and biases in the IMU, the double-integration algorithm will accumulate large errors over time, resulting in a large direction offset in a short time and a large difference from the true trajectory. The methods based on neural networks have effectively reduced the error accumulation and can well estimate the motion trajectory. The model proposed in this embodiment is more in line with the actual situation in trajectory prediction. The speed branch model makes the overall trajectory size closer to the actual, and the direction branch model makes the model have a higher fitting degree when going straight and turning. This also shows the higher accuracy and stronger robustness of the combination of ResNet and Transformer networks in trajectory prediction.

[0055] Therefore, the indoor pedestrian positioning method based on the ResNet and Transformer networks in this embodiment can effectively improve the accuracy and reliability of pedestrian positioning in the indoor environment, while reducing manual participation and lowering costs.

[0056] Inspired by the above ideal embodiments of the present invention, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. An indoor pedestrian positioning method based on ResNet and Transformer network, characterized in that: The steps include: S1. Collect IMU data from smartphones, obtain pedestrians’ real motion trajectories through devices equipped with Tango technology, and pre-process the collected data; S2, extract the motion features of IMU data through the ResNet network, and predict the instantaneous speed based on the residual mapping between the learning input data; S3, input the predicted instantaneous speed sequence into the Transformer model, capture the time series characteristics of the IMU data through the self-attention mechanism, and optimize the trajectory of the motion direction in combination with the speed prediction; S4. Evaluate the accuracy of navigation positioning based on the predicted trajectory and the actual trajectory.

2. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 1, characterized in that: The step S1 of collecting the IMU data of the smartphone specifically includes collecting the data of three-axis acceleration and three-axis angular velocity through the IMU sensor on the smartphone, and the collection frequency is 200 Hz.

3. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 2 is characterized in that: The device equipped with Tango technology adopts a smartphone with model number ASUS_A002A, which integrates Tango technology developed by Google. The Tango technology can obtain the real movement trajectory of a person through visual recognition technology.

4. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 2 is characterized in that: During IMU data collection, the smartphone was held in a natural way to simulate daily usage scenarios.

5. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 3 is characterized in that: The preprocessing in step S1 is specifically as follows: S11, performing denoising on the acceleration and angular velocity data by using Kalman filtering to filter out high-frequency interference signals; S12, realize the time synchronization of IMU data and real trajectory through Bluetooth module: S13. First, convert the IMU data into a coordinate system and unify it into the global inertial coordinate system to eliminate the influence of the device posture change on the data. The calculation formula is as follows: In this formula, and Indicated in i The acceleration vector and angular velocity vector in the coordinate system of the device equipped with Tango technology at all times, and Indicated in The acceleration vector and angular velocity vector in the IMU device coordinate system at this moment; Indicated in i The rotation matrix that is converted from the IMU device coordinate system to the device coordinate system equipped with Tango technology at all times is calculated using the quaternion representation method as follows: in, Indicated in i The quaternion that is converted from the IMU device coordinate system to the device coordinate system equipped with Tango technology at all times. and Respectively represent the attitude quaternions of the IMU and the device equipped with Tango technology in their device coordinate systems at the initial moment; Represents the rotation quaternion from the IMU device coordinate system to the Tango-enabled device coordinate system at the initial moment, which is determined by fixing the initial relative posture of the two devices.

6. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 1, characterized in that: In step S2, the motion features of the IMU data are extracted through the ResNet network, specifically: S21, input the preprocessed IMU acceleration and angular velocity data into the ResNet network, and map the raw data into a high-dimensional feature space through the input layer; S22. Construct multiple residual blocks in the residual network. Each residual block extracts local motion features in the IMU data by combining a convolutional layer, a batch normalization layer, and an activation function, and alleviates the gradient vanishing problem through residual connections. S23, using multiple residual blocks in cascade to generate a global feature vector to capture the comprehensive motion characteristics of the IMU data; S24, reducing the dimension of the feature vector through a global average pooling layer to reduce redundant information; S25. Through the fully connected network of the output layer, the reduced-dimensional feature vector is mapped to the predicted value of the instantaneous speed.

7. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 6, characterized in that: Based on the predicted value of instantaneous speed, the mean square error loss function is used to optimize the model parameters to minimize the error between the predicted speed and the actual speed. The formula is: In this formula, n represents the number of samples, The model is The predicted speed value of the sample, and is the corresponding actual speed value.

8. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 1, characterized in that: The time series characteristics of the IMU data captured in step S3 are specifically: S31, the velocity sequence predicted by the ResNet network is used as input and input into the encoder structure of the Transformer network together with the time series features of the IMU data; S32, mapping the input data into a high-dimensional feature vector through an embedding layer and adding position encoding information to clarify the order of each data point in the time series; S33. Apply the multi-head self-attention mechanism in the encoder to calculate the correlation between the input data and capture the long-term dependent motion features in the time series. The calculation formula is: in Q , K , V Represents query, key and value respectively, d k is the dimension of the key matrix; S34. Use a multi-layer encoder structure to further extract deep features, and map the feature vector to a predicted value of the motion direction through a fully connected layer.

9. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 8, characterized in that: According to the size of the speed sequence, the data window for direction prediction is dynamically adjusted. The specific method is as follows: When the speed exceeds the set threshold, some data points in the window are discarded to reduce the weight of historical data. The adaptive adjustment function is: in, m is the number of data points in the adjusted window, v low and v high are the lower and upper limits of the speed respectively; the direction prediction is optimized by the cosine similarity loss function, and its formula is: in, is the angle between the predicted direction and the true direction, is the magnitude of the predicted speed; Combine the optimized direction prediction and speed prediction results to generate a complete motion trajectory.

10. The indoor pedestrian positioning method based on ResNet and Transformer network according to claim 1, characterized in that: The step S4 is specifically as follows: S41, aligning the predicted trajectory generated by the speed sequence and the direction sequence with the real trajectory recorded using Tango technology to ensure consistency in the time dimension; S42, using an integral algorithm to calculate the position information of each trajectory point in the predicted trajectory, and based on the evaluation index of the trajectory error, respectively calculating the absolute trajectory error between the predicted trajectory and the actual trajectory ATE and relative trajectory error RTE , where the absolute trajectory error ATE The calculation formula is: in, is the i-th predicted position, is the corresponding real position, N is the total number of trajectory points; relative trajectory error RTE The calculation formula is: Among them, k is the length of the evaluation time window, and are the i+jth predicted position and the true position respectively; By using the absolute trajectory error ATE and relative trajectory error RTE Conduct statistical analysis to evaluate the positioning accuracy of the model at different speeds and in different environments.

Citation Information

Cited By

  • IMU (Inertial Measurement Unit) online self-adaptive calibration method, calibration device and automatic driving system

    CN121657442A