Real-time attitude estimation and tracking system and method for tourists in AR (Augmented Reality) scenic spot
By designing a real-time pose estimation and tracking system on mobile devices, combining inertial measurement data and deep learning models, dynamically adjusting the estimation strategy, the problems of insufficient real-time and strong hardware dependence in mobile devices are solved, and efficient and accurate pose estimation and tracking are achieved.
Patent Information
- Application Number
- CN202510035426.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The real-time application of existing pose estimation and tracking technologies on mobile devices has limited performance, strong hardware dependency, accuracy problems in dynamic scenarios and complex environments, as well as multi-object overlap and complex background processing limitations.
A real-time pose estimation and tracking system for tourists visiting scenic spots was designed. Inertial measurement data and frame images were obtained through the I/O data management module, and combined with the preliminary estimation module, the tracking and detection module, the equipment detection module, the internal estimation module and the external estimation module, the pose estimation strategy was dynamically adjusted, and the parallel processing capabilities of the deep learning model and the GPU were used to achieve efficient pose estimation.
Improves real-time and accuracy of pose estimation, enhances compatibility and tolerance of mobile devices, can operate effectively on low-end devices, and maintains stability and efficiency in dynamic scenarios and complex environments.
Smart Images

Figure CN119942645A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of posture estimation and tracking, and in particular to a real-time posture estimation and tracking system and method for tourists visiting scenic spots using AR. Background Art
[0002] In the fields of artificial intelligence, computer vision, and computer graphics, object pose estimation and tracking technology is one of the key technologies for realizing augmented reality (AR) applications. Existing pose estimation and tracking methods mainly rely on images or sensor data, and use algorithms to calculate the pose information of objects in three-dimensional space, such as position, rotation angle, etc. These technologies play an important role in the alignment of virtual objects with real scenes, the naturalness of user interaction, and the real-time rendering of dynamic scenes. For example, technologies such as AMD's TressFX and NVIDIA's HairWorks provide virtual objects with realistic appearance and dynamic effects through complex physical simulation and graphics rendering technology. In addition, basic models such as residual networks (ResNet) and U-Net have also achieved remarkable results in image processing and feature extraction, providing strong support for the development of pose estimation and tracking technology.
[0003] Although existing posture estimation and tracking technologies have made progress in some aspects, real-time applications on mobile devices still face many challenges. First, the performance of mobile devices is often limited, resulting in poor real-time performance of existing posture estimation and tracking algorithms. Second, existing methods usually need to rely on high-performance hardware, such as GPUs, which limits their application on low-end mobile devices. In addition, existing technologies are prone to tracking loss or inaccurate posture estimation when dealing with dynamic scenes and complex environments. For example, in dynamic environments such as scenic spot tours, tourists move and change their perspectives frequently. Existing posture estimation and tracking technologies are difficult to accurately update the rendering of virtual scenes in real time, affecting the user's immersive experience. Finally, existing technologies also have certain limitations when dealing with multiple overlapping objects and complex backgrounds. They cannot effectively merge and distinguish different virtual objects, reducing the practicality and user experience of AR applications. Summary of the invention
[0004] In order to solve the above technical problems, the present invention proposes a real-time posture estimation and tracking system and method for tourists visiting scenic spots in AR, so as to solve the problems existing in the above-mentioned prior art.
[0005] To achieve the above objectives, in a first aspect, the present invention provides a real-time posture estimation and tracking system for tourists visiting scenic spots in AR, comprising:
[0006] I / O data management module, used to obtain inertial measurement data and frame images from mobile devices of tourists visiting scenic spots through AR;
[0007] A preliminary estimation module, used to fine-tune the posture of the previous frame through the inertial measurement data to obtain updated posture data;
[0008] A tracking detection module, used to detect whether the mobile device is in violent motion or whether the tracking object is lost based on the updated posture data, and determine whether to trigger the external estimation module according to the first detection result obtained;
[0009] A device detection module, used for detecting whether a GPU exists and whether a device is overheated based on the updated posture data, and determining whether to enable an internal estimation module according to a second detection result obtained;
[0010] An internal estimation module, configured to use a deep learning model to perform a first posture estimation on the frame image to obtain a first posture estimation result;
[0011] An external estimation module, used for performing a second posture estimation on a high-performance server to obtain a second posture estimation result;
[0012] Among them, the first posture estimation result and the second posture estimation result are updated and output to the display terminal of the mobile device through the posture state management unit of the I / O data management module.
[0013] Preferably, the I / O data management module comprises:
[0014] An inertial measurement unit, for measuring inertial data and transmitting the inertial data to the preliminary estimation module;
[0015] Built-in camera for outputting sequence frame images;
[0016] The posture state management unit is used to store the posture states of all objects and continuously update them.
[0017] Preferably, in the inertial measurement unit, inertial data is obtained by measuring a built-in gyroscope, and the inertial data includes velocity, angular velocity, and acceleration.
[0018] Preferably, the posture fine-tuning formula in the preliminary estimation module is:
[0019]
[0020] Among them, P t is the camera pose at time t, P t+δ is the camera posture at time t+δ, is the camera transformation matrix;
[0021]
[0022] Among them, Ro t+δrepresents the rotation from time t to time t+δ, Tr t+δ It represents the translation from time t to time t+δ, where δ represents a small time interval.
[0023] Preferably, the tracking and detection module comprises:
[0024] An object loss unit, used to calculate the projection area of the bounding box of the object on the screen and set a first threshold. When the projection area is smaller than the first threshold, it is considered that the tracked object is lost;
[0025] The device motion unit is used to calculate the average offset between the bounding box vertices and the vertices of the previous frame and set a second threshold. If the average offset is greater than the second threshold, it is considered that the device has undergone a large-scale movement.
[0026] Preferably, the deep learning model network in the internal estimation module includes: a plurality of inverted residual blocks, a thermal head, and a displacement head, wherein the inverted residual block includes dimensionality increase, depthwise separable convolution, and dimensionality reduction steps;
[0027] The deep learning model outputs a single-channel thermal map and a displacement map.
[0028] Preferably, the mobile device display terminal is a tablet computer, mobile phone, Switch, VR headset, AR glasses or other mobile devices.
[0029] In a second aspect, the present invention further provides a method for real-time posture estimation and tracking of tourists visiting a scenic spot in AR, characterized in that the method is used to implement the real-time posture estimation and tracking system for tourists visiting a scenic spot in AR, and comprises the following steps:
[0030] Obtain inertial measurement data and frame images through the I / O data management module;
[0031] Through the preliminary estimation module, the posture of the previous frame is fine-tuned to obtain the updated posture data;
[0032] Detecting, by means of a tracking detection module, whether the mobile device is in violent motion or whether the tracking object is lost, and determining whether to trigger an external estimation module according to the first detection result obtained;
[0033] Detecting, by means of a device detection module, whether a GPU exists and whether a device is overheated, and determining whether to enable an internal estimation module according to a second detection result obtained;
[0034] Using a deep learning model to perform a first posture estimation on the frame image through an internal estimation module to obtain a first posture estimation result;
[0035] Performing a second posture estimation on a high-performance server through an external estimation module to obtain a second posture estimation result;
[0036] Among them, the first posture estimation result and the second posture estimation result are updated and output to the display terminal of the mobile device through the posture state management unit of the I / O data management module.
[0037] In a third aspect, the present invention further discloses a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method described in the second aspect when executed by a processor.
[0038] In a fourth aspect, the present invention further discloses a computer program product, comprising a computer program, which implements the steps of the method described in the second aspect when executed by a processor.
[0039] Compared with the prior art, the present invention has the following advantages and technical effects:
[0040] The present invention provides a real-time posture estimation and tracking system for tourists visiting a scenic spot in an AR environment, comprising: an I / O data management module, used for obtaining inertial measurement data and frame images through a mobile device of the tourists visiting the scenic spot in an AR environment; a preliminary estimation module, used for fine-tuning the posture of a previous frame through the inertial measurement data to obtain updated posture data; a tracking detection module, used for detecting whether a mobile device is in violent motion or a tracking object is lost based on the updated posture data, and determining whether to trigger an external estimation module according to a first detection result obtained; a device detection module, used for detecting whether a GPU exists and whether a device is overheated based on the updated posture data, and determining whether to enable an internal estimation module according to a second detection result obtained; an internal estimation module, used for performing a first posture estimation on the frame image using a deep learning model to obtain a first posture estimation result; an external estimation module, used for performing a second posture estimation on a high-performance server to obtain a second posture estimation result; wherein the first posture estimation result and the second posture estimation result are updated and output to a display end of a mobile device through a posture state management unit of the I / O data management module.
[0041] The technology of the present invention is lightweight and can be used for real-time estimation and tracking on mobile phones. It has strong compatibility and tolerance for mobile devices, and can run on mobile devices of any performance, and can run in environments without "inertial measurement units", without GPUs, or without Internet access. The system is modular in design and can be replaced by the latest models.
[0042] The present invention combines the high-frequency data of the inertial measurement unit and the posture information of the previous frame, and utilizes the deep learning model and the parallel processing capability of the GPU to achieve efficient posture estimation on mobile devices, thereby improving the real-time performance of posture estimation and solving the problem of insufficient real-time performance of the prior art on mobile devices.
[0043] The present invention takes into account the hardware configuration of the mobile device, such as the presence or absence of a GPU, and the temperature condition of the device, and can dynamically adjust the posture estimation strategy according to the actual performance of the device, thereby improving the compatibility of the mobile device and solving the problem of limited application of the prior art on low-end mobile devices.
[0044] The "tracking detection" module is used to promptly detect and handle abnormal situations in the tracking process, such as violent movement of the device or loss of the object, and when the device performance is limited or abnormalities occur, posture estimation is performed using a high-performance server. The present invention can ensure the accuracy and stability of posture estimation, and solves the accuracy problem of the prior art in dynamic scenes and complex environments.
[0045] The present invention can effectively manage and update the posture states of multiple objects through the "posture state management" module, solves the limitations of the prior art in dealing with multiple object overlaps and complex backgrounds, and improves the practicality and user experience of AR applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0047] Figure 1 is a system schematic diagram of an embodiment of the present invention;
[0048] Figure 2 A network diagram of a deep learning model according to an embodiment of the present invention;
[0049] Figure 3 4 is an internal network diagram of an inverted residual block according to an embodiment of the present invention. DETAILED DESCRIPTION
[0050] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0051] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0052] Embodiment 1
[0053] like Figure 1 As shown, this embodiment provides a real-time posture estimation and tracking system for tourists visiting scenic spots in AR, including:
[0054] I / O data management module, used to obtain inertial measurement data and frame images from mobile devices of tourists visiting scenic spots through AR;
[0055] A preliminary estimation module, used to fine-tune the posture of the previous frame through the inertial measurement data to obtain updated posture data;
[0056] A tracking detection module, used to detect whether the mobile device is in violent motion or whether the tracking object is lost based on the updated posture data, and determine whether to trigger the external estimation module according to the first detection result obtained;
[0057] A device detection module, used for detecting whether a GPU exists and whether a device is overheated based on the updated posture data, and determining whether to enable an internal estimation module according to a second detection result obtained;
[0058] An internal estimation module, configured to use a deep learning model to perform a first posture estimation on the frame image to obtain a first posture estimation result;
[0059] An external estimation module, used for performing a second posture estimation on a high-performance server to obtain a second posture estimation result;
[0060] Among them, the first posture estimation result and the second posture estimation result are updated and output to the display terminal of the mobile device through the posture state management unit of the I / O data management module.
[0061] In this embodiment, the system specifically includes the following:
[0062] Mobile device display: tablets, mobile phones, Switch, VR headsets, AR glasses, and other mobile devices. Different brands or models of the same device can also lead to huge differences in performance. These devices must have cameras (otherwise they cannot acquire images for AR processing), but do not need to have an "inertial measurement unit" or a GPU.
[0063] I / O data management module: responsible for managing the input and output data of the system, which includes three units: inertial measurement unit, built-in camera or system screenshot, and attitude state management.
[0064] Furthermore, the inertial measurement unit measures inertial data (including speed, angular velocity, acceleration, etc.) through the device's built-in gyroscope and passes the data to the "preliminary estimation" module. The sampling rate of the "inertial measurement unit" of various devices is different. If calculated at 200Hz, each sampling takes 5ms.
[0065] Furthermore, the built-in camera or system screenshot unit will continuously output sequence frame images, where the "built-in camera" refers to the device's built-in camera (or camera), which can be processed through the operation of the camera's return image. "System screenshot" means that for devices without a camera, this frame of the picture can be saved (screenshot saving) when outputting the rendered image. Some devices have this system, and you can also develop a system software with this function yourself. The speed, resolution, and compression ratio of various devices are different. If calculated at 120FPS, the resolution of each frame is 720x480, using JPEG lossy compression, and the compression ratio is 10:1.
[0066] Furthermore, the posture state management unit is responsible for storing the posture states of all objects, continuously updating them, and continuously outputting the postures to the "mobile device display terminal".
[0067] The preliminary estimation module has two inputs, one is the posture of the previous frame (or initial frame), and the other is the inertial data. It is responsible for fine-tuning the posture of the previous frame through the inertial data to obtain the updated posture data (that is, the correct posture in the above figure). However, there are three prerequisites for doing so. One is that the mobile device cannot move violently, one is that the mobile device must have an "inertial measurement unit", and one is that there must be status data of the previous frame. These three conditions are indispensable.
[0068] As an innovative implementation method, the working principle of the initial estimation module is as follows:
[0069] Posture fine-tuning formula: Among them, P t is the camera posture at time t (i.e., the posture transmitted from “Posture State Management”), P t+δ is the camera pose at time t+δ (i.e., the pose to be estimated in the current frame), is the camera transformation matrix, where Ro t+δ represents the rotation from time t to time t+δ, Tr t+δ It represents the translation from t to t+δ, where δ represents a small time interval. It can be calculated by the "inertial measurement unit" on the mobile phone. Because Ro t+δ It can be calculated by the angular velocity rotation change of the mobile phone gyroscope, and Tr t+δ The acceleration can be obtained by calculating the movement offset of the acceleration through the accelerometer.
[0070] The above attitude fine-tuning formula will be affected by factors such as noise, deviation, and inaccurate initial velocity estimation of the system, resulting in inaccurate attitude. In order to eliminate this part of the error, we designed an attitude correction method, including angular velocity correction and velocity correction:
[0071] Angular velocity correction formula:
[0072]
[0073] in, is the estimated rotation bias, is the rotation estimated by the inertial measurement unit, is the estimated true rotation, is the estimated noise rotation (ignored for ease of calculation), is the estimated angular velocity bias, is the estimated Euler angle deviation, and Δt is the time from the last time the backend estimation was triggered to now.
[0074] Speed correction formula:
[0075]
[0076] in, is the estimated velocity deviation, is the estimated acceleration bias, is the velocity estimated by the inertial measurement unit, is the average speed of the backend attitude estimation, Δt is the time interval between two consecutive frame attitudes, is the time interval estimated by the backend.
[0077] Through the above-mentioned attitude correction method, the deviation generated by the inertial measurement unit (i.e. or ), regularly compensate for speed deviation (i.e. or ).
[0078] The tracking detection module is used to detect whether the mobile device is moving violently, whether the tracking object is lost, etc. This module pursues speed rather than accuracy. It can detect quickly and return the result (correct or wrong). If the return is "correct", the posture output by the "preliminary estimation" is stored in the "posture state management" as the "correct posture"; if it returns "wrong", the "device detection" is triggered, and the posture output by the "preliminary estimation" is regarded as the "wrong posture" and discarded. Note that this module is synchronized with the "preliminary estimation" module (more precisely, if all modules are in the idle state, all modules are synchronized).
[0079] As an innovative implementation method, the working principle of the tracking detection module is as follows:
[0080] When an object is beyond the frustum of the mobile device camera, or is too small on the screen, the tracking object is considered lost, and tracking will no longer be performed, and "error" will be returned. Or when the "inertial measurement unit" detects that the device is moving too fast, "error" will also be returned.
[0081] Furthermore, the detection algorithm calculates the projection area of the object's bounding box on the screen and designs a threshold (such as 4 pixels). If it is less than this threshold, the tracking object is considered lost and an "error" is returned. In addition, the "inertial measurement unit" is used to calculate the average offset between the bounding box vertices and the vertices of the previous frame. If it is greater than the set threshold (such as 15), it is considered that the device has undergone a large range of movement and an "error" is also returned.
[0082] The device detection module mainly performs two tests, one is to detect whether the GPU exists, and the other is to detect whether the device is overheated. Many low-end devices do not have GPUs, so it is necessary to detect whether the GPU exists. This is only detected once during initialization. If the GPU exists, the "internal estimation" module is activated, otherwise it is invalid (because the "internal estimation" module must use the GPU for estimation). To detect whether the device is overheated, the "device detection" module is triggered every time it is detected, and the temperature returned by the device's temperature sensor is over 45 degrees, then the device is considered overheated. When the device is overheated, the use of the "internal estimation" module should be suspended and turned on again when the temperature drops to 37 degrees.
[0083] The internal estimation module is a method that uses a deep learning model and the internal GPU of the mobile phone to perform posture estimation. The input parameter of "internal estimation" is the frame image. The triggering conditions of "internal estimation" are: ① Initialization, and the GPU exists, without overheating, trigger "internal estimation", and store the generated posture in "posture state management", and use it as one of the input conditions for initializing the "preliminary estimation" module. ② In the non-initialized state, there is a GPU, there is no overheating, and the "tracking detection" module returns "error", then trigger "internal estimation", and store the generated posture in "posture state management".
[0084] As an innovative implementation method, the working principle of the internal estimation module is:
[0085] Deep learning model network diagram, such as Figure 2 As shown. The "input image" is the "frame image" mentioned above, with a resolution of 720x480 and RGB three-channel JPEG format. The network consists of 2 360x240x16 inverted residual blocks, 3 90x60x32 inverted residual blocks, 2 45x30x64 inverted residual blocks, 2 90x60x2 inverted residual blocks, and 1 thermal head and 1 displacement head, which output a single-channel thermal map with a resolution of 45x30 and a 16-channel displacement map with a resolution of 45x30. The internal network diagram of the inverted residual block is shown in Figure 3 shown.
[0086] All network modules inside the internal estimation module are composed of inverted residual blocks. The purpose is to be able to run on mobile devices. Compared with convolution blocks, the number of parameters can be greatly reduced, which is conducive to model lightweight. Each inverted residual block needs: ① Dimension increase: The inverted residual block first uses 1×1 convolution to increase the dimension of the input feature map, that is, to increase the number of channels. Although this step seems to increase the parameters, it is actually to provide more information and representation capabilities in the subsequent depthwise separable convolution, thereby improving the performance of the model at a relatively low computational cost; ② Depthwise separable convolution: On the feature map after dimensionality increase, the inverted residual block applies depthwise separable convolution for feature extraction. Depthwise separable convolution is an efficient convolution method that decomposes the standard convolution into depthwise convolution and pointwise convolution (1×1 convolution), thereby significantly reducing the number of parameters and the amount of computation; ③ Dimension reduction: Finally, the inverted residual block uses another 1×1 convolution to reduce the dimension of the feature map, that is, to reduce the number of channels to restore or adjust the dimension of the feature map. This step also helps to reduce the number of parameters.
[0087] The thermal head outputs a 45x30x1 thermal map:
[0088]
[0089] Where p is the pixel of the input image, is the set of all object instances in the image, μ i is the mass center position of object i, σ i is the kernel size proportional to the object size. This heat map conforms to a bivariate normal distribution. When there are multiple objects in the image, we select the maximum heat for each pixel. The L2 loss function (i.e., mean square error) is used. The first function of the heat map: if there are multiple overlapping objects, they are merged according to the heat map.
[0090] The displacement head is used to estimate the displacement field of the bounding box vertices (similar to the vertex offset), the displacement vector (also called the displacement field)
[0091]
[0092] Where p is the pixel of the input image, x i is one of the projections of the 8 vertices of the object bounding box on the image plane. The regression head outputs a 40×30×16 feature vector, where each bounding box vertex contributes two displacement channels (8 vertices means 16 channels). The loss function is Using L1 loss (i.e., mean absolute error), we average the difference between the predicted displacement and the actual displacement within the range of the heat map to minimize this average value (thus making Close to reality, this is the second function of the heat map), Y is 0.2 in the experiment.
[0093] The external estimation module is a high-performance server outside the mobile device, which is connected to the mobile device through the network. It does not limit the specific posture estimation method or model, but it should be a fast and high-precision posture estimation and tracking method. The input parameter of the "internal estimation" is also a frame image, and the frame image is transmitted through the network through the "tracking detection". The triggering conditions of the "external estimation" module are the same as those of the "internal estimation". In addition, when the device is overheated, the "internal estimation" module is disabled, and the posture can only be estimated through the "external estimation" module. Generally, the "external estimation" is more accurate than the "internal estimation", but the speed is slower. In addition to the model estimation time, there is also the time for network transmission back and forth. Therefore, the posture estimated by the "external estimation" will be updated to the "posture state management" as the final posture state.
[0094] In this embodiment, the operation process of the system is as follows:
[0095] S1. When the program starts, it is initialized first, and the presence of an "inertial measurement unit" (assuming it exists) and a GPU (assuming it exists) are detected. Frame images are sent to the "internal estimation" and "external estimation" respectively. Generally, the "internal estimation" estimates the posture first and stores it in the "posture state management". Then the "external estimation" estimates the posture and stores it in the "posture state management" (if the "external estimation" estimates it first and stores it in the "posture state management", the posture update of the "internal estimation" is no longer accepted at this time, because the "external estimation" is more accurate than the "internal estimation"). The "posture state management" will decide to send the stored posture to the "mobile device display end" and the "preliminary estimation" module (for the posture estimation of the next frame) within 5ms (the reason for 5ms is that the update speed of the "inertial measurement unit" is 5ms, and the state needs to be fine-tuned at this time). If within 5ms, the "posture state management" only receives the "correct posture" sent by the "internal estimation" but does not receive the "correct posture" of the "external estimation", it will no longer wait and send the posture to the "mobile device display" and "preliminary estimation" modules respectively. After the "correct posture" of the "external estimation" comes out, the "posture state management" can also update the stored posture and send the posture to the mobile device display (no longer passing it to the "preliminary estimation" module).
[0096] S2, "Preliminary estimation" preliminarily estimates a posture based on the posture of the previous frame and the inertial data from the "Inertial Measurement Unit", and sends this posture to the "Tracking Detection" module. At the same time, the "Built-in Camera or System Screenshot" sends the frame image to the "Tracking Detection" and "Internal Estimation" synchronously, and the two modules and "Device Detection" perform data processing at the same time (the "Internal Estimation" at this time is not a necessary estimation, and it will only be started when the "Internal Estimation" is in an idle state, otherwise it will not be started).
[0097] S3. When the "tracking detection" output is "correct", the posture transmitted by the "preliminary estimation" is stored in the "posture state management" as the "correct posture". If the "tracking detection" output is "error", two things are done. One is to request the "internal estimation" based on whether the "device detection" is overheated. If the "internal estimation" is executing the "non-essential estimation" of the previous frame at this time, the estimation stops immediately and executes the posture estimation of the current frame; if the "internal estimation" is executing the "non-essential estimation" of the current frame at this time, it continues to execute. The other is to send an estimation request (and send a frame image) to the "external estimation" through the network. When the posture estimation is completed, the "internal estimation" and "external estimation" will update the posture to the "posture state management" respectively (with the "external estimation" as the final state). Note that as long as the "Posture State Management" has the correct posture, it will send the posture to the "Mobile Device Display" and "Preliminary Estimation" modules. If the posture is updated later, the "Posture State Management" will only send the updated posture to the "Mobile Device Display" (this is because the display needs to update the latest posture information in real time, while the "Preliminary Estimation" module only needs the posture information once, and then fine-tunes it in combination with the inertial data without waiting for the final state).
[0098] It should also be noted that both "Internal Estimation" and "External Estimation" will perform estimation based on frame images when idle, but this is "non-essential estimation" (the system will mark it). If the "tracking detection" output is "error" at this time, the estimation request sent to "Internal Estimation" and "External Estimation" is "mandatory" (the system will mark it). Requests with the "mandatory" label must have a higher level than "non-essential estimation". At this time, the estimation task other than this time (or current frame) will be stopped and the current estimation task will be initiated.
[0099] This example provides special case handling:
[0100] (1) If the mobile device does not have an "inertial measurement unit" (although most mobile devices are currently equipped with this sensor), the "initial estimation" and "tracking detection" modules will fail, and the system will directly run the "device detection", "internal estimation" and "external estimation".
[0101] (2) If the mobile device does not have a GPU, the "Device Detection" and "Internal Estimation" modules will fail. The system will only perform "Preliminary Estimation" and will be highly dependent on "External Estimation". In this case, the network must be kept open.
[0102] (3) If you are in a confined space and the network signal is poor or blocked, the "external estimation" module will fail and must be run on a device with a GPU.
[0103] If the mobile device does not have an "Inertial Measurement Unit" and a GPU, the "Preliminary Estimation", "Tracking Detection", "Device Detection" and "Internal Estimation" modules will all fail, and the system can only track through a network connection "External Estimation".
[0104] (4) If the mobile device does not have an "Inertial Measurement Unit" and the network is blocked, the "Preliminary Estimation", "Tracking Detection", "Device Detection", and "External Estimation" modules will all fail, and the system can only track through "Internal Estimation".
[0105] (5) If the mobile device does not have a GPU and the network is blocked, the system cannot run (due to the lack of initialized state information, the "preliminary estimation" module cannot estimate the posture based on inertial data alone).
[0106] Embodiment 2
[0107] Based on the same inventive concept, this embodiment also provides a method for real-time posture estimation and tracking of tourists visiting a scenic spot in AR, which is used to implement the real-time posture estimation and tracking system for tourists visiting a scenic spot in AR, and the method includes:
[0108] Obtain inertial measurement data and frame images through the I / O data management module;
[0109] Through the preliminary estimation module, the posture of the previous frame is fine-tuned to obtain the updated posture data;
[0110] Detecting, by means of a tracking detection module, whether the mobile device is in violent motion or whether the tracking object is lost, and determining whether to trigger an external estimation module according to the first detection result obtained;
[0111] Detecting, by means of a device detection module, whether a GPU exists and whether a device is overheated, and determining whether to enable an internal estimation module according to a second detection result obtained;
[0112] Using a deep learning model to perform a first posture estimation on the frame image through an internal estimation module to obtain a first posture estimation result;
[0113] Performing a second posture estimation on a high-performance server through an external estimation module to obtain a second posture estimation result;
[0114] Among them, the first posture estimation result and the second posture estimation result are updated and output to the display terminal of the mobile device through the posture state management unit of the I / O data management module.
[0115] The real-time posture estimation and tracking method for AR tourists visiting scenic spots provided in this embodiment has all the advantages of the real-time posture estimation and tracking system for AR tourists visiting scenic spots provided in the first embodiment.
[0116] Embodiment 3
[0117] This embodiment further discloses a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first embodiment are implemented.
[0118] Embodiment 4
[0119] This embodiment also discloses a computer program product, including a computer program, which implements the steps of the method described in the first embodiment when executed by a processor.
[0120] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A real-time posture estimation and tracking system for tourists visiting scenic spots in AR, characterized in that: include: I / O data management module, used to obtain inertial measurement data and frame images from mobile devices of tourists visiting scenic spots through AR; A preliminary estimation module, used to fine-tune the posture of the previous frame through the inertial measurement data to obtain updated posture data; A tracking detection module, used to detect whether the mobile device is in violent motion or whether the tracking object is lost based on the updated posture data, and determine whether to trigger the external estimation module according to the first detection result obtained; A device detection module, used for detecting whether a GPU exists and whether a device is overheated based on the updated posture data, and determining whether to enable an internal estimation module according to a second detection result obtained; An internal estimation module, configured to use a deep learning model to perform a first posture estimation on the frame image to obtain a first posture estimation result; An external estimation module, used for performing a second posture estimation on a high-performance server to obtain a second posture estimation result; Among them, the first posture estimation result and the second posture estimation result are updated and output to the display terminal of the mobile device through the posture state management unit of the I / O data management module.
2. The system according to claim 1, characterized in that The I / O data management module includes: An inertial measurement unit, for measuring inertial data and transmitting the inertial data to the preliminary estimation module; Built-in camera for outputting sequence frame images; The posture state management unit is used to store the posture states of all objects and continuously update them.
3. The system according to claim 2, characterized in that In the inertial measurement unit, inertial data is measured by a built-in gyroscope, and the inertial data includes velocity, angular velocity, and acceleration.
4. The system according to claim 1, characterized in that The posture fine-tuning formula in the preliminary estimation module is: Among them, P t is the camera pose at time t, P t+δ is the camera posture at time t+δ, is the camera transformation matrix; Among them, Ro t+δ represents the rotation from time t to time t+δ, Tr t+δ It represents the translation from time t to time t+δ, where δ represents a small time interval.
5. The system according to claim 1, characterized in that The tracking and detection module comprises: An object loss unit, used to calculate the projection area of the bounding box of the object on the screen and set a first threshold. When the projection area is smaller than the first threshold, it is considered that the tracked object is lost; The device motion unit is used to calculate the average offset between the bounding box vertices and the vertices of the previous frame and set a second threshold. If the average offset is greater than the second threshold, it is considered that the device has undergone a large-scale movement.
6. The system according to claim 1, characterized in that The deep learning model network in the internal estimation module includes: a plurality of inverted residual blocks, a thermal head, and a displacement head, wherein the inverted residual block includes dimensionality increase, depthwise separable convolution, and dimensionality reduction steps; The deep learning model outputs a single-channel thermal map and a displacement map.
7. The system according to claim 1, characterized in that The mobile device display end is a tablet computer, mobile phone, Switch, VR headset, AR glasses or other mobile devices.
8. A method for real-time posture estimation and tracking of tourists visiting scenic spots in AR, characterized in that: The method for implementing the real-time posture estimation and tracking system for AR tourists visiting scenic spots as described in any one of claims 1 to 7 comprises the following steps: Obtain inertial measurement data and frame images through the I / O data management module; Through the preliminary estimation module, the posture of the previous frame is fine-tuned to obtain the updated posture data; Detecting, by means of a tracking detection module, whether the mobile device is in violent motion or whether the tracking object is lost, and determining whether to trigger an external estimation module according to the first detection result obtained; Detecting, by means of a device detection module, whether a GPU exists and whether a device is overheated, and determining whether to enable an internal estimation module according to a second detection result obtained; Using a deep learning model to perform a first posture estimation on the frame image through an internal estimation module to obtain a first posture estimation result; Performing a second posture estimation on a high-performance server through an external estimation module to obtain a second posture estimation result; Among them, the first posture estimation result and the second posture estimation result are updated and output to the display terminal of the mobile device through the posture state management unit of the I / O data management module.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to claim 8 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claim 8 are implemented.
Citation Information
Patent Citations
AR tracking method and device, AR equipment and storage medium
CN116703963A
Tracking augmented reality devices
CN118052960A
Edge device, storage medium, and method of controlling edge device
US20220180118A1
Electronic device for providing image for training of artificial intelligence model and operation method thereof
US20240242486A1
Visual-inertial odometry implementation method and system
WO2019157925A1