AR 6DoF tracking and large space identification method based on WeChat applet
By combining AR 6DoF tracking technology, point cloud data processing and optimization alignment algorithm, the existing AR technology has been solved, and the high-precision docking and real-time update of virtual objects and real scenes has been achieved, improving the efficiency and user experience of AR applications.
Patent Information
- Application Number
- CN202510319506.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing AR technology based on WeChat mini-programs has problems such as poor compatibility, low accuracy and insufficient stability in 6DoF tracking and large-space recognition, which limits the wide deployment and efficient application of AR applications in large-space environments.
It adopts AR 6DoF tracking technology, point cloud data processing technology, global optimization alignment algorithm and improved ICP algorithm, combined with PointNet++ point cloud denoising technology, to achieve high-precision docking and real-time updates between virtual objects and real scenes.
It improves the application accuracy and stability of AR technology in large space environments, solves the error problem of docking between virtual objects and real scenes, and realizes efficient and accurate three-dimensional modeling and real-time synchronization of virtual objects.
Smart Images

Figure CN120236044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of spatial recognition, and particularly to an AR 6DoF tracking and large space recognition method based on a WeChat mini program. Background Art
[0002] With the rapid development of augmented reality (AR) technology, its applications have gradually penetrated into various platforms, especially on the mobile Internet platform. As a widely used development platform, the WeChat mini program has become the first choice for many developers to build AR applications. However, current AR applications based on the WeChat mini program face some technical bottlenecks, especially in 6 degrees of freedom (6DoF) tracking and large space recognition.
[0003] The AR function in the prior art usually relies on the XR-frame plugin of the WeChat mini program. Although this plugin supports certain devices, its compatibility and stability are poor, and it only supports some models. The XR-frame plugin has a limited range of model support and cannot be widely applied to various hardware devices, which restricts the application scope of the AR function. In addition, it is difficult to integrate the XR-frame plugin with the existing mature Unity AR development technology, resulting in heavy adaptation work during the development process. Developers need to invest a lot of time and effort to solve compatibility problems. Due to these technical limitations, developers need to have a higher professional technical threshold to implement high-quality AR functions.
[0004] Although the WeChat mini program can already support basic AR functions, there are still many deficiencies in achieving the accuracy and stability of AR recognition. The user experience is usually limited by system performance and device compatibility, especially in AR applications in large space scenarios. Existing 6DoF tracking technologies, such as the V1 plane interface, although they support a relatively wide range of models, lack accurate measurement of real-world distances. While the V2 plane interface provides a real positioning function for physical distances, its model support range is narrow, and its initialization speed is slow and power consumption is high. Such technical limitations restrict the wide deployment of AR technology in actual large space applications and affect its effectiveness in high-precision positioning and large space scenarios.
[0005] Existing AR technologies also face challenges in large space recognition. Although some technologies such as point cloud recognition technology can perform certain spatial mapping, most of these technologies rely on dedicated support from hardware devices. And in the WeChat mini program platform, how to accurately and real-time process point cloud data and effectively dock it with virtual objects is still a technical problem. The prior art lacks an effective solution to solve the problems of alignment of virtual and real space coordinate systems and large space recognition accuracy.
[0006] Under these technical limitations, existing AR applications have many bottlenecks in terms of tracking accuracy, spatial stability, and development efficiency, making it difficult to meet higher - requirement application scenarios, especially complex AR application scenarios that require precise, real - time tracking and large - space recognition.
[0007] Therefore, how to provide an AR 6DoF tracking and large - space recognition method based on WeChat mini - programs is an urgent problem for those skilled in the art. Summary of the Invention
[0008] An object of the present invention is to propose an AR 6DoF tracking and large - space recognition method based on WeChat mini - programs. The present invention makes full use of AR technology, point cloud data processing technology, global optimization alignment algorithm, and improved ICP algorithm, and details the implementation process of the precise docking of virtual objects and the real - world scene. Through this method, high - precision synchronization and stable docking of virtual objects and the large - space environment can be achieved, with the advantages of good device compatibility, high precision, strong stability, and real - time update.
[0009] An AR 6DoF tracking and large - space recognition method based on WeChat mini - programs according to an embodiment of the present invention includes the following steps:
[0010] S1. Use AR 6DoF tracking technology to achieve spatial docking between virtual objects and the real world, obtain the position and attitude data of the device in three - dimensional space, and construct a data set;
[0011] S2. Based on the data set, construct a virtual - real space coordinate system alignment standard, and initially align the coordinate system of the virtual object with the coordinate system of the real world;
[0012] S3. Through point cloud data acquisition, use PointNet++ for point cloud denoising and simplification processing, optimize and identify the objects and structural features in the large - space environment, and combine point cloud recognition technology to optimize the three - dimensional modeling of the environment to obtain a three - dimensional space feature model;
[0013] S4. Based on the obtained three - dimensional space feature model, combine the ICP algorithm with the global optimization alignment algorithm of the least - squares method to precisely dock the virtual object with the real - world scene, and obtain the alignment relationship between the optimized virtual object and the real - world scene;
[0014] S5. Update the alignment relationship between the virtual object and the real - world scene through dynamic sensor data to maintain the real - time stability of the virtual object in the entire large - space environment;
[0015] S6. Combine the updated alignment relationship to perform spatial re - positioning of the virtual object, and the virtual object continues to be synchronized with the real environment in the dynamic scene.
[0016] Optionally, S1 specifically includes:
[0017] S11. Collect the acceleration, angular velocity, and magnetic field information of the device in three-dimensional space through a sensor to obtain the original motion data of the device;
[0018] S12. Use the original motion data of the device and calculate the real-time position and attitude change of the device through the Kalman filtering algorithm;
[0019] S13. Based on the pose change of the device, calculate the position and rotation matrix of the device in three-dimensional space to obtain the pose data of the device.
[0020] Optionally, S2 specifically includes:
[0021] S21. Based on the pose data of the device, calculate the spatial position and rotation of the virtual object relative to the real world, preliminarily align the coordinate system of the virtual object with the coordinate system of the real world to obtain the initial position of the virtual object, and transfer it to the subsequent steps for further docking;
[0022] S23. Based on the real-time spatial position of the device, dynamically adjust the alignment relationship of the virtual object and continuously update the pose data of the virtual object.
[0023] Optionally, S3 specifically includes:
[0024] S31. The original point cloud data P cloud = {p1, p2, …, p n} collected from the device, where each point p i = [x i , y i , z i represents the three-dimensional coordinates of the i-th point in the point cloud data, and collect the three-dimensional point cloud data of the large space environment through a sensor;
[0025] S32. Preprocess and remove noise from the collected point cloud data, and use the K-nearest neighbor algorithm to calculate the neighborhood set N i of each point p i . The neighborhood point set is calculated by the following formula:
[0026] N i = {p j ∣ ‖p i - p j ‖2 < ∈};
[0027] Among them, p i and p j are points in the point cloud, and ‖p i - p j‖2 represents the Euclidean distance, ∈ is the radius threshold of the neighborhood, the distance between the control point and the neighborhood. Through the selection of neighborhood points, each point in the point cloud will be assigned a neighborhood set N i ;
[0028] S33. Use the multi-layer perceptron network of PointNet++ to extract features from the neighborhood data N of each point i ; For each point p in the point cloud i and its neighborhood point set N i , by calculating the local feature f i and applying weighted averaging, the local feature vector of this point is obtained. The local feature f of each point p i is calculated by the following formula: i f
[0029] = MLP({p i | p j ∈ N j}); i}
[0030] where f i is the local feature of point p i , N i is the neighborhood set of this point. This local feature extraction layer processes the neighborhood points through a multi-layer perceptron network to extract local geometric information;
[0031] S34. Use the global feature aggregation mechanism of PointNet++ to aggregate the local feature f i to obtain the global feature vector g, thereby representing the global geometric feature of the entire point cloud data. The specific global feature aggregation formula is:
[0032]
[0033] where g is the aggregated global feature, containing the global geometric information of the point cloud data;
[0034] S35. Calculate the denoising weight w of each point i , this weight measures the importance of each point in the denoising process. After calculating the weighted sum and concatenation of the feature f of each point i and the global feature g, use a fully connected layer for calculation to obtain the denoising weight of each point:
[0035] w i = σ(W w ·(f i ⊕ g) + b w );
[0036] where W w is the learned weight matrix, b wis the bias term, σ is the activation function, ⊕ represents the concatenation operation, and f i is the local feature, g is the global feature, and w i is the denoising weight of point p i . This step calculates the importance of each point in the denoising process and adjusts the contribution of each point through the learned weight;
[0037] S36. According to the calculated denoising weight w i , set the denoising threshold w thresh . Remove those points whose weight values are lower than this threshold, regard them as noise points, and the point cloud data P ′ after removing noise can be obtained through the following formula:
[0038] P ′ = {p i | w i ≥ w thresh , p i ∈ P cloud};
[0039] where P ′ is the point cloud data after denoising, w thresh is the denoising weight threshold, which determines which points will be removed. Points with weights lower than the threshold are considered noise and removed from the point cloud data;
[0040] S37. Through steps S31 to S36, obtain the point cloud data P ′ . The data has removed noise points and retained representative structural feature points in the large - scale spatial environment. The denoised and simplified point cloud data constructs a three - dimensional spatial feature model.
[0041] Optionally, the specific steps of S4 include:
[0042] S41. Based on the obtained three - dimensional spatial feature model, define an optimization error function ε optimized to measure the docking error between the virtual object and the real scene. The error function includes position error, rotation error, and scale error, and the error function is defined as follows:
[0043]
[0044] where P i and P ′ virtual,i respectively represent the positions of the i - th point in the real scene and the virtual object, R i and R ′ virtual,i respectively represent the rotation matrices of the i - th point in the real scene and the virtual object, S real and S ′ virtualThe scale matrices representing the real scene and the virtual object respectively, λ1 and λ2 are parameters that adjust the relative importance among the position error, rotation error, and scale error, ‖·‖2 is the Euclidean distance, which measures the spatial distance between points;
[0045] S42. Through the global optimization alignment algorithm of the least squares method, minimize the error function ε optimized , by calculating the partial derivatives of the error function with respect to the position P ′ virtual,i of the virtual object and the rotation matrix R ′ virtual,i to update them. The specific steps are as follows:
[0046] S421. Calculate the gradients of the error function with respect to the position and the rotation matrix. Calculate the gradients of the error function with respect to the position and the rotation matrix of the virtual object. The position gradient obtained by taking the derivative of the position P ′ virtual,i is:
[0047]
[0048] The rotation matrix gradient obtained by taking the derivative of the rotation matrix R ′ virtual,i is:
[0049]
[0050] S422. Through the gradient descent method, update the position P ′ virtual,i of the virtual object and the rotation matrix R ′ virtual,i according to the calculated gradients:
[0051] Position update:
[0052]
[0053] Rotation matrix update:
[0054]
[0055] where η is the learning rate, which controls the step size of each update;
[0056] S43. Based on the preliminary docking result provided by the global optimization algorithm, use the ICP algorithm to perform local docking optimization on the virtual object and the real scene. The specific steps are as follows:
[0057] S431. Calculate the error between each pair of points (P i , P ′ virtual,i ) between the virtual object and the real scene. The local error function ε ICPAs follows:
[0058]
[0059] Wherein, R ′ virtual and T virtual are respectively the rotation matrix and translation vector of the virtual object;
[0060] S432. Calculate the update amounts of the position ΔT ICP and the rotation matrix ΔR ICP by minimizing the local error function. The update formulas are as follows:
[0061] Position update amount:
[0062]
[0063] Rotation matrix update amount:
[0064]
[0065] Wherein, arg min represents the variable that minimizes a certain function;
[0066] S433. Update the position P ′ virtual and the rotation matrix R ′ virtual of the virtual object according to the minimization result of the local error function:
[0067] Position update:
[0068] P ′ virtual,optimized = P ′ virtual + ΔT ICP ;
[0069] Rotation matrix update:
[0070] R ′ virtual,optimized = R ′ virtual ·ΔR ICP ;
[0071] Wherein, P ′ virtual,optimized represents the position of the updated virtual object, and R ′ virtual,optimized represents the updated rotation matrix;
[0072] S434. Optimize the position and rotation matrix through multiple iterations until the error converges. The iteration stop condition is:
[0073] ‖ΔP′ virtual ‖2 < ∈1, ‖ΔR ′ virtual ‖2 < ∈2;
[0074] Wherein, ΔP ′ virtual is the correction amount of the position, and ΔR ′ virtual is the correction amount of the rotation matrix. ∈1 and ∈2 are the error tolerance thresholds, indicating when the optimization process stops;
[0075] S44. Weightedly combine the update results of the global optimization algorithm in S42 and the update results of the ICP algorithm in S43 to obtain a more optimized position and rotation matrix.
[0076] The beneficial effects of the present invention are as follows:
[0077] Through the AR 6DoF tracking and large space recognition method based on WeChat mini-program proposed by the present invention, the defects of the existing technology in the docking accuracy and stability between virtual objects and the real scene in a large space environment are successfully overcome. This method realizes the precise docking between virtual objects and the real scene by combining the self-developed improved ICP algorithm and the global optimization alignment algorithm of the least squares method. In traditional AR tracking technologies, due to poor hardware compatibility and low positioning accuracy, there are large errors in the docking between virtual objects and the real environment, while the present invention effectively improves the application accuracy of AR technology in a large space environment and solves the technical bottleneck in the existing technology.
[0078] The present invention not only greatly improves the applicability of AR technology in a large space environment, but also updates the docking relationship between virtual objects and the real scene through dynamic sensor data to ensure the real-time stability of virtual objects in the entire large space environment. In addition, by combining the PointNet++ point cloud denoising technology and the large space recognition algorithm, it can accurately identify and model the objects and structural features in the environment, thereby realizing efficient and accurate three-dimensional modeling optimization and further improving the docking accuracy between virtual objects and the real scene.
[0079] Through the method of the present invention, virtual objects and the real environment can be synchronized more precisely and stably, which not only improves the application performance of the system in a large space scene, but also provides users with a better augmented reality experience. Due to the good device compatibility and high accuracy of this method, developers can more efficiently realize the tracking and docking of virtual objects in actual applications, reducing technical problems and resource consumption in the development process. Generally speaking, the present invention provides an efficient, stable and highly accurate AR technology solution, which has strong practicability and broad application prospects. Description of the Drawings
[0080] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the accompanying drawings:
[0081] Figure 1 is a flowchart of a method for AR 6DoF tracking and large space recognition based on a WeChat mini-program proposed by the present invention;
[0082] Figure 2 is a schematic diagram of the combination of the improved ICP algorithm and the least squares global optimization alignment algorithm in the present invention. Detailed implementation manners
[0083] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0084] Referring to Figure 1 and Figure 2 , a method for AR 6DoF tracking and large space recognition based on a WeChat mini-program includes the following steps:
[0085] S1. Using AR 6DoF tracking technology, realize the spatial docking between virtual objects and the real world, obtain the position and attitude data of the device in the three-dimensional space, and construct a data set;
[0086] S2. Based on the data set, construct a standard for aligning the virtual and real space coordinate systems, and initially align the coordinate system of the virtual object with the coordinate system of the real world;
[0087] S3. Through point cloud data acquisition, use PointNet++ for point cloud denoising and simplification processing, optimize and identify the objects and structural features in the large space environment, and combine point cloud recognition technology to optimize the three-dimensional modeling of the environment to obtain a three-dimensional space feature model;
[0088] S4. Based on the obtained three-dimensional space feature model, combine the improved ICP algorithm with the least squares global optimization alignment algorithm to accurately dock the virtual object with the real scene, and obtain the alignment relationship between the optimized virtual object and the real scene;
[0089] S5. Update the alignment relationship between the virtual object and the real scene through dynamic sensor data to maintain the real-time stability of the virtual object in the entire large space environment;
[0090] S6. Combine the updated alignment relationship to perform spatial repositioning of the virtual object, and the virtual object continues to be synchronized with the real environment in the dynamic scene.
[0091] In this embodiment, the specific content of S1 includes:
[0092] S11. Collect the acceleration, angular velocity, and magnetic field information of the device in three-dimensional space through the sensor to obtain the original motion data of the device;
[0093] S12. Use the original motion data of the device and calculate the real-time position and attitude change of the device through the Kalman filtering algorithm;
[0094] S13. Based on the pose change of the device, calculate the position and rotation matrix of the device in three-dimensional space to obtain the pose data of the device.
[0095] In this embodiment, the specific steps of S2 include:
[0096] S21. Based on the pose data of the device, calculate the spatial position and rotation of the virtual object relative to the real world, preliminarily align the coordinate system of the virtual object with the coordinate system of the real world to obtain the initial position of the virtual object, and transfer it to the subsequent steps for further docking;
[0097] S23. Based on the real-time spatial position of the device, dynamically adjust the alignment relationship of the virtual object and continuously update the pose data of the virtual object.
[0098] In this embodiment, the specific steps of S3 include:
[0099] S31. The original point cloud data P cloud = {p1, p2,..., p n} collected from the device, where each point p i = [x i , y i , z i represents the three-dimensional coordinates of the i-th point in the point cloud data, and collect the three-dimensional point cloud data of the large space environment through the sensor;
[0100] S32. Preprocess and remove noise from the collected point cloud data, and use the K-nearest neighbor algorithm to calculate the neighborhood set N i of each point p i . The neighborhood point set is calculated by the following formula:
[0101] N i = {p j ∣ ‖p i - p j ‖2 < ∈};
[0102] where, p i and p j are points in the point cloud, and ‖p i - p j‖2 represents the Euclidean distance, ∈ is the radius threshold of the neighborhood, and controls the distance between the point and the neighborhood. By selecting the neighborhood points, each point in the point cloud will be assigned a neighborhood set N i ;
[0103] S33, use PointNet++ multi-layer perception network to collect the neighborhood data N of each point i Perform feature extraction; for each point p in the point cloud i and its neighborhood point set N i , by calculating the local feature f i And apply weighted average to obtain the local feature vector of the point. For each point p i The local features f i Calculated by the following formula:
[0104] f i =MLP({p j ∣p j ∈N i});
[0105] Among them, f i For point p i The local features of N i is the neighborhood set of the point, and the local feature extraction layer processes the neighborhood points through a multi-layer perception network to extract local geometric information;
[0106] S34, using the global feature aggregation mechanism of PointNet++, local features f i Aggregate to obtain the global feature vector g, which represents the global geometric features of the entire point cloud data. The specific global feature aggregation formula is:
[0107]
[0108] Among them, g is the aggregated global feature, which contains the global geometric information of the point cloud data;
[0109] S35. Calculate the denoising weight w of each point i , which measures the importance of each point in the denoising process and calculates the feature f of each point i After weighting and concatenating with the global feature g, the fully connected layer is used to calculate and obtain the denoising weight of each point:
[0110] w i =σ(W w ·(f i ⊕g)+b w );
[0111] Among them, W w is the learned weight matrix, b wis the bias term, σ is the activation function, ⊕ represents the concatenation operation, and f i is the local feature, g is the global feature, and w i is the denoising weight of point p i . This step calculates the importance of each point in the denoising process and adjusts the contribution of each point through the learned weight;
[0112] S36. According to the calculated denoising weight w i , set the denoising threshold w thresh . Remove the points whose weight values are lower than this threshold and regard them as noise points. The point cloud data P ′ after removing noise can be obtained through the following formula:
[0113] P ′ = {p i | w i ≥ w thresh , p i ∈ P cloud};
[0114] where P ′ is the point cloud data after denoising, and w thresh is the denoising weight threshold, which determines which points will be removed. The points with weights lower than the threshold are considered noise and removed from the point cloud data;
[0115] S37. Through steps S31 to S36, obtain the point cloud data P ′ . The data has removed the noise points and retained the representative structural feature points in the large - scale space environment. The denoised and simplified point cloud data constructs a three - dimensional space feature model.
[0116] In this embodiment, the specific steps of S4 include:
[0117] S41. Based on the obtained three - dimensional space feature model, define an optimization error function ε optimized to measure the docking error between the virtual object and the real scene. The error function includes position error, rotation error, and scale error, and the error function is defined as follows:
[0118]
[0119] where P i and P ′ virtual,i respectively represent the positions of the i - th point in the real scene and the virtual object, R i and R ′ virtual,i respectively represent the rotation matrices of the i - th point in the real scene and the virtual object, S real and S ′ virtualThe scale matrices representing the real scene and the virtual object respectively, λ1 and λ2 are parameters that regulate the relative importance among the position error, rotation error, and scale error, ‖·‖2 is the Euclidean distance, which measures the spatial distance between points;
[0120] S42. Through the global optimization alignment algorithm of the least squares method, minimize the error function ε optimized , by calculating the partial derivatives of the error function with respect to the position P ′ virtual,i of the virtual object and the rotation matrix R ′ virtual,i to update them. The specific steps are as follows:
[0121] S421. Calculate the gradients of the error function with respect to the position and rotation matrix, calculate the gradients of the error function with respect to the position and rotation matrix of the virtual object, and the position gradient obtained by taking the derivative of the position P ′ virtual,i is:
[0122]
[0123] The rotation matrix gradient obtained by taking the derivative of the rotation matrix R ′ virtual,i is:
[0124]
[0125] S422. Through the gradient descent method, update the position P ′ virtual,i of the virtual object and the rotation matrix R ′ virtual,i according to the calculated gradients:
[0126] Position update:
[0127]
[0128] Rotation matrix update:
[0129]
[0130] where η is the learning rate, which controls the step size of each update;
[0131] S43. Based on the preliminary docking result provided by the global optimization algorithm, use the ICP algorithm to perform local docking optimization on the virtual object and the real scene. The specific steps are as follows:
[0132] S431. Calculate the error between each pair of points (P i , P ′ virtual,i ) between the virtual object and the real scene. The local error function ε ICPAs follows:
[0133]
[0134] Wherein, R ′ virtual and T virtual are respectively the rotation matrix and translation vector of the virtual object;
[0135] S432. Calculate the update amounts of the position ΔT ICP and the rotation matrix ΔR ICP by minimizing the local error function. The update formulas are as follows:
[0136] Position update amount:
[0137]
[0138] Rotation matrix update amount:
[0139]
[0140] Wherein, arg min represents the variable that minimizes a certain function;
[0141] S433. Update the position P ′ virtual and the rotation matrix R ′ virtual of the virtual object according to the minimization result of the local error function:
[0142] Position update:
[0143] P ′ virtual,optimized = P ′ virtual + ΔT ICP ;
[0144] Rotation matrix update:
[0145] R ′ virtual,optimized = R ′ virtual ·ΔR ICP ;
[0146] Wherein, P ′ virtual,optimized represents the position of the updated virtual object, and R ′ virtual,optimized represents the updated rotation matrix;
[0147] S434. Optimize the position and rotation matrix through multiple iterations until the error converges. The iteration stop condition is:
[0148] ‖ΔP′ virtual ‖2 < ∈1, ‖ΔR ′ virtual ‖2 < ∈2;
[0149] where ΔP ′ virtual is the correction amount of the position, and ΔR ′ virtual is the correction amount of the rotation matrix. ∈1 and ∈2 are error tolerance thresholds, indicating when the optimization process stops;
[0150] S44. Weightedly combine the update results of the global optimization algorithm in S42 and the update results of the ICP algorithm in S43 to obtain a more optimized position and rotation matrix.
[0151] Example 1:
[0152] To verify the feasibility of the present invention in implementation, the present invention is applied to the warehouse management system of a large industrial park, and the AR 6DoF tracking and large space recognition method based on WeChat mini-program is used to improve the accuracy and efficiency of the storage location of items in the warehouse. The area of the warehouse site is about 5000 square meters, and there are more than 1000 different items in the warehouse that need to be tracked, located and managed. The number of items is large and they are scattered. The traditional manual tracking method cannot meet the requirements of modern warehouse management due to insufficient accuracy and low efficiency.
[0153] In the warehouse of the industrial park, the item storage area is divided into multiple shelves, and each shelf has different types and quantities of items. Most of the shelves are stacked in different positions of the warehouse, and the light is poor in some areas, and the item labels may also be blocked. The traditional warehouse management system relies on manual barcode scanning and the use of handheld devices for positioning, with low work efficiency and easy to make mistakes, especially in the area with high-density item stacking.
[0154] To optimize the storage and retrieval process of items, the warehouse management system introduces AR 6DoF tracking technology and large space recognition method. Through this method, warehouse staff can view the location information of items through a smartphone or AR glasses, and the AR interface will display the docking situation of virtual items and real items in real time, helping the staff to find items more quickly and accurately.
[0155] The application scenarios of the present invention mainly focus on the real-time tracking and positioning of warehouse items. Warehouse staff wear AR glasses or use intelligent devices supporting WeChat mini-programs. The devices obtain the real-time position and attitude data of the devices in the warehouse through AR 6DoF tracking technology and conduct a preliminary docking with the real-world coordinate system. With the device position and attitude data collected by sensors, the virtual interface on the AR glasses can display the position of the items in real time and mark the spatial docking situation with virtual objects. The staff can quickly find the position of the target item just by following the guidance of the virtual object.
[0156] In actual operation, through point cloud data acquisition technology, the present invention can obtain the three-dimensional model of the large-space environment in the warehouse and accurately model the spatial position of the items. Through point cloud denoising and simplification processing, the system can update the spatial features of the item positions in real time and optimize the docking with virtual objects, making the item positioning in the warehouse more accurate. Whenever the staff finds the target item, the system will dynamically update the docking relationship between the virtual object and the real scene according to the real-time sensor data to ensure the accuracy and timeliness of the information.
[0157] In addition, through the combination of the global optimization alignment algorithm and the improved ICP algorithm in the present invention, the docking relationship between the virtual object and the real scene in the warehouse has been accurately optimized. The position and rotation of the virtual object are adjusted in real time to ensure that the staff can quickly and accurately find the items at any time and in any environment, improving the efficiency of warehouse management.
[0158] To verify the beneficial effects of the present invention, we compared the work efficiency and accuracy before and after introducing the technology of the present invention during the implementation process. The test scenario was a typical item storage area in the warehouse, with an area of about 500 square meters and 200 different items stored.
[0159] Table 1: Comparison Table of Warehouse Item Positioning Efficiency
[0160]
[0161]
[0162] According to the above data, we can clearly see that after introducing the technology of the present invention, the efficiency of warehouse management has been significantly improved. First, the item positioning time has been shortened from 15 minutes to 5 minutes, the accuracy has been improved from the original 1.5 meters to 0.3 meters, and the positioning error rate has been greatly reduced. Second, due to the increase in the real-time data update frequency, the staff can obtain the latest position information of the items more timely, further optimizing the warehouse management process. In addition, the system response time has also been shortened from 10 seconds to 2 seconds, improving the user experience.
[0163] Through these data, we can clearly prove the effectiveness of the present invention in large space recognition, item tracking and positioning. Combining point cloud denoising, 3D modeling, global optimization alignment and ICP algorithm, the present invention greatly improves the accuracy and efficiency of warehouse management, meeting the requirements of high precision and real-time performance of modern warehouse management systems.
[0164] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, should be covered by the protection scope of the present invention.
Claims
1. An AR 6DoF tracking and large space recognition method based on WeChat applet, characterized in that: The steps include: S1. Use AR 6DoF tracking technology to achieve spatial docking between virtual objects and the real world, obtain the position and posture data of the device in three-dimensional space, and build a data set; S2. Based on the data set, a virtual-real space coordinate system alignment standard is constructed to preliminarily align the coordinate system of the virtual object with the coordinate system of the real world; S3. Through point cloud data collection, PointNet++ is used to perform point cloud denoising and simplification processing, optimize the identification of objects and structural features in large space environments, and combine point cloud recognition technology to optimize the three-dimensional modeling of the environment to obtain a three-dimensional spatial feature model; S4. Based on the obtained three-dimensional spatial feature model, the ICP algorithm is combined with the global optimization alignment algorithm of the least squares method to accurately align the virtual object with the real scene, and obtain the optimized alignment relationship between the virtual object and the real scene; S5. Update the alignment relationship between the virtual object and the real scene through dynamic sensor data to maintain the real-time stability of the virtual object in the entire large space environment; S6. Based on the updated alignment relationship, the virtual object is spatially repositioned, and the virtual object is continuously synchronized with the real environment in the dynamic scene.
2. According to claim 1, the AR 6DoF tracking and large space recognition method based on WeChat applet is characterized in that: The S1 specifically includes: S11, collecting acceleration, angular velocity and magnetic field information of the device in three-dimensional space through sensors to obtain original motion data of the device; S12, using the original motion data of the device and the Kalman filter algorithm to calculate the real-time position and posture change of the device; S13. Based on the posture change of the device, calculate the position and rotation matrix of the device in the three-dimensional space to obtain the posture data of the device.
3. According to claim 1, the AR 6DoF tracking and large space recognition method based on WeChat applet is characterized in that: The S2 specifically includes: S21. Based on the posture data of the device, calculate the spatial position and rotation of the virtual object relative to the real world, preliminarily align the coordinate system of the virtual object with the coordinate system of the real world, obtain the initial position of the virtual object, and pass it to the subsequent steps for further docking; S23. Based on the real-time spatial position of the device, dynamically adjust the alignment relationship of the virtual object and continuously update the position and posture data of the virtual object.
4. The AR 6DoF tracking and large space recognition method based on WeChat applet of claim 1, characterized in that: The S3 specifically includes: S31, original point cloud data P collected from the device cloud ={p1,p2,…,p n }, where each point p i =[x i ,y i ,z i ] represents the three-dimensional coordinates of the i-th point in the point cloud data, and the three-dimensional point cloud data of the large space environment is collected by the sensor; S32, pre-processing and noise removal of the collected point cloud data, using the K nearest neighbor algorithm to calculate the p of each point i The neighborhood set N i , the neighborhood point set is calculated by the following formula: N i ={p j ∣‖p i -p j ‖2<∈}; Among them, p i and p j is a point in the point cloud, ‖p i -p j ‖2 represents the Euclidean distance, ∈ is the radius threshold of the neighborhood, and controls the distance between the point and the neighborhood. By selecting the neighborhood points, each point in the point cloud will be assigned a neighborhood set N i ; S33, use PointNet++ multi-layer perception network to collect the neighborhood data N of each point i Perform feature extraction; for each point p in the point cloud i and its neighborhood point set N i , by calculating the local feature f i And apply weighted average to obtain the local feature vector of the point. For each point p i The local features f i Calculated by the following formula: f i =MLP({p j ∣p j ∈N i }); Among them, f i For point p i The local features of N i is the neighborhood set of the point, and the local feature extraction layer processes the neighborhood points through a multi-layer perception network to extract local geometric information; S34, using the global feature aggregation mechanism of PointNet++, local features f i Aggregate to obtain the global feature vector g, which represents the global geometric features of the entire point cloud data. The specific global feature aggregation formula is: Among them, g is the aggregated global feature, which contains the global geometric information of the point cloud data; S35. Calculate the denoising weight w of each point i , which measures the importance of each point in the denoising process and calculates the feature f of each point i After weighting and concatenating with the global feature g, the fully connected layer is used to calculate and obtain the denoising weight of each point: Among them, W w is the learned weight matrix, b w is the bias term, σ is the activation function, represents the splicing operation, f i is a local feature, g is a global feature, and w i For point p i Denoising weights: This step calculates the importance of each point in the denoising process and adjusts the contribution of each point through the learned weights. S36, according to the calculated denoising weight w i , set the denoising threshold w thresh , remove the points whose weight values are lower than the threshold and regard them as noise points. The point cloud data P′ after noise removal can be obtained by the following formula: P′={p i ∣w i ≥w thresh ,p i ∈P cloud }; Among them, P′ is the denoised point cloud data, w thresh The weight threshold for denoising determines which points will be removed. Points with weights lower than the threshold are considered noise and removed from the point cloud data. S37. Through steps S31 to S36, point cloud data P′ is obtained. Noise points have been removed from the data and representative structural feature points in a large spatial environment have been retained. A three-dimensional spatial feature model is constructed using the denoised and simplified point cloud data.
5. The AR 6DoF tracking and large space recognition method based on WeChat applet according to claim 1, characterized in that: The S4 specifically includes: S41. Based on the obtained three-dimensional spatial feature model, define an optimization error function ε optimized , measures the docking error between the virtual object and the real scene. The error function includes position error, rotation error and scale error. The error function is defined as follows: Among them, P i and P′ virtual,i Represents the position of the i-th point in the real scene and the virtual object, R i and R′ virtual,i Represents the rotation matrix of the i-th point in the real scene and the virtual object, S real and S′ virtual denote the scale matrices of the real scene and virtual object, respectively; λ1 and λ2 denote the parameters for adjusting the relative importance of position error, rotation error, and scale error; ‖·‖2 is the Euclidean distance, which measures the spatial distance between points; S42, minimize the error function ε through the global optimization alignment algorithm of the least squares method optimized , by calculating the error function for the virtual object position P′ virtual,i and the rotation matrix R′ virtual,i The specific steps are as follows: S421, calculate the gradient of the error function with respect to the position and rotation matrix, calculate the gradient of the error function with respect to the position and rotation matrix of the virtual object, and calculate the gradient of the error function with respect to the position P′ virtual,i The position gradient obtained by derivation is: For the rotation matrix R′ virtual,i Derivatively obtain the gradient of the rotation matrix: S422: Update the position P′ of the virtual object according to the calculated gradient by using the gradient descent method. virtual,i and the rotation matrix R′ virtual,i : Location Updates: Rotation matrix update: Among them, η is the learning rate, which controls the step size of each update; S43. Based on the preliminary docking results provided by the global optimization algorithm, the ICP algorithm is used to perform local docking optimization on the virtual object and the real scene. The specific steps are as follows: S431, calculate each pair of points (P i ,P′ virtual,i ), the local error function ε ICP As shown below: Among them, R′ virtual and T virtual They are the rotation matrix and translation vector of the virtual object respectively; S432, calculate the position ΔT by minimizing the local error function ICP and the rotation matrix ΔR ICP The update amount is as follows: Location update amount: Rotation matrix update amount: Among them, arg min represents the variable that minimizes a function; S433: Update the position P′ of the virtual object according to the minimization result of the local error function virtual and the rotation matrix R′ virtual : Location Updates: P′ virtual,optimized =P′ virtual +ΔT ICP ; Rotation matrix update: R′ virtual,optimized =R′ virtual ·ΔR ICP ; Among them, P′ virtual,optimized represents the updated position of the virtual object, R′ virtual,optimized represents the updated rotation matrix; S434, optimize the position and rotation matrix through multiple iterations until the error converges and the iteration stops under the condition: ‖ΔP′ virtual ‖2<∈1,‖ΔR′ virtual ‖2<∈2; Among them, ΔP′ virtual is the position correction, ΔR′ virtual is the correction amount of the rotation matrix, ∈1 and ∈2 are the error tolerance thresholds, indicating when the optimization process stops; S44, weightedly merging the updated result of the global optimization algorithm in S42 and the updated result of the ICP algorithm in S43 to obtain a more optimized position and rotation matrix.
Citation Information
Cited By
Unity-based WeChat applet AR development framework and application thereof
CN121635874A