Intelligent multimode AR identification system
By combining hardware devices such as RGB-D cameras and IMUs with multi-module software systems, and dynamically selecting and fusing recognition modes, the robustness and accuracy issues of AR recognition systems in complex environments are solved, enabling adaptive AR recognition and virtual content generation.
Patent Information
- Application Number
- CN202510916631.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-11-14
AI Technical Summary
Existing AR recognition systems suffer from problems such as limited protection types, insufficient robustness, lack of data quality processing mechanisms, and lagging model library updates, making them unable to adapt to dynamic scenarios.
The hardware device, consisting of an RGB-D camera, an inertial measurement unit (IMU), and a computing unit, combined with a data interface unit, a plane recognition module, a point cloud recognition module, and a model recognition module, dynamically selects the recognition mode through an intelligent switching and fusion module, and generates control commands using a data quality assessment unit to achieve multi-source data fusion and adaptive recognition.
It significantly improves positioning accuracy and robustness in complex dynamic scenes, ensures that the system automatically degrades when there is low-quality data, maintains minimum availability, and provides a stable AR experience by updating the model library online to adapt to different lighting and occlusion environments.
Smart Images

Figure CN120953328A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AR recognition, and more specifically to an intelligent multi-mode AR recognition system. Background Technology
[0002] AR (Augmented Reality) is a technology that uses computer vision, sensor fusion, and other technologies to overlay virtual information onto real physical scenes in real time. Its core lies in achieving precise interaction between virtual content and the real world through the recognition and positioning of the real environment, and it is widely used in navigation, industrial maintenance, education and training, and other fields.
[0003] AR recognition systems are the core support for AR technology. Their main function is to perceive and understand real-world scenes through hardware sensors and software algorithms, thereby determining the overlay position and posture of virtual content. Key technical requirements include: multi-source data fusion, environmental adaptability, and dynamic switching capabilities.
[0004] Current AR recognition systems have the following shortcomings in practical applications: limited protection types, insufficient robustness, lack of data quality processing mechanisms, and lagging model library updates, making them unable to adapt to dynamic scenarios. Therefore, an intelligent multi-mode AR recognition system is proposed. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides an intelligent multi-mode AR recognition system, including hardware devices and software systems;
[0006] The hardware device includes:
[0007] An RGB-D camera is used to simultaneously capture RGB images and depth maps of a scene;
[0008] An inertial measurement unit (IMU) is used to acquire the device's acceleration and angular velocity in real time.
[0009] The computing unit is used to process sensor data;
[0010] The software system includes:
[0011] Data interface unit, used for:
[0012] Convert an RGB image to a grayscale image;
[0013] Convert depth maps into point cloud data;
[0014] The initial device pose T of the IMU data is calculated by integration. imu ;
[0015] The planar recognition module is used to achieve planar localization through image feature point detection and matching, and adopts an improved ORB descriptor in the feature matching stage;
[0016] The point cloud recognition module is used to reconstruct the scene through point cloud feature extraction and registration;
[0017] The model recognition module is used to recognize objects through 3D model feature matching.
[0018] The intelligent switching and fusion module is used to dynamically select the recognition mode and fuse multi-source data;
[0019] The upper-layer application and interaction module is used to generate virtual content and respond to user interactions;
[0020] The data quality assessment unit is used to generate data quality control instructions.
[0021] Furthermore, the point cloud data acquisition process of the point cloud recognition module is specifically as follows:
[0022] The system receives raw point cloud data from the data interface unit, performs voxel grid downsampling on the raw point cloud, removes noise points by statistical outlier filtering, and uses a KD-Tree structure to accelerate neighborhood search in order to obtain preprocessed point cloud data for scene reconstruction.
[0023] Furthermore, the intelligent switching and fusion module includes:
[0024] The dynamic evaluation unit is used to automatically activate the optimal combination of recognition modes based on the scene complexity feature vector and calculate the scene complexity score based on the weight matrix W.
[0025] An adaptive fusion engine is used to generate a comprehensive pose matrix T by fusing multi-mode localization results using the entropy weight method. fused ;
[0026] The feedback enhancement unit is used to record the positioning accuracy improvement rate η. When η < η threshold At that time, η threshold The preset threshold triggers the model recognition module to perform online 3D reconstruction of the current scene;
[0027] The reconstructed model is added to the model database and new feature descriptors are generated.
[0028] Furthermore, the dynamic evaluation unit performs the following operations:
[0029] Obtain the following data from the data interface unit:
[0030] The FAST feature point density f1 of a grayscale image;
[0031] The curvature variance f2 of the point cloud data;
[0032] The confidence level f3 of the pose matrix output by the model matching module;
[0033] Next, we construct the scene complexity feature vector, specifically as follows:
[0034] Based on the weight matrix W = [w1, w2, w3] T Computational scenario complexity score:
[0035] Where wi is the i-th element of the weight matrix, which corresponds to the weight of feature fi; fi is the i-th element of the feature vector;
[0036] If S > θ1 > θ2, then the point cloud recognition module and the model recognition module are activated simultaneously for recognition, where θ1 and θ2 are preset thresholds;
[0037] If θ2≤S≤θ1, then the plane recognition module and the model recognition module are activated for recognition.
[0038] If S < θ2, then only the plane recognition module is activated for recognition;
[0039] The weight matrix W is generated through offline reinforcement learning:
[0040] Construct a training dataset containing different lighting and occlusion scenarios;
[0041] With the goal of minimizing positioning error, the weight matrix W is iteratively optimized using the Q-learning algorithm;
[0042] After each identification task is completed, the actual positioning error is fed back to the dynamic evaluation unit to update the weight matrix W;
[0043] The operation of the adaptive fusion engine includes the following process:
[0044] The pose matrix T output by the receiving plane recognition module plane The pose matrix T output by the point cloud recognition module pointcloud The pose matrix T output by the model recognition module model The pose matrix T provided by the data interface unit imu ;
[0045] Calculate the information entropy H of each matrix. k And normalized to obtain The specific process is as follows: It is the sum of the information entropy of the four modules (plane, point cloud, model, IMU), where k∈{plane,pointcloud,model,imu};
[0046] Calculate the fusion weights of each matrix: Hm is a preset entropy value used to adjust the denominator in the weight calculation. This is the normalized entropy value of the kth pose.
[0047] Generate the fused pose matrix: T k Let T be the pose matrix of the k-th recognition module, namely Tplane, Tpointcloud, and Tmodel;
[0048] The adaptive fusion engine responds to data quality commands:
[0049] When a point cloud degradation command is received, T pointcloud The entropy value H k Calculate by magnification of 2 times;
[0050] When an IMU disable command is received, set α imu =0.
[0051] Furthermore, information entropy H k The calculation method is as follows:
[0052] Decompose the pose matrix into rotational components R. k Translation component t k =[t x , t y , t z ] T ;
[0053] Calculate the quaternion angular velocity variance of the rotational component
[0054] Calculate the Shannon entropy of the translation component:
[0055]
[0056] Where p(t) i ) represents the coordinate value t of the translation component. i The probability distribution within the time window, where i is the dimension index of the translation component, i = 1, 2, 3 correspond to the x, y, z axis coordinates in three-dimensional space respectively, and ti is the i-th dimension coordinate value of the translation component of the k-th module;
[0057] Define the overall uncertainty: In the formula, The quaternion angular velocity variance of the rotation component is used to quantify the degree of attitude jitter.
[0058] Furthermore, the process for obtaining the positioning accuracy improvement rate is defined as follows:
[0059]
[0060] Where T gt For the true pose, T prevFor the pose matrix of the previous frame, |·| F Let Frobenius be the matrix norm.
[0061] Furthermore, the online reconstruction operation includes:
[0062] Dense point cloud based on pose matrix output by point cloud recognition module;
[0063] The surface mesh is generated using the moving cube algorithm;
[0064] Key features of the model are extracted using FPFH descriptors and stored in the database.
[0065] Furthermore, the planar recognition module employs an improved ORB descriptor in the feature matching stage and introduces depth map gradient information in the descriptor generation stage.
[0066] The descriptor dimension is expanded to 384 bits, with the first 256 bits being traditional ORB features and the last 128 bits being a deep gradient histogram.
[0067] The virtual content generation of the upper-layer application and interaction module satisfies:
[0068] When T fused When the rate of change of the rotational component exceeds the threshold, a low-precision simplified model is enabled;
[0069] When the translation component is stable for several consecutive frames, switch to the high-precision model.
[0070] Furthermore, the data quality assessment unit is used to generate data quality control instructions, including:
[0071] Image discard command: triggered when the blur level of a grayscale image exceeds a threshold;
[0072] Point cloud degradation command: triggered when the percentage of valid points in the point cloud data is lower than a threshold;
[0073] IMU disable command: Triggered when the variance of IMU readings exceeds a threshold.
[0074] Furthermore, the data quality assessment unit performs the following operations:
[0075] Image blur detection:
[0076] Calculate the sum of squared gradients of a grayscale image Where x and y are the coordinate indices of the image pixels; ΔI(x, y) is the gradient vector of the grayscale image at pixel (x, y), reflecting the rate of change of the pixel value;
[0077] If G sum <τ g Generate an image discard instruction and request re-acquisition;
[0078] Point cloud integrity detection: Calculate the effective point cloud percentage ρ, the specific process is as follows:
[0079] N valid N is the number of points that pass the outlier filter. total This represents the total number of points in the original point cloud, i.e., the number of unfiltered point clouds output by the data interface unit.
[0080] If ρ < τ ρ Generate point cloud degradation instructions;
[0081] IMU anomaly detection: Calculate the accelerometer variance and gyroscope variance. When either the accelerometer variance or the gyroscope variance exceeds the corresponding threshold, generate an IMU disable command.
[0082] The point cloud recognition module performs a degradation processing operation:
[0083] Skip curvature calculation in the feature extraction stage;
[0084] ICP registration was performed using voxel center points instead of the original point cloud.
[0085] In the output pose matrix T pointcloud Add low-confidence markers;
[0086] The data quality assessment unit works in conjunction with the intelligent switching module:
[0087] When the image discard command is triggered for 3 consecutive frames, the dynamic evaluation unit is forced to switch to point cloud recognition mode;
[0088] When the IMU disable command lasts for more than 5 seconds, the static scene matching algorithm of the model recognition module is activated.
[0089] The beneficial effects of this invention are reflected in:
[0090] The system automatically and dynamically selects the optimal combination of recognition modes based on scene complexity, solving the problem of single recognition methods failing in complex environments. By fusing pose information from planes, point clouds, models, and IMUs using the entropy weight method, it significantly improves positioning accuracy and robustness in complex dynamic scenes. During the preprocessing stage, it screens for image blur, point cloud integrity, and abnormal IMU data in real time to prevent invalid data from entering the processing flow and causing system crashes. It automatically triggers degradation processing for low-quality data to maintain minimum system availability and ensure service continuity. Based on positioning accuracy feedback, it triggers online reconstruction of scene models, continuously expanding the model database and achieving autonomous evolution of system performance. When there is intense movement, it uses a low-precision model to ensure smoothness, and switches to a high-precision model when stationary to enhance realism and optimize computing power allocation efficiency. It dynamically optimizes recognition weight parameters through reinforcement learning to adapt to different lighting and occlusion environments, eliminating the need for manual parameter tuning. It integrates multimodal interaction methods and combines stable pose rendering of virtual content to provide a natural and intuitive AR experience. The data quality assessment and intelligent switching module are linked, and a backup recognition mode is forcibly activated when there are continuous anomalies, building a multi-layered fault protection system. Attached Figure Description
[0091] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the accompanying drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.
[0092] Figure 1 This is a system block diagram of the present invention. Detailed Implementation
[0093] The embodiments of the technical solution of the present invention will now be described in detail with reference to the accompanying drawings. These embodiments are merely illustrative of the technical solution of the present invention and are therefore intended to limit the scope of protection of the present invention.
[0094] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application should have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains.
[0095] like Figure 1 As shown, an intelligent multi-mode AR recognition system includes hardware devices and software systems;
[0096] The hardware device includes:
[0097] An RGB-D camera is used to simultaneously capture RGB images and depth maps of a scene;
[0098] An inertial measurement unit (IMU) is used to acquire the device's acceleration and angular velocity in real time.
[0099] The computing unit is used to process sensor data;
[0100] The software system includes:
[0101] Data interface unit, used for:
[0102] Convert an RGB image to a grayscale image;
[0103] Convert depth maps into point cloud data;
[0104] The initial device pose T of the IMU data is calculated by integration. imu ;
[0105] The planar recognition module is used to achieve planar localization through image feature point detection and matching, and adopts an improved ORB descriptor in the feature matching stage;
[0106] The point cloud recognition module is used to reconstruct the scene through point cloud feature extraction and registration;
[0107] The model recognition module is used to recognize objects through 3D model feature matching.
[0108] The intelligent switching and fusion module is used to dynamically select the recognition mode and fuse multi-source data;
[0109] The upper-layer application and interaction module is used to generate virtual content and respond to user interactions;
[0110] The data quality assessment unit is used to generate data quality control instructions.
[0111] The point cloud data acquisition process of the point cloud recognition module is as follows:
[0112] The system receives raw point cloud data from the data interface unit, performs voxel grid downsampling on the raw point cloud, removes noise points by statistical outlier filtering, and uses a KD-Tree structure to accelerate neighborhood search in order to obtain preprocessed point cloud data for scene reconstruction.
[0113] By using voxel grid downsampling technology, millions of point clouds are compressed to tens of thousands of points, effectively solving the computing power bottleneck of mobile devices. Dense point clouds that originally required high-end GPUs to process can now run smoothly on ordinary mobile phone processors, enabling AR systems to achieve fast response on lightweight devices such as smart glasses.
[0114] Statistical outlier filtering can intelligently distinguish between real object surfaces and noise interference. For example, in an industrial workshop environment, it can accurately eliminate instantaneous interference points such as welding sparks and metal reflections, while preserving the complete geometric structure of the equipment outline, providing clean input data for subsequent recognition.
[0115] The efficiency of neighborhood search is improved by more than ten times. When the system needs to calculate the curvature of the point cloud, the brute-force search that used to require traversing all points can now be quickly located using spatial indexing, significantly reducing CPU load and extending device battery life.
[0116] The intelligent switching and fusion module includes:
[0117] The dynamic evaluation unit is used to automatically activate the optimal combination of recognition modes based on the scene complexity feature vector and calculate the scene complexity score based on the weight matrix W.
[0118] An adaptive fusion engine is used to generate a comprehensive pose matrix T by fusing multi-mode localization results using the entropy weight method. fused ;
[0119] The feedback enhancement unit is used to record the positioning accuracy improvement rate η. When η < η threshold At that time, η thredshold The preset threshold triggers the model recognition module to perform online 3D reconstruction of the current scene;
[0120] Add the reconstructed model to the model database and generate new feature descriptors;
[0121] By calculating scene complexity scores in real time, the system automatically identifies changes in environmental features, such as texture loss, motion blur, and low-light interference. It activates the best combination of recognition modes based on the score threshold to avoid the failure of a single mode in complex scenes. It also automatically allocates fusion weights based on pose uncertainty to solve sensor conflict problems.
[0122] When abnormal data is automatically downweighted, such as suppressing the IMU when there is jitter, or reducing its weight when the point cloud is sparse, the stability of virtual and real matching is improved. The recognition effect of each recognition is quantified by the positioning improvement rate, and the model reconstruction is automatically triggered. The new object model is reconstructed online and added to the database, so that the system has continuous learning ability. Multi-mode parallelism is enabled in high-complexity scenarios, while single mode is retained in simple scenarios to save energy. When a module fails, the fusion strategy is automatically adjusted to ensure basic functions.
[0123] The dynamic evaluation unit performs the following operations:
[0124] Obtain the following data from the data interface unit:
[0125] The FAST feature point density f1 of a grayscale image;
[0126] The curvature variance f2 of the point cloud data;
[0127] The confidence level f3 of the pose matrix output by the model matching module;
[0128] Next, we construct the scene complexity feature vector, specifically as follows:
[0129] Based on the weight matrix W = [w1, w2, w3] T Calculate the scenario complexity score, using the weight matrix W = [w1, w2, w3]. T Generated through offline reinforcement learning, used to calculate scene complexity scores, where w1 corresponds to the FAST feature point density f1 of the grayscale image, w2 corresponds to the curvature variance f2 of the point cloud data, w3 corresponds to the confidence score f3 of the model matching module, and T represents matrix transpose.
[0130] Where wi is the i-th element of the weight matrix, corresponding to the weight of feature fi; fi is the i-th element of the feature vector, W T Transpose the weight matrix;
[0131] If S > θ1 > θ2, then the point cloud recognition module and the model recognition module are activated simultaneously for recognition, where θ1 and θ2 are preset thresholds;
[0132] If θ2≤S≤θ1, then the plane recognition module and the model recognition module are activated for recognition.
[0133] If S < θ2, then only the plane recognition module is activated for recognition;
[0134] The weight matrix W is generated through offline reinforcement learning:
[0135] Construct a training dataset containing different lighting and occlusion scenarios;
[0136] With the goal of minimizing positioning error, the weight matrix W is iteratively optimized using the Q-learning algorithm;
[0137] After each identification task is completed, the actual positioning error is fed back to the dynamic evaluation unit to update the weight matrix W;
[0138] The operation of the adaptive fusion engine includes the following process:
[0139] The pose matrix T output by the receiving plane recognition module plane The pose matrix T output by the point cloud recognition module pointcloud The pose matrix T output by the model recognition module model The pose matrix T provided by the data interface unit imu ;
[0140] Calculate the information entropy H of each matrix. k And normalized to obtain The specific process is as follows:
[0141] It is the sum of the information entropy of the four modules (plane, point cloud, model, IMU), where k∈{plane,pointcloud,model,imu};
[0142] Calculate the fusion weights of each matrix: Hm is a preset entropy value used to adjust the denominator in the weight calculation. This is the normalized entropy value of the kth pose.
[0143] Generate the fused pose matrix: T k Let T be the pose matrix of the k-th recognition module, namely Tplane, Tpointcloud, and Tmodel;
[0144] The adaptive fusion engine responds to data quality commands:
[0145] When a point cloud degradation command is received, T pointcloud The entropy value H k Calculate by magnification of 2 times;
[0146] When an IMU disable command is received, set α imu =0;
[0147] The system quantifies scene features in real time and automatically calculates scene scores based on a weight matrix optimized by reinforcement learning. When a weak-texture environment is detected, it immediately switches to a point cloud-dominated mode. In cluttered garage scenes, it employs multi-mode fusion, completely resolving the pain point of traditional AR failing in extreme environments. Weights are dynamically allocated based on the uncertainty of each positioning source. When the phone camera shakes during shooting, IMU data is automatically suppressed. When rain or fog interferes with the point cloud, the model's recognition weights are increased, ensuring that the virtual navigation arrow remains stably aligned with the real road surface on bumpy cycling routes, reducing positioning jitter.
[0148] After each operation, the system automatically verifies the improvement rate of positioning accuracy. When a new object recognition deviation is detected, an online reconstruction mechanism is triggered immediately. The system autonomously improves the model database within three days, enhancing the accuracy of secondary recognition in the same scene. In case of sensor malfunctions, such as continuous IMU jitter, a degradation strategy is automatically initiated, isolating the fault source by resetting weights to zero, ensuring uninterrupted basic navigation functions. Users can still obtain usable guidance in areas with poor signal coverage, such as subway tunnels, and the system crash rate has decreased.
[0149] For example, when a delivery driver enters an older residential area by bicycle and needs to use AR navigation:
[0150] Scene diagnosis: The dappled shadows of the sycamore trees indicate sparse texture features and a decrease in f1.
[0151] The chaotic, illegally constructed iron sheds indicate a sudden increase in the curvature f2 of the point cloud.
[0152] At this point, the system score S > θ1, activating the dual mode of point cloud and model;
[0153] Then perform anti-interference fusion: the bumpy stone road indicates that IMU weights need to be suppressed to prevent positioning jitter;
[0154] Strong sunlight interference from the setting sun indicates a need to increase the model weights;
[0155] The virtual navigation arrow generated at this time is firmly attached to the uneven road surface with an error of less than 5 centimeters.
[0156] Information entropy H k The calculation method is as follows:
[0157] Decompose the pose matrix into rotational components R. k Translation component t k =[t x , t y , t z ] T ;
[0158] Calculate the quaternion angular velocity variance of the rotational component
[0159] Calculate the Shannon entropy of the translation component:
[0160]
[0161] Where p(t) i ) represents the coordinate value t of the translation component. i The probability distribution within the time window, where i is the dimension index of the translation component, i = 1, 2, 3 correspond to the x, y, z axis coordinates in three-dimensional space respectively, and ti is the i-th dimension coordinate value of the translation component of the k-th module;
[0162] Define the overall uncertainty: In the formula, The quaternion angular velocity variance of the rotation component is used to quantify the degree of attitude jitter.
[0163] Attitude jitter is quantified by rotational variance. For example, the noise of the IMU when the equipment shakes. Coordinate stability, such as positioning fluctuations when the point cloud is sparse, is measured by translational entropy, enabling real-time evaluation of the output quality of each module.
[0164] For example: the rotational variance of the IMU when the user's handheld device is rotated rapidly. Increase, leading to H k If the weight of the IMU is increased, the system will automatically reduce its weight in the fusion process to avoid location drift.
[0165] Modules with lower entropy values, such as point cloud recognition in stable scenarios, receive higher weight in fusion, resolving data conflicts between different sensors and improving overall positioning accuracy.
[0166] In a weakly textured environment, the translational entropy H of planar recognition is... trans,k As the weight increases, the system reduces its weight and instead relies on low-entropy data from point cloud recognition to ensure stable virtual object alignment.
[0167] In conjunction with data quality assessment, when data in a certain module is abnormal, the fusion strategy is dynamically adjusted through an entropy amplification mechanism to maintain system robustness.
[0168] If a user uses AR navigation in a subway station, and the handheld device passes over a glass curtain wall area, the glass curtain wall area has a weak texture and is reflective.
[0169] Planar recognition: The glass curtain wall has sparse texture, low FAST feature point density, and large fluctuations in translation components. trans,k High, rotational variance The equipment is shaking slightly.
[0170] Point cloud recognition: Glass reflection results in a low percentage of valid points in the point cloud, triggering a degradation instruction, but depth information is still partially valid, and the translation entropy H... trans,k Medium, low rotational variance.
[0171] IMU: Rotational variance due to device shaking while the user is walking. High, translation entropy H trans,k medium.
[0172] Model recognition: The pre-stored subway station column model was successfully matched, and the translation entropy H was obtained. trans,k Low, rotational variance Low.
[0173] Information entropy calculation and fusion weights:
[0174] H recognized by the model k The lowest value, meaning low uncertainty in both rotation and translation, and low fusion weight α. k Highest;
[0175] IMU's H k The highest value, i.e., significant rotational jitter, has a weight α. k A value close to 0 indicates that it may be disabled;
[0176] H for planar and point cloud recognition k Medium weighting, with weights distributed inversely proportional to entropy.
[0177] At this time, the virtual navigation arrow is mainly rendered based on the model recognition results to avoid positioning drift caused by the lack of texture on the glass curtain wall. Even if the device shakes, it can still stably fit the ground marking line.
[0178] The process for obtaining the positioning accuracy improvement rate is defined as follows:
[0179]
[0180] Where T gt For the true pose, T prev For the pose matrix of the previous frame, |·| F The Frobenius norm of the matrix;
[0181] When η is lower than a preset threshold, online model reconstruction is triggered, enabling the system to dynamically update the 3D model library based on actual errors, thus achieving a closed-loop evolution from identification to evaluation to optimization.
[0182] In complex environments such as changing lighting and occlusion, the effectiveness of the identification strategy is evaluated in real time by η, and the system automatically switches to a better combination of modes to avoid degradation of positioning accuracy.
[0183] If a user uses AR navigation to find products in a supermarket, the system needs to recognize the shelf surface and overlay virtual labels.
[0184] Initial identification:
[0185] During the initial recognition, only the planar recognition module was used, and the pose error was ||T||. prev -T gt || F = 8 centimeters.
[0186] After fusing point cloud data with model recognition, the error is reduced to ||T|| fused -T gt || F = 3 centimeters.
[0187] The improvement rate of positioning accuracy is calculated as follows: η = (8-3) / 8 = 0.625, which is an improvement of 62.5%.
[0188] At this point, η is higher than the threshold, and the system maintains the current model.
[0189] Optimization after environmental changes:
[0190] When the product placement on the shelf changes and the system is re-identified, the error after fusion increases to ||T. prev -T gt || F =6 cm, η = (8-6) / 8 = 0.25, which is below the threshold.
[0191] The system triggers online reconstruction: a new shelf model is generated based on the new point cloud data, and the feature descriptors are updated.
[0192] Upon re-identification, the error decreased to ||T prev -T gt || F=2 cm, η = (8-2) / 8 = 0.75, the positioning accuracy is significantly improved. When η is lower than the preset threshold, the model is reconstructed online, so that the system can dynamically update the 3D model library according to the actual error, and realize the closed-loop evolution from recognition to evaluation to optimization.
[0193] In complex environments such as changing lighting and occlusion, the effectiveness of the identification strategy is evaluated in real time by η, and the system automatically switches to a better combination of modes to avoid degradation of positioning accuracy.
[0194] The online reconstruction operation includes:
[0195] Dense point cloud based on pose matrix output by point cloud recognition module;
[0196] The surface mesh is generated using the moving cube algorithm;
[0197] Key features of the model are extracted and stored in the database using FPFH descriptors; by comparing the pose error before and after fusion, the improvement of positioning accuracy by multi-mode recognition, such as plane, point cloud, and model fusion, is accurately measured, providing a quantitative basis for system optimization.
[0198] Dynamically adapting to scene changes and solving the problem of outdated models: When objects in the scene are updated, such as when goods on the shelf are replaced or new equipment is added, the system automatically reconstructs the 3D model through positioning error feedback, avoiding recognition failure caused by outdated model library.
[0199] The system can independently expand its model library to improve its generalization ability in complex scenes. The online reconstruction model is automatically added to the library and new features are generated, enabling the system to recognize more object types and reduce localization failures caused by unknown objects.
[0200] The closed-loop optimization mechanism reduces the cost of manual operation and maintenance. There is no need for manual model updates. The system evolves autonomously based on the actual positioning effect, which is especially suitable for continuous adaptation in dynamic scenarios.
[0201] For example, if a user uses an AR mobile app to navigate and find products in a chain supermarket, the supermarket updates the shelf display weekly.
[0202] Perform online reconstruction of the newly added shelf model;
[0203] Initial state: The supermarket has added a shelf for imported snacks, but there is no data for this shelf in the APP model library.
[0204] When a user first navigates to this area, the model recognition module has no matching model, resulting in a large fusion positioning error and a positioning accuracy improvement rate η = 0.3, which is lower than the threshold of 0.5.
[0205] Triggering the reconstruction mechanism: When the feedback reinforcement unit detects that η is not up to standard, it uses the pose matrix of the shelf point cloud data output by the point cloud recognition module, such as the metal frame and snack packaging outline, to generate a surface mesh using the moving cube algorithm. It then extracts key shelf features, such as shelf height and corner curvature, using the FPFH descriptor.
[0206] Model import and effect: After the new shelf model is added to the database, when the user navigates again, the model matching confidence improves, the fusion positioning error decreases from 10 cm to 4 cm, η=0.6, and the virtual navigation arrow accurately fits the edge of the shelf.
[0207] The planar recognition module uses an improved ORB descriptor in the feature matching stage and introduces depth map gradient information in the descriptor generation stage.
[0208] The descriptor dimension is expanded to 384 bits, with the first 256 bits being traditional ORB features and the last 128 bits being a deep gradient histogram.
[0209] The virtual content generation of the upper-layer application and interaction module satisfies:
[0210] When T fused When the rate of change of the rotational component exceeds the threshold, a low-precision simplified model is enabled;
[0211] When the translation component is stable for several consecutive frames, switch to the high-precision model;
[0212] Improved ORB descriptor for planar recognition module: Introduces depth map gradient information, expanding the descriptor dimension to 384 bits (the first 256 bits are traditional ORB features, and the last 128 bits are depth gradient histogram).
[0213] The dynamic model switching strategy for upper-layer applications: when the rotation rate of change is high, a low-precision simplified model is used, and when the translation is stable, it is switched to a high-precision model.
[0214] Improved recognition accuracy in weak texture scenes;
[0215] Traditional ORB relies on image texture features, which are prone to failure in areas with weak texture, such as white supermarket shelves and glass freezers. Improved descriptors supplement geometric features with depth gradient information, such as the three-dimensional contours of shelf edges, to enhance matching robustness.
[0216] For example, the glass surface of a supermarket freezer has a sparse texture. The traditional ORB feature point density f1 = 5 points / square centimeter. After improvement, by combining depth gradient, the feature point density is increased to 15 points / square centimeter, and the matching accuracy increases from 30% to 85%.
[0217] Enhanced resistance to environmental interference;
[0218] The depth gradient histogram is not sensitive to changes in lighting, such as reflections caused by direct spotlights in a supermarket, and can effectively distinguish object boundaries, reducing false matching.
[0219] For example, when beverage shelves are exposed to strong light, the traditional ORB (Organic Array Block) suffers from increased false matching rate due to pixel gradient distortion, while the improved ORB reduces the false matching rate.
[0220] A dynamic balance between computing power and user experience;
[0221] When the rotation rate is high, such as when a user moves their phone quickly to find a product, a low-precision model is switched to reduce the amount of rendering computation and ensure a smooth 60fps performance; when the translation is stable, such as when a user stops to view a product, a high-precision model is enabled to present details and enhance the realism of AR.
[0222] For example, when using AR-guided shopping in the refrigerated section of a supermarket:
[0223] Accurate localization in weakly textured environments:
[0224] Customers are looking for yogurt in the refrigerated section, where the shelves are made of stainless steel with glass doors and have a sparse texture.
[0225] Problems with traditional solutions: Traditional ORB cannot stably match the glass surface, and the virtual label drifts frequently with an error exceeding 10 centimeters.
[0226] The improved ORB descriptor captures the 3D contours of glass door edges through depth gradients, such as the abrupt changes in depth at the corners of the door frame. It combines 256-bit traditional texture features with a 128-bit depth gradient histogram to generate more unique feature descriptors.
[0227] Matching accuracy has been improved, and virtual yogurt labels are now precisely aligned with shelf railings, reducing errors.
[0228] Dynamic model switching optimizes the user experience;
[0229] Users quickly move from the left side of the refrigerated display case to the right side, holding their mobile phones, to find the desired product.
[0230] Dynamic strategy execution: When the rate of change of the rotation component is detected to be greater than 150° / second, the upper-layer application switches to a low-precision simplified model, such as using a cube instead of a yogurt box, and the rendering frame rate is maintained at 60fps.
[0231] When the user stops the translation component and it remains stable for 3 consecutive frames, it automatically switches to a high-precision model, which includes yogurt carton label texture and bump details. The model's face count is increased from 800 to 12,000, enhancing realism.
[0232] When moving quickly, the navigation arrows follow smoothly, and when stopped, the product details are clearly visible without any visual gaps.
[0233] The data quality assessment unit is used to generate data quality control instructions, including:
[0234] Image discard command: triggered when the blur level of a grayscale image exceeds a threshold;
[0235] Point cloud degradation command: triggered when the percentage of valid points in the point cloud data is lower than a threshold;
[0236] IMU disable command: Triggered when the variance of IMU readings exceeds a threshold.
[0237] The data quality assessment unit performs the following operations:
[0238] Image blur detection:
[0239] Calculate the sum of squared gradients of a grayscale image Where x and y are the coordinate indices of the image pixels; This is the gradient vector of a grayscale image at pixel (x,y), reflecting the rate of change of the pixel value;
[0240] If G sum <τ g , τ g To set a preset threshold, generate an image discard instruction and request re-acquisition;
[0241] Point cloud integrity detection: Calculate the effective point cloud percentage ρ, the specific process is as follows:
[0242] N valid N is the number of points that pass the outlier filter. total This represents the total number of points in the original point cloud, i.e., the number of unfiltered point clouds output by the data interface unit.
[0243] If ρ < τ p Generate point cloud degradation instructions;
[0244] IMU anomaly detection: Calculate the accelerometer variance and gyroscope variance. When either the accelerometer variance or the gyroscope variance exceeds the corresponding threshold, generate an IMU disable command.
[0245] The point cloud recognition module performs a degradation processing operation:
[0246] Skip curvature calculation in the feature extraction stage;
[0247] ICP registration was performed using voxel center points instead of the original point cloud.
[0248] In the output pose matrix T pointcloud Add low-confidence markers;
[0249] The data quality assessment unit works in conjunction with the intelligent switching module:
[0250] When the image discard command is triggered for 3 consecutive frames, the dynamic evaluation unit is forced to switch to point cloud recognition mode;
[0251] When the IMU disable command lasts for more than 5 seconds, the static scene matching algorithm of the model recognition module is activated.
[0252] Real-time screening of low-quality data such as blurry images and sparse point clouds prevents them from entering the feature extraction stage and avoids localization failures caused by noisy data, such as feature point mismatch and point cloud registration deviation.
[0253] If supermarket lighting reflections cause blurry images of shelves captured by a camera, the sum of squared gradients G sum <τ g The system discards the image frame and re-acquires it to avoid the planar recognition module misjudging the shelf location due to blurry images.
[0254] Adaptive degradation handling maintains minimum system availability;
[0255] Instead of discarding low-quality data directly, simplified algorithms are enabled, such as skipping curvature calculations when point clouds are downgraded. This allows basic positioning services to still be provided even when data quality is insufficient, preventing the AR function from becoming completely ineffective.
[0256] For example, if the reflection from the glass of a supermarket freezer causes the effective point cloud percentage ρ = 25% < τ, then the effective point percentage τ is less than 25%. p After the downgrade command is triggered, the point cloud recognition module skips curvature calculation and uses voxel center point registration. Although the positioning accuracy decreases, it can still maintain the virtual label's approximate fit.
[0257] By linking data quality commands with the intelligent switching module, such as discarding three consecutive frames of images to force a switch to point cloud mode, and activating a backup recognition strategy when a single sensor fails, the system's fault tolerance in complex environments is improved.
[0258] For example, when a user walks quickly in a supermarket, the IMU experiences violent shaking, which increases the variance of the accelerometer.
[0259] After the disable command is triggered, the system automatically reduces the IMU fusion weight to 0, relying on point cloud and model recognition to maintain positioning, thus avoiding large drift of the virtual arrow due to IMU noise.
[0260] For example, when using AR-guided shopping in the deli section of a supermarket:
[0261] When users photographed the food inside the glass display case in the cooked food section, the reflection from the glass caused the images to become blurry.
[0262] Data quality assessment: Calculate the sum of squared gradients G of the grayscale image sum =800, below the threshold τ g =1500, triggering the image discard command.
[0263] The system indicates that the image is blurry and is retaking the image, automatically retrying to acquire a clear image.
[0264] Results: The sum of squared gradients of the re-acquired images increased to 2000, the planar recognition module accurately extracted the edge features of the counter, and the virtual food labels were accurately superimposed with an error of less than 2 cm.
[0265] When performing degradation processing for sparse point clouds:
[0266] The supermarket's strong spotlights shone directly on the cheese shelves, resulting in an increase in invalid points in the reflective areas of the point cloud data.
[0267] Data quality assessment: Percentage of valid points ρ = 5000 / 20000 = 25% < τ p =40%, triggering the point cloud degradation instruction.
[0268] Degradation strategy execution: The point cloud recognition module skips curvature calculation, uses voxel center points instead of the original point cloud for ICP registration, and marks low confidence in the output pose matrix.
[0269] At this point, the positioning accuracy dropped from 3 cm to 6 cm, but the virtual tag was not lost due to the sparse point cloud and could still roughly indicate the cheese area, ensuring basic navigation functions.
[0270] Fault isolation during IMU malfunction:
[0271] When a user holds a mobile phone and moves quickly while pushing a shopping cart in a supermarket, the IMU (Important Detector Unit) readings become abnormal due to vibration.
[0272] Data quality assessment: Gyroscope variance = 0.8 > preset value 0.5, triggering IMU disable command.
[0273] Linkage switching strategy: The adaptive fusion engine sets the IMU weight α=0. At the same time, if the IMU is disabled for more than 5 seconds, the static scene matching algorithm of the model recognition module is activated, such as matching the shelf column model.
[0274] At this point, the virtual navigation arrow no longer vibrates with the device's vibration. It relies on the shelf model for positioning, with the error maintained within 5 centimeters, ensuring that users can follow the guidance normally.
[0275] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.
Claims
1. An intelligent multi-mode AR recognition system, characterized in that, This includes hardware equipment and software systems; The hardware device includes: An RGB-D camera is used to simultaneously capture RGB images and depth maps of a scene; An inertial measurement unit (IMU) is used to acquire the device's acceleration and angular velocity in real time. The computing unit is used to process sensor data; The software system includes: Data interface unit, used for: Convert an RGB image to a grayscale image; Convert depth maps into point cloud data; The initial device pose T of the IMU data is calculated by integration. imu ; The planar recognition module is used to achieve planar localization through image feature point detection and matching, and adopts an improved ORB descriptor in the feature matching stage; The point cloud recognition module is used to reconstruct the scene through point cloud feature extraction and registration; The model recognition module is used to recognize objects through 3D model feature matching. The intelligent switching and fusion module is used to dynamically select the recognition mode and fuse multi-source data; The upper-layer application and interaction module is used to generate virtual content and respond to user interactions; The data quality assessment unit is used to generate data quality control instructions.
2. The intelligent multi-mode AR recognition system according to claim 1, characterized in that: The point cloud data acquisition process of the point cloud recognition module is as follows: The system receives raw point cloud data from the data interface unit, performs voxel grid downsampling on the raw point cloud, removes noise points by statistical outlier filtering, and uses a KD-Tree structure to accelerate neighborhood search in order to obtain preprocessed point cloud data for scene reconstruction.
3. The intelligent multi-mode AR recognition system according to claim 1, characterized in that: The intelligent switching and fusion module includes: The dynamic evaluation unit is used to automatically activate the optimal combination of recognition modes based on the scene complexity feature vector and calculate the scene complexity score based on the weight matrix W. An adaptive fusion engine is used to generate a comprehensive pose matrix T by fusing multi-mode localization results using the entropy weight method. fused ; The feedback reinforcement unit is used to record the positioning accuracy improvement rate η. When η < η threshold At that time, η threshold The preset threshold triggers the model recognition module to perform online 3D reconstruction of the current scene; The reconstructed model is added to the model database and new feature descriptors are generated.
4. The intelligent multi-mode AR recognition system according to claim 3, characterized in that: The dynamic evaluation unit performs the following operations: Obtain the following data from the data interface unit: The FAST feature point density f1 of a grayscale image; The curvature variance f2 of point cloud data; The confidence level f3 of the pose matrix output by the model matching module; Next, construct the scene complexity feature vector, specifically F = [f1, f2, f3]; Based on the weight matrix W = [w1, w2, w3] T Computational scenario complexity score: Where wi is the i-th element of the weight matrix, which corresponds to the weight of feature fi; fi is the i-th element of the feature vector; If S > θ1 > θ2, then the point cloud recognition module and the model recognition module are activated simultaneously for recognition, where θ1 and θ2 are preset thresholds; If θ2≤S≤θ1, then the plane recognition module and the model recognition module are activated for recognition. If S < θ2, then only the plane recognition module is activated for recognition; The weight matrix W is generated through offline reinforcement learning: Construct a training dataset containing different lighting and occlusion scenarios; With the goal of minimizing positioning error, the weight matrix W is iteratively optimized using the Q-learning algorithm; After each identification task is completed, the actual positioning error is fed back to the dynamic evaluation unit to update the weight matrix W; The operation of the adaptive fusion engine includes the following process: The pose matrix T output by the receiving plane recognition module plane The pose matrix T output by the point cloud recognition module pointcloud The pose matrix T output by the model recognition module model The pose matrix T provided by the data interface unit imu ; Calculate the information entropy H of each matrix. k And normalized to obtain The specific process is as follows: The sum of the information entropy of the plane recognition module, point cloud recognition module, model recognition module and data interface unit, where k∈{plane,pointcloud,model,imu}; Calculate the fusion weights of each matrix: Hm is a preset entropy value used to adjust the denominator in the weight calculation; Generate the fused pose matrix: T k Let T be the pose matrix of the k-th recognition module, namely Tplane, Tpointcloud, and Tmodel; The adaptive fusion engine responds to data quality commands: When a point cloud degradation command is received, T pointcloud The entropy value H k Calculate by magnification of 2 times; When an IMU disable command is received, set α imu =0.
5. The intelligent multi-mode AR recognition system according to claim 3, characterized in that: Information entropy H k The calculation method is as follows: Decompose the pose matrix into rotational components R k Translation component t k =[t x , t y , t z ] T ; Calculate the quaternion angular velocity variance of the rotational component Calculate the Shannon entropy of the translation component: Where p(t) i ) represents the coordinate value t of the translation component. i The probability distribution within the time window, where i is the dimension index of the translation component, i = 1, 2, 3 correspond to the x, y, z axis coordinates in three-dimensional space respectively, and ti is the i-th dimension coordinate value of the k-th module translation component; Define the overall uncertainty:
6. The intelligent multi-mode AR recognition system according to claim 3, characterized in that: The process for obtaining the positioning accuracy improvement rate is defined as follows: Where Tgt is the true pose, Tprev is the pose of the previous frame, |·| F Let Frobenius be the matrix norm.
7. The intelligent multi-mode AR recognition system according to claim 6, characterized in that: The online reconstruction operation includes: Dense point cloud based on pose matrix output by point cloud recognition module; The surface mesh is generated using the moving cube algorithm; Key features of the model are extracted using FPFH descriptors and stored in the database.
8. The intelligent multi-mode AR recognition system according to claim 1, characterized in that: The planar recognition module uses an improved ORB descriptor in the feature matching stage and introduces depth map gradient information in the descriptor generation stage. The descriptor dimension is expanded to 384 bits, with the first 256 bits being traditional ORB features and the last 128 bits being a deep gradient histogram. The virtual content generation of the upper-layer application and interaction module satisfies: When T fused When the rate of change of the rotational component exceeds the threshold, a low-precision simplified model is enabled. When the translation component is stable for several consecutive frames, switch to the high-precision model.
9. The intelligent multi-mode AR recognition system according to claim 1, characterized in that: The data quality assessment unit is used to generate data quality control instructions, including: Image discard command: triggered when the blur level of a grayscale image exceeds a threshold; Point cloud degradation command: triggered when the percentage of valid points in the point cloud data is lower than a threshold; IMU disable command: Triggered when the variance of IMU readings exceeds a threshold.
10. The intelligent multi-mode AR recognition system according to claim 9, characterized in that: The data quality assessment unit performs the following operations: Image blur detection: Calculate the sum of squared gradients of a grayscale image Where x and y are the coordinate indices of the image pixels; This is the gradient vector of a grayscale image at pixel (x,y), reflecting the rate of change of the pixel value; If G sum <τ g Generate an image discard instruction and request reacquisition, τ g The preset threshold; Point cloud integrity detection: Calculate the percentage of valid point clouds N valid N is the number of points that pass the outlier filter. total This represents the total number of points in the original point cloud, i.e., the number of unfiltered point clouds output by the data interface unit. If ρ < τ p Generate point cloud degradation instructions, τ p The preset threshold; IMU anomaly detection: Calculate the accelerometer variance and gyroscope variance. When either the accelerometer variance or the gyroscope variance exceeds the corresponding threshold, generate an IMU disable command. The point cloud recognition module performs a degradation processing operation: Skip curvature calculation in the feature extraction stage; ICP registration was performed using voxel center points instead of the original point cloud. In the output pose matrix T pointcloud Add low-confidence markers; The data quality assessment unit works in conjunction with the intelligent switching module: When the image discard command is triggered for 3 consecutive frames, the dynamic evaluation unit is forced to switch to point cloud recognition mode; When the IMU disable command lasts for more than 5 seconds, the static scene matching algorithm of the model recognition module is activated.
Citation Information
Patent Citations
RGBD visual inertia simultaneous localization and mapping based on point-line feature fusion
CN113763470A
Multi-source data fusion scene space model adaptive modeling method
CN119339007A
AR game camera system based on combination of real scene and 3D game elements
CN120163717A
Point cloud segmentation method and system, and computer storage medium
WO2021097618A1
Binocular vision and IMU-based underwater scene three-dimensional reconstruction method, and device
WO2024045632A1