A method and system for three-dimensional object recognition visualization based on augmented reality environments
Patent Information
- Application Number
- CN202610691176.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]综上,相关技术中存在的技术问题有待得到改善
[0008]The embodiments of this application include at least the following beneficial effects: First, the embodiments of this application acquire object image data acquired by the image acquisition unit, inertial measurement data acquired by the inertial measurement unit, and user interaction information. Then, based on the object image data, a visual tracking stability score is calculated, and based on the inertial measurement data, the cumulative deviation value of the inertial measurement unit is calculated. Then, based on the user interaction information, the user's operation intention is identified. If the user's operation intention is to observe only, the posture of the 3D object model is adjusted based on the visual tracking stability score and the cumulative deviation value. If the user's operation intention is to perform fine interaction, the rendering parameters of the 3D object model in the interaction area are adjusted based on the visual tracking stability score. Thus, the 3D object model can be adjusted accordingly for different user operation intentions by combining the visual tracking stability score and the cumulative deviation value of the inertial measurement unit.
Smart Images

Figure CN122597723A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of augmented reality technology, and in particular to a method and system for recognizing and visualizing three-dimensional objects based on augmented reality environments. Background Technology
[0002] Augmented Reality (AR) technology is widely used in the field of cultural heritage, providing visitors with an immersive interactive experience by overlaying three-dimensional models of cultural relics onto physical entities. These systems rely on the camera, image sensor, and main processor of AR devices to capture environmental images, identify cultural relics, and calculate spatial position and pose, achieving precise alignment between virtual models and physical cultural relics.
[0003] However, in practical applications, the fusion of virtual and real elements often faces challenges. When visitors slowly move around an artifact over a long period, the camera pose estimation program heavily relies on the artifact's own feature points. However, many artifacts, such as bronzes, have uniformly oxidized surfaces, subtle relief undulations, and lack high-contrast textures or sharp corners. This makes it difficult for standard feature point extraction programs to obtain sufficiently stable and evenly distributed feature points, leading to decreased feature point extraction and matching performance and persistently low matching confidence. Simultaneously, prolonged slow movement causes the inertial measurement unit within the AR device to accumulate drift errors. Due to the low confidence in visual feature point matching, the camera pose estimation program cannot effectively correct these errors using visual information. These errors gradually accumulate, eventually causing a slight misalignment in the virtual 3D model, which should perfectly align with the real artifact, resulting in low visualization accuracy. When visitors perform detailed interactions such as virtual touch and detail highlighting, the sense of misalignment is noticeable, leading to a poor user experience.
[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention
[0005] The main objective of this invention is to propose a three-dimensional object recognition and visualization method and system based on augmented reality environment, which can combine visual tracking stability score and cumulative deviation value of inertial measurement unit to adjust the corresponding three-dimensional object model according to different user operation intentions.
[0006] On one hand, embodiments of the present invention provide a method for recognizing and visualizing three-dimensional objects based on augmented reality environments, including the following steps: The system acquires object image data collected by the image acquisition unit, inertial measurement data collected by the inertial measurement unit, and user interaction information. The object image data includes multiple image frames, the inertial measurement data includes angular velocity and linear acceleration, and the user interaction information includes screen touch information, gesture operation information, and gaze focus information. Both the image acquisition unit and the inertial measurement unit are integrated into the augmented reality device. Calculate the visual tracking stability score based on the object image data; Based on the inertial measurement data, calculate the cumulative deviation value of the inertial measurement unit; Based on the user interaction information, identify the user's operational intent; If the user's intention is to observe only, the pose of the 3D object model is adjusted according to the visual tracking stability score and the cumulative deviation value to smooth the presentation of the 3D object model, which is used to represent the virtual projection of a real 3D object in the augmented reality environment. If the user's intention is a fine-grained interaction, then the rendering parameters of the 3D object model in the interaction area are adjusted according to the visual tracking stability score so that the 3D object model remains aligned with the real 3D object.
[0007] On the other hand, embodiments of the present invention provide a three-dimensional object recognition and visualization system based on augmented reality environment, including: The data acquisition module is used to acquire object image data acquired by the image acquisition unit, inertial measurement data acquired by the inertial measurement unit, and user interaction information. The object image data includes multiple image frames, the inertial measurement data includes angular velocity and linear acceleration, and the user interaction information includes screen touch information, gesture operation information, and gaze focus information. Both the image acquisition unit and the inertial measurement unit are integrated into the augmented reality device. The stability score calculation module is used to calculate the visual tracking stability score based on the object image data; The cumulative deviation calculation module is used to calculate the cumulative deviation value of the inertial measurement unit based on the inertial measurement data. The user operation intent recognition module is used to recognize the user operation intent based on the user interaction information. The model pose adjustment module is used to adjust the pose of the three-dimensional object model according to the visual tracking stability score and the cumulative deviation value if the user's operation intention is to observe only, so as to smooth the presentation of the three-dimensional object model. The three-dimensional object model is used to represent the virtual projection of a real three-dimensional object in the augmented reality environment. The rendering parameter adjustment module is used to adjust the rendering parameters of the 3D object model in the interaction area according to the visual tracking stability score if the user's operation intention is a fine interaction, so as to keep the 3D object model aligned with the real 3D object.
[0008] The embodiments of this application include at least the following beneficial effects: First, the embodiments of this application acquire object image data acquired by the image acquisition unit, inertial measurement data acquired by the inertial measurement unit, and user interaction information. Then, based on the object image data, a visual tracking stability score is calculated, and based on the inertial measurement data, the cumulative deviation value of the inertial measurement unit is calculated. Then, based on the user interaction information, the user's operation intention is identified. If the user's operation intention is to observe only, the posture of the 3D object model is adjusted based on the visual tracking stability score and the cumulative deviation value. If the user's operation intention is to perform fine interaction, the rendering parameters of the 3D object model in the interaction area are adjusted based on the visual tracking stability score. Thus, the 3D object model can be adjusted accordingly for different user operation intentions by combining the visual tracking stability score and the cumulative deviation value of the inertial measurement unit.
[0009] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the description and the drawings. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0011] Figure 1 This is a flowchart illustrating a three-dimensional object recognition and visualization method based on augmented reality environment, according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a three-dimensional object recognition and visualization system based on augmented reality environment according to an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.
[0013] Among related technologies, augmented reality (AR) technology, with its unique advantage of blending virtual and real elements, creates an immersive interactive experience for visitors in the field of cultural heritage. Its system typically consists of devices such as handheld tablets or head-mounted displays, with core components including cameras, image sensors, and a main processor. In operation, the device first captures environmental images, identifies the target artifact, accurately calculates its spatial position and posture relative to the artifact, and then renders and aligns the virtual 3D model onto the physical artifact, presenting it to the user. However, in actual scenarios of cultural heritage digitization and virtual display, this blending of virtual and real elements encounters several specific challenges.
[0014] Take, for example, visitors using AR devices to observe bronze artifacts in an exhibition. While bronze artifacts possess rich textures and decorative patterns on a macroscopic scale, upon close inspection, their main surface colors are uniform and their reflective properties are subtle. For an AR system to achieve stable overlay between a virtual model and the real artifact, it relies on camera pose estimation, which in turn depends heavily on visual tracking calculations. When visitors slowly move around the bronze artifact at close range for an extended period, the parallax of background environmental feature points changes little, while the positions of feature points on the bronze artifact's surface change drastically. Based on the principle of motion parallax, the calculation program will place greater emphasis on relying on the artifact's own feature points to infer the camera's three-dimensional motion, as these feature points provide richer positional information.
[0015] However, the visual characteristics of bronze artifacts pose a significant challenge to feature point extraction. At close range, their intricate decorations are mostly subtle relief undulations rather than high-contrast textures or sharp corners, and most surfaces exhibit large areas of uniform oxidation. Standard feature point extraction algorithms, such as scale-invariant feature transformation algorithms, struggle to detect a sufficient number of stable and evenly distributed strong feature points on the artifact's surface; most are weak, transient points. The diffuse reflection characteristics of bronze surfaces further reduce the uniqueness of feature points. As camera pose estimation becomes increasingly reliant on the artifact's own feature points, and the artifact's surface struggles to provide stable, high-quality feature points, the image feature point extraction and matching performance of AR systems drops significantly. In a large number of image frames, the number of identified feature points on the artifact is limited and unevenly distributed, resulting in low feature point matching confidence. The visual tracking component of the pose estimation system cannot accurately update the camera pose when input information is insufficient and unreliable.
[0016] Meanwhile, as visitors move slowly for extended periods, the inertial measurement units (IMUs) of AR devices (including gyroscopes and accelerometers) accumulate drift errors. IMUs estimate attitude and position through integrated measurements; small deviations or noise can lead to error accumulation over time. When the confidence level of visual feature point matching is low, the camera attitude estimation program lacks reliable visual information to correct these drifts, and can only rely more on the deteriorating IMU data, creating a vicious cycle.
[0017] These factors work together to cause a continuous accumulation of asymptotic errors in camera pose estimation, ultimately leading to a misalignment of the virtual 3D model on the surface of the real artifact, resulting in low visualization accuracy. This misalignment is particularly noticeable when visitors engage in detailed interactions, as the virtual model may lag behind or deviate from its alignment, severely undermining the core objective of seamless AR overlay and impacting the user experience.
[0018] The embodiments of this application will be explained in detail below with reference to the accompanying drawings: Figure 1 This is an optional flowchart of a three-dimensional object recognition and visualization method based on augmented reality environment provided in the embodiments of this application. Figure 1The method may include, but is not limited to, steps S101 to S106.
[0019] Step S101: Acquire object image data acquired by the image acquisition unit, inertial measurement data acquired by the inertial measurement unit, and user interaction information. The object image data includes multiple image frames, the inertial measurement data includes angular velocity and linear acceleration, and the user interaction information includes screen touch information, gesture operation information, and gaze focus information. Both the image acquisition unit and the inertial measurement unit are integrated in the augmented reality device. Step S102: Calculate the visual tracking stability score based on the object image data; Step S103: Calculate the cumulative deviation value of the inertial measurement unit based on the inertial measurement data; Step S104: Identify the user's operation intent based on user interaction information; Step S105: If the user's operation intention is to observe only, then adjust the pose of the 3D object model according to the visual tracking stability score and cumulative deviation value to smooth the presentation of the 3D object model. The 3D object model is used to represent the virtual projection of the real 3D object in the augmented reality environment. Step S106: If the user's intention is fine interaction, adjust the rendering parameters of the 3D object model in the interaction area according to the visual tracking stability score so that the 3D object model is aligned with the real 3D object.
[0020] Steps S101 to S106 as shown in the embodiments of this application can combine the visual tracking stability score and the cumulative deviation value of the inertial measurement unit to adjust the corresponding three-dimensional object model according to different user operation intentions.
[0021] In some embodiments, steps S101-S106 may first acquire object image data collected by the image acquisition unit, inertial measurement data collected by the inertial measurement unit, and user interaction information. The object image data includes multiple image frames, and the user interaction information includes screen touch information, gesture operation information, and gaze focus information. The image acquisition unit can continuously acquire object image data, which consists of multiple consecutive image frames, to capture visual information in the real world. For example, when a user wears AR glasses to observe an artifact, the built-in camera of the AR glasses continuously captures images of the artifact. Simultaneously, the inertial measurement unit can acquire inertial measurement data, including the device's angular velocity and linear acceleration, which reflect the device's motion state. For example, the gyroscope and accelerometer in the AR glasses record the speed and direction of the user's head rotation and movement in real time. Furthermore, the system also acquires user interaction information, including user touch information on the screen, gesture operation information recognized by the gesture recognition module, and gaze focus information acquired through eye-tracking technology. For example, a user may express their intention by touching the side of the AR glasses, making a pinch gesture, or staring at a specific area of the artifact for a long time. The image acquisition unit and the inertial measurement unit are both integrated into the augmented reality device, ensuring synchronous data acquisition and processing.
[0022] As we can understand, augmented reality devices refer to portable devices that integrate image acquisition units and inertial measurement units, such as AR glasses, AR helmets, or smartphones. These devices can perceive their surroundings in real time and overlay virtual information. The image acquisition unit typically refers to the device's built-in camera, used to capture image data of objects in the real world; this data exists in the form of continuous image frames. The inertial measurement unit typically includes a gyroscope and accelerometer, used to measure the device's angular velocity and linear acceleration to provide motion information. User interaction information refers to the data generated when users interact with virtual content through augmented reality devices. Screen touch information records the user's touch behavior on the device screen, gesture operation information records the user's hand gestures performed through the gesture recognition module, and gaze focus information obtains the area the user is looking at through eye-tracking technology.
[0023] Then, a visual tracking stability score is calculated based on the object image data. The visual tracking stability score is used to evaluate the reliability of visual tracking. For example, it can be calculated by analyzing the number and distribution of feature points in the image, as well as the matching stability between consecutive frames. If the image has abundant feature points and stable matching, the visual tracking stability score will be high; conversely, if the feature points are sparse or the matching is unstable, the score will be low. Understandably, the visual tracking stability score is a metric for measuring the performance of a visual tracking system, reflecting the stability and accuracy of matching feature points extracted from the object image data with the 3D object model.
[0024] Then, based on the inertial measurement data, the cumulative deviation value of the inertial measurement unit (IMU) is calculated. The IMU will drift during long-term operation, leading to errors in attitude estimation. For example, these deviations can be estimated and accumulated by integrating the raw data from the gyroscope and accelerometer, combined with a certain filtering algorithm (such as Kalman filtering). It can be understood that the cumulative deviation value of the IMU refers to the accumulated error in attitude and position estimation caused by factors such as sensor noise and integration errors during long-term operation.
[0025] Based on user interaction information, the system identifies the user's operational intent. It can analyze screen touch information, gesture operation information, and gaze focus information to determine whether the user is merely observing an object or desires more refined interaction. For example, if a user stares at an object for a long time without obvious touch or gesture operations, it can be determined that they are merely observing; if the user performs a pinch gesture or touches a specific area of the screen, it can be determined that they are engaging in refined interaction.
[0026] If the user's intention is observation only, the pose of the 3D object model is adjusted based on the visual tracking stability score and cumulative deviation value to smooth the presentation of the 3D object model, which represents the virtual projection of a real 3D object in the augmented reality environment. In observation-only mode, the user is primarily concerned with the overall presentation of the model and is sensitive to minor jitters or jumps. Therefore, a smoothing algorithm can be used to adjust the pose of the 3D object model, taking into account both the stability of visual tracking and the cumulative deviation of the inertial measurement unit. For example, when the visual tracking stability score is low and the cumulative deviation value of the inertial measurement unit is high, the smoothing coefficient can be increased to make the changes in the model's pose more gradual and continuous, thereby reducing visual discomfort.
[0027] If the user's intention is fine-grained interaction, the rendering parameters of the 3D object model in the interaction area are adjusted based on the visual tracking stability score to ensure alignment between the 3D object model and the real 3D object. In fine-grained interaction mode, users have extremely high requirements for the alignment accuracy between the model and the real object; any slight misalignment can affect the interactive experience. Therefore, emphasis can be placed on the accuracy of visual tracking. For example, when the visual tracking stability score is high, the model can be aligned with the real object with high precision as much as possible, even if the model's pose may undergo slight and rapid adjustments. When the visual tracking stability score is low, the rendering parameters of the interaction area can be adjusted, such as increasing transparency or highlighting, to alert the user that the current alignment may be off and guide the user to make corrections.
[0028] In this embodiment, when the user is simply observing, the system smoothly adjusts the pose of the 3D object model based on the visual tracking stability score and cumulative deviation value. This means that even in cases of unstable visual tracking or drift in the inertial measurement unit (IMU), the system can use weighted fusion and other methods to make the virtual model presentation smoother and more stable, avoiding the discomfort caused by model jitter or slow drift. For example, when the visual tracking stability score is low, the system relies more on IMU data for pose estimation, but also considers the cumulative deviation value of the IMU and uses smoothing filtering to avoid drastic model jumps. When the user performs detailed interactions, the system adjusts the rendering parameters of the 3D object model in the interactive area based on the visual tracking stability score to ensure high-precision alignment between the model and the real object. For example, when the user tries to "touch" a detail on an artifact, if the visual tracking stability score is high, the system ensures that the virtual model is perfectly aligned with the real object in that area, providing accurate interactive feedback. Even if the visual tracking stability score is low, the system can guide the user to perform more accurate interactions by adjusting rendering parameters (such as highlighting the interactive area or providing visual guide lines), rather than simply misaligning the model.
[0029] Through the above technical solution, this embodiment effectively solves the problem of inaccurate alignment between 3D object models and real objects in augmented reality environments by introducing the recognition of user operation intentions and adopting differentiated model adjustment strategies for different intentions. This method not only improves the visual comfort of users in the observation-only mode, but also ensures a high-precision virtual-real fusion experience in the refined interaction mode. It is significantly superior to the single or passive model adjustment schemes in existing technologies, providing a more stable, accurate, and immersive solution for the application of augmented reality technology in fields such as cultural heritage.
[0030] In some embodiments, the step S102, calculating the visual tracking stability score based on the object image data, may include, but is not limited to, the following steps: Step S201: Extract visual feature points from the object image data. The visual feature points include edge midpoints and corner points. Step S202: Match the visual feature points with the 3D object model and calculate the number of interior points. Interior points are used to represent visual feature points that match the preset reference points in the 3D object model. Step S203: Match the visual feature points with the 3D object model and calculate the average positional deviation of the non-interior points. The non-interior points are used to represent visual feature points that do not match the preset reference points in the 3D object model. Step S204: Calculate the proportion of interior points based on the number of interior points and the number of preset reference points in the 3D object model; Step S205: Calculate the visual tracking stability score based on the ratio of inliers and the average positional deviation of non-inliers.
[0031] In some embodiments, visual feature points can be extracted from object image data first. Visual feature points are pixels in an image that have unique and recognizable patterns. They exhibit good stability across different image frames and are commonly used for image matching, object recognition, and pose estimation. Visual feature points include edge midpoints and corner points. Edge midpoints are typically located in the middle of the object's contour, while corner points are located at the intersections of the object's contour. Both are easily detected and tracked due to their high local contrast and structural information in the image, serving as a reliable benchmark for subsequent feature matching.
[0032] Then, the visual feature points are matched with the 3D object model, and the number of inliers is calculated. Inliers represent visual feature points that match preset reference points in the 3D object model. Inliers typically represent the correct projection of the 3D object model into the current image frame and are fundamental for accurate pose estimation and tracking.
[0033] The visual feature points are then matched with the 3D object model, and the average positional deviation of the non-interior points is calculated. Non-interior points represent visual feature points that do not match a preset reference point in the 3D object model. These non-interior points may be caused by background noise, occlusion, lighting variations, or incorrect matching. By distinguishing between interior and non-interior points, the accuracy of visual tracking can be effectively evaluated. The average spatial distance between all non-interior points and their nearest preset reference point can be calculated as the average positional deviation of the non-interior points. A smaller average positional deviation may indicate that even with some mismatches, these mismatched points are relatively close to the object model, and the tracking quality is acceptable; while a larger deviation may indicate serious tracking problems.
[0034] The inlier ratio is calculated based on the number of inliers and the number of preset reference points in the 3D object model. The ratio of the number of inliers to the total number of preset reference points in the 3D object model can be used as the inlier ratio. A higher inlier ratio generally indicates better accuracy and stability in visual tracking.
[0035] Finally, based on the ratio of interior points and the average positional deviation of non-interior points, a visual tracking stability score is calculated to quantify the reliability and accuracy of the current visual tracking. This aims to provide a comprehensive evaluation so that the rendering of the 3D object model can be adjusted accordingly.
[0036] This embodiment evaluates the stability of visual tracking from multiple dimensions through refined analysis of object image data. First, by extracting visual feature points from the object image data and matching them with a preset 3D object model, it is possible to distinguish between inliers aligned with the model and non-inliers that are not aligned. The number of inliers directly reflects the degree of fit between the model and the actual object in the image, while the average positional deviation of non-inliers quantifies the dispersion of the mismatched portion. By calculating the proportion of inliers, the proportion of effective matches can be intuitively understood. Finally, these indicators are combined to form a visual tracking stability score, which comprehensively reflects the accuracy and reliability of the current visual tracking. This multi-dimensional evaluation mechanism enables the system to more accurately determine the tracking status of the 3D object model in the current augmented reality environment, providing a reliable basis for subsequent pose adjustments or rendering parameter adjustments.
[0037] Through the above technical solution, this embodiment can accurately and comprehensively quantify the stability of visual tracking of 3D object models in augmented reality environments. By comprehensively considering the proportion of interior points and the average positional deviation of non-interior points, this embodiment can more accurately reflect the actual quality of current visual tracking and avoid misjudgments that may be caused by a single indicator. Therefore, it provides a more reliable and refined basis for subsequent adjustments to model pose or rendering parameters based on user operation intentions, thereby improving the user experience and system robustness of 3D object recognition and visualization in augmented reality environments.
[0038] In some embodiments, step S205, calculating the visual tracking stability score based on the inlier ratio and the average positional deviation of non-inlier points, may include, but is not limited to, the following steps: Extract descriptor feature vectors corresponding to visual feature points on different image frames. The descriptor feature vectors include local brightness, contrast, and gradient direction. Calculate the feature vector distance between different image frames based on the descriptor feature vectors corresponding to visual feature points on different image frames; The number of stable frames is determined based on the feature vector distance between different image frames. Stable frames are used to indicate that the feature vector distance between adjacent image frames is less than a preset distance threshold. The temporal stability score is calculated based on the total number of frames and the number of stable frames. The visual tracking stability score is calculated based on the temporal stability score, the proportion of inliers, and the average positional deviation of non-inliers.
[0039] In some embodiments, relying solely on feature matching results from a single frame may not adequately reflect the continuity and stability of visual tracking. For example, when an object is moved rapidly in an augmented reality environment or tracked against a complex background, even with high-quality feature matching in a single frame, drastic changes in feature point positions or descriptors between adjacent frames may indicate jitter or discontinuity in the tracking, affecting the smoothness or alignment accuracy of the 3D object model. Failure to address these issues could lead to unstable jitter or misalignment of virtual objects in the augmented reality experience.
[0040] To achieve this, we can first extract descriptor feature vectors corresponding to visual feature points on different image frames. Descriptor feature vectors refer to feature information describing the local image region surrounding the visual feature point, which can include local brightness, contrast, and gradient direction. These features can capture information such as image texture, edges, and corners, thereby enabling the identification and matching of the same visual feature point across different image frames. For example, algorithms such as SIFT (Scale Invariant Feature Transform), SURF (Speed-Up Robust Feature Transform), or ORB (Oriented Fast and Rotational BRIEF) can be used to extract these descriptor feature vectors.
[0041] Then, based on the descriptor feature vectors corresponding to visual feature points in different image frames, the feature vector distance between different image frames is calculated. The feature vector distance between different image frames is an indicator that measures the similarity of the descriptor feature vectors of the same visual feature point in adjacent image frames. The smaller the distance, the smaller the change in the feature point over time, and the more stable the tracking. For example, Euclidean distance, Hamming distance, or cosine similarity can be used to calculate the feature vector distance.
[0042] The number of stable frames is then determined based on the feature vector distance between different image frames. A stable frame indicates that the feature vector distance between adjacent image frames is less than a preset distance threshold. This preset distance threshold can be set according to the actual application scenario and the required tracking stability, for example, through experiments or empirical values. When the feature vector distance is less than this threshold, the visual tracking of that frame is considered relatively stable. A temporal stability score is then calculated based on the total number of frames and the number of stable frames. For example, the ratio of the number of stable frames to the total number of frames can be used as the temporal stability score. This score reflects the overall stability of visual tracking over a period of time.
[0043] Finally, the visual tracking stability score is calculated based on the temporal stability score, the proportion of inliers, and the average positional deviation of non-inliers. For example, these three indicators can be combined using a weighted average, where the weights can be adjusted according to the contribution of each indicator to visual tracking stability.
[0044] This embodiment effectively addresses the problem that relying solely on single-frame feature matching results may not adequately reflect the continuity and stability of visual tracking by introducing the analysis of descriptor feature vectors corresponding to visual feature points across different image frames. Specifically, by extracting descriptor feature vectors and calculating the feature vector distance between different image frames, the degree of change of visual feature points in the time dimension can be quantified. When the feature vector distance is less than a preset distance threshold, it indicates that the change in visual feature points between adjacent frames is small, thus identifying them as stable frames. By counting the number of stable frames and combining it with the total number of frames to calculate a temporal stability score, this embodiment can evaluate the smoothness and continuity of visual tracking from a time series perspective. Finally, by combining this temporal stability score with the proportion of inliers and the average positional deviation of non-inliers based on single-frame matching, a more comprehensive and robust visual tracking stability score can be provided, thus more accurately reflecting the actual tracking quality of 3D object models in augmented reality environments.
[0045] To illustrate this technical solution more clearly, a specific example is used below. Assume an augmented reality device is tracking a real 3D object and continuously acquiring image frames. First, from the continuously acquired image frame sequence, for example, from frame N to frame N+K, descriptor feature vectors are extracted for visual feature points (such as edge midpoints and corner points) in each frame. These descriptor feature vectors can be calculated using the ORB algorithm, which includes information such as local brightness, contrast, and gradient direction. Next, for each pair of adjacent image frames (e.g., frame N and frame N+1), the Hamming distance between the descriptor feature vectors of the same visual feature points in the two frames is calculated. Then, a preset distance threshold is set; for example, if the Hamming distance is less than this threshold, the feature point is considered stable between the two frames. The number of pairs of adjacent frames whose feature point matches satisfy the stability condition throughout the entire frame sequence (N to N+K) is counted, thus determining the number of stable frames.
[0046] Subsequently, a temporal stability score is calculated based on the number of stable frames and the total number of frames K. For example, if K = 100 frames, and 90 frames are determined to be stable frames, the temporal stability score is 0.9. Finally, this temporal stability score is weighted and fused with the static stability score calculated based on the inlier ratio and the average positional deviation of non-inliers. For example, the weight of the temporal stability score can be set to 0.4, and the combined weight of the inlier ratio and the average positional deviation of non-inliers can be set to 0.6, thus obtaining the final visual tracking stability score. In this way, even in cases where the matching quality of some individual frames is high but the overall tracking has jitter, the temporal stability score can effectively reduce the final visual tracking stability score, prompting the system to adopt a more conservative pose adjustment or rendering strategy to maintain the smoothness of the user experience.
[0047] Through the above technical solution, this embodiment overcomes the limitation of inaccurate visual tracking stability evaluation caused by relying solely on single-frame feature matching results and ignoring temporal continuity. By introducing a temporal stability score, this application can more precisely capture jitter, drift, or discontinuity in the visual tracking process, so that the calculated visual tracking stability score not only reflects the matching quality of the current frame but also embodies the smoothness and consistency of tracking in the temporal dimension. This comprehensive evaluation method enables augmented reality devices to more accurately determine the current state of visual tracking, thereby making more reasonable and timely responses when adjusting the pose of 3D object models or rendering parameters in subsequent processes. This significantly improves the smoothness and alignment accuracy of 3D object models in augmented reality environments, providing users with a more stable and immersive augmented reality experience.
[0048] In some embodiments, in step S103, calculating the cumulative deviation value of the inertial measurement unit based on the inertial measurement data may include, but is not limited to, the following steps: Acquire user observation time; Based on the visual tracking stability score, adjust the number of diagonal elements in the covariance matrix, which is a parameter in the extended Kalman filter; Based on the user observation time and the number of diagonal elements of the covariance matrix, an extended Kalman filter is used to analyze inertial measurement data and object image data to calculate the cumulative deviation value.
[0049] In some embodiments, the user's observation time can be acquired first. User observation time refers to the duration for which a user observes a specific 3D object in an augmented reality environment. This duration can be monitored and recorded in real time by the augmented reality device through eye tracking, head pose estimation, and other methods. The purpose is to provide a temporal reference for subsequent calculations of cumulative deviation values, especially since errors in the inertial measurement unit may gradually accumulate when a user observes the same object for an extended period.
[0050] Then, based on the visual tracking stability score, the number of diagonal elements in the covariance matrix is adjusted. The covariance matrix is a key parameter in the extended Kalman filter, used to describe the uncertainty of state estimation. When the visual tracking stability score is high, it indicates that the visual data is relatively reliable. In this case, the number of diagonal elements in the covariance matrix can be appropriately reduced, making the filter more inclined to trust the visual data. Conversely, when the visual tracking stability score is low, the number of diagonal elements is increased to reduce the weight of visual data in the fusion, thereby relying more on inertial measurement data and avoiding excessive errors introduced due to the instability of visual data.
[0051] Then, based on the user observation time and the number of diagonal elements in the covariance matrix, an extended Kalman filter is used to analyze the inertial measurement data and object image data to calculate the cumulative bias value. The extended Kalman filter is a nonlinear filter widely used in state estimation and data fusion. It can be used to fuse inertial measurement data (such as angular velocity and linear acceleration) and object image data to estimate the attitude and position of augmented reality devices, calculating the cumulative bias value of the inertial measurement unit in the process. By using the user observation time and the number of diagonal elements in the covariance matrix adjusted according to the visual tracking stability score as input parameters, the extended Kalman filter can process sensor data more intelligently and robustly, resulting in a more accurate cumulative bias value.
[0052] This embodiment incorporates user observation time, enabling the system to perceive the user's level of attention and duration to 3D objects. Simultaneously, by dynamically adjusting the number of diagonal elements in the covariance matrix of the extended Kalman filter based on the visual tracking stability score, an adaptive allocation of trust between visual and inertial measurement data is achieved. A high visual tracking stability score indicates reliable visual information, and the filter will place more emphasis on visual data to correct inertial measurement unit (IMU) drift. Conversely, a low visual tracking stability score reduces the weight of visual data, relying more on IMU data and combining it with user observation time to more accurately estimate and compensate for the IMU's cumulative bias. This adaptive fusion strategy allows the calculation of cumulative bias values to fully utilize the advantages of different sensors and effectively cope with complex and ever-changing augmented reality environments.
[0053] Through the above technical solution, this embodiment can more accurately calculate the cumulative deviation value of the inertial measurement unit (IMU). Specifically, by considering the user's observation time, the system can better understand the cumulative characteristics of IMU errors, especially in long-term observation scenarios. Furthermore, dynamically adjusting the covariance matrix parameters in the extended Kalman filter based on the visual tracking stability score makes the sensor data fusion process more robust and adaptive, effectively avoiding the negative impact of unstable data from a single sensor on the calculation of the cumulative deviation value. Therefore, the calculated cumulative deviation value is more accurate, providing a more reliable basis for subsequent attitude adjustment or rendering parameter adjustment of the 3D object model, thereby improving the overall stability and user experience of 3D object recognition and visualization in augmented reality environments.
[0054] In some embodiments, step S104, identifying the user's operational intent based on user interaction information, may include, but is not limited to, the following steps: Extract touch duration and touch location from screen touch information; Extract the current gesture from the gesture operation information. The current gesture includes pinch, grab, swipe, or no gesture. Extract gaze duration from gaze focus information; If the touch duration exceeds the preset touch threshold and the touch location is within a 3D object model, the touch determination result is determined to be a touch; otherwise, the touch determination result is determined to be no touch. If the current gesture is no gesture, the gesture determination result is determined to be no interactive gesture; otherwise, the gesture determination result is determined to be an interactive gesture. If the touch detection result indicates the presence of touch, the gesture detection result indicates the presence of interactive gestures, or the gaze duration exceeds the preset gaze threshold, then the user's operation intention is determined to be fine interaction; otherwise, the user's operation intention is determined to be observation only.
[0055] In some embodiments, touch duration and touch location can be extracted from screen touch information first. The user's touch behavior on the screen can be monitored in real time using touch sensors or screen input modules integrated into the augmented reality device. Touch duration refers to the length of time the user's finger or stylus contacts the screen, and touch location refers to the specific position of the user's touch point in the screen coordinate system. This information is used to initially determine whether the user has directly interacted with the virtual object on the screen.
[0056] The current gesture is then extracted from the gesture operation information, including pinch, grasp, swipe, or no gesture. The user's hand movements can be captured and analyzed using depth sensors, cameras, or gesture recognition modules integrated into the augmented reality device. The current gesture can include various predefined gesture types; for example, pinch typically represents grasping or scaling a virtual object, grasping may represent moving an object, swiping may be used to rotate or pan a view, and no gesture indicates that the user's hand is stationary or in a non-interactive state. This gesture information provides clues about the user's non-contact interaction with the augmented reality environment.
[0057] Next, gaze duration is extracted from the gaze focus information. This can be achieved by using eye-tracking sensors integrated into augmented reality devices to monitor the user's eye movement trajectory and gaze point. Gaze duration refers to the length of time a user's gaze lingers on a specific area or object. A longer gaze duration usually indicates a higher level of user attention to that area or object, potentially foreshadowing an interaction intent.
[0058] If the touch duration exceeds the preset touch threshold and the touch location is within a 3D object model, the touch determination result is determined to be a touch. Otherwise, the touch determination result is determined to be no touch. The preset touch threshold can be set according to the actual application scenario and user experience requirements, such as 0.1 seconds or 0.2 seconds. If the current gesture is no gesture, the gesture determination result is determined to be no interactive gesture. Otherwise, regardless of the type of interactive gesture detected (such as pinch, grab, swipe), the gesture determination result is determined to be an interactive gesture.
[0059] If the touch detection result indicates the presence of a touch, the gesture detection result indicates the presence of an interactive gesture, or the gaze duration exceeds a preset gaze threshold, then the user's intention is determined to be fine-grained interaction. Otherwise, the user's intention is determined to be observation only. The final user intention identification can be achieved by combining the touch detection result, gesture detection result, and gaze duration. If the touch detection result indicates the presence of a touch, the gesture detection result indicates the presence of an interactive gesture, or the gaze duration exceeds the preset gaze threshold, then the user's current intention is comprehensively determined to be "fine-grained interaction." The preset gaze threshold can be determined based on experience or user research results, such as 1 second or 2 seconds. Conversely, if none of the above conditions are met—that is, no touch, no interactive gesture, and the gaze duration does not reach the threshold—the user's intention is determined to be "observation only."
[0060] This embodiment integrates multi-source user interaction information such as screen touch, gesture operation, and gaze focus, and performs comprehensive judgment based on preset logical rules to more comprehensively and accurately identify the user's true operational intentions in the augmented reality environment. This multimodal fusion recognition method avoids the misjudgment or insufficient information problems that may be caused by a single interaction mode, thus providing a reliable basis for subsequent 3D object model presentation or rendering parameter adjustment. For example, when the user performs fine interaction, the system can provide more accurate alignment and rendering effects; while when the user only observes, the system can prioritize ensuring the smooth presentation of the model.
[0061] Through the above technical solution, this embodiment can accurately identify the user's operational intent, effectively distinguishing whether the user wishes to perform detailed interactive operations or simply observe the 3D object model in the augmented reality environment. This accurate intent recognition helps the system intelligently adjust the presentation of the 3D object model according to the user's actual needs, thereby significantly improving the user's interactive experience and immersion in the augmented reality environment.
[0062] In some embodiments, step S105, adjusting the pose of the 3D object model based on the visual tracking stability score and cumulative deviation value, may include, but is not limited to, the following steps: Obtain the raw pose data at the current moment and the smoothed pose data at the previous moment; The first smoothing coefficient is determined based on the visual tracking stability score and the cumulative deviation value; Based on the first smoothing coefficient, the original pose data at the current moment and the smoothed pose data at the previous moment are weighted and summed to calculate the smoothed pose data at the current moment. Adjust the pose of the 3D object model based on the smooth pose data at the current moment.
[0063] In some embodiments, simple pose adjustments may not adequately guarantee the smoothness and stability of the 3D object model, especially when there are fluctuations in visual tracking or inertial measurement data, which may cause the virtual object to jitter or have unnatural transitions.
[0064] To achieve this, we can first acquire the raw pose data and the smoothed pose data from the previous moment. The raw pose data at the current moment refers to the pose information of the 3D object model directly acquired by the augmented reality device through sensors (such as visual sensors and inertial sensors) at the current moment, without any smoothing processing. It may contain some noise or instantaneous fluctuations. The smoothed pose data from the previous moment refers to the pose information of the 3D object model that has been processed and smoothed by this method before the current moment. It represents the stable pose at historical moments.
[0065] Then, based on the visual tracking stability score and the cumulative deviation value, a first smoothing coefficient is determined. The first smoothing coefficient is a weight value between 0 and 1, used to balance the influence of the current original pose data and the smoothed pose data from the previous moment on the final smoothed pose data. For example, when the visual tracking stability score is high and the cumulative deviation value is low, indicating good tracking quality, the first smoothing coefficient can be increased, giving more weight to the original pose data and making the model more responsive; conversely, when the tracking quality is poor, the first smoothing coefficient can be decreased, giving more weight to the smoothed pose data from the previous moment to maintain model stability and reduce jitter.
[0066] Then, based on the first smoothing coefficient, a weighted sum is performed on the original pose data at the current moment and the smoothed pose data at the previous moment to calculate the smoothed pose data at the current moment. Various methods can be used, such as linear interpolation or exponential smoothing. For example, smoothed pose data = (1 - first smoothing coefficient) * smoothed pose data at the previous moment + first smoothing coefficient * original pose data at the current moment. This weighted summation effectively fuses historical smoothing information and current real-time information, thereby generating smoothed pose data that reflects the latest state and possesses good stability.
[0067] Finally, based on the smooth pose data at the current moment, the pose of the 3D object model is adjusted to make its presentation in the augmented reality environment smoother and more natural, reducing model jitter caused by sensor noise or environmental changes.
[0068] Through the above technical solution, this embodiment can adaptively adjust the pose smoothness of the 3D object model based on the real-time status of visual tracking and inertial measurement. By introducing historical smooth pose data and dynamically adjusted first smoothing coefficient, this embodiment effectively avoids jittering or unnatural movement of the 3D object model caused by sensor data fluctuations or tracking instability in augmented reality environments. Therefore, it significantly improves the smoothness and stability of the 3D object model in observation-only mode, providing users with a more immersive and comfortable augmented reality experience and reducing visual fatigue.
[0069] In some embodiments, step S106, adjusting the rendering parameters of the 3D object model in the interactive area based on the visual tracking stability score, may include, but is not limited to, the following steps: Step S301: If the visual tracking stability score is less than the preset stability threshold, then obtain the user's interaction points on the 3D object model, the global pose information of the augmented reality device relative to the real 3D object, and the high-precision 3D model data of the real 3D object that has been stored in advance. Step S302: Determine the geometric position of the interaction point on the high-precision 3D model data; Step S303: Adjust the rendering parameters of the 3D object model in the interactive area based on the global pose information and geometric position.
[0070] In some embodiments, if the visual tracking stability score is low, i.e. the visual tracking effect is poor, relying solely on the visual tracking stability score to adjust rendering parameters may not guarantee high-precision alignment between the 3D object model and the real 3D object. Especially when the user is performing fine-grained interactions, such inaccurate alignment may affect the user experience and interaction effects.
[0071] Therefore, the visual tracking stability score can be assessed first. If the score is lower than the preset stability threshold, it indicates that the current visual tracking may be unstable or lack accuracy. In this case, more precise information needs to be introduced to assist in adjusting the rendering parameters. This can be achieved by acquiring user interaction points on the 3D object model, the global pose information of the augmented reality device relative to the real 3D object, and pre-stored high-precision 3D model data of the real 3D object. Interaction points refer to specific locations on the 3D object model specified by the user through touch, gestures, or gazing; these points are the focus of the user's fine-grained interaction. Global pose information refers to the position and orientation of the augmented reality device relative to the real 3D object in the world coordinate system, which is usually obtained through the device's own positioning and tracking system. High-precision 3D model data of the real 3D object refers to the precise geometric representation of the real 3D object in the digital world, such as CAD models or high-precision point clouds or mesh models obtained through 3D scanning, which contain precise edge and surface information of the object. The preset stability threshold can be set based on historical visual tracking data and analysis of the visual tracking effect.
[0072] Then, the geometric position of the interaction point on the high-precision 3D model data is determined. The user's interaction points on the virtual 3D object model can be mapped to their precise geometric coordinates on the high-precision 3D model data of the real 3D object through coordinate transformation and model matching. This can be achieved by aligning the virtual 3D object model with the high-precision 3D model and transforming the interaction points from the virtual model coordinate system to the high-precision model coordinate system.
[0073] Then, based on the global pose information and geometric position, the rendering parameters of the 3D object model in the interactive area are adjusted. The global pose information of the augmented reality device can be used, combined with the precise geometric position of the interactive point on the high-precision 3D model, to calculate the precise rendering position and orientation of the virtual 3D object model in the augmented reality environment. For example, a high-precision 3D model can be placed in the augmented reality environment based on the global pose information, and the rendering parameters of the virtual 3D object model in the interactive area (such as position, rotation, scaling, or local deformation) can be fine-tuned based on the geometric position of the interactive point on the high-precision model to ensure high-precision alignment between the virtual model and the real object in that area.
[0074] To illustrate this technical solution more clearly, a specific example is used below. Assume a user is using an augmented reality device to observe and finely interact with a real-world car model. When the user touches a specific area of the car model (e.g., a door handle) for fine interaction, the system continuously calculates the visual tracking stability score. If the visual tracking stability score falls below a preset stability threshold (e.g., due to changes in lighting or rapid movement causing loss of visual feature points), the system acquires the location of the door handle touched by the user as the interaction point. Simultaneously, it acquires the global pose information relative to the real car obtained by the augmented reality device through SLAM or other positioning technologies, and retrieves pre-stored high-precision 3D model data of the car. Next, the system maps the user's touch interaction point to the precise geometric position of the door handle on the high-precision 3D model data. Finally, based on the global pose information of the augmented reality device and the geometric position of the door handle on the high-precision 3D model, the system precisely adjusts the rendering parameters of the virtual car model in the door handle area, ensuring that the virtual door handle maintains high-precision alignment with the real door handle in the augmented reality environment. Even if visual tracking is temporarily unstable, the user can still experience a smooth and precise interactive experience.
[0075] Through the above technical solution, this embodiment can effectively solve the problem of insufficient alignment accuracy between the 3D object model and the real 3D object when visual tracking is unstable. By combining global pose information and high-precision 3D model data, the system can provide more stable and accurate rendering parameter adjustments. Especially when users perform fine interactions, it significantly improves the interactive experience and realism in augmented reality environments, ensuring a high degree of integration between virtual content and the real world.
[0076] In some embodiments, in step S303, adjusting the rendering parameters of the 3D object model in the interaction area based on global pose information and geometric position may include, but is not limited to, the following steps: Get the user's sliding operation state on the surface of the 3D object model; If the swipe operation is a fast swipe, then the swipe duration is monitored as the duration of visual feature point loss within the interactive area. Based on the duration of visual feature point loss, the second smoothing coefficient used to align global pose information with the 3D object model is adjusted to extend the smoothing process cycle. Based on the second smoothing coefficient, global pose information, and geometric position, adjust the rendering parameters of the 3D object model in the interactive area.
[0077] In some embodiments, when a user performs dynamic interactive operations such as rapid swiping, visual feature points within the interactive area may be temporarily lost, leading to a decrease in visual tracking stability. If the system immediately adjusts the rendering parameters quickly at this time, the rendering of the 3D object model may become jittery or inconsistent, affecting the user experience.
[0078] To achieve this, the system can first acquire the user's swiping operation state on the surface of the 3D object model. The swiping operation state refers to the movement pattern of the user's hand or fingers as recognized by the system when the user interacts with the virtual 3D object model in an augmented reality environment via touchscreen, gesture recognition, or other interactive methods. For example, this can be determined by analyzing the touch trajectory, speed, and acceleration in screen touch information, or the gesture type and amplitude in gesture operation information.
[0079] If the swipe operation is a rapid swipe, the swipe duration is monitored as the duration of visual feature point loss within the interaction area. When the system detects this rapid swipe state, it monitors the duration of the swipe operation. This swipe duration is considered the duration of visual feature point loss within the interaction area because rapid swipes often lead to image blurring, feature points quickly moving out of the field of view, or being occluded, making it difficult for the visual tracking system to stably extract and match feature points.
[0080] The system then adjusts the second smoothing coefficient, used to align global pose information with the 3D object model, based on the duration of the lost visual feature points, to extend the smoothing processing cycle. The second smoothing coefficient is a weighted factor that controls the smoothing degree of the 3D object model's rendering parameters. When the duration of lost visual feature points is long, the system increases the second smoothing coefficient accordingly, thus extending the smoothing processing cycle. Extending the smoothing processing cycle means that for a period of time, the system will rely more on historical pose information and rendering parameters, rather than immediately responding to potentially unstable current visual tracking data, to avoid model jitter caused by brief feature point losses.
[0081] Finally, based on the second smoothing coefficient, global pose information, and geometric position, the rendering parameters of the 3D object model in the interactive area are adjusted. This adjustment combines the second smoothing coefficient, the global pose information of the augmented reality device relative to the real 3D object, and the geometric position of the interactive point on pre-stored high-precision 3D model data. This method makes the changes in rendering parameters smoother and more gradual, maintaining visual alignment between the 3D object model and the real 3D object even when visual feature points are temporarily lost.
[0082] To illustrate this technical solution more clearly, a specific example is used below. Suppose a user is observing a virtual 3D car model in an augmented reality environment and attempts to rotate the model by rapidly swiping across the screen. When the user's finger rapidly swipes across the screen, the system first acquires the state of the user's swipe operation on the surface of the 3D object model and identifies it as a rapid swipe. Since rapid swiping may cause visual feature points (such as edges and textures) in local areas of the car model's surface to blur or disappear from view for a short period, the system monitors the duration of this swipe operation and uses it as the duration of lost visual feature points within the interactive area. For example, if the rapid swipe lasts for 0.5 seconds, the system dynamically adjusts the second smoothing coefficient used to align global pose information with the 3D object model based on this duration. Specifically, the system might adjust the second smoothing coefficient from the default value of 0.7 to 0.9 to extend the smoothing processing cycle of the rendering parameters. This means that in the next 0.5 seconds, when adjusting the rendering parameters of the car model, the system will focus more on maintaining the model's stability rather than immediately responding to potentially unstable visual tracking data. Ultimately, the system adjusts the rendering parameters of the car model in the interactive area based on the adjusted second smoothing coefficient, the global pose information of the augmented reality device relative to the real environment, and the geometric position of the user interaction point on the pre-stored high-precision 3D model data of the car. In this way, even if the user's rapid swiping causes a brief loss of visual feature points, the rotation and presentation of the car model in the augmented reality environment can remain smooth and stable, avoiding sudden jitter or jumps in the model, thus providing users with a more natural and immersive interactive experience.
[0083] Through the above technical solution, this embodiment effectively solves the problem of uneven adjustment of rendering parameters and model jitter in 3D object models caused by the temporary loss of visual feature points during dynamic interactive operations such as rapid swiping. By monitoring the swiping operation status and associating it with the duration of feature point loss, and adaptively adjusting the smoothing coefficient accordingly, this embodiment can extend the smoothing processing cycle. Thus, even when visual tracking stability temporarily decreases, it can still maintain stable alignment and smooth visual presentation between the 3D object model and the real 3D object, significantly improving the user interaction experience in augmented reality environments.
[0084] In some embodiments, after extracting visual feature points from the object image data in step S201, the method may further include, but is not limited to, the following steps: Acquire ambient lighting information, global pose information of augmented reality devices relative to real 3D objects, and high-precision 3D model data of pre-stored real 3D objects; Based on ambient lighting information, brightness distribution analysis is performed on the interactive area to identify areas of high light reflection; In the high-reflection area, the object edge contour is fitted based on global pose information and high-precision 3D model data to obtain the edge contour fitting result; Based on the edge contour fitting results, the visual feature points are corrected.
[0085] In some embodiments, when there are areas of high gloss reflection on the object's surface, the local brightness of the image changes drastically, which may lead to inaccurate or unstable extraction of visual feature points (such as edge midpoints and corner points). This inaccuracy will affect the accuracy of subsequent visual tracking stability score calculations, which may result in poor alignment between the 3D object model and the real 3D object in the augmented reality environment, especially in scenarios requiring fine interaction.
[0086] To achieve this, we can first acquire ambient lighting information, the global pose information of the augmented reality device relative to the real 3D object, and pre-stored high-precision 3D model data of the real 3D object. Ambient lighting information refers to parameters such as the light intensity, direction, and color of the environment in which the augmented reality device is located, which can be obtained through the device's built-in light sensor or by analyzing image data. Global pose information refers to the position and orientation of the augmented reality device relative to the real 3D object in the world coordinate system, which is usually obtained through the device's own positioning and tracking system. High-precision 3D model data of the real 3D object refers to the precise geometric representation of the real 3D object in the digital world, such as a CAD model or a high-density point cloud or mesh model obtained through 3D scanning, which contains precise edge and surface information of the object.
[0087] Then, based on ambient lighting information, brightness distribution analysis is performed on the interactive area to identify specular reflection regions. The purpose is to identify areas in the image that are too bright or too dim, potentially leading to inaccurate feature point extraction. Specular reflection regions refer to high-brightness areas on an object's surface caused by specular reflection. The pixel values in these areas are typically saturated or near saturation, resulting in the loss of texture information and thus affecting the accurate extraction of visual feature points.
[0088] In the high-brightness reflection region, object edge contours are fitted based on global pose information and high-precision 3D model data to obtain the edge contour fitting results. The aim is to accurately reconstruct or estimate the true edge contour of the object in the high-brightness region by combining prior geometric knowledge of the object (from high-precision 3D model data) and the device's current viewing angle (from global pose information). The edge contour fitting results are a more accurate representation of the object's edges obtained through this combined analysis.
[0089] Finally, the visual feature points are corrected based on the edge contour fitting results. The initially extracted visual feature points can be compared and adjusted with the more accurate edge contour obtained by fitting. For example, if a preliminarily extracted visual feature point is located in a highlight area and deviates from the fitted edge contour, its position can be adjusted to the closest fitted edge contour, or the inaccurate feature point can be directly removed, and a more reliable feature point can be regenerated from the fitted edge contour.
[0090] To illustrate this technical solution more clearly, a specific example is used below. Suppose a user wearing an augmented reality (AR) device observes a smooth, metallic-looking car model. When the ambient light is strong and the light shines on the car model's surface from a specific angle, distinct areas of high-gloss reflection will form on the model. First, the AR device acquires image data of the object from its image acquisition unit and initially extracts visual feature points. Simultaneously, the device acquires current ambient lighting information; for example, it detects high ambient brightness through a light sensor and analyzes the brightness distribution of the image to identify the bright spots on the car model's surface caused by high-gloss reflection.
[0091] Next, the system acquires the global pose information of the augmented reality device relative to the car model, as well as pre-stored high-precision 3D model data of the car model. Within the identified specular reflection area, the system uses the global pose information to project the high-precision 3D model onto the current image frame, and combines this with the image data to fit the edge contour of the car model in that specular area. For example, a method combining model edge projection and image gradient information can be used to accurately determine the actual edge position of the car model within the specular area.
[0092] Finally, the initially extracted visual feature points are corrected based on the more accurate edge contour results obtained from the fitting. For example, if a corner point initially extracted is located within a highlight area and its position deviates significantly from the fitted edge contour, the system will adjust it to the closest fitted edge contour or directly replace it with a more reliable corner point recalculated from the fitted edge contour. In this way, high-quality visual feature points can be obtained even under challenging conditions of high-reflectivity, thereby ensuring the accuracy and stability of subsequent 3D object recognition and visualization processes.
[0093] Through the above technical solution, this embodiment significantly improves the accuracy and robustness of visual feature point extraction in complex lighting environments, especially in areas with high gloss reflection. This embodiment combines ambient lighting information, global pose information, and high-precision 3D model data to achieve accurate fitting of the object's edge contour within the glossy area, and uses this as a basis to correct the visual feature points. This correction mechanism effectively avoids the negative impact of gloss reflection on feature point extraction, ensuring that the extracted visual feature points more accurately reflect the geometric features of the real object. This improves the reliability of subsequent visual tracking stability score calculations, thereby ensuring that the 3D object model maintains more precise alignment with the real 3D object in the augmented reality environment, providing a more stable and accurate visual experience, especially during detailed user interactions.
[0094] The beneficial effects of implementing the embodiments of the present invention include: the embodiments of this application first acquire object image data acquired by the image acquisition unit, inertial measurement data acquired by the inertial measurement unit, and user interaction information. Then, based on the object image data, a visual tracking stability score is calculated, and based on the inertial measurement data, the cumulative deviation value of the inertial measurement unit is calculated. Then, based on the user interaction information, the user's operation intention is identified. If the user's operation intention is to observe only, the posture of the three-dimensional object model is adjusted based on the visual tracking stability score and the cumulative deviation value. If the user's operation intention is to perform fine interaction, the rendering parameters of the three-dimensional object model in the interaction area are adjusted based on the visual tracking stability score. Thus, the three-dimensional object model can be adjusted according to different user operation intentions by combining the visual tracking stability score and the cumulative deviation value of the inertial measurement unit.
[0095] like Figure 2 As shown, this embodiment of the invention also provides a three-dimensional object recognition and visualization system based on augmented reality environment, including: The data acquisition module 401 is used to acquire object image data acquired by the image acquisition unit, inertial measurement data acquired by the inertial measurement unit, and user interaction information. The object image data includes multiple image frames, the inertial measurement data includes angular velocity and linear acceleration, and the user interaction information includes screen touch information, gesture operation information, and gaze focus information. Both the image acquisition unit and the inertial measurement unit are integrated into the augmented reality device. The stability score calculation module 402 is used to calculate the visual tracking stability score based on the object image data. The cumulative deviation calculation module 403 is used to calculate the cumulative deviation value of the inertial measurement unit based on the inertial measurement data. The user operation intent recognition module 404 is used to recognize the user's operation intent based on user interaction information; The model pose adjustment module 405 is used to adjust the pose of the 3D object model based on the visual tracking stability score and cumulative deviation value if the user's operation intention is to only observe, so as to smooth the presentation of the 3D object model. The 3D object model is used to represent the virtual projection of the real 3D object in the augmented reality environment. The rendering parameter adjustment module 406 is used to adjust the rendering parameters of the 3D object model in the interaction area according to the visual tracking stability score if the user's operation intention is fine interaction, so as to keep the 3D object model aligned with the real 3D object.
[0096] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0097] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
Claims
1. A method for recognizing and visualizing three-dimensional objects based on augmented reality environments, characterized in that, Includes the following steps: The system acquires object image data collected by the image acquisition unit, inertial measurement data collected by the inertial measurement unit, and user interaction information. The object image data includes multiple image frames, the inertial measurement data includes angular velocity and linear acceleration, and the user interaction information includes screen touch information, gesture operation information, and gaze focus information. Both the image acquisition unit and the inertial measurement unit are integrated into the augmented reality device. Calculate the visual tracking stability score based on the object image data; Based on the inertial measurement data, calculate the cumulative deviation value of the inertial measurement unit; Based on the user interaction information, identify the user's operational intent; If the user's intention is to observe only, the pose of the 3D object model is adjusted according to the visual tracking stability score and the cumulative deviation value to smooth the presentation of the 3D object model, which is used to represent the virtual projection of a real 3D object in the augmented reality environment. If the user's intention is a fine-grained interaction, then the rendering parameters of the 3D object model in the interaction area are adjusted according to the visual tracking stability score so that the 3D object model remains aligned with the real 3D object.
2. The method according to claim 1, characterized in that, The step of calculating the visual tracking stability score based on the object image data includes: Visual feature points are extracted from the object image data, including edge midpoints and corner points; The visual feature points are matched with the three-dimensional object model, and the number of interior points is calculated. The interior points are used to represent visual feature points that match preset reference points in the three-dimensional object model. The visual feature points are matched with the three-dimensional object model, and the average positional deviation of the non-interior points is calculated. The non-interior points are used to represent visual feature points that do not match the preset reference points in the three-dimensional object model. Calculate the proportion of the interior points based on the number of interior points and the number of preset reference points in the three-dimensional object model; The visual tracking stability score is calculated based on the inlier ratio and the average positional deviation of the non-inlier points.
3. The method according to claim 2, characterized in that, The step of calculating the visual tracking stability score based on the inlier ratio and the average positional deviation of the non-inlier points includes: Extract the descriptor feature vectors corresponding to the visual feature points on different image frames. The descriptor feature vectors include local brightness, contrast, and gradient direction. Based on the descriptor feature vectors corresponding to the visual feature points in different image frames, the feature vector distance between different image frames is calculated. The number of stable frames is determined based on the feature vector distance between the different image frames. The stable frames are used to indicate that the feature vector distance between adjacent image frames is less than a preset distance threshold. The temporal stability score is calculated based on the total number of frames and the number of stable frames. The visual tracking stability score is calculated based on the temporal stability score, the proportion of inliers, and the average positional deviation of non-inliers.
4. The method according to claim 1, characterized in that, The step of calculating the cumulative deviation value of the inertial measurement unit based on the inertial measurement data includes: Acquire user observation time; Based on the visual tracking stability score, adjust the number of diagonal elements in the covariance matrix, where the covariance matrix is a parameter in the extended Kalman filter; Based on the user observation time and the number of diagonal elements of the covariance matrix, an extended Kalman filter is used to analyze the inertial measurement data and the object image data to calculate the cumulative deviation value.
5. The method according to claim 1, characterized in that, The step of identifying the user's operation intent based on the user interaction information includes: Extract the touch duration and touch location from the screen touch information; Extract the current gesture from the gesture operation information, where the current gesture includes pinching, grasping, swiping, or no gesture; Extract gaze duration from the gaze focus information; If the touch duration exceeds a preset touch threshold and the touch location is within a 3D object model, then the touch determination result is determined to be that a touch exists; otherwise, the touch determination result is determined to be that no touch occurs. If the current gesture is no gesture, the gesture determination result is determined to be no interactive gesture; otherwise, the gesture determination result is determined to be an interactive gesture. If the touch determination result indicates that a touch exists, the gesture determination result indicates that an interactive gesture exists, or the gaze duration exceeds a preset gaze threshold, then the user's operation intention is determined to be fine interaction; otherwise, the user's operation intention is determined to be observation only.
6. The method according to claim 1, characterized in that, The step of adjusting the pose of the 3D object model based on the visual tracking stability score and the cumulative deviation value includes: Obtain the raw pose data at the current moment and the smoothed pose data at the previous moment; A first smoothing coefficient is determined based on the visual tracking stability score and the cumulative deviation value; Based on the first smoothing coefficient, the original pose data at the current moment and the smoothed pose data at the previous moment are weighted and summed to calculate the smoothed pose data at the current moment. The pose of the 3D object model is adjusted based on the smooth pose data at the current moment.
7. The method according to claim 1, characterized in that, The step of adjusting the rendering parameters of the 3D object model in the interactive area based on the visual tracking stability score includes: If the visual tracking stability score is less than the preset stability threshold, then the user's interaction points on the 3D object model, the global pose information of the augmented reality device relative to the real 3D object, and the high-precision 3D model data of the real 3D object stored in advance are obtained. Determine the geometric position of the interaction point on the high-precision 3D model data; Based on the global pose information and the geometric position, adjust the rendering parameters of the 3D object model in the interactive area.
8. The method according to claim 7, characterized in that, The step of adjusting the rendering parameters of the 3D object model in the interactive area based on the global pose information and the geometric position includes: Get the user's sliding operation state on the surface of the 3D object model; If the sliding operation state is a fast sliding, then the sliding duration is monitored as the duration of visual feature point loss in the interactive area; Based on the duration of the loss of the visual feature points, the second smoothing coefficient used to align the global pose information with the 3D object model is adjusted to extend the smoothing process cycle. Based on the second smoothing coefficient, the global pose information, and the geometric position, adjust the rendering parameters of the 3D object model in the interactive area.
9. The method according to claim 2, characterized in that, After extracting visual feature points from the object image data, the method further includes: Acquire ambient lighting information, global pose information of augmented reality devices relative to real 3D objects, and high-precision 3D model data of pre-stored real 3D objects; Based on the ambient lighting information, a brightness distribution analysis is performed on the interactive area to identify areas of high light reflection. In the high-reflectivity region, the object edge contour is fitted based on the global pose information and the high-precision three-dimensional model data to obtain the edge contour fitting result; The visual feature points are corrected based on the edge contour fitting results.
10. A three-dimensional object recognition and visualization system based on augmented reality environment, characterized in that, include: The data acquisition module is used to acquire object image data acquired by the image acquisition unit, inertial measurement data acquired by the inertial measurement unit, and user interaction information. The object image data includes multiple image frames, the inertial measurement data includes angular velocity and linear acceleration, and the user interaction information includes screen touch information, gesture operation information, and gaze focus information. Both the image acquisition unit and the inertial measurement unit are integrated into the augmented reality device. The stability score calculation module is used to calculate the visual tracking stability score based on the object image data; The cumulative deviation calculation module is used to calculate the cumulative deviation value of the inertial measurement unit based on the inertial measurement data. The user operation intent recognition module is used to recognize the user operation intent based on the user interaction information. The model pose adjustment module is used to adjust the pose of the three-dimensional object model according to the visual tracking stability score and the cumulative deviation value if the user's operation intention is to observe only, so as to smooth the presentation of the three-dimensional object model. The three-dimensional object model is used to represent the virtual projection of a real three-dimensional object in the augmented reality environment. The rendering parameter adjustment module is used to adjust the rendering parameters of the 3D object model in the interaction area according to the visual tracking stability score if the user's operation intention is a fine interaction, so as to keep the 3D object model aligned with the real 3D object.