A smart control method and system for AR devices
By establishing a geometric mapping model and head pose tracking, the camera orientation of the AR device is dynamically controlled, solving the problem of deviation between the user's line of sight and the camera, and improving the accurate matching of AR content and user experience.
Patent Information
- Application Number
- CN202411902524.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-23
AI Technical Summary
When users wear existing AR devices, the difference in facial structure and the complexity of the environment can cause a deviation between the line of sight and the camera, affecting the accuracy of AR content and user experience, especially exacerbating dizziness and discomfort in dynamic scenes.
By acquiring users' facial biometric data and camera data, a geometric mapping model is established. Combined with head pose tracking and salient region analysis, the camera shooting direction of the AR device is dynamically controlled to achieve precise matching between the user's gaze and the device's camera.
It improves the alignment accuracy between AR content and the user's visual focus, enhances immersion and interactive experience, adapts to the needs of different users and environments, and reduces line-of-sight matching errors and interference.
Smart Images

Figure CN119847330B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of AR device technology, and in particular to an intelligent control method and system for AR devices. Background Technology
[0002] In the rapid development of augmented reality (AR) technology, precise matching between the user's gaze and the AR device's camera is crucial for enhancing the user experience. However, in practical applications, this technical challenge has yet to be perfectly resolved. When users wear AR devices, individual differences in facial structure, such as interpupillary distance, eye size, and nose bridge height, lead to significant discrepancies between the image captured by the device's camera and the user's actual gaze. This discrepancy is already sufficient to affect the user experience in static environments, and in dynamic scenes, the user's head movements and eye rotations further exacerbate this problem.
[0003] Traditional AR devices typically rely on a fixed camera position and orientation to capture images and render AR content. However, this fixed approach cannot track and compensate for the deviation between the user's gaze and the camera in real time. As the user's head and eyes move, the matching accuracy between the AR content and the user's gaze gradually decreases, leading to inaccurate overlap between the user's visual focus and the content when observing AR content. This not only affects the user's perception and understanding of AR content but may also cause discomfort symptoms such as dizziness and headaches, significantly reducing user comfort and device usability.
[0004] Furthermore, the complexity of the usage environment poses a greater challenge to the accuracy of eye-matching. Under different lighting conditions, the quality and contrast of camera images are affected, thus impacting the accuracy of eye-matching. For example, in low-light environments, the camera may fail to capture a clear image, causing the eye-matching algorithm to fail; while under direct sunlight, the reduced image contrast also increases the difficulty of eye-matching. In complex environments such as outdoors, rapid background changes and the presence of interference further severely test the stability and robustness of the eye-matching algorithm.
[0005] To address these issues, existing AR devices and technologies have attempted various methods to improve the accuracy of gaze matching. However, these methods often have limitations and cannot fully adapt to the needs of different users and usage environments. Therefore, achieving precise matching between the user's gaze and the AR device's camera, thereby improving the presentation accuracy of AR content and the user experience, remains a critical problem that urgently needs to be solved in the current development of AR technology. Summary of the Invention
[0006] The purpose of this invention is to provide an intelligent control method and system for AR devices, which achieves precise alignment between AR content and the user's visual focus, greatly improving the immersiveness and interactive experience of AR applications, thereby solving at least one of the aforementioned problems in the prior art.
[0007] In a first aspect, the present invention provides an intelligent control method for an AR device, the method specifically comprising:
[0008] Acquire facial biometric data of at least one user and camera data of an AR device, and establish a geometric mapping model between the user's gaze and the device's camera based on the facial biometric data and the camera data;
[0009] At least one user eye image is acquired. Based on the user eye image, the coordinates of the user's pupil center position and the coordinates of the corneal reflection point position are determined. The coordinates of the pupil center position and the coordinates of the corneal reflection point position are input into the geometric mapping model for matching to obtain the user's gaze direction information. The user's gaze direction information includes the user's gaze movement trajectory and the user's fixation point.
[0010] The image brightness value distribution in the region near the user's gaze point is obtained, and several candidate salient regions are determined based on the image brightness value distribution. The user's attention focus is determined by comparing the candidate salient regions with a preset salient region template.
[0011] A head pose tracking algorithm is used to acquire the user's head pose information in real time. Based on the head pose information, the user's gaze direction information, and the user's attention focus, the camera shooting direction of the AR device is dynamically controlled.
[0012] Secondly, the present invention provides an intelligent control system for an AR device, the system specifically comprising:
[0013] The first control module is used to acquire facial biometric data of at least one user and camera data of an AR device, and to establish a geometric mapping model between the user's gaze and the device's camera based on the facial biometric data and the camera data.
[0014] The second control module is used to acquire at least one user eye image, determine the user's pupil center position coordinates and corneal reflection point position coordinates based on the user eye image, input the pupil center position coordinates and corneal reflection point position coordinates into the geometric mapping model for matching, and obtain user gaze direction information, the user gaze direction information including user gaze movement trajectory and user fixation point;
[0015] The third control module is used to acquire the image brightness value distribution in the area near the user's gaze point, determine several candidate salient regions based on the image brightness value distribution, and determine the user's attention focus by comparing the candidate salient regions with a preset salient region template.
[0016] The fourth control module is used to acquire the user's head posture information in real time using a head posture tracking algorithm, and dynamically control the camera shooting direction of the AR device based on the head posture information, the user's gaze direction information and the user's attention focus.
[0017] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements the intelligent control method of the AR device as described in any of the above methods.
[0018] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the intelligent control method for an AR device as described in any of the above methods.
[0019] Compared with the prior art, the present invention has at least one of the following technical effects:
[0020] 1. This invention achieves precise alignment between AR content and the user's visual focus, significantly improving the immersiveness and interactive experience of AR applications.
[0021] 2. This invention not only improves the matching accuracy between AR content and the user's gaze, but also enhances the user's immersion and interactive experience when using AR devices. It can effectively solve the problems of gaze matching in traditional AR devices and promote the further development of AR technology.
[0022] 3. By acquiring the user's facial biometric data and the camera data of the AR device, this invention can establish a precise geometric mapping model between the user's gaze and the device's camera, achieving real-time and accurate matching between the user's gaze and AR content, thereby improving the user's immersion and interactive experience.
[0023] 4. By analyzing the distribution of image brightness values in the area near the user's gaze point, this invention can identify candidate salient regions and compare them with a preset salient region template to determine the user's focus of attention, which helps AR devices to present content that the user is interested in more intelligently.
[0024] 5. This invention acquires the user's head posture information in real time, and combines it with the user's gaze direction information and attention focus to dynamically control the camera shooting direction of the AR device, further enhancing the presentation effect of AR content and improving the user experience.
[0025] 6. By establishing the mapping relationship data between the gaze direction and the camera imaging plane coordinates, and using the support vector machine algorithm for modeling and training, this invention can construct a high-precision geometric mapping model between the user's gaze and the device's camera, thereby improving the accuracy and stability of gaze tracking.
[0026] 7. This invention performs grayscale processing and histogram equalization on the user's eye image, and uses the Hough circle detection algorithm to quickly and accurately determine the coordinates of the user's pupil center position and the corneal reflection point position. This results in a clearer and more standardized user eye image, which helps to improve the accuracy of determining the coordinates of the pupil center position and the corneal reflection point position.
[0027] 8. This invention acquires real-time ambient light intensity data near the AR device and pupil size change data of the user, and performs correlation analysis to build a pupil size change prediction model. This helps the AR device adjust retinal imaging data in real time according to environmental changes and user pupil changes, thereby improving imaging quality and user experience.
[0028] 9. This invention uses Gaussian mixture model and K-means algorithm to filter and cluster gaze regions, which can construct a gaze category tree model with hierarchical relationships, helping to more accurately predict users' future attention shift trends.
[0029] 10. This invention calculates the probability of user attention shift between different gaze categories based on gaze category tree model and hidden Markov model, which can generate attention prediction list, providing strong support for intelligent rendering and presentation of AR content, and helping to improve user interaction experience and satisfaction. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 This is a flowchart illustrating an intelligent control method for an AR device according to an embodiment of the present invention;
[0032] Figure 2 This is a schematic diagram of the structure of an intelligent control system for an AR device according to an embodiment of the present invention;
[0033] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0034] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0035] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0036] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0037] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0038] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0039] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0040] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating an embodiment of the intelligent control method for an AR device disclosed in this invention is shown below, detailed in the following description:
[0041] S101, acquire facial biometric data of at least one user and camera data of the AR device, and establish a geometric mapping model between the user's gaze and the device's camera based on the facial biometric data and the camera data.
[0042] In this embodiment, high-resolution facial scanning devices, such as 3D facial scanners, are used to scan the faces of several users. These devices can capture detailed geometric information of the face, including the three-dimensional coordinates of key feature points such as facial contours, eyes, nose, and mouth. Furthermore, deep learning algorithms can be used to process the facial images and extract facial feature descriptors, such as Histogram of Oriented Gradients (HOG) and Local Binary Patterns (LBP).
[0043] AR devices, such as AR glasses or AR helmets, use cameras to capture real-time video streams of users. These cameras have high frame rates and high resolutions, enabling them to capture facial details and subtle eye movements. Through video stream processing and analysis, user eye features, such as pupil position, size, and eye movement patterns, can be extracted.
[0044] The acquired facial biometric and camera data undergo preprocessing, including denoising, image enhancement, and feature extraction. A geometric mapping model is trained using machine learning or deep learning algorithms, such as convolutional neural networks (CNNs). This model takes the user's facial biometric and camera data as input and outputs the geometric mapping relationship between the user's gaze and the device's camera. During training, a large amount of labeled data can be used to optimize the model's parameters and improve its generalization ability. Cross-validation and regularization are used to optimize the model and reduce the risk of overfitting and underfitting. Furthermore, data augmentation techniques, such as rotation, scaling, and translation, can be used to increase data diversity and further improve the model's accuracy.
[0045] By using a trained geometric mapping model, the user's gaze direction can be tracked in real time. When a user wears an AR device, the camera captures an image of the user's face, and the model calculates the user's gaze direction. This information can be used for interactive control in AR applications, such as selecting menu items and moving virtual objects. The geometric mapping model also allows for a more personalized AR experience. For example, based on the user's gaze direction, parameters such as the display position, size, and angle of virtual objects can be adjusted to better suit the user's visual habits and needs.
[0046] In this embodiment, by combining facial biometric data and camera data, the user's gaze direction can be tracked more accurately, reducing errors and interference. Utilizing a geometric mapping model, more natural and fluid AR interactive control can be achieved, enhancing user immersion and engagement.
[0047] S102, acquire at least one user eye image, determine the user's pupil center position coordinates and corneal reflection point position coordinates based on the user eye image, input the pupil center position coordinates and corneal reflection point position coordinates into the geometric mapping model for matching, and obtain user gaze direction information, the user gaze direction information including user gaze movement trajectory and user fixation point.
[0048] In this embodiment, a high-precision camera (such as an infrared camera or a high-resolution visible light camera) is used to capture images of at least one user's eyes. The camera should be placed in a position that can clearly capture details of the user's eyes, such as a built-in camera on AR glasses. The acquired eye images are preprocessed, including noise reduction, contrast enhancement, and grayscale conversion, to improve the accuracy of subsequent processing. In addition, preliminary contours of the pupil and corneal reflective points are extracted using image processing techniques (such as edge detection and morphological processing).
[0049] Ellipse fitting algorithms or Hough transforms are used to accurately fit the pupil region in the preprocessed image, thereby determining the coordinates of the pupil center. The position coordinates of corneal reflection points generated on the cornea by a light source (such as an infrared source) are located using image processing techniques (such as template matching and feature point detection). Corneal reflection points are often used to help determine the direction of vision because their relative position to the pupil center reflects the tilt angle of the gaze.
[0050] A geometric mapping model is pre-trained, capable of calculating the user's gaze direction based on the coordinates of the pupil center and the corneal reflector. This model can be constructed using machine learning algorithms (such as support vector machines or neural networks) or geometric optics principles. The real-time coordinates of the pupil center and the corneal reflector are input into the geometric mapping model, which then calculates the user's gaze direction information, including the trajectory of the gaze (i.e., how the gaze changes over time) and the fixation point (i.e., the spatial location the gaze points to at a given moment).
[0051] In this embodiment, by accurately locating the pupil center and corneal reflection point, and combining this with a geometric mapping model, the accuracy and stability of gaze tracking can be significantly improved.
[0052] S103, obtain the image brightness value distribution in the area near the user's gaze point, determine several candidate salient regions based on the image brightness value distribution, and determine the user's attention focus by comparing the candidate salient regions with a preset salient region template.
[0053] In this embodiment, firstly, the user is placed in a stable lighting environment, and a high-precision camera (such as the camera on AR glasses) is used to capture an image of the area near the user's gaze point. The camera should have high resolution and fast response capabilities to ensure that the captured image is clear and real-time. The captured image is preprocessed, including noise reduction and contrast enhancement, to improve the accuracy of subsequent processing. Then, the image is converted into a grayscale image for brightness value analysis. Using image processing software or algorithms, the brightness value of each pixel in the grayscale image is calculated, and a brightness value distribution map is generated. This map can visually display the brightness changes in the area near the user's gaze point.
[0054] Based on the luminance value distribution map, set one or more luminance thresholds. These thresholds are used to distinguish between salient and insignificant areas. Generally, areas with higher luminance values are more likely to attract the user's attention. Areas with luminance values exceeding the set thresholds are extracted as candidate salient areas. These areas contain important information or objects that the user is currently focusing on.
[0055] Based on application scenarios and user needs, a series of salient region templates are pre-designed and stored. These templates can be regions of specific shapes, colors, or textures, used for comparison with candidate salient regions near the user's gaze point. The extracted candidate salient regions are compared with the pre-set salient region templates to find the best-matching template. The comparison process can be implemented using algorithms such as similarity calculation and feature matching. Based on the comparison results, the salient region template that best matches the area near the user's gaze point is determined and used as the user's current focus of attention. This focus area contains the key information or object that the user is paying attention to.
[0056] In this embodiment, by accurately calculating the distribution of image brightness values and extracting candidate salient regions, and then comparing them with a preset template, the user's current focus of attention can be identified more accurately. In applications such as AR, adjusting the interaction method or displayed content according to changes in the user's focus of attention can provide the user with a more personalized and smooth experience.
[0057] S104, a head posture tracking algorithm is used to obtain the user's head posture information in real time, and the camera shooting direction of the AR device is dynamically controlled based on the head posture information, the user's gaze direction information and the user's attention focus.
[0058] In this embodiment, the AR device includes: a head posture tracking module, which uses technologies such as an optical barcode tracking system, sensor array, infrared camera, or laser scanner to acquire the user's head posture information in real time, including the head rotation angle and tilt degree; a gaze direction detection module, which captures the user's eye image, determines the coordinates of the pupil center position and the corneal reflection point position, and inputs them into a geometric mapping model to acquire the user's gaze direction information in real time; and an attention focus recognition module, which determines candidate saliency regions based on the image brightness value distribution in the area near the user's gaze point, and compares them with a preset saliency region template to determine the user's attention focus.
[0059] When a user wears an AR device, the head pose tracking module captures the user's head movements in real time and converts them into digital signals. These signals are processed to obtain precise head pose information. Simultaneously, the gaze direction detection module captures the user's eye images and calculates the user's gaze direction in real time using image processing algorithms and geometric mapping models. The attention focus recognition module determines the user's focus based on the image brightness distribution in the area near the user's gaze point. This is typically achieved by comparing the similarity between candidate salience regions and preset templates. The AR device's built-in control algorithm dynamically adjusts the camera's shooting direction based on head pose information, gaze direction information, and attention focus. For example, if the user's head tilts to the right, their gaze is directed to the right, and their attention focus is also on the right, the camera will adjust its shooting direction accordingly to the right to ensure the user can clearly see the content they are focusing on.
[0060] In this embodiment, by dynamically adjusting the camera's shooting direction, it is ensured that users can clearly see the content they are interested in, thereby enhancing the interactive experience and immersion of the AR device. Combining head posture tracking and gaze direction detection allows users to control the camera's shooting direction through natural head movements and gaze shifts, enhancing the naturalness and intuitiveness of the interaction.
[0061] In some embodiments, step S101 above, which involves establishing a geometric mapping model between the user's gaze and the device camera based on the facial biometric data and the camera data, specifically includes:
[0062] Based on the spatial relationship between the facial biometric data and the camera data, establish a mapping relationship between the gaze direction and the coordinates of the camera imaging plane;
[0063] Using the mapping relationship data as input, a support vector machine algorithm is used for modeling and training to construct a geometric mapping model between the user's line of sight and the device's camera.
[0064] In this embodiment, facial biometric data, including facial key point coordinates and facial pose angles, is acquired; coordinate system parameters of the camera imaging plane are acquired, and a camera coordinate system is established; based on the spatial position of the facial biometric data in the camera coordinate system, the three-dimensional coordinates of facial feature points relative to the camera are calculated; based on the facial pose angle and the three-dimensional coordinates of the feature points, the vector representation of the gaze direction in the camera coordinate system is calculated; the gaze direction vector is intersected with the camera imaging plane to obtain the two-dimensional coordinates of the gaze on the imaging plane; mathematical methods such as least squares are used to fit the mapping relationship between the gaze direction vector and the imaging plane coordinates to obtain mapping relationship data; the mapping relationship data is used as training samples and input into a support vector machine algorithm for modeling training to obtain a geometric mapping model between the user's gaze and the camera.
[0065] For example, the spatial relationship between facial biometric data and camera data forms the basis for building a user gaze tracking system. This relationship can be represented by the two-dimensional coordinates of facial feature points (such as the corners of the eyes and the tip of the nose) on the camera's imaging plane. For instance, assuming the coordinates of the outer corner of the left eye are detected as (100, 150) pixels and the coordinates of the outer corner of the right eye are (300, 150) pixels, these data constitute the spatial location information of the facial features. Establishing the mapping relationship between the gaze direction and the coordinates of the camera's imaging plane is the next crucial step. This mapping relationship can be modeled using methods such as polynomial functions or neural networks. For example, a quadratic polynomial function can be used to represent the relationship between the coordinates (x, y) of the pupil center point on the camera's imaging plane and the gaze direction angle (θ, φ): x = a1θ² + b1θ + c1φ² + d1φ + e1, y = a2θ² + b2θ + c2φ² + d2φ + e2. Where a1, b1, c1, d1, e1, a2, b2, c2, d2, and e2 are undetermined coefficients. Support Vector Machine (SVM) is a powerful machine learning method suitable for building geometric mapping models between a user's gaze and a device's camera. SVM can handle high-dimensional feature spaces and achieve nonlinear mappings through kernel function tricks. In this example, a Radial Basis Function (RBF) kernel can be used, which can capture complex nonlinear relationships. Input features can include the coordinates of facial feature points, head pose angles, etc., and the output is the angle of the gaze direction or the coordinates of the gaze point on the screen. During training, a large amount of calibration data needs to be collected. For example, the user gazes at different locations on the screen (such as a 9-point grid), while recording facial feature data and the actual gaze location. Assume 1000 sets of sample data are collected, each containing the coordinates of 10 facial feature points and the corresponding gaze direction angle. This data is used to train the SVM model, and the model performance is optimized by adjusting kernel function parameters (such as the γ value of the RBF kernel) and the regularization parameter C. The constructed geometric mapping model enables real-time gaze tracking. When a new facial image is input, the model can quickly predict the user's gaze direction or fixation point on the screen. This machine learning-based gaze tracking method has significant advantages over traditional geometric methods. It can adapt to differences in facial features among different users and is more robust to changes in lighting and subtle head movements. Through continuous learning and model updates, the system can continuously improve its accuracy and adaptability, providing users with a more natural and intuitive interactive experience.
[0066] In some embodiments, step S102 above, which involves determining the user's pupil center coordinates and corneal reflector coordinates based on the user's eye image, and inputting the pupil center coordinates and corneal reflector coordinates into the geometric mapping model for matching to obtain the user's gaze direction information, specifically includes:
[0067] The user's eye image is processed by grayscale conversion and histogram equalization to obtain a standard user eye image;
[0068] The Hough circle detection algorithm is used to analyze the standard user eye image to determine the coordinates of the user's pupil center and the coordinates of the corneal reflection point.
[0069] Calculate the offset between the coordinates of the user's pupil center and the coordinates of the corneal reflection point to obtain the offset vector;
[0070] The offset vector is matched with the geometric mapping model to obtain the user's line-of-sight direction information.
[0071] In this embodiment, a user's eye image is acquired and converted to grayscale to reduce color interference. Histogram equalization is then applied to the grayscale eye image to enhance contrast and clarify details, facilitating subsequent feature extraction and analysis. The preprocessed, standardized eye image is input into a Hough circle detection algorithm to determine the pupil's position by detecting circular features. Based on the pupil position obtained from the Hough circle detection algorithm, a coordinate system is established with the pupil center as the origin to determine the pupil center's coordinates. Corneal reflectance features are extracted from the standardized eye image, and the position of these features is determined by analyzing the high-brightness reflectance points on the corneal surface. The coordinates of these corneal reflectance points are then transformed into a coordinate system with the pupil center as the origin to obtain the coordinates of the corneal reflectance points relative to the pupil center.
[0072] The offset between the pupil center coordinates and the corneal reflector coordinates is calculated to obtain an offset vector, which includes horizontal and vertical offset components. A machine learning algorithm is used to establish a geometric mapping model, modeling the correspondence between eye feature parameters and gaze direction, resulting in a mapping function. Eye feature parameters include pupil center coordinates, corneal reflector coordinates, and the offset vector. The calculated offset vector is input into the established geometric mapping model, and the corresponding user gaze direction information is obtained through the mapping function.
[0073] For example, firstly, the user's eye image is converted to grayscale, reducing computation and highlighting eye structure features. Histogram equalization further enhances image contrast, making the pupil and iris boundaries clearer. Taking a 640x480 pixel eye image as an example, after grayscale conversion, the brightness value of each pixel ranges from 0 to 255. Histogram equalization might stretch the brightness value of dark areas (such as the pupil) from 10-30 to 5-50, while stretching bright areas (such as the sclera) from 200-220 to 180-250, thus increasing the contrast between the pupil and the surrounding area. Next, the Hough circle detection algorithm is used to locate the pupil. This algorithm detects circular structures by accumulating votes in the parameter space. Assume the detected pupil radius is 20 pixels, with center coordinates (320, 240). Meanwhile, corneal reflective points typically appear as small, high-brightness areas, possibly located at (325, 235). Calculating the offset vector between the pupil center and the corneal reflector is a crucial step. In this example, the offset vector is (5, -5) pixels. This vector contains rich information about the gaze direction. For example, when the user looks to the right, the corneal reflector shifts to the left relative to the pupil center, producing an offset vector similar to (-8, 0). Finally, this offset vector is input into a pre-trained geometric mapping model. This model, which may be a multinomial function or a neural network, learns the complex nonlinear relationship between the offset vector and the actual gaze direction. For example, an offset of (5, -5) might correspond to the user's gaze being slightly to the upper right, about 10 degrees, relative to the center of the screen. The advantage of this method lies in its robustness and adaptability. Through grayscale conversion and histogram equalization, the system can adapt to different lighting conditions. The Hough circle detection algorithm can handle partial occlusion, such as ptosis or eyelash interference. The machine learning-based geometric mapping model can adapt to the differences in eye structure among different users, providing a personalized gaze tracking experience. In practical applications, this technology can achieve gaze tracking accuracy to 1-2 degrees.
[0074] In some embodiments, step S102 above, which involves obtaining the image brightness value distribution in the region near the user's gaze point and determining several candidate salient regions based on the image brightness value distribution, specifically includes:
[0075] A first image containing the user's gaze point is obtained, and the first image is subjected to Gaussian filtering for noise reduction to obtain a second image;
[0076] The second image is divided into blocks by using a fixed-size sliding window that slides across the second image with a certain step size to obtain several image blocks.
[0077] For each image block, calculate the arithmetic mean of the brightness values of all pixels within that image block to obtain the average brightness value of that image block.
[0078] Based on the average brightness value of each image block, the Sobel operator is used to calculate the brightness gradient between two adjacent image blocks to obtain the gradient distribution map of the second image;
[0079] The gradient distribution map is binarized by setting pixels with gradient values greater than or equal to a preset gradient value threshold to 1 and pixels with gradient values less than the preset gradient value threshold to 0, thereby obtaining a binarized gradient distribution map.
[0080] Connectivity analysis is performed on the binarized gradient distribution map to extract connected regions with a pixel value of 1, and the geometric center of each connected region is calculated to obtain candidate salient regions.
[0081] In this embodiment, a first image containing the user's gaze point is acquired; the first image is subjected to Gaussian filtering for noise reduction to obtain a second image; a sliding window of a preset size is used to slide across the second image with a preset step size to obtain multiple image blocks; for each image block, the arithmetic mean of the brightness values of all pixels within the image block is calculated to obtain the average brightness value of the image block; based on the average brightness values of two adjacent image blocks, the Sobel operator is used to calculate the brightness gradient between them to obtain the gradient distribution map of the second image. According to a preset gradient threshold, the gradient distribution map is binarized, with pixels whose gradient values are greater than or equal to the threshold set to 1, and pixels whose gradient values are less than the threshold set to 0, resulting in a binarized gradient map. A connected component analysis algorithm is used to extract connected components from the binarized gradient map to obtain connected regions in the image where all pixels have a value of 1. For each extracted connected region, the coordinates of the geometric center point of the connected region are calculated, and the coordinates of the geometric center point are used as the representative point of the connected region. Based on the size, shape, and other characteristics of the connected regions, it is determined whether each connected region meets the conditions for a candidate salient region; if it does, the connected region is identified as a candidate salient region. All candidate salient regions are merged by combining adjacent candidate regions that meet certain conditions, resulting in merged candidate salient regions. A saliency metric is calculated for each merged candidate salient region. Based on the saliency metric values, the candidate regions are sorted to obtain a sorted candidate region list. The top N regions with the highest saliency metric values in the sorted candidate region list are determined as the final salient regions, where N is a preset threshold for the number of salient regions.
[0082] For example, the first image is first subjected to Gaussian filtering for noise reduction to obtain the second image. Gaussian filtering effectively removes high-frequency noise from the image while preserving important edge information. For instance, for a 640x480 pixel eye image, a 5x5 Gaussian kernel with a standard deviation of 1 might be used to smooth out minor noise without excessively blurring the pupil edges. Next, the second image is segmented. This step aims to divide the image into smaller units, facilitating subsequent local feature analysis. Assuming a 32x32 pixel sliding window is used, sliding across the image with a step size of 16 pixels, approximately 600 overlapping image blocks can be obtained. The average brightness value is calculated for each image block as its brightness distribution feature. This method effectively captures local brightness changes, which is particularly important for identifying pupil and iris boundaries. Then, the Sobel operator is used to calculate the brightness gradient between adjacent image blocks, generating a gradient distribution map. The Sobel operator is a commonly used edge detection operator that can effectively identify regions with abrupt changes in brightness in an image. In eye images, the boundaries between the pupil and iris, and between the iris and sclera, often exhibit significant brightness gradients. For example, at the pupil edge, the gradient value might suddenly increase from 10 to 100. Binarizing the gradient distribution map further highlights salient regions. Setting an appropriate gradient value threshold, such as 50, can effectively separate potential regions of interest. Pixels above this threshold are set to 1, and those below are set to 0, forming a binarized gradient distribution map. This step effectively removes background noise while preserving the contours of key structures such as the pupil and iris. Finally, connected component analysis is performed on the binarized gradient distribution map to extract connected regions with a pixel value of 1. These regions are likely to correspond to important features such as the pupil, iris edge, or corneal reflection points. Calculating the geometric center of each connected region yields candidate salient regions. For example, if an approximately circular connected region with an area of about 1000 pixels and a geometric center near the image center is detected, this is likely the pupil region. The advantage of this method lies in its robustness and adaptability. Through multi-step image processing and feature extraction, key eye structures can be accurately located under different lighting conditions and individual differences. For example, even in low-contrast images, gradient analysis and connected component analysis can effectively identify the pupil position.
[0083] In some embodiments, step S103 above, which involves determining the user's attention focus by comparing the candidate salience region with a preset salience region template, specifically includes:
[0084] Collect massive amounts of image data, label salient regions in the massive amounts of image data, and form an image training dataset;
[0085] The image training dataset is used as input, and a convolutional neural network model is used for modeling and training to obtain a salient region detection model.
[0086] The candidate saliency regions are input into the saliency region detection model for matching to obtain the saliency probability value of each candidate saliency region;
[0087] The candidate saliency region with the highest saliency probability value is taken as the user attention concentration region, and the geometric center coordinates of the user attention concentration region are calculated. The geometric center coordinates of the user attention concentration region are taken as the user attention focus.
[0088] In this embodiment, massive amounts of image data are acquired and preprocessed, including image size normalization and noise removal, to obtain preprocessed image data. For the preprocessed image data, salient regions in the images are labeled using manual or semi-automatic annotation methods, resulting in image data with salient region annotations. The image data with salient region annotations is divided into a training set and a test set, where the training set is used for model training and the test set is used for model evaluation. Based on the characteristics of the salient region detection task, the architecture of the convolutional neural network model is designed, including convolutional layers, pooling layers, and fully connected layers, and the model's hyperparameters are determined. The training set is input into the convolutional neural network model, and the model parameters are continuously updated through iterative forward and backward propagation until the model's loss function on the training set converges or reaches a preset number of iterations. The trained convolutional neural network model is used to predict the test set, obtaining the salient region detection results for the test set images, and the model's evaluation metrics on the test set, such as precision and recall, are calculated. If the model's performance on the test set meets the expected requirements, the trained convolutional neural network model is saved as a salient region detection model; otherwise, the model architecture or hyperparameters are adjusted, and the training and testing are repeated until a detection model that meets the requirements is obtained.
[0089] A set of candidate salient regions is obtained and input into a pre-trained salient region detection model. The model extracts and matches features for each candidate region and calculates the salient probability value of each candidate region. Based on the salient probability value, the candidate region with the highest probability value is selected and determined as the user attention focus region. For the determined user attention focus region, a geometric calculation method is used to obtain the geometric center coordinates of the region. It is determined whether the geometric center coordinates are located in the center region of the image. If so, it is taken as the user attention focus; otherwise, the next step is continued. Based on the relative positional relationship between the salient region and the image center, the geometric center coordinates are adjusted by weighted averaging to obtain the corrected position coordinates. The corrected geometric center coordinates are determined as the user attention focus, and the focus position coordinates are output as input for subsequent business processing.
[0090] For example, firstly, a massive amount of image data is collected and salient regions are labeled. These images may include eye images in various scenarios, such as different lighting conditions and when wearing glasses. The labeling process is usually done by professionals who mark key areas such as the pupil and iris edge in the images. For example, in a 640x480 pixel eye image, the pupil region might be labeled as a circular area with a diameter of about 100 pixels. Next, a convolutional neural network model is trained using the labeled image training dataset. For example, classic network architectures such as VGG16 or ResNet50 can be used as a base, and then fine-tuned according to the specific task. During training, the network learns how to extract key features from the input images and ultimately outputs a probability map of salient regions. After training, candidate salient regions are input into the model for matching. The model outputs a saliency probability value for each candidate region, typically ranging from 0 to 1. For example, the true pupil region might receive a high probability value of 0.95, while the background region might only have a low probability value of around 0.1. Finally, the region with the highest saliency probability value is selected as the user's attention focus area. This region is likely to correspond to the pupil or other important eye structures. Calculating the geometric center of this region yields the precise coordinates of the user's attention focus. For example, if the detected high-probability region is a 100x100 pixel square with its top-left corner coordinates (200, 150), then the coordinates of the attention focus are (250, 200).
[0091] In some embodiments, the method further includes the following steps in steps S101-S104:
[0092] Real-time acquisition of ambient light intensity data near the AR device and pupil size change data of the user;
[0093] The ambient light intensity data and the pupil size change data are correlated and analyzed to form the first training dataset;
[0094] Using the first training dataset as input, a machine learning algorithm is used for modeling and training to obtain a pupil size change prediction model.
[0095] The pupil size change prediction model determines the user's pupil size change prediction data, and the retinal imaging data of the AR device is updated in real time using the pupil size change prediction data.
[0096] In this embodiment, real-time data on ambient light intensity and real-time changes in the user's pupil size are acquired. The acquired ambient light intensity and pupil size change data undergo preprocessing, including data cleaning and normalization. Based on the preprocessed data, correlation analysis is used to determine the degree of correlation between the two. If the correlation analysis shows a correlation higher than a preset threshold, the ambient light intensity and pupil size change data are combined to form a training dataset. A supervised learning algorithm, such as a support vector machine or neural network, is used as input to train a pupil size change prediction model. The user's current pupil image data is acquired, pupil size feature parameters are extracted, and input into the pupil size change prediction model for prediction processing, resulting in predicted pupil size changes over a future period. Based on the predicted pupil size changes, the AR device's retinal imaging parameters that need adjustment are determined, including focal length and resolution, generating retinal imaging adjustment parameters. These parameters are then transmitted to the AR device's display module and displayed in real-time in the user's field of vision, completing the dynamic adjustment of the AR retinal imaging.
[0097] For example, the light intensity in an indoor office environment is approximately 500 lux, while direct sunlight can reach 100,000 lux. Changes in pupil size can be captured by a built-in eye-tracking camera, recording real-time changes in pupil diameter. The normal pupil diameter varies between 2 and 8 millimeters. Correlation analysis of these two sets of data forms a training dataset that reveals the relationship between light intensity and pupil size. Generally, the stronger the light, the smaller the pupil, in order to regulate the amount of light entering the eye. For example, in a 500-lux indoor environment, the pupil diameter might be 5 millimeters; while under 10,000-lux outdoor sunlight, the pupil might shrink to 3 millimeters. Using machine learning algorithms to model and train on this data can yield a more accurate model for predicting pupil size changes. Commonly used algorithms include Support Vector Machines (SVM) and Random Forests. The model's input features might include current light intensity, the rate of light change, and the user's age, with the output being the predicted pupil diameter. For example, for a 30-year-old user, when the light intensity rapidly increases from 500 lux to 5000 lux, the model might predict that the pupil diameter will shrink from 5 mm to 3.5 mm within 2 seconds. The advantage of this predictive model is its ability to anticipate pupil changes, allowing for faster adjustments to the AR device's retinal imaging parameters. Traditional methods require waiting for the actual pupil change before adjustments can be made, potentially causing temporary visual discomfort. The predictive model, however, can adjust the image the instant the light intensity changes, significantly improving the user experience. The AR device's retinal imaging data is updated in real-time based on the predicted pupil size change data, primarily involving adjustments to image brightness, contrast, and sharpness. For instance, when the system predicts the pupil is about to constrict, it can preemptively reduce image brightness and increase contrast to ensure clear visibility in bright light. Conversely, when the system predicts the pupil is about to dilate, it can increase brightness and decrease contrast to prevent excessive glare in low light. This prediction-based real-time adjustment not only improves visual comfort but also reduces eye fatigue, which is particularly important for users who use AR devices for extended periods. At the same time, it also provides a foundation for personalized visual experiences. The system can provide image adjustment strategies that are more tailored to individual needs based on each user's unique pupil response to light.
[0098] In some embodiments, the method further includes the following steps in steps S101-S104:
[0099] Obtain the user's historical interaction behavior data, and extract the user's eye movement data during browsing based on the historical interaction behavior data;
[0100] A Gaussian mixture model is used to calculate the probability values of the distribution points in the eye movement data, and a set of gaze regions whose probability values of the distribution points are higher than a preset probability threshold are selected.
[0101] The K-means algorithm is used to cluster all gaze regions in the gaze region set. Euclidean distance is set as the standard for calculating the distance between centroids. If the distance between the centroids of two gaze regions is lower than the preset distance between centroids threshold, they are classified as gaze regions of the same gaze category, thus constructing a gaze category tree model with hierarchical relationship.
[0102] Based on the gaze category tree model, a hidden Markov model is used to calculate the user attention shift probability between different gaze categories and generate an attention prediction list.
[0103] The system acquires real-time user interaction behavior data and performs visual rendering processing on the current AR content of the AR device based on the attention prediction list and the real-time interaction behavior data.
[0104] In this embodiment, historical user interaction data is acquired, and eye-tracking data during browsing is extracted based on this data. A convolutional neural network is used to extract features from the eye-tracking data, resulting in feature vectors. Based on these feature vectors, a Gaussian mixture model is used to calculate the probability values of the distribution points in the eye-tracking data. Eye-tracking data points with probability values higher than a preset probability threshold are selected and grouped into gaze regions. These gaze regions are then clustered to obtain multiple gaze region sets, each containing multiple gaze regions.
[0105] Obtain a set of gaze regions. For each gaze region in the set, extract its centroid coordinates. Based on the centroid coordinates, calculate the centroid distance between any two gaze regions using the Euclidean distance formula. Set a preset centroid distance threshold. By comparing the centroid distance between gaze regions with the threshold, determine whether they belong to the same gaze category. If the centroid distance is less than the preset threshold, the two gaze regions are classified into the same gaze category; if the centroid distance is greater than or equal to the preset threshold, the two gaze regions are classified into different gaze categories. During the clustering process, dynamically adjust the centroid coordinates of each gaze category, using the mean of the centroid coordinates of all gaze regions within the same category as the new centroid coordinates for that category, until the clustering results converge. Based on the clustering results of the gaze regions, construct a gaze category tree model. By analyzing the hierarchical and inclusion relationships between different gaze categories, determine the structure and depth of the category tree. In the gaze category tree model, each gaze category obtained from clustering is used as a node in the tree. Based on the hierarchical relationships between categories, a hierarchical gaze category tree is formed through the links between parent and child nodes.
[0106] For nodes in the gaze category tree model, a Hidden Markov Model (HMM) is used to calculate the state transition probability matrix between adjacent nodes, obtaining the probability of user attention shifting between different gaze categories. Based on the calculated state transition probability matrix, an attention prediction list is generated, containing the gaze categories the user is likely to focus on in the future and their probabilities. Real-time user interaction data collected by the AR device's sensors, including head movement data and gesture interaction data, is acquired. This real-time user interaction data is matched with the attention prediction list to determine the gaze category the user is most likely to focus on, identifying the AR content that needs to be prioritized for rendering. Based on the determined AR content to be prioritized, the rendering parameters of the AR device are dynamically adjusted to perform visual enhancement rendering on the relevant AR content, improving rendering quality and frame rate. The changes in the user's real-time interaction data are continuously tracked, and the rendering strategy of the AR content is dynamically adjusted in real time based on the updates to the attention prediction list, maintaining synchronization between the AR content and the user's attention.
[0107] For example, the system first acquires historical user interaction data, including browsing trajectories, dwell time, and interactive operations within the AR environment. For instance, in an AR museum tour application, the system can record the time a user spends viewing each exhibit and their gaze path. Gaussian mixture models are a commonly used probability density estimation method suitable for analyzing the distribution characteristics of eye-tracking data. In AR scenarios, they can help identify the areas of greatest interest to the user. Suppose that in an AR tour, a user's eye-tracking data shows the highest probability of fixation in the "Mona Lisa" painting area, exceeding a preset threshold of 0.8; this area would then be selected as an important gaze region. K-means clustering can further categorize these gaze regions. In the AR museum example, if a user frequently moves back and forth between the "Mona Lisa" and the nearby "Last Supper," the centroid distance between these two areas might be less than a preset threshold of 50 pixels, thus being categorized into the same gaze category, forming the "Da Vinci works" hierarchy. This hierarchical gaze category tree model helps understand the user's interest structure. Hidden Markov models are used to predict the shift in user attention. In AR guided tours, the system can calculate the probability of a user shifting from the "Leonardo da Vinci" category to the "Impressionist" category. For example, if this probability is 0.6, higher than other categories, the system can predict that the user might be interested in Impressionist works next. Finally, visual rendering is performed based on real-time interactive behavior and attention prediction lists. For instance, when a user is viewing the "Mona Lisa," and the system predicts they might be interested in Impressionist works, the AR device can subtly highlight nearby Monet works at the edge of the user's field of vision, guiding the user's attention naturally. This intelligent rendering based on user behavior and interests can greatly enhance the personalization and smoothness of the AR experience, making information presentation more in line with users' cognitive habits and interests. The integrated application of this series of technologies not only improves the presentation of AR content but also helps content creators better understand user behavior, thereby optimizing AR experience design. For example, by analyzing a large number of users' gaze category tree models, museums may discover a strong attention shift correlation between Leonardo da Vinci and Monet works, and can adjust exhibit layouts or design new themed guided tour routes accordingly. This data-driven approach opens up vast opportunities for innovation in AR applications, transforming it not only into a display technology but also into an intelligent platform that deeply understands and meets user needs.
[0108] Reference Figure 2 An embodiment of the present invention provides an intelligent control system 2 for an AR device, the system 2 specifically comprising:
[0109] The first control module 201 is used to acquire facial biometric data of at least one user and camera data of an AR device, and to establish a geometric mapping model between the user's gaze and the device's camera based on the facial biometric data and the camera data.
[0110] The second control module 202 is used to acquire at least one user eye image, determine the user's pupil center position coordinates and corneal reflection point position coordinates based on the user eye image, input the pupil center position coordinates and corneal reflection point position coordinates into the geometric mapping model for matching, and obtain user gaze direction information, the user gaze direction information including user gaze movement trajectory and user fixation point;
[0111] The third control module 203 is used to acquire the image brightness value distribution in the area near the user's gaze point, determine several candidate salient regions based on the image brightness value distribution, and determine the user's attention focus by comparing the candidate salient regions with a preset salient region template.
[0112] The fourth control module 204 is used to acquire the user's head posture information in real time using a head posture tracking algorithm, and dynamically control the camera shooting direction of the AR device based on the head posture information, the user's gaze direction information and the user's attention focus.
[0113] It is understandable that, such as Figure 1 The content of the intelligent control method embodiment of the AR device shown is applicable to the intelligent control system embodiment of this AR device. The specific functions implemented by the intelligent control system embodiment of this AR device are the same as those shown below. Figure 1 The intelligent control method of the AR device shown is the same as that implemented in this embodiment, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the intelligent control method embodiment of the AR device shown are also the same.
[0114] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0115] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0116] Reference Figure 3 The present invention also provides a computer device 3, including: a memory 302 and a processor 301, and a computer program 303 stored in the memory 302. When the computer program 303 is executed on the processor 301, it implements the intelligent control method of the AR device as described in any of the above methods.
[0117] The computer device 3 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that... Figure 3 The computer device 3 is merely an example and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0118] The processor 301 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0119] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may be an external storage device of the computer device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 3. Furthermore, the memory 302 may include both internal and external storage units of the computer device 3. The memory 302 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 302 can also be used to temporarily store data that has been output or will be output.
[0120] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the intelligent control method for an AR device as described in any of the above methods.
[0121] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0122] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0123] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0124] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A smart control method for an AR device, characterized in that, The method specifically includes: Acquire facial biometric data of at least one user and camera data of an AR device, and establish a geometric mapping model between the user's gaze and the device's camera based on the facial biometric data and the camera data; At least one user eye image is acquired. Based on the user eye image, the coordinates of the user's pupil center position and the coordinates of the corneal reflection point position are determined. The coordinates of the pupil center position and the coordinates of the corneal reflection point position are input into the geometric mapping model for matching to obtain the user's gaze direction information. The user's gaze direction information includes the user's gaze movement trajectory and the user's fixation point. The image brightness value distribution in the region near the user's gaze point is obtained, and several candidate salient regions are determined based on the image brightness value distribution. The user's attention focus is determined by comparing the candidate salient regions with a preset salient region template. A head pose tracking algorithm is used to acquire the user's head pose information in real time. Based on the head pose information, the user's gaze direction information, and the user's attention focus, the camera shooting direction of the AR device is dynamically controlled. Real-time acquisition of ambient light intensity data near the AR device and pupil size change data of the user; The ambient light intensity data and the pupil size change data are correlated and analyzed to form the first training dataset; Using the first training dataset as input, a machine learning algorithm is used for modeling and training to obtain a pupil size change prediction model. The pupil size change prediction model determines the user's pupil size change prediction data, and the retinal imaging data of the AR device is updated in real time using the pupil size change prediction data.
2. The method according to claim 1, characterized in that, The step of establishing a geometric mapping model between the user's gaze and the device's camera based on the facial biometric data and the camera data specifically includes: Based on the spatial relationship between the facial biometric data and the camera data, establish a mapping relationship between the gaze direction and the coordinates of the camera imaging plane; Using the mapping relationship data as input, a support vector machine algorithm is used for modeling and training to construct a geometric mapping model between the user's line of sight and the device's camera.
3. The method according to claim 1, characterized in that, The step of determining the user's pupil center coordinates and corneal reflector coordinates based on the user's eye image, and inputting the pupil center coordinates and corneal reflector coordinates into the geometric mapping model for matching to obtain the user's gaze direction information, specifically includes: The user's eye image is processed by grayscale conversion and histogram equalization to obtain a standard user eye image; The Hough circle detection algorithm is used to analyze the standard user eye image to determine the coordinates of the user's pupil center and the coordinates of the corneal reflection point. Calculate the offset between the coordinates of the user's pupil center and the coordinates of the corneal reflection point to obtain the offset vector; The offset vector is matched with the geometric mapping model to obtain the user's line-of-sight direction information.
4. The method according to claim 1, characterized in that, The step of obtaining the image brightness value distribution in the region near the user's gaze point and determining several candidate salient regions based on the image brightness value distribution specifically includes: A first image containing the user's gaze point is obtained, and the first image is subjected to Gaussian filtering for noise reduction to obtain a second image; The second image is divided into blocks by using a fixed-size sliding window that slides across the second image with a certain step size to obtain several image blocks. For each image block, calculate the arithmetic mean of the brightness values of all pixels within that image block to obtain the average brightness value of that image block. Based on the average brightness value of each image block, the Sobel operator is used to calculate the brightness gradient between two adjacent image blocks to obtain the gradient distribution map of the second image; The gradient distribution map is binarized by setting pixels with gradient values greater than or equal to a preset gradient value threshold to 1 and pixels with gradient values less than the preset gradient value threshold to 0, thereby obtaining a binarized gradient distribution map. Connectivity analysis is performed on the binarized gradient distribution map to extract connected regions with a pixel value of 1, and the geometric center of each connected region is calculated to obtain candidate salient regions.
5. The method according to claim 4, characterized in that, The step of determining the user's attention focus by comparing the candidate salience regions with a preset salience region template specifically includes: Collect massive amounts of image data, label salient regions in the massive amounts of image data, and form an image training dataset; The image training dataset is used as input, and a convolutional neural network model is used for modeling and training to obtain a salient region detection model. The candidate salient regions are input into the salient region detection model for matching to obtain the salient probability value of each candidate salient region; The candidate saliency region with the highest saliency probability value is taken as the user attention concentration region, and the geometric center coordinates of the user attention concentration region are calculated. The geometric center coordinates of the user attention concentration region are taken as the user attention focus.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain the user's historical interaction behavior data, and extract the user's eye movement data during browsing based on the historical interaction behavior data; A Gaussian mixture model is used to calculate the probability values of the distribution points in the eye movement data, and a set of gaze regions whose probability values of the distribution points are higher than a preset probability threshold are selected. The K-means algorithm is used to cluster all gaze regions in the gaze region set. Euclidean distance is set as the standard for calculating the distance between centroids. If the distance between the centroids of two gaze regions is lower than the preset distance between centroids threshold, they are classified as gaze regions of the same gaze category, thus constructing a gaze category tree model with hierarchical relationship. Based on the gaze category tree model, a hidden Markov model is used to calculate the user attention shift probability between different gaze categories and generate an attention prediction list. The system acquires real-time user interaction behavior data and performs visual rendering processing on the current AR content of the AR device based on the attention prediction list and the real-time interaction behavior data.
7. An intelligent control system for an AR device, characterized in that, The system specifically includes: The first control module is used to acquire facial biometric data of at least one user and camera data of an AR device, and to establish a geometric mapping model between the user's gaze and the device's camera based on the facial biometric data and the camera data. The second control module is used to acquire at least one user eye image, determine the user's pupil center position coordinates and corneal reflection point position coordinates based on the user eye image, input the pupil center position coordinates and corneal reflection point position coordinates into the geometric mapping model for matching, and obtain user gaze direction information, the user gaze direction information including user gaze movement trajectory and user fixation point; The third control module is used to acquire the image brightness value distribution in the area near the user's gaze point, determine several candidate salient regions based on the image brightness value distribution, and determine the user's attention focus by comparing the candidate salient regions with a preset salient region template. The fourth control module is used to acquire the user's head posture information in real time using a head posture tracking algorithm, and dynamically control the camera shooting direction of the AR device based on the head posture information, the user's gaze direction information and the user's attention focus. The fifth control module is used to acquire real-time data on ambient light intensity near the AR device and changes in the user's pupil size. The ambient light intensity data and the pupil size change data are correlated and analyzed to form the first training dataset; Using the first training dataset as input, a machine learning algorithm is used for modeling and training to obtain a pupil size change prediction model. The pupil size change prediction model determines the user's pupil size change prediction data, and the retinal imaging data of the AR device is updated in real time using the pupil size change prediction data.
8. A computer device, characterized in that, include: The memory and processor, and the computer program stored in the memory, which, when executed on the processor, implements the intelligent control method for the AR device as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the intelligent control method for the AR device as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Electronic devices comprising camera module and method for controlling photographing direction thereof
WO2022035166A1