Method and system for adjusting camera of intelligent glasses
By integrating multi-source information and employing precise electromechanical control through smart glasses, intelligent adaptive adjustment of the camera's field of view is achieved. This solves the problems of interactive interruption and insufficient dynamic response in existing technologies, thereby improving user experience and system stability.
Patent Information
- Application Number
- CN202511737015.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-11-25
AI Technical Summary
Existing smart glasses rely on manual operation or fixed modes to adjust the camera's viewing angle, resulting in interrupted interaction, insufficient dynamic response capabilities, misaligned virtual registration, and increased cognitive load on users, making it difficult to meet the needs of application scenarios with high real-time and accuracy requirements.
It employs an eye-tracking module, a head posture perception module, an environment perception module, an intent parsing module, and a perspective decision module working together to achieve intelligent adaptive adjustment of the camera's perspective through multi-source information fusion and precise electromechanical control.
It achieves intelligent adjustment of camera perspective without manual user intervention, improves augmented reality registration accuracy, reduces user cognitive load, enhances interactive immersion, and gradually adapts to individual user behavior patterns, thereby improving the system's robustness and stability in complex environments.
Smart Images

Figure CN121209705A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart wearable device technology, specifically relating to a method and system for adjusting the camera of smart glasses. Background Technology
[0002] In the field of wearable computing devices, smart glasses, as an important carrier of augmented reality and mixed reality technologies, are gradually being applied to various scenarios such as industrial maintenance, remote collaboration, and daily interaction. As the core sensor for smart glasses to perceive the physical environment and achieve virtual-real fusion, the quality and viewing angle of the images captured directly affect the user's interactive experience and task execution efficiency.
[0003] Adjusting the camera's field of view is a key technological direction for achieving high-quality augmented reality interaction in smart glasses. This technology aims to dynamically adjust the camera's shooting parameters and framing based on the user's current gaze target, head posture, and environmental content, ensuring that virtual information can be accurately and stably superimposed on real-world objects.
[0004] Existing technologies typically rely on manual user operation or preset fixed adjustment modes to control the camera's viewing angle. Manual adjustment requires the user to interrupt the current task and make explicit interactions, severely reducing the continuity and efficiency of the operation process. Fixed mode adjustment lacks the ability to dynamically respond to user intentions and changes in the environment. When the user's gaze is fixed on a rapidly moving target or when the ambient light changes drastically, it can easily lead to virtual registration misalignment and visual discomfort.
[0005] Existing adjustment mechanisms are disconnected from users' natural visual behavior, increasing cognitive load and impairing immersion during interaction, making it difficult to meet the demands of applications requiring high real-time performance and accuracy, such as industrial inspection and remote guidance. Therefore, there is an urgent need to develop a technical solution that can intelligently sense user intent and adaptively adjust the camera's viewing angle. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for adjusting the camera of smart glasses, in order to solve the problems of interaction interruption, insufficient dynamic response capability, virtual registration misalignment and increased user cognitive load caused by reliance on manual operation or fixed adjustment mode in the prior art.
[0007] To achieve the above objectives, the present invention provides a camera adjustment system for smart glasses, the system comprising an eye-tracking module, a head posture perception module, an environment perception module, an intent parsing module, a perspective decision module, and a camera driving module.
[0008] The eye-tracking module is used to collect the user's eye movement data in real time and extract the coordinates of the gaze point.
[0009] The head posture perception module is used to acquire data on the azimuth, pitch, and rotation angles of the user's head in three-dimensional space.
[0010] The environmental perception module is used to collect scene depth information, ambient light intensity, and key object contour data.
[0011] The intent parsing module generates a quantitative description of the user's current visual intent by fusing multi-source information, based on the gaze coordinates output by the eye-tracking module, the head posture data output by the head posture perception module, and the environmental data collected by the environment perception module.
[0012] The perspective decision module receives the visual intent description output by the intent parsing module, combines it with a preset perspective adjustment strategy set, and generates target perspective parameters. Based on the target perspective parameters issued by the perspective decision module, the camera driver module drives the smart glasses' camera to perform corresponding optical zoom, mechanical rotation, or electronic image stabilization operations.
[0013] Furthermore, the eye-tracking module includes an infrared light source array, a corneal reflection image sensor, and a pupil center localization unit. The infrared light source array projects an invisible infrared light spot onto the user's eyeball. The corneal reflection image sensor captures the infrared reflection image of the eyeball surface. The pupil center localization unit performs grayscale and binarization processing on the reflection image, then uses an ellipse fitting algorithm to extract the pixel coordinates of the pupil center, and converts the pixel coordinates into a three-dimensional gaze vector based on a pre-calibrated eyeball model.
[0014] Furthermore, the head posture sensing module integrates a nine-axis inertial measurement unit, which includes a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The gyroscope outputs angular velocity data, the accelerometer outputs linear acceleration data, and the magnetometer outputs geomagnetic field strength data. A Kalman filter algorithm is used to fuse and process the multi-source sensor data, eliminating accumulated errors and outputting a high-precision quaternion representation of the head posture.
[0015] Furthermore, the environmental perception module includes a depth camera, an ambient light sensor, and an edge computing unit. The depth camera uses the time-of-flight principle to acquire physical distance information between each point in the scene and the glasses. The ambient light sensor monitors the ambient light intensity in real time. The edge computing unit performs planar segmentation and object clustering processing on the depth image, extracting the 3D bounding boxes and surface texture features of key objects within the user's gaze area.
[0016] Furthermore, the intent parsing module performs the following processing flow: First, a sliding window mean filter is applied to the fixation point coordinate sequence to eliminate physiological tremor noise; Secondly, by combining head posture data, the gaze point is transformed from the eye coordinate system to the world coordinate system; Then, based on the depth information and object contour data provided by the environmental perception module, it is determined whether the gaze point falls on the surface of a specific object. If it is determined to be object gaze, the object's motion speed, surface reflectivity, and relative distance to the user are extracted as the intent feature vector; If no specific object is detected, the user's browsing or search intent is identified based on the smoothness and speed variation pattern of the gaze point movement trajectory.
[0017] Furthermore, the perspective decision module incorporates a hybrid decision-making logic based on rules and data.
[0018] This module has multiple pre-stored perspective adjustment strategies, each associated with a specific intent type and environmental conditions.
[0019] When the intent parsing module outputs the object's gaze intent, the viewpoint decision module dynamically calculates the angular velocity and acceleration required for the camera to track the object based on the object's motion state, and optimizes the camera's exposure time and gain parameters based on the object's surface reflectivity and ambient light intensity.
[0020] When the user's browsing intent is recognized, the perspective decision module adaptively adjusts the camera's field of view size to match the user's visual scanning rhythm based on the covariance relationship between the head rotation angular velocity and the gaze point movement speed.
[0021] Furthermore, the camera driver module includes a digital signal processor and a miniature stepper motor assembly.
[0022] After receiving the target viewing angle parameters from the viewing angle decision module, the digital signal processor decomposes them into focal length adjustment, gimbal deflection angle, and image stabilization compensation.
[0023] The miniature stepper motor assembly drives the lens zoom ring, gimbal mechanism and optical image stabilization components to work together according to the decomposed instructions, so as to achieve real-time alignment between the camera's optical axis and the user's line of sight vector.
[0024] As one embodiment of the present invention, the system also includes an adaptive learning unit.
[0025] This unit continuously records the intent feature vectors generated by the intent parsing module and the viewpoint parameters finally adopted by the viewpoint decision module in different scenarios, and builds a database of user personalized adjustment preferences.
[0026] Through periodic clustering analysis, the adaptive learning unit identifies the perspective parameter patterns of user habits and dynamically updates the strategy weights in the perspective decision module, enabling the system to adjust its behavior to gradually adapt to individual user differences.
[0027] As another embodiment of the present invention, the perspective decision module introduces a delayed decision-making mechanism in the decision-making process.
[0028] When the confidence level of the intent output by the intent parsing module is lower than the preset threshold, the viewpoint decision module does not immediately drive the camera to make adjustments, but instead continues to collect several frames of perception data and reconfirm the intent.
[0029] Only when the intent classification results of three consecutive frames of data are consistent and the confidence level exceeds the threshold, will the viewpoint decision module generate the final target viewpoint parameters and send them to the camera driver module.
[0030] The present invention also provides a method for adjusting the camera of smart glasses, which uses the above-described camera adjustment system for smart glasses to adjust the camera of smart glasses.
[0031] This mechanism effectively avoids frequent malfunctions of the camera caused by momentary interference or misjudgment of intent.
[0032] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Through the coordinated data acquisition of the eye-tracking module, head posture perception module and environment perception module, the system can accurately capture the user's real-time visual focus and head movement status. 2. The intent parsing module uses multi-source information fusion technology to transform raw perceptual data into a computable visual intent description, providing a quantitative basis for perspective decision-making; 3. The perspective decision module generates target perspective parameters that match user intent and environmental conditions based on a hybrid logic of rules and data-driven approaches, thereby enabling proactive pre-adjustment of the camera perspective. 4. The camera driver module uses precision electromechanical control to ensure that the camera's optical axis is quickly and accurately aligned with the user's line of sight; 5. The entire system forms a closed-loop control from perception, analysis, decision-making to execution, which can realize intelligent adaptive adjustment of the camera perspective without manual intervention from the user, improve the accuracy of augmented reality registration, reduce the user's cognitive load and enhance the interactive immersion. 6. The introduction of adaptive learning units enables the system to gradually adapt to individual user behavior patterns and improve the level of personalized services; the delayed decision-making mechanism effectively suppresses noise interference and improves the robustness and stability of the system's decision-making in complex environments. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the overall technical solution architecture of the camera adjustment method for smart glasses proposed in this invention; Figure 2 This is a schematic diagram of the core principle framework of multi-source information fusion and intent parsing in this invention; Figure 3This is a flowchart of the hybrid decision-making logic process of the perspective decision-making module in this invention, which is based on rules and data-driven approaches. Figure 4 This is a schematic diagram of the multi-level interaction relationship and data flow of the eye tracking, head posture perception and environment perception modules in this invention; Figure 5 This is a flowchart illustrating the logical process of the adaptive learning unit in this invention for constructing personalized user preferences. Detailed Implementation Please refer to the attached document. Figure 1 This embodiment details the specific hardware configuration, module connection relationships, and data flow process of a camera adjustment system for smart glasses.
[0034] The system uses the temples and frames of smart glasses as the main physical carriers, and its core processing unit is integrated into a miniature main control board inside the temples.
[0035] After the system is powered on, each module completes self-test and parameter calibration according to the preset initialization process, and then enters continuous operation.
[0036] The eye-tracking module projects an invisible infrared spot with a wavelength of 850nm onto the user's eyeball through its infrared light source array. This spot is reflected by the cornea to form a specific bright pupil image.
[0037] The corneal reflection image sensor is a complementary metal-oxide-semiconductor image sensor with a global shutter mode. Its pixel array is 640×480 and the frame rate is set to 120fps. It is responsible for capturing infrared reflection images of the eye surface, including the bright and dark pupil areas.
[0038] The pupil center localization unit is integrated into a dedicated image signal processor within the eye-tracking module. This processor performs the following processing flow for each received frame of reflection image: First, it performs image grayscale conversion, converting the original Bayer array image into an 8-bit grayscale image; then, it performs adaptive threshold binarization to distinguish between the pupil region and the iris region; next, it uses the Canny edge detection algorithm to extract the pupil contour pixel set; finally, it applies the least squares ellipse fitting algorithm to calculate the pupil center pixel coordinates.
[0039] The coordinates are then transformed based on the eye model established in advance using the nine-point calibration method. The eye model parameters include a corneal curvature radius of 7.8 mm and a distance of 3.5 mm from the pupil center to the corneal apex. The final output is a three-dimensional gaze vector of the user in the current image coordinate system. Its data format is an array containing three floating-point numbers, which represent the components of the gaze vector in the X-axis, Y-axis and Z-axis directions, respectively.
[0040] Please continue to refer to the appendix. Figure 1The head posture sensing module is closely fitted to the inside of the nose bridge support of the smart glasses, and its nine-axis inertial measurement unit is connected to the main control board through an integrated circuit bus.
[0041] The three-axis gyroscope has a measurement range of ±2000° / s. The original value of the output angular velocity data is a 16-bit signed integer, and the converted physical quantity unit is ° / s.
[0042] The triaxial accelerometer has a measurement range of ±16 times the acceleration due to gravity. The raw value of the output linear acceleration data is a 16-bit signed integer, and the converted physical quantity unit is m / s². 2 .
[0043] The triaxial magnetometer has a measurement range of ±8G. The original value of the output geomagnetic field strength data is a 16-bit signed integer, and the converted physical quantity unit is μT.
[0044] The raw data from the nine-axis inertial measurement unit is sampled at a frequency of 1000Hz and transmitted to the microcontroller built into the head attitude sensing module via an integrated circuit bus.
[0045] The microcontroller runs a Kalman filter algorithm, and its state vector contains four components of a quaternion and three components of the gyroscope bias, for a total of seven state variables.
[0046] In the prediction phase of the Kalman filter algorithm, the quaternion state is updated based on the gyroscope angular velocity data, and the update formula is a discretized form of the quaternion differential equation.
[0047] In the correction phase of the Kalman filter algorithm, accelerometer data and magnetometer data are fused sequentially. Accelerometer data is used to correct pitch and roll angles, while magnetometer data is used to correct yaw angles.
[0048] After Kalman filtering and fusion calculation, the output is a quaternion representation of the head attitude, which is an array containing four floating-point numbers. At the same time, the output is the Euler angle representation of the head attitude, including azimuth, pitch and rotation angles, with each angle having an accuracy of 0.01°.
[0049] The depth camera of the environmental perception module is mounted in the upper center of the smart glasses frame, with its optical axis parallel to the default direction of the user's line of sight.
[0050] The depth camera uses the time-of-flight principle. Its infrared laser diode emits a modulated light signal with a wavelength of 940nm. The light signal is reflected by objects in the scene and then received by an avalanche photodiode array.
[0051] Depth cameras calculate the time of flight of light signals by measuring the phase difference between emitted and received light, and then deduce the physical distance information between each point in the scene and the glasses based on the speed of light constant.
[0052] The depth camera outputs a depth image with a resolution of 320×240, where each pixel has a 16-bit unsigned integer value representing a distance value in millimeters.
[0053] An ambient light sensor is integrated into the outer side of the temple of the smart glasses. Its spectral response range covers the visible light band from 380nm to 780nm, and the output ambient light intensity value is a 32-bit floating-point number in lux.
[0054] The edge computing unit uses a low-power system-on-a-chip, whose built-in graphics processor performs real-time processing of depth images.
[0055] The processing flow first performs bilateral filtering for noise reduction, with a filter kernel size of 5×5 pixels; then, a region-growing-based planar segmentation algorithm is executed to extract the main planes in the scene; finally, a density clustering algorithm is used to cluster objects in non-planar regions, extracting the 3D bounding boxes and surface texture features of key objects within the user's gaze area.
[0056] The data structure of the 3D bounding box contains the 3D coordinates of eight vertices, and the surface texture features include average reflectance, color histogram, and local binary pattern feature vector.
[0057] Please refer to the attached document. Figure 2 The intent parsing module runs on the multi-core processor of the main control board, and its software thread priority is set to real-time level.
[0058] This module receives the 3D gaze vector output by the eye-tracking module, the head pose quaternion output by the head pose perception module, and the depth image, ambient light intensity, and object feature data output by the environment perception module.
[0059] The intent parsing module first performs sliding window mean filtering on a sequence of 20 consecutive gaze coordinates. The sliding window size is 5 frames and the step size is 1 frame to eliminate high-frequency noise caused by physiological tremors.
[0060] The filtered gaze coordinates, combined with the head pose quaternion, are transformed from the eye coordinate system to the world coordinate system using a coordinate transformation matrix.
[0061] The coordinate transformation matrix is composed of the rotation matrix derived from the head posture quaternion and the preset installation position offset of the smart glasses.
[0062] In the world coordinate system, the intent resolution module uses the depth information provided by the environment perception module and a ray casting algorithm to determine whether the gaze point intersects with the 3D bounding box of a specific object.
[0063] If an intersection is detected, it is determined that the user is currently gazing at an object, and the object's motion speed, surface reflectivity, and relative distance to the user are extracted as the intention feature vector.
[0064] The velocity of the object is calculated by the displacement difference between the center points of the object's three-dimensional bounding box in three consecutive depth images, and the unit is m / s.
[0065] Surface reflectance is calculated based on the mapping relationship between ambient light intensity and the gray value of the object's surface, and the unit is %.
[0066] The relative distance is taken directly from the distance value of the pixel corresponding to the gaze point in the depth image, and the unit is meters.
[0067] If no specific object is detected, the intent parsing module analyzes the smoothness and speed change pattern of the gaze point movement trajectory over the past 1 second.
[0068] Smoothness is assessed by calculating the rate of change of curvature of the gaze point trajectory, while the velocity variation pattern is assessed by the ratio of the standard deviation to the mean of the gaze point movement velocity.
[0069] When the smoothness is higher than 0.8 and the speed change ratio is lower than 0.3, it is identified as user browsing intent; when the smoothness is lower than 0.5 and the speed change ratio is higher than 0.6, it is identified as user search intent.
[0070] The final output of the intent parsing module is a structured data object containing intent type encoding, intent confidence score, and intent feature vector.
[0071] Please refer to the attached document. Figure 3 The perspective decision-making module, as the core decision-making unit of the system, has pre-stored a hybrid decision-making logic based on rules and data.
[0072] After receiving the visual intent description output by the intent parsing module, this module first queries the pre-stored view adjustment strategy set.
[0073] The strategy set is stored in the form of a hash table, where the key is a combination of intent type and environmental conditions, and the value is the corresponding view parameter adjustment rule.
[0074] When the intent type in the visual intent description is object gaze intent, the viewpoint decision module dynamically calculates the angular velocity and acceleration required for the camera to track the object based on the motion state of the target object.
[0075] The formula for calculating angular velocity is the ratio of the object's velocity to the relative distance, multiplied by a proportionality coefficient of 0.85.
[0076] The acceleration is calculated as the rate of change of angular velocity, which is obtained by the difference between the angular velocity values of the past three frames.
[0077] Meanwhile, the perspective decision module optimizes the camera's exposure time and gain parameters based on the object's surface reflectivity and ambient light intensity.
[0078] The exposure time adjustment rules are as follows: when the ambient light intensity is below 100 lux and the surface reflectivity is below 30%, the exposure time is set to 33 ms; when the ambient light intensity is above 1000 lux and the surface reflectivity is above 70%, the exposure time is set to 8 ms; in other cases, the exposure time is calculated through linear interpolation. The gain parameter adjustment rules are as follows: for every 1 ms decrease in exposure time, the gain increases by 3 dB, but the maximum gain does not exceed 24 dB.
[0079] When the intent type in the visual intent description is user browsing intent, the viewpoint decision module adaptively adjusts the camera's field of view based on the covariance relationship between the head rotation angular velocity and the gaze point movement speed.
[0080] The calculation window for covariance is the data of the past 10 frames. When the covariance value is greater than 0.7, it is determined that the head rotation and the movement of the gaze point are highly synchronized, and the field of view is set to 75°. When the covariance value is between 0.3 and 0.7, the field of view is set to 60°. When the covariance value is less than 0.3, the field of view is set to 45°.
[0081] The target viewpoint parameters ultimately generated by the viewpoint decision module are structured data, including focal length, gimbal deflection angle, gimbal pitch angle, exposure time, and gain value.
[0082] Please continue to refer to the appendix. Figure 1 After the camera driver module receives the target view parameters from the view decision module, its digital signal processor first decomposes the parameters into instructions.
[0083] The focal length adjustment is calculated based on the focal length value in the target viewpoint parameters, combined with the camera lens calibration curve, and converted into the number of pulses for the stepper motor. The gimbal tilt and pitch angles are calculated based on the angle values in the target viewpoint parameters, combined with the gimbal mechanism's reduction ratio and step angle, and converted into the number of pulses for each of the two stepper motors.
[0084] The image stabilization compensation amount is calculated by a proportional-integral-derivative controller based on the angular velocity data output in real time by the head posture perception module to determine the compensation angle of the image stabilization component.
[0085] The digital signal processor generates four pulse width modulation signals, which respectively drive the lens zoom ring motor, gimbal deflection motor, gimbal pitch motor and optical image stabilization component motor in the micro stepper motor group.
[0086] The lens zoom ring motor uses a two-phase four-wire stepper motor with a step angle of 1.8° and a drive pulse frequency of 1000Hz.
[0087] The gimbal's yaw and pitch motors are micro-geared stepper motors with a reduction ratio of 1:64, a step angle of 0.9°, and a drive pulse frequency of 500Hz.
[0088] The optical image stabilization component uses a voice coil motor, driven by an analog voltage signal with a voltage range of ±5V and a response frequency of 200Hz. All motors work in tandem to ensure that the camera's optical axis remains aligned with the user's line-of-sight vector in real time within three-dimensional space, with an alignment error controlled within 0.5°.
[0089] Please refer to the attached document. Figure 5 The system in this embodiment also includes an adaptive learning unit, which runs independently on the coprocessor of the main control board, and its data storage area is non-volatile memory.
[0090] The adaptive learning unit continuously records the intent feature vector generated by the intent parsing module and the view parameters finally adopted by the view decision module in different scenarios. The timestamp accuracy of each record is 1ms.
[0091] The recorded data is stored in a circular buffer with a capacity of 10,000 records.
[0092] The adaptive learning unit performs periodic clustering analysis every 24 hours, using a density-based noise-based spatial clustering algorithm to cluster the stored records. The neighborhood radius parameter of the clustering algorithm is set to 0.3, and the minimum number of samples is set to 50. Through clustering analysis, the adaptive learning unit identifies user-preferred viewpoint parameter patterns, with each pattern represented by the viewpoint parameter value of its cluster center.
[0093] The number of identified patterns is typically 3 to 5, such as quick browsing mode, detailed observation mode, and dynamic tracking mode. The adaptive learning unit then dynamically updates the weight values of the policy hash table in the viewpoint decision module based on the frequency of occurrence and recent usage time of each pattern.
[0094] The weight update formula is: the new weight equals the old weight multiplied by 0.9 plus the current frequency factor multiplied by 0.1.
[0095] The frequency factor is the ratio of the number of times a pattern appears in the most recent 1000 records to the total number of records.
[0096] By updating the weights, the decision-making behavior of the perspective decision-making module gradually adapts to individual user differences, making the system more in line with user habits.
[0097] The system in this embodiment also integrates a delayed decision-making mechanism, which is implemented by a dedicated state machine within the viewpoint decision-making module.
[0098] When the intent confidence level output by the intent parsing module is lower than the preset threshold of 0.75, the viewpoint decision module enters a waiting state.
[0099] While in the waiting state, the perspective decision module continuously collects subsequent perception data frames and reconfirms the intent for each frame of data.
[0100] The viewpoint decision module generates the final target viewpoint parameters and sends them to the camera driver module only when the intent classification results of three consecutive frames of data are consistent and the confidence level of each frame exceeds the threshold of 0.75.
[0101] The maximum duration of the waiting state is 500ms. If the conditions are not met within the timeout period, the viewpoint decision module will use the default viewpoint parameters, namely a focal length of 35mm, a field of view of 60°, and an exposure time of 16ms.
[0102] The delayed decision-making mechanism effectively avoids frequent camera malfunctions caused by momentary interference or misjudgment of intent, improving the system's decision-making robustness and stability in complex environments.
[0103] Data communication between the various modules of the system adopts a unified communication protocol. The data frame format includes a frame header, data length, module identifier, timestamp, data payload, and cyclic redundancy check code.
[0104] The frame header is a fixed 2-byte value 0xAA55, the data length is a 2-byte unsigned integer, the module identifier is a 1-byte enumeration value, the timestamp is a 4-byte millisecond-level time, the data payload is a variable-length byte array, and the cyclic redundancy check code is a 2-byte CRC16 check value.
[0105] The communication rate is set to 1 Mbit / s, and the error retransmission mechanism adopts the automatic repeat request protocol, with a maximum of 3 retransmissions.
[0106] The system power management adopts dynamic voltage and frequency adjustment technology, which dynamically adjusts the operating voltage and clock frequency according to the real-time calculated load of each module. Under typical usage scenarios, the overall power consumption of the system is less than 350mW, which can support the smart glasses to work continuously for more than 8 hours.
[0107] Please refer to the attached document. Figure 4 This embodiment further elaborates on the multi-level interaction relationship and data flow details between the eye-tracking module, the head posture perception module, and the environment perception module.
[0108] After extracting the pixel coordinates of the pupil center, the pupil center localization unit of the eye-tracking module not only performs eyeball model transformation, but also outputs pupil diameter change data.
[0109] The pupil diameter was calculated by averaging the major and minor axes of an ellipse obtained through an ellipse fitting algorithm, with a sampling frequency of 120Hz.
[0110] Pupil diameter data is transmitted to the intent parsing module in 32-bit floating-point format as an auxiliary assessment indicator of user cognitive load.
[0111] When the pupil diameter changes by more than 15% within 1 second, the intent parsing module lowers the current intent confidence by 0.1 as a signal of increased decision uncertainty.
[0112] The nine-axis inertial measurement unit of the head posture sensing module not only outputs the head posture quaternion, but also provides the jerk value of the head motion acceleration.
[0113] The Jerk value is the derivative of the acceleration, calculated using the central difference method from five consecutive frames of acceleration data, and its unit is m / s². 3 .
[0114] When the jerk value of head movement exceeds 10 m / s 3 At this time, the head posture perception module sends a motion change flag to the intent parsing module. After receiving the flag, the intent parsing module temporarily adjusts the window size of its sliding window mean filter from 5 frames to 3 frames to speed up the response to sudden head movements.
[0115] When extracting the 3D bounding boxes of key objects, the edge computing unit of the environment perception module adds a surface material classification function.
[0116] Surface material classification is based on infrared reflection intensity captured by a depth camera and texture features captured by a visible light camera, and is achieved through a support vector machine classifier.
[0117] The classifier was pre-trained on datasets of 10 common materials, including metal, plastic, glass, wood, and fabric.
[0118] The material classification results are output as 1-byte enumeration values, serving as an additional dimension to the intent feature vector.
[0119] When the surface material of an object is detected to be highly reflective, the viewpoint decision module introduces an additional compensation coefficient of 0.8 when calculating the exposure parameters to prevent overexposure.
[0120] The intent parsing module uses a multi-level collision detection algorithm to determine whether the gaze point falls on the surface of a specific object.
[0121] The first level is a coarse axis-aligned bounding box detection to quickly eliminate obviously irrelevant objects; the second level is directional bounding box detection to improve detection accuracy; the third level is a precise ray triangular facet intersection detection based on depth images to ensure accurate determination of the relationship between the gaze point and the object surface.
[0122] The multi-level detection algorithm ensures accuracy while keeping the average processing time within 5ms.
[0123] The perspective decision-making module introduces a reinforcement learning mechanism into its rule-based and data-driven hybrid decision-making logic.
[0124] The reinforcement learning agent takes the intent feature vector and the environmental state as input, adjusts the viewpoint parameters as actions, and uses the stability of the user's gaze point in the next 3 seconds as a reward signal.
[0125] The reward function is defined as the reciprocal of the variance of the fixation point location; the smaller the variance, the higher the reward.
[0126] The reinforcement learning agent employs a deep Q-network structure. The network input layer consists of 10 dimensions of the intent feature vector plus 5 dimensions of the environment state. The hidden layers are two fully connected layers, each with 128 neurons. The output layer consists of 5 discretized viewpoint adjustment actions. The deep Q-network updates its weights every 24 hours, using data from the most recent 1000 decision records stored in the adaptive learning unit.
[0127] The camera driver module employs adaptive microstepping control technology when driving the micro stepper motor assembly. It dynamically adjusts the microstepping subdivision factor based on the deviation between the target angle and the current angle.
[0128] When the deviation is greater than 5°, the full step mode is used with a subdivision factor of 1; when the deviation is between 1° and 5°, the 4-subdivision microstep mode is used; when the deviation is less than 1°, the 16-subdivision microstep mode is used.
[0129] Adaptive microstepping control technology ensures fast response while keeping the operating noise of the stepper motor below 35dB.
[0130] When building a database of user-personalized adjustment preferences, the adaptive learning unit adds scene context labels.
[0131] Scene context is obtained through joint analysis of depth images and visible light images from the environment perception module, and scene classification is performed using a convolutional neural network.
[0132] Scene categories include five main categories: indoor office, outdoor street, indoor home, outdoor nature, and transportation.
[0133] During cluster analysis, the adaptive learning unit prioritizes clustering records within the same scene context to form a scene-specific viewpoint parameter pattern library.
[0134] When the system detects that a user has entered a specific scene, the viewpoint decision module prioritizes loading the viewpoint parameter mode corresponding to that scene, enabling rapid adjustment for scene adaptation.
[0135] The delayed decision-making mechanism introduces multiple hypothesis tracking technology while waiting for confirmation of intent.
[0136] When the intent confidence output by the intent parsing module is lower than the threshold but there are multiple candidate intents, the viewpoint decision module tracks the development trend of all candidate intents in parallel.
[0137] Each candidate intent is assigned a hypothesis weight, which is dynamically updated based on the confidence level changes in consecutive frames.
[0138] When the weight of a certain hypothesis continues to increase over three consecutive frames and eventually exceeds 0.8, the viewpoint decision module will adopt that hypothesis as the final decision, even if the confidence levels of other hypotheses do not fully meet the conditions.
[0139] Multihypothesis tracking technology reduces the system's decision latency in fuzzy intent scenarios from an average of 500ms to 300ms.
[0140] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0141] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A camera adjustment system for smart glasses, characterized in that, include: The eye-tracking module is used to collect the user's eye movement data in real time and extract the coordinates of the gaze point; The head posture perception module is used to acquire data on the azimuth, pitch, and rotation angles of the user's head in three-dimensional space. The environmental perception module is used to collect scene depth information, ambient light intensity, and key object contour data; The intent parsing module, based on the gaze coordinates output by the eye-tracking module, the head posture data output by the head posture perception module, and the environmental data collected by the environment perception module, generates a quantitative description of the user's current visual intent through multi-source information fusion calculation. The perspective decision module receives the visual intent description output by the intent parsing module, combines it with a preset perspective adjustment strategy set, and generates target perspective parameters. The camera driver module, based on the target viewing angle parameters sent by the viewing angle decision module, drives the camera of the smart glasses to perform corresponding optical zoom, mechanical rotation, or electronic image stabilization operations.
2. The camera adjustment system for smart glasses according to claim 1, characterized in that, The eye-tracking module includes an infrared light source array, a corneal reflection image sensor, and a pupil center positioning unit. The infrared light source array projects an invisible infrared light spot onto the user's eyeball. The corneal reflection image sensor captures the infrared reflection image on the surface of the eyeball. The pupil center positioning unit performs grayscale and binarization processing on the reflection image, and then uses an ellipse fitting algorithm to extract the pixel coordinates of the pupil center. Based on a pre-calibrated eyeball model, the pixel coordinates are converted into a three-dimensional gaze vector.
3. The camera adjustment system for smart glasses according to claim 1, characterized in that, The head posture sensing module integrates a nine-axis inertial measurement unit, which includes a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The gyroscope outputs angular velocity data, the accelerometer outputs linear acceleration data, and the magnetometer outputs geomagnetic field strength data. The multi-source sensor data is fused and calculated using a Kalman filter algorithm to eliminate accumulated errors and output a high-precision head posture quaternion representation.
4. The camera adjustment system for smart glasses according to claim 1, characterized in that, The environmental perception module includes a depth camera, an ambient light sensor, and an edge computing unit. The depth camera uses the time-of-flight principle to acquire the physical distance information between each point in the scene and the glasses. The ambient light sensor monitors the ambient light intensity value in real time. The edge computing unit performs planar segmentation and object clustering processing on the depth image to extract the three-dimensional bounding boxes and surface texture features of key objects within the user's gaze area.
5. The camera adjustment system for smart glasses according to claim 1, characterized in that, The intent parsing module performs the following processing flow: First, a sliding window mean filter is applied to the fixation point coordinate sequence to eliminate physiological tremor noise; Secondly, by combining head posture data, the gaze point is transformed from the eye coordinate system to the world coordinate system; Then, based on the depth information and object contour data provided by the environmental perception module, it is determined whether the gaze point falls on the surface of a specific object. If it is determined to be object gaze, the object's motion speed, surface reflectivity, and relative distance to the user are extracted as the intent feature vector; If no specific object is detected, the user's browsing or search intent is identified based on the smoothness and speed variation pattern of the gaze point movement trajectory.
6. The camera adjustment system for smart glasses according to claim 1, characterized in that, The perspective decision module is embedded with a hybrid decision logic based on rules and data. The module has multiple perspective adjustment strategies, each of which is associated with a specific intent type and environmental conditions. When the intent parsing module outputs the object's gaze intent, the viewpoint decision module dynamically calculates the angular velocity and acceleration required for the camera to track the object based on the object's motion state, and optimizes the camera's exposure time and gain parameters based on the object's surface reflectivity and ambient light intensity. When the user's browsing intent is recognized, the perspective decision module adaptively adjusts the camera's field of view size to match the user's visual scanning rhythm based on the covariance relationship between the head rotation angular velocity and the gaze point movement speed.
7. The camera adjustment system for smart glasses according to claim 1, characterized in that, The camera driver module includes a digital signal processor and a micro stepper motor assembly; After receiving the target viewing angle parameters from the viewing angle decision module, the digital signal processor decomposes them into focal length adjustment, gimbal deflection angle and image stabilization compensation. The miniature stepper motor assembly drives the lens zoom ring, gimbal mechanism and optical image stabilization components to work together according to the decomposed instructions, so as to achieve real-time alignment between the camera's optical axis and the user's line of sight vector.
8. The camera adjustment system for smart glasses according to claim 1, characterized in that, The system also includes an adaptive learning unit; This unit continuously records the intent feature vectors generated by the intent parsing module and the view parameters finally adopted by the view decision module in different scenarios, and builds a database of users' personalized adjustment preferences. Through periodic clustering analysis, the adaptive learning unit identifies the perspective parameter patterns of user habits and dynamically updates the strategy weights in the perspective decision module, enabling the system to adjust its behavior to gradually adapt to individual user differences.
9. A camera adjustment system for smart glasses according to claim 8, characterized in that, The periodic clustering analysis employs a density-based noise-based spatial clustering algorithm; the clustering analysis is performed once at fixed intervals, and the identified viewpoint parameter patterns include a fast browsing mode, a fine observation mode, and a dynamic tracking mode.
10. A method for adjusting the camera of smart glasses, characterized in that, The camera adjustment of the smart glasses is achieved using the camera adjustment system of the smart glasses as described in any one of claims 1 to 9.
Citation Information
Patent Citations
An adaptive adjusting method, a pair of AR intelligent learning glasses and an AR intelligent learning system
CN108933916A
Intelligent control method and system for AR equipment
CN119847330A
Synchronous control system for shooting visual angle of intelligent glasses and intelligent glasses
CN220419683U
Eye tracking combiner having multiple perspectives
US10852817B1
Adaptive visual focus and tracking headgear
US20250291179A1