A method and system for adjusting the camera of smart glasses

By fusing multi-source information from eye-tracking, head posture, and environmental perception modules, and combining intent parsing and perspective decision-making, the smart glasses camera achieves adaptive adjustment, solving the problems of interaction interruption and cognitive load caused by perspective adjustment in existing technologies, and improving user experience and system stability.

CN121209705BActive Publication Date: 2026-01-30NINGBO JINSHENGXIN IMAGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511737015.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-01-30
Estimated Expiration
2045-11-25

AI Technical Summary

Technical Problem

The camera viewing angle adjustment of existing smart glasses relies on manual operation or fixed modes, which leads to interrupted interaction, insufficient dynamic response capability, virtual registration misalignment, and increased cognitive load on users, making it difficult to meet the needs of application scenarios with high requirements for real-time performance and accuracy.

Method used

It employs an eye-tracking module, a head posture perception module, an environment perception module, an intent parsing module, and a perspective decision module working together to achieve intelligent adaptive adjustment of the camera's perspective through multi-source information fusion and precise electromechanical control.

Benefits of technology

It achieves intelligent adjustment of camera perspective without manual user intervention, improves augmented reality registration accuracy, reduces user cognitive load and enhances interactive immersion, adapts to individual user differences and improves the system's robustness in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121209705B_ABST
    Figure CN121209705B_ABST
Patent Text Reader

Abstract

This invention relates to the field of smart wearable device technology, specifically disclosing a camera adjustment method and system for smart glasses. The system includes an eye-tracking module, a head posture perception module, an environment perception module, an intent parsing module, a perspective decision module, and a camera driving module. Through multi-source information fusion, it analyzes the user's visual intent in real time, generates target perspective parameters, and drives the camera to perform optical zoom, mechanical rotation, or electronic image stabilization, achieving intelligent adaptive adjustment of the camera's perspective and improving augmented reality registration accuracy and interactive immersion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of intelligent wearable devices, and particularly relates to a camera adjustment method and system for intelligent glasses. BACKGROUND

[0002] In the field of wearable computing devices, smart glasses are gradually applied to multiple scenarios such as industrial maintenance, remote collaboration and daily interaction as an important carrier of augmented reality and mixed reality technology. As the core sensor for smart glasses to perceive the physical environment and realize the fusion of virtual and real, the quality and angle of view of the images captured by the camera directly affect the user's interactive experience and task execution efficiency.

[0003] Camera angle adjustment of smart glasses is a key technology direction to realize high-quality augmented reality interaction. This technology aims to dynamically adjust the shooting parameters and viewfinder range of the camera according to the user's current gaze target, head pose and environmental content, to ensure that virtual information can be accurately and stably superimposed on real-world objects.

[0004] Existing technologies usually rely on manual operation or use preset fixed adjustment modes to control the camera angle. Manual adjustment requires the user to interrupt the current task and perform explicit interaction, which seriously reduces the continuity and efficiency of the operation process. Fixed mode adjustment lacks dynamic response capability to user intentions and environmental changes, and when the user's gaze target moves quickly or the environmental light changes dramatically, it is easy to cause virtual registration misalignment and visual discomfort.

[0005] The existing adjustment mechanism is out of line with the user's natural visual behavior, causing increased cognitive load and impaired immersion in the interaction process, making it difficult to meet the needs of application scenarios such as industrial inspection and remote guidance that require real-time and precision. Therefore, it is urgent to develop a technical solution that can intelligently perceive user intentions and adaptively adjust the camera angle. SUMMARY

[0006] The present application aims to provide a camera adjustment method and system for intelligent glasses to solve the problems of interaction interruption, insufficient dynamic response capability, virtual registration misalignment and increased user cognitive load caused by relying on manual operation or fixed adjustment mode in the prior art.

[0007] To achieve the above-mentioned purpose, the present application provides a camera adjustment system for intelligent glasses, which includes an eye movement tracking module, a head pose perception module, an environment perception module, an intention analysis module, an angle decision module and a camera driving module.

[0008] The eye movement tracking module is used to collect user eye movement data in real time and extract the gaze point coordinates.

[0009] The head pose perception module is configured to obtain azimuth, pitch and roll angle data of the user's head in a three-dimensional space.

[0010] The environment perception module is configured to collect scene depth information, ambient light intensity and key object contour data.

[0011] The intent analysis module is configured to generate a quantitative description of the user's current visual intent based on the gaze point coordinates output by the eye tracking module, the head pose data output by the head pose perception module, and the environmental data collected by the environment perception module through multi-source information fusion calculation.

[0012] The view angle decision module receives the visual intent description output by the intent analysis module, combines a preset view angle adjustment strategy set, and decides to generate target view angle parameters. The camera driving module drives the camera of the smart glasses to perform corresponding optical zoom, mechanical rotation or electronic anti-shake operations according to the target view angle parameters issued by the view angle decision module.

[0013] Further, the eye tracking module includes an infrared light source array, a corneal reflection image sensor, and a pupil center positioning unit. The infrared light source array projects invisible infrared spots to the user's eyeball. The corneal reflection image sensor captures the infrared reflection image on the surface of the eyeball. The pupil center positioning unit performs grayscale and binary processing on the reflection image, then extracts the pupil center pixel coordinates using an ellipse fitting algorithm, and converts the pixel coordinates into a three-dimensional gaze vector according to a pre-calibrated eyeball model.

[0014] Further, the head pose perception module integrates a nine-axis inertial measurement unit, which includes a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The gyroscope outputs angular velocity data, the accelerometer outputs linear acceleration data, and the magnetometer outputs geomagnetic field intensity data. Through Kalman filtering algorithm, the above multi-source sensing data are fused and solved to eliminate cumulative errors and output high-precision head pose quaternion representation.

[0015] Further, the environment perception module includes a depth camera, an ambient light sensor, and an edge computing unit. The depth camera uses the time-of-flight principle to obtain the physical distance information of each point in the scene from the glasses. The ambient light sensor monitors the ambient light intensity value in real time. The edge computing unit performs plane segmentation and object clustering processing on the depth image, and extracts the three-dimensional bounding box and surface texture features of the key objects in the user's gaze area.

[0016] Further, the intent analysis module performs the following processing flow:

[0017] First, the gaze point coordinate sequence is subjected to sliding window mean filtering to eliminate physiological tremor noise;

[0018] Second, the gaze point is converted from the eye coordinate system to the world coordinate system in combination with the head pose data;

[0019] Then, according to the depth information and the object contour data provided by the environment perception module, it is determined whether the gaze point falls on the surface of a specific object;

[0020] If it is determined that the object is gazed, the motion speed, surface reflectivity and relative distance of the object to the user are extracted as the intention feature vector;

[0021] If no specific object is detected, according to the smoothness and speed change pattern of the gaze point movement trajectory, the user is identified as a browsing intention or a search intention.

[0022] Further, the view angle decision module is embedded with a hybrid decision logic based on rules and data driving.

[0023] The module pre-stores a plurality of view angle adjustment strategies, each of which is associated with a specific intention type and environmental condition.

[0024] When the intention analysis module outputs an object gazing intention, the view angle decision module dynamically calculates the angular velocity and acceleration required for the camera to track the object according to the motion state of the target object, and optimizes the camera exposure time and gain parameters according to the object surface reflectivity and environmental light intensity.

[0025] When the user's browsing intention is identified, the view angle decision module adaptively adjusts the size of the camera field of view according to the covariance relationship between the head rotation angular velocity and the gaze point movement speed to match the user's visual scanning rhythm.

[0026] Further, the camera driving module includes a digital signal processor and a micro stepping motor set.

[0027] After receiving the target view angle parameters issued by the view angle decision module, the digital signal processor decomposes them into focal length adjustment amount, gimbal deflection angle and image stabilization compensation amount.

[0028] The micro stepping motor set drives the lens zoom ring, gimbal mechanism and optical anti-shake component to act cooperatively according to the decomposed instructions, so as to realize the real-time alignment of the camera optical axis and the user's line of sight vector.

[0029] As an embodiment of the present application, the system further has an adaptive learning unit.

[0030] The unit continuously records the intention feature vectors generated by the intention analysis module and the view angle parameters finally adopted by the view angle decision module under different scenarios, and constructs a user personalized adjustment preference database.

[0031] Through periodic cluster analysis, the adaptive learning unit identifies the view angle parameter mode of the user's habits, and dynamically updates the strategy weight in the view angle decision module, so that the system adjustment behavior gradually adapts to individual differences of the user.

[0032] As another embodiment of the present application, the view angle decision module introduces a delayed decision mechanism in the decision process.

[0033] When the intention confidence output by the intention analysis module is lower than the preset threshold, the view angle decision module does not immediately drive the camera to perform adjustment, but continues to collect subsequent frames of perception data and performs intention reconfirmation.

[0034] Only when the intention classification results of the continuous 3 frames of data are consistent and the confidence is all above the threshold, the view angle decision module generates the final target view angle parameter and issues it to the camera driving module.

[0035] The present application also provides a camera adjustment method of intelligent glasses, which uses the above-mentioned camera adjustment system of intelligent glasses to realize the camera adjustment of intelligent glasses.

[0036] This mechanism effectively avoids the frequent misoperation of the camera caused by transient interference or intention misjudgment.

[0037] Compared with the prior art, the present application has the beneficial effects that:

[0038] 1. Through the cooperative data acquisition of the eye movement tracking module, the head posture perception module and the environment perception module, the system can accurately capture the real-time visual attention target and head movement state of the user;

[0039] 2. The intention analysis module uses multi-source information fusion technology to convert the original perception data into a computable visual intention description, providing a quantitative basis for view angle decision;

[0040] 3. The view angle decision module generates target view angle parameters that match the user's intention and environmental conditions based on the mixed logic of rules and data driving, realizing the active pre-adjustment of the camera view angle;

[0041] 4. The camera driving module ensures that the optical axis of the camera is quickly and accurately aligned with the user's line of sight through precise electromechanical control;

[0042] 5. The entire system forms a closed-loop control from perception, analysis, decision to execution, and can realize intelligent adaptive adjustment of the camera view angle without manual intervention of the user, improving the augmented reality registration accuracy, reducing the user's cognitive load and enhancing the interactive immersion;

[0043] 6. The introduction of the adaptive learning unit enables the system to gradually adapt to the individual behavior patterns of the user, improving the level of personalized service; the delayed decision mechanism effectively suppresses noise interference and improves the decision robustness and stability of the system in complex environments. BRIEF DESCRIPTION OF DRAWINGS

[0044] Figure 1is the overall technical scheme architecture schematic diagram of the camera adjustment method of the intelligent glasses proposed in the present application;

[0045] Figure 2 is the core principle framework schematic diagram of multi-source information fusion and intention analysis in the present application;

[0046] Figure 3 is the mixed decision logic flow framework diagram of the view angle decision module based on rules and data driving in the present application;

[0047] Figure 4 is the multi-level interaction relationship and data flow schematic diagram of the eye movement tracking, head posture perception and environment perception module in the present application;

[0048] Figure 5 is the logic flow framework diagram of the self-adaptive learning unit constructing user individualized adjustment preference in the present application; DETAILED DESCRIPTION

[0049] Please refer to the attached Figure 1 The embodiment details the specific hardware composition, module connection relationship and data flow process of the camera adjustment system of the intelligent glasses.

[0050] The system takes the temple and frame of the intelligent glasses as the main physical carrier, and the core processing unit is integrated on the micro main control board inside the temple.

[0051] After the system is powered on and started, each module completes self-checking and parameter calibration according to the preset initialization process, and then enters a continuous running state.

[0052] The eye movement tracking module projects an invisible infrared light spot with a wavelength of 850nm to the user's eyeball through its infrared light source array. The light spot forms a specific bright pupil image after being reflected by the cornea.

[0053] The corneal reflection image sensor adopts a global shutter mode complementary metal oxide semiconductor image sensor, with a pixel array of 640x480 and a frame rate of 120fps, responsible for capturing the infrared reflection image of the eyeball surface containing the bright pupil and dark pupil regions.

[0054] The pupil center positioning unit is integrated in the dedicated image signal processor inside the eye movement tracking module. The processor performs the following processing procedures on each received frame of reflection image: first, perform image grayscale conversion to convert the original Bayer array image to an 8-bit grayscale image; then perform adaptive threshold binarization processing to distinguish the pupil region and the iris region; then extract the pupil contour pixel set using the Canny edge detection algorithm; finally, apply the least squares ellipse fitting algorithm to calculate the pupil center pixel coordinates.

[0055] The coordinates are then transformed based on the eye model established in advance using the nine-point calibration method. The eye model parameters include a corneal curvature radius of 7.8 mm and a distance of 3.5 mm from the pupil center to the corneal apex. The final output is a three-dimensional gaze vector of the user in the current image coordinate system. Its data format is an array containing three floating-point numbers, which represent the components of the gaze vector in the X-axis, Y-axis and Z-axis directions, respectively.

[0056] Please continue to refer to the appendix. Figure 1 The head posture sensing module is closely fitted to the inside of the nose bridge support of the smart glasses, and its nine-axis inertial measurement unit is connected to the main control board through an integrated circuit bus.

[0057] The three-axis gyroscope has a measurement range of ±2000° / s. The original value of the output angular velocity data is a 16-bit signed integer, and the converted physical quantity unit is ° / s.

[0058] The triaxial accelerometer has a measurement range of ±16 times the acceleration due to gravity. The raw value of the output linear acceleration data is a 16-bit signed integer, and the converted physical quantity unit is m / s². 2 .

[0059] The triaxial magnetometer has a measurement range of ±8G. The original value of the output geomagnetic field strength data is a 16-bit signed integer, and the converted physical quantity unit is μT.

[0060] The raw data from the nine-axis inertial measurement unit is sampled at a frequency of 1000Hz and transmitted to the microcontroller built into the head attitude sensing module via an integrated circuit bus.

[0061] The microcontroller runs a Kalman filter algorithm, and its state vector contains four components of a quaternion and three components of the gyroscope bias, for a total of seven state variables.

[0062] In the prediction phase of the Kalman filter algorithm, the quaternion state is updated based on the gyroscope angular velocity data, and the update formula is a discretized form of the quaternion differential equation.

[0063] In the correction phase of the Kalman filter algorithm, accelerometer data and magnetometer data are fused sequentially. Accelerometer data is used to correct pitch and roll angles, while magnetometer data is used to correct yaw angles.

[0064] After Kalman filtering and fusion calculation, the output is a quaternion representation of the head attitude, which is an array containing four floating-point numbers. At the same time, the output is the Euler angle representation of the head attitude, including azimuth, pitch and rotation angles, with each angle having an accuracy of 0.01°.

[0065] The depth camera of the environmental perception module is mounted in the upper center of the smart glasses frame, with its optical axis parallel to the default direction of the user's line of sight.

[0066] The depth camera uses the time-of-flight principle. Its infrared laser diode emits a modulated light signal with a wavelength of 940nm. The light signal is reflected by objects in the scene and then received by an avalanche photodiode array.

[0067] Depth cameras calculate the time of flight of light signals by measuring the phase difference between emitted and received light, and then deduce the physical distance information between each point in the scene and the glasses based on the speed of light constant.

[0068] The depth camera outputs a depth image with a resolution of 320×240, where each pixel has a 16-bit unsigned integer value representing a distance value in millimeters.

[0069] An ambient light sensor is integrated into the outer side of the temple of the smart glasses. Its spectral response range covers the visible light band from 380nm to 780nm, and the output ambient light intensity value is a 32-bit floating-point number in lux.

[0070] The edge computing unit uses a low-power system-on-a-chip, whose built-in graphics processor performs real-time processing of depth images.

[0071] The processing flow first performs bilateral filtering for noise reduction, with a filter kernel size of 5×5 pixels; then, a region-growing-based planar segmentation algorithm is executed to extract the main planes in the scene; finally, a density clustering algorithm is used to cluster objects in non-planar regions, extracting the 3D bounding boxes and surface texture features of key objects within the user's gaze area.

[0072] The data structure of the 3D bounding box contains the 3D coordinates of eight vertices, and the surface texture features include average reflectance, color histogram, and local binary pattern feature vector.

[0073] Please refer to the attached document. Figure 2 The intent parsing module runs on the multi-core processor of the main control board, and its software thread priority is set to real-time level.

[0074] This module receives the 3D gaze vector output by the eye-tracking module, the head pose quaternion output by the head pose perception module, and the depth image, ambient light intensity, and object feature data output by the environment perception module.

[0075] The intent parsing module first performs sliding window mean filtering on a sequence of 20 consecutive gaze coordinates. The sliding window size is 5 frames and the step size is 1 frame to eliminate high-frequency noise caused by physiological tremors.

[0076] The filtered gaze coordinates, combined with the head pose quaternion, are transformed from the eye coordinate system to the world coordinate system using a coordinate transformation matrix.

[0077] The coordinate transformation matrix is ​​composed of the rotation matrix derived from the head posture quaternion and the preset installation position offset of the smart glasses.

[0078] In the world coordinate system, the intent resolution module uses the depth information provided by the environment perception module and a ray casting algorithm to determine whether the gaze point intersects with the 3D bounding box of a specific object.

[0079] If an intersection is detected, it is determined that the user is currently gazing at an object, and the object's motion speed, surface reflectivity, and relative distance to the user are extracted as the intention feature vector.

[0080] The velocity of the object is calculated by the displacement difference between the center points of the object's three-dimensional bounding box in three consecutive depth images, and the unit is m / s.

[0081] Surface reflectance is calculated based on the mapping relationship between ambient light intensity and the gray value of the object's surface, and the unit is %.

[0082] The relative distance is taken directly from the distance value of the pixel corresponding to the gaze point in the depth image, and the unit is meters.

[0083] If no specific object is detected, the intent parsing module analyzes the smoothness and speed change pattern of the gaze point movement trajectory over the past 1 second.

[0084] Smoothness is assessed by calculating the rate of change of curvature of the gaze point trajectory, while the velocity variation pattern is assessed by the ratio of the standard deviation to the mean of the gaze point movement velocity.

[0085] When the smoothness is higher than 0.8 and the speed change ratio is lower than 0.3, it is identified as user browsing intent; when the smoothness is lower than 0.5 and the speed change ratio is higher than 0.6, it is identified as user search intent.

[0086] The final output of the intent parsing module is a structured data object containing intent type encoding, intent confidence score, and intent feature vector.

[0087] Please refer to the attached document. Figure 3 The perspective decision-making module, as the core decision-making unit of the system, has pre-stored a hybrid decision-making logic based on rules and data.

[0088] After receiving the visual intent description output by the intent parsing module, this module first queries the pre-stored view adjustment strategy set.

[0089] The strategy set is stored in the form of a hash table, where the key is a combination of intent type and environmental conditions, and the value is the corresponding view parameter adjustment rule.

[0090] When the intent type in the visual intent description is object gaze intent, the viewpoint decision module dynamically calculates the angular velocity and acceleration required for the camera to track the object based on the motion state of the target object.

[0091] The formula for calculating angular velocity is the ratio of the object's velocity to the relative distance, multiplied by a proportionality coefficient of 0.85.

[0092] The acceleration is calculated as the rate of change of angular velocity, which is obtained by the difference between the angular velocity values ​​of the past three frames.

[0093] Meanwhile, the perspective decision module optimizes the camera's exposure time and gain parameters based on the object's surface reflectivity and ambient light intensity.

[0094] The exposure time adjustment rules are as follows: when the ambient light intensity is below 100 lux and the surface reflectivity is below 30%, the exposure time is set to 33 ms; when the ambient light intensity is above 1000 lux and the surface reflectivity is above 70%, the exposure time is set to 8 ms; in other cases, the exposure time is calculated through linear interpolation. The gain parameter adjustment rules are as follows: for every 1 ms decrease in exposure time, the gain increases by 3 dB, but the maximum gain does not exceed 24 dB.

[0095] When the intent type in the visual intent description is user browsing intent, the viewpoint decision module adaptively adjusts the camera's field of view based on the covariance relationship between the head rotation angular velocity and the gaze point movement speed.

[0096] The calculation window for covariance is the data of the past 10 frames. When the covariance value is greater than 0.7, it is determined that the head rotation and the movement of the gaze point are highly synchronized, and the field of view is set to 75°. When the covariance value is between 0.3 and 0.7, the field of view is set to 60°. When the covariance value is less than 0.3, the field of view is set to 45°.

[0097] The target viewpoint parameters ultimately generated by the viewpoint decision module are structured data, including focal length, gimbal deflection angle, gimbal pitch angle, exposure time, and gain value.

[0098] Please continue to refer to the appendix. Figure 1 After the camera driver module receives the target view parameters from the view decision module, its digital signal processor first decomposes the parameters into instructions.

[0099] The focal length adjustment is calculated based on the focal length value in the target viewpoint parameters, combined with the camera lens calibration curve, and converted into the number of pulses for the stepper motor. The gimbal tilt and pitch angles are calculated based on the angle values ​​in the target viewpoint parameters, combined with the gimbal mechanism's reduction ratio and step angle, and converted into the number of pulses for each of the two stepper motors.

[0100] The image stabilization compensation amount is calculated by a proportional-integral-derivative controller based on the angular velocity data output in real time by the head posture perception module to determine the compensation angle of the image stabilization component.

[0101] The digital signal processor generates four pulse width modulation signals, which respectively drive the lens zoom ring motor, gimbal deflection motor, gimbal pitch motor and optical image stabilization component motor in the micro stepper motor group.

[0102] The lens zoom ring motor uses a two-phase four-wire stepper motor with a step angle of 1.8° and a drive pulse frequency of 1000Hz.

[0103] The gimbal's yaw and pitch motors are micro-geared stepper motors with a reduction ratio of 1:64, a step angle of 0.9°, and a drive pulse frequency of 500Hz.

[0104] The optical image stabilization component uses a voice coil motor, driven by an analog voltage signal with a voltage range of ±5V and a response frequency of 200Hz. All motors work in tandem to ensure that the camera's optical axis remains aligned with the user's line-of-sight vector in real time within three-dimensional space, with an alignment error controlled within 0.5°.

[0105] Please refer to the attached document. Figure 5 The system in this embodiment also includes an adaptive learning unit, which runs independently on the coprocessor of the main control board, and its data storage area is non-volatile memory.

[0106] The adaptive learning unit continuously records the intent feature vector generated by the intent parsing module and the view parameters finally adopted by the view decision module in different scenarios. The timestamp accuracy of each record is 1ms.

[0107] The recorded data is stored in a circular buffer with a capacity of 10,000 records.

[0108] The adaptive learning unit performs periodic clustering analysis every 24 hours, using a density-based noise-based spatial clustering algorithm to cluster the stored records. The neighborhood radius parameter of the clustering algorithm is set to 0.3, and the minimum number of samples is set to 50. Through clustering analysis, the adaptive learning unit identifies user-preferred viewpoint parameter patterns, with each pattern represented by the viewpoint parameter value of its cluster center.

[0109] The number of identified patterns is typically 3 to 5, such as quick browsing mode, detailed observation mode, and dynamic tracking mode. The adaptive learning unit then dynamically updates the weight values ​​of the policy hash table in the viewpoint decision module based on the frequency of occurrence and recent usage time of each pattern.

[0110] The weight update formula is: the new weight equals the old weight multiplied by 0.9 plus the current frequency factor multiplied by 0.1.

[0111] The frequency factor is the ratio of the number of times a pattern appears in the most recent 1000 records to the total number of records.

[0112] By updating the weights, the decision-making behavior of the perspective decision-making module gradually adapts to individual user differences, making the system more in line with user habits.

[0113] The system in this embodiment also integrates a delayed decision-making mechanism, which is implemented by a dedicated state machine within the viewpoint decision-making module.

[0114] When the intent confidence level output by the intent parsing module is lower than the preset threshold of 0.75, the viewpoint decision module enters a waiting state.

[0115] While in the waiting state, the perspective decision module continuously collects subsequent perception data frames and reconfirms the intent for each frame of data.

[0116] The viewpoint decision module generates the final target viewpoint parameters and sends them to the camera driver module only when the intent classification results of three consecutive frames of data are consistent and the confidence level of each frame exceeds the threshold of 0.75.

[0117] The maximum duration of the waiting state is 500ms. If the conditions are not met within the timeout period, the viewpoint decision module will use the default viewpoint parameters, namely a focal length of 35mm, a field of view of 60°, and an exposure time of 16ms.

[0118] The delayed decision-making mechanism effectively avoids frequent camera malfunctions caused by momentary interference or misjudgment of intent, improving the system's decision-making robustness and stability in complex environments.

[0119] Data communication between the various modules of the system adopts a unified communication protocol. The data frame format includes a frame header, data length, module identifier, timestamp, data payload, and cyclic redundancy check code.

[0120] The frame header is a fixed 2-byte value 0xAA55, the data length is a 2-byte unsigned integer, the module identifier is a 1-byte enumeration value, the timestamp is a 4-byte millisecond-level time, the data payload is a variable-length byte array, and the cyclic redundancy check code is a 2-byte CRC16 check value.

[0121] The communication rate is set to 1 Mbit / s, and the error retransmission mechanism adopts the automatic repeat request protocol, with a maximum of 3 retransmissions.

[0122] The system power management adopts dynamic voltage and frequency adjustment technology, which dynamically adjusts the operating voltage and clock frequency according to the real-time calculated load of each module. Under typical usage scenarios, the overall power consumption of the system is less than 350mW, which can support the smart glasses to work continuously for more than 8 hours.

[0123] Please refer to the attached document.Figure 4 This embodiment further elaborates on the multi-level interaction relationship and data flow details between the eye-tracking module, the head posture perception module, and the environment perception module.

[0124] After extracting the pixel coordinates of the pupil center, the pupil center localization unit of the eye-tracking module not only performs eyeball model transformation, but also outputs pupil diameter change data.

[0125] The pupil diameter was calculated by averaging the major and minor axes of an ellipse obtained through an ellipse fitting algorithm, with a sampling frequency of 120Hz.

[0126] Pupil diameter data is transmitted to the intent parsing module in 32-bit floating-point format as an auxiliary assessment indicator of user cognitive load.

[0127] When the pupil diameter changes by more than 15% within 1 second, the intent parsing module lowers the current intent confidence by 0.1 as a signal of increased decision uncertainty.

[0128] The nine-axis inertial measurement unit of the head posture sensing module not only outputs the head posture quaternion, but also provides the jerk value of the head motion acceleration.

[0129] The Jerk value is the derivative of the acceleration, calculated using the central difference method from five consecutive frames of acceleration data, and its unit is m / s². 3 .

[0130] When the jerk value of head movement exceeds 10 m / s 3 At this time, the head posture perception module sends a motion change flag to the intent parsing module. After receiving the flag, the intent parsing module temporarily adjusts the window size of its sliding window mean filter from 5 frames to 3 frames to speed up the response to sudden head movements.

[0131] When extracting the 3D bounding boxes of key objects, the edge computing unit of the environment perception module adds a surface material classification function.

[0132] Surface material classification is based on infrared reflection intensity captured by a depth camera and texture features captured by a visible light camera, and is achieved through a support vector machine classifier.

[0133] The classifier was pre-trained on datasets of 10 common materials, including metal, plastic, glass, wood, and fabric.

[0134] The material classification results are output as 1-byte enumeration values, serving as an additional dimension to the intent feature vector.

[0135] When the surface material of an object is detected to be highly reflective, the viewpoint decision module introduces an additional compensation coefficient of 0.8 when calculating the exposure parameters to prevent overexposure.

[0136] The intent parsing module uses a multi-level collision detection algorithm to determine whether the gaze point falls on the surface of a specific object.

[0137] The first level is a coarse axis-aligned bounding box detection to quickly eliminate obviously irrelevant objects; the second level is directional bounding box detection to improve detection accuracy; the third level is a precise ray triangular facet intersection detection based on depth images to ensure accurate determination of the relationship between the gaze point and the object surface.

[0138] The multi-level detection algorithm ensures accuracy while keeping the average processing time within 5ms.

[0139] The perspective decision-making module introduces a reinforcement learning mechanism into its rule-based and data-driven hybrid decision-making logic.

[0140] The reinforcement learning agent takes the intent feature vector and the environmental state as input, adjusts the viewpoint parameters as actions, and uses the stability of the user's gaze point in the next 3 seconds as a reward signal.

[0141] The reward function is defined as the reciprocal of the variance of the fixation point location; the smaller the variance, the higher the reward.

[0142] The reinforcement learning agent employs a deep Q-network structure. The network input layer consists of 10 dimensions of the intent feature vector plus 5 dimensions of the environment state. The hidden layers are two fully connected layers, each with 128 neurons. The output layer consists of 5 discretized viewpoint adjustment actions. The deep Q-network updates its weights every 24 hours, using data from the most recent 1000 decision records stored in the adaptive learning unit.

[0143] The camera driver module employs adaptive microstepping control technology when driving the micro stepper motor assembly. It dynamically adjusts the microstepping subdivision factor based on the deviation between the target angle and the current angle.

[0144] When the deviation is greater than 5°, the full step mode is used with a subdivision factor of 1; when the deviation is between 1° and 5°, the 4-subdivision microstep mode is used; when the deviation is less than 1°, the 16-subdivision microstep mode is used.

[0145] Adaptive microstepping control technology ensures fast response while keeping the operating noise of the stepper motor below 35dB.

[0146] When building a database of user-personalized adjustment preferences, the adaptive learning unit adds scene context labels.

[0147] Scene context is obtained through joint analysis of depth images and visible light images from the environment perception module, and scene classification is performed using a convolutional neural network.

[0148] Scene categories include five main categories: indoor office, outdoor street, indoor home, outdoor nature, and transportation.

[0149] During cluster analysis, the adaptive learning unit prioritizes clustering records within the same scene context to form a scene-specific viewpoint parameter pattern library.

[0150] When the system detects that a user has entered a specific scene, the viewpoint decision module prioritizes loading the viewpoint parameter mode corresponding to that scene, enabling rapid adjustment for scene adaptation.

[0151] The delayed decision-making mechanism introduces multiple hypothesis tracking technology while waiting for confirmation of intent.

[0152] When the intent confidence output by the intent parsing module is lower than the threshold but there are multiple candidate intents, the viewpoint decision module tracks the development trend of all candidate intents in parallel.

[0153] Each candidate intent is assigned a hypothesis weight, which is dynamically updated based on the confidence level changes in consecutive frames.

[0154] When the weight of a certain hypothesis continues to increase over three consecutive frames and eventually exceeds 0.8, the viewpoint decision module will adopt that hypothesis as the final decision, even if the confidence levels of other hypotheses do not fully meet the conditions.

[0155] Multihypothesis tracking technology reduces the system's decision latency in fuzzy intent scenarios from an average of 500ms to 300ms.

[0156] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0157] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A camera adjustment system for smart glasses, characterized in that, Comprise: Eye tracking module for real-time acquisition of user eye movement data and extraction of gaze point coordinates; Head pose perception module for obtaining user head orientation, pitch and roll angle data in three-dimensional space; Environment perception module for collecting scene depth information, ambient light intensity and key object contour data; Intention analysis module based on gaze point coordinates output by eye tracking module, head pose data output by head pose perception module and environment data collected by environment perception module, through multi-source information fusion calculation to generate a quantitative description of the user's current visual intent; View decision module receives the visual intent description output by the intention analysis module, combines the preset view adjustment strategy set, and decides to generate target view parameters; Camera driving module drives the camera of the intelligent glasses to perform corresponding optical zoom, mechanical rotation or electronic anti-shake operation according to the target view parameters issued by the view decision module; The view decision module has embedded hybrid decision logic based on rules and data driving. The view decision module pre-stores multiple view adjustment strategies, each strategy is associated with a specific intent type and environmental condition. When the object gaze intent is output by the intention analysis module, the view decision module dynamically calculates the angular velocity and acceleration required for the camera to track the target object according to the motion state of the target object, and optimizes the camera exposure time and gain parameters according to the target object reflectivity and ambient light intensity. When the user's browsing intent is recognized, the view decision module adaptively adjusts the camera field of view angle to match the user's visual scanning rhythm according to the covariance relationship between the head rotation angular velocity and the gaze point movement speed.

2. The camera adjustment system of claim 1, wherein, The eye tracking module includes an infrared light source array, a corneal reflection image sensor, and a pupil center positioning unit. The infrared light source array projects invisible infrared spots onto the user's eyeball. The corneal reflection image sensor captures the infrared reflection image on the eyeball surface. The pupil center positioning unit performs grayscale and binary processing on the reflection image, then uses an ellipse fitting algorithm to extract the pupil center pixel coordinates, and converts the pixel coordinates into a three-dimensional gaze vector based on a pre-calibrated eyeball model.

3. The camera adjustment system of claim 1, wherein, The head pose perception module integrates a nine-axis inertial measurement unit, which includes a three-axis gyroscope, a three-axis accelerometer, and a three-axis magnetometer. The gyroscope outputs angular velocity data, the accelerometer outputs linear acceleration data, and the magnetometer outputs geomagnetic field intensity data. Through Kalman filtering algorithm, the above data are fused and solved to eliminate cumulative errors and output high-precision head pose quaternion representation.

4. The camera adjustment system of claim 1, wherein, The environment perception module includes a depth camera, an ambient light sensor, and an edge computing unit. The depth camera uses the time-of-flight principle to obtain the physical distance information of each point in the scene from the glasses. The ambient light sensor monitors the ambient light intensity value in real time. The edge computing unit performs plane segmentation and object clustering processing on the depth image, and extracts the three-dimensional bounding box and surface texture features of the key objects in the user's gaze area.

5. The camera adjustment system of claim 1, wherein, The intention analysis module performs the following processing flow: First, the gaze point coordinate sequence is subjected to sliding window mean filtering to eliminate physiological tremor noise; Secondly, the gaze point is converted from the eye coordinate system to the world coordinate system combined with the head pose data; Then, according to the depth information and the object contour data provided by the environment perception module, it is determined whether the gaze point falls on the surface of a specific object; If the object is determined to be gazed, the motion speed, surface reflectivity and relative distance of the object to the user are extracted as the intention feature vector; If no specific object is detected, the user is identified as a browsing intention or a search intention according to the smoothness and speed change pattern of the gaze point movement trajectory.

6. The camera adjustment system of smart glasses of claim 1, wherein, The camera driving module includes a digital signal processor and a micro stepping motor set; After receiving the target view angle parameters issued by the view angle decision module, the digital signal processor decomposes the parameters into focal length adjustment, gimbal deflection angle and image stabilization compensation; The micro stepping motor set drives the lens zoom ring, gimbal mechanism and optical anti-shake component to move cooperatively according to the decomposed instructions, so as to realize the real-time alignment of the camera optical axis and the user's line of sight vector.

7. The camera adjustment system of smart glasses of claim 1, wherein, The system is also provided with an adaptive learning unit; The unit continuously records the intention feature vectors generated by the intention analysis module and the final adopted view angle parameters by the view angle decision module in different scenarios, and constructs a user personalized adjustment preference database; Through periodic clustering analysis, the adaptive learning unit identifies the user's habit view angle parameter mode and dynamically updates the strategy weight in the view angle decision module, so that the system adjustment behavior gradually adapts to the individual differences of the user.

8. The camera adjustment system of smart glasses of claim 7, wherein, The periodic clustering analysis adopts a density-based noise application spatial clustering algorithm; the clustering analysis is executed once at a fixed time, and the identified view angle parameter mode includes a fast browsing mode, a fine observation mode and a dynamic tracking mode. 9.A camera adjustment method of smart glasses, characterized in that, The camera adjustment system of the smart glasses is used to realize the camera adjustment of the smart glasses.

Citation Information

Patent Citations

  • An adaptive adjusting method, a pair of AR intelligent learning glasses and an AR intelligent learning system

    CN108933916A

  • Synchronous control system for shooting visual angle of intelligent glasses and intelligent glasses

    CN220419683U