Eye using habit monitoring method and system

By using a camera to collect facial images in eye habit monitoring and combining Sobel template enhancement, perspective projection and PnP algorithms, the problem of insufficient accuracy of eye habit monitoring in the prior art is solved, high-precision and real-time eye habit monitoring is achieved, and early warning capabilities for eye use health are enhanced.

CN120014690APending Publication Date: 2025-05-16SHENZHEN SUNCHIP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510082504.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing eye habit monitoring method roughly detects the head contour through the object detection algorithm, which leads to the misjudgment of the continuous eye use state when the head is gaze or moves, resulting in insufficient monitoring accuracy and frequent false alarms.

Method used

Facial images are collected using a camera and image enhancement is performed through Sobel templates with curvature weight factors to accurately determine the eye distance and head posture. Combining the nonlinear least squares optimization algorithm and the PnP algorithm, the head posture and gaze direction are estimated in real time, the gaze point is generated and its maintenance duration is monitored.

Benefits of technology

It improves the accuracy and real-time nature of eye habit monitoring, reduces false alarm situations, provides multi-dimensional monitoring (distance, attitude, and gaze time), greatly enhances the accuracy of early warning for eye health, and ensures eye safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014690A_ABST
    Figure CN120014690A_ABST
Patent Text Reader

Abstract

The invention provides an eye using habit monitoring method and system, and relates to the technical field of data processing, and the method comprises the steps: collecting a face image of a target user through a camera located on a desktop; performing image enhancement on the face image; determining an eye using distance between the target user and a desktop in the enhanced face image by taking a real face parameter of the target user as a reference object; extracting real-time facial features of the target user from the enhanced facial image; based on the real-time facial features, determining a real-time head posture of the target user in combination with a nonlinear least square optimization algorithm and a PnP algorithm; determining a gazing direction of the target user according to the head posture; in combination with the eye using distance, projecting the annotation direction to a desktop, generating a fixation point, and monitoring the maintenance duration of the fixation point in a preset range; and when the eye using distance exceeds the preset eye using distance or the maintaining duration exceeds the preset maintaining duration, sending out early warning information. And high-precision monitoring of eye using habits can be automatically completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to an eye habit monitoring method and system. Background Art

[0002] Eye habits refer to the way people use their eyes in daily life, study or work, including the distance of use, duration of continuous use of eyes, gaze direction and posture, etc. These habits directly affect vision health and eye comfort.

[0003] Bad eye habits (such as long-term close-up use of the eyes or incorrect gaze posture) are the main causes of vision problems such as myopia and dry eyes, especially among teenagers. Through eye habit monitoring, eye behavior can be monitored in real time, and incorrect eye posture or distance can be corrected in time, thereby effectively preventing vision damage and helping to develop a healthy eye use method. This not only helps to protect eye health, but also improves learning and work efficiency, which has important practical significance.

[0004] However, existing eye habit monitoring methods often track the target user's eye habits by roughly detecting the head contour through target detection algorithms. This method can easily be judged as the same behavior during head gaze or activity, that is, the eyes are always in use. This leads to insufficient accuracy in eye habit monitoring, frequent false alarms, and insufficient practicality. Summary of the invention

[0005] In order to solve the technical problems that the existing eye habit monitoring methods in the prior art often use target detection algorithms to roughly detect the head contour to track the eye habits of the target user, which can easily be judged as the same behavior during head gaze or activity, that is, the eyes are always in use, which leads to insufficient accuracy of eye habit monitoring, frequent false alarms and insufficient practicality, the present invention provides an eye habit monitoring method and system.

[0006] The technical solution provided by the embodiment of the present invention is as follows:

[0007] First aspect

[0008] An embodiment of the present invention provides an eye habit monitoring method, comprising:

[0009] S1: Use a camera located on the desktop to capture the target user's facial image;

[0010] S2: Perform image enhancement on the facial image using a Sobel template with a curvature weight factor to obtain an enhanced facial image;

[0011] S3: Using the target user's real facial parameters as a reference, the eye distance between the target user and the desktop is determined in the enhanced facial image based on the perspective projection principle;

[0012] S4: extracting real-time facial features of the target user from the enhanced facial image in sequence, wherein the facial features include a left eye position, a right eye position, and a mouth position;

[0013] S5: Based on real-time facial features, the real-time head posture of the target user is determined by combining the nonlinear least squares optimization algorithm and the PnP algorithm;

[0014] S6: Determine the gaze direction of the target user based on the head posture;

[0015] S7: Combined with the eye distance, the annotation direction is projected onto the desktop to generate the fixation point, and the duration of the fixation point within the preset range is monitored;

[0016] S8: When the eye use distance exceeds the preset eye use distance or the maintenance time exceeds the preset maintenance time, a warning message is issued.

[0017] Second aspect

[0018] An embodiment of the present invention provides an eye habit monitoring system, comprising:

[0019] processor;

[0020] A memory having computer-readable instructions stored therein, which, when executed by a processor, implements the eye-use habit monitoring method of the first aspect.

[0021] The third aspect

[0022] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the eye habit monitoring method of the first aspect is implemented.

[0023] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0024] In an embodiment of the present invention, a facial image is collected by a camera, and the image is enhanced in combination with a Sobel template with a curvature weight factor, thereby improving the detection accuracy of the curved lines on the face. On this basis, the eye distance is calculated using perspective projection, and the head posture and gaze direction are accurately estimated based on real-time facial features including the left eye position, the right eye position, and the mouth position, combined with the PnP algorithm and the nonlinear least squares optimization algorithm, effectively avoiding frequent false alarms and low-precision monitoring results caused by only detecting the head posture, and generating the gaze point in real time and monitoring its maintenance duration. It can automatically complete high-precision monitoring of the target user's eye habits, significantly improve real-time and adaptability, and provide multi-dimensional monitoring (distance, posture, gaze time), which greatly enhances the accuracy of early warning of eye health and ensures eye safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0026] Figure 1 A schematic diagram of a flow chart of an eye habit monitoring method provided by an embodiment of the present invention;

[0027] Figure 2 A schematic diagram of the structure of an eye habit monitoring system provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0029] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0030] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0031] Reference Manual Attached Figure 1 , showing a flow chart of an eye habit monitoring method provided by an embodiment of the present invention.

[0032] The embodiment of the present invention provides an eye habit monitoring method, which can be implemented by an eye habit monitoring device, which can be a terminal or a server. The processing flow of the eye habit monitoring method may include the following steps:

[0033] S1: Use a camera located on the desktop to capture the facial image of the target user.

[0034] The desktop refers to a flat table used for study or work, which is usually used to place books, electronic devices (such as computers, tablets) and other learning tools, and is also used to place cameras to capture the user's facial image. The camera is placed on the desktop. The camera collects the user's facial image when studying or working, providing basic data for subsequent image processing and eye behavior monitoring. The position and angle of the camera are reasonably designed to ensure that facial features are clearly visible, thereby improving the accuracy and comprehensiveness of monitoring.

[0035] In a possible implementation, the camera is located at an end of the desktop far away from the target user.

[0036] It is understandable that placing the camera on the end of the desktop away from the user can effectively capture the complete image of the user's face and avoid the viewing angle limitation and feature deformation caused by too close distance, thereby improving the precision of image acquisition and the accuracy of monitoring, while reducing the user's sense of interference and improving the user experience.

[0037] S2: The facial image is enhanced by using a Sobel template with a curvature weight factor to obtain an enhanced facial image.

[0038] Among them, the curvature weight factor is a parameter that is dynamically adjusted according to the degree of curvature of the local area of ​​the image, and is used to enhance the detection ability of curved lines. By calculating the image brightness change rate (gradient) and the local curvature (curvature), the weight factor is increased in the area with significant curvature, so that the image processing is more sensitive to the curved features and can more accurately monitor the curved lines. The Sobel template is an operator used for image edge detection, which is often used to extract the horizontal and vertical edge information of the image. The traditional Sobel template is sensitive to straight line edges, but not accurate enough for arc features. This method combines the curvature weight factor to modify the Sobel template, making it more suitable for the enhancement of arc features. By introducing the curvature weight factor into the Sobel template, the collected facial image is enhanced, making the arc features (such as facial contours, eyes, etc.) more prominent. This enhancement process provides high-precision basic data for the subsequent extraction of key facial features and accurate calculation of head posture, which helps to improve the reliability and accuracy of eye behavior monitoring.

[0039] In a possible implementation manner, S2 is specifically:

[0040] S201: Determine a curvature weight factor by combining the image brightness change rate of the facial image and the local curvature degree of the image:

[0041]

[0042] Wherein, w represents the curvature weight factor, x' and y' represent the image brightness change rate of the facial image on the x-axis and y-axis respectively, and x" and y" represent the local curvature degree of the facial image on the x-axis and y-axis respectively.

[0043] Among them, the image change rate is the rate of change of the brightness value with the position, reflecting the gradient direction and describing the location of the edge. The local curvature is the curvature of the edge or contour, reflecting whether the edge is a straight line or an arc. The numerator describes the combined changes of the horizontal and vertical gradients, which is used to judge the curved shape of the edge. It captures the contribution of the horizontal and vertical gradient changes to the arc. The denominator is used for normalization to avoid the influence of the brightness value itself on the calculated curvature. It ensures that the curvature value is independent of the absolute size of the image brightness, but reflects the shape characteristics of the edge itself. The curvature w represents the shape characteristics of the local edge or line. The larger the value, the more significant the curvature at that point.

[0044] S202: Determine a Sobel template having a curvature weight factor using the curvature weight factor:

[0045] S arc =S sobel (1+w)

[0046] Among them, S arc represents the Sobel template with curvature weight factor, S sobel Represents a horizontal template.

[0047] The horizontal template is as follows:

[0048]

[0049] S203: Perform image enhancement on the facial image using a Sobel template with a curvature weight factor to obtain an enhanced facial image:

[0050] I enhanced =I+G

[0051]

[0052] G x =I·S arc ,G y =I·S arc

[0053] Among them, I and I enhanced represent the facial image and enhanced facial image respectively, G represents the enhanced gradient amplitude, Gx and G y They represent the horizontal enhancement gradient and the vertical enhancement gradient respectively.

[0054] It should be noted that this process achieves facial image enhancement through three steps. First, the curvature weight factor is calculated using the image brightness change rate and the local curvature degree to dynamically reflect the curvature characteristics of the edge, making the detection of arcs more sensitive. Then, based on the calculated w, the traditional Sobel template is modified to obtain a Sobel template with a curvature weight factor to improve the detection ability of arc features. Finally, the image is convolved to calculate the horizontal enhancement gradient and the vertical enhancement gradient to generate an enhanced facial image. The advantage of this process is that it can significantly improve the detection accuracy of arc features (such as eyes and facial contours), avoid the defect of traditional methods that are insensitive to arc detection, and provide higher quality data support for subsequent feature extraction and head posture estimation.

[0055] S3: Using the target user's real facial parameters as a reference, the eye distance between the target user and the desktop is determined in the enhanced facial image based on the perspective projection principle.

[0056] Among them, the real facial parameters refer to the physical size parameters of the target user's face (such as facial height, width, etc.). These parameters are known fixed values ​​and are used as reference standards for measurement in the image to help deduce the distance. Perspective projection refers to the imaging law of a three-dimensional object on a two-dimensional plane. Its characteristic is that the farther the object is, the smaller the image is. Based on this principle, the actual object distance can be calculated through parameters such as the facial size (pixel value) and camera focal length in the image. Eye distance refers to the actual distance between the target user's eyes and the desktop, which is an important parameter affecting eye health. This method measures this distance by combining facial parameters with enhanced images.

[0057] By combining the user's real facial parameters with the facial size in the enhanced image and based on the principle of perspective projection, the user's eye distance from the desktop is accurately calculated. This step provides basic data for real-time monitoring of the user's eye health, ensuring that the measurement results are accurate and reliable, thereby better assisting subsequent warning functions.

[0058] In a possible implementation, the real facial parameter is specifically the facial height, and the eye distance is calculated specifically as follows:

[0059]

[0060] Where d is the eye distance, h is real Indicates the real face height, h img Represents the face height pixel value in the enhanced face image, H imgrepresents the total height of the enhanced facial image, f represents the focal length of the camera, H sensor Indicates the camera sensor height.

[0061] The camera sensor height refers to the physical height of the camera sensor (in millimeters). It is used to calculate the physical size of the pixel in the real world, thereby converting the measured value (pixel) on the image into the real physical size. This is one of the key parameters for calculating the distance or object size, which can be found in conjunction with the camera specifications. More specifically, when you want to get the pixel height H from the image, img When calculating the height or distance of an object in the real world, you must know the size of each pixel in the physical world. The physical height H of the sensor sensor and the image resolution H img There is a direct relationship between (the pixel height of the image) and can be used to calculate: the physical height of each pixel is specifically

[0062] It should be noted that this process uses the perspective projection principle of the camera, combined with the real height of the face and the height of the facial pixels in the image, to infer the user's eye distance. The pixel values ​​in the image are converted into actual physical sizes through the focal length of the camera and the physical height of the sensor to ensure the accuracy of the measurement. The advantage is that this method makes full use of camera parameters and image information, can accurately calculate the eye distance, does not require additional complex equipment, has strong adaptability, and provides reliable basic data for monitoring user eye behavior.

[0063] S4: extracting real-time facial features of the target user from the enhanced facial image in chronological order, wherein the real-time facial features include a left eye position, a right eye position, and a mouth position.

[0064] It should be noted that by extracting features from the enhanced facial image, the user's left eye, right eye and mouth positions are identified in sequence. These facial features are key geometric points that can provide accurate basic data for subsequent head posture estimation and gaze direction calculation, ensuring the stability and accuracy of the monitoring process.

[0065] In a possible implementation, S4 specifically includes:

[0066] S401: Detect the left eye position and the right eye position of the target user through a key point detection model.

[0067] Optionally, the key point detection algorithm is the Dlib 68-point model.

[0068] Among them, the Dlib 68-point model is an algorithm model widely used for facial key point detection, developed based on machine learning and statistical methods. The model can locate 68 specific facial key points in a given facial image, including the eyes, eyebrows, nose, mouth, and feature points of the facial contour. The position of each point is learned by the model through training data, and can accurately mark standard facial features.

[0069] Optionally, the left eye position and the right eye position can also be located by the valley field. Specifically, first, the valley points in the enhanced facial image are detected. The grayscale value of the human iris is usually lower (darker) than the surrounding area, which appears as a valley in the image. Morphological operations are used to extract the valley field and find the low grayscale area in the image. After that, the eye candidate areas are screened. The grayscale value and valley field value of the candidate points must meet certain threshold conditions. The first threshold condition is that the grayscale value of the candidate point must be lower than the first threshold, that is, it must be a "dark point". This is because the grayscale value of the human iris is lower than the surrounding area. The second threshold condition is that the valley field value must be greater than the second threshold, that is, the point must show a strong "valley" feature (representing a morphological depression). The candidate areas are further scored by functions and. These functions combine weighted grayscale values ​​and morphological calculation results. The functions are specifically:

[0070]

[0071] Among them, f(x-2,y), f(x+2,y), f(x-3,y) and f(x+3,y) represent the pixel grayscale values ​​on both sides of the candidate point (x,y), at a distance of 2 pixels and 3 pixels respectively. They are used to estimate whether a point is located in a local extreme area (such as the bottom of a valley). They represent the local window average grayscale value of (x,y), calculated in the range of 3x3 and 5x5 windows respectively, and are used to represent the local background grayscale. The purpose is to compare whether the grayscale of the candidate point is darker than the surrounding background. Φ 1,1 (x,y) and Φ 2,1 (x, y) represent the average values ​​within the 3x3 and 5x5 windows, respectively, representing the local structural morphological features. They capture the morphological “valley” characteristics of the candidate points (i.e., possible eye region features). 1,1 , W 1,2 , W 2,1 and W 2,2 They all represent corresponding weighting factors, which are used to adjust the contribution ratio of each feature to the scoring result.

[0072] Then, the facial candidate region is generated. Based on the geometric head model (such as the distance between the two eyes and other information) and the position of the eye candidate points, the possible face area is inferred, and the candidate area is compared with the face template to verify whether it belongs to a face. Finally, the face in the image is identified, and the key feature eyes are extracted to provide a basis for further face recognition or expression analysis. The left and right eye positions are extracted by the valley field, and the low grayscale characteristics and morphological concave features of the human iris can be used to more accurately locate the eye area. The combination of grayscale value and structural features for screening reduces background interference, improves the robustness of the algorithm and its adaptability to complex lighting and occlusion conditions, and provides a more stable feature basis for subsequent face analysis.

[0073] S402: Under the facial geometry constraint, determine the mouth position according to the left eye position and the right eye position:

[0074]

[0075] in, and They represent the left eye position coordinates, right eye position coordinates and mouth position coordinates in the enhanced facial image, respectively. y Indicates the vertical offset of the mouth position.

[0076] Optionally, the offset may be 1.5 times the interocular distance.

[0077] Specifically, the method of extracting the left and right eye features in sequence and then using geometric constraints to infer the mouth features has multiple advantages. First, the eye position is a relatively stable and obvious area in the facial features. It can be quickly and accurately located through key point detection or valley field methods, providing a basis for the subsequent inference of mouth features. Secondly, the distance and position relationship between the eyes provides a reliable reference for geometric constraints, making the calculation of the mouth position more accurate. Through this sequential processing, the error accumulation of feature extraction can be effectively reduced, and the accuracy and robustness of the overall detection can be improved, especially in complex scenes (such as uneven lighting or partial occlusion), the reliability of the mouth positioning results can be better ensured. Finally, this method can provide high-quality feature data support for subsequent head posture estimation and gaze direction judgment.

[0078] S5: Based on real-time facial features, the real-time head posture of the target user is determined by combining the nonlinear least squares optimization algorithm and the PnP algorithm.

[0079] Among them, the nonlinear least squares optimization algorithm is an algorithm that minimizes the error between the actual measurement value and the model prediction value by iterative solution. In this method, it is used to optimize the rotation matrix and translation vector in the head posture calculation to ensure that the error is minimized and improve the accuracy of posture estimation. The PnP algorithm is a three-dimensional point pose calculation method based on perspective projection, which is used to infer the pose (rotation and translation) between the camera and the object from the relationship between the known three-dimensional points and the corresponding two-dimensional image points. In this method, the PnP algorithm combines the three-dimensional model of standard facial features with the two-dimensional facial features extracted in real time to calculate the spatial pose of the user's head. Real-time head pose refers to the instantaneous state of the direction (rotation) and position (translation) of the user's head in three-dimensional space, which describes the dynamic changes of the user's head, such as head orientation and tilt angle.

[0080] It should be noted that the nonlinear least squares optimization algorithm and the PnP algorithm are combined to match the two-dimensional position of facial feature points with the standard three-dimensional model to calculate the head posture of the target user in real time. This step can accurately capture the user's head direction and position, providing reliable data support for further analysis of gaze direction and eye behavior, thereby improving the accuracy and real-time performance of monitoring.

[0081] In a possible implementation, S5 specifically includes:

[0082] S501: A conversion formula describing the conversion of facial features of a standard face model from a three-dimensional model coordinate system to a camera coordinate system is established through a PnP algorithm:

[0083] p i =R·P i +T

[0084] Among them, P i represents the 3D model point of the i-th facial feature point defined by the standard face model in the 3D model coordinate system, p i Indicates P i In the three-dimensional position in the camera coordinate system, R represents the rotation matrix describing the head posture, and T represents the translation vector describing the position of the head in the camera coordinate system.

[0085] Among them, the standard face model is an idealized face structure based on a three-dimensional coordinate system, which defines a set of fixed facial feature points (such as eyes, nose, mouth, etc.) and their positions in three-dimensional space. The distribution of these feature points is based on the average value of a large amount of face data, reflecting the geometric shape of a typical face, and providing a unified reference framework for face posture estimation, such as 3D Morphable Model, Basel Face Model or 68-point standard face model. The three-dimensional feature points in the standard face model are mapped to the camera coordinate system through the PnP algorithm to generate a rotation matrix and translation vector to describe the direction and position of the head respectively. This process establishes a transformation relationship between the three-dimensional model and the camera view, which can accurately reflect the posture of the head in the camera coordinate system, and provides an accurate data basis for the subsequent gaze direction and gaze point calculation.

[0086] S502: Establishing a projection formula for transforming facial features from the camera coordinate system to the enhanced facial image:

[0087]

[0088] Among them, (x i ,y i ,z i ) represents the 3D model point p in the camera coordinate system i The coordinates of x and f y They represent the horizontal focal length and vertical focal length of the camera, which describe the camera internal parameters. x and c y Respectively represent the horizontal pixel coordinates and vertical pixel coordinates of the camera optical center, u i and v i Represents the 2D pixel coordinates of the i-th facial feature in the enhanced facial image.

[0089] S503: Calculate the two-dimensional pixel coordinate error between the real-time facial features and the facial features in the enhanced facial image:

[0090]

[0091] Where E represents the two-dimensional pixel coordinate error value, represents the two-dimensional pixel coordinates of the i-th real-time facial feature in the enhanced facial image, n represents the number of facial features, and || || represents the two-norm operator.

[0092] Optionally, n is 3, namely the left eye position, the right eye position and the mouth position.

[0093] S504: Using a nonlinear least squares optimization algorithm, with the goal of eliminating the two-dimensional pixel coordinate error value, update the rotation matrix and translation vector describing the real-time head posture to obtain the real-time head posture:

[0094] R new =R·exp(Δθ)

[0095] T new =T+ΔT

[0096]

[0097] Among them, R new and T new They represent the updated rotation matrix and translation vector respectively, exp represents the natural exponential function, J represents the Jacobian matrix of E for R and T, and the subscript T represents the transpose. T J) -1 Indicates J T The inverse of J, Δθ and ΔT represent the rotation increment and translation increment respectively.

[0098] It is understandable that in head pose estimation, the real-time pose of the head is described by a rotation matrix and a translation vector, which together form a complete rigid body transformation that maps the standard model coordinate system to the camera coordinate system. The rotation matrix determines the orientation of the head in 3D space. The translation vector determines the position of the head in 3D space. Together, they map the standard 3D model to the camera coordinate system and ultimately project it to the image plane. The rotation matrix and the translation vector are continuously adjusted by optimizing the error, and ultimately can accurately represent the real-time head pose.

[0099] Specifically, the process achieves accurate estimation of head posture through the PnP algorithm and the nonlinear least squares optimization algorithm. First, the PnP algorithm maps the three-dimensional feature points of the standard face model to the camera coordinate system, generates the initial rotation matrix and translation vector, which are used to describe the direction and position of the head. Then, the projection formula from the camera coordinate system to the image plane is established to project the three-dimensional points into two-dimensional pixel points. Subsequently, the two-dimensional pixel error between the actual detected real-time feature points and the projection points is calculated. Finally, the nonlinear least squares optimization algorithm is used to iteratively adjust R and T to minimize the error and obtain accurate real-time head posture. The advantage of this method is that it combines the three-dimensional geometric model and the actual image data, improves the accuracy and robustness of posture estimation through error optimization, can adapt to posture changes in complex scenes, and provide reliable basic data for gaze direction and eye monitoring.

[0100] S6: Determine the gaze direction of the target user according to the head posture.

[0101] Among them, the gaze direction refers to the direction of sight corresponding to the direction of the user's head, which is calculated by the head posture (rotation matrix and translation vector) and the center point of the eye position. It reflects the extended trajectory of the user's sight and can be used to determine the user's focus point or the area of ​​​​attention. By using the head posture information to calculate the user's gaze direction, the user's sight trajectory can be accurately predicted. This step can identify the specific location of the user's gaze, provide key data support for eye distance analysis and gaze point monitoring, and ensure the accuracy and dynamic adaptability of eye behavior monitoring.

[0102] In a possible implementation manner, the gaze direction is determined as follows:

[0103] L(t)=P eye +t·d head ,t>0

[0104]

[0105] Among them, d head represents the orientation vector of the head, and t represents the control annotation ray along d head Extension parameter of the extension scale, P eye represents the position of the eye center point, and L(t) represents the gaze ray associated with t representing the annotation direction.

[0106] It should be noted that by combining the updated rotation matrix and translation vector to determine the gaze direction, the spatial orientation of the head and the line of sight trajectory of the center of the eye position can be accurately reflected. The head orientation vector provides the direction of the line of sight, while the center of the eye determines the starting position of the line of sight. This method comprehensively considers the head posture and eye position, can dynamically adapt to the rotation and translation changes of the head, improves the accuracy and real-time performance of the gaze direction calculation, and provides a more stable foundation for gaze point projection and eye behavior analysis.

[0107] S7: Based on the eye distance, the annotation direction is projected onto the desktop to generate the fixation point, and the duration of the fixation point within the preset range is monitored.

[0108] Among them, the gaze point refers to the intersection of the extension line of the user's gaze direction and the desktop, reflecting the point where the user's line of sight falls on the actual desktop. It is a specific coordinate calculated in combination with the eye distance and the gaze direction, and is used to track the user's gaze behavior. The user's gaze direction is projected in combination with the eye distance, the user's gaze point on the desktop is calculated, and the time the point stays within the preset range is monitored in real time. Through this process, it can be determined whether the user has been staring at a certain area for a long time or the distance is too close, providing a scientific basis for the analysis of eye health behavior and laying the foundation for subsequent early warnings.

[0109] It should be noted that those skilled in the art can set the size of the preset range according to actual needs, and the present invention does not limit this. Optionally, the preset range can be set according to the size of the desktop, and can be the range of the desktop or a range smaller than the desktop.

[0110] In a possible implementation, the gaze point is specifically:

[0111]

[0112] Among them, P gaze represents the gaze point coordinates, d head,z Indicates d head The z component of , d represents the eye distance.

[0113] It should be noted that the gaze point is determined by combining the position of the center of the eye and the z component of the head orientation vector to achieve accurate positioning of the gaze point. This method fully considers the changes in head posture and the three-dimensional characteristics of the gaze direction, and extends the line of sight to the desktop for accurate projection. The advantage is that it can dynamically calculate the specific position of the user's gaze, adapt to different eye distances and changes in head posture, ensure the accuracy and reliability of monitoring, and provide a scientific basis for eye behavior analysis and health intervention.

[0114] S8: When the eye use distance exceeds the preset eye use distance or the maintenance time exceeds the preset maintenance time, a warning message is issued.

[0115] It should be noted that those skilled in the art can set the preset eye distance and the preset maintenance time according to actual needs, and the present invention is not limited thereto.

[0116] In the actual application process, first, the user's facial image is collected through a camera, and the Sobel template with curvature weight factor is used to enhance the recognition of arc features in the image. Then, based on the principle of perspective projection, the eye distance between the user and the desktop is calculated, key facial features are extracted in sequence, and the head posture is accurately estimated by combining the PnP algorithm and the nonlinear least squares optimization algorithm. Next, the gaze direction is determined according to the head posture, and the gaze direction is projected onto the desktop to generate the gaze point and its maintenance duration is monitored. Finally, when the eye distance or gaze duration exceeds the preset range, an early warning message is issued. This method has the advantages of high precision, dynamic adaptability and multi-dimensional monitoring, which can effectively help users improve their eye habits and protect their vision health.

[0117] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0118] In an embodiment of the present invention, a facial image is collected by a camera, and the image is enhanced in combination with a Sobel template with a curvature weight factor, thereby improving the detection accuracy of the curved lines on the face. On this basis, the eye distance is calculated using perspective projection, and the head posture and gaze direction are accurately estimated based on real-time facial features including the left eye position, the right eye position, and the mouth position, combined with the PnP algorithm and the nonlinear least squares optimization algorithm, effectively avoiding frequent false alarms and low-precision monitoring results caused by only detecting the head posture, and generating the gaze point in real time and monitoring its maintenance duration. It can automatically complete the high-precision monitoring of the target user's eye habits, and the real-time and adaptability are significantly improved. At the same time, it provides multi-dimensional monitoring (distance, posture, gaze time), which greatly enhances the accuracy of early warning for eye health and ensures eye safety.

[0119] Reference Manual Attached Figure 2 , showing a structural schematic diagram of an eye habit monitoring system provided by the present invention.

[0120] The present invention further provides an eye habit monitoring system 20, which is applied to the above-mentioned eye habit monitoring method, and comprises:

[0121] Processor 201.

[0122] The memory 202 stores computer-readable instructions, and when the computer-readable instructions are executed by the processor 201, the eye-use habit monitoring method of the method embodiment is implemented.

[0123] The eye habit monitoring system 20 provided by the present invention can execute the above-mentioned eye habit monitoring method and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.

[0124] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0125] In an embodiment of the present invention, a facial image is collected by a camera, and the image is enhanced in combination with a Sobel template with a curvature weight factor, thereby improving the detection accuracy of the curved lines on the face. On this basis, the eye distance is calculated using perspective projection, and the head posture and gaze direction are accurately estimated based on real-time facial features including the left eye position, the right eye position, and the mouth position, combined with the PnP algorithm and the nonlinear least squares optimization algorithm, effectively avoiding frequent false alarms and low-precision monitoring results caused by only detecting the head posture, and generating the gaze point in real time and monitoring its maintenance duration. It can automatically complete the high-precision monitoring of the target user's eye habits, and the real-time and adaptability are significantly improved. At the same time, it provides multi-dimensional monitoring (distance, posture, gaze time), which greatly enhances the accuracy of early warning for eye health and ensures eye safety.

[0126] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0127] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).

[0128] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0129] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0130] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0131] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0132] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0133] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0134] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0135] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0137] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0138] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the eye habit monitoring method of the method embodiment is implemented.

[0139] A computer-readable storage medium provided by the present invention can implement the steps and effects of the eye habit monitoring method of the above method embodiment. To avoid repetition, the present invention will not go into details.

[0140] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0141] In an embodiment of the present invention, a facial image is collected by a camera, and the image is enhanced in combination with a Sobel template with a curvature weight factor, thereby improving the detection accuracy of the curved lines on the face. On this basis, the eye distance is calculated using perspective projection, and the head posture and gaze direction are accurately estimated based on real-time facial features including the left eye position, the right eye position, and the mouth position, combined with the PnP algorithm and the nonlinear least squares optimization algorithm, effectively avoiding frequent false alarms and low-precision monitoring results caused by only detecting the head posture, and generating the gaze point in real time and monitoring its maintenance duration. It can automatically complete high-precision monitoring of the target user's eye habits, significantly improve real-time and adaptability, and provide multi-dimensional monitoring (distance, posture, gaze time), which greatly enhances the accuracy of early warning of eye health and ensures eye safety.

[0142] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0143] There are a few points to note:

[0144] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.

[0145] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.

[0146] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.

[0147] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A method for monitoring eye habits, characterized in that the method include: S1: Use a desktop camera to capture the target user's facial image; S2: performing image enhancement on the facial image by using a Sobel template with a curvature weight factor to obtain an enhanced facial image; S3: using the real facial parameters of the target user as a reference, determining the eye distance between the target user and the desktop in the enhanced facial image based on the perspective projection principle; S4: extracting real-time facial features of the target user from the enhanced facial image in chronological order, wherein the real-time facial features include a left eye position, a right eye position, and a mouth position; S5: Based on the real-time facial features, a nonlinear least squares optimization algorithm and a PnP algorithm are combined to determine the real-time head posture of the target user; S6: Determine the gaze direction of the target user according to the head posture; S7: combining the eye distance, projecting the annotation direction onto the desktop, generating a fixation point, and monitoring the duration of the fixation point within a preset range; S8: When the eye use distance exceeds the preset eye use distance or the maintenance time exceeds the preset maintenance time, a warning message is issued.

2. The eye habit monitoring method according to claim 1, characterized in that: The camera is specifically located at an end of the desktop far away from the target user.

3. The eye habit monitoring method according to claim 1, characterized in that: The S2 is specifically: S201: Determine the curvature weight factor by combining the image brightness change rate of the facial image and the local curvature degree of the image: Wherein, w represents the curvature weight factor, x' and y' represent the image brightness change rate of the facial image on the x-axis and y-axis respectively, and x" and y" represent the local curvature degree of the facial image on the x-axis and y-axis respectively; S202: Determine a Sobel template having a curvature weight factor using the curvature weight factor: S arc =S sobel ·(1+w) Among them, S arc represents the Sobel template with curvature weight factor, S sobel Indicates the horizontal template; S203: Perform image enhancement on the facial image using a Sobel template with a curvature weight factor to obtain the enhanced facial image: I enhanced =I+G G x =I·S arc ,G y =I·S arc Among them, I and I enhanced represent the facial image and enhanced facial image respectively, G represents the enhanced gradient amplitude, G x and G y They represent the horizontal enhancement gradient and the vertical enhancement gradient respectively.

4. The eye habit monitoring method according to claim 1, characterized in that: The real facial parameter is specifically the facial height, and the eye distance is calculated in the following manner: Where d is the eye distance, h is real Indicates the real face height, h img Represents the face height pixel value in the enhanced face image, H img represents the total height of the enhanced facial image, f represents the focal length of the camera, H sensor Indicates the camera sensor height.

5. The eye habit monitoring method according to claim 1, characterized in that: The S4 specifically includes: S401: Detecting the left eye position and the right eye position of the target user by using a key point detection model; S402: Under facial geometry constraints, determine the mouth position according to the left eye position and the right eye position: in, and They represent the left eye position coordinates, right eye position coordinates and mouth position coordinates in the enhanced facial image, respectively. y Indicates the vertical offset of the mouth position.

6. The eye habit monitoring method according to claim 1, characterized in that: The S5 specifically includes: S501: A conversion formula describing the conversion of facial features of a standard face model from a three-dimensional model coordinate system to a camera coordinate system is established through a PnP algorithm: p i =R·P i +T Among them, P i represents the 3D model point of the i-th facial feature point defined by the standard face model in the 3D model coordinate system, p i Indicates P i The three-dimensional position in the camera coordinate system, R represents the rotation matrix describing the head posture, and T represents the translation vector describing the position of the head in the camera coordinate system; S502: Establishing a projection formula for transforming facial features from the camera coordinate system to the enhanced facial image: Among them, (x i ,y i ,z i ) represents the 3D model point p in the camera coordinate system i The coordinates of x and f y They represent the horizontal focal length and vertical focal length of the camera, which describe the camera internal parameters. x and c y Respectively represent the horizontal pixel coordinates and vertical pixel coordinates of the camera optical center, u i and v i represents the 2D pixel coordinates of the i-th facial feature in the enhanced facial image; S503: Calculate the two-dimensional pixel coordinate error between the real-time facial feature and the facial feature in the enhanced facial image: Where E represents the two-dimensional pixel coordinate error value, represents the two-dimensional pixel coordinates of the i-th real-time facial feature in the enhanced facial image, n represents the number of facial features, and |||| represents the two-norm operator; S504: Using a nonlinear least squares optimization algorithm, with the goal of eliminating the two-dimensional pixel coordinate error value, updating the rotation matrix and translation vector describing the real-time head posture to obtain the real-time head posture: R new =R·exp(Δθ) T new =T+ΔT Among them, R new and T new They represent the updated rotation matrix and translation vector respectively, exp represents the natural exponential function, J represents the Jacobian matrix of E with respect to R and T, and the subscript T represents the transpose, (JTJ) -1 Indicates J T The inverse of J, Δθ and ΔT represent the rotation increment and translation increment respectively.

7. The eye habit monitoring method according to claim 6, characterized in that: The method for determining the gaze direction is specifically as follows: L(t)=P eye +t·d head ,t>0 Among them, d head represents the orientation vector of the head, and t represents the control annotation ray along d head Extension parameter of the extension scale, P eye represents the position of the eye center, and L(t) represents the gaze ray associated with t representing the annotation direction.

8. The eye habit monitoring method according to claim 7, characterized in that: The focus points are specifically: Among them, P gaze represents the gaze point coordinates, d head,z Indicates d head The z component of , d represents the eye distance.

9. An eye habit monitoring system, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the eye habit monitoring method as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the eye habit monitoring method as described in any one of claims 1 to 8 is implemented.