Intelligent glasses zooming method for preventing and treating eye fatigue and related device

Through multimodal data processing and deep learning technology, smart glasses can accurately identify user visual fatigue status and adjust lens focal length in real time, solving the problem of inaccurate visual fatigue recognition in the prior art, and improving myopia prevention and control effect and user visual experience.

CN120375459APending Publication Date: 2025-07-25SHANGHAI WEICON OPTICAL CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510477482.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing smart zoom glasses are not accurate enough when identifying the user's visual fatigue level, resulting in untimely or inaccurate focal length adjustments, affecting the effect of myopia prevention and control, especially in scenarios where indoor office and outdoor activities are frequent for a long time.

Method used

By obtaining the user's multimodal data in real time, including eye images, eye movement data and ambient lighting data, visual and time series features are extracted using convolutional neural networks and long and short-term memory networks, dimensionality reduction and clustering are combined with autoencoders, user behavior patterns are identified using deep embedded clustering, and visual fatigue classification is performed through pre-trained classification models, and finally the curvature of the lens liquid crystal zoom layer is dynamically adjusted.

Benefits of technology

It realizes accurate identification and real-time response to the user's visual fatigue state, improves the recognition accuracy and classification efficiency of the visual fatigue state, and significantly improves the user's visual comfort and eye health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375459A_ABST
    Figure CN120375459A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent glasses, and discloses an intelligent glasses zooming method for preventing and treating eye fatigue and a related device. The method comprises the following steps: acquiring multi-modal data of a user in real time; extracting visual features of the eye image, extracting time sequence features of the eye movement data and the environment data, and fusing the visual features and the time sequence features to form a user behavior feature vector; performing automatic clustering on the user behavior feature vectors, and identifying a user behavior mode; on the basis of the current user mode obtained through automatic clustering, the visual features and the time sequence features are combined and input to a pre-trained classification model, and accurate visual fatigue classification of the current user state is carried out; and determining and dynamically adjusting the curvature of the liquid crystal zoom layer of the lens in real time according to the asthenopia classification result and the current user use mode. The visual fatigue state of the user is accurately monitored in real time, and the focal length is accurately regulated and controlled according to the visual fatigue state, so that active adaptation and accurate intervention of myopia prevention and control are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of smart glasses, and in particular, to a smart zoom method and related device for glasses to prevent and treat eye fatigue. Background Art

[0002] Although existing smart zoom glasses have the function of dynamically adjusting the lens focal length according to the eye use distance and ambient light, in actual applications, the user's visual fatigue level is often not accurately recognized and responded to, resulting in untimely or inaccurate adjustment of the focal length, affecting the effect of preventing and controlling myopia. Especially in two typical scenarios: users who work indoors for a long time, have a stable distance but use their eyes frequently and briefly, and users who have frequent outdoor activities and rapid changes in ambient light, due to the lack of accurate perception of the user's visual fatigue degree by traditional algorithms, it is difficult to achieve highly accurate and adaptive adjustment, affecting the effect of smart glasses actively intervening in myopia. Summary of the Invention

[0003] In order to accurately and real-time monitor the user's visual fatigue state and accordingly perform precise focal length control to achieve active adaptation and precise intervention in myopia prevention and control, this application provides a smart zoom method and related device for glasses to prevent and treat eye fatigue.

[0004] In a first aspect, this application provides a smart zoom method for glasses to prevent and treat eye fatigue, adopting the following technical solutions:

[0005] A smart zoom method for glasses to prevent and treat eye fatigue includes the following steps:

[0006] S1. Real-time obtain multi-modal data of the user, where the multi-modal data includes eye image data, eye movement data, ambient light data, and eye use distance data;

[0007] S2. Extract the visual features of the eye image, extract the time series features of the eye movement data and environmental data, and fuse the visual features and time series features to form a user behavior feature vector;

[0008] S3. Automatically cluster the user behavior feature vector to identify the user behavior pattern;

[0009] S4. Based on the user's current mode obtained from the automatic clustering, and combined with the visual features and time series features, input them into a pre-trained classification model for accurate visual fatigue classification of the user's current state;

[0010] S5. According to the visual fatigue classification result and the current user usage mode, determine and dynamically adjust the curvature of the liquid crystal zoom layer of the lens in real time.

[0011] Optionally, the S2 includes the following steps:

[0012] S21. Use a CNN network to extract image spatial features from the eye image data;

[0013] S22. Use an LSTM network to extract time series features from the eye movement data;

[0014] S23. Extract time series features from the environmental light data and the eye use distance data separately or in combination;

[0015] S24. Fuse the image spatial features and the time series features to form a user behavior feature vector, and use a pre-trained autoencoder for dimensionality reduction to obtain a low-dimensional feature representation of the user behavior.

[0016] Optionally, the pre-trained autoencoder is a deep stacked autoencoder, including:

[0017] An encoder, including a three-layer fully connected network. The first layer reduces the input features to 64 dimensions and uses a ReLU activation function. The second layer reduces from 64 dimensions to 32 dimensions and uses a ReLU activation function. The third layer further reduces the dimensions to 16 dimensions and uses a linear activation function, thereby obtaining a 16-dimensional low-dimensional feature representation z;

[0018] A decoder, including a three-layer structure symmetric to the encoder, for reconstructing the 16-dimensional low-dimensional feature representation z back to the original input feature space, thereby realizing the reconstruction of the original input features;

[0019] A loss function, using the mean squared error loss function;

[0020] After the pre-trained autoencoder is trained, the output z of the encoder is extracted as the output.

[0021] Optionally, the S3 includes the following steps:

[0022] S31. Perform deep embedding clustering on the dimensionality-reduced features to automatically discover the user's behavior patterns;

[0023] S32. Based on the clustering results, adopt a dynamic threshold mechanism to automatically evaluate and adjust the cluster centers in real time, merge similar behavior patterns, automatically identify new patterns, and update the user pattern set.

[0024] Optionally, the S31 includes the following steps:

[0025] S311. Initialize the cluster centers using the output z of the pre-trained autoencoder, where the cluster center vector is defined as: μ j , j = 1, 2,, K;

[0026] S312. Based on each user behavior feature z i For the cluster center μ j Soft assignment probability:

[0027]

[0028] S313. Set the target probability

[0029] S314. Set the clustering loss function

[0030] S315. Minimize the clustering loss function to optimize the feature representation and the cluster centers, and automatically discover user behavior clusters.

[0031] Optionally, the S32 includes:[[]]

[0032] S321. Regularly calculate the distance matrix D of all cluster centers cluster , where the current cluster set is represented as: C = {μ1, μ2,, μ K} The distance is the Euclidean distance between cluster centers, D cluster (i,j) = ||μ i - μ j ||2, i,j = 1,2,, K, i≠j; the distance matrix D cluster is used to characterize the similarity between clusters;

[0033] S322. Calculate the statistical indicators of the distances between all pairs of cluster centers and the standard deviation of the distances to determine the merging threshold T merge = μ Cdist - γσ Cdist ; where γ is a hyperparameter;

[0034] S323. Compare the relationship between the cluster center D cluster and the merging threshold T merge to merge the cluster pairs and update the cluster pair centers; if there are multiple groups of cluster pairs to be merged, preferentially merge the nearest cluster pairs and update the cluster pair centers.

[0035] Optionally, the S323 includes the following sub-steps:[[]]

[0036] S3231. Calculate the new cluster center after merging where n i and n j respectively represent the number of data points that have been matched by the two original clusters in history;

[0037] S3232. Remove the original two clusters n i and n j from the cluster set, and add the new cluster μ new to the cluster set;

[0038] S3233. Update the data volume of the new clustering cluster to n new = n i + n j 。

[0039] Optionally, the S4 includes the following steps:

[0040] S41. Concatenate and fuse the CNN features, LSTM features, and the clustering result of the user's current behavior pattern to form a complete user state feature vector F user ;

[0041] S42. Automatically select a pre-trained classification model suitable for the current pattern according to the user clustering pattern;

[0042] S43. Input the user state feature F user into the selected pre-trained classification model, and the model outputs the classification probability or label of the user's current visual fatigue state.

[0043] In a second aspect, the present application provides an intelligent zoom system for glasses for preventing and treating eye fatigue, adopting the following technical solution:

[0044] An intelligent zoom system for glasses for preventing and treating eye fatigue, including a processor, and a program of the method for preventing and treating eye fatigue of glasses as described in any one of the above is run in the processor.

[0045] In a third aspect, the present application provides a storage medium, adopting the following technical solution:

[0046] A storage medium stores a program of the method for preventing and treating eye fatigue of glasses as described in any one of the above.

[0047] In summary, the present application includes at least one of the following beneficial technical effects:

[0048] 1. The present application proposes an intelligent glasses visual fatigue dynamic monitoring and prevention and control method based on user multi-modal data. By real-time acquiring multi-dimensional data such as user eye images, eye movement data, environmental light, and eye use distance, using a convolutional neural network and a long short-term memory network to accurately extract visual features and time series features respectively, and using an autoencoder to effectively fuse and reduce dimensions, the accuracy and robustness of user behavior features are significantly improved.

[0049] 2. The present application uses deep embedded clustering to automatically discover and continuously optimize user behavior patterns, effectively avoiding the subjective errors and inaccurate pattern recognition brought by manual classification, and realizing the fine capture and accurate representation of users' personalized usage habits. Combining with a classification model to accurately identify the user's current visual fatigue state further improves the recognition accuracy and classification efficiency of the visual fatigue state.

[0050] 3. In addition, the present application uses reinforcement learning to control the curvature of the liquid crystal zoom lens in real time and accurately, enabling the smart glasses to quickly respond to the visual needs of users in different scenarios, effectively preventing or alleviating visual fatigue, and enhancing visual comfort. At the same time, the liquid crystal zoom lens has the advantages of rapid response and no mechanical movement, providing a long-term stable, accurate and efficient dynamic visual adaptation effect, and significantly improving the user's eye health and visual experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 It is a flowchart of a method for intelligent zooming of glasses for preventing and treating eye fatigue in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The following details the embodiments of the present application, and the examples of the embodiments are shown in the drawings.

[0053] In the description of this specification, the description with reference to the terms "certain embodiments", "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0054] The embodiment of the present application discloses a method for intelligent zooming of glasses for preventing and treating eye fatigue, with reference to Figure 1 , including the following steps S1 - S5.

[0055] S1. Real - time acquisition of multi - modal data of the user, where the multi - modal data includes eye image data, eye movement data, ambient light data, and eye - using distance data.

[0056] The real - time collection and acquisition of the user's multi - modal data are realized through multiple micro - sensors integrated in the frame of the smart glasses, specifically including a micro - camera, an infrared eye - movement tracking device, an ambient light sensor, and a distance sensor. These sensors work together to construct a multi - modal data stream containing eye image data, eye movement data, ambient light data, and eye - using distance data in real time. Multi - modal data refers to a data set obtained by multiple different types of sensors and capable of describing the user's current eye - using state from multiple dimensions.

[0057] Eye image data is usually captured by a micro camera, clearly recording visual information such as the pupil diameter of the user's eyes, the eyelid closure state, and the eye posture. These data can directly reflect the fatigue state of the user's current visual system. For example, when working at a close distance for a long time, the pupil diameter of the user may gradually contract or dilate abnormally, and the frequency and duration of eyelid closure increase significantly. By capturing and analyzing these visual features, early signs of eye fatigue can be accurately identified.

[0058] Eye movement data is obtained through an infrared eye movement tracking device, specifically recording the user's blink frequency, gaze duration, and the stability of the fixation point. A decrease in blink frequency is often associated with long-term gazing and eye muscle fatigue, while an extended gaze duration and a decrease in the stability of the fixation point usually indicate a high level of concentration, which is likely to induce visual fatigue.

[0059] At the same time, the light intensity of the environment where the user is located is detected in real time through an environmental light sensor, and the precise distance between the user's line of sight and the viewing target is measured in real time through a distance sensor. These two environmental data are crucial for evaluating the quality of the visual environment and eye use habits. For example, when the user is in a weak light environment indoors for a long time, it is easy to cause pupil dilation and exacerbate the symptoms of visual fatigue; under strong outdoor light conditions, the eyes need to adapt to light changes frequently, which is also likely to cause eye fatigue. In addition, the precise measurement of the user's eye use distance directly reflects whether the user works at a close distance for a long time or frequently changes the eye use distance.

[0060] S2. Extract the visual features of the eye image, extract the time series features of the eye movement data and environmental data, and fuse the visual features and time series features to form a user behavior feature vector; among them, the environmental data includes environmental light data and eye use distance data.

[0061] First, use a convolutional neural network (CNN) to extract visual features from the eye image data. CNN can effectively capture the spatial local features in the image, such as the diameter change of the user's pupil, the degree of eyelid closure, and the eye posture. These visual features are crucial for identifying visual fatigue. For example, when the user continuously works at a close distance for a long time, the degree of eyelid closure will gradually increase, and the change in pupil diameter becomes unstable. The spatial visual features extracted by CNN can accurately reflect these subtle changes and help the system identify the early fatigue state.

[0062] Meanwhile, in this step, long short-term memory network (LSTM) is used to extract time series features from eye movement data, environmental light data, and eye use distance data. The LSTM network is particularly good at capturing the internal laws and long-term dependencies of data changes over time, and can deeply explore the trends of the user's blink frequency, gaze duration, and fixation point stability evolving over time. For example, when the user is in a fatigued state, the blink frequency will gradually decrease, the gaze time will lengthen, and the stability of the fixation point will also decrease. These time series changes in eye movement features can sensitively and directly reflect the process of the user's visual fatigue. At the same time, LSTM is also applied to environmental light data and eye use distance data to capture the time characteristics of the user's eye use behavior changing with environmental conditions. Taking outdoor activities as an example, the outdoor light intensity changes drastically, and the eye use distance also changes frequently. The environmental time series features extracted by LSTM can timely reflect the dynamic change trends of the environment and eye use habits, helping the system to adapt to the user's current scenario in real time.

[0063] After separately extracting visual features and time series features, this step effectively fuses these two types of features to form a unified user behavior feature vector. The significance of fusing visual features and time series features lies in integrating information from different dimensions to more comprehensively and accurately describe the user's real-time eye use state. Visual features represent the user's current eye space state, while time series features represent the dynamic trend of the user's eye use state changing over time. The combination of these two can form a more accurate user behavior representation. For example, when the visual features show that the user's current pupil is dilated, and the time series features of LSTM also show that the eye use distance has been continuously approaching recently and the blink frequency has decreased, the fused user behavior features can clearly reflect that the user is currently in a relatively severe visual fatigue state.

[0064] Specifically, in one embodiment, S2 includes the following steps S21 - S24.

[0065] S21. Use the CNN network to extract image space features from the eye image data.

[0066] CNN first preliminarily processes the real-time captured eye image, including image denoising, size normalization, and enhancement processing, to improve the accuracy and stability of subsequent feature extraction. Then, CNN gradually extracts the feature information in the image through the convolutional layer. For example, the degree of shrinkage or expansion of the pupil diameter and the subtle changes in the eyelid closure state. These changes can accurately represent the user's visual fatigue degree. For example, when the user continuously reads at a close distance for more than a certain period of time, they usually show a gradually abnormal expansion or contraction of the pupil diameter and an increased frequency and amplitude of eyelid closure.

[0067] S22. Use the LSTM network to extract time series features from the eye movement data.

[0068] In this step, the eye movement data captured by the smart glasses in real time, such as blinking frequency, gaze duration, and stability of the gaze point, are first constructed into a continuous time series form, and the data is preprocessed, such as outlier filtering and data normalization, to ensure data quality and model stability. Subsequently, these eye movement data sequences are fed into a pre-trained LSTM network, which automatically captures and learns important time patterns and change trends in the data sequence through its unique gating mechanism. For example, when a user works at a close distance for a long time, his blinking frequency gradually decreases, the gaze duration is significantly prolonged, and the stability of the gaze point gradually decreases. This time series feature is captured and encoded by LSTM, accurately reflecting the process of the user's transition from a normal state to a fatigue state.

[0069] S23. Extract time series features from the ambient light data and the eye distance data separately or in combination.

[0070] This step first preprocesses the ambient light intensity data (in Lux) collected in real time by the light sensor and the eye distance data measured by the distance sensor, including data smoothing, outlier removal, and data normalization, to improve data quality. Subsequently, these processed environmental and distance data are input into the pre-trained LSTM network in the form of time series. Through the unique gating structure of LSTM, the network can effectively learn and extract the inherent laws and long-term trends of these data over time. For example, when working indoors for a long time, the ambient light data usually shows stable intensity and slow changes, while the eye distance is mostly maintained in a relatively fixed close range; when the user is in an outdoor activity state, the light intensity may change dramatically in a short period of time, and the eye distance also changes frequently and significantly. The LSTM network captures these different patterns and trends over time and encodes them into high-quality time series features that can be used for subsequent analysis.

[0071] The importance of extracting the time series features of ambient light and eye distance data lies in that it effectively makes up for the limitations of simple eye data and more comprehensively describes the user's real eye environment and eye habits. For example, when the ambient light gradually dims, the user's pupil diameter will naturally expand, and the risk of visual fatigue will increase; similarly, when the eye distance gradually shortens and is maintained at a close distance for a long time, the user is more likely to experience symptoms of visual fatigue.

[0072] S24. The image spatial features and time series features are integrated to form a user behavior feature vector, and the pre-trained autoencoder is used for dimensionality reduction to obtain a low-dimensional feature representation of user behavior.

[0073] In this step, first, the visual features extracted in step S21 (such as spatial features like pupil diameter change and eyelid closure degree) are concatenated and fused with the time series features of the eye movement data, environmental light data, and eye use distance data extracted in steps S22 and S23 to form a complete, rich, and high-dimensional user behavior feature vector. Such a fused feature vector can comprehensively depict the spatial features of the user's eye state and the dynamic features of the user's eye use behavior and environmental state evolving over time. For example, when the user is in a stable indoor working state, the visual features show that the pupil is relatively stable, but the time series features show that the blink frequency gradually decreases and the eye use distance gradually approaches. These pieces of information combined can accurately describe that the user is about to enter a state of visual fatigue.

[0074] To further improve the efficiency and accuracy of subsequent data analysis, this step also uses a pre-trained autoencoder to reduce the dimension of the above-mentioned fused high-dimensional feature vector. As a special neural network structure, the autoencoder realizes the dimension reduction and reconstruction of the input data through the encoder and decoder structures. In the present invention, the encoder of the autoencoder consists of three fully connected networks, which gradually compress the fused features from a high-dimensional space to a low-dimensional space (eventually to a 16-dimensional low-dimensional feature representation), while the decoder reconstructs the low-dimensional features to ensure that important information is not lost during the dimension reduction process. For example, a fused feature originally with a dimension as high as hundreds of dimensions can be reduced to a 16-dimensional low-dimensional feature representation after being processed by the autoencoder, significantly reducing the computational complexity of subsequent clustering and classification models while retaining the key feature information.

[0075] Specifically, in one embodiment, the pre-trained autoencoder is a deep stacked autoencoder, including:

[0076] An encoder, including three fully connected networks. The first layer reduces the input features to 64 dimensions and uses the ReLU activation function. The second layer reduces from 64 dimensions to 32 dimensions and uses the ReLU activation function. The third layer further reduces the dimension to 16 dimensions and uses the linear activation function, thereby obtaining a 16-dimensional low-dimensional feature representation z;

[0077] A decoder, including a three-layer structure symmetric to the encoder, for reconstructing the 16-dimensional low-dimensional feature representation z back to the original input feature space, thereby realizing the reconstruction of the original input features;

[0078] A loss function, using the mean squared error loss function;

[0079] After the pre-trained autoencoder is trained, the output z of the encoder is extracted as the output.

[0080] Specifically, the encoder gradually maps the input high-dimensional data into a latent space feature representation with a lower dimension, and this process is called feature compression or dimensionality reduction. The low-dimensional feature representation after dimensionality reduction retains the most important and core parts of the original data while removing irrelevant or redundant information. The decoder then attempts to reconstruct the original high-dimensional data as accurately as possible from this low-dimensional feature representation, and this reconstruction process actually serves to supervise the effect of dimensionality reduction. If too much important information is lost during the dimensionality reduction process, the decoder cannot accurately reconstruct the input data, and thus a large reconstruction error will occur during training. Therefore, by continuously training and optimizing the structures of the encoder and decoder to minimize the reconstruction error, it can be ensured that the finally obtained low-dimensional feature representation has the characteristics of high efficiency, compactness, and minimal information loss.

[0081] Taking the present invention as an example, the initial dimension of the fused user behavior feature vectors may be relatively high. Directly using these high-dimensional features for clustering or classification will cause a heavy computational burden, and some of these features may be redundant or noisy data, which is not conducive to subsequent analysis. Dimensionality reduction is performed through the encoder part of the autoencoder, which can compress features from a high-dimensional space to a more representative low-dimensional space, making subsequent clustering and classification tasks easier to execute, faster, and more accurate. The reconstruction process of the decoder ensures the effectiveness and quality of the encoder compression process, avoiding the problem of information loss caused by over-compression.

[0082] S3. Automatically cluster the user behavior feature vectors to identify user behavior patterns.

[0083] Clustering is an unsupervised machine learning method. Its main idea is to automatically divide data into multiple intrinsically related clusters or groups according to the similarity between data, so that the data within the same cluster has as high a similarity as possible, while the data between different clusters has as large a difference as possible. In the present invention, the role of clustering is to automatically identify and divide the behavior patterns of users without relying on pre-defined or marked specific patterns to which users belong, thereby improving the self-adaptability and personalization of the system.

[0084] In the specific implementation process, this step utilizes the low-dimensional user behavior feature representation obtained in the aforementioned step S24 to automatically cluster and identify user behavior patterns through the Deep Embedded Clustering (DEC) algorithm. By combining the advantages of deep neural networks and clustering algorithms, DEC automatically discovers the potential relationships and pattern differences among user behavior features, and continuously optimizes the feature representation and cluster centers during the training process. For example, when a user works indoors at a short distance for a long time, the DEC algorithm can automatically identify the stable features of the user in terms of pupil diameter change, blink frequency, gaze time, and eye use distance, forming a clear indoor office mode cluster; while when the user is frequently active outdoors, the algorithm can also automatically identify the dynamic features of rapid changes in light intensity and continuous adjustment of eye use distance, forming an outdoor activity mode cluster that is significantly different from the indoor mode.

[0085] The significant advantage of automatically clustering user behavior patterns is that the system does not need to rely on artificially preset mode categories or strict classification criteria, but automatically forms mode classifications based on the characteristics of the data itself. The benefit of this approach is that it can reflect the personalized usage habits and behavior differences of different users in real scenarios, and as the data accumulates, gradually refine and update the user mode classifications. For example, when a user is in a mixed eye use state for a long time (both indoor office and outdoor activities), the DEC algorithm can dynamically identify and establish the corresponding mode clusters to ensure that the subsequent steps can adjust the visual fatigue classification strategy and lens zoom scheme according to the precise mode the user is in.

[0086] Specifically, in one embodiment, the S3 includes the following steps S31 - S32.

[0087] S31. Perform deep embedded clustering on the dimensionality-reduced features to automatically discover the user's behavior patterns.

[0088] The DEC first initializes the cluster centers using the output features of the pre-trained autoencoder. The features output by the autoencoder contain the most representative low-dimensional information of user behavior. Based on these features, the DEC initializes the cluster centers, thus improving the stability and accuracy of clustering. Subsequently, by calculating the soft assignment probabilities of user behavior features to each cluster center, the DEC algorithm automatically evaluates the possibility of each user behavior belonging to each cluster, and further calculates the target probability and the clustering loss function based on this. By optimizing this clustering loss function, the feature representation and the cluster centers gradually tend to be optimal. For example, when a user is engaged in a reading activity for a long time in an indoor environment, the DEC can quickly recognize that this behavior feature highly matches the existing indoor office mode cluster center, and thus classify this user behavior feature into the corresponding mode; while when the user switches to the outdoor fast movement state, the DEC can also keenly recognize the difference between this behavior feature and the existing modes and quickly establish a corresponding new mode cluster.

[0089] Specifically, in one embodiment, S31 includes the following steps S311-S315.

[0090] S311. Initialize the cluster centers using the output z of the pre-trained autoencoder, where the cluster center vector is defined as: μ j , j = 1, 2,, K.

[0091] S312. Based on each user behavior feature z i calculate the soft assignment probability to the cluster center μ j :

[0092]

[0093] After the initialization of the cluster centers is completed, the algorithm calculates the soft assignment probability based on the distance between each user behavior feature and each cluster center, that is, measures the degree to which each user behavior feature belongs to each cluster. For example, a user who has been working stably indoors for a long time has the smallest distance between his behavior feature vector and the cluster center of the "indoor close-range working mode", so he obtains a higher soft assignment probability, while the distance from the cluster center of the "outdoor sports mode" is larger, so the soft assignment probability is lower.

[0094] S313. Set the target probability

[0095] Subsequently, through these soft assignment probabilities, the algorithm sets the target probability to guide the optimization of the clustering model.

[0096] S314. Set the clustering loss function

[0097] S315. Minimize the clustering loss function to optimize the feature representation and the cluster centers and automatically discover the user behavior clusters.

[0098] To optimize the cluster centers and feature representations, the algorithm defines a corresponding clustering loss function, which is calculated based on the difference between the soft assignment probability and the target probability. By continuously iterating to optimize the model parameters and gradually minimizing this loss function, the algorithm can continuously adjust the correspondence between the cluster center positions and the feature representations, making the user behavior features in the same cluster more concentrated and the differences between different clusters more significant.

[0099] For example, when a user changes from a long-term indoor stable working state to a state of frequent outdoor activities, their behavior features will gradually move away from the original cluster center and approach the outdoor mode cluster. Through the above process, the algorithm automatically updates the cluster center positions and the attribution of user behavior features, achieving the adaptive adjustment of the clusters.

[0100] S32. Based on the clustering results, adopt a dynamic threshold mechanism to automatically evaluate and adjust the cluster centers in real time, merge similar behavior patterns, automatically identify new patterns, and update the user pattern set.

[0101] Specifically, in one embodiment, S32 includes the following steps S321 - S323.

[0102] S321. Regularly calculate the distance matrix D of all cluster centers cluster , where the current cluster set is represented as: C = {μ1, μ2,, μ K}}, The distance is the Euclidean distance between the cluster centers, D cluster (i,j) = ||μ i - μ j ||2, i,j = 1,2,, K, i ≠ j; The distance matrix D cluster is used to characterize the similarity between clusters.

[0103] S322. Calculate the statistical index of the distances between all pairs of cluster centers and the standard deviation of the distances to determine the merging threshold T merge = μ Cdist - γσ Cdist ; where γ is a hyperparameter.

[0104] When the distances between the cluster centers are generally far (the pattern differences are obvious), σ Cdist is large, and the threshold should become more strict to avoid mismerging patterns with obvious differences; when the distances between the cluster centers are generally close (the clusters are more similar), σ Cdist is small, and the threshold is appropriately relaxed to more sensitively identify and merge highly similar patterns.

[0105] S323. Compare the cluster centers D cluster with the merging threshold T merge to merge the cluster pairs and update the cluster pair centers; if there are multiple groups of cluster pairs to be merged, the nearest cluster pairs are preferentially merged and the cluster pair centers are updated.

[0106] After obtaining the distance matrix of the cluster centers, the algorithm further calculates statistical metrics of the distances between all clusters, such as the mean distance and standard deviation, and dynamically determines the threshold for pattern merging based on these statistical metrics. When the distance between the centers of two user behavior pattern clusters is less than this dynamic threshold, the system determines that these two patterns are highly similar and automatically triggers the merging mechanism. For example, the long-term close-reading pattern and the long-term computer-office pattern of a user in an indoor environment may gradually narrow the gap between their cluster centers as the user's usage habits change. When the merging threshold is reached, the system automatically merges them into a single indoor long-term eye-using pattern. Conversely, when the distance between the new user behavior characteristics and the existing pattern clusters exceeds the dynamic threshold, the system automatically identifies it as a new user pattern and adds a corresponding cluster.

[0107] The cluster merging and new pattern recognition process is achieved by preferentially merging the closest cluster pairs. During the merging process, the algorithm updates the new cluster center after merging by weighting it according to the historical data volume to more accurately represent the true behavior trend of the user. For example, assuming that the indoor office pattern and the reading pattern are merged into one pattern, the position of the new cluster center will consider the historical data volumes of the previous two clusters simultaneously to ensure that the new cluster center can truly and accurately represent the user's long-term usage habits.

[0108] Specifically, in one embodiment, S323 includes the following sub-steps S3231 - S3233.

[0109] S3231. Calculate the new cluster center after merging where n i and n j respectively represent the number of data points matched historically by the two original clusters.

[0110] S3232. Remove the original two clusters n i and n j from the cluster set, and add the new cluster μ new to the cluster set.

[0111] S3233. Update the data volume of the new cluster to n new = n i + n j .

[0112] In the specific implementation process, the system first automatically determines the cluster pairs that need to be merged based on the distance matrix obtained from dynamic evaluation and the pattern merging threshold, that is, the cluster pairs where the distance between the cluster centers is less than the merging threshold. These cluster pairs represent that the user behavior patterns have high similarity, and the system will automatically merge these patterns to reduce pattern redundancy.

[0113] After determining the merged cluster pairs, the system updates the center of the newly merged cluster by weighting historical data. Specifically, the position of the newly merged cluster center is the weighted average of the positions of the original two cluster centers, and the weights are the number of data points accumulated in the respective histories of the two original clusters. For example, when a user has two behavior patterns in an indoor environment: one is a long - time computer office mode, and the other is a long - time paper document reading mode. When the two patterns gradually become very close with the evolution of the user's eye - using habits, the system automatically merges these two patterns into a unified "long - term indoor near - distance eye - using mode". At this time, the new cluster center will comprehensively consider the number of historical data points of the original two clusters to ensure that the newly formed pattern better fits the user's actual long - term behavior characteristics.

[0114] Finally, after the merging process is completed, the system updates the cluster set, deletes the original two cluster centers, adds the newly formed cluster center to the cluster set, and simultaneously updates the corresponding data statistics information.

[0115] S4. Based on the user's current pattern obtained from automatic clustering, and combined with visual features and time - series features, input it into a pre - trained classification model for accurate visual fatigue classification of the user's current state;

[0116] Specifically, in one embodiment, S4 includes the following steps S41 - S43.

[0117] S41. Concatenate and fuse the CNN features, LSTM features, and the clustering result of the user's current behavior pattern to form a complete user state feature vector F user 。

[0118] S42. Automatically select a pre - trained classification model suitable for the current pattern according to the user's clustering pattern.

[0119] S43. Input the user state feature F user into the selected pre - trained classification model, and the model outputs the classification probability or label of the user's current visual fatigue state.

[0120] The system first fuses the visual features, time series features, and the encoding of the user behavior pattern (such as one-hot encoding or embedding vector) obtained in the previous steps to form a unified user state feature vector. This fusion method can comprehensively consider the user's current eye space state, dynamic eye movement and environmental change features, as well as the overall user behavior pattern, making the description of the user state more accurate and comprehensive.

[0121] Subsequently, according to the user's current behavior pattern, the system automatically selects the corresponding pre-trained Transformer-Capsule classification model for visual fatigue classification. The Transformer-Capsule classification model is a deep classification model that combines the advantages of the Transformer network and the Capsule Network. The Transformer is good at capturing the long-range dependencies and global features of the data, while the Capsule Network can capture the tiny spatial pose changes in the visual data.

[0122] Specifically, when the user's current mode is recognized by the clustering model as the indoor long-term close-range mode, the system will automatically call the high-sensitivity fatigue classification model designed for the indoor long-term close-range environment to timely capture the subtle changes in the user's visual fatigue state. For example, when the user is working indoors for a long time, the user state feature vector shows that the user's blink frequency decreases and the abnormal change in pupil diameter gradually increases, and the system classifies the user as mildly fatigued. When the user's mode is recognized as the outdoor frequent light change mode, the system automatically calls the fast-response fatigue classification strategy optimized for the outdoor scene to quickly respond to the visual discomfort caused by the frequent light changes for the user. For example, when the user enters a bright environment from a dark place, the system immediately judges the change in the user's visual fatigue state to ensure that the subsequent focal length and light transmittance adjustments are timely and accurate. In addition, for other newly automatically recognized user modes, the system can also automatically match and use the corresponding optimized fatigue classification strategies to ensure accurate classification effects for the visual fatigue classification of each user mode.

[0123] S5. According to the visual fatigue classification result and the current user usage mode, determine and dynamically adjust the curvature of the liquid crystal zoom layer of the lens in real time.

[0124] The liquid crystal zoom layer of the lens is an optical element composed of liquid crystal materials. By controlling the arrangement state of liquid crystal molecules through an electric field, the dynamic continuous change of the optical focal length is realized. This technology enables the lens to quickly and accurately adjust the optical power of the lens only by relying on an electric signal without mechanical movement, so as to achieve the purpose of optimizing the user's visual comfort in real time.

[0125] In the specific implementation process, the lens zoom parameters output by the reinforcement learning algorithm (such as DDPG) drive the rearrangement of liquid crystal molecules in the lens in the form of electrical signals, and the curvature of the lens is adjusted precisely in real time, so as to achieve continuous and fine focal length changes. For example, when the user is reading or working indoors at a short distance for a long time, when the system detects that the user enters a mild or moderate fatigue state, the curvature of the lens zoom layer will automatically increase or decrease by a certain amount, appropriately adjust the visual focus position, and relieve the visual burden of the user's continuous short-distance focusing. When the user goes outdoors and the environmental light changes frequently and violently, the curvature of the lens will quickly respond to the environmental changes for corresponding adjustments. For example, when the user enters a shaded area from a bright environment, the focal length and light transmittance of the lens will be automatically adjusted immediately.

[0126] Based on the aforementioned accurate visual fatigue classification results and the user's current behavior pattern, the curvature of the liquid crystal zoom layer of the smart glasses lens is dynamically adjusted through the reinforcement learning algorithm to achieve precise intervention in the user's visual state, and significantly improve the visual comfort and vision protection effect. Reinforcement learning is a machine learning method that optimizes decision-making strategies by an agent learning in continuous interaction with the environment to maximize the long-term cumulative reward, and is suitable for real-time dynamic precise adjustment tasks.

[0127] In the specific implementation process, the system uses the visual fatigue state classification results obtained in S4 (such as normal, mild fatigue, moderate fatigue, severe fatigue) and the encoding of the current user mode as the input state of the reinforcement learning model. The reinforcement learning model adopts the Deep Deterministic Policy Gradient (DDPG) algorithm, which can effectively handle the characteristics of continuous change of the lens curvature and achieve high-precision dynamic adjustment of the lens curvature. DDPG calculates and outputs the optimal action, that is, the specific lens zoom parameter, through the continuous mapping relationship between the state space and the action space. For example, when the user is in a mild fatigue state in an indoor environment, the state input of the reinforcement learning model is the mild fatigue classification result and the indoor short-distance mode encoding, and DDPG outputs an accurate action value to finely adjust the curvature of the liquid crystal zoom layer of the lens to an appropriate degree to help the user relieve visual fatigue and reduce the risk of further deterioration.

[0128] The embodiment of the present application also discloses a smart zoom system for glasses for preventing and treating eye fatigue, including a processor, and a program of the smart zoom method for glasses for preventing and treating eye fatigue described in any one of the above is run in the processor.

[0129] The embodiment of the present application also discloses a storage medium storing a program of the smart zoom method for glasses for preventing and treating eye fatigue described in any one of the above.

[0130] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. An intelligent zoom method for glasses to prevent and treat eye fatigue, characterized in that, It includes the following steps: S1. Obtain the user's multi-modal data in real time, where the multi-modal data includes eye image data, eye movement data, ambient light data, and eye-using distance data; S2. Extract the visual features of the eye image, extract the time series features of the eye movement data and environmental data, and fuse the visual features and time series features to form a user behavior feature vector; among them, the environmental data includes ambient light data and eye-using distance data; S3. Automatically cluster the user behavior feature vector to identify the user behavior pattern; S4. Based on the user's current pattern obtained by automatic clustering, and combined with the visual features and time series features, input them into a pre-trained classification model for accurate visual fatigue classification of the user's current state; S5. According to the visual fatigue classification result and the current user usage pattern, determine and dynamically adjust the curvature of the lens liquid crystal zoom layer in real time.

2. The intelligent zoom method of glasses for preventing and treating eye fatigue according to claim 1, characterized in that, The S2 includes the following steps: S21. Use a CNN network to extract image spatial features from the eye image data; S22. Use an LSTM network to extract time series features from the eye movement data; S23. Extract time series features from the ambient light data and eye-using distance data separately or combined; S24. Fuse the image spatial features and time series features to form a user behavior feature vector, and use a pre-trained autoencoder for dimensionality reduction to obtain a low-dimensional feature representation of the user behavior.

3. The intelligent zoom method for glasses for preventing and treating eye fatigue according to claim 2, characterized in that, The pre-trained autoencoder is a deep stacked autoencoder, including: An encoder, including three layers of fully connected networks. The first layer reduces the input features to 64 dimensions and uses a ReLU activation function. The second layer reduces from 64 dimensions to 32 dimensions and uses a ReLU activation function. The third layer further reduces the dimensions to 16 dimensions and uses a linear activation function to obtain a 16-dimensional low-dimensional feature representation z; A decoder, including a three-layer structure symmetric to the encoder, used to reconstruct the 16-dimensional low-dimensional feature representation z back to the original input feature space, thereby realizing the reconstruction of the original input features; A loss function, using a mean square error loss function; After the pre-trained autoencoder is trained, extract the output z of the encoder as the output.

4. The intelligent zooming method for glasses for preventing and treating eye fatigue according to claim 3, characterized in that, The S3 includes the following steps: S31. Perform deep embedded clustering on the dimensionality-reduced features to automatically discover the user's behavior patterns; S32. Based on the clustering results, adopt a dynamic threshold mechanism to automatically evaluate and adjust the clustering cluster centers in real time, merge similar behavior patterns, automatically identify new patterns, and update the user pattern set.

5. The intelligent zoom method of glasses for preventing and treating eye fatigue according to claim 4, wherein, The S31 includes the following steps: S311. Initialize the cluster centers with the output z of the pre-trained autoencoder, where the cluster center vector is defined as: μ j , j = 1, 2,, K; S312. Based on each user behavior feature z i For the cluster center μ j Soft assignment probability: S313. Set the target probability S314. Set the clustering loss function S315. Minimize the clustering loss function to optimize the feature representation and clustering cluster centers, and automatically discover the user behavior clusters.

6. The intelligent zoom method for glasses for preventing and treating eye fatigue according to claim 5, characterized in that, The S32 includes the following steps: S321. Regularly perform the distance matrix D of all cluster centers cluster , where the current cluster set is represented as: The distance mentioned is the Euclidean distance of the cluster center, The distance matrix D cluster is used to characterize the similarity between clusters; S322. Calculate the statistical metrics of the distances between all pairs of cluster centers and the standard deviation of the distances to determine the merging threshold T merge = μ Cdist - γσ Cdist ; where γ is a hyperparameter; S323. Compare the cluster center D cluster with the merging threshold T merge to merge cluster pairs and update the cluster pair centers; if there are multiple groups of cluster pairs to be merged, give priority to merging the closest cluster pairs and update the cluster pair centers.

7. The intelligent zooming method for glasses for preventing and treating eye fatigue according to claim 6, wherein The S323 includes the following sub-steps: S3231. Calculate the new cluster center after merging where n i and n j respectively represent the number of data points that have been matched in the histories of the two original clusters; S3232. Remove the original two clustering clusters n i and n j from the cluster set, and add the new clustering cluster μ new to the cluster set; S3233. Update the data volume of the new clustering cluster to n new = n i + n j .

8. The intelligent zoom method for glasses for preventing and treating eye fatigue according to claim 7, wherein The S4 includes the following steps: S41. Concatenate and fuse the CNN features, LSTM features, and the clustering result of the user's current behavior pattern to form a complete user state feature vector F user ; S42. Automatically select a pre-trained classification model suitable for the current pattern according to the user clustering pattern; S43. Input the user status feature F user into the selected pre-trained classification model, and the model outputs the classification probability or label of the user's current visual fatigue status.

9. An intelligent zoom system for glasses used to prevent and treat eye fatigue, characterized in that, It includes a processor, and a program for the intelligent zoom method of glasses for preventing and treating eye fatigue as described in any one of claims 1-8 runs in the processor.

10. A storage medium, characterized in that, Store a program for the intelligent zoom method of glasses for preventing and treating eye fatigue as described in any one of claims 1-8.

Citation Information

Cited By

  • Eye fatigue detection method and system for brain wave glasses

    CN120753584A

  • Method for detecting and relieving asthenopia and electronic equipment

    CN122025113A