Line-of-sight detection device and setting processing method thereof
By combining a gaze detection device with an eye-tracking and target detection module, the problem of gaze detection glasses lacking direct assistance is solved, enabling real-time assistance for users and providing the function of quickly finding images of targets of interest.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JUJIA UNITED TECHNOLOGY CO LTD
- Filing Date
- 2026-03-12
- Publication Date
- 2026-06-05
AI Technical Summary
Existing gaze detection glasses lack direct auxiliary applications and are mainly used for subsequent statistical analysis, rather than providing direct real-time assistance to the user.
Design a gaze detection device that combines an eye-tracking module and a target detection module. It can determine the gaze state by sensing eye information and detect targets of interest, and store relevant image information for users to review and refer to.
It enables direct auxiliary applications of the line-of-sight detection device, helping users quickly find image information of targets of interest, reducing storage burden, and supporting practical applications in daily life.
Smart Images

Figure CN122157343A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of gaze detection, and in particular to a gaze detection device and its setting processing method that stores images for users to review and refer to. Background Technology
[0002] A gaze detection device is an intelligent device capable of instantly detecting and analyzing a user's gaze direction and duration, such as gaze detection glasses. Currently, Artificial Intelligence (AI) object detection, achieved through small module chips, can perform fast and highly accurate object recognition and capture with low power consumption, making it widely applicable; most imaging devices have integrated object detection functionality. However, the common application of gaze detection glasses currently involves long-term recording of the user's gaze points for subsequent medical physiological or psychological analysis, or for commercial purposes such as statistical analysis of areas of interest for consumers or visitors. Therefore, current applications are mostly for providing post-analysis records, lacking direct assistance for the user. In view of this, it is necessary to provide a gaze detection device that can provide direct assistance to solve the aforementioned technical problems. Summary of the Invention
[0003] To address the problems of the prior art, the present invention aims to provide a gaze detection device that can provide direct auxiliary applications. In a first aspect, the present invention provides a gaze detection device, comprising: an eye tracking module and a target detection module. The eye tracking module includes a sensing device that senses eye information to estimate and determine the gaze direction and its temporal changes, thereby distinguishing between a saccade state and a fixation state. The gaze direction is the direction of eye fixation, and the temporal change refers to the change in the gaze direction over time. The target detection module includes multiple target categories, allowing the setting of at least one target category of interest. In the fixation state, the target detection module can detect (or identify) a target. When the target detection module determines that image information used for comparison matches the target category of interest, it captures and stores the image information for user review and reference.
[0004] In some embodiments, the gaze detection device may include or may be an eyeglass device with two inward-facing detection lenses for facing the user's eyes respectively. The eye-tracking module senses the gaze through the two detection lenses. The eyeglass device has a scene camera lens facing outward, which provides the image information for comparison. The eye information includes eye position information and eye image information.
[0005] In some embodiments, the target detection module includes a first neural network processor, which is used to determine whether the user is in a saccade state or a gaze state. If it is determined to be in a gaze state and the target detection module performs target detection, and the image information used for comparison matches the target category of interest, then a partial image is extracted from the image information and the partial image is connected (or linked) and stored to a smart device.
[0006] In some embodiments, the gaze detection device includes the smart device and a first communication module. The smart device includes a second communication module, a detection management software program, and a storage unit. The target detection module includes a central processing unit (CPU), a graphics processing unit (GPU), and a second neural network processor (NNF). The first and second communication modules can communicate wirelessly or via wired data transmission. The detection management software program allows the user to pre-set at least one target category of interest. The CPU, GPU, and NNF are used to perform target detection calculations. The CPU controls the gaze detection device to display the previously categorized and stored image information of the target category via instructions. In some embodiments, the smart device is a mobile device, such as a smartphone, tablet, or laptop.
[0007] In some embodiments, the gaze detection device is combined with the smart device as a laptop or a tablet computer, or the gaze detection device and the smart device are integrated into a laptop computer. The laptop computer or tablet computer has at least one detection lens facing the user to correspond to the user's eyes. The eye-tracking module senses the gaze through the at least one detection lens. The user provides the image information for comparison by browsing the screen display of data on the laptop computer or tablet computer.
[0008] In addition, in a second aspect, the present invention provides a gaze detection setting processing method, the steps of which include, in sequence: Step 1: Define at least one target category that you are interested in; Step Two: Eye information is obtained through a gaze detection device, and the temporal changes in gaze patterns are estimated to distinguish between saccades and fixations; and Step 3: If it is determined to be gaze, target detection is performed. If the image information used for comparison matches at least one target category of interest, the image information is captured and stored in a smart device via a connection for the user to review and refer to.
[0009] In some embodiments, the step of setting at least one target category of interest further includes setting the retention time (or storage period) of the image information, wherein the capture is performed by cropping the image and storing it after capturing the target selection range (or target detection box) output by a neural network processor.
[0010] In some embodiments, the gaze detection setting process further includes a user review instruction control step, in which the user uses an instruction to cause the gaze detection device or the smart device to display the image information of the previously categorized and stored target category.
[0011] In some embodiments, the instructions used by the user control include voice control, text input, gestures, actions, touch or button press control, and the target category includes objects, events, people, behaviors, expressions or text. Attached Figure Description
[0012] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.
[0013] Figure 1 This diagram shows the architecture of the line-of-sight detection device according to the first embodiment of the present invention.
[0014] Figure 2 This diagram illustrates the gaze detection device of the first embodiment of the present invention used for visual review.
[0015] Figure 3 This diagram illustrates the gaze detection device of the second embodiment of the present invention used for visual review.
[0016] Figure 4 This diagram illustrates the setup and operation flow of a line-of-sight detection device according to a feasible embodiment of the present invention.
[0017] Figure 5 This diagram illustrates the setup and operation process of a line-of-sight detection device according to another embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0019] Please refer to Figure 1 , 2As shown in Figure 5, in a first aspect, the present invention provides a gaze detection device, comprising: an eye movement tracking module 10 and a target detection module 20. The eye movement tracking module 10 includes a sensing device 11, which estimates and judges gaze and its temporal changes by sensing information from the eyes 12, thereby distinguishing between a saccade state and a fixation state. The gaze is the direction of eye fixation, and the temporal change of gaze refers to the change in gaze direction over time. The target detection module 20 includes multiple target categories, allowing the user to set at least one target category of interest. For example, the user can pre-set a target category 13, such as a car. Furthermore, for example, during daily use, through the user's brief fixation behavior detection or fixation detection, in the fixation state, the target detection module 20 can detect a target, such as a car. When the target detection module 20 determines that image information used for comparison matches the target category of interest, it retrieves and stores the image information for the user's review and reference. This invention uses AI object detection to capture and store images and timestamps in the background (or in the background), and can remove them after a set period. If a user remembers seeing a certain car at a certain time, they can quickly find the stored image information by reviewing it. The user can review the image by keyword search or user control commands.
[0020] In some embodiments, such as Figure 1 , 2 As shown in Figure 5, the sensing device 11 may include or be an eyeglass device. The eyeglass device has two or more detection lenses 111 facing inward. The following description uses two detection lenses 111 as an example. The two detection lenses 111 are respectively facing the user's two eyes 12. The eye-tracking module 10 senses the line of sight 14 through the two detection lenses 111. The eyeglass device has a scene camera lens 112 facing outward. The scene camera lens 112 provides the image information for comparison. The information of the eyes 12 includes, for example, eye position information and eye image information. In some embodiments, the target detection module 20 includes a first neural network processor 21 (NPU). The first neural network processor 21 can perform AI gaze inference to determine or infer whether the user is in a saccade or gaze state. If it is determined to be a gaze state and the target detection module 20 performs target detection, and the image information used for comparison matches the target category of interest, then a partial image or all images are extracted from the image information and stored in a smart device 30. In feasible embodiments, if only a partial image is extracted, the storage burden can be effectively reduced and the focus can be more concentrated on the target category of interest, excluding other background information or irrelevant information. Furthermore, as... Figure 1 and Figure 2 As shown, in some embodiments, the smart device 30 is a mobile device, such as a smartphone, tablet, or laptop.
[0021] Furthermore, in some embodiments, such as Figure 1 , 2 As shown in Figure 5, the gaze detection device includes the smart device 30 and a first communication module 22. The smart device 30 includes a second communication module 31, a detection management software program 32, and a storage unit 33. The target detection module 20 includes a computing device 23, which includes a central processing unit (CPU), a graphics processing unit (GPU), and a second neural processing unit (NPU). The computing device 23 can be used to perform AI object inference and execute calculations. The first communication module 22 and the second communication module 31 can communicate via data transmission 41, including wireless or wired data transmission. Data transmission 41 may include, but is not limited to, data transmission required for updating or setting related software or firmware, and data transmission of images, pictures, or text information acquired by the scene camera 112 or the detection camera 111. The detection management software program 32 can be used by the user to pre-set at least one target category of interest. The central processing unit, the graphics processing unit, and the second neural network processor are used to perform target detection calculations. The central processing unit controls the gaze detection device to display the image information of the previously categorized and stored target category through instruction control.
[0022] Furthermore, the application process and related details of feasible embodiments of the present invention are specifically illustrated below. (Refer to...) Figure 1 , 2 As shown in Figure 5, the present invention may include a gaze-detecting glasses device, a smart device 30 capable of wired storage, and multi-type object (or event) detection artificial intelligence (AI) mounted on the glasses device or the smart device 30. The multi-type object (or event) detection AI includes common AI models such as the YOLO model, which have at least 80 preset target categories, including people, cats, dogs, horses, cows, sheep, cars, motorcycles, buses, etc. Suppliers can modify or add additional target categories, and can also add event type AI models. These AI models can run on the computing units of the smart device, such as NPUs and GPUs.
[0023] The multi-object detection AI can be integrated into eyeglasses or smart devices, and suppliers can configure it according to computing requirements. The smart device 30 can be a mobile phone, tablet, or laptop, and can be connected to the eyeglasses for gaze detection via wired or wireless means.
[0024] like Figure 1 , 2 As shown in Figure 5, this invention develops an application program (APP) for a smart device system, which is connected to the software development kit (SDK) of the gaze detection glasses device. The APP runs in the background on the smart device, unidirectionally receiving notifications or information from the gaze detection glasses device. The gaze detection glasses device is only used for gaze detection estimation. The supplier provides preset target category object (or event) detection AI models in the smart device APP. If users desire more diverse or advanced object (or event) detection target categories, the supplier can provide different versions of the service for updating and running on the smart device.
[0025] Upon first use, users preset the target categories they wish to aid in memorization and the data retention period. The provider gradually expands the target category service, increasing the diversity of target categories, and users can choose from preset categories of interest. After the user sets the target category, the computing device (such as a smart device or gaze-detecting glasses) receives images from the camera in front of the gaze-detecting glasses and identifies the user based on their gaze behavior. Users can set the length of time the application retains information to reduce the storage burden on smart devices.
[0026] Furthermore, human eye movements are divided into saccades and fixations. During use, the system continuously monitors the user's gaze. Even if the user is not paying particular attention, a brief fixation will trigger the AI to detect the area being gazed upon. The gaze detection glasses device obtains eye information through sensors to estimate real-time gaze, analyzes saccades and fixations based on temporal changes in gaze, and if it determines that the gaze is momentarily fixed, it will issue a fixation notification to subsequent processes.
[0027] In addition, if the AI detects an object (or event) that matches the user's preset target category, it will capture a partial or full image and store it on the user's smart device via a connection. The stored information includes the time and other recordable feature information. If it does not match, no further processing will be performed.
[0028] Taking partial image capture as an example, after receiving a gaze notification, this invention performs AI object (or event) detection on the image from the camera in front of the glasses device, obtains the position and range of the object (or event) in the image, and makes a judgment based on the image landing point of the gaze vector and the target category set by the user. If it is determined to be an object (or event) of interest and being gazed at, the position and range of the object (or event) in the image are cropped and captured, and saved in the background to the time path specified by the smart device. It can be stored in the lowest file size format (or file format), or the user can choose different resolutions or compression formats as needed.
[0029] Furthermore, if a user sends a review description, this invention can review the data stored on the smart device and present the most relevant image to the user. Since the auxiliary memory information is stored in the background, the user only makes a fleeting glance and may not have a clear memory of it, nor be aware of the application's actions. Later, if the user mentions related terms verbally or searches for them on the smart device, this invention can effectively present relevant video information that the user has previously viewed for reference and review. Additionally, the user can directly view all stored information for further confirmation or comparison.
[0030] In addition, such as Figure 3 As shown, in another embodiment of the present invention, the gaze detection device is integrated with the smart device 50 into a laptop, tablet, or desktop computer. The laptop, tablet, or desktop computer has at least one detection lens 51 facing the user, corresponding to the user's eyes 12. The eye-tracking module senses the gaze through the at least one detection lens 51, and the user provides the image information for comparison by browsing the screen display of the aforementioned computer data. This embodiment pertains to different application scenarios; for example, gaze detection on a screen can be performed using an external camera, not limited to eyeglasses, but since the gaze corresponds to the screen display image, similar applications can be achieved.
[0031] In addition, regarding the second aspect, please refer to... Figure 4 As shown, the present invention provides a gaze detection setting processing method, the steps of which include, in sequence: Step 1 S100: Define at least one target category of interest; Step 2 S200: Eye information is obtained by sensing through the gaze detection device and the temporal changes of gaze are estimated and judged in order to distinguish between saccades and fixations; Step 3 S300: If it is determined to be gaze, target detection is performed. If the image information used for comparison matches the set at least one target category of interest, the image information is captured and stored in a smart device via a connection for the user to review and refer to.
[0032] In some embodiments, the step of setting the target category further includes setting the retention time (or storage period) of the image information, wherein the image is cropped and stored after being captured by the target selection range output by the neural network processor.
[0033] In some embodiments, the method further includes a user review instruction control step, in which the user instructs the gaze detection device or smart device to display previously categorized and stored image information of the target category.
[0034] In some embodiments, the user control commands include voice control, text input, gestures, actions, touch or button presses, and the target categories include objects, events, people, behaviors, expressions or text.
[0035] Further, please refer to Figure 5 As shown, the specific process and related details of the feasible embodiments of the present invention are illustrated. In the system initialization and setting stage, the present invention can first set the target category (or interest category) and retention time (or storage period) that the user is interested in: the user or system administrator first sets the target category (e.g., specific person, object, keyword (or key words), sound content, etc.) and the data retention time (how long the image, event or record is retained before it is automatically deleted). This setting will serve as the basis for all subsequent judgments and data filtering.
[0036] Secondly, regarding the image (visual) processing flow (left side flow) step S2: Gaze detection operation: The system activates the gaze tracking module to continuously detect the user's eye gaze direction and determine whether the user is currently paying attention to a specific area or object in the image. Regarding step S3: Is the user looking at the image? If not, the system continues gaze detection and returns to gaze detection operation. If yes, it means the user has paid attention to a certain area in the image, and the process proceeds to the next stage. Regarding step S4: Trigger object detection: The system activates the object recognition (or target recognition) model (e.g., AI image recognition) to analyze and classify objects in the image area being looked at by the user. Next, regarding step S5: Is the object a target category of interest? The identified object is compared with a pre-set target category of interest. If not, it is determined to be a non-target object, and the system does not store it, returning to the gaze detection flow. If yes, it is determined to be an object of interest to the user, and the process proceeds to data (or information) saving. Next, regarding step S6: cropping and storing object images: the system crops the image area related to the object of interest in the screen and stores the cropped image as an event, avoiding the retention of redundant background images. Furthermore, regarding step S7: removing expired image information: the system periodically checks the stored image data. This invention can be designed so that image data exceeding the "retention time" (or storage period) is automatically deleted to save storage space and meet privacy or data management needs. Additionally, it should be noted that in other embodiments, object images may not be cropped; instead, the captured full-screen image may be stored directly. In other words, in another embodiment, step S6 does not perform image cropping; instead, the full-screen image captured when the user's gaze occurs is stored directly.
[0037] In this case, the system can associate and store the full-screen image with at least one of the following information: the user's gaze area coordinates, object recognition results, object category tags, timestamps, or event identifiers, to facilitate subsequent retrieval, analysis, or reprocessing of image data based on the user's attention behavior. Furthermore, whether to crop the image can be dynamically selected based on system resource limitations, application scenarios, or privacy requirements. For example, in situations where storage space is limited or background interference needs to be reduced, cropped object images can be stored; while in applications where complete scene information needs to be preserved or for subsequent reanalysis, the full-screen image can be stored.
[0038] Secondly, regarding the speech and text processing workflow (right-hand flow), step S8: Speech and Text Recognition: The system simultaneously performs speech-to-text recognition, text content analysis, and converts speech into analyzable text data. Next, regarding step S9: Whether it relates to the target category of interest: Analyze whether the speech or text content contains keywords (or key terms), semantics, or themes related to the target category. If so, ignore the content and continue speech or text recognition. If yes, proceed to the information integration stage. Next, regarding step S10: Category (or target category) and time information integration (or summarization): Summarize the identified valid information, including: the target category of interest, the time of occurrence, the source (video or speech), and its association with video events (if they occurred simultaneously). Finally, regarding step S11: Information Output: Output the integrated information, for example: providing it to users for querying, displaying it on the interface, transmitting it to the backend system or database, or for subsequent analysis, reminders, or recording purposes. Furthermore, the user control commands include voice control, text input, gestures, actions, touch or button press control, and the target categories include objects, events, people, behaviors, expressions or text.
[0039] In summary, the features of the feasible embodiments of this invention are: 1. Combining gaze, image, voice, and text recognition. 2. Event (or object) driven, storing data only when the user is looking at the data and it matches their interest. 3. Privacy and efficiency oriented, saving only necessary images and reducing storage burden. 4. Automatically clearing expired data. 5. High scalability, with target categories and retention time (or storage period) conditions that can be freely set.
[0040] In short, this invention provides users with substantial and immediate assistance through gaze detection, going beyond simply using gaze information for statistical research. As AI applications grow, the target categories that users can review are no longer limited to objects, but can be expanded to people, events, behaviors, expressions, text, etc., becoming a secretary for visual images. It supports the practical application of gaze detection glasses in daily life, becoming a good personal assistant for users' AI applications.
[0041] The foregoing has provided a detailed description of a gaze detection device that can act as a visual imaging secretary, along with its setting and processing method. Specific embodiments have been used to illustrate the principles and implementation methods of this application. The descriptions of these embodiments are solely for the purpose of helping to understand the technical solutions and core ideas of this application. Those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions in the embodiments of this application.
[0042] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A line-of-sight detection device, characterized in that, include: An eye-tracking module includes a sensing device that senses eye information to estimate and determine the temporal changes in gaze, thereby distinguishing between saccades and fixations; and A target detection module includes multiple target categories for setting at least one target category of interest. In the gaze state, the target detection module can detect targets. When the target detection module determines that the image information used for comparison matches the target category of interest, it captures and stores the image information for the user to review and refer to.
2. The line-of-sight detection device as described in claim 1, characterized in that, The sensing device includes an eyeglasses device with two inward-facing detection lenses that are respectively directed toward the user's eyes. The eye-tracking module senses the line of sight through the two detection lenses. The eyeglasses device also has a scene camera lens facing outward, which provides image information for comparison. The eye information includes eye position information and eye image information.
3. The line-of-sight detection device as described in claim 1, characterized in that, The target detection module includes a first neural network processor, which is used to determine whether the user is in a saccade state or a gaze state. If it is determined to be in a gaze state and the target detection module performs target detection, and the image information used for comparison matches the target category of interest, then a partial image is extracted from the image information and the partial image is connected and stored in a smart device.
4. The line-of-sight detection device as described in claim 3, characterized in that, The gaze detection device includes a smart device and a first communication module. The smart device includes a second communication module, a detection management software program, and a storage unit. The target detection module includes a central processing unit, a graphics processing unit, and a second neural network processor. The first communication module and the second communication module can communicate wirelessly or via wired data transmission. The detection management software program can be used to allow the user to pre-set at least one target category of interest. The central processing unit, the graphics processing unit, and the second neural network processor are used to perform target detection calculations. The central processing unit controls the gaze detection device to display the previously categorized and stored image information of the target category through instruction control.
5. The line-of-sight detection device as described in claim 3, characterized in that, The smart device is a mobile device, which is a smartphone, tablet, or laptop.
6. The line-of-sight detection device as described in claim 3, characterized in that, The gaze detection device is integrated with the smart device as a laptop or tablet computer. The laptop or tablet computer has at least one detection lens facing the user to correspond to the user's eyes. The eye-tracking module senses the gaze through the at least one detection lens. The user provides the image information for comparison by browsing the screen display of data on the laptop or tablet computer.
7. A method for setting and processing line for gaze detection, characterized in that, The steps, in sequence, include: Specify at least one target category of interest; Eye information is obtained by sensing with a gaze detection device and the temporal changes of gaze are estimated and judged in order to distinguish between saccades and fixations. as well as If it is determined to be gaze, target detection is performed. If the image information used for comparison matches at least one target category of interest, the image information is captured and stored in a smart device via a connection for the user to review and refer to.
8. The gaze detection setting processing method as described in claim 7, characterized in that, The step of setting at least one target category of interest further includes setting the retention time (or storage period) of the image information, which is obtained by cropping the image and storing it after the target selection range output by a neural network processor.
9. The gaze detection setting processing method as described in claim 7, characterized in that, The gaze detection setting process further includes a user review instruction control step, in which the user uses an instruction to cause the gaze detection device or the smart device to display the image information of the previously categorized and stored target category.
10. The gaze detection setting processing method as described in claim 9, characterized in that, The user controls the instructions used, including voice control, text input, gestures, actions, touch or button presses, and the target categories include objects, events, people, behaviors, expressions or text.