Eye movement information extraction method and device and electronic equipment

By extracting and fusing the features of frame images and event images in eye movement tests, the problem of inability to effectively extract eye movement information in the prior art is solved, and a more accurate cognitive impairment assessment is achieved.

CN120093236AActive Publication Date: 2025-06-06HEBEI UNIVERSITY
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510592544.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-06
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The prior art cannot effectively extract valuable and accurate eye movement information from a large number of eye movement data, which affects the accuracy of individual cognitive impairment assessment.

Method used

By acquiring the frame image and event image of the subject during eye movement test, geometric feature extraction and gaze depth prediction value determination are performed, and fusion feature extraction is performed based on this information to obtain accurate eye movement information.

Benefits of technology

It realizes the effective extraction of accurate and valuable eye movement information from eye movement data, and improves the accuracy of individual cognitive impairment detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120093236A_ABST
    Figure CN120093236A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of data processing, and provides an eye movement information extraction method and device and electronic equipment. The method comprises the following steps: acquiring a frame image and an event image of an eye of a subject during an eye movement test; performing geometric feature extraction on the frame image, and determining a fixation depth prediction value according to the extracted geometric features; based on the fixation depth predicted value, performing fusion feature extraction on the frame image and the event image to obtain a fusion feature; and obtaining eye movement information according to the fusion features. According to the method, the frame image capable of providing stable visual information is obtained, the event image which is more sensitive for rapidly changing eye movement information is captured, the fixation depth predicted value is extracted, the fusion features are extracted from the frame image and the event image on the basis, accurate and valuable eye movement information can be effectively extracted from eye movement data, and the visual effect of the eye movement information is improved. And the accuracy of individual cognitive impairment detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, device and electronic device for extracting eye movement information. Background Art

[0002] Eye movement is coordinated and controlled by multiple brain regions, including the frontal eye movement area, parietal lobe, basal ganglia, thalamus and brainstem, and is an important external representation of brain function. Therefore, eye movement data obtained through eye movement tests can reflect specific brain functions, reveal individual problems in cognition, emotion and nervous system function, and be used to detect whether an individual has cognitive impairment.

[0003] A large amount of real-time data is generated during the eye movement test, including eye movement trajectory, gaze point, and other data. The data processing methods in related technologies cannot fully adapt to the complex and changeable eye movement data, and cannot effectively extract valuable and accurate eye movement information from a large amount of eye movement data, which in turn affects the accuracy of individual cognitive impairment assessment. Summary of the invention

[0004] In view of this, the embodiments of the present application provide an eye movement information extraction method, device and electronic device to solve the technical problem that related methods cannot effectively extract valuable and accurate eye movement information from a large amount of eye movement data, thereby affecting the accuracy of individual cognitive impairment assessment.

[0005] In a first aspect, an embodiment of the present application provides a method for extracting eye movement information, comprising: Acquire eye frame images and event images of the subject during the eye movement test; Extracting geometric features from the frame image, and determining a gaze depth prediction value according to the extracted geometric features; Based on the gaze depth prediction value, extracting fusion features from the frame image and the event image to obtain fusion features; Eye movement information is obtained according to the fusion features.

[0006] In a possible implementation of the first aspect, the extracting fusion features of the frame image and the event image based on the gaze depth prediction value to obtain the fusion features includes: Extracting eye movement features from the frame image to obtain a first eye movement feature vector; Based on the gaze depth prediction value, extracting eye movement features from the event image to obtain a second eye movement feature vector; A fusion feature is obtained according to the first eye movement feature vector and the second eye movement feature vector.

[0007] In a possible implementation of the first aspect, extracting eye movement features from the event image based on the gaze depth prediction value to obtain a second eye movement feature vector includes: According to the event image, a positive event point set and a negative event point set are extracted, and a time feature of the event is obtained by recursive filtering; wherein the positive event point set is determined based on a pixel position with increased brightness in the event image, and the negative event point set is determined based on a pixel position with decreased brightness in the event image; Constructing a three-channel event frame according to the positive event point set, the negative event point set and the time feature; Based on the gaze depth prediction value, a second eye movement feature vector is extracted from the three-channel event frame.

[0008] In a possible implementation of the first aspect, extracting a second eye movement feature vector from the three-channel event frame based on the gaze depth prediction value includes: Based on the gaze depth prediction value, dividing the three-channel event frame into different gaze areas; Extract eye movement features from each gaze area to obtain corresponding second eye movement features; A second eye movement feature vector is obtained according to the second eye movement feature corresponding to each gaze area.

[0009] In a possible implementation manner of the first aspect, the second eye movement feature vector includes a plurality of second sub-eye movement feature vectors; The step of extracting eye movement features from the event image based on the gaze depth prediction value to obtain a second eye movement feature vector comprises: Dividing the event image into a plurality of regions; the plurality of regions include a pupil region, an upper eyelid region and a light spot region; For each region, extract the sub-positive event point set and the sub-negative event point set in the region, and obtain the sub-time features of the events in the region through recursive filtering; Constructing a three-channel event frame corresponding to the region according to the sub-positive event point set, the sub-negative event point set and the sub-time feature; Based on the gaze depth prediction value, a second sub-eye movement feature vector is extracted from the three-channel event frame corresponding to the region.

[0010] In a possible implementation manner of the first aspect, obtaining a fusion feature according to the first eye movement feature vector and the second eye movement feature vector includes: Performing weighted fusion on the second sub-eye movement feature vectors to obtain a weighted fused feature vector; The weighted fused feature vector is concatenated with the first eye movement feature vector to obtain a fused feature.

[0011] In a possible implementation of the first aspect, extracting geometric features from the frame image and determining a gaze depth prediction value according to the extracted geometric features includes: Extracting geometric features from the frame image to obtain geometric features of the pupil in the frame image; A gaze depth prediction value is determined according to the geometric features of the pupil in the frame image and the geometric features of the pupil of the subject.

[0012] In a possible implementation manner of the first aspect, obtaining eye movement information according to the fusion feature includes: Eye movement information is obtained according to the fusion features and a preset eye movement model; wherein the eye movement model is obtained by training the fusion features of frame images and event images of different eyes, and the corresponding eye movement information.

[0013] In a second aspect, an embodiment of the present application provides an eye movement information extraction device, comprising: The acquisition module is used to acquire the frame images and event images of the eyes of the subject during the eye movement test.

[0014] The determination module is used to extract geometric features from the frame image and determine a gaze depth prediction value according to the extracted geometric features.

[0015] An extraction module is used to extract fusion features from the frame image and the event image based on the gaze depth prediction value to obtain fusion features.

[0016] The obtaining module is used to obtain eye movement information according to the fusion features.

[0017] In a third aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the eye movement information extraction method as described in any one of the first aspects is implemented.

[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the eye movement information extraction method as described in any one of the first aspects is implemented.

[0019] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here.

[0020] The eye movement information extraction method, device and electronic device provided in the embodiments of the present application can obtain frame images that can provide stable visual information and event images that are more sensitive to rapidly changing eye movement information when the subject is undergoing an eye movement test, and then extract fusion features from the above two images. By combining the above two images and obtaining eye movement features from different angles, it is possible to more comprehensively and accurately reflect the actual situation of eye movement and obtain effective and accurate eye movement information. At the same time, the present application extracts the gaze depth prediction value, and extracts fusion features from the frame image and the event image based on this. The gaze depth prediction value provides an additional information dimension for feature extraction, so that the extracted fusion features can better adapt to different eye movement scenes and individual differences, further improving the effectiveness and pertinence of the features. In this way, the present application can effectively extract accurate and valuable eye movement information from the eye movement data, thereby improving the accuracy of individual cognitive impairment detection.

[0021] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application; Figure 2 is a flow chart of an eye movement information extraction method provided in an embodiment of the present application; Figure 3 is a structural schematic diagram of an eye movement information extraction device provided in an embodiment of the present application; Figure 4 It is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0024] The present application is described more clearly below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the effects of the present application, but are not intended to limit the present application in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present application. These all fall within the scope of protection of the present application.

[0025] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0026] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0027] In the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0028] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0029] In addition, the “plurality” mentioned in the embodiments of the present application should be interpreted as two or more.

[0030] Eye movement data can reflect specific brain functions. Subjects with different cognitive and neurological functions may have different eye movement characteristics. The eye movement data obtained through eye movement tests can be used to evaluate the cognitive impairment of the subjects, provide data support for the detection and determination of cognitive impairment, and provide a basis for the formulation of subsequent personalized rehabilitation plans.

[0031] Eye movement tests include saccade tasks, positive saccade tasks, negative saccade tasks, and visual search tasks. During the eye movement test, the movement trajectory, fixation point, and movement speed of the subject's eyes can be tracked, generating a large amount of real-time data. However, related technologies cannot fully adapt to complex and changeable eye movement data, and cannot effectively extract valuable and accurate eye movement information from a large amount of eye movement data, which in turn affects the accuracy of individual cognitive impairment assessment.

[0032] Based on the above defects, when the subject is undergoing an eye movement test, the present application obtains a frame image that can provide stable visual information, and an event image that is more sensitive to capturing rapidly changing eye movement information, and then extracts fusion features from the above two images. By combining the above two images and obtaining eye movement features from different angles, the actual situation of eye movement can be more comprehensively and accurately reflected, and effective and accurate eye movement information can be obtained. At the same time, the present application extracts the gaze depth prediction value, and extracts fusion features from the frame image and the event image based on this. The gaze depth prediction value provides an additional information dimension for feature extraction, so that the extracted fusion features can better adapt to different eye movement scenarios and individual differences, further improving the effectiveness and pertinence of the features.

[0033] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below in conjunction with the accompanying drawings.

[0034] First reference Figure 1 , Figure 1 The application scenario of the present application is schematically shown, which may include an eye tracker 11, a display screen 12 and an electronic device 400.

[0035] When conducting an eye movement test, the subject wears an eye tracker 11, and the display screen 12 displays the corresponding stimulation points and controls the movement of the stimulation points according to the set content of the eye movement test, such as a saccade task, a positive saccade task, an anti-saccade task, and a visual search task, so as to guide the subject's gaze behavior. In the process of the subject looking at the stimulation point as required, the eye tracker 11 collects frame images and event images of the subject's eyes at a certain frequency. Taking the content of the eye movement test being a positive saccade task as an example, at the beginning of the eye movement test, a fixation point is displayed in the center of the display screen 12, and the subject needs to stare at the fixation point. After that, the fixation point disappears and the stimulation point appears, and the subject needs to quickly look at the stimulation point. During this process, the eye tracker 11 collects frame images and event images of the subject's eyes at a certain frequency.

[0036] The electronic device 400 obtains the above-mentioned frame images and event images collected by the eye tracker 11, extracts geometric features of the frame images, and determines a gaze depth prediction value based on the extracted geometric features. Based on the gaze depth prediction value, the electronic device 400 extracts fusion features of the frame images and the event images to obtain fusion features, and then obtains eye movement information based on the fusion features.

[0037] Here, the eye tracker 11 has a high sampling rate (such as 1000Hz), high precision (such as spatial resolution less than 0.006°, spatial precision 0.15°) and low latency (millisecond level), and can track and sample the subject's eyes in real time. The eye tracker 11 includes at least two cameras, which can be near-infrared enhanced cameras, wherein the camera with higher resolution is used as an event camera to collect event images, and the camera with lower resolution is used as a frame image acquisition device. The electronic device 400 can be a hardware device with data storage, processing and analysis functions. Among them, the electronic device 400 can also be a controller in the eye tracker 11.

[0038] Combine the following Figure 1 ,refer to Figure 2 To describe the eye movement information extraction method provided according to an exemplary embodiment of the present application.

[0039] Figure 2 FIG. 1 is a flow chart of a method for extracting eye movement information provided by an embodiment of the present application. Figure 2 As shown, the method in the embodiment of the present application may include: Step S201: Acquire frame images and event images of the subject's eyes during an eye movement test.

[0040] As mentioned above, the contents of the eye movement test may include a saccade task, a positive saccade task, an anti-saccade task, and a visual search task, etc. In this embodiment, there is no specific limitation on which one or several test contents are performed.

[0041] Frame images can include images of the subject's eyes, providing a static picture of the subject's eyes and periorbital area at a certain moment, containing rich texture, shape and other information, which can be used to analyze the overall state and relatively stable characteristics of the eyes. Event images asynchronously record brightness change events generated when the eyes move, which can include brightness changes of pixels when the eyes move. Event images are more sensitive to rapid eye movements and small changes, and can capture details that may be missed by frame images.

[0042] Step S202: extract geometric features from the frame image, and determine a gaze depth prediction value according to the extracted geometric features.

[0043] Exemplarily, the frame image contains geometric features of the eye, such as eye contour, pupil size and position, etc. Geometric features can be extracted by image processing algorithms. For example, in this embodiment, an edge detection algorithm can be used to determine the eye contour, and then based on the principle of similar triangles or other geometric models, combined with relevant size information of the camera that collects the frame image and relevant size information of the subject's eyes, the gaze depth prediction value is calculated.

[0044] Here, the gaze depth prediction value can reflect the depth of the subject's eye gaze point in space, providing an additional information dimension as a reference for subsequent fusion feature extraction, which is helpful to extract the movement characteristics of the eyes at different depths.

[0045] In some embodiments, when determining the gaze depth prediction value based on the extracted geometric features, geometric features of the frame image can be extracted to obtain the geometric features of the pupil in the frame image. Thereafter, the gaze depth prediction value is determined based on the geometric features of the pupil in the frame image and the geometric features of the subject's pupil.

[0046] Exemplarily, in this embodiment, an image segmentation algorithm, such as a threshold segmentation method, a semantic segmentation method based on deep learning, etc., can be used to accurately separate the pupil from the frame image. Then, the geometric features of the pupil in the frame image are calculated, such as the pupil diameter, area, ellipticity and / or pupil center coordinates. These geometric features can directly reflect the position and size changes of the pupil in the eye, and when the subject observes objects at different distances, the pupil size will change adaptively. Therefore, the change of the pupil is closely related to the gaze depth. Extracting the above geometric features can provide key data for the subsequent prediction of the gaze depth.

[0047] In this embodiment, based on the principle of similar triangles, according to the geometric features of the pupil in the frame image, the geometric features of the subject's pupil, and the focal length of the camera that collects the frame image, the distance from the subject's eyes to the camera can be obtained, and then the distance from the subject's eyes to the camera is used as an approximate value of the gaze depth, that is, the gaze depth prediction value is obtained.

[0048] Optionally, the diameter of the pupil in the frame image is taken as the geometric feature of the pupil in the frame image, and the diameter of the subject's pupil is taken as the geometric feature of the subject's pupil, then the expression of the gaze depth prediction value is: dimage / dreal=f / Z Where dimage is the diameter of the pupil in the frame image, dreal is the diameter of the subject's pupil, f is the focal length of the camera, that is, the distance from the optical center of the camera to the image plane, and Z is the distance from the subject's eyes to the camera, which is the predicted value of the gaze depth.

[0049] Step S203: Based on the gaze depth prediction value, extract fusion features from the frame image and the event image to obtain fusion features.

[0050] In one possible implementation, when obtaining the fusion feature, this embodiment can perform eye movement feature extraction on the frame image to obtain a first eye movement feature vector, and based on the gaze depth prediction value, perform eye movement feature extraction on the event image to obtain a second eye movement feature vector, and then obtain the fusion feature based on the first eye movement feature vector and the second eye movement feature vector.

[0051] In this embodiment, the eye movement features extracted from the frame image include appearance features and motion features. Among them, image processing algorithms, such as edge detection methods, morphological operation methods, etc., can be used to extract appearance features such as eye contour, pupil size and position. For example, in this embodiment, the Hough transform method is used to detect the circular contour of the pupil, calculate the radius and center coordinates of the pupil, and the threshold segmentation method is used to determine the boundary of the eye, that is, the eye contour, and then obtain the aspect ratio of the eye. These appearance features can reflect the basic shape and state of the eye, which are of great significance for analyzing the starting and ending positions of eye movement, the rotation direction of the eyeball, etc.

[0052] Optionally, in this embodiment, the changes of the eyes in adjacent frame images are compared to calculate the motion characteristics of the eyeball, such as the motion speed and acceleration. For example, the motion vector of the eyeball on the frame image plane can be calculated by the optical flow method, and then the motion speed, acceleration and direction of the eyeball can be calculated based on the size and direction of the motion vector. In addition, the blinking action can be detected based on the frame image, and the opening and closing degree of the eyelids and the blinking frequency can be calculated as motion characteristics.

[0053] After the appearance features and motion features are extracted from the frame image, the appearance features and motion features are normalized, and the normalized features are combined into a vector, that is, a first eye movement feature vector is obtained.

[0054] In some embodiments, when obtaining the second eye movement feature vector, a positive event point set and a negative event point set can be extracted according to the event image, and the time characteristics of the event can be obtained through recursive filtering. Thereafter, a three-channel event frame is constructed according to the positive event point set, the negative event point set and the time characteristics, and the second eye movement feature vector is extracted from the three-channel event frame based on the gaze depth prediction value.

[0055] The positive event point set is determined based on the pixel positions with increased brightness in the event image, and the negative event point set is determined based on the pixel positions with decreased brightness in the event image.

[0056] Exemplarily, the brightness change events generated when the eye moves are recorded in the event image. When the brightness of a certain pixel position increases, a positive event is generated, and when the brightness of a certain pixel position decreases, a negative event is generated. By processing the event image, the pixel positions corresponding to all positive events can be extracted to form a positive event point set, and the pixel positions corresponding to all negative events can be extracted to form a negative event point set. The above-mentioned positive event point set and negative event point set reflect the specific position information of the scene brightness change during eye movement. For example, when the subject's eyes scan quickly, a large number of positive and negative events will be generated in the corresponding area. For example, when the subject's eyes scan quickly from left to right, the brightness of the corresponding pixel position on the left side of the event image decreases, while the brightness of the corresponding pixel position on the right side increases.

[0057] Optionally, when analyzing the change pattern of an event over time, recursive filtering can be used to obtain the time characteristics of the event. For example, the time characteristics of the event are calculated by a recursive filtering algorithm. Among them, the recursive filtering algorithm is a filtering method that can adaptively track time changes. Through continuous iteration, events close to the current moment have a greater impact on the results, thereby retaining the time sequence of events and better reflecting the dynamic process of eye movement.

[0058] Afterwards, a three-channel event frame is constructed based on the positive event point set, the negative event point set and the time feature. Specifically, the positive event channel shows the distribution of positive events, and the pixel value is used to represent the distribution of positive events. The pixel position where the positive event occurs is assigned a higher pixel value, and other positions are assigned a lower value or even 0, so that the distribution of the area with increased brightness in the scene can be intuitively observed. Similarly, the negative event channel shows the distribution of negative events, and the pixel value is used to represent the distribution of negative events. The pixel position where the negative event occurs is assigned a lower pixel value, and other positions are assigned a higher value, so that the area with reduced brightness in the scene can be intuitively observed. The time feature channel records the time features of the event, and reflects the time accumulation information of events at different locations through pixel values. For example, the later the time occurs, the higher the pixel value of the pixel position. Afterwards, the positive event channel, the negative event channel and the time feature channel constitute a three-channel event frame. The three-channel event frame integrates the spatial and temporal information of the event, providing a rich data foundation for the subsequent extraction of more effective and representative eye movement features.

[0059] In some embodiments, when extracting the second eye movement feature vector from the three-channel event frame, the three-channel event frame can be divided into different gaze areas based on the gaze depth prediction value, and eye movement features are extracted for each gaze area to obtain the corresponding second eye movement features. Then, the second eye movement feature vector is obtained according to the corresponding second eye movement features of each gaze area.

[0060] Exemplarily, according to the gaze depth prediction value, the three-channel frame can be divided into gaze areas with different gaze depths. Among them, different gaze depths may correspond to different eye movement patterns and areas of interest. For example, when the subject looks at a closer object, the micro-movement of the eye and the change in focal length are more obvious. In this way, different gaze areas also correspond to different eye movement patterns and areas of interest due to different gaze depths. Therefore, the gaze depth prediction value provides another information dimension for the extraction of the second eye movement feature, so that the extracted second eye movement feature can better adapt to different eye movement scenes and individual differences, thereby improving the effectiveness and pertinence of the second eye movement feature.

[0061] Optionally, a corresponding second eye movement feature is extracted from each gaze area, and a second eye movement feature vector is constructed based on all the second eye movement features. In this way, the second eye movement feature vector contains feature information of the event image in different areas and at different time scales, and can more comprehensively describe the features and patterns of the eye movement.

[0062] Afterwards, a fusion feature can be obtained according to the first eye movement feature vector and the second eye movement feature vector. For example, according to the timestamp of the first eye movement feature vector and the timestamp of the second eye movement feature vector, the first eye movement feature vector and the second eye movement feature vector are spliced ​​to finally obtain the fusion feature of the two. In this way, the frame image that can provide stable visual information and the event image that is more sensitive to capturing rapidly changing eye movement information can be combined to obtain eye movement features from different angles, thereby obtaining effective and accurate eye movement information.

[0063] Optionally, since the eye movement features in different regions have different physiological and behavioral meanings, in order to more effectively extract the second eye movement feature, the event image may be divided into regions, and the eye movement features may be extracted for each region.

[0064] In some embodiments, the second eye movement feature vector includes multiple second sub-eye movement feature vectors. When obtaining the second eye movement feature vector, the event image can be first divided into multiple regions, and the multiple regions can include a pupil region, an upper eyelid region, and a light spot region. Then, for each region, a sub-positive event point set and a sub-negative event point set in the region are extracted, and the sub-time features of the events in the region are obtained through recursive filtering. According to the sub-positive event point set, the sub-negative event point set and the sub-time features, a three-channel event frame corresponding to the region is constructed, and based on the gaze depth prediction value, the second sub-eye movement feature vector is extracted from the three-channel event frame corresponding to the region.

[0065] In this embodiment, the event image is divided into multiple areas such as the pupil area, the upper eyelid area and the light spot area because the eye movement features of the above different areas have different physiological and behavioral meanings. For example, the size and position changes of the pupil directly reflect the gaze direction and the degree of concentration of the eye. In the event image, the events in the pupil area are mainly related to the rotation of the eyeball and the scaling of the pupil. The movement of the upper eyelid, such as blinking, is an important eye movement feature. The change in blinking frequency can reflect the fatigue level and mental state of the subject. In the event image, the events in the upper eyelid area are mainly related to the opening and closing movement of the eyelid. The light spot is usually formed by the reflected light generated by the external light source irradiating the eye. There is a relatively stable positional relationship between the light spot and the pupil. By tracking the position change of the light spot, the position and movement trajectory of the pupil can be assisted. In the event image, the events in the light spot area are mainly related to the movement and brightness change of the light spot. Therefore, the event image is divided into multiple areas with different physiological and behavioral meanings, and the eye movement feature extraction is performed on each area respectively, so that the second eye movement feature can be extracted more effectively.

[0066] Exemplarily, in this embodiment, for each area in the pupil area, upper eyelid area and light spot area, the sub-positive event point set and the sub-negative event point set in the area are extracted, and the specific implementation process and principle of obtaining the sub-time characteristics of the events in the area through recursive filtering processing can refer to the steps of extracting the positive event point set and the negative event point set according to the event image, and obtaining the time characteristics of the event through recursive filtering processing in the aforementioned embodiment, and will not be repeated here.

[0067] Similarly, in the present embodiment, for each region, a three-channel event frame corresponding to the region is constructed, and the specific implementation process and principle of extracting the second sub-eye movement feature vector from the three-channel event frame corresponding to the region based on the gaze depth prediction value can refer to the steps of constructing the three-channel event frame according to the positive event point set, the negative event point set and the time feature, and extracting the second eye movement feature vector from the three-channel event frame based on the gaze depth prediction value in the aforementioned embodiment, which will not be repeated here.

[0068] Exemplarily, when obtaining the fused feature, each second sub-eye movement feature vector may be weightedly fused to obtain a weighted fused feature vector, and the weighted fused feature vector may be concatenated with the first eye movement feature vector to obtain the fused feature.

[0069] In this embodiment, the weights corresponding to the pupil area, upper eyelid area and light spot area can be set as needed. For example, the light spot area is mainly used to assist in determining the position and movement trajectory of the pupil, so the weight of the light spot area can be lower than the weights of the pupil area and the upper eyelid area, that is, the weight of the second sub-eye movement feature vector corresponding to the light spot area is lower than the weight of the second sub-eye movement feature vector corresponding to the pupil area and the weight of the second sub-eye movement feature vector corresponding to the upper eyelid area. Of course, the setting of the corresponding weights of each area can also be set according to other needs, and no specific restrictions are made here. Afterwards, the feature vector after weighted fusion is spliced ​​with the first eye movement feature vector to obtain a fused feature.

[0070] In this way, the event image is first divided into regions to obtain multiple regions with eye movement features of different physiological and behavioral significances, and then eye movement features are extracted for each region, so that the second eye movement feature can be extracted more effectively and accurately.

[0071] Step S204: Obtain eye movement information based on the fusion features.

[0072] Exemplarily, eye movement information can be obtained based on the fusion features and the preset eye movement model. Here, the preset eye movement model can be based on a hybrid neural network architecture, have multi-modal parallel processing capabilities, and achieve real-time multi-task output. The eye movement model inputs the fusion features and outputs the eye movement information, and the eye movement model is obtained by training the fusion features of frame images and event images of different eyes and the corresponding eye movement information.

[0073] Optionally, the eye movement information may include saccade amplitude, saccade duration, saccade latency, eyeball acceleration, etc. It should be noted that the frame image and the event image are fused to extract features to obtain fusion features, and the eye movement information is obtained based on the fusion features. After multiple frame images and multiple event images are collected, multiple eye movement information can be obtained accordingly. The same type of eye movement information can be averaged to prevent accidental abnormalities caused by environmental factors from affecting the eye movement information, thereby improving the accuracy of the eye movement information. The fusion features obtained by combining frame images that can provide stable visual information with event images that are more sensitive to capturing rapidly changing eye movement information can more comprehensively and accurately reflect the actual situation of eye movement, and then based on the above fusion features and the preset eye movement model, effective and accurate eye movement information can be obtained.

[0074] As mentioned above, subjects with impaired cognitive and neurological functions usually show characteristic abnormalities in eye movement tests, which are then reflected in the eye movement information, that is, eye movement information can represent brain function. For example, a prolonged saccade latency indicates that there may be functional degeneration or damage in the frontal lobe eye movement area, resulting in slow saccade initiation. A reduced saccade amplitude indicates that the dopaminergic neurons in the basal ganglia may be degenerated, resulting in short saccades.

[0075] After obtaining the eye movement information, each eye movement information can be compared with the corresponding judgment threshold, and then the comparison results of each eye movement information can be combined to evaluate the cognitive function and neural function of the subject to determine whether the subject has cognitive impairment. For example, the normal range of eye saccade latency for healthy adults is 150~250ms, then the judgment threshold corresponding to the eye saccade latency can be set to 250ms. When the eye saccade latency of the subject is greater than the judgment threshold, it is confirmed that the subject may have cognitive impairment. It should be noted that the normal range of eye saccade latency for healthy adults mentioned above is only for example.

[0076] The eye movement information extraction method provided in the embodiment of the present application obtains frame images and event images of the subject's eyes during an eye movement test, extracts geometric features from the frame images, determines a gaze depth prediction value based on the extracted geometric features, and based on the gaze depth prediction value, extracts fusion features from the frame images and the event images to obtain fusion features, and then obtains eye movement information based on the fusion features.

[0077] When the subject is undergoing an eye movement test, the present application obtains a frame image that can provide stable visual information, and an event image that is more sensitive to capturing rapidly changing eye movement information, and then extracts fusion features from the above two images. By combining the above two images and obtaining eye movement features from different angles, the actual situation of eye movement can be more comprehensively and accurately reflected, and effective and accurate eye movement information can be obtained. At the same time, the present application extracts the gaze depth prediction value, and extracts fusion features from the frame image and the event image based on this. The gaze depth prediction value provides an additional information dimension for feature extraction, so that the extracted fusion features can better adapt to different eye movement scenes and individual differences, further improving the effectiveness and pertinence of the features. In this way, the present application can effectively extract accurate and valuable eye movement information from the eye movement data, thereby improving the accuracy of individual cognitive impairment detection.

[0078] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0079] Figure 3 Schematic diagram of the structure of an eye movement information extraction device provided by an embodiment of the present application. Figure 3 As shown, the eye movement information extraction device provided in this embodiment may include: an acquisition module 301 , a determination module 302 , an extraction module 303 and a obtaining module 304 .

[0080] The acquisition module 301 is used to acquire the frame images and event images of the eyes of the subject during the eye movement test.

[0081] The determination module 302 is used to extract geometric features from the frame image and determine a gaze depth prediction value according to the extracted geometric features.

[0082] The extraction module 303 is used to extract fusion features from the frame image and the event image based on the gaze depth prediction value to obtain fusion features.

[0083] The obtaining module 304 is used to obtain eye movement information according to the fusion feature.

[0084] Optionally, the extraction module 303 is further used for: Extracting eye movement features from the frame image to obtain a first eye movement feature vector; Based on the gaze depth prediction value, extracting eye movement features from the event image to obtain a second eye movement feature vector; A fusion feature is obtained according to the first eye movement feature vector and the second eye movement feature vector.

[0085] Optionally, the extraction module 303 is further used for: According to the event image, a positive event point set and a negative event point set are extracted, and a time feature of the event is obtained by recursive filtering; wherein the positive event point set is determined based on a pixel position with increased brightness in the event image, and the negative event point set is determined based on a pixel position with decreased brightness in the event image; Constructing a three-channel event frame according to the positive event point set, the negative event point set and the time feature; Based on the gaze depth prediction value, a second eye movement feature vector is extracted from the three-channel event frame.

[0086] Optionally, the extraction module 303 is further used for: Based on the gaze depth prediction value, dividing the three-channel event frame into different gaze areas; Extract eye movement features from each gaze area to obtain corresponding second eye movement features; A second eye movement feature vector is obtained according to the second eye movement feature corresponding to each gaze area.

[0087] Optionally, the second eye movement feature vector includes a plurality of second sub-eye movement feature vectors; and the extraction module 303 is further used for: Dividing the event image into a plurality of regions; the plurality of regions include a pupil region, an upper eyelid region and a light spot region; For each region, extract the sub-positive event point set and the sub-negative event point set in the region, and obtain the sub-time features of the events in the region through recursive filtering; Constructing a three-channel event frame corresponding to the region according to the sub-positive event point set, the sub-negative event point set and the sub-time feature; Based on the gaze depth prediction value, a second sub-eye movement feature vector is extracted from the three-channel event frame corresponding to the region.

[0088] Optionally, the extraction module 303 is further used for: Performing weighted fusion on the second sub-eye movement feature vectors to obtain a weighted fused feature vector; The weighted fused feature vector is concatenated with the first eye movement feature vector to obtain a fused feature.

[0089] Optionally, the determination module 302 is further configured to: Extracting geometric features from the frame image to obtain geometric features of the pupil in the frame image; A gaze depth prediction value is determined according to the geometric features of the pupil in the frame image and the geometric features of the pupil of the subject.

[0090] Optionally, the obtaining module 304 is further used for: Eye movement information is obtained according to the fusion features and a preset eye movement model; wherein the eye movement model is obtained by training the fusion features of frame images and event images of different eyes, and the corresponding eye movement information.

[0091] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0092] Figure 4 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 4 As shown, the electronic device 400 of this embodiment includes: a processor 410 and a memory 420, wherein the memory 420 stores a computer program 421 that can be run on the processor 410. When the processor 410 executes the computer program 421, the steps in any of the above method embodiments are implemented, for example Figure 2 Alternatively, when the processor 410 executes the computer program 421, the functions of each module / unit in the above-mentioned device embodiments are implemented, for example Figure 3 Functions of modules 301 to 304 are shown.

[0093] Exemplarily, the computer program 421 may be divided into one or more modules / units, one or more modules / units are stored in the memory 420, and are executed by the processor 410 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 421 in the electronic device 400.

[0094] Those skilled in the art will understand that Figure 4 These are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those shown in the figure, or a combination of certain components, or different components, such as input and output devices, network access devices, buses, etc.

[0095] The processor 410 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0096] The memory 420 may be an internal storage unit of the electronic device, such as a hard disk or memory of the electronic device, or an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. The above-mentioned memory 420 may also include both an internal storage unit of the electronic device and an external storage device. The above-mentioned memory 420 is used to store computer programs and other programs and data required by the electronic device. The memory 420 may also be used to temporarily store data that has been output or is to be output.

[0097] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0098] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0099] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0100] In the embodiments provided by the present invention, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0101] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0103] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0104] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for extracting eye movement information, characterized in that: include: Acquire eye frame images and event images of the subject during the eye movement test; Extracting geometric features from the frame image, and determining a gaze depth prediction value according to the extracted geometric features; Based on the gaze depth prediction value, extracting fusion features from the frame image and the event image to obtain fusion features; Eye movement information is obtained according to the fusion features.

2. The eye movement information extraction method according to claim 1, characterized in that: The step of extracting fusion features from the frame image and the event image based on the gaze depth prediction value to obtain fusion features includes: Extracting eye movement features from the frame image to obtain a first eye movement feature vector; Based on the gaze depth prediction value, extracting eye movement features from the event image to obtain a second eye movement feature vector; A fusion feature is obtained according to the first eye movement feature vector and the second eye movement feature vector.

3. The eye movement information extraction method according to claim 2, characterized in that: The step of extracting eye movement features from the event image based on the gaze depth prediction value to obtain a second eye movement feature vector comprises: According to the event image, a positive event point set and a negative event point set are extracted, and a time feature of the event is obtained by recursive filtering; wherein the positive event point set is determined based on a pixel position with increased brightness in the event image, and the negative event point set is determined based on a pixel position with decreased brightness in the event image; Constructing a three-channel event frame according to the positive event point set, the negative event point set and the time feature; Based on the gaze depth prediction value, a second eye movement feature vector is extracted from the three-channel event frame.

4. The eye movement information extraction method according to claim 3, characterized in that: The step of extracting a second eye movement feature vector from the three-channel event frame based on the gaze depth prediction value comprises: Based on the gaze depth prediction value, dividing the three-channel event frame into different gaze areas; Extract eye movement features from each gaze area to obtain corresponding second eye movement features; A second eye movement feature vector is obtained according to the second eye movement feature corresponding to each gaze area.

5. The eye movement information extraction method according to claim 2, characterized in that: The second eye movement feature vector includes a plurality of second sub-eye movement feature vectors; The step of extracting eye movement features from the event image based on the gaze depth prediction value to obtain a second eye movement feature vector comprises: Dividing the event image into a plurality of regions; the plurality of regions include a pupil region, an upper eyelid region and a light spot region; For each region, extract the sub-positive event point set and the sub-negative event point set in the region, and obtain the sub-time features of the events in the region through recursive filtering; Constructing a three-channel event frame corresponding to the region according to the sub-positive event point set, the sub-negative event point set and the sub-time feature; Based on the gaze depth prediction value, a second sub-eye movement feature vector is extracted from the three-channel event frame corresponding to the region.

6. The eye movement information extraction method according to claim 5, characterized in that: The obtaining of a fusion feature according to the first eye movement feature vector and the second eye movement feature vector comprises: Performing weighted fusion on the second sub-eye movement feature vectors to obtain a weighted fused feature vector; The weighted fused feature vector is concatenated with the first eye movement feature vector to obtain a fused feature.

7. The eye movement information extraction method according to any one of claims 1 to 6, characterized in that: The step of extracting geometric features from the frame image and determining a gaze depth prediction value according to the extracted geometric features includes: Extracting geometric features from the frame image to obtain geometric features of the pupil in the frame image; A gaze depth prediction value is determined according to the geometric features of the pupil in the frame image and the geometric features of the pupil of the subject.

8. The eye movement information extraction method according to any one of claims 1 to 6, characterized in that: Obtaining eye movement information according to the fusion feature includes: Eye movement information is obtained according to the fusion features and a preset eye movement model; wherein the eye movement model is obtained by training the fusion features of frame images and event images of different eyes, and the corresponding eye movement information.

9. An eye movement information extraction device, characterized in that: include: An acquisition module, used for acquiring frame images and event images of the subject's eyes during an eye movement test; A determination module, used for extracting geometric features from the frame image, and determining a gaze depth prediction value according to the extracted geometric features; An extraction module, configured to extract fusion features from the frame image and the event image based on the gaze depth prediction value to obtain fusion features; The obtaining module is used to obtain eye movement information according to the fusion features.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the eye movement information extraction method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Matching feature-based three-dimensional fixation point estimation and three-dimensional eye movement model establishment method

    CN113158879A

  • Pose estimation method and related device

    CN115997234A

  • Eye image processing method and device, eye movement tracking system and electronic device

    CN116434317A

  • Eye movement tracking method, device, equipment, medium and program

    CN117971040A

  • Sight line estimation method and electronic equipment

    CN118301314A