Eye movement video quality detection method and system
By analyzing the physical characteristics and content of eye movement videos, combining Haar cascade detector and deep learning model, the quality of eye movement videos is evaluated and low-quality videos are automatically screened out, which solves the problem of uneven video quality in eye movement tracking, and improves the accuracy of tracking results and the efficiency of data acquisition.
Patent Information
- Application Number
- CN202510726723.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
In the existing eye movement tracking methods, the quality of eye movement videos is uneven, which makes the accuracy and stability of the tracking results unable to be guaranteed, affecting the efficiency and reliability of data acquisition.
By extracting physical characteristics such as frame rate, resolution and light conditions of eye movement videos, combining Haar cascade detector and deep learning model to detect the position of the face and eye, calculate factors such as head position, orientation and opening, comprehensively evaluate the user's seriousness, and automatically screen out poor quality videos.
It realizes an accurate and comprehensive evaluation of the quality of eye movement videos, ensures the accuracy and stability of eye movement tracking results, and improves the efficiency and reliability of data acquisition.
Smart Images

Figure CN120260136A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of eye movement tracking technology, and particularly relates to a method and system for detecting the quality of eye movement videos. Background Art
[0002] In the field of eye movement tracking, existing video quality assessment methods include focus image / video compression algorithms, methods based on autocorrelation analysis, objective video quality assessment algorithms, "MOSp metric" video quality measurement methods, subjective video quality assessment methods, etc.
[0003] The focus image / video compression algorithm is a commonly used video quality assessment method. This method evaluates the quality by quantifying the distortion degree of the image or video. Commonly used metrics include PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index). These metrics only consider the physical characteristics of the image and ignore the accurate simulation of human subjective perception. Therefore, they cannot fully and accurately evaluate the quality of eye movement videos.
[0004] The video quality assessment method based on autocorrelation analysis analyzes the spatial autocorrelation in the image or video to evaluate the quality, mainly focusing on temporal and spatial distortions and ignoring more cognitive factors such as color information and factors such as eye movements that need to be concerned about in eye movement tracking.
[0005] The objective video quality assessment algorithm is a method based on machine learning and artificial intelligence. It trains a model to map subjective scores to predict video quality. However, this method has certain subjectivity and subjective errors, and there has been no in-depth study on the quality assessment of eye movement videos for eye movement tracking applications.
[0006] The "MOSp metric" method is a method for measuring the quality of eye movement videos based on human subjective evaluation. It uses subjects to subjectively score eye movement videos and converts the subjective scores into objective quality indicators by establishing a model. However, this method requires a large amount of time and financial investment, and there are certain subjectivity and subjective errors in the evaluation results.
[0007] In summary, although existing video quality assessment methods provide some references, they still have problems in not being able to fully and accurately evaluate the quality of eye movement videos. The quality of eye movement videos recorded by existing eye movement tracking methods is unstable, including problems such as low video resolution, unstable frame rate, inconsistent exposure, etc., resulting in the inability to guarantee the accuracy of eye movement tracking results.
[0008] During the video data acquisition process, there are some uncontrollable factors. Some users fail to accurately align their heads with the camera when recording eye movement videos, resulting in a restricted head in the camera's field of view; some users' heads are not fully exposed within the camera's acquisition range, making eye movement tracking impossible; in addition, the problem that some users have a relatively small eye opening degree further affects the accuracy of eye movement tracking; during the eye movement experiment, some users do not complete the tasks seriously, and all these factors will affect the subsequent data analysis of the experiment.
[0009] The accuracy of eye movement tracking is directly affected by the video quality. When the facial video quality is poor, it will lead to a decrease in the accuracy of eye movement tracking results and cannot meet the application requirements.
[0010] In response to the above problems, this application proposes a method and system for detecting the quality of eye movement videos. Summary of the Invention
[0011] In view of one or more technical defects in the above prior art, this application proposes the following technical solutions.
[0012] Based on the first aspect of this application, a method for detecting the quality of eye movement videos is proposed, including: S1: Extract the frame rate and resolution of the eye movement video, analyze the basic information of the eye movement video for video basic information evaluation, and the evaluation formula is: Score1=(F×W f +R×W r +S×W s ) / (W f +W r +W s ); Wherein, Score1 represents the evaluation score of the video basic information, F, R, and S respectively represent the frame rate score, resolution score, and frame rate stability score, and W f , W r , W s respectively represent the frame rate score weight, resolution score weight, and frame rate stability score weight; S2: Convert the eye movement video into a grayscale image, perform brightness analysis on the grayscale image to obtain the actual light value, and then perform numerical mapping to obtain the actual brightness value; Use the Haar cascade detector to detect the face position, divide the frontal face area by locating the face key points, and obtain the brightness of the frontal face area and the background brightness through image processing algorithms; Evaluate the lighting conditions according to the actual brightness value, the brightness of the frontal face area, and the background brightness, and the calculation formula is: Score2=(L×W l +M×W m) / (W l +W m ); L = 1 - |L s - L0| / L0; M = MIN(L zm / 2 × L bj , 1); Among them, Score2 represents the light condition score, L represents the brightness score, L s represents the actual brightness value, L0 represents the standard brightness, L zm represents the brightness of the frontal face area, L bj represents the background brightness, M represents the facial backlight score, W l and W m respectively represent the brightness score weight and the facial backlight score weight; S3: Analyze the head position, eye position, head orientation, head integrity, and eye opening degree in the eye movement video, calculate the head position score, head integrity score, head orientation score, and eye opening degree score respectively, and obtain the user recording process score; S4: Calculate the user's seriousness score according to the time when the user completes the eye movement; S5: Calculate the comprehensive video quality score. If the comprehensive video quality score is higher than the pre-set excellent score line, it is determined as a high-quality video. If the comprehensive video quality score is lower than the pre-set passing score line, it is determined as a low-quality video and the video is automatically screened out.
[0013] Furthermore, the calculation formula for the frame rate score is: F = (F s - F min ) / (F max - F min ); Among them, F represents the frame rate score, F s represents the actual frame rate, F min represents the lowest standard frame rate, which is the lowest frame rate threshold that the eye movement analysis system can accept, F max represents the highest standard frame rate, which is the frame rate threshold that the eye movement analysis system expects to reach perfectly; The calculation formula for the frame rate stability score is: S = (S0 - S min ) / (S max - S min ); Among them, S represents the frame rate stability score, S0 represents the frame rate stability, calculated according to the fluctuation of the video frame rate, S min represents the lowest standard stability, which is the worst stability threshold that the eye movement analysis system can accept, Smax Indicates the highest standard of stability, which is the perfect stability threshold expected for the eye movement analysis system; The calculation formula for the resolution score is: R=(R s -R min ) / (R max -R min ); Wherein, R represents the resolution score, R s represents the actual resolution, R min represents the lowest standard resolution, which is the lowest resolution threshold that the eye movement analysis system can accept, and R max represents the highest standard resolution, which is the perfect resolution threshold expected for the eye movement analysis system.
[0014] Furthermore, a Haar cascade classifier based on the Viola-Jones framework, Histogram of Oriented Gradients (HOG), or a deep learning model is used to detect the face position, face key points, eye key points, and eye position in the eye movement video.
[0015] Furthermore, the calculation formulas for the user recording process score are respectively: Score3=(A1×W A1 +A2×W A2 +A3×W A3 +A4×W A4 ) / (W A1 +W A2 +W A3 +W A4 )θ represents; A1 = 1 - |the proportion of the head occupied - the standard proportion| / the standard proportion; A2 = the actual head area occupied on the screen / the entire head area; A3 = 1 - |the actual head orientation - the correct orientation| / the correct orientation; A4 = (the actual opening degree - the lowest standard opening degree) / (the highest standard opening degree - the lowest standard opening degree); Wherein, Score3 represents the user recording process score, A1, A2, A3, and A4 respectively represent the head position score, head integrity score, head orientation score, and eye opening degree score, and W A1 , W A2 , W A3 and W A4 respectively represent the head position score weight, head integrity score weight, head orientation score weight, and eye opening degree score weight; The proportion of the head occupied is the ratio of the number of pixels in the head area to the total number of pixels in the eye movement video.
[0016] Furthermore, the head integrity is obtained by the following method: Use the Haar cascade algorithm to initially screen the face positions and verify the face positions through the MTCNN model. Then, combined with the video resolution, adopt the dynamic anchor box scaling technology to adjust the scale of the detection window, select the largest one from the detected face positions as the head position, and convert the area of the head position into a rectangular bounding box as the head region; Extract the boundary of the head region, fit the actually occupied head region on the screen in combination with the geometric projection model, and dynamically correct the estimated value of the area of the actually occupied head region on the screen to obtain the final actually occupied head region on the screen; The head integrity is the ratio of the estimated value of the area of the actually occupied head region to the area of the entire head region.
[0017] Introducing the MTCNN model for fine verification on the basis of initially screening with the Haar cascade algorithm solves the problem of false detection of a single algorithm in complex lighting or occlusion scenarios.
[0018] Adopting the dynamic anchor box scaling technology to adjust the scale of the detection window can avoid missed detection caused by a fixed window.
[0019] Dynamically correcting the estimated value of the area of the actually occupied head region on the screen can reduce the error accumulation caused by pose changes in the traditional method and significantly improve the measurement stability in complex scenarios.
[0020] Furthermore, for the actual head orientation, determine the position of the tip of the nose according to the face key points, calculate the difference in the abscissa and the difference in the ordinate between the coordinates of the tip of the nose and the center coordinates of the head region, and calculate the head orientation angle. The calculation formula is: θ = atan2(Δy, Δx) × 180 / π; where θ represents the head orientation angle, Δy is the difference in the ordinate between the coordinates of the tip of the nose and the center coordinates of the head region, and Δx is the difference in the abscissa between the coordinates of the tip of the nose and the center coordinates of the head region.
[0021] Furthermore, detect the positions of the eye key points according to the positions of the human eyes. The eye key points include the inner corner, outer corner, upper eyelid, and lower eyelid of the eyes. Take the distance between the upper eyelid and the lower eyelid as the height of the eyes, and take the distance between the inner corner and the outer corner as the opening width of the eyes. The actual opening degree is the ratio of the opening width to the height of the eyes.
[0022] Furthermore, the calculation formula for the score of the user's attentiveness is: Score4 = 1 - |D s - D0| / D0; Among them, Score4 represents the score of the user's conscientiousness, and D s represents the actual time spent by the user to complete the eye movement. D0 represents the standard completion time, which is reasonably set according to the difficulty and complexity of the task; The calculation formula for the comprehensive video quality score is as follows: Score=(Score1×W1+Score2×W2+Score3×W3+Score4×W4) / (W1+W2+W3+W4); Among them, Score is the comprehensive video quality score, Score1 represents the evaluation score of the basic video information, Score2 represents the score of the light condition, Score3 represents the score of the user's recording process, Score4 represents the score of the user's conscientiousness, and W1, W2, W3, and W4 respectively represent the weights of the evaluation score of the basic video information, the score of the light condition, the score of the user's recording process, and the score of the user's conscientiousness.
[0023] Based on the second aspect of this application, a detection system for the quality of eye movement videos is also proposed, including: Basic video information evaluation module: Extract the frame rate and resolution of the eye movement video, analyze the basic information of the eye movement video for basic video information evaluation, and the evaluation formula is: Score1=(F×W f +R×W r +S×W s ) / (W f +W r +W s ); Among them, Score1 represents the evaluation score of the basic video information, F, R, and S respectively represent the frame rate score, the resolution score, and the frame rate stability score, and W f , W r , W s respectively represent the frame rate score weight, the resolution score weight, and the frame rate stability score weight; Light condition evaluation module: Convert the eye movement video into a grayscale image, perform brightness analysis on the grayscale image to obtain the actual light value, and then perform numerical mapping to obtain the actual brightness value; Use the Haar cascade detector to detect the face position, divide the frontal face area by locating the face key points, and obtain the brightness of the frontal face area and the background brightness through the image processing algorithm; Evaluate the light condition according to the actual brightness value, the brightness of the frontal face area, and the background brightness, and the calculation formula is: Score2=(L×W l +M×W m ) / (Wl +W m ); L = 1 - |L s - L0| / L0; M = MIN(L zm / 2 × L bj ,1); Wherein, Score2 represents the light condition score, L represents the brightness score, L s represents the actual brightness value, L0 represents the standard brightness, L zm represents the brightness of the frontal face area, L bj represents the background brightness, M represents the face backlight score, W l and W m respectively represent the brightness score weight and the face backlight score weight; Recording process evaluation module: Analyze the head position, eye position, head orientation, head integrity, and eye opening degree in the eye movement video, calculate the head position score, head integrity score, head orientation score, and eye opening degree score respectively, and obtain the user recording process score; User seriousness scoring module: Calculate the user seriousness score according to the time taken by the user to complete the eye movement; Comprehensive evaluation module: Calculate the comprehensive video quality score. If the comprehensive video quality score is higher than the pre-set excellent score line, it is determined as a high-quality video. If the comprehensive video quality score is lower than the pre-set passing score line, it is determined as a low-quality video and the video is automatically screened out.
[0024] This application saves a large amount of human resources, improves the efficiency and reliability of data collection, and effectively solves the problem of uneven eye movement video quality in existing eye movement tracking methods. By automatically removing eye movement videos with poor quality, the accuracy and stability of eye movement tracking results are ensured.
[0025] Based on the third aspect of this application, a computer program product is also proposed, which has one or more computer programs. When the computer program product is executed by a computer processor, the method described in any one of the above is implemented.
[0026] The technical effect of this application is as follows: By analyzing the physical characteristics and content of the eye movement video, this application more accurately and comprehensively evaluates the video quality, effectively solves the problem of uneven eye movement video quality in existing eye movement tracking methods. By automatically removing eye movement videos with poor quality, the accuracy and stability of eye movement tracking results can be ensured, and the efficiency and reliability of data collection can be improved. Description of the Drawings
[0027] Other features, objects, and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments read in conjunction with the accompanying drawings.
[0028] Figure 1 is an overall flowchart of a method for detecting the quality of an eye movement video provided according to an embodiment of the present application; Figure 2 is an overall framework diagram of a system for detecting the quality of an eye movement video provided according to an embodiment of the present application; Figure 3 is a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed Embodiments
[0029] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only for explaining the relevant invention and not for limiting the invention. Additionally, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings.
[0030] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0031] Reference is made below to Figure 1 , Figure 1 which shows a method for detecting the quality of an eye movement video, including: S1: Extract the frame rate and resolution of the eye movement video, analyze the basic information of the eye movement video for video basic information evaluation, and the evaluation formula is: Score1=(F×W f +R×W r +S×W s ) / (W f +W r +W s ); where Score1 represents the evaluation score of the video basic information, F, R, and S respectively represent the frame rate score, resolution score, and frame rate stability score, and W f , W r , W s respectively represent the frame rate score weight, resolution score weight, and frame rate stability score weight; S2: Convert the eye movement video into a grayscale image, perform brightness analysis on the grayscale image to obtain the actual light value, and then perform numerical mapping to obtain the actual brightness value; Use the Haar cascade detector to detect the face position, divide the frontal face area by locating the face key points, and obtain the brightness of the frontal face area and the background brightness through the image processing algorithm; Evaluate the lighting condition according to the actual brightness value, the brightness of the frontal face area and the background brightness. The calculation formula is: Score2=(L×W l +M×W m ) / (W l +W m ); L=1-|L s -L0| / L0; M=MIN(L zm / 2×L bj ,1); Among them, Score2 represents the lighting condition score, L represents the brightness score, L s represents the actual brightness value, L0 represents the standard brightness, L zm represents the brightness of the frontal face area, L bj represents the background brightness, M represents the backlight score of the face, W l and W m respectively represent the brightness score weight and the backlight score weight of the face; S3: Analyze the head position, eye position, head orientation, head integrity, and eye opening degree in the eye movement video, calculate the head position score, head integrity score, head orientation score, and eye opening degree score respectively, and obtain the user recording process score; S4: Calculate the user's seriousness score according to the time when the user completes the eye movement; S5: Calculate the comprehensive video quality score. If the comprehensive video quality score is higher than the pre-set excellent score line, it is determined as a high-quality video. If the comprehensive video quality score is lower than the pre-set passing score line, it is determined as a low-quality video and the video is automatically screened out.
[0032] It should be noted that the calculation formula for the frame rate score is: F=(F s -F min ) / (F max -F min ); Among them, F represents the frame rate score, F s represents the actual frame rate, F min represents the lowest standard frame rate, which is the lowest frame rate threshold that the eye movement analysis system can accept, F max represents the highest standard frame rate, which is the frame rate threshold that the eye movement analysis system expects to reach perfectly; The calculation formula for the frame rate stability score is: S = (S0 - S min ) / (S max - S min ); Among them, S represents the frame rate stability score, S0 represents the frame rate stability, which is calculated according to the fluctuation of the video frame rate, and S min represents the lowest standard stability, which is the worst stability threshold that the eye movement analysis system can accept, and S max represents the highest standard stability, which is the perfect stability threshold that the eye movement analysis system expects to achieve; The calculation formula for the resolution score is as follows: R = (R s - R min ) / (R max - R min ); Among them, R represents the resolution score, and R s represents the actual resolution, and R min represents the lowest standard resolution, which is the lowest resolution threshold that the eye movement analysis system can accept, and R max represents the highest standard resolution, which is the perfect resolution threshold that the eye movement analysis system expects to achieve.
[0033] It should be noted that the Haar cascade classifier based on the Viola-Jones framework, the histogram of oriented gradients HOG, or the deep learning model is used to detect the face position, face key points, eye key points, and eye position in the eye movement video.
[0034] It should be noted that the calculation formulas for the user recording process score are as follows: Score3 = (A1 × W A1 + A2 × W A2 + A3 × W A3 + A4 × W A4 ) / (W A1 + W A2 + W A3 + W A4 ); A1 = 1 - |the proportion occupied by the head - the standard proportion| / the standard proportion; A2 = the actual head area occupied on the screen / the entire head area; A3 = 1 - |the actual head orientation - the correct orientation| / the correct orientation; A4 = (the actual opening degree - the lowest standard opening degree) / (the highest standard opening degree - the lowest standard opening degree); Among them, Score3 represents the score of the user's recording process. A1, A2, A3, and A4 respectively represent the head position score, head integrity score, head orientation score, and eye opening degree score. W A1 、W A2 、W A3 and W A4 respectively represent the head position score weight, head integrity score weight, head orientation score weight, and eye opening degree score weight; The proportion occupied by the head is the ratio of the number of pixels in the head area to the total number of pixels in the eye movement video.
[0035] It should be noted that the head integrity is obtained through the following method: Use the Haar cascade algorithm to initially screen the face position and verify the face position through the MTCNN model. Then, combined with the video resolution, adopt the dynamic anchor box scaling technology to adjust the scale of the detection window, select the largest one from the detected face positions as the head position, and convert the area of the head position into a rectangular bounding box as the head area; Extract the boundary of the head area, fit the actually occupied head area on the screen in combination with the geometric projection model, and dynamically correct the estimated value of the area of the actually occupied head area on the screen to obtain the final actually occupied head area on the screen; The head integrity is the ratio of the estimated value of the area of the actually occupied head area to the area of the entire head area.
[0036] It should be noted that introducing the MTCNN model for fine verification on the basis of initially screening with the Haar cascade algorithm solves the problem of false detection of a single algorithm in complex lighting or occlusion scenarios. Adopting the dynamic anchor box scaling technology to adjust the scale of the detection window can avoid missed detection caused by a fixed window; dynamically correcting the estimated value of the area of the actually occupied head area on the screen can reduce the error accumulation caused by posture changes in the traditional method and significantly improve the measurement stability in complex scenarios.
[0037] It should be noted that the actual head orientation determines the tip position according to the face key points, calculates the horizontal and vertical differences between the tip position coordinates and the center coordinates of the head area, and calculates the head orientation angle. The calculation formula is: θ = atan2(Δy, Δx) × 180 / π; Among them, θ represents the head orientation angle, Δy is the vertical difference between the tip position coordinates and the center coordinates of the head area, and Δx is the horizontal difference between the tip position coordinates and the center coordinates of the head area.
[0038] It should be noted that the positions of the eye key points are detected according to the position of the human eye. The eye key points include the inner corner, outer corner, upper eyelid and lower eyelid of the eye. The distance between the upper eyelid and the lower eyelid is used as the height of the eye, and the distance between the inner corner and the outer corner is used as the opening width of the eye. The actual opening degree of the eye is the ratio of the opening width to the eye height.
[0039] It should be noted that the calculation formula for the score of the user's earnestness is: Score4 = 1 - |D s - D0| / D0; where Score4 represents the score of the user's earnestness, D s represents the actual time taken by the user to complete the eye movement, and D0 represents the standard completion time, which is reasonably set according to the difficulty and complexity of the task; The calculation formula for the comprehensive score of the video quality is: Score = (Score1 × W1 + Score2 × W2 + Score3 × W3 + Score4 × W4) / (W1 + W2 + W3 + W4); where Score is the comprehensive score of the video quality, Score1 represents the evaluation score of the basic video information, Score2 represents the score of the light condition, Score3 represents the score of the user's recording process, Score4 represents the score of the user's earnestness, and W1, W2, W3, and W4 respectively represent the weights of the evaluation score of the basic video information, the score of the light condition, the score of the user's recording process, and the score of the user's earnestness.
[0040] In a specific embodiment, a three-stage cascaded network combines a classification loss function (used to distinguish faces and non-faces), a bounding box loss function (used to regress the position and size of the face bounding box), and a key point localization loss function (used to localize face key points) to train the MTCNN deep learning model, thereby realizing face detection; The formula for the total loss function of training the MTCNN deep learning model is: , where, is the classification loss function, is the bounding box regression loss, is the key point localization loss, and are weight coefficients used to balance different tasks; , where, is the true label ( represents a face, indicates non-face), is the face probability predicted by the model; ,
[0041] wherein, is the number of samples, are the predicted bounding box parameters, are the true bounding box parameters, and the SmoothL1 loss is defined as: , wherein, is the number of key points, are the predicted key point coordinates, are the true key point coordinates.
[0042] In a specific embodiment, the Histogram of Oriented Gradients (HOG) is used to detect faces. The local face gradient direction histogram of the image is extracted to calculate the gradient magnitude and direction of each pixel point of the image. When extracting image features, the image is divided into several cell units and the histogram of the gradient direction within each interval is statistically calculated. Finally, a support vector machine classifier is used to distinguish faces from non-faces, and negative sample mining is introduced to optimize the model and perform face detection; The formula for calculating the gradient magnitude of each pixel point of the image is: , The formula for calculating the gradient direction of each pixel point of the image is: , wherein, and are the gradients of the image in the x and y directions respectively, which can be calculated by methods such as the Sobel operator; The formula for statistically calculating the histogram of the gradient direction within each interval is: , wherein, p represents a single pixel point in the image, cell represents a local region unit, is the th histogram value of the direction interval, is the th direction interval, is the indicator function.
[0043] In a specific embodiment, a Haar cascade classifier based on the Viola-Jones framework is trained to extract a large number of Haar-like rectangular features, and the integral image is used to calculate the sum of pixels in each region. Finally, the AdaBoost algorithm is used to iteratively select key features, the weights of each weak classifier are corrected according to the error rate, and multiple weak classifiers are cascaded into a strong classifier for face detection; The calculation formula of the Haar-like rectangular feature is: , where Haar Feature represents the Haar-like rectangular feature; The calculation formula of the integral image is: , , where Integral(x, y) represents the sum of the grayscale values of all pixels within the rectangular region formed from the image origin (0,0) to the current point (x, y); Sum represents the total grayscale value of the pixels within any rectangular region calculated through the integral image; The weight update formula of the weak classifier is: , , where represents the grayscale value of pixel point , represents the grayscale value located at pixel point, is the th error rate of the weak classifier, is the weight of this classifier, is the th weak classifier, is the total number of weak classifiers, and H(x) represents the strong classifier formed by the weighted combination of multiple weak classifiers.
[0044] In a specific embodiment, the Haar cascade classifier is used to extract the local texture differences in the eye region, and rectangular Haar templates are designed for the eye features of the collected positive eye sample images and negative background sample images. For example, the upper and lower eyelids are captured according to the horizontal edge features, and the eye corners are captured according to the vertical features; Calculate the pixel difference of the human eye region using the integral image, screen the key eye Haar features of the human eye and the background through the AdaBoast algorithm, and construct a Haar cascade classifier. The Haar cascade classifier excludes non-eye regions based on the preliminary eye features, determines the pupil position using the key eye Haar features, and then locates the coordinates of the human eye position through a sliding window.
[0045] In a specific embodiment, identify the human eye region through the gradient direction distribution of the eye region, normalize the human eye region image to a fixed size of 32x32 pixels, calculate the gradient magnitude and direction of each pixel in the human eye region image, divide the human eye region image into 8x8 cells, count the gradient histograms of 9 directions in each cell of the human eye region image, perform L2-Hys normalization in blocks of 2x2 cells to generate HOG feature vectors, and input them into a linear vector machine to classify and train the positive human eye sample images and negative background sample images. Then, eliminate the overlapping detection frames through non-maximum suppression, and adjust the sliding window step size and scale to determine the human eye position.
[0046] In a specific embodiment, mark the binocular bounding boxes and class labels in the human eye region image, use the DarkNet-53 backbone network to extract the multi-scale feature maps of the human eye region image and preset anchor boxes for the feature maps. Determine the anchor box size according to the human eye size, and use the multi-task loss function, CloU Loss function, and Focal Loss function to train the deep learning model. The deep learning model detects the final human eye bounding box according to the set confidence level and performs threshold filtering and non-maximum suppression post-processing to determine the human eye position, and outputs the human eye position coordinates; Among them, the calculation formula of the CloU function is as follows: , Among them, CIoU Loss represents the CloU loss function, which is a function used to measure the difference between the model prediction value and the true value. IoU is the intersection over union of the predicted bounding box and the true bounding box, is the square of the Euclidean distance between the centers of the predicted bounding box and the true bounding box, is the square of the diagonal length of the smallest rectangle enclosing the two bounding boxes, is a measure of the consistency of the aspect ratios of the predicted bounding box and the true bounding box; , Among them, and h gt are the width and height of the true bounding box, and h pred are the width and height of the predicted bounding box, is a coefficient used to balance the influence of the center point distance and the aspect ratio; The calculation formula of the Focal Loss function is: , where L focal (p, y) represents the weighted cross-entropy loss value for a single sample, is the true label ( ), when it represents a positive sample, when it represents a negative sample, is the balance coefficient used to adjust the weights of positive and negative samples, usually taking the value of 0.25, is the adjustment factor used to control the focus intensity of the Focal Loss, usually taking the value of 2.
[0047] It should be noted that this application analyzes the physical characteristics and content of the eye movement video to more accurately and comprehensively evaluate the video quality, effectively solving the problem of uneven eye movement video quality in existing eye movement tracking methods. By automatically removing low-quality eye movement videos, the accuracy and stability of the eye movement tracking results can be ensured, and the efficiency and reliability of data collection can be improved.
[0048] Next, referring to Figure 2 , Figure 2 shows a detection system for the quality of eye movement videos, including a video basic information evaluation module a, a light condition evaluation module b, a recording process evaluation module c, a user attentiveness scoring module d, and a comprehensive evaluation module e.
[0049] In a specific embodiment, the video basic information evaluation module a is configured to: extract the frame rate and resolution of the eye movement video, analyze the basic information of the eye movement video for video basic information evaluation, and the evaluation formula is: Score1=(F×W f +R×W r +S×W s ) / (W f +W r +W s ); where Score1 represents the evaluation score of the video basic information, F, R, and S respectively represent the frame rate score, the resolution score, and the frame rate stability score, and W f , W r , W s respectively represent the frame rate score weight, the resolution score weight, and the frame rate stability score weight.
[0050] In a specific embodiment, the light condition evaluation module b is configured to: convert the eye movement video into a grayscale image, perform brightness analysis on the grayscale image to obtain the actual light value, and then perform numerical mapping to obtain the actual brightness value; Use the Haar cascade detector to detect the face position, divide the frontal face area by locating the key points of the face, and obtain the brightness of the frontal face area and the background brightness through an image processing algorithm; Evaluate the light condition based on the actual brightness value, the brightness of the frontal face area, and the background brightness. The calculation formula is: Score2=(L×W l +M×W m ) / (W l +W m ); L=1-|L s -L0| / L0; M=MIN(L zm / 2×L bj ,1); Wherein, Score2 represents the light condition score, L represents the brightness score, L s represents the actual brightness value, L0 represents the standard brightness, L zm represents the brightness of the frontal face area, L bj represents the background brightness, M represents the backlight score of the face, W l and W m represent the brightness score weight and the backlight score weight of the face respectively.
[0051] In a specific embodiment, the recording process evaluation module c is configured to: analyze the head position, eye position, head orientation, head integrity, and eye opening degree in the eye movement video, calculate the head position score, head integrity score, head orientation score, and eye opening degree score respectively, and obtain the user recording process score.
[0052] In a specific embodiment, the user seriousness scoring module d is configured to: calculate the user seriousness score according to the time taken by the user to complete the eye movement.
[0053] In a specific embodiment, the comprehensive evaluation module e is configured to: calculate the comprehensive video quality score. If the comprehensive video quality score is higher than a preset excellent score line, it is determined as a high-quality video. If the comprehensive video quality score is lower than a preset passing score line, it is determined as a low-quality video and the video is automatically screened out.
[0054] It should be noted that, compared with the existing technical methods that only focus on the physical characteristics of the video, the eye movement video quality detection method of the present application provides a more comprehensive evaluation. In addition to considering basic information such as the resolution and frame rate of the video, it also focuses on the evaluation factors inside the video content, such as the lighting conditions, the position and orientation of the face, etc., as well as the capture of eye information. The present application also takes into account the degree of attentiveness of the user when completing the task. By comprehensively evaluating these factors, the method of the present application can provide a more accurate, comprehensive and objective evaluation of the eye movement video quality, thereby improving the accuracy and reliability of the eye movement tracking results.
[0055] It should be noted that by adopting computer vision and image processing technologies, poor-quality eye movement videos can be quickly and accurately detected and automatically removed; compared with the traditional manual screening method, the automated operation of the present application saves a large amount of human resources, improves the efficiency and reliability of data collection, and effectively solves the problem of uneven quality of eye movement videos in the existing eye movement tracking methods. By automatically removing poor-quality eye movement videos, the accuracy and stability of the eye movement tracking results are ensured, which is very important for applications that rely on eye movement tracking data, such as user experience research, medical diagnosis, etc. The present application can provide high-quality eye movement tracking data.
[0056] The following refers to Figure 3 , which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. Figure 3 The shown electronic device is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0057] As Figure 3 shown, the computer system includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage section 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, ROM 302, and RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0058] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a liquid crystal display (LCD) and the like, as well as a speaker and the like; a storage section 308 including a hard disk and the like; and a communication section 309 including a network interface card such as a LAN card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as needed. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 310 as needed so that a computer program read therefrom is installed into the storage section 308 as needed.
[0059] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 309 and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, the above-described functions defined in the methods of the present application are performed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0060] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0061] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0062] The modules described in the embodiments of this application can be implemented in software or in hardware.
[0063] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: extract the frame rate and resolution of the eye movement video, analyze the basic information of the eye movement video and perform video basic information evaluation, convert the eye movement video into a grayscale image, perform brightness analysis on the grayscale image to obtain the actual light value and then perform numerical mapping to obtain the actual brightness value; use the Haar cascade detector to detect the face position, divide the frontal face area by locating the face key points, and obtain the brightness of the frontal face area and the background brightness through an image processing algorithm; evaluate the light conditions according to the actual brightness value, the brightness of the frontal face area and the background brightness, analyze the head position, eye position, head orientation, head integrity and eye opening degree in the eye movement video, and calculate the head position score, head integrity score, head orientation score and eye opening degree score respectively, and further calculate the user recording process score; calculate the user's seriousness score according to the time when the user completes the eye movement; calculate the comprehensive video quality score, and if the comprehensive video quality score is higher than the preset excellent score line, determine that the video is a high-quality video, and if the comprehensive video quality score is lower than the preset passing score line, determine that the video is a low-quality video and automatically screen out the video.
[0064] Finally, it should be noted that the above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solution formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. A method for detecting the quality of eye movement videos, characterized in that, Including: S1: Extract the frame rate and resolution of the eye movement video, analyze the basic information of the eye movement video for video basic information evaluation, and the evaluation formula is: Score1=(F×W f +R×W r +S×W s ) / (W f +W r +W s ); Among them, Score1 represents the evaluation score of the video basic information, and F, R, and S respectively represent the frame rate score, resolution score, and frame rate stability score, and W f 、W r 、W s respectively represent the frame rate score weight, resolution score weight, and frame rate stability score weight; S2: Convert the eye movement video into a grayscale image, perform brightness analysis on the grayscale image to obtain the actual light value, and then perform numerical mapping to obtain the actual brightness value; Use the Haar cascade detector to detect the face position, divide the frontal face area by locating the face key points, and obtain the brightness of the frontal face area and the background brightness through the image processing algorithm; Evaluate the light condition according to the actual brightness value, the brightness of the frontal face area and the background brightness, and the calculation formula is: Score2=(L×W l +M×W m ) / (W l +W m ); L = 1 - |L s - L0| / L0; M = MIN(L zm / 2 × L bj , 1); Among them, Score2 represents the light condition score, L represents the brightness score, L s represents the actual brightness value, L0 represents the standard brightness, L zm represents the brightness of the frontal face area, L bj represents the background brightness, M represents the face backlight score, W l and W m respectively represent the brightness score weight and the face backlight score weight; S3: Analyze the head position, eye position, head orientation, head integrity, and eye opening degree in the eye movement video, calculate the head position score, head integrity score, head orientation score, and eye opening degree score respectively, and obtain the user recording process score; S4: Calculate the user's attentiveness score according to the time when the user completes the eye movement; S5: Calculate the comprehensive video quality score. If the comprehensive video quality score is higher than the pre-set excellent score line, it is determined as a high-quality video. If the comprehensive video quality score is lower than the pre-set passing score line, it is determined as a low-quality video and the video is automatically screened out.
2. The method according to claim 1, wherein The calculation formula of the frame rate score is: F=(F s -F min ) / (F max -F min ); Among them, F represents the frame rate score, F s represents the actual frame rate, F min represents the minimum standard frame rate, which is the lowest frame rate threshold that the eye movement analysis system can accept, F max represents the maximum standard frame rate, which is the frame rate threshold that the eye movement analysis system expects to achieve perfectly; The calculation formula of the frame rate stability score is: S=(S0 - S min ) / (S max - S min ); Among them, S represents the frame rate stability score, and S0 represents the frame rate stability, which is calculated according to the fluctuation of the video frame rate. S min represents the lowest standard stability, which is the worst stability threshold that the eye movement analysis system can accept. S max represents the highest standard stability, which is the perfect stability threshold that the eye movement analysis system expects to achieve; The calculation formula of the resolution score is: R=(R s -R min ) / (R max -R min ); Among them, R represents the resolution score, R s represents the actual resolution, R min represents the lowest standard resolution, which is the lowest resolution threshold that the eye movement analysis system can accept, R max represents the highest standard resolution, which is the resolution threshold that the eye movement analysis system expects to achieve perfectly.
3. The method according to claim 1, characterized in that Use the Haar cascade classifier based on the Viola-Jones framework, Histogram of Oriented Gradients (HOG), or deep learning model to detect the face position, face key points, eye key points, and eye position in the eye movement video.
4. The method according to claim 1, wherein The calculation formulas of the user recording process score are respectively: Score3=(A1×W A1 +A2×W A2 +A3×W A3 +A4×W A4 ) / (W A1 +W A2 +W A3 +W A4 ); A1 = 1 - |the proportion of the head occupied - the standard proportion| / the standard proportion; A2 = the actual head area occupied on the screen / the entire head area; A3 = 1 - |the actual head orientation - the correct orientation| / the correct orientation; A4 = (the actual opening degree - the minimum standard opening degree) / (the maximum standard opening degree - the minimum standard opening degree); Among them, Score3 represents the score of the user's recording process, and A1, A2, A3, and A4 represent the head position score, head integrity score, head orientation score, and eye opening degree score respectively. W A1 , W A2 , W A3 , and W A4 represent the head position score weight, head integrity score weight, head orientation score weight, and eye opening degree score weight respectively; The proportion of the head occupied is the ratio of the number of pixels in the head area to the total number of pixels in the eye movement video.
5. The method according to claim 4, characterized in that The head integrity is obtained through the following method: Use the Haar cascade algorithm to preliminarily screen the face position and verify the face position through the MTCNN model. Then, combined with the video resolution, adopt the dynamic anchor box scaling technology to adjust the detection window scale, select the largest one from the detected face positions as the head position, and convert the area of the head position into a rectangular bounding box as the head area; Extract the boundary of the head area, combine the geometric projection model to fit the actual head area occupied on the screen, and perform dynamic correction on the estimated area value of the actual head area occupied on the screen to obtain the final actual head area occupied on the screen; The head integrity is the ratio of the estimated area value of the actual head area occupied to the area of the entire head area.
6. The method according to claim 4, characterized in that, The actual head orientation determines the position of the tip of the nose based on the facial key points, calculates the difference in the abscissa and the difference in the ordinate between the coordinates of the tip of the nose position and the center coordinates of the head region, and calculates the head orientation angle. The calculation formula is as follows: θ = atan2(Δy, Δx) × 180 / π; where θ represents the head orientation angle, Δy is the difference in the ordinate between the coordinates of the tip of the nose position and the center coordinates of the head region, and Δx is the difference in the abscissa between the coordinates of the tip of the nose position and the center coordinates of the head region.
7. The method according to claim 4, wherein Detect the positions of the eye key points according to the positions of the eyes. The eye key points include the inner corner, outer corner, upper eyelid, and lower eyelid of the eyes. Take the distance between the upper eyelid and the lower eyelid as the height of the eyes, and take the distance between the inner corner and the outer corner as the opening width of the eyes. The actual opening degree is the ratio of the opening width to the eye height.
8. The method according to claim 1, wherein The calculation formula for the score of the user's attentiveness is as follows: Score4 = 1 - |D s - D0| / D0; Among them, Score4 represents the user's conscientiousness score, and D s represents the time actually spent by the user to complete the eye movement, and D0 represents the standard completion time, which is reasonably set according to the difficulty and complexity of the task. The calculation formula for the comprehensive score of the video quality is as follows: Score = (Score1 × W1 + Score2 × W2 + Score3 × W3 + Score4 × W4) / (W1 + W2 + W3 + W4); where Score is the comprehensive score of the video quality, Score1 represents the evaluation score of the basic information of the video, Score2 represents the score of the light condition, Score3 represents the score of the user's recording process, Score4 represents the score of the user's attentiveness, and W1, W2, W3, and W4 respectively represent the weights of the evaluation score of the basic information of the video, the weight of the light condition score, the weight of the user's recording process score, and the weight of the user's attentiveness score.
9. A detection system for the quality of eye movement videos, characterized in that, including: Video basic information evaluation module: Extract the frame rate and resolution of the eye movement video, analyze the basic information of the eye movement video for video basic information evaluation. The evaluation formula is as follows: Score1=(F×W f +R×W r +S×W s ) / (W f +W r +W s ); Among them, Score1 represents the evaluation score of the video basic information. F, R, and S respectively represent the frame rate score, resolution score, and frame rate stability score. W f 、W r 、W s respectively represent the frame rate score weight, resolution score weight, and frame rate stability score weight; Light condition evaluation module: Convert the eye movement video into a grayscale image, perform brightness analysis on the grayscale image to obtain the actual light value, and then perform numerical mapping to obtain the actual brightness value; Use the Haar cascade detector to detect the face position, divide the frontal face region by locating the facial key points, and obtain the brightness of the frontal face region and the background brightness through the image processing algorithm; Evaluate the light condition according to the actual brightness value, the brightness of the frontal face region, and the background brightness. The calculation formula is as follows: Score2=(L×W l +M×W m ) / (W l +W m ); L = 1 - |L s - L0| / L0; M = MIN(L zm / 2 × L bj , 1); Among them, Score2 represents the light condition score, L represents the brightness score, L s represents the actual brightness value, L0 represents the standard brightness, L zm represents the brightness of the frontal face area, L bj represents the background brightness, M represents the face backlight score, W l and W m respectively represent the brightness score weight and the face backlight score weight; Recording process evaluation module: Analyze the head position, eye position, head orientation, head integrity, and eye opening degree in the eye movement video, calculate the head position score, head integrity score, head orientation score, and eye opening degree score respectively, and obtain the user's recording process score; User attentiveness scoring module: Calculate the score of the user's attentiveness according to the time taken by the user to complete the eye movement; Comprehensive evaluation module: Calculate the comprehensive score of the video quality. If the comprehensive score of the video quality is higher than the pre-set excellent score line, it is determined as a high-quality video. If the comprehensive score of the video quality is lower than the pre-set passing score line, it is determined as a low-quality video and the video is automatically screened out.
10. A computer program product having one or more computer programs thereon, characterized in that, When the computer program product is executed by a computer processor, the method described in any one of claims 1-8 is implemented.
Citation Information
Patent Citations
Multi-source information fusion driving safety early warning method based on fuzzy evaluation
CN117274959A
Face recognition method and face recognition system
CN117409451A
A method for suppressing reflection in infrared imaging glasses based on image processing
CN119741684A
Electronic eye protector
CN207473227U
A dry eye syndrome prevention method and its performance model
IN201721043982A