All-day monitoring sight line estimation method
Through the light intensity detection and dynamic adjustment of the line of sight estimation model, combined with the LowLight and NormalLight models, the problem of line of sight estimation under unsatisfactory lighting conditions is solved, stable and accurate line of sight estimation under different lighting environments is achieved, and the accuracy of the monitoring system is improved.
Patent Information
- Application Number
- CN202510390587.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
Existing line of sight estimation systems are difficult to maintain high accuracy when lighting conditions are not ideal, especially in low-light environments, where the stability and accuracy of line of sight estimation are insufficient.
The light intensity detection algorithm is used for brightness analysis, the line of sight estimation model is dynamically adjusted, the face detection is performed using the MTCNN network, and the LowLight and NormalLight line of sight estimation models are combined to automatically switch under different lighting environments. Facial features are extracted through the ResNet18 network, and ELA attention mechanism and multi-layer perceptron are introduced for line of sight estimation.
Ensure the accuracy and stability of line of sight estimation under different lighting environments, improve the accuracy and practicality of the monitoring system, and adapt to the needs of line of sight estimation under various lighting conditions.
Smart Images

Figure CN120340088A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and artificial intelligence, and in particular, to a method for estimating line of sight for all-day monitoring. Background Art
[0002] With the increasing demand for behavior monitoring systems in the modern education field, monitoring technologies based on computer vision and artificial intelligence have gradually become the mainstream. In the prior art, many systems rely on facial recognition or line-of-sight estimation of cameras to judge people's attention and behavior. However, most existing line-of-sight estimation systems have certain limitations. Especially in the case of unsatisfactory lighting conditions, it is often difficult to maintain high-precision line-of-sight estimation. For example: in the intelligent observation system for learning behaviors in the classroom, the camera captures and analyzes the line of sight of students to form a classroom behavior observation report. The lighting conditions in the student area of the classroom are relatively low, which poses a technical difficulty. However, the prior art often lacks sufficient flexibility and does not fully consider the impact of lighting changes on the accuracy of the model. Therefore, how to ensure the stability and accuracy of line-of-sight estimation under low-light and normal-light conditions remains a major problem in the current technology. Summary of the Invention
[0003] Object of the Invention: In order to overcome the deficiencies in the prior art, the present invention provides a method for estimating line of sight for all-day monitoring, which can automatically adjust the line-of-sight estimation model in different lighting environments, obtain the line-of-sight estimation result, ensure the accuracy and stability of line-of-sight estimation, and improve the accuracy and practicality of the monitoring system.
[0004] Technical Solution: To achieve the above object, a method for estimating line of sight for all-day monitoring of the present invention includes: real-time acquiring video stream data of a monitoring area, using a light intensity detection algorithm to perform quantitative analysis on the brightness information of the acquired frame image to obtain the result of quantitative analysis; using the MTCNN network to perform face detection and localization on the acquired frame image to obtain a face image, and comparing the light intensity of the result of quantitative analysis with a preset threshold.
[0005] When the light intensity is lower than the preset threshold, enable the LowLight line-of-sight estimation model for low-light compensation, input the face image into the LowLight line-of-sight estimation model for low-light compensation to obtain the line-of-sight estimation result.
[0006] When the light intensity is not lower than the preset threshold, enable the NormalLight line-of-sight estimation model for standard lighting, input the face image into the NormalLight line-of-sight estimation model for standard lighting to obtain the line-of-sight estimation result.
[0007] The gaze estimation result is mapped through a three-dimensional eyeball coordinate system to generate gaze direction data, and the attention concentration of the gaze is evaluated; the real-time processing pipeline dynamically analyzes the video stream data, dynamically presents the evaluation result through a human-computer interaction interface, and provides storage and structured export of visual data.
[0008] Further, the quantitative analysis of the brightness information of the acquired frame image by using the light intensity detection algorithm includes image brightness calculation, local light intensity analysis, global light intensity analysis, time series light intensity analysis, and multi-channel light intensity fusion analysis;
[0009] For the image brightness calculation, for a color image, the brightness value L of the picture is converted from the RGB color space to the brightness component by weighted average; the calculation process is as follows:
[0010] L = 0.299×R + 0.587×G + 0.114×B
[0011] In the formula, L is the brightness value of this pixel point, and R, G, and B are the red, green, and blue components of this pixel point respectively;
[0012] For the local light intensity analysis, the image is divided into several sub-regions, and the local light intensity of each sub-region; the local light intensity L region The calculation process is as follows:
[0013]
[0014] In the formula, L(i, j) is the brightness value of the (i, j)th pixel, and m and n are the width and height of the sub-region respectively;
[0015] For the global light intensity analysis, the global light intensity of the image is calculated according to the local brightness of all sub-regions; the global light intensity L global The calculation process is as follows:
[0016]
[0017] In the formula, M is the total number of sub-regions, and L region (k) is the brightness value of the kth sub-region.
[0018] Further, for the time series light intensity analysis, time series analysis technology is used to track the change trend of light. The brightness value L avg (t) of each frame image is calculated in k consecutive frame images; the change trend of light in the time series, that is, the light change amount ΔL avg (t) is expressed as:
[0019] ΔL avg (t) = L avg (t) - Lavg ((t - 1)
[0020] The multi-channel light intensity fusion analysis introduces a multi-channel light intensity analysis method and combines the brightness data of RGB images, infrared images, and ultraviolet images; the global light intensity L global is used as the brightness data of the RGB image, and the comprehensive light intensity calculation process is as follows:
[0021]
[0022] In the formula, L RGB is the light intensity of the RGB image, L IR and L UV are the light intensities of the infrared and ultraviolet channels respectively, and W1, W2, and W3 are the weight coefficients of the RGB image, infrared, and ultraviolet channels respectively, is the multiplication operation.
[0023] Furthermore, the light intensity L total of the result of the quantitative analysis is compared with the preset threshold L threshold ;
[0024] When the light intensity is lower than the preset threshold L total < L threshold , it is determined that the image brightness is a low-light image, and the low-light compensation LowLight gaze estimation model is enabled;
[0025] When the light intensity is not lower than the preset threshold L total ≥ L threshold , it is determined that the image brightness is a standard-light image, and the standard-light NormalLight gaze estimation model is enabled;
[0026] According to the real-time light data and the historical light change trend, the system dynamically adjusts the preset threshold L threshold ; the calculation formula is as follows:
[0027] L threshold (t) = L threshold (t - 1) + αΔL avg (t)
[0028] In the formula, α is the dynamic adjustment coefficient, αΔL avg (t) is the light change amount, and L threshold (t) is the preset threshold at time t.
[0029] Further, for the standard illumination NormalLight gaze estimation model, feature extraction is performed on the input face image through the ResNet18 network of the deep convolutional neural network model to obtain a high-dimensional facial feature vector; the extracted high-dimensional facial feature vector is processed through two fully connected layers to output a three-dimensional feature vector; the three-dimensional feature vector is converted into the angular error of gaze estimation through a tangent transformation; the mean squared error loss function MSELoss is used as the loss function of the model; MSELoss measures the error by calculating the expected value of the squared difference between the predicted value and the true value; the mean squared error loss function MSELoss is defined as follows:
[0030]
[0031] In the formula, ξ gt is the true gaze value, ξ pred is the predicted value of gaze estimation, and n is the number of samples.
[0032] Further, for the low illumination compensation LowLight gaze estimation model, the face image sequentially passes through basic convolutional layers, batch normalization BN, activation functions, and max pooling MaxPool operations; then it is input into the residual network and processed through the residual modules in the first stage, second stage, third stage, and fourth stage to extract global features; then it is processed through the SSD module to finally form facial features, where the residual modules in the second stage and third stage are improved residual modules with the ELA attention mechanism module introduced; both the left eye image and the right eye image in the face image sequentially pass through basic convolutional layers, batch normalization BN, activation functions, and max pooling MaxPool operations, and then are sequentially processed through three improved residual modules with the ELA attention mechanism module introduced to extract the detailed features of the eyes; then local features are extracted through adaptive average pooling and an additional convolutional module to finally form left eye features and right eye features; the facial features, left eye features, and right eye features are equally fused through a multi-layer perceptron MLP, and then pass through two fully connected layers in sequence to obtain the final gaze estimation result.
[0033] Further, the improved residual module with the ELA attention mechanism module includes two improved residual sub-modules; the output end of the first improved residual sub-module is connected to the input end of the second improved residual sub-module in a transmission manner; the improved residual sub-module includes a residual block and an ELA attention mechanism module; the residual block processes the input of the improved residual sub-module to obtain the output of the residual block, and the output of the residual block is processed through the ELA attention mechanism module to obtain the output of the ELA attention mechanism module. The output obtained by multiplying and fusing the output of the residual block and the output of the ELA attention mechanism module is used as the output of the improved residual sub-module.
[0034] Furthermore, in the ELA attention mechanism module, strip pooling is adopted in the spatial dimension to extract feature vectors in the horizontal and vertical directions; the input feature map is X h ∈R B×C×H×W ; taking the average of the width dimension, a feature map with the size of B×C×H×1 is obtained Then the feature map passes through a 1D convolution, group normalization, and Sigmoid activation function in sequence to obtain a feature map X with the size of B×C×H×1 h ; taking the average of the height dimension, a feature map with the size of B×C×1×W is obtained Then the feature map passes through a 1D convolution, group normalization, and Sigmoid activation function in sequence to obtain a feature map X with the size of B×C×1×W w ;
[0035] Finally, the product of the input feature map X and the two attention feature maps X h and X w obtains the spatial attention feature map Y; the calculation process is as follows:
[0036]
[0037] In the formula, represents element-wise multiplication.
[0038] Furthermore, the multi-layer perceptron includes an input layer, two hidden layers, a batch normalization layer, a ReLU activation function, a Dropout layer, and an output layer; the final output feature obtained by comprehensively processing the multi-stream input features is used as the gaze estimation result;
[0039] ξ pred = MLP(F eye , F face )
[0040] In the formula, F eye is the eye feature, including the left eye feature and the right eye feature; F face is the face feature; MLP is the multi-layer perceptron operation.
[0041] Beneficial effects: The gaze estimation method for all-day monitoring of the present invention can automatically adjust the gaze estimation model in different lighting environments to ensure the accuracy and stability of gaze estimation, solve the problem of unstable accuracy in low-light and normal-lighting environments in the prior art, and improve the accuracy and practicality of the monitoring system; by fusing the improved residual network and the ELA attention mechanism, a gaze estimation model that can make full use of the long-distance dependence relationship and key region features in eye and face images is constructed, and this model can be flexibly selected according to the set threshold size to adapt to the gaze estimation requirements in different lighting environments. Description of the Drawings
[0042] Figure 1 It is a flowchart of the line-of-sight estimation method for all-day monitoring;
[0043] Figure 2 It is a network structure diagram of the normal light line-of-sight estimation model;
[0044] Figure 3 It is a diagram of the dimensional change of the normal light line-of-sight estimation model;
[0045] Figure 4 It is a network structure diagram of the low light compensation line-of-sight estimation model;
[0046] Figure 5 It is a schematic structural diagram of the improved residual module introducing the ELA attention mechanism module;
[0047] Figure 6 It is a schematic structural diagram of the ELA attention mechanism module;
[0048] Figure 7 It is a schematic diagram of the dimensional change of the SSD module;
[0049] Figure 8 It is a schematic structural diagram of the line-of-sight estimation system for all-day monitoring. Detailed Implementation Manner
[0050] The present invention will be further described below with reference to the accompanying drawings.
[0051] As Figure 1 shown, a line-of-sight estimation method for all-day monitoring includes: acquiring video stream data of a monitoring area in real time, and quantifying and analyzing the brightness information of the acquired frame image in real time by using a light intensity detection algorithm to obtain the result of the quantification analysis; using the MTCNN network to perform face detection and positioning on the acquired frame image, obtaining a face image by detecting the face area in the image and providing the facial position and bounding box, and providing a positioning basis for subsequent facial feature point extraction; comparing the light intensity of the result of the quantification analysis with a preset threshold,
[0052] when the light intensity is lower than the preset threshold, enabling the low light compensation line-of-sight estimation model, inputting the face image into the low light compensation line-of-sight estimation model to obtain the line-of-sight estimation result;
[0053] when the light intensity is not lower than the preset threshold, enabling the normal light line-of-sight estimation model, inputting the face image into the normal light line-of-sight estimation model to obtain the line-of-sight estimation result;
[0054] The gaze estimation result is mapped through a three-dimensional eyeball coordinate system to generate gaze direction data, and the attention concentration degree of the gaze is evaluated; the real-time processing pipeline dynamically analyzes the video stream data, dynamically presents the evaluation result through a human-computer interaction interface, and provides the functions of storing and structurally exporting the visual data.
[0055] The MTCNN network is a multi-task cascaded network based on deep learning, which can accurately locate faces at different scales and calculate the bounding boxes of faces. This method is applicable to face detection under various lighting and pose conditions, and is especially suitable for applications in dynamic scenarios. In the system, both the LowLight gaze estimation model and the NormalLight gaze estimation model for low-light compensation use the ResNet18 residual network for feature extraction, accurately obtain facial features, and provide data support for subsequent gaze estimation. According to the output, which is the gaze estimation angle error, that is, the gaze estimation result. When the gaze estimation angle error exceeds the threshold (30°) and this error exceeds 10 times, the system will determine that the user's attention is dispersed. Through the above steps, the system can monitor and accurately evaluate the user's attention state in real time, and determine the dispersion of attention when the gaze error exceeds the set threshold.
[0056] As Figure 8 shown, it includes an all-weather gaze estimation system; the all-weather gaze estimation system includes a user login and management module, a system management module, a real-time monitoring module, and a historical data and export module; the user login and management module is used for user login to the system, password modification, and administrator user maintenance; the system management module is used for the management of cameras and the setting management of gaze thresholds; the system management module is also used to process each frame of the video image for gaze estimation to obtain the gaze estimation result; the real-time monitoring module is used for viewing the real-time monitoring screen and viewing the real-time data of the scene; the historical data and export module is used for viewing and exporting historical data. The system processes camera data in real time, evaluates the user's gaze and behavior; the results can be viewed through the interface, and the save and export functions are supported.
[0057] The all-weather monitoring gaze estimation method is based on the Python programming language and the PyTorch framework; Python is also selected as the programming language for the system implementation part; the core algorithm is developed through the OpenCV image processing library and the Python language; to ensure the efficiency and scalability of the system, the backend service uses the Flask framework, which is a lightweight Web framework written in Python; Flask is lightweight and easy to use, and has powerful Web service capabilities, suitable for rapid development and deployment; in terms of data management and query, the system uses the MySQL database, which ensures the reliability of data operations with its stability and maturity; the overall structural design diagram of the system is as Figure 8 shown.
[0058] The quantization analysis of the brightness information of the acquired frame images using the light intensity detection algorithm includes image brightness calculation, local illumination intensity analysis, global illumination intensity analysis, time series illumination intensity analysis, and multi-channel illumination intensity fusion analysis; the brightness value of the image reflects the overall illumination intensity of the image. By calculating the brightness of each frame of the image, the illumination conditions of the current environment can be estimated, which is the image brightness calculation; for color images, in the image brightness calculation, the brightness value L of the picture is converted from the RGB color space to the brightness component by weighted average; the calculation process is as follows:
[0059] L = 0.299×R + 0.587×G + 0.114×B
[0060] In the formula, L is the brightness value of the pixel point, and R, G, and B are the red, green, and blue components of the pixel point respectively;
[0061] In practical applications, the brightness of the overall image may not fully reflect the illumination differences in local areas. Especially in dynamic scenes, it is necessary to calculate the local illumination intensity of the image; in the local illumination intensity analysis, the image is divided into several sub-regions, each sub-region is a sub-image of m×n, and the local illumination intensity of each sub-region; the local illumination intensity L region The calculation process is as follows:
[0062]
[0063] In the formula, L(i, j) is the brightness value of the (i, j)th pixel, and m and n are the width and height of the sub-region respectively; by calculating the average brightness of each sub-region, the system can more finely evaluate the illumination intensity of different regions.
[0064] To further measure the fluctuation of the regional brightness, calculate the variance of the brightness within a certain region to evaluate the degree of local illumination change. The calculation process is as follows:
[0065]
[0066] In the formula, σ 2 represents the illumination intensity variance within the region, and L region is the average brightness value of the region, that is, the illumination intensity; the larger the variance value, the greater the illumination difference within the region; the regional brightness and variance comprehensively reflect the local illumination difference, which helps to detect the uneven illumination regions.
[0067] In the global illumination intensity analysis, the global illumination intensity of the image is calculated according to the local brightness of all sub-regions; the global illumination intensity L global The calculation process is as follows:
[0068]
[0069] where M is the total number of sub-regions, and L region (k) is the luminance value of the k-th sub-region, which is the illumination intensity of the k-th sub-region; this formula obtains the global luminance value of the entire image by weighted averaging the luminances of multiple regions, providing a basis for subsequent illumination mode determination.
[0070] To measure the luminance contrast of the image, the system also calculates the contrast value of the image:
[0071]
[0072] where L max is the maximum luminance in the image, and L min is the minimum luminance in the image. The contrast value can help the system determine whether the illumination is uniform and further improve the evaluation of the illumination conditions.
[0073] For the time-series illumination intensity analysis, since the illumination conditions may change over time, time-series analysis techniques are used to track the change trend of illumination. The luminance value L avg (t) of each frame of image is calculated in consecutive k frames of images, which can represent the illumination intensity of each frame of image; the change trend of illumination in the time series, that is, the illumination change amount ΔL avg (t) is expressed as:
[0074] ΔL avg (t) = L avg (t) - L avg (t - 1)
[0075] In this way, the system can capture the illumination fluctuations in a timely manner, especially in the case of sudden illumination changes, improving the accuracy of illumination estimation.
[0076] To smooth short-term fluctuations, the system also uses a moving average method to calculate the smoothed value of illumination:
[0077]
[0078] where L smooth (t) represents the smoothed illumination value at time t, and n is the size of the moving window; L avg (i) is the luminance value of the image at time i; this method effectively eliminates small short-term fluctuations and helps the system evaluate illumination changes more stably.
[0079] For the multi-channel illumination intensity fusion analysis, a multi-channel illumination intensity analysis method is introduced, combining the luminance data of RGB images, infrared images, and ultraviolet images; the global illumination intensity L global is used as the luminance data of the RGB image, and the comprehensive light intensity calculation process is as follows:
[0080]
[0081] Wherein, L RGB is the illumination intensity of the RGB image, and L IR and L UV are the illumination intensities of the infrared and ultraviolet channels respectively. For the fusion of multi-channel illumination intensities, in most cases, the weights of infrared L IR and ultraviolet L UV are generally 0; W1, W2, and W3 are the weight coefficients of the RGB image, infrared, and ultraviolet channels respectively, used to adjust the contributions of different channels. In most cases, W2 and W3 are 0; is the multiplication operation. By fusing the data of multiple illumination channels, the system can provide more accurate illumination assessment in complex environments such as at night and against the light.
[0082] Based on the analysis result of the illumination intensity, the system compares the current illumination state with a preset threshold to determine the current illumination mode and select an appropriate line-of-sight estimation model; the light intensity of the result of the quantitative analysis, that is, the comprehensive light intensity L total is compared with the preset threshold L threshold ;
[0083] When the light intensity is lower than the preset threshold L total < L threshold , it is determined that the image brightness is a low-illumination image, and the LowLight line-of-sight estimation model for low-illumination compensation is enabled;
[0084] When the light intensity is not lower than the preset threshold L total ≥ L threshold , it is determined that the image brightness is a standard-illumination image, and the NormalLight line-of-sight estimation model for standard illumination is enabled;
[0085] To cope with the rapid change of environmental illumination, the system introduces an adaptive adjustment mechanism; and according to the real-time illumination data and the historical illumination change trend, the system dynamically adjusts the preset threshold L threshold ; through this mechanism, the system can automatically adapt to the change of illumination and ensure accurate line-of-sight estimation under various environmental conditions. The calculation formula is as follows:
[0086] L threshold (t) = L threshold (t - 1) + αΔL avg (t)
[0087] Wherein, α is the dynamic adjustment coefficient, and αΔL avg (t) is the illumination change amount, and L threshold (t) is the preset threshold at time t.
[0088] As Figures 2-3 shown, for the standard illumination NormalLight gaze estimation model, feature extraction is performed on the input face image through the ResNet18 network of the deep convolutional neural network model to obtain a high-dimensional facial feature vector; the face image of each frame is detected through the deep convolutional neural network model, and the system extracts features from the face region through the ResNet18 network; the system extracts features from the input face image through the ResNet18 network to obtain a feature vector; and a linear transformation layer is used to perform dimensionality reduction processing on the obtained feature vector. The calculation process is as follows:
[0089] Z = W4·X + b1
[0090] In the formula, Z is the feature vector after dimensionality reduction processing, X is the feature vector extracted from the ResNet18 network, and W4 and b1 are the weight matrix and bias vector of the first linear transformation layer respectively; these two parameters are learned through an optimization algorithm during the training process. The feature vector Z after dimensionality reduction processing is further processed through a fully connected layer and then undergoes a non-linear transformation to obtain a hidden state H;
[0091] H = σ(W5·Z + b2)
[0092] In the formula, W5 and b2 are the weight matrix and bias vector of the fully connected layer respectively, and σ is the Sigmoid activation function; the fully connected layer minimizes the loss function of the network by optimizing W5 and b2, and the finally generated H is the hidden state after non-linearly transforming the image features; this hidden state provides a deep feature representation to help the subsequent gaze estimation module further analyze the user's attention state; the hidden state converts the original features into an abstract representation that is more conducive to task processing, extracts high-level features, and finally the ResNet18 network extracts and obtains the output high-dimensional facial feature vector.
[0093] As Figures 2-3 shown, the extracted high-dimensional facial feature vector is processed through two fully connected layers to output a three-dimensional feature vector; the three-dimensional feature vector is converted into the angular error of gaze estimation through a tangent transformation; the mean squared error loss function MSELoss is used as the loss function of the model; MSELoss measures the error by calculating the expected value of the squared difference between the predicted value and the true value; the smaller MSELoss is, the smaller the prediction error of the model and the better the performance of the model; the mean squared error loss function MSELoss is defined as follows:
[0094]
[0095] In the formula, ξ gt is the true value of the gaze, ξ predis the predicted value of gaze estimation, and n is the number of samples.
[0096] As Figure 4 shown, in the LowLight gaze estimation model for low-light compensation, the face image sequentially passes through basic convolutional layers, batch normalization (BN), activation functions, and max pooling operations; then it is input into the residual network and processed by residual modules in the first stage, second stage, third stage, and fourth stage to extract global features; then it is processed by the SSD module to finally form facial features, where the residual modules in the second stage and third stage are improved residual modules with the ELA attention mechanism module introduced; both the left-eye image and the right-eye image in the face image pass through basic convolutional layers, batch normalization (BN), activation functions, and max pooling operations, and then are sequentially processed by three improved residual modules with the ELA attention mechanism module introduced to extract the detailed features of the eyes; then local features are extracted through adaptive average pooling and an additional convolutional module, and finally left-eye features and right-eye features are formed; the facial features, left-eye features, and right-eye features are equally fused through a multi-layer perceptron (MLP), and then pass through two fully connected layers in sequence to obtain the final gaze estimation result, and finally the angular error of gaze estimation is obtained. The loss function of the LowLight gaze estimation model for low-light compensation also uses the mean squared error loss function (MSELoss).
[0097] As Figure 5 shown, the network structures of the three improved residual modules with the ELA attention mechanism module introduced and the residual modules in the second stage and third stage are the same; the improved residual module with the ELA attention mechanism module includes two improved residual sub-modules; the output end of the first improved residual sub-module is connected to the input end of the second improved residual sub-module in a transmission manner, the input end of the first improved residual sub-module serves as the input end of the improved residual module, and the output end of the second improved residual sub-module serves as the output end of the improved residual module; both of the two improved residual sub-modules include a residual block and an ELA attention mechanism module; the residual block processes the input of the improved residual sub-module to obtain the output of the residual block, the output of the residual block is processed by the ELA attention mechanism module to obtain the output of the ELA attention mechanism module, and the output obtained by multiplying and fusing the output of the residual block and the output of the ELA attention mechanism module serves as the output of the improved residual sub-module; the ELA attention mechanism is introduced after each residual block, and the output of the residual block is processed by the ELA attention mechanism and then passed to the next layer to improve the effectiveness of feature extraction. The output of the first improved residual sub-module serves as the input of the second improved residual sub-module. Each residual block contains two convolutional blocks; the input end of the first convolutional block serves as the input end of the residual block, the output of the first convolutional block serves as the input of the second convolutional block, and the output of the second convolutional block is added and fused with the input of the residual block to obtain the output of the residual block.
[0098] As shown Figures 4-5 in the figure, the LowLight gaze estimation model is divided into two parts. One is the extraction of facial features, and the other is the extraction of left and right eye features. For feature extraction, the ResNet18 network is also used to extract features, but the residual modules in the ResNet18 network are improved. The ResNet18 network in the extraction of left and right eye features contains three improved residual modules, namely Res2, Res3, and Res4. In Res2, the network contains two residual blocks. The first residual block increases the number of channels from 64 to 128 through a convolutional layer with a stride of 2 to achieve downsampling. The second residual block keeps the number of channels at 128 and the stride at 1 without downsampling. In Res3, there are also two residual blocks. The first residual block increases the number of channels from 128 to 256 and uses a convolutional layer with a stride of 2 for downsampling. The second residual block keeps the number of channels at 256 and the stride at 1. In Res4, there are also two residual blocks, and the last two residual blocks of the network further increase the number of channels. The first residual block increases the number of channels from 256 to 512 while performing downsampling. The second residual block keeps the number of channels at 512 and the stride at 1. The designs of the residual modules Res2, Res3, and Res4 are as shown Figure 5 in the figure.
[0099] As shown Figure 6 in the figure, in the ELA attention mechanism module, strip pooling is used in the spatial dimension to extract feature vectors in the horizontal and vertical directions instead of global spatial pooling to maintain an elongated kernel shape to capture long-range dependencies. The input feature map is X h ∈R B×C×H×W , where B is the batch size of the features, C is the number of feature channels, H is the height dimension of the features, and W is the width dimension of the features. The average is taken over the width dimension to obtain a feature map of size B×C×H×1 Then the feature map passes through a 1D convolution, group normalization, and Sigmoid activation function in sequence to obtain a feature map of size B×C×H×1, denoted as X h ;
[0100]
[0101] In the formula, σ is the Sigmoid activation function; X[:, :, :, j] will select a sub-tensor of shape (B, C, H), and only the i-th element is selected in the W dimension. GN represents group normalization, with each group containing 16 channels. Conv1d represents a 1D convolution operation with a kernel size of 7 to capture long-range dependencies.
[0102] The average is taken over the height dimension to obtain a feature map of size B×C×1×W Then the feature map successively passes through a 1D convolution, group normalization, and a Sigmoid activation function to obtain a feature map X of size B×C×1×W w ;
[0103]
[0104] where σ is the Sigmoid activation function; GN represents group normalization; Conv1d represents the 1D convolution operation; X[:, :, :, i] will select a sub-tensor of shape (B, C, W), and only the j-th element is selected in the H dimension.
[0105] Finally, the product of the input feature map X and two attention feature maps X h and X w results in a spatial attention feature map Y, which is the output of the ELA attention mechanism module; the calculation process is as follows:
[0106]
[0107] where denotes element-wise multiplication.
[0108] As Figure 7 shown, the SSD module includes a convolutional layer (with 512 input channels and 64 output channels), followed by a max pooling layer, a batch normalization layer, and an activation function; then, it goes through two convolutional layers (both with 64 channels), and each convolutional layer is followed by batch normalization and an activation function; no pooling is performed after the first convolutional layer, and max pooling is performed after the second convolutional layer. The processed features are flattened into a one-dimensional vector (64-dimensional), passed through a fully connected layer (outputting 256 dimensions), and then through a batch normalization layer, a ReLU activation function, and a Dropout layer. Finally, facial features are extracted through the fully connected layer for subsequent regression or classification tasks. The dimensionality changes in the SSD module are as Figure 7 shown.
[0109] The multi-layer perceptron described above includes an input layer, two hidden layers, a batch normalization layer, a ReLU activation function, a Dropout layer, and an output layer; the final output feature obtained through the comprehensive processing of the multi-stream input features is used as the gaze estimation result;
[0110] ξ pred = MLP(F eyw , F face )
[0111] where F eye is the eye feature, including the left eye feature and the right eye feature; F faceare facial features; MLP is the operation of a multi-layer perceptron.
[0112] Example 1
[0113] Before comparing the light intensity L of the result of quantitative analysis total with the preset threshold L threshold First, it is judged whether the contrast value of the image is equal to or lower than the set contrast threshold. At the same time, the variances of all local regions are combined into a sequence where z is the image divided into z sub-regions, and the sequence composed of the variances of the light intensities in several regions Each sequence in is compared with the set variance threshold; only when the contrast value of the image is equal to or lower than the set contrast threshold and the sequence composed of the variances of the light intensities in several regions has less than 5 variances equal to or lower than the set variance threshold, can the image be passed by comparing the light intensity L total of the result of quantitative analysis with the preset threshold L threshold to make a judgment on the LowLight line-of-sight estimation model or the NormalLight line-of-sight estimation model;
[0114] When the contrast value of the image is higher than the set contrast threshold and the sequence composed of the variances of the light intensities in several regions has 5 or more variances higher than the set variance threshold, the picture is directly input into the LowLight line-of-sight estimation model to estimate the line of sight in the face in the image.
[0115] Example 2
[0116] Select the experimental environment; hardware environment, NVIDIA GeForce RTX 3090 GPU graphics card; software environment, developed based on the Windows 11 system; development environment, Pycharm development platform, Python programming language, Pytorch framework.
[0117] This experiment is set to use the MPIIFaceGaze dataset and the Gaze360 dataset for training and testing; mainly tested the NormalLight model under normal lighting mode. Training settings, for the two datasets, the training settings are the same. The model was trained for 60 iterations, the batch size was set to 32, the learning rate was set to 0.001, the weight decay was set to 0.1. The decay step was set to 8000.
[0118] The present invention uses the angular error, which is the mainstream evaluation index of gaze estimation, to compare the performance with other gaze estimation models, that is, the deviation angle between the predicted value and the true value of gaze estimation, and numerically represents the performance. The smaller this index, the better the effect. Assume that the actual gazing direction is g and the estimated gazing direction is Then the angular error can be calculated as:
[0119]
[0120] where, |||| represents taking the norm, ||g|| is the scalar value of the vector of the actual gazing direction, is the scalar value of the vector of the estimated gazing direction.
[0121] The comparison models adopt advanced methods for gaze estimation such as Mnist, a method for gaze direction estimation called GazeNet, a multi-task face analysis method called FullFace, a real-time generative adversarial network called RT-Gene, and a neural network called Dilated-Net constructed by dilated convolution. The experimental settings for each method adopt the settings in the corresponding papers, including the model framework and hyperparameters, so as to reproduce their network performance. The experimental results are shown in Table 1, and the optimal results are shown in bold.
[0122] Table 1 Experimental results of the NormalLight / LowLight network proposed by the present invention and other advanced networks
[0123]
[0124] It can be seen from the experimental data in Table 1 that through experimental verification, it can be seen that the angular error in this solution is lower than that of other methods. Therefore, the accuracy of gaze estimation in this solution is better than that of other methods; the method of the present invention can effectively solve the problem of reduced accuracy of gaze estimation in low-light environments and has strong practical value.
[0125] Example 3
[0126] This example will introduce an applicable scenario of the present invention: The gaze estimation technology has a wide range of application scenarios, and one of them is the monitoring of learning concentration in the online education environment. The all-day gaze estimation monitoring system and method proposed by the present invention can be used to detect the attention state of learning personnel in front of the computer and provide intelligent reminders when the attention is distracted, so as to improve learning efficiency.
[0127] First, the system captures the learner's facial images in real time through a camera, and uses the lighting condition detection module to judge the lighting situation of the current environment, automatically adjusting the gaze estimation model to ensure stable operation under different lighting conditions. Then, the system uses ResNet18 for facial feature extraction and combines the corresponding gaze estimation model to calculate the learner's gaze angle and determine whether their gaze is focused on the learning screen. If it is detected that the gaze deviates from the learning content and the deviation time exceeds the set threshold, the system records the relevant data and conducts a comprehensive analysis in combination with the historical learning behavior. Finally, when the system detects that the learner's attention is distracted (for example, the gaze deviates from the screen for a long time or the head is frequently lowered), it will intervene by means of pop-up reminders, sound prompts, or learning data feedback to guide the learner to refocus their attention. In addition, the system can generate a learning concentration report to provide data support for teachers or parents to optimize personalized learning strategies.
[0128] The above is only a description of the preferred embodiments of the present invention. Those of ordinary skill in the art can make several modifications and optimizations based on the above disclosure without departing from the basic principle content, and these improvements and optimizations should be regarded as the protection scope understood by the present invention.
Claims
1. A line-of-sight estimation method for all-day monitoring, characterized in that: It includes real-time acquisition of video stream data in the monitoring area, quantification and analysis of the brightness information of the acquired frame image using a light intensity detection algorithm to obtain the result of the quantification analysis; using the MTCNN network to perform face detection and localization on the acquired frame image to obtain a face image, and comparing the light intensity of the result of the quantification analysis with a preset threshold. When the light intensity is lower than the preset threshold, enable the LowLight line-of-sight estimation model for low-light compensation, input the face image into the LowLight line-of-sight estimation model for low-light compensation to obtain the line-of-sight estimation result. When the light intensity is not lower than the preset threshold, enable the NormalLight line-of-sight estimation model for standard illumination, input the face image into the NormalLight line-of-sight estimation model for standard illumination to obtain the line-of-sight estimation result. Map the line-of-sight estimation result through a three-dimensional eyeball coordinate system to generate line-of-sight direction data, and evaluate the attention concentration of the line of sight. The real-time processing pipeline dynamically analyzes the video stream data, dynamically presents the evaluation result through a human-computer interaction interface, and provides storage and structured export of the visualized data.
2. The all-day monitoring line-of-sight estimation method according to claim 1, wherein: The quantification and analysis of the brightness information of the acquired frame image using the light intensity detection algorithm includes image brightness calculation, local light intensity analysis, global light intensity analysis, time series light intensity analysis, and multi-channel light intensity fusion analysis. For the image brightness calculation, for a color image, convert the brightness value L of the picture from the RGB color space to the brightness component by weighted average. The calculation process is as follows: L = 0.299×R + 0.587×G + 0.114×B In the formula, L is the brightness value of this pixel point, and R, G, and B are the red, green, and blue components of this pixel point respectively. For the local light intensity analysis, the image is divided into several sub-regions, and the local light intensity of each sub-region; the local light intensity L region The calculation process is as follows: In the formula, L(i, j) is the brightness value of the (i, j)th pixel, and m and n are the width and height of the sub-region respectively. The global illumination intensity analysis calculates the global illumination intensity of an image based on the local brightness of all sub-regions; the global illumination intensity L global The calculation process is as follows: where M is the total number of sub-regions, and L region (k) is the luminance value of the k-th sub-region.
3. The method for estimating line of sight for all-day monitoring according to claim 2, wherein: For the above time - series light intensity analysis, time - series analysis techniques are used to track the changing trend of light. The brightness value L of each frame image is calculated in k consecutive frame images. avg (t); The changing trend of light in the time series, that is, the light change amount ΔL avg (t) is expressed as: ΔL avg (t) = L avg (t) - L avg (t - 1) The multi-channel light intensity fusion analysis introduces a multi-channel light intensity analysis method and combines the brightness data of RGB images, infrared images, and ultraviolet images; the global light intensity L global is used as the brightness data of the RGB image, and the comprehensive light intensity calculation process is as follows: Where, L RGB is the illumination intensity of the RGB image, L IR and L UV are the illumination intensities of the infrared and ultraviolet channels respectively, and W1, W2, and W3 are the weight coefficients of the RGB image, the infrared, and the ultraviolet channels respectively, is the multiplication operation.
4. The all-day monitoring line-of-sight estimation method according to claim 3, characterized in that: The light intensity L of the result of the quantitative analysis total is compared with a preset threshold L threshold for comparison When the light intensity is lower than the preset threshold L total <L threshold then it is determined that the image brightness is a low-light image, and the LowLight visual estimation model for low-light compensation is enabled; When the light intensity is not lower than the preset threshold L total ≥L threshold then it is determined that the image brightness is a standard illumination image, and the standard illumination NormalLight line-of-sight estimation model is enabled; The system dynamically adjusts the preset threshold L according to real-time illumination data and historical illumination change trends threshold ; The calculation formula is as follows: L threshold L(t) = threshold L(t - 1)+αΔL avg (t) Where α is the dynamic adjustment coefficient, and αΔL avg (t) is the amount of light change, and L threshold (t) is the preset threshold at time t.
5. The all-day monitoring line-of-sight estimation method according to claim 1, characterized in that: For the NormalLight line-of-sight estimation model for standard illumination, extract high-dimensional facial feature vectors through the ResNet18 network of the deep convolutional neural network model according to the input face image; process the extracted high-dimensional facial feature vectors through two fully connected layers to output a three-dimensional feature vector; convert the three-dimensional feature vector into an angular error of line-of-sight estimation through a tangent transformation; use the mean squared error loss function MSELoss as the loss function of the model; MSELoss measures the error by calculating the expected value of the square of the difference between the predicted value and the true value; the mean squared error loss function MSELoss is defined as follows: where ξ gt is the true value of the line of sight, and ξ pred is the predicted value of the line of sight estimate, and n is the number of samples.
6. A method for estimating line of sight for all-day monitoring according to claim 1, characterized in that: For the LowLight line-of-sight estimation model for low-light compensation, the face image passes through basic convolutional layers, batch normalization BN, activation functions, and max pooling MaxPool operations in sequence; then it is input into the residual network and processed through the residual modules in the first stage, second stage, third stage, and fourth stage to extract global features. Finally, the face features are formed through the processing of the SSD module. The residual modules in the second and third stages are improved residual modules with the ELA attention mechanism module introduced. Both the left-eye image and the right-eye image in the face image sequentially pass through the basic convolutional layer, batch normalization BN, activation function, and max pooling MaxPool operations, and then are processed through three improved residual modules with the ELA attention mechanism module introduced to extract the detailed features of the eyes. Then, local features are extracted through adaptive average pooling and an additional convolutional module, and finally, the left-eye features and the right-eye features are formed. The face features, left-eye features, and right-eye features are equally fused through a multi-layer perceptron MLP, and then pass through two fully connected layers in sequence to obtain the final gaze estimation result.
7. A line-of-sight estimation method for all-day monitoring according to claim 1, characterized in that: The improved residual module with the ELA attention mechanism module includes two improved residual sub-modules. The output end of the first improved residual sub-module is connected to the input end of the second improved residual sub-module for transmission. The improved residual sub-module includes a residual block and an ELA attention mechanism module. The residual block processes the input of the improved residual sub-module to obtain the output of the residual block. The output of the residual block is processed through the ELA attention mechanism module to obtain the output of the ELA attention mechanism module. The output obtained by multiplying and fusing the output of the residual block and the output of the ELA attention mechanism module is used as the output of the improved residual sub-module.
8. A method for estimating line of sight for all-day monitoring according to claim 7, wherein: The ELA attention mechanism module extracts feature vectors in the horizontal and vertical directions by using strip pooling in the spatial dimension; the input feature map is X h ∈R B×C×H×W ; Average over the width dimension to obtain a feature map of size B×C×H×1 Then the feature map Passes through a 1D convolution, group normalization, and Sigmoid activation function in sequence to obtain a feature map X of size B×C×H×1 h ; Average over the height dimension to obtain a feature map of size B×C×1×W Then the feature map Passes through a 1D convolution, group normalization, and Sigmoid activation function in sequence to obtain a feature map X of size B×C×1×W w ; Finally, multiply the input feature map X by the product of the two attention feature maps X h and X w to obtain the spatial attention feature map Y; The calculation process is as follows: In the formula, represents element-wise multiplication.
9. The all-day monitoring line-of-sight estimation method according to claim 6, characterized in that: The multi-layer perceptron includes an input layer, two hidden layers, a batch normalization layer, a ReLU activation function, a Dropout layer, and an output layer. The final output feature obtained through the comprehensive processing of the multi-stream input features is used as the gaze estimation result. ξ pred = MLP(F eye , F face ) where F eye is an eye feature, including a left-eye feature and a right-eye feature; F face is a facial feature; MLP is the operation of the multi-layer perceptron.
Citation Information
Patent Citations
Line-of-sight Detection Apparatus, Method, Image Capturing Apparatus And Control Method
CN104063709A
Fatigue driving detection method based on illumination detection
CN109635664A
Data processing system, method and device for pupil positioning
CN117935343A
Line-of-sight estimation method based on low-light image enhancement technology
CN118762393A
Facial part detection apparatus
JP2010134867A