Classroom teacher behavior spatiotemporal feature fusion tracking method
By locking the alias frequency of the classroom environment and performing fractional delay phase equalization, the problem of decreased eye tracking accuracy caused by interference from the projection light source and indoor lighting was solved, enabling stable monitoring of the teacher's eye trajectory and accurate analysis of teaching behavior.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-03-31
AI Technical Summary
In a real classroom environment, the specular reflection and periodic light intensity changes produced by the projection light source and indoor lighting on the whiteboard or interactive whiteboard surface lead to a decrease in eye tracking accuracy. The teacher's eye trajectory oscillates periodically, affecting the spatial positioning accuracy of teaching behavior and the reliability of teaching analysis models based on eye data.
By acquiring the pupil-corneal reflection difference signal of the teacher's eye movement, locking the alias frequency of the classroom environment lighting modulation frequency and the eye movement sampling rate, calculating the whiteboard glare coefficient and performing fractional delay phase equalization, generating a de-twisting vector, and solving the coordinates of the teacher's gaze point on the whiteboard plane, stable monitoring of the teacher's blackboard writing behavior is achieved.
It effectively overcomes optical interference, ensures the stability and accuracy of teachers' eye trajectory, and improves the reliability of eye tracking and the credibility of teaching behavior analysis in classroom scenarios.
Smart Images

Figure CN121305689B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of eye-tracking technology, and more specifically, to a method for tracking the spatiotemporal characteristics of classroom teacher behavior. Background Technology
[0002] With the development of educational informatization and intelligent teaching, eye-tracking technology based on video and infrared is widely used to study teachers' attention distribution and lecturing behavior in the classroom. Existing research is mostly conducted in experimental environments with uniform lighting and controlled reflection. However, in real classrooms, teachers often lecture in front of whiteboards or interactive whiteboards with direct projection light, and these surfaces typically have high gloss and strong reflectivity. Furthermore, lighting equipment and projectors generally use AC-powered LEDs or fluorescent lamps, which exhibit periodic brightness fluctuations (temporal light modulation) within the visually invisible range. These factors collectively cause continuous changes in the lighting conditions of eye-tracking images, introducing complex interference to eye-tracking monitoring in classroom scenarios.
[0003] When a teacher writes or explains near the whiteboard, the specular reflection area moves with the teacher's stance and head posture, causing strong reflected light to create localized overexposure in the camera or eye tracker lens. Simultaneously, the temporal modulation of illumination and projection causes periodic fluctuations in image brightness at a frequency of approximately 100 Hz. When the camera uses a rolling shutter, the exposure time for different scan lines varies slightly; this inter-line time difference translates the light intensity modulation into brightness bands at the image level. Because gaze-tracking algorithms rely on the precise extraction of the pupil center and corneal highlight positions, the periodic changes in brightness and localized overexposure cause a regular shift in the estimated position of the pupil boundary and reflection point over time. This shift manifests at the signal level as a stable phase difference between the horizontal and vertical directions, causing the system's calculated gaze to exhibit temporal cross-twisting.
[0004] This phase misalignment, caused by both whiteboard reflection and light source modulation, persists throughout the teacher's explanation. Even if the teacher maintains a stable head posture or focuses on fixed content, the system will still misjudge the point where their gaze lands on the whiteboard. This results in the gaze trajectory exhibiting periodic oscillations on the image, with the gaze point jumping between adjacent lines of text or constantly shifting between adjacent illustrations. This distortion not only compromises the spatial positioning accuracy of teaching behavior but also renders the teaching analysis model based on gaze data unreliable. Summary of the Invention
[0005] This invention provides a method for tracking the spatiotemporal features of classroom teacher behavior, which solves the technical problems mentioned in the background.
[0006] The first aspect is the method for tracking the spatiotemporal characteristics of classroom teacher behavior, including:
[0007] Acquire the pupil-corneal reflection difference signal of the teacher's eye movement, and lock the alias frequency based on the aliasing relationship between the classroom ambient lighting modulation frequency and the eye movement sampling rate;
[0008] Calculate the whiteboard glare coefficient of the difference signal at the alias frequency. The whiteboard glare coefficient is used to quantify the degree of phase distortion caused by the coupling of whiteboard specular reflection and rolling shutter imaging.
[0009] The whiteboard glare coefficient is analyzed as the effective time misalignment between axes, and the effective time misalignment between axes is used to perform fractional delay phase equalization on the image line scanning axis component of the difference signal to generate a de-twisting slant vector.
[0010] The coordinates of the teacher's gaze point on the whiteboard plane are calculated based on the de-torsion oblique vector to monitor the teacher's blackboard writing behavior.
[0011] Secondly, the classroom teacher behavior spatiotemporal feature fusion tracking system, applied in any of the classroom teacher behavior spatiotemporal feature fusion tracking methods described above, includes:
[0012] The frequency locking module acquires the pupil-corneal reflection difference signal of the teacher's eye movement and locks the alias frequency based on the aliasing relationship between the classroom ambient lighting modulation frequency and the eye movement sampling rate.
[0013] The distortion quantization module calculates the whiteboard glare coefficient of the difference signal at an alias frequency. The whiteboard glare coefficient is used to quantify the degree of phase distortion caused by the coupling between the whiteboard specular reflection and the rolling shutter imaging.
[0014] The distortion analysis module resolves the whiteboard glare coefficient as an effective time misalignment between axes, and uses the effective time misalignment between axes to perform fractional delay phase equalization on the image line scan axis components of the difference signal, generating a de-distortion slant vector;
[0015] The gaze point monitoring module calculates the coordinates of the teacher's gaze point on the whiteboard plane based on the de-twisted oblique vector, so as to monitor the teacher's blackboard writing behavior.
[0016] The beneficial effects of this invention include: in a real classroom environment, it can effectively overcome the interference of specular reflection and periodic light intensity changes on the surface of a whiteboard or interactive whiteboard caused by projection light sources and indoor lighting, thereby achieving stable, continuous, and accurate monitoring of the teacher's gaze direction and explanation behavior. By real-time identification and correction of time misalignment caused by light source modulation in eye movement signals, this invention ensures that the teacher's gaze trajectory remains consistent and accurate under strong light, reflective light, and dynamic lighting conditions, thus significantly improving the reliability of eye tracking and the credibility of teaching behavior analysis in classroom scenarios. Attached Figure Description
[0017] Figure 1 This is an alternative to the Lissajous trajectory diagram of the present invention;
[0018] Figure 2 This is a stability diagram of the gaze point on the whiteboard surface according to the present invention;
[0019] Figure 3 This is a flowchart of the classroom teacher behavior spatiotemporal feature fusion tracking method of the present invention;
[0020] Figure 4 This is a block diagram of the classroom teacher behavior spatiotemporal feature fusion tracking system of the present invention. Detailed Implementation
[0021] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0022] like Figure 1 As shown, Figure 1 This shows the pupil-corneal reflectance difference vector after bandpass filtering at the locked alias frequency (i.e., Figure 1 The title shows a schematic diagram comparing the Lissajous trajectories of the horizontal and vertical components of the PCCR vector components after bandpass filtering. The orange curve represents the original input signal trajectory before processing. Due to the coupling between classroom ambient lighting modulation (TLM) and the camera's rolling shutter effect, there is a significant phase lag between the horizontal and vertical components (i.e., a non-zero whiteboard glare coefficient), resulting in a broad elliptical shape on the phase plane, corresponding to a significant quarter-phase error. The blue straight line represents the processed output signal trajectory after fractional delay phase equalization based on the inter-axis effective time misalignment. It can be seen that the original elliptical trajectory has collapsed into an approximately straight line, indicating that the phase lag of the vertical component has been compensated, achieving phase alignment. This comparison intuitively verifies the effectiveness of this invention in eliminating classroom-specific optical-sampling alias interference at the signal level.
[0023] like Figure 2 As shown, Figure 2This diagram illustrates the spatial distribution stability of the gaze point calculated by the system on the physical plane of the whiteboard when a teacher is simulating prolonged gaze at a fixed target point on the whiteboard. The black plus signs represent the actual center position of the gaze target. The orange scatter clusters represent the gaze point distribution without de-skewing processing. Driven by continuous periodic phase noise at alias frequencies, the gaze point coordinates fluctuate significantly, exhibiting a divergent distribution with directional stretching, severely affecting the judgment of the writing location. The blue scatter clusters represent the gaze point distribution calculated after generating the de-skewed vector using the method of this invention. The scatter points converge tightly and cover the area around the actual target center, with a significantly smaller distribution range than before processing. This demonstrates that by eliminating underlying phase distortion, this invention can effectively suppress gaze jitter caused by whiteboard glare and illumination modulation, significantly improving the stability and positioning accuracy of gaze tracking in monitoring teacher writing behavior.
[0024] Example 1: As Figure 3 As shown, the method for fusion and tracking the spatiotemporal characteristics of classroom teacher behavior includes:
[0025] Acquire the pupil-corneal reflection difference signal of the teacher's eye movement, and lock the alias frequency based on the aliasing relationship between the classroom ambient lighting modulation frequency and the eye movement sampling rate;
[0026] Calculate the whiteboard glare coefficient of the difference signal at the alias frequency. The whiteboard glare coefficient is used to quantify the degree of phase distortion caused by the coupling of whiteboard specular reflection and rolling shutter imaging.
[0027] The whiteboard glare coefficient is analyzed as the effective time misalignment between axes, and the effective time misalignment between axes is used to perform fractional delay phase equalization on the image line scanning axis component of the difference signal to generate a de-twisting slant vector.
[0028] The coordinates of the teacher's gaze point on the whiteboard plane are calculated based on the de-torsion oblique vector to monitor the teacher's blackboard writing behavior.
[0029] In one embodiment of the present invention, the method of acquiring the pupil-corneal reflection difference signal of the teacher's eye movement and locking the alias frequency based on the aliasing relationship between the classroom ambient lighting modulation frequency and the eye movement sampling rate includes:
[0030] The original pupil-corneal reflection difference vector is obtained by subtracting the corneal reflection center coordinates from the pupil center coordinates. The original pupil-corneal reflection difference vector is then subjected to detrending and bandpass filtering to generate a preprocessed pupil-corneal reflection difference signal.
[0031] Traverse the preset set of illumination modulation frequencies, calculate the minimum absolute value of the difference between each frequency value in the set and an integer multiple of the eye-tracking sampling rate, and obtain the set of candidate alias frequencies;
[0032] Power spectral density estimation is performed on the preprocessed pupil-corneal reflection difference signal to obtain the horizontal component self-power spectral density, vertical component self-power spectral density, and horizontal-vertical component cross-power spectral density at each alias frequency candidate value.
[0033] The imaginary part of the cross-power spectral density is calculated and divided by the square root of the product of the transverse component self-power spectral density and the longitudinal component self-power spectral density to obtain the whiteboard glare coefficient value corresponding to each alias frequency candidate value.
[0034] Compare the absolute values of the whiteboard glare coefficients corresponding to each candidate alias frequency, and determine the candidate alias frequency corresponding to the one with the largest absolute value as the locked alias frequency.
[0035] Pupil center coordinates are two-dimensional coordinates of the geometric center point of the pupil located in the eye image, reflecting the real-time position of the pupil in the image, and are extracted by the image recognition algorithm of the eye tracker.
[0036] The corneal reflection center coordinates are the two-dimensional coordinates of the geometric center of the reflective bright spot (corneal reflection point) formed by the light source illuminating the cornea, extracted by the highlight detection algorithm of the eye tracker.
[0037] The original difference vector of pupil-corneal reflection is a two-dimensional vector obtained by subtracting the corresponding components of the pupil center coordinates from the corneal reflection center coordinates. It directly reflects the relative positional relationship between the two and contains basic information about the direction of vision, but it is subject to noise and drift interference. Specifically, the horizontal component of the original difference vector of pupil-corneal reflection is equal to the horizontal coordinate of the pupil center minus the horizontal coordinate of the corneal reflection center, and the vertical component is equal to the vertical coordinate of the pupil center minus the vertical coordinate of the corneal reflection center.
[0038] Detrending is a process that removes the slow changing trends caused by head movements, body swaying, etc., from the original difference vector of pupillary and corneal reflexes, thus avoiding low-frequency interference from affecting subsequent signal analysis. Specifically, a second-order polynomial fitting method is used. First, the horizontal and vertical components of the original difference vector are fitted with second-order polynomials to obtain trend curves. Then, the values of the corresponding trend curves are subtracted from the original component data to obtain the detrended component data.
[0039] Bandpass filtering is a process that retains the narrowband signal related to the classroom lighting modulation frequency in the original difference vector while filtering out high-frequency noise and low-frequency redundant signals. Specifically, the filtering frequency band is set to 0.5 Hz to (half of the eye-tracking sampling rate minus 5 Hz), and a zero-phase bandpass filter is used to filter the detrended horizontal and vertical components respectively to obtain the filtered component data.
[0040] The preprocessed pupil-corneal reflection difference signal is a clean signal obtained after detrending and bandpass filtering, which eliminates most of the interference and retains the effective information related to phase distortion.
[0041] The preset lighting modulation frequency set is a set of frequencies pre-set according to the characteristics of common classroom lighting equipment, covering the typical modulation frequencies of equipment such as LED lights, fluorescent lights and projectors; specifically, the preset lighting modulation frequency set is set to include 90 Hz, 100 Hz, 110 Hz, 120 Hz and 130 Hz.
[0042] Eye sampling rate is the frequency at which an eye tracker acquires images of the eye, i.e., the number of image frames acquired per unit time, measured in Hertz. It is a key device parameter for calculating the alias frequency and is determined by the hardware specifications of the eye tracker used.
[0043] The alias frequency candidate value set is an effective frequency set formed by sampling and aliasing the preset illumination modulation frequency. Each candidate value is within the Nyquist frequency range corresponding to the eye-tracking sampling rate and is a candidate range for subsequently locking the target alias frequency. Specifically, for each frequency in the preset illumination modulation frequency set, the absolute value of the difference between the frequency and an integer multiple of the eye-tracking sampling rate is calculated, and the minimum value is taken as the alias frequency candidate value corresponding to the frequency. Candidate values less than or equal to 0 Hz and greater than half of the eye-tracking sampling rate are eliminated, and the remaining candidate values form the alias frequency candidate value set.
[0044] Power spectral density estimation is a method for analyzing the frequency domain characteristics of a signal. It is used to extract the energy distribution and phase correlation characteristics of the preprocessed signal at each candidate alias frequency. Specifically, a sliding window with a length of T seconds and a step size of T / 2 seconds is used to window the preprocessed signal in segments, calculate the Fourier transform of each segment, and then average the frequency domain results of each segment to obtain the power spectral density estimate at each frequency.
[0045] The transverse component power spectral density is the power spectral density of the transverse component of the preprocessed signal at a specific frequency, reflecting the energy intensity of the transverse component at that frequency.
[0046] The longitudinal component power spectral density is the power spectral density of the longitudinal component of the preprocessed signal at a specific frequency, reflecting the energy intensity of the longitudinal component at that frequency.
[0047] The cross-power spectral density of the horizontal and vertical components is an indicator that reflects the degree of phase correlation between the horizontal and vertical components of the preprocessed signal at a specific frequency. Its imaginary part is the core data for calculating the whiteboard glare coefficient.
[0048] The whiteboard glare coefficient is a numerical value that quantifies the degree of phase distortion caused by the coupling between the specular reflection of the whiteboard and the rolling shutter imaging at a specific alias frequency. The value ranges from -1 to 1. Specifically, the whiteboard glare coefficient is equal to the imaginary part of the cross power spectral density of the horizontal and vertical components at that alias frequency, divided by the square root of the product of the horizontal component's self-power spectral density and the vertical component's self-power spectral density.
[0049] The locked alias frequency is the frequency with the most significant phase distortion among the candidate alias frequencies. This frequency reflects the coupling interference between optics and sampling in the classroom environment.
[0050] In one embodiment of the present invention, calculating the whiteboard glare coefficient of the difference signal at an alias frequency includes:
[0051] The horizontal and vertical components of the preprocessed pupil-corneal reflection difference signal are segmented and windowed using a time window function. Fourier transform is then performed on each segmented signal after windowing to extract the frequency domain response value at the locked alias frequency.
[0052] The average of the squared modulus of the frequency domain response values of each segment's transverse component is used to obtain the transverse component's self-power spectral density. The average of the squared modulus of the frequency domain response values of each segment's longitudinal component is used to obtain the longitudinal component's self-power spectral density. Finally, the average of the complex conjugate product of the frequency domain response values of each segment's transverse component and longitudinal component is used to obtain the cross-power spectral density.
[0053] Obtain the imaginary part of the cross-power spectral density, and calculate the square root of the product of the transverse component self-power spectral density and the longitudinal component self-power spectral density. Divide the imaginary part by the square root to obtain the whiteboard glare coefficient.
[0054] The time window function is used for signal segmentation processing. It can reduce spectral leakage when segmenting signals and adapt to the frequency domain analysis requirements of eye-tracking signals. Specifically, the Hanning window is selected as the time window function, and the window length is set to the time length corresponding to one-tenth of the eye-tracking sampling rate, that is, the window length is equal to the reciprocal of the eye-tracking sampling rate multiplied by ten.
[0055] The horizontal component of the preprocessed pupil-corneal reflection difference signal is the signal component along the horizontal direction in the preprocessed pupil-corneal reflection difference signal, which retains effective information on the relative positional changes of the pupil and corneal reflection point in the horizontal direction.
[0056] The longitudinal component of the preprocessed pupil-corneal reflection difference signal is the signal component along the vertical direction in the preprocessed pupil-corneal reflection difference signal, which retains effective information on the relative positional changes of the pupil and corneal reflection point in the vertical direction.
[0057] The segmentation operation divides the horizontal and vertical component signals into continuous signal segments according to the time window length. The windowing operation multiplies each signal segment by the time window function to reduce spectral leakage. Specifically, the horizontal and vertical component signals are slidably segmented according to the set time window length, with the sliding step size set to half the window length. Each segmented signal is multiplied by the corresponding point of the Hanning window function to obtain the windowed segmented signal.
[0058] The Fourier transform is a mathematical operation that converts a windowed segmented signal in the time domain into a signal in the frequency domain, used to obtain the response characteristics of a signal at different frequencies. Specifically, a fast Fourier transform is performed on each windowed segmented signal to convert the time domain signal into a complex frequency domain signal, thus obtaining the complete frequency domain result of each segmented signal.
[0059] The frequency response value is the complex value of a signal at a locked alias frequency, containing the amplitude and phase information of the signal at that frequency. It is divided into the transverse component frequency response value and the longitudinal component frequency response value. Specifically, in the frequency domain signal obtained by Fourier transform, the complex value corresponding to the locked alias frequency is found. This value is the frequency domain response value of the corresponding segmented signal at that frequency.
[0060] The squared modulus of the transverse component frequency domain response is the result of squaring the amplitude of the transverse component frequency domain response (complex number), reflecting the energy intensity of the transverse signal segment at the target frequency; specifically, the real and imaginary parts of the transverse component frequency domain response of each segment are squared and then added together to obtain the squared modulus of the transverse component frequency domain response of that segment.
[0061] The squared modulus of the longitudinal component frequency domain response is the result of squaring the amplitude of the longitudinal component frequency domain response (complex number), reflecting the energy intensity of the segmented longitudinal signal at the target frequency; specifically, the real and imaginary parts of the longitudinal component frequency domain response of each segment are squared and then added together to obtain the squared modulus of the longitudinal component frequency domain response of that segment.
[0062] The complex conjugate of the longitudinal component frequency domain response value is the result of performing a conjugate operation on the longitudinal component frequency domain response value (complex number), which is used to reflect the phase correlation between the horizontal and vertical signals in the frequency domain. Specifically, the real part of the longitudinal component frequency domain response value is kept unchanged, and the imaginary part is changed to its opposite value to obtain the complex conjugate of the longitudinal component frequency domain response value.
[0063] The imaginary part of the cross power spectral density is the imaginary part of the cross power spectral density (complex number), which directly reflects the phase difference characteristics of the horizontal and vertical signals at the target frequency.
[0064] The square root of the product of the horizontal component power spectral density and the vertical component power spectral density is a parameter used for normalization, which can eliminate the interference of the difference in amplitude between the horizontal and vertical signals on the phase distortion normalization result. Specifically, the product of the horizontal component power spectral density and the vertical component power spectral density is first calculated, and then the arithmetic square root of the product is taken to obtain the normalization factor.
[0065] In one embodiment of the present invention, resolving the whiteboard glare coefficient as an interaxial effective time misalignment includes:
[0066] A numerical truncation operation is performed on the whiteboard glare coefficient to limit its value within the preset domain range of the arcsine function, thus obtaining the limited coefficient value.
[0067] Calculate the arcsine of the coefficient value after limiting to obtain the phase angle value;
[0068] Calculate the product of twice pi and the locked alias frequency to obtain the angular frequency value;
[0069] Dividing the phase angle value by the angular frequency value yields the effective time misalignment between axes.
[0070] The numerical truncation operation is a method of forcibly adjusting the whiteboard glare coefficient that exceeds the preset range to the interval boundary value to avoid domain errors in subsequent arcsine calculations. Specifically, if the whiteboard glare coefficient is greater than the upper limit of the preset interval, it is adjusted to the upper limit value; if it is less than the lower limit of the preset interval, it is adjusted to the lower limit value; if it is within the interval, the original value remains unchanged.
[0071] The preset domain interval of the arcsine function is the numerical range within which the arcsine function can be effectively calculated, specifically adapted to the value characteristics of the whiteboard glare coefficient; specifically, the preset domain interval of the arcsine function is set to a closed interval from -0.999 to 0.999.
[0072] The coefficient value after limiting is the effective value obtained after the whiteboard glare coefficient is truncated, and it is completely within the domain of the arcsine function.
[0073] The arcsine value is the result of performing an arcsine mathematical operation on the coefficient values after amplitude limiting. This result directly corresponds to the phase angle of the horizontal and vertical signals at the target frequency. Specifically, the angle result obtained by performing the operation on the coefficient values after amplitude limiting using the standard arcsine calculation method is the arcsine value.
[0074] The phase angle value is an angular quantity that reflects the magnitude of the phase difference between the horizontal and vertical signals at a locked alias frequency. It is a key intermediate quantity connecting phase distortion and time misalignment.
[0075] The angular frequency value is the result of converting the locked alias frequency from Hertz units to angular frequency units, and is used to establish the quantization relationship between phase angle and time misalignment; specifically, the angular frequency value is equal to pi multiplied by two, and then multiplied by the locked alias frequency.
[0076] Inter-axis effective time misalignment is a physical quantity that quantifies the time delay between one image line scan axis component and another axis component, directly reflecting the degree of signal misalignment caused by the coupling of the rolling shutter and illumination modulation.
[0077] In one embodiment of the present invention, fractional delay phase equalization is performed on the image line scan axis components of the difference signal using the effective time misalignment between axes to generate a de-twisting slant vector, including:
[0078] The horizontal component of the preprocessed pupil-corneal reflection difference signal is assigned as the horizontal component of the de-twisted slant vector.
[0079] Determine the order of the Lagrange interpolation filter, and use the inverse of the effective time misalignment between axes as the delay parameter. Calculate the product of the delay parameter minus the difference of the accumulated iteration variable and the ratio of the filter tap index minus the difference of the accumulated iteration variable to obtain the filter coefficients.
[0080] The longitudinal component of the preprocessed pupil-corneal reflection difference signal is used as the image line scan axis component. The sum of the product of the filter coefficient and the sampled value of this component at the current sampling time minus the sampled value at the corresponding time of the filter tap index is calculated to obtain the longitudinal component of the de-twisted slant vector.
[0081] The transverse component of the preprocessed pupil-corneal reflection difference signal is the horizontal signal component after detrending and bandpass filtering. It has no significant phase distortion and can be directly used as the transverse component of the de-distorted slant vector.
[0082] The horizontal component of the de-twisted slant vector is the signal part along the horizontal direction in the de-twisted slant vector. The pre-processed horizontal component is used directly to ensure that the original effective information of the horizontal signal is not lost.
[0083] The order of the Lagrange interpolation filter is a key parameter that determines the accuracy and complexity of the Lagrange interpolation filter. The higher the order, the higher the interpolation accuracy, but the greater the computational load. It is necessary to balance accuracy and efficiency in the setting. Specifically, the order of the Lagrange interpolation filter is set to an even number, with a value range from the fourth to the eighth order.
[0084] The inverse of the effective time misalignment between axes is the result of taking the opposite sign of the effective time misalignment between axes. It is used to offset the time delay of the image line scanning axis components and achieve phase balance. Specifically, the inverse of the effective time misalignment between axes is equal to zero minus the effective time misalignment between axes.
[0085] The delay parameter is a core parameter used to calculate the coefficients of the Lagrange interpolation filter, and it is directly related to the time misalignment compensation of the image line scan axis components.
[0086] The cumulative iteration variable is used to calculate the filter coefficients. Its value range corresponds to the order of the Lagrange interpolation filter, and it iterates through all integers from zero to the filter order.
[0087] The filter tap index is an index value that identifies the position of each tap in the Lagrange interpolation filter. The value is consistent with the cumulative iteration variable, starting from zero and increasing sequentially to the filter order.
[0088] The filter coefficients are the core parameters of the Lagrange interpolation filter, used to perform fractional delay compensation on the image line scan axis components. Each tap corresponds to a coefficient. Specifically, for each filter tap index, all non-self cumulative iteration variables are traversed, the difference between the delay parameter and the cumulative iteration variable is calculated, and the ratio of this difference to the difference between the filter tap index and the cumulative iteration variable is multiplied together to obtain the filter coefficient corresponding to that tap index.
[0089] The longitudinal component of the preprocessed pupil-corneal reflection difference signal is the vertical signal component after detrending and bandpass filtering. It has phase distortion caused by the coupling of rolling shutter and illumination modulation, and is the target of fractional delay phase equalization.
[0090] The image line scan axis component is the signal component corresponding to the scanning direction of the camera's rolling shutter. In a classroom scene, the rolling shutter scans along the vertical direction, so the vertical component is defined as this component, which is the main component carrying phase distortion.
[0091] The current sampling time refers to the time node at which the current data point of the image row scan axis component is acquired.
[0092] The filter tap index corresponds to a time node that is the current sampling time shifted backward by the filter tap index by sampling intervals. It is used to obtain the historical sampling data required for interpolation operations. Specifically, the filter tap index corresponds to a time node that is the current sampling time minus the filter tap index multiplied by the reciprocal of the eye-tracking sampling rate.
[0093] The sampled value is the signal value of the image line scan axis component at the time corresponding to the filter tap index.
[0094] The longitudinal component of the de-torsion vector is the vertical signal component obtained after fractional delay phase equalization, which eliminates the phase distortion caused by the effective time misalignment between axes.
[0095] The de-distorted slant vector is a two-dimensional vector composed of phase-corrected transverse and longitudinal components, which eliminates the optical and sampling coupling distortion unique to classrooms.
[0096] In one embodiment of the present invention, the coordinates of the teacher's gaze point on the whiteboard plane are calculated based on the de-twisted slant vector to monitor the teacher's blackboard writing behavior, including:
[0097] The de-distorted skew vector is substituted into the pre-calibrated Pupil Center Corneal Reflection (PCCR) mapping function for calculation to obtain a normalized camera gaze vector containing horizontal and vertical coordinate components.
[0098] The input homogeneous coordinate vector is constructed by using the horizontal and vertical coordinate components of the normalized camera line-of-sight vector and the numerical value. The output homogeneous coordinate vector is obtained by performing matrix multiplication on the predetermined homography matrix and the input homogeneous coordinate vector.
[0099] The third component of the output homogeneous coordinate vector is obtained as the homogeneous scale factor. The first and second components of the output homogeneous coordinate vector are divided by the homogeneous scale factor to obtain the coordinates of the gaze point on the whiteboard plane.
[0100] The pre-calibrated PCCR mapping function was established through previous calibration experiments. It is used to convert the de-distorted slant vector into the line-of-sight vector in the camera coordinate system. After calibration, it has stable conversion accuracy. Specifically, the nine-point calibration method is used. Nine marker points with known three-dimensional coordinates are set in the camera field of view. The de-distorted slant vector corresponding to each marker point is recorded. The parameters of the PCCR mapping function are obtained by fitting with the least squares method to determine the conversion relationship from vector to line of sight.
[0101] The normalized camera gaze vector is a unit vector in the camera coordinate system that reflects the teacher's gaze direction. It contains two components, horizontal and vertical, eliminating the influence of distance on the gaze direction.
[0102] The lateral component of the normalized camera gaze vector is the value along the horizontal direction of the camera in the normalized camera gaze vector, reflecting the horizontal offset of the gaze direction.
[0103] The vertical component of the normalized camera gaze vector is the value along the vertical direction of the camera in the normalized camera gaze vector, reflecting the vertical offset of the gaze direction.
[0104] The numerical value is a fixed constant used to construct homogeneous coordinate vectors, which is used to extend two-dimensional vectors into three-dimensional homogeneous vectors.
[0105] The input homogeneous coordinate vector is a three-dimensional vector obtained by expanding the two-dimensional normalized camera view vector. It serves as an intermediate carrier for mapping the camera view to the whiteboard plane. Specifically, the first component of the input homogeneous coordinate vector is the horizontal component of the normalized camera view vector, the second component is the vertical component, and the third component is fixed to the value of one. The input homogeneous coordinate vector is obtained by combining these components in this order.
[0106] The pre-determined homography matrix is a three-dimensional matrix describing the transformation relationship between the camera coordinate system and the whiteboard plane coordinate system. It is determined through prior calibration and is a key parameter for realizing the mapping of the line-of-sight vector to the whiteboard coordinates. Specifically, the checkerboard calibration method is used. A standard checkerboard is placed on the whiteboard plane, and multiple checkerboard images from different angles are taken. The camera coordinates and physical coordinates of the checkerboard corner points are extracted, and the homography matrix is obtained by solving a system of linear equations.
[0107] Matrix multiplication is a mathematical operation that applies the transformation relationship of the homography matrix to the input homogeneous coordinate vector, realizing the transformation from the camera's line of sight vector to the homogeneous vector of the whiteboard plane. Specifically, each row element of the homography matrix is multiplied by the three components of the input homogeneous coordinate vector and then summed. The three results are arranged in order to form the output homogeneous coordinate vector.
[0108] The output homogeneous coordinate vector is the result of matrix multiplication. It is a three-dimensional homogeneous vector that contains the homogeneous coordinate information of the gaze point on the whiteboard plane.
[0109] The homogeneous scale factor is the third component value of the output homogeneous coordinate vector. It is used to eliminate scale ambiguity of homogeneous coordinates and convert the homogeneous vector into two-dimensional Cartesian coordinates.
[0110] The gaze point coordinates on the whiteboard plane are the two-dimensional coordinates of the teacher's gaze point on the physical plane of the whiteboard. They directly reflect the teacher's gaze position when writing on the whiteboard and are the final result of teacher behavior monitoring. Specifically, the horizontal coordinate of the gaze point on the whiteboard plane is equal to the first component of the output homogeneous coordinate vector divided by the homogeneous scale factor, and the vertical coordinate is equal to the second component of the output homogeneous coordinate vector divided by the homogeneous scale factor.
[0111] Example 2, as Figure 4 As shown, the classroom teacher behavior spatiotemporal feature fusion tracking system, applied in any of the classroom teacher behavior spatiotemporal feature fusion tracking methods described herein, includes:
[0112] The frequency locking module acquires the pupil-corneal reflection difference signal of the teacher's eye movement and locks the alias frequency based on the aliasing relationship between the classroom ambient lighting modulation frequency and the eye movement sampling rate.
[0113] The distortion quantization module calculates the whiteboard glare coefficient of the difference signal at an alias frequency. The whiteboard glare coefficient is used to quantify the degree of phase distortion caused by the coupling between the whiteboard specular reflection and the rolling shutter imaging.
[0114] The distortion analysis module resolves the whiteboard glare coefficient as an effective time misalignment between axes, and uses the effective time misalignment between axes to perform fractional delay phase equalization on the image line scan axis components of the difference signal, generating a de-distortion slant vector;
[0115] The gaze point monitoring module calculates the coordinates of the teacher's gaze point on the whiteboard plane based on the de-twisted oblique vector, so as to monitor the teacher's blackboard writing behavior.
[0116] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.
Claims
1. A method for tracking the spatiotemporal characteristics of classroom teacher behavior, characterized in that, include: Acquire the pupil-corneal reflex difference signal of the teacher's eye movement, and lock the alias frequency based on the aliasing relationship between the classroom ambient lighting modulation frequency and the eye movement sampling rate, including: The original pupil-corneal reflection difference vector is obtained by subtracting the corneal reflection center coordinates from the pupil center coordinates. The original pupil-corneal reflection difference vector is then subjected to detrending and bandpass filtering to generate a preprocessed pupil-corneal reflection difference signal. Traverse the preset set of illumination modulation frequencies, calculate the minimum absolute value of the difference between each frequency value in the set and an integer multiple of the eye-tracking sampling rate, and obtain the set of candidate alias frequencies; Power spectral density estimation is performed on the preprocessed pupil-corneal reflection difference signal to obtain the horizontal component self-power spectral density, vertical component self-power spectral density, and horizontal-vertical component cross-power spectral density at each alias frequency candidate value. The imaginary part of the cross-power spectral density is calculated and divided by the square root of the product of the transverse component self-power spectral density and the longitudinal component self-power spectral density to obtain the whiteboard glare coefficient value corresponding to each alias frequency candidate value. Compare the absolute values of the whiteboard glare coefficients corresponding to each candidate alias frequency, and determine the candidate alias frequency corresponding to the one with the largest absolute value as the locked alias frequency; Calculate the whiteboard glare coefficient at the alias frequency of the difference signal, including: The horizontal and vertical components of the preprocessed pupil-corneal reflection difference signal are segmented and windowed using a time window function. Fourier transform is then performed on each segmented signal after windowing to extract the frequency domain response value at the locked alias frequency. The average of the squared modulus of the frequency domain response values of each segment's transverse component is used to obtain the transverse component's self-power spectral density. The average of the squared modulus of the frequency domain response values of each segment's longitudinal component is used to obtain the longitudinal component's self-power spectral density. Finally, the average of the complex conjugate product of the frequency domain response values of each segment's transverse component and longitudinal component is used to obtain the cross-power spectral density. Obtain the imaginary part of the cross-power spectral density, and calculate the square root of the product of the transverse component self-power spectral density and the longitudinal component self-power spectral density. Divide the imaginary part by the square root to obtain the whiteboard glare coefficient. The whiteboard glare coefficient is used to quantify the degree of phase distortion caused by the coupling of specular reflection from the whiteboard and rolling shutter imaging. The whiteboard glare coefficient is analyzed as an inter-axis effective time misalignment, including: A numerical truncation operation is performed on the whiteboard glare coefficient to limit its value within the preset domain range of the arcsine function, thus obtaining the truncation coefficient value. Calculate the arcsine of the coefficient value after limiting to obtain the phase angle value; Calculate the product of twice pi and the locked alias frequency to obtain the angular frequency value; Dividing the phase angle value by the angular frequency value yields the effective time misalignment between axes; Furthermore, fractional delay phase equalization is performed on the image line scan axis components of the difference signal using the effective time misalignment between axes, generating a de-twisting slant vector, including: The horizontal component of the preprocessed pupil-corneal reflection difference signal is assigned as the horizontal component of the de-twisted slant vector. The order of the Lagrange interpolation filter is determined, and a delay parameter is obtained by taking the inverse of the inter-axis effective time offset. A product of the delay parameter minus an iteration variable multiplied by a difference between a filter tap index minus the iteration variable is calculated, and a ratio of the product to a difference between the filter tap index minus the iteration variable is obtained. The filter coefficient is obtained by multiplying the ratio by the filter tap index minus the iteration variable. The iteration variable is an iteration variable used to calculate the filter coefficient, and the iteration variable ranges from zero to all integers of the Lagrange interpolation filter; The longitudinal component of the pre-processed pupil corneal reflection difference signal is taken as an image row scanning axis component, and a sum of a product of the filter coefficient and a sampling value of the component at a current sampling time minus a sampling value at a filter tap index corresponding time is calculated to obtain a longitudinal component of the deskew vector; The gaze point coordinates of the teacher on the whiteboard plane are calculated based on the deskew vector to monitor the teacher's writing behavior on the whiteboard.
2. The classroom teacher behavior spatio-temporal feature fusion tracking method according to claim 1, characterized in that, The gaze point coordinates of the teacher on the whiteboard plane are calculated based on the deskew vector to monitor the teacher's writing behavior on the whiteboard, comprising: The deskew vector is substituted into the pre-calibrated PCCR mapping function to obtain a normalized camera line-of-sight vector containing horizontal and vertical coordinate components; The horizontal and vertical coordinate components of the normalized camera line-of-sight vector are used to construct an input homogeneous coordinate vector, and a pre-determined homography matrix is multiplied by the input homogeneous coordinate vector to obtain an output homogeneous coordinate vector; The third component of the output homogeneous coordinate vector is taken as a homogeneous scale factor, and the first and second components of the output homogeneous coordinate vector are divided by the homogeneous scale factor to obtain the gaze point coordinates on the whiteboard plane.
3. The classroom teacher behavior spatiotemporal feature fusion tracking system applied in the classroom teacher behavior spatiotemporal feature fusion tracking method of any one of claims 1-2, characterized in that, Comprising: A frequency locking module acquires a pupil corneal reflection difference signal of the teacher's eye movement, and locks an alias frequency according to the aliasing relationship between the classroom environment lighting modulation frequency and the eye movement sampling rate; A distortion degree quantification module calculates a whiteboard glare coefficient of the difference signal at the alias frequency, and the whiteboard glare coefficient is used to quantify the degree of phase distortion caused by the coupling of whiteboard mirror reflection and rolling shutter imaging; A distortion analysis module analyzes the whiteboard glare coefficient into an inter-axis effective time offset, and uses the inter-axis effective time offset to implement fractional delay phase equalization on the image row scanning axis component of the difference signal to generate a deskew vector; A gaze point monitoring module calculates the gaze point coordinates of the teacher on the whiteboard plane based on the deskew vector to monitor the teacher's writing behavior on the whiteboard.
Citation Information
Patent Citations
Eye movement tracking method, eye movement tracking device, equipment, computer equipment and medium
CN120610616A
Naked eye 3D display optimization method based on real-time eyeball tracking
CN120751110A