Frequency-Domain-Based Early Fatigue Detection Method for Airport Security X-Ray Machine Security Inspectors

The frequency domain-based method extracts the eye-wide ratio and gaze direction sequence characteristics of airport security inspectors, and uses deep learning networks to perform early fatigue detection, which solves the problem of difficulty in detecting early fatigue of security inspectors in the prior art, and improves security accuracy and safety.

CN117058576BActive Publication Date: 2025-07-29XIDIAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310953973.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-31
Publication Date
2025-07-29
Estimated Expiration
2043-07-31

AI Technical Summary

Technical Problem

There is a lack of effective early fatigue detection methods for airport security X-ray machine security inspectors in the prior art, which leads to high risk of missed detection and missed detection. The existing methods are mainly aimed at middle and late fatigue and it is difficult to detect early fatigue.

Method used

Using a frequency domain-based method, by obtaining the frequency domain characteristics of the surveillance video's eye width ratio and the gaze-estimated direction sequence, a deep learning layered multi-scale long and short-term memory network is used to identify the fatigue state of the security inspector.

Benefits of technology

The security inspection quality of security inspectors has been improved, the missed and missed inspections have been reduced, the safety of passengers has been effectively maintained, the original information is retained through frequency domain characteristics, and the accuracy of early fatigue detection has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058576B_ABST
    Figure CN117058576B_ABST
Patent Text Reader

Abstract

The present invention discloses an early fatigue detection method for airport security X-ray machine security inspectors based on the frequency domain, which relates to the technical field of image processing, and includes: obtaining a monitoring video to be detected; reading images from the monitoring video to be detected frame by frame in sequence; obtaining an eye width ratio sequence; obtaining a gaze estimation direction sequence; respectively extracting the frequency domain features of the eye width ratio sequence and the frequency domain features of the gaze estimation direction sequence; splicing the frequency domain features of the eye width ratio sequence and the frequency domain features of the gaze estimation direction sequence to obtain spliced features; performing normalization processing on the spliced features; using a trained deep learning hierarchical multi-scale long short-term memory network to process the normalized spliced features to obtain an output result; obtaining a discretized category corresponding to the output result to obtain the state of the security inspector; and making corresponding reminders according to the state of the security inspector. The present invention can reduce the occurrence of misdetection, missed detection and other situations, and safeguard the safety of passengers to a greater extent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method for early fatigue detection of airport security X-ray machine security inspectors based on the frequency domain. Background Art

[0002] In 2018, the International Air Transport Association (IATA) predicted that, based on the comprehensive calculation of international and domestic passenger traffic volume, China will surpass the United States to become the world's largest aviation market around 2025, and the passenger flow of China's civil aviation market will reach 1.3 billion around 2035. Among them, airport security is an important checkpoint for aviation security work. As the executors of the work, airport security inspectors are responsible for checking whether dangerous items are carried in passengers' luggage. Their work tasks are heavy and boring, and it is a profession that is extremely prone to psychological and physiological fatigue. The Civil Aviation Safety Authority of Australia conducted a survey among security inspectors, and the results showed that: 77% of the security inspectors working on the front line had deteriorated fatigue levels in the past 3 - 5 years, 60% of the security inspectors worked more than 55 hours per week, and 22% of the personnel worked more than 9 hours per day. In addition, the contradiction between the increasing number of flights and the shortage of airport security personnel is becoming increasingly prominent, which is extremely likely to lead to security personnel working in fatigue, and fatigue is one of the important factors causing "wrong", "forgotten", and "omitted" inspections during the work of security inspectors.

[0003] Fatigue detection is an important issue. Currently, it has been applied in fields such as driving and workplaces, but there are few studies on the fatigue detection method for airport luggage X-ray security inspectors. Most of the current fatigue detection methods are aimed at driver fatigue detection and mainly focus on detecting extreme fatigue. There are few studies on early fatigue detection, and most of them focus on simulated fatigue research, where the signs of drowsiness are exaggerated and can be clearly seen.

[0004] Therefore, it is urgent to improve the defects existing in the prior art. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a method for early fatigue detection of airport security X-ray machine security inspectors based on the frequency domain. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0006] In a first aspect, the present invention provides a method for early fatigue detection of airport security X-ray machine security inspectors based on the frequency domain, including:

[0007] Obtain the monitoring video to be detected;

[0008] Read the images of the monitoring video to be detected frame by frame in sequence; obtain the eye width ratio of each frame of image to form an eye width ratio sequence; estimate the gaze direction of each frame of image to form a gaze estimation direction sequence;

[0009] Extract the frequency domain features of the eye width ratio sequence and the gaze estimation direction sequence respectively;

[0010] Splicing the frequency domain features of the eye width ratio sequence and the frequency domain features of the gaze estimation direction sequence to obtain splicing features;

[0011] Normalize the splicing features;

[0012] Use the trained deep learning hierarchical multi-scale long short-term memory network to process the normalized splicing features to obtain the output result; obtain the discretized category corresponding to the output result and obtain the security inspector status; and make corresponding reminders based on the security inspector status.

[0013] Beneficial effects of the present invention:

[0014] The present invention provides a frequency-domain-based method for detecting early fatigue of airport X-ray security inspectors. On the one hand, the method fuses the eye width ratio sequence and the gaze estimation direction sequence, so that the multimodal information has different expressions and observation angles, with some overlap or complementarity. By rationally processing the multimodal information and fusing the complementary information, richer feature information can be obtained, effectively improving the final recognition effect. On the other hand, considering that time domain features are inherently irreversible and too much information is lost during the feature expression process, resulting in limited time domain feature expression capabilities, the present invention uses the eye width ratio sequence and the gaze estimation direction sequence for fatigue detection based on frequency domain features, which can view the problem from an objective perspective. The reversibility of the original information enables the frequency domain features to reflect the essence of things or phenomena, and can retain the original information to a great extent. Thus, the obtained early fatigue detection method focuses on security inspectors, effectively improving the quality of security inspections, reducing the occurrence of false detections and missed detections, and maintaining passenger safety to a greater extent.

[0015] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flow chart of a frequency-domain-based method for detecting early fatigue of airport security X-ray inspectors provided by an embodiment of the present invention;

[0017] Figure 2 This is a schematic diagram of the positions of key points of the eyes provided by an embodiment of the present invention;

[0018] Figure 3 This is a schematic diagram of a two-layer bidirectional LSTM network provided by an embodiment of the present invention;

[0019] Figure 4 This is a schematic diagram of a division sequence provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.

[0021] Currently, airport X-ray security personnel must stare at computer screens for extended periods, often while seated, which can easily lead to physical and psychological fatigue. This makes traditional fatigue detection solutions based on body movements infeasible. Most existing fatigue detection solutions target the mid- to late-stage of fatigue, when noticeable changes in facial expression and body movements, such as drowsiness, yawning, and frequent eye closure, have already occurred. By this point, it's too late to detect fatigue.

[0022] In view of this, in order to solve the technical problem in the prior art that airport security X-ray machine inspectors may have incorrect, forgotten or missed inspections due to fatigue, the present invention proposes an early fatigue detection method for airport security X-ray machine inspectors based on frequency domain, which transforms the feature processing mode from the traditional spatial and time domain-based facial feature detection mode to a frequency domain-based facial feature detection mode, and performs frequency domain processing on the original facial feature signal to detect the fatigue state of airport X-ray machine security inspectors.

[0023] See Figure 1 As shown, Figure 1 This is a flow chart of a frequency-domain-based method for detecting early fatigue of X-ray inspectors for airport security inspections, provided by an embodiment of the present invention. The frequency-domain-based method for detecting early fatigue of X-ray inspectors for airport security inspections provided by the present invention includes:

[0024] S101: Obtain a surveillance video to be detected.

[0025] Specifically, in this embodiment, a surveillance video in front of the airport X-ray machine security inspector is obtained as the surveillance video to be detected.

[0026] S102 , dividing the surveillance video to be detected into frames and reading images in sequence; obtaining the eye width ratio of each frame image to form an eye width ratio sequence; estimating the gaze direction of each frame image to form a gaze estimation direction sequence.

[0027] Specifically, in this embodiment, the process of acquiring the eye width ratio sequence includes:

[0028] Divide the surveillance video to be detected into frames and read the images in sequence to obtain a series of images;

[0029] Preprocessing the series of images; for example, removing noise from the series of images, or performing normalization operations on the series of images;

[0030] Divide the preprocessed image into multiple units, calculate the gradient and direction of each unit, and project each unit into a histogram of different directions;

[0031] Normalize the histogram of each unit to reduce the influence of factors such as illumination and shadow, and obtain the corresponding feature descriptor;

[0032] Concatenate the feature descriptors of all units to obtain the feature vector of the facial image;

[0033] Use a cascade classifier to classify the feature vector of the facial image to obtain the position information of the facial key points. In this embodiment, the position information of 68 facial key points is obtained. It should be noted that a cascade classifier is usually composed of multiple weak classifiers, and the number of negative samples is gradually reduced in a cascaded manner to improve the detection speed and accuracy;

[0034] According to the position information of the eye key points included in the facial key points, calculate the eye width ratio of each frame of image to form an eye width ratio sequence; where the expression of the eye width ratio is:

[0035]

[0036] Among them, P1, P2, P3, P4, P5, and P6 are the position information of the eye key points.

[0037] Please refer to Figure 2 as shown in Figure 2 is a schematic diagram of the position of the eye key points provided by the embodiment of the present invention. In this embodiment, the eye position is located by 6 eye key points. In formula (1), the numerator represents the distance of the eye feature points in the vertical direction, and the denominator represents the distance of the eye feature points in the horizontal direction. Through calculation, the eye width ratio sequence E={e1, e2,..., e n} in the n-frame sequence images is obtained.

[0038] In this embodiment, the process of obtaining the gaze estimation direction sequence includes:

[0039] Read the images of the monitoring video to be detected frame by frame in sequence to obtain a series of images;

[0040] Use Densepose to detect the pose of the object in the sequence images, and use Faster-RCNN to detect the human body region from the sequence images; according to the obtained pose of the object and the detected human body region, use the trained neural network to perform chunking, and convert the chunking points into a heat map (IUV map) and a segmentation map (IDNS map);

[0041] Use a head detector to obtain the human head bounding box image in the segmentation map;

[0042] Please refer to Figure 3 as shown in Figure 3It is a schematic diagram of the two - layer bidirectional LSTM network provided by an embodiment of the present invention. The trained ResNet - 18 network is used to process 7 consecutive frames of human head bounding box images to obtain 256 - dimensional facial features. The two - layer bidirectional LSTM network is used to learn and mine the forward and backward sequence patterns of the facial features. Finally, the vectors are concatenated and mapped through a fully - connected layer to obtain the estimated gaze direction for each frame, forming a sequence of estimated gaze directions.

[0043] In this embodiment, through estimation, the sequence of estimated gaze directions G = {g1, g2, …, g n} is obtained in the n - frame sequence image.

[0044] It should be noted that during the training process of the two - layer bidirectional LSTM network, the model is optimized to make the difference between the gaze direction and the true labels θ gt and φ gt become smaller, and the output f(I)=(θ, φ, σ) is obtained, where d=(θ, φ) is the predicted value of the gaze direction in the spherical coordinate system, where θ is the yaw angle, φ is the pitch angle, σ is the offset of the gaze direction prediction, τ is the target quantile, τ = 0.1 or τ = 0.9, and the loss function is defined as:

[0045]

[0046]

[0047] Among them, the expression for calculating the true value of the gaze direction is:

[0048]

[0049] S103. Extract the frequency - domain features of the eye - width ratio sequence and the frequency - domain features of the sequence of estimated gaze directions respectively.

[0050] Specifically, in this embodiment, extracting the frequency - domain features of the eye - width ratio sequence includes:

[0051] Please refer to Figure 4 as shown. Figure 4 It is a schematic diagram of dividing the sequence provided by an embodiment of the present invention. The eye - width ratio sequence is divided into multiple overlapping segments. Among them, the window wlen is set, and the window slides with the first step distance inc. The window slides once for each frame, and each frame corresponds to a segment. The segment is the time - domain feature x i (n).

[0052] According to the first step distance, the number of frames nf for dividing the eye - width ratio sequence is obtained, and its expression is:

[0053]

[0054] Among them, frame_num is the length of the eye width ratio sequence;

[0055] Perform Fourier transform on the time-domain feature x i (n) of the i-th frame to obtain the linear spectral feature x i (k) of the i-th frame, and its expression is:

[0056]

[0057] where n is the number of points of the n-th Fourier transform, N is the total number of points of the Fourier transform, e is the natural base, j is the imaginary unit, π is the pi, and k is the k-th linear sequence;

[0058] According to the linear spectral feature x i (k) of the i-th frame, calculate the power spectrum P i (k) of the i-th frame, and its expression is:

[0059]

[0060] Use multiple filter banks to process the power spectrum P i (k) of the i-th frame respectively to obtain the smoothed spectrum H m (k), and its expression is:

[0061]

[0062]

[0063] where m is the m-th group of filters, f(m) is the center frequency of the m-th group of filters, and f is the center frequency of the filter;

[0064] Multiply the linear spectral feature x i (k) of the i-th frame by the smoothed spectrum H m (k) obtained by each filter bank to obtain the logarithmic spectrum s(m), and its expression is:

[0065]

[0066] Accumulate the logarithmic spectra corresponding to multiple groups of filter banks and perform discrete cosine transform to obtain the Mel frequency cepstrum coefficients (FMCC coefficients), and its expression is:

[0067]

[0068] where M is the total number of filter banks, and L is the L-order Mel spectrum cepstrum coefficients;

[0069] Successively obtain the Mel frequency cepstrum coefficients corresponding to nf frames, construct the Mel frequency cepstrum coefficient matrix, and thus obtain the frequency-domain feature of the eye width ratio sequence.

[0070] In this embodiment, extracting the gaze estimation direction sequence includes:

[0071] Please continue to refer to Figure 4 As shown, divide the gaze estimation direction sequence into multiple overlapping segments; among them, set the window wlen, and the window slides with the first step distance inc, and the window slides once for each frame. Each frame corresponds to a segment, and the segment is the time-domain feature x i (n);

[0072] According to the first distance, obtain the number of frames nf for dividing the gaze estimation direction sequence, and its expression is:

[0073]

[0074] Among them, frame_num is the length of the gaze estimation direction sequence;

[0075] Perform Fourier transform on the time-domain feature x i (n) of the i-th frame to obtain the linear spectral feature x i (k) of the i-th frame, and its expression is:

[0076]

[0077] Among them, n is the number of points of the n-th Fourier transform, N is the total number of points of the Fourier transform, e is the natural base, j is the imaginary unit, π is the pi, and k is the k-th linear sequence;

[0078] According to the linear spectral feature x i (k) of the i-th frame, calculate the power spectrum P i (k) of the i-th frame, and its expression is:

[0079]

[0080] Use multiple filter banks to process the power spectrum P i (k) of the i-th frame respectively to obtain the smoothed spectrum H m (k), and its expression is:

[0081]

[0082]

[0083] Among them, m is the m-th group of filters, f(m) is the center frequency of the m-th group of filters, and f is the center frequency of the filter;

[0084] Compare the linear spectral feature x i (k) of the i-th frame with the smoothed spectrum H m(k) is multiplied to obtain the logarithmic spectrum s(m), which is expressed as:

[0085]

[0086] The logarithmic spectra corresponding to multiple filter groups are accumulated and subjected to discrete cosine transform to obtain the Mel-frequency cepstral coefficient (MFCC), which is expressed as follows:

[0087]

[0088] Where M is the total number of filter banks, L is the L-order Mel-spectrum cepstral coefficient;

[0089] The Mel-frequency cepstral coefficients corresponding to the nf frames are obtained in sequence, and a Mel-frequency cepstral coefficient matrix is constructed to obtain the frequency domain features of the gaze estimation direction sequence. Optionally, wavelet and CQCC can also be used for feature extraction.

[0090] It should be noted that this implementation also includes: if the amount of data in the last window cannot meet the amount of data corresponding to a complete window, the last window is padded with data, and the amount of padded data is:

[0091] numfill=nf×inc+wlen-frame_num.

[0092] Specifically, in this embodiment, considering that during the process of window segmentation processing in a sequence, the amount of data in the last window may not be able to meet the data amount corresponding to a window, therefore, corresponding data padding is performed for the case where the amount of data in the last window is insufficient, and 0 is used for padding.

[0093] It should be noted that the filter bank used in the above is a Mel frequency filter bank, which can also be understood as a group of triangular bandpass filters.

[0094] S104 : Splicing the frequency domain features of the eye width ratio sequence and the frequency domain features of the gaze estimation direction sequence to obtain splicing features, thereby achieving feature fusion.

[0095] S105: normalize the splicing features.

[0096] Specifically, in this embodiment, the expression for normalizing the splicing features is:

[0097]

[0098] Among them, μ p,q and σ p,q is the mean and standard deviation of the order q in the first third of the alert state videos of subject p, is the value of the p-th order q of the object after normalization, Nor_Mfcc p,q is the value of the p-th order q of the original object.

[0099] S106. Process the concatenated features after normalization using the trained deep learning hierarchical multi-scale long short-term memory network to obtain an output result; obtain the discretized category corresponding to the output result to obtain the state of the security inspector; and perform corresponding reminders according to the state of the security inspector.

[0100] Specifically, in this embodiment, the discretized categories include:

[0101]

[0102] where out is the output result.

[0103] Specifically, according to the output result, if the output result is a number between 0 and 10, corresponding to the discretized category, determine which state the security inspector is in. When the output result is between 0.0 and 3.3, including 0.0 and 3.3, the security inspector is in a sober state. When the output result is between 3.3 and 6.6, not including 3.3 but including 6.6, the security inspector is in the early fatigue state. When the output result is between 6.6 and 10, not including 6.6 but including 10, the security inspector is in the sleepy state; perform corresponding reminders according to the different states of the security inspector.

[0104] In summary, the present invention provides an early fatigue detection method for security inspectors of airport security X-ray machines based on the frequency domain. On the one hand, the eye width ratio sequence and the gaze estimation direction sequence are fused in terms of features, so that the multi-modal information has different expression methods and observation angles, and there are some overlapping or complementary phenomena. By reasonably processing the multi-modal information and fusing the complementary information, richer feature information can be obtained and the final recognition effect can be effectively improved. On the other hand, considering that the time-domain features are irreversible and too much information is lost in the feature expression process, resulting in limited time-domain feature expression ability, the present invention performs fatigue detection on the eye width ratio sequence and the gaze estimation direction sequence based on the frequency-domain features, which can view problems from an objective perspective. The reversibility of the original information enables the frequency-domain features to reflect the essence of things or phenomena and can retain the original information to a great extent. In this way, the obtained early fatigue detection method focuses on the security inspector, effectively improves the security inspection quality of the security inspector, reduces the occurrence of misdetection, missed detection, etc., and safeguards the safety of passengers to a greater extent.

[0105] It should be noted that, in this document, relational terms such as first and second are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that an article or device comprising a list of elements includes not only those elements but also other elements not explicitly listed. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of additional identical elements in the article or device comprising the element. Terms such as "connected" or "connected" are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. References to orientations or positional relationships, such as "upper," "lower," "left," and "right," are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate description and simplify the description of the present invention. They do not indicate or imply that the device or element referred to must have, be constructed, or operate in a specific orientation, and are therefore not to be construed as limiting the present invention.

[0106] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.

[0107] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, and all of these should be considered to fall within the scope of protection of the present invention.

Claims

1. An early fatigue detection method for airport security X-ray machine security inspectors based on the frequency domain, characterized in that, Including: Obtain the surveillance video to be detected; Read the images of the surveillance video to be detected frame by frame in sequence; Obtain the eye width ratio of each frame of image to form an eye width ratio sequence; Estimate the gaze direction of each frame of image to form a gaze estimation direction sequence; Extract the frequency domain features of the eye width ratio sequence and the frequency domain features of the gaze estimation direction sequence respectively; The extraction of the frequency domain features of the eye width ratio sequence includes: Divide the eye width ratio sequence into multiple overlapping segments; wherein, a window is set such that the window slides with a first step size , and the window slides once per frame. Each frame corresponds to a segment, and the segment is a time-domain feature ; Obtain the number of frames for dividing the eye width ratio sequence according to the first step distance , and its expression is: ; Among them, is the length of the eye width ratio sequence; Perform a Fourier transform on the time-domain characteristics of the frame to obtain the linear spectral characteristics of the frame, and its expression is: frame , and its expression is: ; Among them, is the number of points of the th Fourier transform, is the total number of points of the Fourier transform, is the natural base, is the imaginary unit, is pi, is the th linear sequence; According to the linear spectral characteristics of the th frame, calculate the power spectrum of the th frame, and its expression is: ; Use multiple filter banks to separately process the frame power spectrum to obtain a smoothed frequency spectrum , and its expression is: ; ; Among them, is the group of filters, is the center frequency of the group of filters, is the center frequency of the filter; Multiply the linear spectral features of the th frame by the smoothed spectrum obtained from each filter bank to obtain the log spectrum , and its expression is: ; Accumulate the logarithmic spectra corresponding to multiple groups of filter banks and perform discrete cosine transform to obtain the Mel frequency cepstral coefficients, and its expression is: ; Among them, is the total number of filter banks, is order Mel-frequency cepstral coefficients; Obtain successively the Mel-frequency cepstral coefficients corresponding to the frames, construct a Mel-frequency cepstral coefficient matrix, and thus obtain the frequency domain features of the eye width ratio sequence; Concatenate the frequency domain features of the eye width ratio sequence and the frequency domain features of the gaze estimation direction sequence to obtain concatenated features; Perform normalization processing on the concatenated features; Use the trained deep learning hierarchical multi-scale long short-term memory network to process the concatenated features after normalization processing to obtain an output result; obtain the discretized category corresponding to the output result to obtain the security inspector status; and give corresponding reminders according to the security inspector status.

2. The early fatigue detection method for security inspectors of airport security X-ray machines based on the frequency domain according to claim 1, characterized in that, The extraction of the gaze estimation direction sequence includes: Divide the sequence of gaze estimation directions into multiple overlapping segments; wherein, a window is set , and the window slides with a first step size . The window slides once per frame, and each frame corresponds to a segment, and the segment is a time-domain feature ; Obtain the number of frames for dividing the sequence of gaze estimation directions according to the first-step distance , and its expression is: ; Among them, is the length of the sequence of gaze estimation directions; Perform a Fourier transform on the time-domain characteristics of the frame to obtain the linear spectral characteristics of the frame, and its expression is: frame , and its expression is: ; Among them, is the number of points of the th Fourier transform, is the total number of points of the Fourier transform, is the natural base, is the imaginary unit, is the pi, is the th linear sequence; According to the linear spectral characteristics of the th frame, calculate the power spectrum of the th frame, and its expression is: ; Using multiple filter banks to process the frame power spectrum respectively, to obtain a smoothed spectrum , and its expression is: ; ; Among them, is the group of filters, is the center frequency of the group of filters, is the center frequency of the filter; Multiply the linear spectral features of the th frame with the smoothed spectrum obtained from each filter bank to obtain the log spectrum , whose expression is: ; Accumulate the logarithmic spectra corresponding to multiple groups of filter banks and perform discrete cosine transform to obtain the Mel frequency cepstral coefficients, and its expression is: ; Among them, is the total number of filter banks, is order Mel-frequency cepstral coefficients; Obtain successively the Mel-frequency cepstral coefficients corresponding to the frames, construct a Mel-frequency cepstral coefficient matrix, and thus obtain the frequency-domain features of the gaze estimation direction sequence.

3. The early fatigue detection method for airport security X-ray machine security inspectors based on the frequency domain according to claim 2, characterized in that Also including: For the data volume in the last window that cannot meet the data volume corresponding to a complete window, fill the data in the last window, and the number of filled data is: 。 4. The early fatigue detection method for airport security X-ray machine security inspectors based on the frequency domain according to claim 1, characterized in that, The expression for performing normalization processing on the concatenated features is: ; Among them, and are the average value and standard deviation of the order in the alert state video of the first third of the object . is the value of the order of the object after normalization . is the value of the order of the original object . . .

5. The early fatigue detection method for security inspectors of airport security X-ray machines based on the frequency domain according to claim 1, characterized in that The output result is a number between 0 and 10.

6. The early fatigue detection method for airport security X-ray machine security inspectors based on the frequency domain according to claim 1, characterized in that, The discretized categories include: ; Among them, is the output result.

7. The method for early fatigue detection of airport security X-ray machine security inspectors based on the frequency domain according to claim 1, characterized in that The acquisition process of the eye width ratio sequence includes: Read the images of the surveillance video to be detected frame by frame in sequence to obtain a series of images; Perform preprocessing on the series of images; Segment the preprocessed images into multiple units, calculate the gradient and direction of each unit, and project each unit into histograms in different directions; Perform normalization processing on the histogram of each unit to obtain the corresponding feature descriptor; Concatenate the feature descriptors of all units to obtain the feature vector of the facial image; Use a cascade classifier to classify the feature vector of the facial image to obtain the position information of the facial key points; According to the position information of the eye key points included in the facial key points, calculate the eye width ratio of each frame of image to form the eye width ratio sequence; where the expression of the eye width ratio is: ; Among them, , , , , and are the position information of the key points of the eye.

8. The method for early fatigue detection of airport security X-ray machine security inspectors based on the frequency domain according to claim 1, characterized in that, The acquisition process of the gaze estimation direction sequence includes: Read the images of the surveillance video to be detected frame by frame in sequence to obtain a series of images; Use Densepose to detect the pose of the object in the sequence of images, and use Faster-RCNN to detect the human body area from the sequence of images; according to the obtained pose of the object and the detected human body area, use the trained neural network to perform chunking, and convert the chunking points into heat maps and segmentation maps; Use a head detector to obtain the human head bounding box image in the segmentation map; Process the images of the human head bounding boxes for 7 consecutive frames using the trained ResNet-18 network to obtain 256-dimensional facial features; use a two-layer bidirectional LSTM network to learn and mine the sequential patterns in the forward and backward directions of the facial features, obtain the estimated gaze direction for each frame, and form a sequence of estimated gaze directions.

Citation Information

Patent Citations

  • Early fatigue detection method and system based on fine eye movement features

    CN112434611A

  • A navigation-attrite-controlled system and method for adjusting the cabin comfort settings of a vehicle

    DE102022127286A1