Edge fatigue detection method and device based on adaptive threshold adjustment and medium

Through the edge fatigue detection method with adaptive threshold adjustment, facial feature detection and dynamic threshold decision model are used to solve the problems of insufficient environmental adaptability and individual difference adaptability in existing technologies, realize efficient fatigue detection in complex driving scenarios, and improve detection accuracy and robustness.

CN120681147AActive Publication Date: 2025-09-23GUANGDONG TOBACCO HEYUAN CITY CO LTD

Patent Information

Application Number
CN202510695990.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-23
Estimated Expiration
2045-05-28

Smart Images

  • Figure CN120681147A_ABST
    Figure CN120681147A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent detection, in particular to an edge fatigue detection method and device based on adaptive threshold adjustment and a medium. The method comprises the following steps: acquiring facial image data and facial infrared data of a driver in a real-time state; preprocessing the face image data and the face infrared data; performing facial mark point positioning on the preprocessed facial image data and facial infrared data through a facial feature detection model to obtain an enhanced 3D mark point image; comparing the enhanced 3D mark point image with a physiological feature database to obtain physiological data; and calculating the physiological data through a dynamic threshold decision model and a multi-index fusion engine to obtain a fatigue decision. According to the method, the environmental adaptability can be improved, individualized precise adaptation is realized, precise adaptation is realized for individual differences, and the multi-working-condition robustness is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent detection technology, and in particular to an edge fatigue detection method, device and medium based on adaptive threshold adjustment. Background Art

[0002] With the rapid development of road traffic, fatigued driving has become a major cause of traffic accidents, posing a serious threat to public transportation safety. Existing technologies, such as NeuroSky's MindWave device, use bioelectrical signals such as electrocardiogram (ECG) and electroencephalogram (EEG) to assess fatigue. However, these devices require wearing specialized electrodes, which can affect driver comfort. Excessive sweating and motion artifacts can easily cause signal distortion, and they are unable to cope with electromagnetic interference in complex driving conditions. Volvo's Driver AlertControl system indirectly infers fatigue status using vehicle parameters such as steering wheel angle and lane departure distance. However, this system has a high false alarm rate for experienced drivers, and its sensitivity decreases by 42% in scenarios such as curves and traffic jams. The requirement for multi-sensor fusion increases system complexity by more than three times. The current mainstream solution uses the OpenCV+Dlib framework for fatigue assessment based on facial feature point detection. However, this framework relies on 68 facial landmarks and uses a fixed threshold to determine the PERCLOS metric. The processing speed on a single-core CPU is ≤20fps, and its accuracy drops to 63% under varying lighting conditions.

[0003] In summary, existing technologies have technical problems such as insufficient environmental adaptability, poor adaptability to individual differences, limited edge computing resources, and lack of robustness in multiple working conditions. These technical problems in related technologies need to be improved. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose an edge fatigue detection method, device and medium based on adaptive threshold adjustment, which can improve environmental adaptability, achieve individualized precise adaptation, accurately adapt to individual differences and enhance multi-working condition robustness.

[0005] To achieve the above objectives, an embodiment of the present application provides an edge fatigue detection method based on adaptive threshold adjustment, the method comprising:

[0006] Acquire the driver's real-time facial image data and facial infrared data;

[0007] Preprocessing the facial image data and the facial infrared data;

[0008] Performing facial landmark location on the pre-processed facial image data and the facial infrared data using a facial feature detection model to obtain an enhanced 3D landmark image;

[0009] Comparing the enhanced 3D landmark image with a physiological feature database to obtain physiological data;

[0010] The physiological data is calculated through a dynamic threshold decision model and a multi-index fusion engine to obtain a fatigue decision.

[0011] In some embodiments, the facial feature detection model includes a contrast-limited adaptive histogram equalization model and a MediaPipe face detection model.

[0012] In some embodiments, the data processing process of the contrast-limited adaptive histogram equalization model includes the following steps:

[0013] Dividing the pre-processed facial image data and the facial infrared data into a plurality of non-overlapping small blocks;

[0014] Calculating the grayscale value of each pixel in the small image block to obtain a histogram of each small image block;

[0015] Clipping the histogram according to a pixel threshold to obtain a contrast-limited histogram;

[0016] Normalizing the clipped contrast-limited histogram, calculating the cumulative distribution function of each gray level, and obtaining a new gray level mapping relationship;

[0017] Smoothing the adjacent small image blocks by overlapping blocks or fusing information of the contrast-limited histograms of adjacent areas during interpolation;

[0018] By using the new grayscale mapping relationship, the grayscale value of each pixel in the small image block is mapped back to the original image coordinate system to obtain an enhanced image.

[0019] In some embodiments, the data processing process of the MediaPipe face detection model includes the following steps:

[0020] Performing multi-scale feature extraction on the enhanced image through a lightweight convolutional neural network to obtain a real-time feature map;

[0021] performing accelerated processing on the real-time feature map according to heterogeneous acceleration hardware;

[0022] The real-time feature map is predicted through a graph convolutional network to obtain an enhanced 3D landmark image.

[0023] In some embodiments, the process of constructing the physiological characteristics database includes the following steps:

[0024] The probability density of the historical enhanced 3D landmark image data obtained after the historical data processing is estimated by using the Gaussian mixture model to obtain the basic Gaussian mixture model parameters;

[0025] Incrementally training the Gaussian mixture model by continuously inputting new historical enhanced 3D marker point image data, and updating parameters of the Gaussian mixture model until the confidence reaches a threshold;

[0026] When the confidence reaches the threshold, the normal physiological range boundary is calculated based on the Gaussian mixture model to obtain a physiological feature threshold table;

[0027] By performing feature fusion on the physiological feature threshold table and the attached physiological information through association analysis, a multi-dimensional physiological feature vector is obtained;

[0028] The multi-dimensional physiological feature vectors are stored according to the spatiotemporal index structure to obtain the physiological feature database.

[0029] In some embodiments, data validity verification is also included, including the following steps:

[0030] Acquiring the historical enhanced 3D marker point image data;

[0031] Fatigue-free verification of the physiological data is performed through fatigue probability and multi-dimensional cross-validation;

[0032] When the verification is passed, the physiological data is updated in the physiological characteristics database;

[0033] When the verification fails, an exception process is triggered, the collection time of the physiological data is extended and the fatigue-free verification is repeated. If it still fails, a review request is sent to the customer service end.

[0034] In some embodiments, the dynamic threshold decision model includes a dual-layer dynamic threshold and a Kalman filter dynamic adjustment; the calculation of the physiological data using the dynamic threshold decision model and the multi-index fusion engine to obtain a fatigue decision includes the following steps:

[0035] Obtaining a first fatigue judgment by comparing the physiological data with the dual-layer dynamic threshold;

[0036] Smoothing the first fatigue judgment through the Kalman filter dynamic adjustment to obtain a second fatigue judgment;

[0037] performing preprocessing on the second fatigue judgment;

[0038] The pre-processed second fatigue judgment is subjected to an “OR” logical combination and a dynamic weight adjustment to obtain the fatigue decision.

[0039] In some embodiments, an edge computing optimization model is further included, wherein the edge computing optimization model includes a lightweight neural network architecture, 4-bit hybrid quantization, spatial dynamic ROI focusing, and a parallel processing pipeline;

[0040] The lightweight neural network architecture is used to compress the volume of each model;

[0041] The 4-bit hybrid quantization is used to reduce memory usage and computational complexity;

[0042] The spatial dynamic ROI focusing is used to dynamically adjust the detection area according to the position of the face;

[0043] The parallel processing pipeline is used to implement overlapping execution of multiple levels of tasks.

[0044] To achieve the above objectives, another aspect of the present application provides an edge fatigue detection device based on adaptive threshold adjustment, the device comprising:

[0045] A data acquisition module, the data acquisition module is used to obtain facial image data and facial infrared data of the driver in real time;

[0046] a data processing module, the data processing module being configured to pre-process the facial image data and the facial infrared data, and locate facial landmarks on the pre-processed facial image data and the facial infrared data using a facial feature detection model to obtain an enhanced 3D landmark image;

[0047] A data decision module is used to compare the enhanced 3D landmark point image with the physiological feature database to obtain physiological data, and calculate the physiological data through a dynamic threshold decision model and a multi-index fusion engine to obtain a fatigue decision.

[0048] To achieve the above objectives, another aspect of an embodiment of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements an edge fatigue detection method based on adaptive threshold adjustment.

[0049] The embodiments of the present application include at least the following beneficial effects: The present application provides a method, device, and medium for edge fatigue detection based on adaptive threshold adjustment. This solution obtains facial image data and facial infrared data of the driver in real time, and locates facial landmarks on the pre-processed data using a facial feature detection model to obtain an enhanced 3D landmark image. The facial feature detection model greatly improves the adaptability to the environment in which the data is acquired, and solves the problem of detection failure of traditional methods in extreme lighting, head posture changes, or occlusion scenarios. Finally, a dynamic threshold decision model and a multi-index fusion engine are used to calculate the data obtained by comparing the enhanced 3D landmark image with the physiological feature database to obtain a fatigue decision. The dynamic threshold decision model achieves precise adaptation to individual differences, eliminates the problem of false detection caused by fixed thresholds, and adapts to differences in drivers' physiological characteristics. The multi-index fusion engine enhances the robustness of this solution in multiple working conditions and solves the problems of false alarms and missed alarms in complex driving scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a flow chart of an edge fatigue detection method based on adaptive threshold adjustment provided by an embodiment of the present application;

[0051] Figure 2 It is a flow chart of data processing of contrast-limited adaptive histogram equalization model;

[0052] Figure 3 It is an algorithm diagram of the contrast-limited adaptive histogram equalization model enhancement process;

[0053] Figure 4 This is a flowchart of the data processing of the MediaPi pe face detection model;

[0054] Figure 5 It is a flow chart of the process of building a physiological characteristics database;

[0055] Figure 6 is a schematic diagram of the source code of the physiological characteristic threshold table;

[0056] Figure 7 This is a schematic diagram of the source code of fatigue probability;

[0057] Figure 8 It is a structural diagram of an edge fatigue detection device based on adaptive threshold adjustment provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are merely examples of devices and methods consistent with some aspects of the embodiments of the present application.

[0059] It will be understood that the terms "first", "second", etc. used in this application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0060] The terms "at least one", "plurality", "each", "any", etc. used in this application include "at least one", "two" or more, "plurality" or "each", "any" or "any one", "each" or "any one" in the context of the present invention, and "at least one" or "at least one" includes one, two or more, "plurality" or "any one" includes two or more, "each" or "each one" in the context of the present invention, and "any" or "any one

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0062] Before explaining the embodiments of the present application in detail, some of the nouns and terms involved in the embodiments of the present application are first explained. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0063] ECR: Eye Closure Rate, an eye movement-related indicator, measures the frequency or duration of eye closure, reflecting conditions such as eye muscle relaxation and abnormal blinking during fatigue, such as excessive continuous eye closure for a long time.

[0064] MOR: Mouth Open Rate, a mouth-related indicator, monitors the frequency of mouth opening and closing or the duration of mouth opening. When fatigue occurs, yawning, jaw relaxation, etc. may occur, leading to abnormal mouth status.

[0065] HNFR: Head Nodding / Frequency Rate, a head posture-related indicator, tracks the frequency of head nodding, tilt angle, or shaking amplitude. When fatigue occurs, the head control ability decreases, resulting in frequent nodding, drooping, and other postures.

[0066] CLAHE: Contrast-limited adaptive histogram equalization model, which avoids excessive noise enhancement by limiting the local contrast amplification factor (such as setting the threshold to 40), processes the image in blocks (such as 8×8 pixel blocks), and enhances details in low-light areas. It is suitable for illumination normalization of facial images.

[0067] MobileNetV3: A lightweight convolutional neural network that combines depthwise separable convolution with the SE attention mechanism. It optimizes the structure through NAS search, significantly reducing the amount of computation while ensuring accuracy. It is often used for real-time facial feature extraction tasks.

[0068] MediaPipe Face Mesh model: MediaPipe face detection model, Google's open source real-time facial mesh rendering solution, based on binocular vision or monocular camera, can detect 468 facial landmarks and achieve high-precision facial pose and expression analysis.

[0069] AdaptiveBaseline: Dynamic baseline correction class, which uses a Gaussian mixture model (GMM) to learn the distribution of individual physiological characteristics and combines it with a confidence mechanism to dynamically update the baseline to adapt to the fatigue detection benchmark of different drivers.

[0070] partial_fit: An incremental learning method used to update Gaussian mixture model parameters online. It directly fits new data when the confidence level is low, and updates the mean through a sliding window smoothing when the confidence level is high, improving model stability.

[0071] EAR: Eye aspect ratio, an indicator that measures the degree of eye openness. The smaller the value, the longer the eyes are closed. It is one of the core visual features of fatigue detection.

[0072] MAR: Mouth opening and closing ratio, an indicator to monitor the degree of mouth opening. A larger value may indicate yawning or jaw relaxation, reflecting a state of nervous fatigue.

[0073] reset: Reset method, used to clear baseline model parameters and confidence, adapt to user switching or environmental mutation scenarios, and reinitialize the fatigue detection benchmark.

[0074] In related technologies, a technical report from the Society of Automotive Engineers (SAE) International indicates that traditional vision solutions have a facial feature extraction completeness rate of less than 45% in strong backlight or nighttime scenes with a contrast ratio of less than 0.3, resulting in a false detection rate of up to 28% (SAE Technical Paper 2021-01-0732). Commonly used fixed threshold strategies cannot adapt to the physiological differences of drivers. A 2019 article published in a leading journal in computer graphics noted that individuals with ptosis have a baseline EAR value 23% lower than that of normal individuals. Myopia also results in reduced pupil dynamic range, leading to MAR detection failures. Furthermore, drivers in specialized occupations, such as long-distance truckers, experience baseline head posture deviations of up to 18°±5° (IEEE TVCG 2019). When running traditional solutions on automotive-grade embedded platforms like the Qualcomm Snapdragon 8155 chip, edge computing resources are limited, with peak memory usage reaching 512MB, three times that of typical in-vehicle terminal configurations. Single-frame processing latency is ≥80ms, making it difficult to meet 100ms real-time requirements. Furthermore, the floating-point computing workload exceeds the computing power limit of the ARM Cortex-A76 core. The 28th World Congress on Intelligent Transport Systems held in 2022 pointed out that in complex scenarios with coverage ≥30%, such as mask occlusion or rapid head rotation with an angular velocity ≥30° / s, the comprehensive detection accuracy of existing solutions drops to 58% (ITS World Congress2022 data).

[0075] In view of this, an edge fatigue detection method based on adaptive threshold adjustment is provided in an embodiment of the present application, which relates to the field of intelligent detection technology. The edge fatigue detection method based on adaptive threshold adjustment provided in an embodiment of the present application can be applied to a terminal, can be applied to a server, or can be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network; the software can be an application that implements an edge fatigue detection method based on adaptive threshold adjustment, etc., but is not limited to the above forms.

[0076] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0077] Figure 1 This is an optional flow chart of an edge fatigue detection method based on adaptive threshold adjustment provided in an embodiment of the present application. Figure 1 The method may include but is not limited to steps S110 to S150:

[0078] Step S110: Acquire facial image data and facial infrared data of the driver in real time;

[0079] Step S120: pre-processing the facial image data and facial infrared data;

[0080] Step S130: locating facial landmarks on the pre-processed facial image data and facial infrared data using a facial feature detection model to obtain an enhanced 3D landmark image;

[0081] Step S140: Comparing the enhanced 3D landmark image with the physiological feature database to obtain physiological data;

[0082] Step S150: Calculate the physiological data through the dynamic threshold decision model and the multi-index fusion engine to obtain a fatigue decision.

[0083] In some embodiments, sensors capture real-time facial image data and infrared facial data of the driver in the cockpit. The captured data is preprocessed and processed using a facial feature detection model. This data is then processed to ensure clarity and accurately reflect the driver's facial landmarks for fatigue detection, resulting in an enhanced 3D landmark image. This facial feature detection model addresses the detection failures of traditional solutions in extreme lighting conditions, head posture variations, and occlusion. It enhances local contrast in low-light environments, improving accuracy from 37.5% to 82.6%. It also supports precise tracking within a ±50° head posture angle. Combining infrared thermal imaging with visible light images, the detection rate remains at 76.4% even in mask occlusion. The enhanced 3D landmark image is then compared with a pre-trained physiological feature database to generate physiological data for fatigue decision-making. The physiological feature database integrates three dimensional metrics: ECR, MOR, and HNFR, effectively addressing false positives and negatives in complex driving scenarios. A dynamic threshold decision model and a multi-metric fusion engine are used to calculate the physiological data and generate fatigue decisions. Fatigue decision directly reflects whether the driver's current state is fatigue driving, so as to ensure the driver's road safety. The dynamic threshold decision model eliminates the false detection problem caused by fixed thresholds and adapts to the differences in drivers' physiological characteristics. In some embodiments, in step S110, a camera and an infrared thermal imaging module are set at the driver's seat. Among them, the camera is a visible light camera with a resolution of ≥1080P@30fps to obtain camera RGB data, that is, facial image data. The frame rate of the infrared thermal imaging module is ≥25fps to obtain infrared thermal imaging data, that is, facial infrared data. The installation position and angle must ensure that the driver's facial image can be obtained at all times.

[0084] In some embodiments, in step S120 , the facial image data and the facial infrared data are preprocessed, where the preprocessing includes time synchronization and spatial alignment.

[0085] Specifically, facial image data and facial infrared data are accurately timestamped and aligned using a dynamic time warping algorithm. Facial landmark detection is used to locate reference points such as the tip of the nose and the corners of the eyes. Affine transformation is then used to unify the coordinates of the facial image data and facial infrared data, ensuring consistent feature space positions.

[0086] In some embodiments, in step S130, the facial feature detection model includes a contrast-limited adaptive histogram equalization model and a MediaPipe face detection model. Among them, the adaptive histogram equalization model is CLAHE preprocessing enhancement, which is used to perform histogram equalization on the input image blocks to improve the contrast in dark environments. Experiments show that CLAHE preprocessing enhancement improves the accuracy of feature point detection in low-light scenes by 45%. The MediaPipe face detection model is the MediaPipeFace Mesh model, which extracts features through a lightweight convolutional neural network backbone network and combines a graph convolutional network to predict 3D facial landmarks. In this way, sub-millimeter facial landmark positioning can be achieved under complex lighting and posture conditions. The core formula is:

[0087] Q = GraphConvolution(f MobileNetV3 (I));

[0088] Among them, I represents the input image, f() represents the feature extraction function composed of the MobileNetV3 network, GraphConvolution() represents the graph convolution operation, and Q∈R468×3 represents the three-dimensional key point coordinates.

[0089] In some embodiments, as Figure 2 As shown, Figure 2 It is a flowchart of the data processing of the contrast-limited adaptive histogram equalization model. Figure 2 The method may include but is not limited to steps S131 to S136:

[0090] Step S131: Divide the pre-processed facial image data and facial infrared data into a plurality of non-overlapping small blocks;

[0091] Step S132: Calculate the grayscale value of each pixel in each small image block to obtain a histogram of each small image block;

[0092] Step S133: clipping the histogram according to the pixel threshold to obtain a contrast-limited histogram;

[0093] Step S134: normalizing the clipped contrast-limited histogram, calculating the cumulative distribution function of each gray level, and obtaining a new gray level mapping relationship;

[0094] Step S135: smoothing the adjacent small blocks by overlapping the blocks or fusing the information of the contrast-limited histogram of the adjacent areas during the interpolation process;

[0095] Step S136: Map the pixel grayscale values ​​in each small image block back to the original image coordinate system through the new grayscale mapping relationship to obtain an enhanced image.

[0096] Specifically, if Figure 3 As shown, Figure 3 This is the algorithm flow chart for the contrast-limited adaptive histogram equalization model enhancement process. Based on the algorithm flow chart, the following process is obtained. The input image is divided into several non-overlapping small blocks, typically rectangular areas of 8×8 or 16×16 pixels. Blocking helps to locally adjust the brightness distribution of different image regions, avoiding the damage to local details caused by global histogram equalization. A histogram is calculated for the pixel grayscale values ​​within each block, and the number of pixels at each grayscale level is counted. The histogram range is typically limited to 0-255 (8-bit images), but can be adjusted according to actual needs to obtain a histogram for each small block. The histogram is clipped to limit the maximum contrast enhancement within each block. If the number of pixels at a certain grayscale level exceeds a preset threshold, the excess pixels are evenly distributed to other grayscale levels to prevent overexposure or noise amplification, resulting in a contrast-limited histogram. The contrast-limited histogram is normalized, and pixel values ​​are redistributed to expand local contrast. The original grayscale values ​​are mapped to a new grayscale range using the cumulative distribution function (CDF), enhancing details in dark or bright areas and generating a new grayscale mapping. To reduce the blocking artifacts caused by block division, the histograms of adjacent blocks are smoothed. Common methods include overlapping blocks or fusing histogram information from adjacent regions during interpolation. The equalized pixel values ​​of all blocks are remapped back to the original image coordinate system, and all blocks are merged to generate the final enhanced image, preserving local contrast while maintaining overall visual consistency.

[0097] In some embodiments, as Figure 4 As shown, Figure 4 This is a flowchart of MediaPipe face detection model data processing. Figure 4 The method may include but is not limited to steps S137 to S139:

[0098] Step S137: performing multi-scale feature extraction on the enhanced image through a lightweight convolutional neural network to obtain a real-time feature map;

[0099] Step S138: Accelerate the real-time feature map using heterogeneous acceleration hardware;

[0100] Step S139: Predict the real-time feature map through the graph convolutional network to obtain an enhanced 3D landmark point image.

[0101] In some embodiments, in step S137, a 3-channel color image with a resolution of 192*192 pixels, i.e., an enhanced image, is input as the original input data of the model to provide visual information for subsequent feature extraction. By reducing the amount of computation or adapting the hardware acceleration fixed resolution, the input size can be unified to facilitate efficient processing by the neural network. Among them, the channels correspond to red, green, and blue pixel values, which are used to capture the color and texture features of the image. Feature extraction is performed on the input image using a lightweight convolutional neural network MobileNetV3, which is used to extract basic features such as edges and corners in the image and high-level semantic features such as facial contours and organ structures through convolution layers, activation functions, and pooling layers. Through optimization methods such as deep separable convolution and bottleneck structure, the feature expression capability is maintained while reducing the amount of computation to adapt to the real-time reasoning requirements of mobile terminals or embedded devices. A feature map containing spatial position and semantic information, i.e., a real-time feature map, is generated for further processing by subsequent network layers.

[0102] In some embodiments, in step S138, model reasoning is accelerated by heterogeneous computing hardware, such as a graphics processor GPU, a digital signal processor DSP, and a neural network processor NPU. Computing tasks are allocated according to the hardware characteristics of the device. For example, the GPU is good at parallel computing, and the NPU is optimized for neural networks. The model running speed is greatly improved, and real-time processing with millisecond-level response is achieved. Avoiding the performance bottleneck of a single CPU processing, reducing power consumption, and extending device life are especially important for mobile devices such as mobile phones and AR glasses. For cross-platform compatibility, different hardware architectures are supported so that the model can be deployed on multiple devices, such as mobile phones, computers, edge computing devices, etc.

[0103] In some embodiments, in step S139, 468 3D key points are output. Specifically, the feature information is processed by a graph convolutional network to predict the three-dimensional coordinates (X, Y, Z) of the 468 key points on the face. The facial key points are regarded as "nodes" of the graph, and the spatial relationship between the nodes is captured by the graph convolutional network, such as geometric constraints such as the distance between the eyes and the height of the bridge of the nose, to infer the three-dimensional structure of the face. Based on the 2D image features extracted by MobileNetV3, that is, the feature map is implemented, and the spatial reasoning capability of the graph convolutional network is combined to generate 3D key point coordinates with depth information to achieve three-dimensional modeling of the facial mesh. Outputting 468 3D key points can accurately describe details such as facial contours, corners of the eyes, corners of the mouth, etc., that is, enhance the 3D landmark image. The MediaPipe face detection model combines the feature extraction capability of the convolutional neural network and the structural reasoning capability of the graph convolutional network to achieve efficient and real-time facial 3D key point detection through hardware acceleration.

[0104] In some embodiments, as Figure 5 As shown, Figure 5 It is a flowchart of the process of building a physiological feature database.

[0105] Figure 5 The method may include but is not limited to steps S210 to S250:

[0106] Step S210: performing probability density estimation on the historical enhanced 3D landmark image data obtained after the historical data processing using a Gaussian mixture model to obtain basic Gaussian mixture model parameters;

[0107] Step S220: incrementally training the Gaussian mixture model by continuously inputting new historical enhanced 3D landmark image data, and updating the parameters of the Gaussian mixture model until the confidence reaches a threshold;

[0108] Step S230: When the confidence reaches the threshold, the normal physiological range boundary is calculated based on the Gaussian mixture model to obtain a physiological feature threshold table;

[0109] Step S240: performing feature fusion on the physiological feature threshold table and the attached physiological information through association analysis to obtain a multi-dimensional physiological feature vector;

[0110] Step S250: storing the multi-dimensional physiological feature vectors according to the spatiotemporal index structure to obtain a physiological feature database.

[0111] In some embodiments, in steps S210 to S250, as Figure 6 As shown, Figure 6 This is a schematic diagram of the source code for the physiological feature threshold table. Specifically, facial image data and facial infrared data obtained from historical data are preprocessed, and the CLAHE algorithm is used to enhance contrast and eliminate illumination and temperature noise. MobileNetV3 is used to extract facial texture features, and MediaPipe Face Mesh is used to locate 468 key points (such as the corners of the eyes, the tip of the nose, and the mandible) to obtain historically enhanced 3D landmark image data. Core metrics such as EAR (eye movement aspect ratio), MAR (mouth opening / closing ratio), and HNFR (head gesture frequency) are calculated from the key point coordinates.

[0112] Eigenvalues ​​such as EAR and MAR, along with driver information, are extracted, and physiological signals such as heart rate and steering wheel grip force are simultaneously collected. These eigenvalues ​​are standardized and then input into a Gaussian mixture model. The model is initialized with three Gaussian components, corresponding to states of alertness, mild fatigue, and severe fatigue. An expectation-maximization algorithm is used to iteratively estimate the mean vector, covariance matrix, and weight coefficients of each component to form the basic model parameters.

[0113] New historically enhanced 3D landmark image data is continuously collected and, after feature extraction, fed into a Gaussian mixture model for incremental training. Each time the partial_fit method is called, the system checks the current confidence level. If it is below 0.9, the Gaussian mixture model parameters are directly updated and the confidence level is increased. If it reaches 0.9, a sliding window mean filter mechanism is used, retaining the historical mean with a weight of 0.9 and incorporating the new data mean with a weight of 0.1 to prevent overfitting due to short-term abnormal fluctuations. Once the model parameters converge and the confidence level stabilizes, the normal range boundaries of the physiological characteristics are calculated based on the mean ±2.5 times the standard deviation of each Gaussian component to generate a table of physiological characteristic thresholds.

[0114] The historical physiological feature threshold table is sorted by timestamp to form a queue. The thresholds in the queue are weighted and aggregated using exponentially decaying weights: the most recent threshold is assigned the highest weight, while older thresholds have exponentially decreasing weights. This mechanism enables the threshold table to dynamically track slow changes in the driver's physiological characteristics, such as increased fatigue tolerance, while maintaining sensitivity to sudden anomalies. Correlation analysis is performed on features such as EAR, MAR, heart rate variability, and blink rate, and correlation coefficients are calculated between them. Strongly correlated features are selected to construct a multidimensional feature vector space. For example, EAR, MAR, and HRV are combined into a three-dimensional vector with a weighting ratio of 4:3:3. The threshold for each dimension is derived from the physiological feature threshold table generated in step S230. Z-score normalization is performed on each dimension to eliminate dimensional differences, forming a unified feature representation space. This allows data from different modalities to be compared and matched in the same space. A spatial index is used for the physiological feature database, creating a KD tree index on the spatial coordinates of 3D landmarks. This supports fast spatial neighborhood queries and is used to locate historical data most similar to the current posture. Through a three-level index structure, the physiological feature vectors corresponding to the enhanced 3D landmark point images are organized into an efficient and queryable database.

[0115] The feature vector is extracted from the 3D landmark image obtained by real-time acquisition and processing, and the Mahalanobis distance between it and the nearest N vectors in the database is calculated. If the distance is less than 2.0, it is judged as "awake"; between 2.0-3.0 is "fatigue"; and if it exceeds 3.0, an "abnormal" warning is triggered. At the same time, the system will return the feature dimension with the largest contribution to the abnormality (such as "EAR deviates from the baseline by 2.8σ") and mark the corresponding 3D landmark position (such as the key point of the corner of the eye). The comparison results are stored in the result table together with metadata such as timestamp and confidence level to form a complete record of physiological state changes, which supports subsequent analysis and model optimization.

[0116] In some embodiments, a redundant verification process is also included in the abnormal state self-check mechanism, which is triggered when the detection results of five consecutive frames are inconsistent. Specifically, the statistical distance between the current physiological characteristics and the baseline characteristics is calculated through Mahalanobis distance monitoring. Gradient limitation is implemented by setting a threshold of maximum gradient change ≤ 5% to prevent parameter mutations. When the Mahalanobis distance exceeds 2.5, the threshold update is frozen and the manual review process is initiated. This mechanism ensures system robustness through both Mahalanobis distance and gradient limitation.

[0117] In some embodiments, when the camera and infrared thermal imaging module are damaged, causing sensor abnormalities, feature extraction failures, or data pollution causing data abnormalities, sudden physiological abnormalities, dangerous driving behaviors, or system abnormalities, resulting in abnormal judgments in a non-fatigue state, it is necessary to determine the specific situation through data validity verification. For data validity judgment, specifically: obtain the driver's enhanced 3D landmark image data through the camera and infrared thermal imaging module, and synchronously collect timing signals of devices such as steering wheel torque sensors or inertial measurement units. Extract visual features such as eye movement aspect ratio (EAR), mouth opening and closing ratio (MAR), head pitch or yaw angle from the image, extract facial skin temperature distribution (such as the temperature change trend around the eyes and nose wings) from the infrared data, and parse the time series of steering wheel torque fluctuations and head posture changes from the sensor data.

[0118] like Figure 7 As shown, Figure 7 This is a schematic diagram of the source code for fatigue probability. For the data collected in the first 30 minutes, the initial_validation function is called to calculate the fatigue probability score P f , respectively calculate the normalized values ​​of EAR and MAR (range 0-1, the larger the value, the higher the fatigue level), and the head posture abnormality index (such as the proportion of duration when the pitch angle exceeds 15°), and calculate the comprehensive score according to the weight formula. The weight of each index can be determined according to the pre-set setting. If P f >0.3, it is considered invalid data (indicating fatigue or abnormal interference may exist during the acquisition period); if P f If ≤0.3, enter multi-dimensional cross-validation.

[0119] When P fIf the EAR or MAR value is ≤0.3, the system enters multi-dimensional cross-validation. The correlation between periods of abnormal EAR or MAR (e.g., EAR <0.2) and facial temperature changes is analyzed. If an abnormal EAR is accompanied by a sudden drop in periocular temperature (possibly due to reduced heat dissipation from closing the eyes), it is considered a valid fatigue signal; otherwise, it is considered interference (e.g., sudden change in illumination causing EAR miscalculation). The system checks whether a head dip (pitch angle >20°) coincides with a sudden drop in steering wheel torque (e.g., grip force <5N). If this occurs simultaneously, it may reflect fatigue-induced decreased steering ability. If the head posture is abnormal but the torque is normal (e.g., the driver adjusts their sitting position), the abnormality is considered non-fatigue. If the multi-dimensional cross-validation finds no evidence of fatigue, such as if the visual abnormality is caused by environmental interference, the validation passes. The current physiological data (EAR, MAR, temperature, torque, etc.) are added to the physiological feature database, triggering a partial_fit update of the baseline model. If the validation fails, the acquisition time is increased exponentially using a backoff strategy, and the fatigue probability calculation and cross-validation are repeated. When the cumulative effective data percentage (effective duration / total collection duration) is ≥70%, the data quality is considered to be qualified and the baseline modeling process (such as Gaussian mixture model initialization) is initiated. If the effective percentage is still <70% after three consecutive extensions, a manual review request (with raw data fragments, characteristic curves, and verification results) is sent to the customer service end, and the data validity is manually annotated with factors such as "equipment failure" or "abnormal driver behavior." Understandably, if the driver is fatigued during the initialization phase (such as driving in the early morning or working all night), this will lead to baseline data contamination, and the subsequent dynamic threshold calculation will be completely ineffective.

[0120] When the verification process is passed and the accumulated valid data meets the modeling requirements (such as ≥1 hour of continuous fatigue-free data), the Gaussian mixture model is initialized based on the cleaned feature data, and three Gaussian components are set to fit different sub-scenarios of the awake state (such as the baseline difference between daytime / nighttime driving). The mean, covariance and other parameters are iteratively calculated through the expectation maximization algorithm, and finally an individual initial baseline threshold table is generated (such as the normal range of EAR [0.3, 0.5]).

[0121] In some embodiments, in step S140. The 3D landmark point image is enhanced to calculate geometric eigenvalues ​​such as the eye movement aspect ratio (EAR), mouth opening and closing ratio (MAR), and head posture angle, and the temperature distribution data of the infrared thermal imaging is simultaneously extracted to form a multidimensional feature vector. The feature vector is input into the physiological feature database, and the matching historical data is quickly retrieved based on the spatiotemporal index structure of the database. By calculating the Mahalanobis distance between the real-time feature vector and the baseline model parameters in the database (such as the mean and covariance of the Gaussian mixture model), the deviation degree data of each physiological indicator is obtained, such as the difference between the EAR and the baseline mean, the fluctuation range of the head posture angle, etc. This process only completes the data comparison and generates physiological data containing eigenvalues, deviations, and timestamps.

[0122] In some embodiments, in step S150, the real-time physiological data obtained from the physiological characteristics database is compared with the dual-layer dynamic threshold stored in the physiological characteristics database. The specific formula of the dual-layer dynamic threshold is:

[0123] T adaptive =α·T global +(1-α)·(μ personal +2.5σ personal );

[0124] Among them, α represents the global threshold weight, μ personal represents the individual baseline threshold, σ personal is the standard deviation, T global represents the global threshold, T adaptive Indicates a dynamic threshold.

[0125] A global threshold is a baseline standard applied to all drivers, while an individual baseline threshold is calculated based on the parameters of a Gaussian mixture model trained using historical driver data. For each physiological indicator, if real-time data exceeds both the global and individual thresholds, the indicator is labeled "abnormal." If only the global threshold is exceeded, the indicator is labeled "suspicious." If neither threshold is exceeded, the indicator is labeled "normal." Ultimately, the labeled results of all indicators are integrated to form a preliminary fatigue assessment, providing a preliminary assessment of whether the driver is currently fatigued.

[0126] The results of the first fatigue judgment are smoothed using a Kalman filter to eliminate short-term fluctuations and misjudgments of the data. The Kalman filter regards the fatigue judgment results as state variables of a dynamic system. By establishing a state transition model and an observation model, it predicts the fatigue state at the next moment and makes corrections based on the real-time observation data (i.e., the first fatigue judgment results). For example, if the EAR indicator briefly triggers an anomaly at a certain moment, but the Kalman filter predicts that it is within the normal fluctuation range, the anomaly mark is corrected to "suspicious" or "normal." Through continuous iterative prediction and correction, a more stable and reliable second fatigue judgment is output, reducing false alarms caused by noise or accidental factors.

[0127] The second fatigue judgment result is normalized and standardized to prepare for the subsequent multi-index fusion. The pre-processed second fatigue judgment result is input into the multi-index fusion engine and integrated according to the three dimensions of eye movement (ECR), mouth position (MOR), and head posture (HNFR). The formula of the multi-index fusion engine is:

[0128]

[0129] The multi-indicator fusion engine determines whether to trigger a fatigue warning for each dimension (such as ECR) based on a preset threshold (such as ECR ≥ 50%). The logical "or" strategy is used to integrate the judgment results of the three dimensions, that is, as long as any one of the dimensions of ECR, MOR, and HNFR reaches the warning threshold, it is judged as a suspected fatigue state. In the experimental group scenario, the weight coefficients of ECR, MOR, and HNFR are dynamically adjusted based on the driver's baseline data (such as the correlation between historical fatigue status and each indicator). For example, if a driver is more likely to droop his head when tired (abnormal HNFR), the weight of HNFR is increased (such as from 30% to 40%). The logical "or" judgment result is combined with the weight to calculate the final fatigue score. If the comprehensive score ≥ the threshold, it is judged to be a fatigue state; otherwise, it is a non-fatigue state, that is, a fatigue decision. The multi-indicator fusion engine adopts a spatiotemporal optimization algorithm, through parallel computing and pipeline processing, such as decomposing the detection process into a three-level cache queue of preprocessing, feature extraction, and decision-making, as well as a lightweight model, to achieve a real-time inference speed of 54.3FPS, which is 37% higher than the accuracy of traditional single-indicator methods.

[0130] In some embodiments, an edge computing optimization model is also included for the overall architecture to overcome the computing power limitations of vehicle terminals and meet real-time requirements. The lightweight MobileNetV3 architecture reduces the compressed model size to 3.6MB, 5.8% of the original 60MB model. 4-bit hybrid quantization technology reduces computational complexity by 70% while maintaining 95% accuracy. A spatial dynamic ROI focusing algorithm reduces the face detection area from the full image to 42%. Parallel processing pipeline technology decomposes the detection process into a three-level cache queue: image preprocessing, feature extraction, and decision reasoning, boosting processing speed to 54.3FPS.

[0131] In some embodiments, the method can be deployed across platforms and implemented through the MediaPipe framework, model quantization technology, and standardized interface design. Specifically, the MediaPipe framework natively supports heterogeneous computing of GPU (OpenGL / Vulkan), DSP (Hexagon), and NPU (Ascend). The model quantization technology uses 4-bit hybrid quantization to reduce the model size by 54%. The standardized interface design provides a unified API interface, which is compatible with Android, iOS, and embedded systems. The measured data in the paper shows that 55ms end-to-end latency is achieved on the Snapdragon 8155 chip, verifying that cross-platform OTA updates rely on the following technologies: MobileNet architecture + hybrid quantization to compress the model from 62MB to 3.6MB, lightweight model, decoupling detection logic from the model, modular design that supports dynamic loading of new model files, and only transmitting model parameter increments to reduce the amount of OTA transmission data. Experiments show that after optimization, the GPU video memory usage is reduced to 186MB, which meets the differential upgrade conditions of the OTA update of the vehicle terminal.

[0132] like Figure 8 As shown, an edge fatigue detection device based on adaptive threshold adjustment includes:

[0133] A data acquisition module is used to obtain facial image data and facial infrared data of the driver in real time;

[0134] A data processing module is used to pre-process facial image data and facial infrared data, locate facial landmarks on the pre-processed facial image data and facial infrared data using a facial feature detection model, and obtain an enhanced 3D landmark image;

[0135] The data decision module is used to compare the enhanced 3D landmark point image with the physiological feature database to obtain physiological data, calculate the physiological data through the dynamic threshold decision model and multi-index fusion engine, and obtain fatigue decision.

[0136] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0137] The present application also provides an electronic device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned edge fatigue detection method based on adaptive threshold adjustment. The electronic device can be any smart terminal, such as a tablet computer or an in-vehicle computer.

[0138] It can be understood that the contents of the above method embodiments are applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0139] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned edge fatigue detection method based on adaptive threshold adjustment.

[0140] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiment, the functions specifically implemented by the present storage medium embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0141] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0142] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0143] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0144] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0145] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0146] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A method for edge fatigue detection based on adaptive threshold adjustment, characterized in that: The method comprises: Acquire the driver's real-time facial image data and facial infrared data; Preprocessing the facial image data and the facial infrared data; Performing facial landmark location on the pre-processed facial image data and the facial infrared data using a facial feature detection model to obtain an enhanced 3D landmark image; Comparing the enhanced 3D landmark image with a physiological feature database to obtain physiological data; The physiological data is calculated through a dynamic threshold decision model and a multi-index fusion engine to obtain a fatigue decision.

2. The method according to claim 1, characterized in that The facial feature detection model includes a contrast-limited adaptive histogram equalization model and a MediaPipe face detection model.

3. The method according to claim 2, characterized in that The data processing process of the contrast-limited adaptive histogram equalization model includes the following steps: Dividing the pre-processed facial image data and the facial infrared data into a plurality of non-overlapping small blocks; Calculating the grayscale value of each pixel in the small image block to obtain a histogram of each small image block; Clipping the histogram according to a pixel threshold to obtain a contrast-limited histogram; Normalizing the clipped contrast-limited histogram, calculating the cumulative distribution function of each gray level, and obtaining a new gray level mapping relationship; Smoothing the adjacent small image blocks by overlapping blocks or fusing information of the contrast-limited histograms of adjacent areas during interpolation; By using the new grayscale mapping relationship, the grayscale value of each pixel in the small image block is mapped back to the original image coordinate system to obtain an enhanced image.

4. The method according to claim 3, characterized in that The data processing process of the MediaPipe face detection model includes the following steps: Performing multi-scale feature extraction on the enhanced image through a lightweight convolutional neural network to obtain a real-time feature map; performing accelerated processing on the real-time feature map according to heterogeneous acceleration hardware; The real-time feature map is predicted through a graph convolutional network to obtain an enhanced 3D landmark image.

5. The method according to claim 1, wherein The process of constructing the physiological characteristics database includes the following steps: The probability density of the historical enhanced 3D landmark image data obtained after historical data preprocessing is estimated by using the Gaussian mixture model to obtain the basic Gaussian mixture model parameters; Incrementally training the Gaussian mixture model by continuously inputting new historical enhanced 3D marker point image data, and updating parameters of the Gaussian mixture model until the confidence reaches a threshold; When the confidence reaches the threshold, the normal physiological range boundary is calculated based on the Gaussian mixture model to obtain a physiological feature threshold table; By performing feature fusion on the physiological feature threshold table and the attached physiological information through association analysis, a multi-dimensional physiological feature vector is obtained; The multi-dimensional physiological feature vectors are stored according to the spatiotemporal index structure to obtain the physiological feature database.

6. The method according to claim 5, characterized in that It also includes data validity verification, including the following steps: Acquiring the historical enhanced 3D marker point image data; Fatigue-free verification of the physiological data is performed through fatigue probability and multi-dimensional cross-validation; When the verification is passed, the physiological data is updated in the physiological characteristics database; When the verification fails, an exception process is triggered, the collection time of the physiological data is extended and the fatigue-free verification is repeated. If it still fails, a review request is sent to the customer service end.

7. The method according to claim 1, characterized in that The dynamic threshold decision model includes a double-layer dynamic threshold and a Kalman filter dynamic adjustment; the dynamic threshold decision model and the multi-index fusion engine are used to calculate the physiological data to obtain fatigue decision, including the following steps: Obtaining a first fatigue judgment by comparing the physiological data with the dual-layer dynamic threshold; Smoothing the first fatigue judgment through the Kalman filter dynamic adjustment to obtain a second fatigue judgment; performing preprocessing on the second fatigue judgment; Performing an OR logic combination and a dynamic weight adjustment on the preprocessed second fatigue judgment to obtain the fatigue decision.

8. The method according to claim 1, characterized in that Also included is an edge computing optimization model comprising a lightweight neural network architecture, 4-bit hybrid quantization, spatial dynamic ROI focusing, and a parallel processing pipeline; The lightweight neural network architecture is used to compress the volume of each model; The 4-bit hybrid quantization is used to reduce memory usage and computational complexity; The spatial dynamic ROI focusing is used to dynamically adjust the detection area according to the position of the face; The parallel processing pipeline is used to implement overlapping execution of multiple levels of tasks.

9. An edge fatigue detection device based on adaptive threshold adjustment, characterized in that: The device comprises: A data acquisition module, the data acquisition module is used to obtain facial image data and facial infrared data of the driver in real time; a data processing module, the data processing module being configured to pre-process the facial image data and the facial infrared data, and locate facial landmarks on the pre-processed facial image data and the facial infrared data using a facial feature detection model to obtain an enhanced 3D landmark image; A data decision module is used to compare the enhanced 3D landmark point image with the physiological feature database to obtain physiological data, and calculate the physiological data through a dynamic threshold decision model and a multi-index fusion engine to obtain a fatigue decision.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • A fatigue state detection method and device

    CN109670421A

  • Fatigue driving detection method based on self-adaptive facial action feature threshold

    CN115171083A

  • Driving state detection method and device and computer equipment

    CN116513202A

  • Multi-person real-time fatigue detection method based on multi-modal information fusion

    CN116612429A

  • Abnormal driving behavior judgment method and system based on multiple modes

    CN118928425A

Cited By

  • Multi-channel one-core multi-detection integrated detection method

    CN121075692A

  • Pilot visual fatigue detection system and method based on artificial intelligence

    CN121505577A

  • A pilot visual fatigue detection system and method based on artificial intelligence

    CN121505577B