Face anonymization fatigue detection method and system based on generative model
Through face anonymization technology based on generative models, the problem that existing driver fatigue detection methods are difficult to balance between privacy protection and detection accuracy is solved, and the key detection information is retained while effectively protecting driver privacy and improving detection accuracy and robustness.
Patent Information
- Application Number
- CN202510065487.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing driver fatigue detection methods are difficult to balance between privacy protection and detection accuracy, especially the vision-based methods have the risk of privacy leakage and the detection accuracy is not high.
The face anonymization technology based on the generative model is adopted to generate potential encodings through the Generative Adversarial Network (GAN), and a new facial image is generated in combination with key facial features and pose geometric heatmaps to achieve anonymization processing while retaining key information for fatigue detection.
On the premise of ensuring the availability of detection data, it effectively protects drivers' privacy, improves the robustness and accuracy of the detection system, significantly reduces false alarms, and makes it better adaptability and stability in complex driving environments.
Smart Images

Figure CN120032352A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent transportation systems and artificial intelligence technology, and specifically relates to a method and system for fatigue detection based on face anonymization and generative model; the invention aims to perform fatigue detection without leaking the driver's identity information based on video anonymization and face replacement technology of the generative model. Background Art
[0002] Road traffic safety has always been a key issue in the field of global public safety. Although countries have made extensive efforts to improve road safety, the number of casualties caused by traffic accidents remains high, and huge economic losses have been caused. Many studies have shown that driver fatigue driving is one of the important causes of traffic accidents. Therefore, accurate and effective detection of driver fatigue is of great significance to improving traffic safety.
[0003] Among the existing methods for detecting driver fatigue, the mainstream technologies can be divided into two categories: vision-based methods and non-visual methods. Non-visual methods usually rely on the driver's physiological signals, such as electroencephalogram (EEG), electrocardiogram (ECG), electromyogram (EMG), and electrooculogram (EOG), or use vehicle status data to analyze the driver's fatigue status. However, the monitoring of physiological signals has certain limitations and requires data to be collected through wearable sensors. This method is invasive and increases the driver's discomfort. In addition, the equipment cost is high and difficult to promote. In addition, although the method of using vehicle data for fatigue detection can easily obtain data, it has poor performance and cannot accurately judge the driver's status.
[0004] Thanks to the rapid development of computer vision technology, vision-based methods have become the mainstream method for fatigue detection. This type of method captures the driver's facial features through a camera and determines the driver's fatigue level based on features such as eye closure status and facial expressions. This method is not physically invasive, is relatively simple to deploy, and can provide high detection accuracy. However, existing visual detection methods involve real-time collection of driver images, which poses obvious privacy issues. Even without physical contact, drivers are subjectively aware that they are being monitored by cameras, which may cause psychological discomfort and there is also a risk of image privacy leakage.
[0005] Currently, research on combining vision-based driver fatigue detection and image and video anonymization has not received enough attention, and even the publicly available technical achievements have their own technical defects.
[0006] For example, patent application CN115908625A proposes a cyclic reversible anonymous face synthesis method to prevent identity leakage, which uses a generative adversarial network to anonymize face images while ensuring the effectiveness of the face after anonymization. When the amount of data is large, the efficiency of the clustering operation used by this method becomes a bottleneck, making it difficult to capture the complex nonlinear features in face images. The acquired features may cause the generated image to be blurred or overly averaged, limiting the realism and detail of the image. The reversibility of the system may affect the user's perception of privacy protection when the image information can still be restored after anonymization.
[0007] Some researchers have tried to protect the driver's privacy information to a certain extent by extracting only the facial parts related to the task. For example, the "de-identified" fatigue detection method only extracts and uses the facial area around the eyes and masks all other content in the background. However, this method is still difficult to promote in practice. The retention of some foreground features may not completely remove the identity information, and excessively blurring the background may affect the model's understanding of the context, resulting in a decrease in detection accuracy; there are limitations in adaptability under different lighting conditions and occlusion conditions, which greatly reduces the data availability while also failing to achieve good detection accuracy.
[0008] Therefore, it is necessary to propose new solutions to better meet the needs of computer vision tasks in specific fields and ensure that anonymized images can be efficiently used in these tasks without compromising their performance. Summary of the invention
[0009] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide a method and system for face anonymization fatigue detection based on a generative model.
[0010] In order to solve the above technical problems, the solution adopted by the present invention is:
[0011] A method for detecting fatigue by anonymous face based on a generative model is provided, the method comprising:
[0012] Obtain dynamic video of the driver's face and segment it according to a preset duration, and use each frame of the video slice within a preset time interval as an original image for input of subsequent processing;
[0013] Synchronously extract facial key features and pose geometry heatmaps from the acquired raw images;
[0014] The generator of the Generative Adversarial Network (GAN) is used to generate a latent code, and a new facial image is further generated based on the latent code and the compressed information contained in the facial key features and posture geometry heat map; by blurring or replacing features related to personal identity and retaining core information related to fatigue detection, anonymization is achieved while retaining key features;
[0015] Preprocess the anonymized image data to obtain the global feature vector and local feature vector of each frame of the image, and form the overall feature representation of each frame of the image through corresponding splicing processing; then arrange the overall feature representation and generate a spatiotemporal feature matrix by combining the feature information in the spatial and temporal dimensions;
[0016] The spatiotemporal attention network is used to extract temporal information and spatial features from the spatiotemporal feature matrix, adaptively capture the dynamic associations between various time points and key features of different spatial positions, and output a one-dimensional spatiotemporal joint feature vector for fatigue judgment.
[0017] Perform feature mapping and probability calculation on the spatiotemporal joint feature vector, make classification decisions on the probability values according to the set threshold, and output the prediction results of the fatigue state; if it is judged that the driver is in a fatigue state, a fatigue alarm is issued; if it is judged that the driver is not in a fatigue state, all the above operations are executed again.
[0018] As a preferred solution of the present invention, a facial key feature extraction module is constructed, and the key point detection and feature extraction model contained in the module is pre-trained, and then used to extract facial key points and related features in the original image; the key points and related features include the key point positions and sizes of the driver's eyes, pupils, and mouth, and the degree of closure of the eyes and mouth; a posture geometry heat map extraction module is constructed, and the posture estimation and key point detection deep learning network contained in the module is pre-trained, and then used to extract the posture geometry heat map; the features of the posture geometry heat map include the driver's head posture, facial orientation and facial geometry information related to fatigue detection.
[0019] As a preferred solution of the present invention, when generating a new facial image, a generator is used to adjust some dimensions in the latent code to blur or replace features related to personal identity; then, the core information related to fatigue detection is controlled to be retained based on the compressed information contained in the input facial key features and posture geometry heat map; the generator has previously learned the rules of feature transformation through pre-training.
[0020] As a preferred solution of the present invention, when preprocessing anonymous image data to obtain a global feature vector, it specifically includes: first using a pre-trained key point detection model to detect facial key points of a face; then performing temporal alignment and normalization processing on all key points; then constructing a feature matrix based on the facial key point data of each frame, and representing the key point information as a two-dimensional matrix; using a dimensionality reduction algorithm to spatially filter the feature matrix, enhance and reduce the input facial key point features, and finally obtain a global feature vector for each frame of the anonymous image.
[0021] As a preferred solution of the present invention, when preprocessing anonymous image data to obtain local feature vectors, the method specifically includes: extracting a bounding box of a key area from each frame image, extracting a corresponding local facial area according to the size and position of the bounding box, and dividing it into multiple small blocks; using a feature embedding algorithm for the local facial area, and mapping each small block to a feature vector through linear mapping; then, splicing the feature vectors of the local facial area in each frame to form a local feature vector of the frame image.
[0022] As a preferred solution of the present invention, a pre-trained spatiotemporal attention network is used to extract the spatiotemporal features of the video image from the spatiotemporal feature matrix, and the temporal information and spatial features of each frame of the image are globally modeled through the spatiotemporal attention mechanism, and the dynamic correlation between each time point and the key features of different spatial positions are adaptively captured based on the attention mechanism; the spatial features include head posture and the closure of the eyes and mouth, and the temporal information includes head movement, head posture change, blinking frequency, yawning and line of sight shift information.
[0023] As a preferred solution of the present invention, a classification module is used to process the spatiotemporal joint feature vector, including: mapping the input spatiotemporal joint features to a new feature space in several fully connected layers of the classification module, and generating a binary classification probability value through a Sigmoid activation function in the output layer of the classification module.
[0024] The present invention further provides a human face anonymization fatigue detection system based on a generative model, comprising: the system comprises a human face anonymization model and a fatigue detection model, the former comprising a facial key feature extraction module, a posture geometry heat map extraction module and an anonymous frame generation module, and the latter comprising a data processing module, a feature extraction module and a fatigue classification module; wherein,
[0025] The facial key feature extraction module includes a key point detection and feature extraction model, which is used to extract facial key features from each frame of the original image of the video slice;
[0026] The pose geometry heat map extraction module contains the pose estimation and key point detection deep learning networks, which is used to extract the pose geometry heat map from each frame of the original image of the video slice;
[0027] Anonymous frame generation module, with a pre-trained generative adversarial network built in, is used to generate latent codes, and further generate new facial images based on the latent codes and the compressed information contained in the facial key features and pose geometry heatmap; by blurring or replacing features related to personal identity and retaining core information related to fatigue detection, anonymization is achieved while retaining key features;
[0028] The data processing module is used to pre-process the anonymized image data, obtain the global feature vector and local feature vector of each frame of the image, and form the overall feature representation of each frame of the image through corresponding splicing processing; then arrange the overall feature representation, and generate a spatiotemporal feature matrix by combining the feature information in the spatial and temporal dimensions;
[0029] The feature extraction module uses the spatiotemporal attention network to extract temporal information and spatial features from the spatiotemporal feature matrix, adaptively captures the dynamic associations between various time points and key features at different spatial locations, and outputs a one-dimensional spatiotemporal joint feature vector for fatigue judgment;
[0030] The fatigue classification module is used to perform feature mapping and probability calculation on the spatiotemporal joint feature vector, make classification decisions on the probability value according to the set threshold and output the prediction result of the fatigue state.
[0031] The present invention also provides a computer device, the computer device comprising a memory and a processor connected in a communication manner; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the aforementioned face anonymization fatigue detection method based on the generation model;
[0032] The present invention further provides a computer-readable storage medium, characterized in that the computer-readable storage medium includes stored program data; when the program data is executed by a processor, the computer device where the storage medium is located is controlled to execute the aforementioned face anonymization fatigue detection method based on the generative model.
[0033] Compared with the prior art, the technical effects of the present invention are:
[0034] 1. By applying face anonymization technology, the present invention can replace the driver's real identity information while retaining the key facial features required for fatigue detection, generating a false portrait without identity features, thereby effectively protecting the driver's privacy while ensuring the availability of detection data.
[0035] 2. The present invention focuses on utilizing the key information retained after anonymization processing, which improves the robustness and accuracy of the detection system, significantly reduces false alarms, and enables it to have better adaptability and stability in complex driving environments.
[0036] 3. Based on the organic combination of anonymization and fatigue detection, the present invention constructs a highly cohesive and low-coupled end-to-end system; achieves a balance between privacy protection and fatigue detection performance, and effectively improves the overall performance and application value of the system.
[0037] 4. The present invention successfully solves the problems of driver privacy leakage and low detection accuracy in traditional fatigue detection methods. Compared with the existing methods based on direct video capture and analysis, the present invention can effectively remove identity-related information and only retain necessary detection features, reducing the risk of privacy leakage.
[0038] 5. The present invention combines a spatio-temporal attention network of deep learning, significantly reducing the false alarm rate while improving the fatigue detection accuracy, and having better robustness especially in complex driving scenarios. Through real-time judgment of the fatigue state, the present invention can timely remind the driver to avoid traffic accidents, thus enhancing driving safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 It is a flowchart for implementing the fatigue detection method of the present invention.
[0040] Figure 2 It is a schematic structural diagram of the fatigue detection system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] First of all, it should be noted that the present invention relates to artificial intelligence technology and is a specific application of computer technology in the field of road traffic safety. In the implementation process of the present invention, the application of multiple software function modules will be involved. The applicant believes that after carefully reading the application documents and accurately understanding the implementation principle and invention purpose of the present invention, and in combination with the existing well-known technologies, those skilled in the art can fully implement the present invention by using their mastered software programming skills. The aforementioned software function modules include but are not limited to: face anonymization model, fatigue detection model, facial key feature extraction module, pose geometry heat map extraction module, anonymized frame generation module, data processing module, feature extraction module, fatigue classification module, generative adversarial network, deep learning network, key point detection and feature extraction model, pose estimation and key point detection model, etc. All those mentioned in the application documents of the present invention belong to this category, and the applicant will not list them one by one.
[0042] Those skilled in the art know that in addition to implementing a part of the system provided by the present invention and its various devices, modules, and units in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system provided by the present invention and its various devices, modules, and units to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system provided by the present invention and its various devices, modules, and units can be regarded as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structures within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as both software modules for implementing the method and the structures within the hardware component.
[0043] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0044] In view of the shortcomings of the existing technology, the present invention proposes a driver fatigue detection method based on face anonymization. The core idea is to anonymize the driver's facial image in the video through a generative model, remove the features related to identity information, retain the key facial features for fatigue detection, and design the downstream fatigue detection algorithm in a targeted manner. The two-part algorithm is jointly optimized to form an end-to-end overall system to ensure that the anonymized image does not affect the detection accuracy. This method effectively reduces the risk of privacy leakage and improves the accuracy and robustness of the detection system.
[0045] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings.
[0046] 1. The present invention proposes a driver fatigue detection method based on face anonymization, and the specific implementation steps are as follows:
[0047] (1) Raw data acquisition: The driver's dynamic video is acquired through the RGB camera installed under the rearview mirror, and the video is segmented according to the preset time length t to divide the video into multiple video slices. Each frame f in the video slice within the preset time interval t is used as the input raw image for subsequent processing.
[0048] (2) Extraction of facial key features: Each original frame f in the video slice is input into the face anonymization model, which has a built-in facial key feature extraction module, which includes a key point detection and feature extraction model. The pre-trained key point detection and feature extraction model is used to extract key points and related features, including the location and size of the key points of the driver's eyes, pupils, and mouth, and the degree of eye and mouth closure.
[0049] (3) Posture geometry heat map extraction: Each original frame f in the video slice is input into the face anonymization model, which has a built-in posture geometry heat map extraction module, including a posture estimation and key point detection deep learning network. The pre-trained posture estimation and key point detection deep learning network is used to extract the driver's posture geometry heat map m, which contains features related to fatigue detection, such as the driver's head posture, facial orientation, and facial geometry information.
[0050] Note that step (3) and step (2) are performed in parallel.
[0051] (4) Face anonymization processing:
[0052] Image and video anonymization technologies can be roughly divided into two categories: based on generative adversarial networks (GANs) and based on diffusion models. Based on the applicant's research summary, the overall performance of the GAN series of models is better than that of the diffusion model in video anonymization and driver fatigue detection tasks. Therefore, in this application, a generative adversarial network is used to anonymize the driver's video. The "based on a generative model" mentioned in the title of the invention refers to a "face anonymization model based on a built-in generative adversarial network."
[0053] As a specific application example, the model for generating anonymized faces can use the pre-trained StyleGAN3, and feature splicing the key facial features and posture geometry heat map extracted in the above steps with the latent code generated by the model, and then input into the generator of StyleGAN3 to generate a new anonymized portrait to ensure that the anonymized key facial features and head posture information are stable and consistent with the original image. The generator of this model generates a new replacement face image by combining the compressed information of the latent code, key facial features and posture geometry heat map, and anonymizes the driver's face while retaining the key features.
[0054] Specifically, a pre-trained generative adversarial network (GAN) generator is used to generate a latent code c in the latent space; then a new facial image g(c, ε, m) is generated based on the code, the input facial key features ε, and the compressed information contained in the posture geometry heat map m. In this process, the generator adjusts some dimensions in the latent code to blur or replace features related to personal identity; at the same time, the core information related to fatigue detection, such as eyes, mouth-related features, and head posture, is controlled to be retained according to the input ε and m. This generation method ensures that the private information in the image is hidden or replaced, but does not affect the subsequent fatigue detection task. In the inference stage, the generator has learned the rules of feature transformation through pre-training, and directly generates anonymized images based on the input features without further adjusting the model parameters. Through the processing of GAN, the generated anonymized image is obtained, ensuring that the private information of facial features is hidden and the key information required for fatigue detection is retained.
[0055] (5) Repeat steps (1)-(4) until all frames in a given video slice are anonymized.
[0056] (6) Entering the fatigue detection stage: the anonymized image is input into the fatigue detection model for subsequent processing.
[0057] (7) Fatigue detection data preprocessing: After the anonymized image is input into the fatigue detection model, the data processing module first preprocesses the video data.
[0058] First, the facial key points of the face are detected using a pre-trained key point detection model. Then, all key points are time-series aligned and normalized to ensure data consistency and availability, and frame loss interpolation is performed when necessary. Next, a feature matrix is constructed based on the facial key point data of each frame, and the key point information of each frame is represented as a two-dimensional matrix. In order to improve the discrimination ability between samples of different fatigue states, a dimensionality reduction algorithm is used to spatially filter the feature matrix, thereby enhancing and reducing the input facial key point features, and obtaining the global feature vector of each frame after processing.
[0059] Meanwhile, the bounding boxes of key regions such as eyes and mouths are extracted from each frame of the video. According to the sizes and positions of these bounding boxes, the corresponding local facial regions are extracted. These regions correspond to the key facial features in step (2) of the method, that is, the key features for fatigue detection that remain consistent after anonymization. At the same time, the extracted local key facial regions (such as the eye and mouth regions) are divided into multiple small blocks. The feature embedding algorithm (FeatureEmbedding) is used for these local facial regions, and these small blocks are mapped into feature vectors through linear mapping. Subsequently, the feature vectors of the local facial regions are concatenated within each frame to form the local feature vector of the frame image.
[0060] Next, the local feature vector and the global feature vector of each frame are concatenated within each frame to form the overall feature representation of the frame image.
[0061] Finally, the overall feature representations of each frame are arranged to combine the feature information in the spatial and temporal dimensions, generating a comprehensive spatio-temporal feature matrix. In this way, the extracted features can better capture the spatio-temporal variation law of the fatigue state, providing more discriminative input features for the subsequent classifier.
[0062] (8) Fatigue detection feature extraction: To enhance the accuracy and robustness of fatigue detection, the present invention takes into account the spatio-temporal information in the video. That is to say, in addition to the spatial information of each frame, the present invention also focuses on the temporal information of the video. Specifically, it involves the changes of the driver between different video frames, manifested as time-related information such as head movement, head pose change, gaze shift, and blink frequency.
[0063] Specifically, the preprocessed spatio-temporal feature matrix in step (7) is input into the feature extraction module. The module uses the spatio-temporal attention network (Spatio-Temporal Attention Network, STAN) to extract the spatio-temporal features of the video. The spatial features include information such as head pose and eye and mouth closure degree, and the temporal features include information such as head movement, head pose change, blink frequency, yawn frequency, and gaze shift. STAN directly extracts spatio-temporal related features from the input sequence through a global spatio-temporal attention mechanism. The temporal information and spatial features of each input frame will be globally modeled through the spatio-temporal attention mechanism. The attention mechanism of the network adaptively captures the dynamic associations between different time points and the key features at different spatial positions, ensuring the capture of fatigue-related behaviors such as yawning, abnormal eye closure degree, blinking, head movement, and gaze change. This module outputs a one-dimensional spatio-temporal joint feature vector for subsequent final fatigue judgment.
[0064] (9) Fatigue status classification: After obtaining the spatiotemporal joint feature vector, it is input into the fatigue classification module. The classification module consists of several fully connected layers, each of which maps the input features to a new feature space, enhancing the feature discrimination capability. In the final output layer, a binary probability value is generated through the Sigmoid activation function to determine the driver's fatigue status. Finally, the system makes a classification decision on the output probability value based on the set threshold, completes the prediction of the fatigue status, and outputs whether the driver is fatigued.
[0065] (10) Feedback and warning: If the detection model determines that the driver is in a fatigued state, the system will trigger a fatigue alarm and send a prompt message to the driver. The prompt content includes sound and light reminders, seat vibration, etc., to help the driver stay alert and prevent dangerous driving behavior.
[0066] (11) If it is determined that the driver is not in a fatigue state, all the above operations are performed again. Steps (1) to (10) are repeated while the driver is driving until the fatigue detection system is turned off.
[0067] Second, in order to implement the above-mentioned face anonymization driver fatigue detection method, the present invention provides a face anonymization driver fatigue detection system formed by two models: a face anonymization model and a fatigue detection model. The former includes a facial key feature extraction module, a posture geometry heat map extraction module and an anonymous frame generation module, and the latter includes a data processing module, a feature extraction module and a fatigue classification module. Among them,
[0068] The facial key feature extraction module includes a key point detection and feature extraction model, which is used to extract facial key features from each frame of the original image of the video slice; the posture geometry heat map extraction module includes a posture estimation and key point detection deep learning network, which is used to extract the posture geometry heat map from each frame of the original image of the video slice; the anonymous frame generation module has a built-in pre-trained generative adversarial network, which is used to generate latent codes and further generate new facial images based on the latent codes and the compressed information contained in the facial key features and posture geometry heat map; by blurring or replacing the features related to personal identity and retaining the core information related to fatigue detection, anonymization is achieved while retaining the key features; the data processing module is used to perform anonymous processing on the anonymously processed The image data is preprocessed to obtain the global feature vector and local feature vector of each frame of the image, and the overall feature representation of each frame of the image is formed after the corresponding splicing processing; then the overall feature representation is arranged, and the feature information in the spatial and temporal dimensions is combined to generate a spatiotemporal feature matrix; the feature extraction module uses the spatiotemporal attention network to extract the timing information and spatial features from the spatiotemporal feature matrix, adaptively capture the dynamic correlation between each time point and the key features of different spatial positions, and output a one-dimensional spatiotemporal joint feature vector for fatigue judgment; the fatigue classification module is used to perform feature mapping and probability calculation on the spatiotemporal joint feature vector, make classification decisions on the probability value according to the set threshold, and output the prediction result of the fatigue state.
[0069] Furthermore, the present invention also provides corresponding computer equipment and computer-readable storage media, which implement the various steps in the above-mentioned face anonymization fatigue detection method by executing the stored computer program.
[0070] In summary, the driver fatigue detection method proposed in the present invention combines privacy protection with high-accuracy detection. By applying face anonymization technology, the present invention can replace the driver's real identity information while retaining the key facial features required for fatigue detection, and generate a false portrait without identity features, thereby effectively protecting the driver's privacy while ensuring the availability of detection data, and reducing psychological intrusion while maintaining detection accuracy. The fatigue detection algorithm proposed in the present invention focuses on the key information retained after anonymization processing, improves the robustness and accuracy of the detection system, significantly reduces false alarms, and makes it more adaptable and stable in complex driving environments.
Claims
1. A method for face anonymization fatigue detection based on a generative model, characterized in that: The method comprises: Obtain dynamic video of the driver's face and segment it according to a preset duration, and use each frame of the video slice within a preset time interval as an original image for input of subsequent processing; Synchronously extract facial key features and pose geometry heatmaps from the acquired raw images; The generator of the Generative Adversarial Network (GAN) is used to generate a latent code, and a new facial image is further generated based on the latent code and the compressed information contained in the facial key features and posture geometry heat map; by blurring or replacing features related to personal identity and retaining core information related to fatigue detection, anonymization is achieved while retaining key features; Preprocess the anonymized image data to obtain the global feature vector and local feature vector of each frame of the image, and form the overall feature representation of each frame of the image through corresponding splicing processing; then arrange the overall feature representation and generate a spatiotemporal feature matrix by combining the feature information in the spatial and temporal dimensions; The spatiotemporal attention network is used to extract temporal information and spatial features from the spatiotemporal feature matrix, adaptively capture the dynamic associations between various time points and key features of different spatial positions, and output a one-dimensional spatiotemporal joint feature vector for fatigue judgment. Perform feature mapping and probability calculation on the spatiotemporal joint feature vector, make classification decisions on the probability values according to the set threshold, and output the prediction results of the fatigue state; if it is judged that the driver is in a fatigue state, a fatigue alarm is issued; if it is judged that the driver is not in a fatigue state, all the above operations are executed again.
2. The method according to claim 1, characterized in that Constructing a facial key feature extraction module, pre-training the key point detection and feature extraction model contained in the module, and then using it to extract facial key points and related features in the original image; the key points and related features include the key point positions and sizes of the driver's eyes, pupils, and mouth, as well as the closure of the eyes and mouth; A posture geometry heat map extraction module is constructed, and the posture estimation and key point detection deep learning networks contained in the module are pre-trained and then used to extract the posture geometry heat map; the features of the posture geometry heat map include the driver's head posture, facial orientation and facial geometry information related to fatigue detection.
3. The method according to claim 1, characterized in that When generating a new facial image, the generator is used to adjust some dimensions in the latent code to blur or replace features related to personal identity; then, the core information related to fatigue detection is controlled to be retained based on the compressed information contained in the input facial key features and posture geometry heat map; the generator has previously learned the rules of feature transformation through pre-training.
4. The method according to claim 1, characterized in that When preprocessing anonymous image data to obtain a global feature vector, it specifically includes: first, using a pre-trained key point detection model to detect facial key points of the face; then performing temporal alignment and normalization on all key points; then constructing a feature matrix based on the facial key point data of each frame, and representing the key point information as a two-dimensional matrix; using a dimensionality reduction algorithm to spatially filter the feature matrix, enhance and reduce the dimensionality of the input facial key point features, and finally obtain the global feature vector of each frame of the anonymous image.
5. The method according to claim 1, characterized in that When preprocessing anonymous image data to obtain local feature vectors, it specifically includes: extracting the bounding box of the key area from each frame image, extracting the corresponding local facial area according to the size and position of the bounding box, and dividing it into multiple small blocks; using a feature embedding algorithm for the local facial area, and mapping each small block to a feature vector through linear mapping; then, splicing the feature vectors of the local facial area in each frame to form a local feature vector of the frame image.
6. The method according to claim 1, characterized in that The spatiotemporal features of the video image are extracted from the spatiotemporal feature matrix using a pre-trained spatiotemporal attention network. The temporal information and spatial features of each frame of the image are globally modeled through the spatiotemporal attention mechanism. The dynamic correlation between each time point and the key features of different spatial positions are adaptively captured based on the attention mechanism. The spatial features include head posture and the degree of closure of the eyes and mouth, and the temporal information includes head movement, head posture changes, blinking frequency, yawning and gaze shift information.
7. The method according to claim 1, characterized in that The spatiotemporal joint feature vector is processed by using a classification module, including: mapping the input spatiotemporal joint features to a new feature space in several fully connected layers of the classification module, and generating a binary classification probability value through a Sigmoid activation function in an output layer of the classification module.
8. A facial anonymization fatigue detection system based on a generative model, characterized in that: include: The system includes a face anonymization model and a fatigue detection model. The former includes a facial key feature extraction module, a posture geometry heat map extraction module and an anonymous frame generation module, and the latter includes a data processing module, a feature extraction module and a fatigue classification module. The facial key feature extraction module includes a key point detection and feature extraction model, which is used to extract facial key features from each frame of the original image of the video slice; The pose geometry heat map extraction module contains the pose estimation and key point detection deep learning networks, which is used to extract the pose geometry heat map from each frame of the original image of the video slice; Anonymous frame generation module, with a pre-trained generative adversarial network built in, is used to generate latent codes, and further generate new facial images based on the latent codes and the compressed information contained in the facial key features and pose geometry heatmap; by blurring or replacing features related to personal identity and retaining core information related to fatigue detection, anonymization is achieved while retaining key features; The data processing module is used to pre-process the anonymized image data, obtain the global feature vector and local feature vector of each frame of the image, and form the overall feature representation of each frame of the image through corresponding splicing processing; then arrange the overall feature representation, and generate a spatiotemporal feature matrix by combining the feature information in the spatial and temporal dimensions; The feature extraction module uses the spatiotemporal attention network to extract temporal information and spatial features from the spatiotemporal feature matrix, adaptively captures the dynamic associations between various time points and key features at different spatial locations, and outputs a one-dimensional spatiotemporal joint feature vector for fatigue judgment; The fatigue classification module is used to perform feature mapping and probability calculation on the spatiotemporal joint feature vector, make classification decisions on the probability value according to the set threshold and output the prediction result of the fatigue state.
9. A computer device, characterized in that: The computer device includes a memory and a processor connected in a communicative manner; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the face anonymization fatigue detection method based on the generative model described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes stored program data; when the program data is executed by a processor, the computer device where the storage medium is located is controlled to execute the face anonymization fatigue detection method based on the generative model described in any one of claims 1 to 7.
Citation Information
Patent Citations
Cyclic reversible anonymous face synthesis method for preventing identity leakage
CN115908625A
Model training and identity anonymization method and device, equipment and storage medium
CN114936377A
Blind face recovery method based on domain alignment GAN prior
CN116362991A
Fatigue detection method based on spatio-temporal information fusion
CN118982859A
Face anonymization using a generative adversarial network
US20230328039A1