A method and system for recognizing and warning abnormal behaviors and pain expressions of drivers

Through improved image processing and deep learning technology, abnormal behaviors and painful expressions of drivers are identified and early warning, and the problem of low manual monitoring efficiency in the prior art is solved, and automated driver status judgment and early warning functions are realized.

CN118658144BActive Publication Date: 2025-05-27HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410686285.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-05-27
Estimated Expiration
2044-05-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify and early warning drivers of abnormal behaviors and painful expressions, especially in private vehicles, where manual monitoring is inefficient and difficult to detect problems in a timely manner.

Method used

The improved RPCA algorithm is used for image noise reduction, and combined with the Laplace operator to enhance image details. Then, the Faster R-CNN model and Soft-NMS algorithm are used for object detection, the GAT attention mechanism and the YOLO-POSE model are introduced for posture recognition, and the GAN-MFER algorithm and the hollow convolution network are used for multi-pose facial expression recognition. The comprehensive results are used to judge the driver's status and issue an alarm.

Benefits of technology

It improves image quality and recognition accuracy, effectively reduces the impact of redundant bounding boxes, improves the accuracy and timeliness of driver status judgments, and realizes an automated early warning function.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118658144B_ABST
    Figure CN118658144B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for identifying and warning abnormal behaviors and painful expressions of drivers, the method comprising: extracting key frames from a video to form an image; using an improved RPCA algorithm to reduce noise on the image; inputting the processed image into a Faster R-CNN model to generate a detection frame of a target to be identified; establishing a posture estimation and recognition model YOLO-POSE model; using a multi-posture face recognition algorithm GAN-MFER of a generative adversarial network to identify the driver's facial expression; accurately judging the driver's current state and giving a corresponding alarm prompt; the system comprises a data acquisition module, an image enhancement module, and an image recognition module. The present invention effectively improves the detection accuracy and efficiency of Faster R-CNN; the results of comprehensive posture recognition and facial expression recognition are used to more accurately judge the driver's current state, effectively improve the recognition efficiency of the model, and effectively ensure the life safety of people inside and outside the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and in particular to a method and system for recognizing and giving early warning of abnormal behavior and painful expressions of a driver. Background Art

[0002] With the rapid development of the economy, the number of cars has increased year by year. Traffic accidents caused by dangerous behaviors or emergencies of drivers are increasing. For example, long-term driving may cause fatigue, slow reaction, misjudgment and other problems, thus causing traffic accidents; or sudden illness may cause loss of control of the vehicle, thus causing traffic accidents.

[0003] To solve this hidden danger at present, most methods use cameras to monitor the driver's status in real time and rely on manual reminders. However, this method mainly relies on manpower, which is inefficient and difficult to detect problems in time. In addition, this method is suitable for public transportation and is not very effective for private vehicles. Summary of the invention

[0004] Purpose of the invention: The purpose of the invention is to provide a method and system for identifying and warning abnormal behavior and painful expressions of drivers.

[0005] Technical solution: The method for identifying and warning abnormal behaviors and painful expressions of drivers of the present invention comprises the following steps:

[0006] (1) Using the video data obtained by the high-definition camera in front of the cab, key frames are extracted from the video to form an image;

[0007] (2) Using an improved RPCA algorithm to reduce the noise of the image; the improved RPCA algorithm introduces the minimization of the weighted Schatten-p norm into RPCA to enhance the performance of RPCA in reducing the noise of the image; and combining the Laplace operator to enhance the details of the image to further improve the clarity of the image;

[0008] (3) The processed image is input into the Faster R-CNN model to generate the target detection box to be identified. At the same time, the Soft-NMS algorithm is introduced to process the target detection box to reduce the impact of redundant bounding boxes on the detection results.

[0009] (4) Based on the YOLO framework, the YOLO-POSE model is established as a pose estimation and recognition model, and the GAT attention mechanism is introduced into the model to improve recognition accuracy;

[0010] (5) The multi-pose face recognition algorithm GAN-MFER based on generative adversarial networks is used to identify the driver's facial expressions, and the dilated convolution DC is integrated into the algorithm to enhance the accuracy of identifying the driver's pathological expressions;

[0011] (6) Based on the results of posture recognition and facial expression recognition, the driver’s current status can be accurately determined and corresponding alarm prompts can be given.

[0012] Furthermore, the step (2) comprises:

[0013] (2.1) The Schatten-p norm is used to achieve low-rank regularization, and the formula is as follows:

[0014]

[0015]

[0016] Where: X is the low-rank matrix of the objective function; W is a non-negative vector; σ i is the i-th eigenvalue of X; n and m are the number of rows and columns of matrix X respectively; ω i is the i-th element of W; p is the weight; L is the augmented Lagrangian function; E is the identity matrix; Z is the Lagrangian multiplier; μ is a positive scalar; Y is the observed data; b is the weight vector;

[0017] (2.2) The Laplace operator is used to enhance the details of the denoised image. The formula is as follows:

[0018]

[0019] Among them, g represents the gradient; (i, j) represents the coordinates of the pixel point; k and l are the grayscale values ​​in the horizontal and vertical directions respectively; (r, s) are the coordinates of the neighboring pixels of (i, j); f is a two-dimensional discrete function; and H is the directional derivative.

[0020] Furthermore, the step (3) comprises:

[0021] (3.1) Use the VGG16 feature extraction network to extract features from the image and obtain the feature map of the image. The loss function is as follows:

[0022]

[0023] Where: L cls is the foreground logarithmic loss; p i is the probability that the i-th pixel is predicted as the target; is the characteristic value of the i-th pixel; λ is the balance ratio; N is the total number of pixels; L reg is the logarithmic loss of the background; t i is the position of the i-th pixel; t i The offset of

[0024] (3.2) The region of interest of the image is obtained through the selective search algorithm SS, that is, the candidate box is generated. The calculation formula is as follows:

[0025]

[0026]

[0027]

[0028]

[0029] s=a 1 s color +a 2 s texture +a 3 s size +a 4 s fill

[0030] Where: s color 、s texture 、s size 、s fill are the color similarity, texture similarity, size similarity and filling similarity of the current region block; s is the candidate box formed by the above four similarities; n is the total number of region blocks divided into the picture; k is the number of the region block; c and t are the color similarity and texture similarity of the current pixel point respectively; r is the region block being detected; B is the region containing r and its neighboring regions; i and j are distinguishing subscripts; a 1 、a 2 、a 3 、a 4 is the adoption weight, which takes 0 or 1;

[0031] (3.3) Perform ROI pooling operation on the candidate box to obtain a feature box map of uniform size;

[0032] (3.4) Perform non-maximum suppression on all candidate boxes, and use the Soft-NMS algorithm to replace the original NMS algorithm. During the execution of the algorithm, Soft-NMS does not simply delete the detection box whose IoU is greater than the threshold, but uses a linear weighted function to calculate and reduce its score. The function formula is as follows:

[0033]

[0034] Where: S i is the confidence score corresponding to the i-th prediction box; M is the candidate box with the highest score; box i is the frame to be detected; N t is the overlap threshold set by the hyperparameter; loU is boxi The ratio of the intersection and union of M.

[0035] Furthermore, the step (4) comprises:

[0036] (4.1) The YOLO-POSE model is established to identify the action recognition of the driver in a relatively fixed posture, which mainly includes the following three parts:

[0037] Mark the key points and confidence levels of the human body in the frame to be detected, and calculate the similarity of the key points. The formula is as follows:

[0038]

[0039] Where: S is the target frame to be detected; C x , C y are the horizontal and vertical coordinates of the anchor point respectively; W is the width of the box to be detected; H is the height of the box to be detected; box conf is the confidence of the box to be detected; class conf is the category confidence of the box to be detected; are the horizontal and vertical coordinates of the i-th human key point; is the confidence of the i-th human key point, Take an integer; n is the number of key points of the human body in the frame to be detected; U is the similarity of the key points; exp is an exponential function with the natural constant e as the base; p is the number of the frame to be detected; p i is the number of the key point of the skeleton of the person; Indicates the visibility of the key point in the image; is the square of the Euclidean distance between the true point and the predicted point; S p Represents the square root of the area occupied by the detection box; σ i is the standard deviation of the key point;

[0040] The YOLO-POSE model uses the HorNet network as the convolution kernel of the model, and uses high-order gated convolution and recursive design to achieve high-order spatial interaction. The relevant formula is as follows:

[0041]

[0042] T k+1 =f k (P k )⊙g s (T k ) / α, k=0, 1, ..., n-1

[0043]

[0044] FLOPs(g n Conv) <HWCc (2M 2 +11 / 3×C c +2)

[0045] x is the input feature of the gated convolution; W is the width of the box to be detected; H is the height of the box to be detected; φ(·) is the linear projection operation; T 0 , P k is the projection feature after linear projection operation of x, k = 0, 1, ..., n-1; f k (·) is the kth convolution operation, k = 0, 1, ..., n-1; ⊙ is the dot product operation; P k Do convolution operation and add it with g s (T k ) and perform dot product operation to get T k+1 , k = 0, 1, ..., n-1; α is the reduction factor; g s is the dimension mapping function; C k is the dimension of the kth operation, k = 0, 1, ..., n-1; n is the number of convolutional layers; C c is the total dimension; FLOPs is the total amount of operations; g n Conv is the n-order spatial interaction capability; M is the scale of the convolution kernel;

[0046] (4.2) The graph attention mechanism GAT is introduced into the YOLO-POSE model to accurately weight and aggregate the neighboring key points of each key point. The GAT formula is as follows:

[0047]

[0048] Where: α ij is the attention coefficient between node i and node j; exp is an exponential function with the natural constant e as the base; LeakyReLU is a function; α T is the transpose of the weight parameter; W is the feature transformation weight matrix shared by each layer; h i is the feature vector of node i; h j is the feature vector of node j; h′ i is the feature of node i; τ is the activation function; N i is the number of neighbor nodes of node i; W k is the attention weight matrix;

[0049] Furthermore, the bounding box loss function defined in step (4.1) is as follows:

[0050]

[0051] Where: L CIoUis the bounding box loss function; IoU is the ratio of the intersection area to the union area of ​​the real box and the predicted box; A is the area of ​​the real box; B is the area of ​​the predicted box; d and d gt are the coordinate positions of the center points of the predicted box and the real box respectively; ρ 2 (d,d gt ) is the square of the Euclidean distance between two center points; j 2 is the length of the diagonal of the minimum closed area formed by the predicted box and the real box; β is the penalty coefficient; v is the penalty term for the aspect ratio of the predicted box and the real box; w z With h z are the width and height of the real box respectively; w and h are the width and height of the predicted box;

[0052] Furthermore, the step (5) comprises:

[0053] (5.1) While performing driver posture recognition, the multi-posture face recognition algorithm GAN-MFER based on the generative adversarial network is used to recognize the driver's facial expression to improve the accuracy of judging the driver's current state. The formula is as follows:

[0054]

[0055] Where: D N The multi-pose face correction result output by the discriminator; The probability that the acquired image can be represented as a real picture; is the identity classification result of the identified image; D E The final output result of face rotation; The amount of loss incurred for correction; To determine the image input source; Image classification results; To classify facial expressions; is the residual attention function of the discriminator when generating face image classification; x is the image distribution data; c is the corresponding posture of the expression;

[0056] The two-way loop optimization function in the generative adversarial network is as follows:

[0057]

[0058]

[0059] Where: L c-1 represents the optimization function in the double loop of face profile-front-profile conversion; L c-2 represents the optimization function in the double loop of face front-side-front conversion; n is the number of loops; x i is the value distributed at node i; G EG is the profile face generator; N is a frontal face image;

[0060] (5.2) The hole convolution network DC is introduced into the multi-pose face recognition algorithm of the generative adversarial network to improve the recognition accuracy and detection efficiency. The formula is as follows:

[0061] f t spatial =E(I t )

[0062] f t temporal =P(f t-1 , f t-2 )

[0063] f t =S(f t spatial , f t temporal )

[0064] Where: f t spatial is the feature map calculated within two frames; E(·) is the spatial information estimator; I t is the input image; f t temporal is the calculated temporal feature map; P(·) is the ConvLSTM predictor; f t-1 、f t-2 is the feature map of two frames; f t is the final feature; S(·) is the fusion of spatiotemporal information.

[0065] Furthermore, the step (6) comprises:

[0066] The results of gesture recognition and facial expression recognition are combined to determine the driver's current state. If the driver nods frequently or blinks for a long time, it is judged that the driver is dozing off and is in a fatigued driving state, and the system will automatically use the radio or speaker to remind the driver.

[0067] When the driver's body shakes violently, beats his chest, falls to the ground, or shows an expression of pain, it is judged that the driver has a sudden illness. The alarms inside and outside the car will sound to remind people inside and outside the car to take relevant measures to reduce the accident rate and casualties.

[0068] The driver's abnormal behavior and pain expression recognition and warning system of the present invention comprises:

[0069] A data acquisition module is used to extract key frames from the video data acquired by the high-definition camera in front of the cab to form an image;

[0070] An image enhancement module is used to reduce the noise of an image using an improved RPCA algorithm; the improved RPCA algorithm introduces the minimization of the weighted Schatten-p norm into RPCA to enhance the performance of RPCA in reducing the noise of an image; and combines the Laplace operator to enhance the details of the image to further improve the clarity of the image;

[0071] The object detection module is used to input the processed image into the Faster R-CNN model to generate the detection box of the target to be identified. At the same time, the Soft-NMS algorithm is introduced to process the target detection box to reduce the impact of redundant bounding boxes on the detection results;

[0072] The image recognition module is used to recognize the driver's facial expression using the multi-pose face recognition algorithm GAN-MFER based on the generative adversarial network, and integrate the dilated convolution DC into the algorithm to enhance the accuracy of recognizing the driver's pathological expression.

[0073] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0074] The RPCA algorithm with weighted Schatten-p norm minimization is used to reduce the noise of the image, which improves the image quality. After noise reduction, the Laplace operator is used to enhance the details of the image, which protects the details of the image and provides a basis for the accuracy of subsequent image recognition.

[0075] When Faster R-CNN marks candidate boxes for images, it is easy to generate multiple candidate boxes that overlap with each other to form redundant candidate boxes, which reduces the accuracy and speed of recognition. For this reason, the Soft-NMS algorithm is introduced to greatly reduce the impact of redundant bounding boxes and effectively improve the detection accuracy and efficiency of Faster R-CNN.

[0076] In order to improve the accuracy of the posture recognition model YOLO-POSE in identifying the driver's head, hand and body posture, the GAT attention mechanism was introduced into the model YOLO-POSE, which effectively improved the recognition efficiency of the model;

[0077] In addition to posture recognition of the driver, an expression recognition channel is added. The multi-posture face recognition algorithm GAN-MFER based on the generative adversarial network is used to recognize the driver's facial expressions, and the atrous convolutional network DC is integrated to enhance the accuracy of GAN-MFER in recognizing the driver's pathological expressions. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 is a flow chart of the present invention;

[0079] Figure 2This is the improved Faster R-CNN algorithm flowchart;

[0080] Figure 3 Schematic diagram of the improved YOLO-POSE network structure. DETAILED DESCRIPTION

[0081] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0082] like Figure 1 As shown, the method for identifying and warning abnormal behaviors and painful expressions of drivers of the present invention specifically comprises the following steps:

[0083] (1) Using the video data obtained by the high-definition camera in front of the cab, key frames are extracted from the video to form an image;

[0084] (2) Using an improved RPCA algorithm to reduce the noise of the image; the improved RPCA algorithm introduces the minimization of the weighted Schatten-p norm into RPCA to enhance the performance of RPCA in reducing the noise of the image; and combining the Laplace operator to enhance the details of the image to further improve the clarity of the image;

[0085] (2.1) The Schauen-p norm is used to achieve low-rank regularization, and its formula is as follows:

[0086]

[0087]

[0088] Where: X is the low-rank matrix of the objective function; W is a non-negative vector; σ i is the i-th eigenvalue of X; n and m are the number of rows and columns of matrix X respectively; ω i is the i-th element of W; p is the weight; L is the augmented Lagrangian function; E is the identity matrix; Z is the Lagrangian multiplier; μ is a positive scalar; Y is the observed data; b is the weight vector;

[0089] (2.2) The Laplace operator is used to enhance the details of the denoised image. The formula is as follows:

[0090]

[0091] Among them, g represents the gradient; (i, j) represents the coordinates of the pixel point; k and l are the grayscale values ​​in the horizontal and vertical directions respectively; (r, s) are the coordinates of the neighboring pixels of (i, j); f is a two-dimensional discrete function; and H is the directional derivative.

[0092] (3) The processed image is input into the Faster R-CNN model to generate the target detection box to be identified. At the same time, the Soft-NMS algorithm is introduced to process the target detection box to reduce the impact of redundant bounding boxes on the detection results.

[0093] (3.1) Use the VGG16 feature extraction network to extract features from the image and obtain the feature map of the image. The loss function is as follows:

[0094]

[0095] Where: L cls is the foreground logarithmic loss; p i is the probability that the i-th pixel is predicted as the target; is the characteristic value of the i-th pixel; λ is the balance ratio; N is the total number of pixels; L reg is the logarithmic loss of the background; t i is the position of the i-th pixel; t i The offset of

[0096] (3.2) The region of interest of the image is obtained through the selective search algorithm SS, that is, the candidate box is generated. The calculation formula is as follows:

[0097]

[0098]

[0099]

[0100]

[0101] s=a 1 s color +a 2 s texture +a 3 s size +a 4 s fill

[0102] Where: S color 、s texture 、s size 、s fill are the color similarity, texture similarity, size similarity and filling similarity of the current region block; s is the candidate box formed by the above four similarities; n is the total number of region blocks divided into the picture; k is the number of the region block; c and t are the color similarity and texture similarity of the current pixel point respectively; r is the region block being detected; B is the region containing r and its neighboring regions; i and j are distinguishing subscripts; a 1 、a2 、a 3 、a 4 is the adoption weight, which takes 0 or 1;

[0103] (3.3) Perform ROI pooling operation on the candidate box to obtain a feature box map of uniform size;

[0104] (3.4) Perform non-maximum suppression on all candidate boxes, and use the Soft-NMS algorithm to replace the original NMS algorithm. During the execution of the algorithm, Soft-NMS does not simply delete the detection box whose IoU is greater than the threshold, but uses a linear weighted function to calculate and reduce its score. The function formula is as follows:

[0105]

[0106] Where: S i is the confidence score corresponding to the i-th prediction box; M is the candidate box with the highest score; box i is the frame to be detected; N t is the overlap threshold set by the hyperparameter; IoU is box i The ratio of the intersection and union of M.

[0107] (4) Based on the YOLO framework, the YOLO-POSE model is established as a pose estimation and recognition model, and the GAT attention mechanism is introduced into the model to improve recognition accuracy;

[0108] (4.1) The YOLO-POSE model is established to identify the action recognition of the driver in a relatively fixed posture, which mainly includes the following three parts:

[0109] Mark the key points and confidence levels of the human body in the frame to be detected, and calculate the similarity of the key points. The formula is as follows:

[0110]

[0111] Where: S is the target frame to be detected; C x , C y are the horizontal and vertical coordinates of the anchor point respectively; W is the width of the box to be detected; H is the height of the box to be detected; box conf is the confidence of the box to be detected; class conf is the category confidence of the box to be detected; are the horizontal and vertical coordinates of the i-th human key point; is the confidence of the i-th human key point, Take an integer; n is the number of key points of the human body in the frame to be detected; U is the similarity of the key points; exp is an exponential function with the natural constant e as the base; p is the number of the frame to be detected; p iis the number of the key point of the skeleton of the person; Indicates the visibility of the key point in the image; is the square of the Euclidean distance between the true point and the predicted point; S p Represents the square root of the area occupied by the detection box; σ i is the standard deviation of the key point;

[0112] The YOLO-POSE model uses the HorNet network as the convolution kernel of the model, and uses high-order gated convolution and recursive design to achieve high-order spatial interaction. The relevant formula is as follows:

[0113]

[0114] T k+1 =f k (P k )⊙g s (T k ) / α, k=0, 1, ..., n-1

[0115]

[0116] FLOPs(g n Conv) <HW C c (2M 2 +11 / 3×C c +2)

[0117] x is the input feature of the gated convolution; W is the width of the box to be detected; H is the height of the box to be detected; φ(·) is the linear projection operation; T 0 , P k is the projection feature after linear projection operation of x, k = 0, 1, ..., n-1; f k (·) is the kth convolution operation, k = 0, 1, ..., n-1; ⊙ is the dot product operation; P k Do convolution operation and add it with g s (T k ) and perform dot product operation to get T k+1 , k = 0, 1, ..., n-1; α is the reduction factor; g s is the dimension mapping function; C k is the dimension of the kth operation, k = 0, 1, ..., n-1; n is the number of convolutional layers; C c is the total dimension; FLOPs is the total amount of operations; g n Conv is the n-order spatial interaction capability; M is the scale of the convolution kernel;

[0118] The bounding box loss function is defined as follows:

[0119]

[0120] Where: L CIoU is the bounding box loss function; IoU is the ratio of the intersection area to the union area of ​​the real box and the predicted box; A is the area of ​​the real box; B is the area of ​​the predicted box; d and d gt are the coordinate positions of the center points of the predicted box and the real box respectively; ρ 2 (d,d gt ) is the square of the Euclidean distance between two center points; j 2 is the length of the diagonal of the minimum closed area formed by the predicted box and the real box; β is the penalty coefficient; v is the penalty term for the aspect ratio of the predicted box and the real box; w z With h z are the width and height of the real box respectively; w and h are the width and height of the predicted box;

[0121] (4.2) The graph attention mechanism GAT is introduced into the YOLO-POSE model to accurately weight and aggregate the neighboring key points of each key point. The GAT formula is as follows:

[0122]

[0123] Where: α ij is the attention coefficient between node i and node j; exp is an exponential function with the natural constant e as the base; LeakyReLU is a function; α T is the transpose of the weight parameter; W is the feature transformation weight matrix shared by each layer; h i is the feature vector of node i; h j is the feature vector of node j; h′ i is the feature of node i; τ is the activation function; N i is the number of neighbor nodes of node i; W k is the attention weight matrix.

[0124] (5) The multi-pose face recognition algorithm GAN-MFER based on generative adversarial networks is used to identify the driver’s facial expressions, and the dilated convolution DC is integrated into the algorithm to enhance the accuracy of identifying the driver’s pathological expressions;

[0125] (5.1) While performing driver posture recognition, the multi-posture face recognition algorithm GAN-MFER based on the generative adversarial network is used to recognize the driver's facial expression to improve the accuracy of judging the driver's current state. The formula is as follows:

[0126]

[0127] Where: D N The multi-pose face correction result output by the discriminator; The probability that the acquired image can be represented as a real picture; is the identity classification result of the identified image; D E The final output result of face rotation; The amount of loss incurred for correction; To determine the image input source; Image classification results; To classify facial expressions; is the residual attention function of the discriminator when generating face image classification; x is the image distribution data; c is the corresponding posture of the expression;

[0128] The two-way loop optimization function in the generative adversarial network is as follows:

[0129]

[0130]

[0131] Where: L c-1 represents the optimization function in the double loop of face profile-front-profile conversion; L c-2 represents the optimization function in the double loop of face front-side-front conversion; n is the number of loops; x i is the value distributed at node i; G E G is the profile face generator; N is a frontal face image;

[0132] (5.2) The hole convolution network DC is introduced into the multi-pose face recognition algorithm of the generative adversarial network to improve the recognition accuracy and detection efficiency. The formula is as follows:

[0133] f t spatial =E(I t )

[0134] f t temporal =P(f t-1 , f t-2 )

[0135] f t =S(f t spatial , f t temporal )

[0136] Where: f t spatial is the feature map calculated within two frames; E(·) is the spatial information estimator; I t is the input image; f t temporal is the calculated temporal feature map; P(·) is the ConvLSTM predictor; ft-1 、f t-2 is the feature map of two frames; f t is the final feature; S(·) is the fusion of spatiotemporal information.

[0137] (6) Based on the results of posture recognition and facial expression recognition, the driver’s current status can be accurately determined and corresponding alarm prompts can be given.

[0138] The results of gesture recognition and facial expression recognition are combined to determine the driver's current state. If the driver nods frequently or blinks for a long time, it is judged that the driver is dozing off and is in a fatigued driving state, and the system will automatically use the radio or speaker to remind the driver.

[0139] When the driver's body shakes violently, beats his chest, falls to the ground, or shows an expression of pain, it is judged that the driver has a sudden illness. The alarms inside and outside the car will sound to remind people inside and outside the car to take relevant measures to reduce the accident rate and casualties.

[0140] The driver's abnormal behavior and pain expression recognition and warning system of the present invention comprises:

[0141] A data acquisition module is used to extract key frames from the video data acquired by the high-definition camera in front of the cab to form an image;

[0142] An image enhancement module is used to reduce the noise of an image using an improved RPCA algorithm; the improved RPCA algorithm introduces the minimization of the weighted Schatten-p norm into RPCA to enhance the performance of RPCA in reducing the noise of an image; and combines the Laplace operator to enhance the details of the image to further improve the clarity of the image;

[0143] The object detection module is used to input the processed image into the Faster R-CNN model to generate the detection box of the target to be identified. At the same time, the Soft-NMS algorithm is introduced to process the target detection box to reduce the impact of redundant bounding boxes on the detection results;

[0144] The image recognition module is used to recognize the driver's facial expression using the multi-pose face recognition algorithm GAN-MFER based on the generative adversarial network, and integrate the dilated convolution DC into the algorithm to enhance the accuracy of recognizing the driver's pathological expression.

[0145] In addition, based on the same inventive concept, an embodiment of the present invention also provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for recognizing and warning abnormal behaviors and painful expressions of the driver are implemented.

[0146] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0147] In addition, based on the same inventive concept, an embodiment of the present invention also provides a storage medium storing a computer program, wherein the computer program is designed to implement the steps of the driver's abnormal behavior, pain expression recognition and warning method when running.

[0148] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0149] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0150] The computer-readable storage medium mentioned above may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0151] The specific description above further illustrates the purpose, technical solutions and beneficial effects of the invention in detail. It should be understood that the above is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for identifying and warning abnormal behaviors and painful expressions of drivers, characterized in that: The steps include: (1) Using the video data obtained by the high-definition camera in front of the cab, key frames are extracted from the video to form an image; (2) Using an improved RPCA algorithm to reduce the noise of the image; the improved RPCA algorithm introduces the minimization of the weighted Schatten-p norm into RPCA to enhance the performance of RPCA in reducing the noise of the image; and combining the Laplace operator to enhance the details of the image to further improve the clarity of the image; (3) The processed image is input into the Faster R-CNN model to generate the target detection box to be identified. At the same time, the Soft-NMS algorithm is introduced to process the target detection box to reduce the impact of redundant bounding boxes on the detection results; (4) Based on the YOLO framework, the YOLO-POSE model is established as a pose estimation and recognition model, and the GAT attention mechanism is introduced into the model to improve recognition accuracy; (5) The multi-pose face recognition algorithm GAN-MFER based on generative adversarial networks is used to identify the driver’s facial expressions, and the dilated convolution DC is integrated into the algorithm to enhance the accuracy of identifying the driver’s pathological expressions; (6) Integrate the results of posture recognition and facial expression recognition to accurately determine the driver's current state and give corresponding alarm prompts. The step (5) comprises: (5.1) While performing driver posture recognition, the multi-posture face recognition algorithm GAN-MFER based on the generative adversarial network is used to recognize the driver's facial expression to improve the accuracy of judging the driver's current state. The formula is as follows: Where: D N The multi-pose face correction result output by the discriminator; The probability that the acquired image can be represented as a real picture; is the identity classification result of the identified image; D E The final output result of face rotation; The amount of loss incurred for correction; To determine the image input source; Image classification results; To classify facial expressions; is the residual attention function of the discriminator when generating face image classification; x is the image distribution data; c is the corresponding posture of the expression; The two-way loop optimization function in the generative adversarial network is as follows: Where: L c-1 represents the optimization function in the double loop of face profile-front-profile conversion; L c-2 represents the optimization function in the double loop of the face front-side-front conversion; n is the number of loops; x i is the value distributed at node i; G E is a profile face generator; G N is a frontal face image; (5.2) The hole convolution network DC is introduced into the multi-pose face recognition algorithm of the generative adversarial network to improve the recognition accuracy and detection efficiency. The formula is as follows: f t spatial =E(I t ) f t temporal =P(f t-1 ,f t-2 ) f t =S(f t spatial ,f t temporal ) Where: f t spatial is the feature map calculated within two frames; E(·) is the spatial information estimator; I t is the input image; f t temporal is the calculated temporal feature map; P(·) is the ConvLSTM predictor; f t-1 、f t-2 is the feature map of two frames; f t is the final feature; S(·) is the fusion of spatiotemporal information.

2. The method for recognizing and warning abnormal behaviors and painful expressions of drivers according to claim 1 is characterized in that: The step (2) comprises: (2.1) The Schatten-p norm is used to achieve low-rank regularization, and the formula is as follows: Where: X is the low-rank matrix of the objective function; W is a non-negative vector; σ i is the i-th eigenvalue of X; n and m are the number of rows and columns of matrix X respectively; ω i is the i-th element of W; p is the weight; L is the augmented Lagrangian function; E is the identity matrix; Z is the Lagrangian multiplier; μ is a positive scalar; Y is the observed data; b is the weight vector; (2.2) The Laplace operator is used to enhance the details of the denoised image. The formula is as follows: Among them, g represents the gradient; (i, j) represents the coordinates of the pixel point; k and l are the grayscale values ​​in the horizontal and vertical directions respectively; (r, s) are the coordinates of the neighboring pixels of (i, j); f is a two-dimensional discrete function; and H is the directional derivative.

3. The method for recognizing and warning abnormal behaviors and painful expressions of drivers according to claim 1, characterized in that: The step (3) comprises: (3.1) Use the VGG16 feature extraction network to extract features from the image and obtain the feature map of the image. The loss function is as follows: Where: L cls is the foreground logarithmic loss; p i is the probability that the i-th pixel is predicted as the target; is the characteristic value of the i-th pixel; λ is the balance ratio; N is the total number of pixels; L reg is the logarithmic loss of the background; t i is the position of the i-th pixel; t i The offset of (3.2) The region of interest of the image is obtained through the selective search algorithm SS, that is, the candidate box is generated. The calculation formula is as follows: s=a1s color +a2s texture +a3s size +a4s fill Where: s color 、s texture 、s size 、s fill are the color similarity, texture similarity, size similarity and filling similarity of the current region block; s is the candidate box formed by the above four similarities; n is the total number of region blocks divided into the picture; k is the number of the region block; c and t are the color similarity and texture similarity of the current pixel point respectively; r is the region block being detected; B is the region containing r and its neighboring region blocks; i and j are distinguishing subscripts; a1, a2, a3, a4 are adoption weights, which are 0 or 1; (3.3) Perform ROIpooling pooling operation on the candidate box to obtain a feature box map of uniform size; (3.4) Perform non-maximum suppression on all candidate boxes, and use the Soft-NMS algorithm to replace the original NMS algorithm. During the execution of the algorithm, Soff-NMS does not simply delete the detection box whose IoU is greater than the threshold, but uses a linear weighted function to calculate and reduce its score. The function formula is as follows: Where: S i is the confidence score corresponding to the i-th prediction box; M is the candidate box with the highest score; box i is the frame to be detected; N t is the overlap threshold set by the hyperparameter; IoU is box i The ratio of the intersection and union of M.

4. The method for recognizing and warning abnormal behaviors and painful expressions of drivers according to claim 1, characterized in that: The step (4) comprises: (4.1) The YOLO-POSE model is established to identify the action recognition of the driver in a relatively fixed posture, which mainly includes the following three parts: Mark the key points and confidence levels of the human body in the frame to be detected, and calculate the similarity of the key points. The formula is as follows: Where: S is the target frame to be detected; C x , C y are the horizontal and vertical coordinates of the anchor point respectively; W is the width of the box to be detected; H is the height of the box to be detected; box conf is the confidence of the box to be detected; class conf is the category confidence of the box to be detected; are the horizontal and vertical coordinates of the i-th human key point; is the confidence of the i-th human key point, Take an integer; n is the number of key points of the human body in the frame to be detected; U is the similarity of the key points; exp is an exponential function with the natural constant e as the base; p is the number of the frame to be detected; p i is the number of the key point of the skeleton of the person; Indicates the visibility of the key point in the image; is the square of the Euclidean distance between the true point and the predicted point; S p Represents the square root of the area occupied by the detection box; σ i is the standard deviation of the key point; The YOLO-POSE model uses the HorNet network as the convolution kernel of the model, and uses high-order gated convolution and recursive design to achieve high-order spatial interaction. The relevant formula is as follows: T k+1 =f k (P k )⊙g s (T k ) / α,k=0,1,…,n-1 FLOPs(g n Conv)<HWC c (2M 2 +11 / 3×C c +2) x is the input feature of the gated convolution; W is the width of the box to be detected; H is the height of the box to be detected; φ(·) is the linear projection operation; T0, P k is the projection feature after linear projection operation of x, k = 0, 1, ..., n-1; f k (·) is the kth convolution operation, k = 0, 1, ..., n-1; ⊙ is the dot product operation; P k Do convolution operation and add it with g s (T k ) and perform dot product operation to get T k+1 , k = 0, 1, ..., n-1; α is the reduction factor; g s is the dimension mapping function; C k is the dimension of the kth operation, k = 0, 1, ..., n-1; n is the number of convolutional layers; C c is the total dimension; FLOPs is the total amount of operations; g n Conv is the n-order spatial interaction capability; M is the scale of the convolution kernel; (4.2) The graph attention mechanism GAT is introduced into the YOLO-POSE model to accurately weight and aggregate the neighboring key points of each key point. The GAT formula is as follows: Where: α ij is the attention coefficient between node i and node j; exp is an exponential function with the natural constant e as the base; LeakyReLU is a function; α T is the transpose of the weight parameter; W is the feature transformation weight matrix shared by each layer; h i is the feature vector of node i; h j is the feature vector of node j; h′ i is the feature of node i; τ is the activation function; N i is the number of neighbor nodes of node i; W k is the attention weight matrix.

5. The method for recognizing and warning abnormal behaviors and painful expressions of drivers according to claim 4 is characterized in that: The bounding box loss function defined in step (4.1) is as follows: Where: L CIoU is the bounding box loss function; IoU is the ratio of the intersection area to the union area of ​​the real box and the predicted box; A is the area of ​​the real box; B is the area of ​​the predicted box; d and d gt are the coordinate positions of the center points of the predicted box and the real box respectively; ρ 2 (d,d gt ) is the square of the Euclidean distance between two center points; j 2 is the length of the diagonal of the minimum closed area formed by the predicted box and the real box; β is the penalty coefficient; v is the penalty term for the aspect ratio of the predicted box and the real box; w z With h z are the width and height of the real box respectively; w and h are the width and height of the predicted box respectively.

6. The method for recognizing and warning abnormal behaviors and painful expressions of drivers according to claim 1, characterized in that: The step (6) comprises: The results of gesture recognition and facial expression recognition are combined to determine the driver's current state. If the driver nods frequently or blinks for a long time, it is judged that the driver is dozing off and is in a fatigued driving state, and the system will automatically use the radio or speaker to remind the driver. When the driver's body shakes violently, beats his chest, falls to the ground, or shows an expression of pain, it is judged that the driver has a sudden illness. The alarms inside and outside the car will sound to remind people inside and outside the car to take relevant measures to reduce the accident rate and casualties.

7. A driver's abnormal behavior and pain expression recognition and warning system, characterized in that: include: A data acquisition module is used to extract key frames from the video data acquired by the high-definition camera in front of the cab to form an image; An image enhancement module is used to reduce the noise of an image using an improved RPCA algorithm; the improved RPCA algorithm introduces the minimization of the weighted Schatten-p norm into RPCA to enhance the performance of RPCA in reducing the noise of an image; and combines the Laplace operator to enhance the details of the image to further improve the clarity of the image; The object detection module is used to input the processed image into the Faster R-CNN model to generate the detection box of the target to be identified. At the same time, the Soft-NMS algorithm is introduced to process the target detection box to reduce the impact of redundant bounding boxes on the detection results; The image recognition module is used to recognize the driver's facial expression using the multi-pose face recognition algorithm GAN-MFER based on the generative adversarial network, and integrates the dilated convolution DC into the algorithm to enhance the accuracy of the recognition of the driver's pathological expression, including: (5.1) While performing driver posture recognition, the multi-posture face recognition algorithm GAN-MFER based on the generative adversarial network is used to recognize the driver's facial expression to improve the accuracy of judging the driver's current state. The formula is as follows: Where: D N The multi-pose face correction result output by the discriminator; The probability that the acquired image can be represented as a real picture; is the identity classification result of the identified image; D E The final output result of face rotation; The amount of loss incurred for correction; To determine the image input source; Image classification results; To classify facial expressions; is the residual attention function of the discriminator when generating face image classification; x is the image distribution data; c is the corresponding posture of the expression; The two-way loop optimization function in the generative adversarial network is as follows: Where: L c-1 represents the optimization function in the double loop of face profile-front-profile conversion; L c-2 represents the optimization function in the double loop of face front-side-front conversion; n is the number of loops; x i is the value distributed at node i; G E is a profile face generator; G N is a frontal face image; (5.2) The hole convolution network DC is introduced into the multi-pose face recognition algorithm of the generative adversarial network to improve the recognition accuracy and detection efficiency. The formula is as follows: f t spatial =E(I t ) f t temporal =P(f t-1 ,f t-2 ) f t =S(f t spatial ,f t temporal ) Where: f t spatial is the feature map calculated within two frames; E(·) is the spatial information estimator; I t is the input image; f t temporal is the calculated temporal feature map; P(·) is the ConvLSTM predictor; f t-1 、f t-2 is the feature map of two frames; f t is the final feature; S(·) is the fusion of spatiotemporal information.

8. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for recognizing and warning abnormal behavior and painful expression of a driver according to any one of claims 1-6 is implemented.

9. A storage medium storing a computer program, characterized in that: The computer program is designed to implement the method for recognizing and warning abnormal behaviors and painful expressions of drivers according to any one of claims 1 to 6 when running.

Citation Information

Patent Citations

  • Facial expression recognition method and device, equipment and computer readable storage medium

    CN111783622A

  • Real-time state monitoring method for elderly living at home based on time deformable attention mechanism

    CN117218709A