Mine car driver facial posture recognition method and system
Through a lightweight algorithm model and dual feature consistency judgment method, the problems of limited computing power resources, high complexity and severe occlusion in facial posture recognition of mine car drivers in mines are solved, and efficient and accurate driver facial posture recognition is achieved.
Patent Information
- Application Number
- CN202211684224.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-12-27
AI Technical Summary
Existing technologies for facial posture recognition of mine car drivers in mines suffer from limited computing resources, high complexity, poor stability, and severe occlusion, resulting in low recognition accuracy.
A lightweight algorithm model combined with a dual feature consistency judgment method is used to obtain image frames, locate the head position, intercept facial key point information and auxiliary information, and use the scale factor and angle cosine value of the eyebrow and eye feature points for judgment to achieve driver facial posture recognition.
It can efficiently, accurately and stably identify the driver's facial posture on terminal devices with limited computing resources, reduce the false alarm rate, and is suitable for the complex environment of underground mines.
Smart Images

Figure CN115731537B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a visual detection technology, and in particular to a method and system for recognizing the facial posture of a mine car driver. Background Art
[0002] To monitor the driver's condition, facial image metrics can be analyzed, such as eye closure, blink frequency, and yawn frequency. These metrics are typically calculated based on a frontal facial image. Therefore, before performing these analysis processes, facial posture estimation is required to improve analysis accuracy.
[0003] Existing technologies use the entire head image area above the shoulders to estimate the Euler angles of the head. However, this technology has the following shortcomings in the actual minecart scene:
[0004] a) If the eye and mouth areas need to be located later, one or more detectors must be cascaded, making the entire technical solution difficult to deploy on devices with limited computing resources.
[0005] b) High complexity, which increases time consumption.
[0006] Other methods use a 3D facial reference model to approximate the rotation vector. However, this method is unstable and suffers from large deviations in real-world scenarios, making it difficult to determine with a fixed threshold. Other methods rely solely on the geometric relationships between key facial points, which can be ineffective when the face is expressed in exaggerated expressions or severely obscured (e.g., when wearing a mask).
[0007] In mines, mine truck drivers typically wear helmets and masks, severely obscuring their faces and complicating facial gesture recognition. Given the mine truck operating environment, improving the accuracy of driver facial gesture recognition in these scenarios, and ensuring reliable analysis of the driver's facial attributes, is a pressing challenge. Summary of the Invention
[0008] To solve the problems in the prior art, the present invention provides a method and system for facial posture recognition of mine car drivers, which combines classification and regression tasks to jointly determine the driver's facial posture, which is more efficient, accurate and stable than existing methods.
[0009] The method for recognizing the facial posture of a mine car driver of the present invention comprises the following steps:
[0010] S1: Acquire image frame;
[0011] S2: Locate the specific position of the driver's head in the image frame;
[0012] S3: Capture the head area image to obtain facial key point information and auxiliary information;
[0013] S4: Based on the dual features of facial key point information and auxiliary information, the driver's facial posture is judged according to the dual feature consistency judgment method.
[0014] The present invention is further improved. In step S1, the imaging device used to obtain the image frame to be detected includes but is not limited to an infrared camera unit based on a CCD sensor or a CMOS sensor.
[0015] The present invention is further improved, and step S1 also includes a preprocessing operation step for the image frame, and the image frame preprocessing operation step includes intercepting sub-images, normalization, denoising, equalization, and scale transformation to obtain an image of a set size.
[0016] The present invention is further improved, and the specific method of the scale transformation operation is:
[0017] a. Select the larger side of the image width and height;
[0018] b. Calculate the width-to-height ratio factor;
[0019] c. scaling the larger side to a preset size, and calculating the new size of the other shorter side based on the aspect ratio factor;
[0020] d. Scale the image to the new scale;
[0021] e. The short side border is filled with 0 pixel value.
[0022] The present invention is further improved. In step S2, the method for locating the specific position of the driver's head is as follows:
[0023] Locate the face position;
[0024] Expand the face position toward the image boundary until it covers the entire head area or reaches the image boundary.
[0025] The present invention is further improved. In step S3, the facial key points include key point information of the face contour, chin, mouth, nose, eyes and eyebrows, and the auxiliary information is a predicted value used for consistency comparison with the result calculated based on the facial key points.
[0026] The present invention is further improved. In step S4, the method for determining the driver's facial posture based on the dual feature consistency determination method is:
[0027] S401: Calculating the scaling factor of the Euler distance between the left and right eyebrow feature points;
[0028] S402: Determine whether the scale factor is greater than a first preset value and is consistent with the auxiliary information; if not, determine it as a profile face; if yes, determine it as a frontal face, and then proceed to the next step;
[0029] S403: Calculate the cosine value of the angle between the left and right eye center feature points and their horizontal projection points;
[0030] S404: Determine whether the cosine value of the angle is less than a second preset value and is consistent with the auxiliary information. If so, determine that it is a frontal face; otherwise, determine that it is a tilted face.
[0031] The present invention is further improved. In steps S401 and S402, the specific method for determining based on the scale factor and the auxiliary information is:
[0032] Select the key points at both ends of the left eyebrow of the face and calculate the Euler distance between them as the length of the left eyebrow; select the key points at both ends of the right eyebrow of the face and calculate the Euler distance between them as the length of the right eyebrow;
[0033] Calculate the proportional factor of eyebrow length, where min is the function of taking the smaller value and max is the function of taking the larger value.
[0034] If the proportional factor is greater than the first preset value and the comparison result is consistent with the auxiliary information, it indicates that the driver's facial posture is a frontal face, otherwise it is a side face.
[0035] The present invention is further improved. In step S403 and step S404, the specific method for determining based on the central feature points of the left and right eyes and the auxiliary information is as follows:
[0036] Select the key points of the left eye of the face, and calculate the minimum circumscribed rectangle of these key points, and then determine the intersection of the main diagonal and the secondary diagonal of the minimum circumscribed rectangle as the left eye center point O left ;
[0037] Select the key points of the right eye of the face, and calculate the minimum bounding rectangle of these key points, and then determine the intersection of the main diagonal and the secondary diagonal of the minimum bounding rectangle as the right eye center O right ;
[0038] Connect the center point O left and O right For line segment O left O right ;
[0039] Line segment O left O right Projected to the horizontal direction, the projection point is O project ;
[0040] Calculate line segment O left O right The cosine of the angle θ with the horizontal projection:
[0041] When it is a left projection:
[0042]
[0043]
[0044] When it is right projection:
[0045]
[0046]
[0047] The angle θ is compared with a second preset value to determine the degree of inclination of the driver's face. If the angle θ is consistent with the auxiliary information, the driver's facial posture is determined to be straight or tilted.
[0048] The present invention also provides a system for implementing the above-mentioned method for recognizing the facial posture of a mine car driver, comprising:
[0049] Image receiving module: used to obtain image frames;
[0050] Positioning module: used to locate the specific position of the driver's head in the image frame;
[0051] Information acquisition module: captures the head area image and obtains facial key point information and auxiliary information;
[0052] Posture determination module: Based on the dual features of facial key point information and auxiliary information, the driver's facial posture is determined according to the dual feature consistency judgment method.
[0053] Compared with the prior art, the present invention has the following beneficial effects:
[0054] 1. The algorithm model provided by this method requires very few parameters and is very lightweight. It occupies few system resources during operation and can be deployed on terminal devices with limited computing resources.
[0055] 2. Through the dual feature consistency judgment method, the driver's facial posture can be accurately and stably identified in the mine car scene. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 Flow chart of the method of the present invention;
[0057] Figure 2 This is a flow chart of a method according to an embodiment of the present invention;
[0058] Figure 3 A schematic diagram of key points of human eyes and eyebrows;
[0059] Figure 4 This is a schematic diagram of calculating the human eye angle θ according to the present invention;
[0060] Figure 5 Schematic diagram of eyebrow length;
[0061] Figure 6 Schematic diagram of the convolutional neural network algorithm model structure. DETAILED DESCRIPTION
[0062] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.
[0063] like Figure 1 As shown, because the mine car onboard terminal has many tasks to perform at the same time, there is usually not enough additional computing power resources to run complex algorithm models. One of the key points of the present invention is to ensure that the algorithm model is lightweight enough; on the other hand, in the mine, the mine car driver usually wears a helmet and a mask, resulting in extremely serious facial occlusion, which increases the difficulty of facial gesture recognition. Therefore, the present invention needs to improve the driver's facial gesture recognition accuracy in the mine scene to provide a reliable guarantee for the subsequent driver's facial attribute analysis. Based on this, a facial gesture recognition method for mine car drivers is proposed, comprising the following steps:
[0064] S1: Acquire image frame;
[0065] S2: Locate the specific position of the driver's head in the image frame;
[0066] S3: Capture the head area image to obtain facial key point information and auxiliary information;
[0067] S4: Based on the dual features of facial key point information and auxiliary information, the driver's facial posture is judged according to the dual feature consistency judgment method.
[0068] The algorithm model provided by this method requires very few parameters and is very lightweight. It occupies few system resources during operation and can be deployed on terminal devices with limited computing resources. Through the dual feature consistency judgment method, the driver's facial posture can be accurately and stably identified in the mine car scenario.
[0069] like Figure 2 As shown in FIG. 1 , as an embodiment of the present invention, the detailed implementation process of the present invention is as follows:
[0070] S1: Get image frame
[0071] This example obtains an image frame from an external infrared camera configured on the vehicle terminal, and then detects whether the image frame contains the driver, that is, whether there is someone in the driver's seat. The judgment conditions for whether there is someone in the driver's seat are:
[0072] Based on the imaging characteristics of the video board and the camera's installation position, the driver's image area occupies more than one-third of the entire image frame. This step has two advantages: first, it can filter out abnormal situations; second, it can obtain a larger and clearer image.
[0073] The imaging device used to obtain the image frame to be detected includes but is not limited to an infrared camera unit based on a CCD (Charge Coupled Device) sensor or a CMOS (Complementary Metal-Oxide-Semiconductor) sensor. Furthermore, the infrared camera unit may be equipped with an element with a photosensitivity function to control the automatic switching of the infrared night vision function; the imaging device can merge images captured by multiple camera units, and the full image resolution is 1280×720.
[0074] Preferably, in order to make the image clearer and reduce redundant information, the acquired image frame is preprocessed, including sub-image capture, normalization, denoising, equalization, and scale transformation.
[0075] In this example, the sub-image is captured by selecting sub-images from one or more camera acquisition units from the full image. Interference noise is removed from the captured image using filters including, but not limited to, a Gaussian low-pass filter and a median filter, which obviously depends on the type of noise present in the image. Optionally, if the image is dark overall, histogram equalization may be performed to improve the image contrast. A normalization operation is performed to limit the pixel value range to between 0 and 1, and the image is scaled to a preset size through a scaling operation. Regarding the scaling operation, since the preset size has equal length and width, direct scaling would result in image content distortion. Therefore, the scaling method used in this embodiment of the present invention is:
[0076] a) Select the larger side of the image width and height;
[0077] b) Calculate the width-to-height ratio factor;
[0078] c) scaling the larger side to a preset size, and calculating a new size of the other shorter side according to the aspect ratio factor;
[0079] d) Scale the image to the new scale;
[0080] e) The short side borders are filled with 0 pixel values.
[0081] S2: Locate the specific position of the driver's head in the image frame
[0082] This example locates the driver's position and then further locates the specific position of the driver's head. This is achieved through the following two steps:
[0083] a) Locate the face position;
[0084] b) Expand the face position toward the image boundary until it covers the entire head area or reaches the image boundary.
[0085] Specifically, in this example, the pre-processed image frame is sent to the algorithm model to determine whether the image frame contains a face. The algorithm model can use but is not limited to machine learning HOG+SVM, HAAR+SVM or deep learning model (HOG is the Histogram of Oriented Gradient directional gradient histogram feature, HAAR is a feature that reflects the grayscale change of the image, and SVM is a support vector machine classifier). Exemplarily, in an embodiment of the present invention, the algorithm model for detecting faces is provided by the libfacedetection library, and the algorithm model is converted on this basis. The conversion operation includes simplification and merging operators to improve the calculation speed of the algorithm model in the vehicle terminal.
[0086] If the image frame contains a face, the algorithm model for detecting faces outputs the location frames of all faces, and each location frame is represented by x, y, w, and h, where x represents the horizontal coordinate of the upper left corner of the location frame, y represents the vertical coordinate of the upper left corner of the location frame, w represents the width of the location frame, and h represents the height of the location frame.
[0087] For example, the camera acquisition unit is typically installed at any location in front of the driver where it can capture the driver's facial image. The output face position frame is divided into a first face position frame and a second face position frame based on the relative imaging position of the camera acquisition unit and the driver. Specifically, the steps include:
[0088] a) Calculate the area of all face location boxes;
[0089] b) sorting all face location frames in descending order according to their areas, and selecting the two with the largest areas;
[0090] c) According to the relative position, the right position frame is selected as the first face position frame, and the remaining original output position frames are all second face position frames. The first face position frame corresponds to the driver's face position frame.
[0091] According to the scale transformation method in step S1, the coordinates of the first face position frame are inversely transformed back to the position of the original RGB image. It should be noted that the inverse transformation is performed on the upper left corner coordinates and the lower right corner coordinates of the first face position frame, and the calculation method of the lower right corner coordinates is (x+w, y+h). The inversely transformed first face position frame is enlarged to cover the entire head area, and then the driver's head image is intercepted in the original RGB image. According to the normalization method and scale transformation method in step S1, the pixel values of the driver's head image are normalized to between 0 and 1, and the width and height are scaled to 112×112.
[0092] S3: Capture the head area image to obtain facial key point information and auxiliary information
[0093] The head area image is intercepted and then fed into the algorithm model to output facial key points and auxiliary information. There are 98 facial key points in total, including facial contour, chin, mouth, nose, eyes and eyebrows. Figure 3 As shown in Figure 2, the auxiliary information in this example is a predicted value, which is used to compare the consistency of the result calculated based on the facial key points. This consistency comparison can improve the accuracy and reduce the false positive rate. This consistency comparison is called the dual feature consistency judgment method.
[0094] like Figure 6 As shown, the algorithm model of this example adopts a convolutional neural network algorithm model, with the input being a normalized driver head image A of size 112×112, a backbone convolutional network B, a classification branch C, and a regression branch D. The regression branch outputs a 196-dimensional vector, with each key point having two values, horizontal and vertical coordinates, and the classification branch outputs a 7-dimensional vector. The backbone convolutional network can use, but is not limited to, lightweight networks such as SqueezeNet, MobileNet, ShuffleNet, and their various variants. The 7-dimensional vectors represent 7 postures of the head, including forward, left turn, right turn, left tilt, right tilt, head down, and head up. The predicted value p in this example is calculated as follows:
[0095]
[0096] p=argmax(U)
[0097] Where V represents the original 7-dimensional vector, U is the transformed 7-dimensional vector, argmax is the function that obtains the index corresponding to the maximum value of the vector element, and the predicted value p is an integer between 0 and 6, which can be used to index the head posture.
[0098] S4: Based on the dual features of facial key point information and auxiliary information, the driver's facial posture is determined according to a dual feature consistency determination method. The method for determining the driver's facial posture according to the dual feature consistency determination method is:
[0099] S401: Calculating the scaling factor of the Euler distance between the left and right eyebrow feature points;
[0100] S402: Determine whether the scale factor is greater than a first preset value and is consistent with the auxiliary information; if not, determine it as a profile face; if yes, determine it as a frontal face, and then proceed to the next step;
[0101] S403: Calculate the cosine value of the angle between the left and right eye center feature points and their horizontal projection points;
[0102] S404: Determine whether the cosine value of the angle is less than a second preset value and is consistent with the auxiliary information. If so, determine that it is a frontal face; otherwise, determine that it is a tilted face.
[0103] like Figure 3 and 5 As shown, the method for judging the frontal face posture in this example is as follows:
[0104] Select the key points with index numbers 46 and 50 on the left eyebrow of the face, and calculate the Euler distance between them as the length L of the left eyebrow left ;
[0105] Select the key points with index numbers 33 and 38 on the right eyebrow of the face, and calculate the Euler distance between them as the length L of the right eyebrow right .
[0106] Calculate the scale factor for eyebrow length That is, divide the smaller of the two by the larger one, where min is the function that takes the smaller value and max is the function that takes the larger value.
[0107] If the proportional factor ω is greater than the first preset value and the comparison result is consistent with the auxiliary information, it indicates that the driver's facial posture is a frontal face, otherwise it is a side face.
[0108] like Figure 3 and Figure 4 As shown, if the driver's facial posture is determined to be a straight face, the degree of facial inclination can be further determined. The specific process is as follows:
[0109] a) Select all key points around the left eye of the face (including key points 68, 72 and two groups of 6 above and below 68 and 72), and calculate the minimum bounding rectangle of these key points, and then determine the intersection of the main diagonal and the secondary diagonal of the minimum bounding rectangle as the left eye center point O left ;
[0110] b) Select all key points around the right eye of the face (including key points 60, 74 and two groups of 6 above and below 64), and calculate the minimum bounding rectangle of these key points, and then determine the intersection of the main diagonal and the secondary diagonal of the minimum bounding rectangle as the right eye center O right ;
[0111] c) Connect the center point O left and O right For line segment O left O right ; d) Line segment O left O right Projected to the horizontal direction, the projection point is O project ; e) Calculate line segment O left O right The cosine of the angle θ with the horizontal projection:
[0112] When it is a left projection:
[0113]
[0114]
[0115] When it is right projection:
[0116]
[0117]
[0118] The angle θ is compared with a second preset value to determine the degree of inclination of the driver's face.
[0119] In this example, if the angle 0 ≤ θ ≤ 30, the facial posture is initially judged to be within the normal range. If it is consistent with the auxiliary information, the driver's facial posture is determined to be straight; otherwise, it is tilted. In practice, the first and second preset values can be adjusted according to the situation to select an appropriate threshold.
[0120] This example also provides a system for implementing the method for recognizing the facial posture of a mine car driver, comprising:
[0121] Image receiving module: used to obtain image frames;
[0122] Positioning module: used to locate the specific position of the driver's head in the image frame;
[0123] Information acquisition module: captures the head area image and obtains facial key point information and auxiliary information;
[0124] Posture determination module: Based on the dual features of facial key point information and auxiliary information, the driver's facial posture is determined according to the dual feature consistency judgment method.
[0125] The module divisions in the embodiments of the present invention are illustrative and represent only one logical functional division. Actual implementations may employ different divisions. Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically as a separate unit, or two or more units may be integrated into a single module. These integrated modules may be implemented in either hardware or software functional modules.
[0126] Preferably, the frame rate of the infrared camera is usually 25 frames per second, and the similarity between adjacent frames is very high. There is no need to process every frame of the image. The present invention processes every 2 to 3 frames of image, which reduces the load on the processor and improves the fluency of the entire system.
[0127] In summary, compared with the prior art, the present invention has the following innovations:
[0128] 1. Efficiency: A single algorithm model can be used to determine the driver's facial posture and the location of key facial organs, with low complexity.
[0129] 2. Accuracy: It has high accuracy in mine car driving scenarios;
[0130] 3. Stability: The first feature is calculated through key points, and the auxiliary information is the second feature. Through the dual features and the dual feature consistency judgment method, the false alarm rate in extreme cases is greatly reduced;
[0131] 4. The present invention has a lightweight design, consumes less resources, and is more suitable for terminal devices with limited computing resources.
[0132] The specific implementation manner described above is a preferred implementation manner of the present invention, and is not intended to limit the specific implementation scope of the present invention. The scope of the present invention includes but is not limited to this specific implementation manner. All equivalent changes made in accordance with the present invention are within the protection scope of the present invention.
Claims
1. A method for recognizing the facial posture of a mine car driver, characterized in that: The steps include: S1: Acquire image frame; S2: Locate the specific position of the driver's head in the image frame; S3: Capture the head area image to obtain facial key point information and auxiliary information; S4: Based on the dual features of facial key point information and auxiliary information, the driver's facial posture is judged according to the dual feature consistency judgment method. In step S4, the method for determining the driver's facial posture based on the dual feature consistency determination method is: S401: Calculating the scaling factor of the Euler distance between the left and right eyebrow feature points; S402: Determine whether the scale factor is greater than a first preset value and is consistent with the auxiliary information; if not, determine it as a profile face; if yes, determine it as a frontal face, and then proceed to the next step; S403: Calculate the cosine value of the angle between the left and right eye center feature points and their horizontal projection points; S404: Determine whether the cosine value of the angle is less than a second preset value and is consistent with the auxiliary information. If so, determine that it is a frontal face; otherwise, determine that it is a tilted face.
2. The method for facial gesture recognition of a mine car driver according to claim 1, wherein: In step S1 , the imaging device used to obtain the image frame to be detected includes but is not limited to an infrared camera unit based on a CCD sensor or a CMOS sensor.
3. The method for facial gesture recognition of a mine car driver according to claim 2, wherein: Step S1 also includes a pre-processing operation step for the image frame, which includes intercepting sub-images, normalization, denoising, equalization, and scale transformation to obtain an image of a set size.
4. The method for facial gesture recognition of a mine car driver according to claim 1, wherein: In step S2, the method for locating the specific position of the driver's head is as follows: Locate the face position; Expand the face position toward the image boundary until it covers the entire head area or reaches the image boundary.
5. The method for facial gesture recognition of a mine car driver according to claim 1, wherein: In step S3, the facial key points include facial contour, chin, mouth, nose, eyes and eyebrows key point information, and the auxiliary information is a predicted value used for consistency comparison with the result calculated based on the facial key points.
6. The method for facial gesture recognition of a mine car driver according to any one of claims 1 to 5, characterized in that: In steps S401 and S402, the specific method for determining based on the scale factor and the auxiliary information is: Select the key points at both ends of the left eyebrow of the face and calculate the Euler distance between them as the length L of the left eyebrow left ; Select the key points at both ends of the right eyebrow of the face and calculate the Euler distance between them as the length L of the right eyebrow right ; Calculate the scale factor for eyebrow length Among them, min is the function of taking the smaller value, and max is the function of taking the larger value. If the proportional factor ω is greater than the first preset value and the comparison result is consistent with the auxiliary information, it indicates that the driver's facial posture is a frontal face, otherwise it is a side face.
7. The method for recognizing the facial posture of a mine car driver according to any one of claims 1 to 5, characterized in that: In step S403 and step S404, the specific method for determining based on the left and right eye center feature points and the auxiliary information is: Select the key points of the left eye of the face, and calculate the minimum circumscribed rectangle of these key points, and then determine the intersection of the main diagonal and the secondary diagonal of the minimum circumscribed rectangle as the left eye center point O left ; Select the key points of the right eye of the face, and calculate the minimum bounding rectangle of these key points, and then determine the intersection of the main diagonal and the secondary diagonal of the minimum bounding rectangle as the right eye center O right ; Connect the center point O left and O right For line segment O left O right ; Line segment O left O right Projected to the horizontal direction, the projection point is O project ; Calculate line segment O left O right The cosine of the angle θ with the horizontal projection: When it is a left projection: When it is right projection: The angle θ is compared with a second preset value to determine the degree of inclination of the driver's face. If the angle θ is consistent with the auxiliary information, the driver's facial posture is determined to be straight or tilted.
8. A system for implementing the method for recognizing the facial gesture of a mine car driver according to any one of claims 1 to 7, characterized in that: include: Image receiving module: used to obtain image frames; Positioning module: used to locate the specific position of the driver's head in the image frame; Information acquisition module: captures the head area image and obtains facial key point information and auxiliary information; Posture determination module: Based on the dual features of facial key point information and auxiliary information, the driver's facial posture is determined according to the dual feature consistency judgment method.
Citation Information
Patent Citations
Human face key point positioning method and device, storage medium and electronic device
CN112257645A
Facial expression recognition method and device and storage medium
CN112560685A