Lightweight driver fatigue monitoring method and system
By monitoring the light intensity in the vehicle and adjusting the light, combining a lightweight training model to identify the driver's eyes and mouth states, the imaging problem of the driver monitoring system in different light environments is solved, the computing resource consumption and false alarm rate are reduced, and monitoring accuracy is improved.
Patent Information
- Application Number
- CN202510792406.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing driver monitoring system has a decline in imaging quality in strong, backlight or low-light environments, which affects driver status judgment, and the computing resources consumed in the deep learning process, resulting in high requirements for on-board equipment.
By monitoring the light intensity in the vehicle, adjusting fill light or reducing light using light intensity control algorithms, combining the "eye diagram" concept to identify the state of the eyes and mouth, building a lightweight training model to reduce the use of computing resources.
Reduce interference and false alarm rates in different light environments, reduce on-board computing resources, and improve the accuracy and efficiency of driver fatigue monitoring.
Smart Images

Figure CN120340004A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of safe driving technology, and in particular to a lightweight driver fatigue monitoring method and system. Background Art
[0002] Driving fatigue refers to the physiological or psychological dysfunction of the driver (i.e., physical or psychological fatigue) caused by various reasons while driving a vehicle, which weakens the driver's ability to perceive the surrounding environment and the ability to control the vehicle, deviating from normal driving behavior. Therefore, how to actively monitor the driver's status and effectively provide driving fatigue warning is a traffic safety issue that needs to be solved urgently.
[0003] Most of the existing driver monitoring systems (DMS) work based on visual signal processing and feature extraction to determine the driver's status. However, there are some inherent problems in practical applications. For example, as the most important data source, the video images captured by the vehicle-mounted camera are easily affected by ambient light. In strong light, backlight or low light environments, the image quality will be reduced to varying degrees, affecting the judgment of the driver's status. The PERCLOS (Percentage of Eyelid Closure over the pupil over time) algorithm is used to subtract two adjacent frames of facial images captured by the camera sensor, perform image binarization, and segment the facial area, which has a certain effect. However, once a person shakes his head or the body position changes a large distance, the recognition effect will be affected. The key point positioning method combined with the mainstream deep neural network algorithm can obtain relatively accurate results, but the entire deep learning process from data set labeling to training and reasoning is very labor-intensive, and after being deployed on the vehicle side, it consumes more computing resources and has higher requirements for vehicle-mounted equipment. Summary of the invention
[0004] The present application is made in view of the above-mentioned problems, and its purpose is to provide a lightweight driver fatigue monitoring method and system. By monitoring the light intensity in the car, a light intensity control algorithm is proposed, and intelligent fine-tuning is performed to control the fill light or dimming light, so as to do "subtraction" for the fatigue detection area, remove the non-strongly related parts, and let the system focus on the parts that need the most attention. In combination with the concept of "eye diagram" in communication, the state of the eyes and mouth is judged, which reduces interference and false alarm rate to a certain extent, reduces the occupancy rate of on-board computing resources, and facilitates more computing resources to be tilted towards vehicle control logic.
[0005] Specifically, the first aspect of the present application provides a lightweight driver fatigue monitoring method, comprising the following steps: Step 1: Collect the light intensity inside the vehicle. Based on the set light intensity threshold range, when the light intensity is not within the range, adjust the light intensity through the adjustment device; Through the light intensity sensor, sense the light intensity of the environment inside and outside the vehicle. Using the dual-threshold mechanism, when the light intensity inside the vehicle is higher than Lux, drive to adjust the angle of the dimming mirror to reduce the amount of light passing through the camera lens; when the light intensity inside the vehicle is lower than Lux, turn on several infrared lights for supplementary lighting to adjust and supplement the lighting for the camera to collect videos. Generally, and should have a relatively obvious difference.
[0006] Both the dimming mirror and the infrared lights are arranged on the image acquisition module. The dimming mirror adjusts the dimming effect by rotating the angle. For each rotation of the angle , the camera correspondingly reduces the amount of light Lux, and the default is 0° without dimming; the number of infrared lights is N. After each one is turned on, it is equivalent to increasing Lux of light for the camera. Control the number of lights turned on according to the light intensity inside the vehicle, and the default is all off without supplementary lighting.
[0007] Step 2: Perform real-time image acquisition and preprocessing on the driver's seat area inside the vehicle; Step 3: Construct a lightweight training model, preprocess the training portrait data and label the key parts, and use the data to train the model; Step 4: Determine the candidate regions in the image data of the driver's seat area based on the mask concept, and identify the key parts through the lightweight training model; Step 5: Set the determination threshold for the change degree of the key parts, and set the over-limit value of the determination index within a unit time. When the determination index reaches the over-limit value, an alarm is issued; Step 6: Perform bypass shielding and local calibration on false alarm situations; Furthermore, in the said Step 1, the light intensity adjustment formula is as follows: ; Where: is the light intensity adjustment decision strategy; is the switch coefficient of the dimming adjustment strategy control, and it is an element in the set {0, 1}; is the switch coefficient of the light enhancement adjustment strategy control, and it is an element in the set {0, 1}; is the actual light intensity inside the vehicle; is the upper limit of the light intensity threshold; is the amount of light reduced when the dimming device adjusts by a unit angle; For adjusting the angle; Is the lower limit of the light intensity threshold; Is the increased light amount when each supplementary light device is lit; Is for floor function; Is for ceiling function.
[0008] If , that is, when the light intensity inside the vehicle is higher than , the image may be overexposed, then in the formula , , do not execute the supplementary light strategy, only adjust the light reduction of the light reduction mirror; If , that is, when the light intensity inside the vehicle is lower than , it may not be possible to see clearly, then in the formula , , do not execute the light reduction strategy, only adjust the supplementary light of the supplementary light lamp; Assume that after the light intensity adjustment strategy, the light intensity inside the vehicle can reach the appropriate range. After adjustment, the situation of not being able to be outside the appropriate range is not considered temporarily; If , that is, when the light intensity inside the vehicle is appropriate, there is no need to adjust the light input amount of the camera. In the formula , , there is no need to adjust the light reduction or supplementary light strategy.
[0009] Furthermore, the preprocessing specifically includes image grayscale processing and binary sparse.
[0010] Obtain the original data, first store it in the local storage module 1, and at the same time, transmit it to the recognition processing module for core function processing. After obtaining a fixed number of video frames, form a batch and perform unified parallel grayscale processing to obtain the corresponding binary image, converting the R, G, B three-channel image into a single-channel image, reducing the data volume by 2 / 3 dimensions.
[0011] The binary sparse is specifically to further sparse the image by the pixel threshold method, obtain the non-blank information in the image, and by adjusting the empirical threshold, using the ability to distinguish the opening and closing states of eyes and mouths in most images as a benchmark, filter out those below the value and retain those above the value, thereby obtaining the binary processed image. The binary processed image can further reduce the memory space occupied and the processing overhead, and at the same time can reduce noise interference.
[0012] Furthermore, the image grayscale processing is specifically to compress and convert the R, G, B three-channel image into a single-channel image. The value-taking formula for each pixel point during the compression process is as follows: ; Wherein: is a single-channel image; are all regulatory factors for extracting key parts; is a red-channel image; is a green-channel image; is a blue-channel image.
[0013] Regulatory factors are added as adjustable parameters to the classical conversion formula, aiming to clearly extract the key parts of the eyes and lips. For parts such as the eyes and lips, which are mainly based on red and green primary colors, the weight value of the regulatory factor can be appropriately increased.
[0014] Furthermore, the annotation of the key parts specifically includes: the elements in the portrait data are relatively concentrated. When the driver is fatigued and sleepy, the eyes and mouth, which are significantly different from the facial expressions in the normal awake state, are selected as the key parts, and the key parts are marked and recognized.
[0015] The eyes and mouth are significantly different from the facial expressions in the normal awake state when the driver is sleepy, and the recognition result has a high reliability. Therefore, the above two parts are focused on for recognition. By marking the eyes and mouth parts of the driver in the figure, supervised learning training is implemented.
[0016] The selection of materials for the portrait data used in training is not limited to driver images. As long as it is a picture containing a human head, it can be adopted as the training set, mainly focusing on the states of the eyes and mouth. The image is binarized mainly to compress the storage space and reduce the processing overhead. When manually annotating, the binarized image may not be as clear and distinct as the original image, which is not conducive to accurate annotation. Therefore, in the training stage, direct annotation is performed on the original image to obtain the position information of the annotation Label box, and the binarized image is used as the corresponding image in the training set Train to complete the mapping of the training set.
[0017] Furthermore, the determination of the candidate region of the image data based on the mask concept specifically includes: according to the installation orientation of the camera, the image collected under the normal driving condition of the driver is used as the mask, and considering the body tilt margin, the position range of the driver is delimited, and the rest is cleared.
[0018] Furthermore, the body tilt margin is specifically to expand the driver image by 5% - 15% around.
[0019] The main acquisition range of the image acquisition module is mainly the upper body of the driver, and also includes the part behind him. If there are passengers sitting in the back row, they will also be acquired, which will interfere with the driver recognition. Therefore, drawing on the commonly used "mask" concept in image processing, according to the installation orientation of the camera, the image acquired under the normal driving condition of the driver is used as the mask, and considering the body tilt margin, the driver's position is roughly delimited. The rest is cleared to reduce the recognition interference caused by the entry of the rear passengers into the frame. It should be noted that the image is not cropped to maintain the consistency of the size before and after, which is convenient for unified processing.
[0020] After being trained by the convolutional neural network, after the system is put into use and the driver's image information is acquired, it can identify the opening and closing states of the driver's eyes and mouth through reasoning to initially determine the driver's state at each moment.
[0021] Furthermore, a determination threshold is set for the degree of change of the key parts, and an overlimit value of the determination index within a unit time is set. Specifically, drawing on the "eye diagram" concept in communication, a determination threshold is set for eye closure and mouth opening. The blink count is preferentially used as the determination index, and the mouth opening count is used as a progressive supplementary determination index, and the overlimit values of these two indexes are set.
[0022] Furthermore, an alarm is issued when the determination index reaches the overlimit value. Specifically, a progressive determination method is adopted. Within a unit time, first, the blink count is judged. If it exceeds the limit, it is considered that the driver may have a fatigue problem and an alarm is issued; if it does not exceed the limit, then the mouth opening count is turned to. If this index exceeds the limit, the system determines that the driver may have a fatigue problem and an alarm is issued.
[0023] After training, the system can identify the states of eyes opening, closing, mouth opening, and closing. Drawing on the "eye diagram" concept in communication, a determination threshold is set for eye closure and mouth opening. Within a unit time, a progressive determination method is adopted. The blink count is preferentially used as the determination index, and the mouth opening count is used as a progressive supplementary determination index. That is, first, the blink count is determined. If it exceeds the limit, it is considered that the driver may have a fatigue problem and an alarm is issued; if it does not exceed the limit, then the mouth opening count is turned to. If this index exceeds the limit, the system determines that the driver may have a fatigue problem and an alarm is issued. For the determination of the overlimit values of the two indexes, there is much research in biology in this regard, and relatively mature recommended reference values can be adopted.
[0024] When it is impossible to detect the mouth or eyes, such as when the driver is eating, drinking, covering the mouth, or wearing sunglasses to cover the eyes, it is recorded as an abnormal situation. In rare cases, the driver has physiological characteristics or habits such as extremely small eyes or habitual mouth opening, etc., which can bypass and shield false alarms and at the same time feedback to the system. During the idle stage, the system uses the data stored during driving, combined with the driver's bypass feedback, to perform secondary training to calibrate the system's determination.
[0025] The second aspect of the present application provides a lightweight driver fatigue monitoring system, including: Light sensing module: used to monitor the light intensity inside the vehicle in real time and transmit the signal to the recognition and processing module; Light supplement and reduction module: used to control the infrared fill light and the light reduction mirror, and adjust the light intensity according to the light intensity control calculation of the recognition and processing module; Image acquisition module: monitors the key features of the driver's face, obtains the original data, and sends it to the recognition and processing module and the local storage module 1; Local storage module 1: stores the original data transmitted by the image acquisition module; Recognition and processing module: used for light intensity control calculation and recognition training of key parts of the face, controls the reminder and warning module when detecting and determining fatigue, and uses the data of the local storage module 2 for local calibration training when feedback false alarms occur; Local storage module 2: used to store the result data and false alarm information data after the recognition and processing module processes the original data; Server: used to remotely store the original detection data, recognition and processing result data, and false alarm information data, perform secondary verification of the decision result, and is electrically connected to the recognition and processing module; Reminder and warning module: used to give voice and vibration prompts when detecting and determining fatigue; Bypass module: used to manually turn off the reminder and warning through the bypass module when the system has false alarms.
[0026] In the second aspect, the present application also provides a computing device, which has the function of implementing the method described in the first aspect above. The beneficial effects can be seen in the description of the first aspect and will not be repeated here. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In a possible design, the structure of the device includes an acquisition module and a training module. Optionally, a construction module may also be included. These modules can implement the functions of the training nodes in the method example of the first aspect above. For specific details, refer to the detailed description in the method example and will not be elaborated here.
[0027] In a third aspect, the present application further provides a computing device, which is used to implement the functions of the method described in the first aspect above. The beneficial effects can be referred to the description in the first aspect and will not be elaborated here. The structure of the computing device includes a processor and a memory. The memory is used to store instructions and / or data. The memory is coupled to the processor. When the processor executes the program instructions stored in the memory, it can implement the functions of the training node in the examples of the first aspect above. The structure of the computing device also includes a communication interface for communicating with other devices.
[0028] In a fourth aspect, the present application further provides a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the methods in the first aspect and all possible designs of the first aspect.
[0029] In a fifth aspect, the present application further provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the methods in the first aspect and all possible designs of the first aspect.
[0030] In a sixth aspect, the present application further provides a computing chip. The chip is connected to the memory. The chip is used to read and execute the software program stored in the memory and execute the methods in the first aspect and all possible implementation manners of the first aspect. Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present drawings or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present drawings. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.
[0032] Figure 1 It is a flowchart of the steps of a lightweight driver fatigue monitoring method; Figure 2 It is a detailed flowchart of a lightweight driver fatigue monitoring method; Figure 3 It is a schematic diagram of data acquisition marking of a lightweight driver fatigue monitoring method; Figure 4 It is a schematic diagram of the discrimination limit of key parts of a lightweight driver fatigue monitoring method; Figure 5 It is a schematic diagram of image change during the process of a lightweight driver fatigue monitoring method; Figure 6 It is a schematic diagram of technical means of a lightweight driver fatigue monitoring method; Figure 7 It is a schematic diagram of the composition structure of a lightweight driver fatigue monitoring system; The realization, functional features, and advantages of this figure will be further described in combination with the embodiments with reference to the figures. Specific Embodiments
[0033] In order to make the objectives, technical solutions, and advantages of this application clearer, the following describes and explains this application in combination with the figures and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application. Based on the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0034] Obviously, the figures in the following description are only some examples or embodiments of this application. For those of ordinary skill in the art, without creative efforts, this application can also be applied to other similar scenarios based on these figures. In addition, it can also be understood that although the efforts made in this development process may be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in this application, some design, manufacturing, or production changes based on the technical content disclosed in this application are only conventional technical means and should not be understood as the content disclosed in this application being insufficient.
[0035] If there is no special instruction, all embodiments and optional embodiments of this application can be combined with each other to form new technical solutions.
[0036] If there is no special instruction, all technical features and optional technical features of this application can be combined with each other to form new technical solutions.
[0037] If there is no special instruction, all steps of this application can be carried out in sequence or randomly, and preferably in sequence. For example, the method includes steps (a) and (b), indicating that the method can include steps (a) and (b) carried out in sequence, or can also include steps (b) and (a) carried out in sequence. For example, it is mentioned that the method may further include step (c), indicating that step (c) can be added to the method in any order. For example, the method can include steps (a), (b), and (c), or can also include steps (a), (c), and (b), or can also include steps (c), (a), and (b), etc.
[0038] If there is no special instruction, the "including" and "comprising" mentioned in this application mean open-ended or can also be closed-ended. For example, the "including" and "comprising" can mean that other components not listed can also be included or comprised, or can only include or comprise the listed components.
[0039] Unless otherwise specified, the term "or" is inclusive in this application. For example, the phrase "A or B" means "A, B, or both A and B". More specifically, any of the following conditions satisfies the condition "A or B": A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or exist).
[0040] To better understand the solutions of the embodiments of this application, some related terms and concepts that may be involved in the embodiments of this application will be introduced below.
[0041] (1) Artificial intelligence (AI), also known as intelligent machinery or machine intelligence, refers to machines made by humans that can exhibit intelligence. Generally, artificial intelligence refers to the technology of presenting human intelligence through ordinary computer programs.
[0042] (2) Machine learning (ML), machine learning is the core of artificial intelligence. Machine learning theory mainly designs and analyzes algorithms that allow computers to learn automatically. Machine learning algorithms are a class of algorithms that automatically analyze and obtain rules from data and use the rules to predict unknown data. Therefore, the core of machine learning is data, algorithms (models), and computing power (computer computing ability). The application fields of machine learning are very extensive, including data mining, data classification, computer vision, natural language processing (NLP), biometric recognition, search engines, medical diagnosis, detection of credit card fraud, securities market analysis, DNA sequence sequencing, speech and handwriting recognition, strategic games, and robot applications. Machine learning is to design an algorithm model to process data and output the results desired by users. Users can continuously optimize the algorithm model to form more accurate data processing capabilities.
[0043] (3) Deep learning (DL). Deep learning is a method of machine learning. Its concept originated from the research on artificial neural networks. A multi-layer perceptron with multiple hidden layers is a deep learning structure. Therefore, deep learning is often also called a deep neural network. Compared with general machine learning, deep learning can automatically extract features, that is, automatically combine simple features into more complex features and use these combinations for multi-layer weight learning to solve problems. The motivation for studying deep learning is to establish a neural network that simulates the human brain for analysis and learning, which imitates the mechanism of the human brain to interpret data, such as images, sounds, and texts. Deep learning first emerged in image recognition. However, in just a few years, deep learning has been extended to various fields of machine learning and has shown excellent performance. It has applications in major fields such as image recognition, speech recognition, audio processing, natural language recognition, robotic bioinformatics processing, search engines, human-computer gaming, online advertising targeted delivery, medical automatic diagnosis, and finance.
[0044] (4) Neural networks (NN). It is an algorithmic mathematical model that imitates the behavioral characteristics of animal neural networks and performs distributed parallel information processing. By adjusting the relationships between a large number of nodes inside the neural network, the purpose of processing information can be achieved. Neural networks have the ability of self-learning and self-adaptation. Neural networks are usually applied in the model training and data derivation processing of artificial intelligence.
[0045] Specifically, a neural network usually can include multiple layers connected end to end, such as a convolutional layer, fully connected layers (FC), an activation layer, or a pooling layer, etc.
[0046] (5) Deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with multiple hidden layers. Dividing the DNN according to the positions of different layers, the neural networks inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected between each other. That is to say, any neuron in the i-th layer must be connected to any neuron in the (i + 1)-th layer. Although the DNN seems very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: where, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called the coefficient), and α() is the activation function. Each layer just performs such a simple operation on the input vector to obtain the output vector.
[0047] (6) A convolutional neural network (CNN) is a deep neural network with a convolutional structure. The convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers, and this feature extractor can be regarded as a filter. The convolutional layer refers to the neuron layer in the convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of the convolutional neural network, a neuron can only be connected to some adjacent layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some neurons arranged in a rectangle. The neurons in the same feature plane share weights, and the shared weights here are the convolutional kernels. Sharing weights can be understood as a way of extracting image information that is independent of position. The convolutional kernels can be initialized in the form of matrices of random sizes, and during the training process of the convolutional neural network, the convolutional kernels can obtain reasonable weights through learning. Additionally, the direct benefit brought by sharing weights is to reduce the connections between the layers of the convolutional neural network while reducing the risk of overfitting.
[0048] In this embodiment, as Figure 1 shown, a lightweight driver fatigue monitoring method includes the following steps: Step 1: Collect the interior light intensity of the vehicle. Based on the set light intensity threshold range, when the light intensity is not within the range, adjust the light intensity through the adjustment device. Through the light intensity sensor, sense the interior and exterior environmental light intensity of the vehicle. Using the dual-threshold mechanism, when the interior light intensity of the vehicle is higher than Lux, drive to adjust the angle of the light reduction mirror to reduce the amount of light passing through the camera lens; when the interior light intensity of the vehicle is lower than Lux, turn on several infrared lights for supplementary lighting to adjust and supplement the lighting for the camera to collect videos. Generally and should have a relatively obvious difference.
[0049] Both the light reduction mirror and the infrared lights are arranged on the image acquisition module. The light reduction mirror adjusts the light reduction effect by rotating the angle. For every rotation angle , the camera correspondingly reduces the light amount by Lux, and the default is no light reduction at 0°; there are N infrared lights, and after each one is turned on, it is equivalent to increasing the light amount of the camera by Lux. Control the number of lights turned on according to the interior light intensity of the vehicle, and the default is all off without supplementary lighting.
[0050] Step 2: Perform real-time image acquisition and preprocessing on the driver's seat area inside the vehicle. Step 3: Build a lightweight training model, preprocess the training portrait data and label the key parts, and use the data to train the model. Step 4: Determine the candidate regions in the image data of the driver's seat area based on the mask concept, and identify the key parts through a lightweight training model; Step 5: Set a determination threshold for the degree of change of the key parts, and set the over-limit value of the determination index within a unit time. When the determination index reaches the over-limit value, an alarm is issued; Step 6: Perform bypass shielding and local calibration for false alarm situations; Further, in the said Step 1, the adjustment formula for the light intensity is as follows: ; If , that is, when the light intensity inside the vehicle is higher than , the image may be overexposed, then in the formula , , the supplementary light strategy is not executed, and only the dimming mirror is adjusted to reduce the light; If , that is, when the light intensity inside the vehicle is lower than , it may be impossible to see clearly, then in the formula , , the dimming strategy is not executed, and only the supplementary light lamp is adjusted to supplement the light; Assume that after the light intensity adjustment strategy, the light intensity inside the vehicle can reach the appropriate range. After adjustment, the situation of not being able to be outside the appropriate range is not considered temporarily; If , that is, when the light intensity inside the vehicle is appropriate, there is no need to adjust the light input amount of the camera. In the formula , , there is no need to adjust the dimming or supplementary light strategy.
[0051] In this embodiment, the detailed flowchart of a lightweight driver fatigue monitoring method is as shown in Figure 2 , which illustrates the specific process of the fatigue detection method provided by this application; Further, the preprocessing specifically includes image grayscale processing and binary sparsification.
[0052] Obtain the original data, first store it in the local storage module 1, and at the same time, transmit it to the recognition processing module for core function processing. After obtaining a fixed number of video frames, form a batch and perform unified parallel grayscale processing to obtain the corresponding binary image, converting the R, G, B three-channel image into a single-channel image, reducing the data volume by 2 / 3 dimensions.
[0053] The binary sparsification is specifically to further sparsify the image through the pixel threshold method, obtain the non-blank information in the image, and adjust the empirical threshold. Based on being able to distinguish the opening and closing states of the eyes and mouth in most images, filter out those lower than this value and retain those higher than this value, thereby obtaining the binary processed image. The binary processed image can further reduce the memory space occupied and the processing overhead, and at the same time can reduce noise interference.
[0054] Further, the image grayscale processing is specifically to compress and transform the R, G, and B channel images into a single-channel image. The value formula for each pixel point during the compression process is as follows: ; In this embodiment, a regulation factor is added as an adjustable parameter to the classical conversion formula. The purpose is to clearly extract the key parts of the eyes and lips. For parts such as the eyes and lips, which are mainly based on the red and green primary colors, the weight value of the regulation factor can be appropriately increased.
[0055] Further, mark the key parts, specifically including: Since the elements in the portrait data are relatively concentrated, select the eyes and mouth, which have significant differences in facial expressions compared to when the driver is normal and awake during driver fatigue and drowsiness, as the key parts, and mark and identify the key parts.
[0056] In this embodiment, the schematic diagram of the changes during the processing of the collected images is as shown in Figure 5 shown, and the specific technical means used during the image processing is as shown in Figure 6 shown.
[0057] The eyes and mouth have significant differences in facial expressions compared to when the driver is normal and awake during driver drowsiness, and the recognition result has a high reliability. Therefore, focus on identifying the above two parts. By marking the eyes and mouth parts of the driver in the figure, supervised learning training is implemented.
[0058] The selection of materials for the portrait data used for training is not limited to driver images. As long as the pictures contain human heads, they can be adopted as the training set, mainly focusing on the states of the eyes and mouth. Binarize the image, mainly to compress the storage space and reduce the processing overhead. When manually marking, the binarized image may not be as clear and distinct as the original image, which is not conducive to accurate marking. Therefore, during the training stage, directly mark on the original image to obtain the position information of the marked Label box, and use the binarized image as the corresponding image in the training set Train to complete the mapping of the training set. The acquisition and marking of the image data are as shown in Figure 3 shown.
[0059] Further, determine the candidate region of the image data based on the mask concept, specifically including: According to the installation orientation of the camera, use the image collected under the normal driving condition of the driver as the mask, and fully consider the body tilt margin to delimit the driver's position range, and clear the rest.
[0060] The main acquisition range of the image acquisition module is mainly the upper body of the driver, and also includes the part behind him. If there are passengers sitting in the back row, they will also be acquired, which will interfere with the driver recognition. Therefore, drawing on the commonly used "mask" concept in image processing, according to the installation orientation of the camera, the image acquired under the normal driving condition of the driver is used as the mask, and the body tilt margin is fully considered.
[0061] In this embodiment, the body tilt margin is 10% of the dimension of the driver extending in all directions, roughly delimiting the driver's position. The rest is cleared to reduce the recognition interference caused by the entry of the rear passengers. It should be noted that the image is not cropped to maintain the consistency of the size before and after, which is convenient for unified processing.
[0062] After being trained by the convolutional neural network, when the system is put into use and the driver's image information is acquired, it can identify the opening and closing states of the driver's eyes and mouth through reasoning to preliminarily determine the driver's state at each moment.
[0063] Further, a determination threshold is set for the degree of change of the key parts, and an over-limit value of the determination index within a unit time is set, specifically including: drawing on the "eye diagram" concept in communication, a determination threshold is set for eye closure and mouth opening. Since both the eyes and the mouth are divided into upper and lower edges and are basically symmetric about the midline, the workload is reduced by only studying the upper edge. The areas where the upper edge of the eye diagram of the eyes is opened the largest (the largest distance from the midline of the eye diagram) and the upper edge is closed the smallest (the smallest distance from the midline of the eye diagram) are respectively detected. If the number of times from the area where the upper edge of the eye diagram is opened the largest to the area where it is closed the smallest within a unit time exceeds c times, it is considered that the blinking frequency is too high and it belongs to fatigue driving. If the upper edge of the eye diagram of the eyes has been in the area where it is closed the smallest for a period of time Δt, it is also considered fatigue driving. Similarly, the judgment of the mouth eye diagram is the same as above, except that if the upper edge of the mouth eye diagram has been in the area where it is opened the largest, it is considered fatigue driving, and the logic is opposite to that of the eye eye diagram; the blinking number is preferentially used as the determination index, and the number of times the mouth is opened is taken as the progressive supplementary determination index, and the over-limit values of these two indexes are set.
[0064] Further, when the determination index reaches the over-limit value, an alarm is issued, specifically including: adopting a progressive determination method. Within a unit time, first judge the blinking number. If it exceeds the limit, it is considered that the driver may have a fatigue problem and an alarm is issued; if it does not exceed the limit, then turn to the number of times the mouth is opened. If this index exceeds the limit, the system determines that the driver may have a fatigue problem and an alarm is issued.
[0065] After training, the system can identify the states of eyes opening, closing, mouth opening, and closing. As Figure 4 shown, this application draws on the "eye diagram" concept in communication and sets determination thresholds for eye closure and mouth opening, that is, when the eyes are closed to reach Figure 4The horizontal line in the middle is the judgment threshold. When the mouth is opened wide to reach Figure 4 The maximum value of the vertical line in the middle is the judgment threshold. Within a unit time, a progressive judgment method is adopted. The number of blinks is preferentially used as the judgment index, and the number of times the mouth is opened wide is taken as the progressive supplementary judgment index. That is, first judge the number of blinks. If it exceeds the limit, it is considered that the driver may have a fatigue problem and an alarm is issued; if it does not exceed the limit, then turn to the number of times the mouth is opened wide. If this index exceeds the limit, the system determines that the driver may have a fatigue problem and an alarm is issued. Regarding the determination of the over-limit values of the two indexes, there is a lot of research in biology in this regard, and relatively mature recommended reference values can be adopted. In this embodiment, the number of blinks adopts 20 times per minute as the over-limit value, and the number of times the mouth is opened wide adopts 2 times per minute as the over-limit value.
[0066] In the case where the driver eats or drinks and covers the mouth, wears sunglasses to cover the eyes, etc., and the mouth and eyes cannot be detected, it is recorded as an abnormal situation. In extremely rare cases, the driver has physiological characteristics or habits such as too small eyes or habitual mouth opening, etc., and false alarms can be bypassed and shielded, and at the same time, it is fed back to the system. During the idle stage, the system uses the data stored during the driving process, combined with the driver's bypass feedback, to perform secondary training to calibrate the system judgment.
[0067] As Figure 7 shown, the second aspect of the present application provides a lightweight driver fatigue monitoring system, including: Light sensing module: used to monitor the light intensity inside the vehicle in real time and transmit the signal to the recognition and processing module; Light supplement and light reduction module: used to control the infrared supplementary light and the light reduction mirror, and adjust the light intensity according to the light intensity control calculation of the recognition and processing module; Image acquisition module: monitors the key features of the driver's face, obtains the original data, and sends it to the recognition and processing module and the local storage module 1; Local storage module 1: stores the original data transmitted by the image acquisition module; Recognition and processing module: used for light intensity control calculation and face key part recognition training, controls the reminder and warning module when detecting and judging fatigue, and uses the data of the local storage module 2 for local calibration training when feedbacking false alarms; Local storage module 2: used to store the result data and false alarm information data processed by the recognition and processing module for the original data; Server: used for remote storage of the original detection data, recognition and processing result data, false alarm information data, performs secondary verification of the decision result, and is electrically connected to the recognition and processing module; Reminder and warning module: used to give voice and vibration prompts when detecting and judging fatigue; Bypass module: used to manually turn off the reminder and warning through the bypass module when the system issues a false alarm.
[0068] It should be noted that the present application is not limited to the above-described embodiments. The above embodiments are merely examples, and embodiments having the same constitution in terms of technical idea and achieving the same effects within the scope of the technical solution of the present application are all included in the technical scope of the present application. In addition, within the scope not departing from the gist of the present application, various modifications that can be conceived by those skilled in the art to the embodiments, and other ways constructed by combining some constituent elements in the embodiments are also included in the scope of the present application.
Claims
1. A lightweight driver fatigue monitoring method, characterized in that, It includes the following steps: Step 1: Collect the light intensity inside the vehicle. When the light intensity is not within the set light intensity threshold range, adjust the light intensity through the adjustment device; Step 2: Collect and preprocess the real-time image of the driving position area inside the vehicle; Step 3: Build a lightweight training model, preprocess the training portrait data and label the key parts, and use the data to train the model; Step 4: Determine the candidate areas in the image data of the driving position area based on the mask concept, and identify the key parts through the lightweight training model; Step 5: Set a determination threshold for the change degree of the key parts, and set the overlimit value of the determination index within a unit time. When the determination index reaches the overlimit value, an alarm is issued; Step 6: Perform bypass shielding and local calibration for false alarm situations.
2. The lightweight driver fatigue monitoring method according to claim 1, wherein The preprocessing specifically includes image grayscale processing and binary sparse.
3. The lightweight driver fatigue monitoring method according to claim 2, characterized in that, The image grayscale processing is specifically to compress and transform the R, G, and B channel images into a single-channel image.
4. The lightweight driver fatigue monitoring method according to claim 1, wherein, The labeling of the key parts specifically includes: since the elements in the portrait data are relatively concentrated, select the eyes and mouth, which are significantly different from the facial expressions in the normal waking state when the driver is fatigued or drowsy, as the key parts, and perform marker recognition on the key parts.
5. A lightweight driver fatigue monitoring method according to claim 1, characterized in that, The determination of the candidate areas of the image data based on the mask concept specifically includes: taking the image collected under the normal driving condition of the driver as the mask according to the installation orientation of the camera, and fully considering the body tilt margin, delimit the driver's position range, and clear the rest.
6. The lightweight driver fatigue monitoring method according to claim 5, characterized in that, The body tilt margin is specifically to expand the driver's image by 5% - 15% around.
7. A lightweight driver fatigue monitoring method according to claim 1, characterized in that, The setting of the determination threshold for the change degree of the key parts is specifically: referring to the "eye diagram" concept in communication, set the determination threshold for eye closure and mouth opening.
8. A lightweight driver fatigue monitoring method according to claim 1, characterized in that, The setting of the overlimit value of the determination index within a unit time is specifically: give priority to taking the number of blinks as the determination index, and take the number of times the mouth opens as the progressive supplementary determination index, and set the overlimit values of these two indexes.
9. A lightweight driver fatigue monitoring method according to claim 1, characterized in that, When the determination index reaches the overlimit value and an alarm is issued, it specifically includes: adopting a progressive determination method. Within a unit time, first judge the number of blinks. If it exceeds the limit, it is considered that the driver may have a fatigue problem and an alarm is issued; if it does not exceed the limit, then turn to the number of times the mouth opens. If this index exceeds the limit, the system determines that the driver may have a fatigue problem and an alarm is issued.
10. A lightweight driver fatigue monitoring system, characterized in that, The system is used to implement the function of a lightweight driver fatigue monitoring method as described in any one of claims 1 - 9.
Citation Information
Patent Citations
Fatigue driving detection method based on eye and mouth states
CN104809445A
Fatigue driving state detection method and device, computer equipment and storage medium
CN108830240A
Fatigue driving recognition method, device and system, vehicle-mounted terminal and server
CN110855934A
Fatigue driving detection method based on image enhancement technology
CN113435415A
Image recognition method and device
CN115424318A