Driving safety multi-dimensional data mining method for physiological and psychological monitoring of driver

By using physiological monitoring modules and gimbal image processing technology, the problem of inaccurate driver fatigue and psychological state in existing technologies has been solved, enabling multi-dimensional monitoring of the driver's physiological and psychological state and improving driving safety.

CN121817819APending Publication Date: 2026-04-10ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-07-28
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies that rely solely on image monitoring cannot accurately determine a driver's fatigue and psychological state, resulting in poor effectiveness in preventing dangerous driving.

Method used

The physiological monitoring module collects the driver's physiological information through a close-range infrared camera, microphone, smart bracelet, and smart seat. Combined with gimbal for image processing and data analysis, it infers the driver's physiological and psychological state and reminds the driver to drive safely via 5G network.

Benefits of technology

It enables multi-dimensional monitoring of the driver's physiological and psychological state, improves the prevention of dangerous driving, and ensures driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121817819A_ABST
    Figure CN121817819A_ABST
Patent Text Reader

Abstract

The invention discloses a driving safety multi-dimensional data mining method for physiological and psychological monitoring of a driver. The method specifically comprises the following steps: (1) monitoring physiological information of the driver through a physiological monitoring module and transmitting the physiological information to a holder; (2) the holder processes the physiological information to determine the physiological state of the driver; (3) the holder deduces the psychological state of the driver according to the physiological state of the driver; (4) the driving state of the driver is obtained in combination with the physiological state and the psychological state, and if the driving state is abnormal, a worker reminds and guides the driver to safely drive the bus to a safe area to have a rest through a holder; and (5) performing short-term prediction and long-term prediction on the physiological state and the psychological state of the driver, and judging whether the driver can continue driving or not. The problems that the physical state of the driver cannot be accurately judged through pure image monitoring in the prior art, and the dangerous driving prevention effect is poor can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of driving safety technology, and more specifically, to a method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring. Background Technology

[0002] Currently, taking public transportation has become an increasingly popular mode of travel. While public transportation brings convenience, it has also led to an increase in traffic accidents. The property damage and social burden caused by numerous traffic accidents are difficult to quantify in monetary terms. According to research by relevant experts abroad, it is estimated that approximately 1-3% of the annual GDP is consumed by medical expenses and additional losses caused by traffic accidents. In particular, traffic accidents caused by unsafe driving behavior of public buses pose a greater threat to people's lives and a greater social burden. The public transportation environment is a major issue that cannot be ignored by the entire society.

[0003] People, vehicles, and roads constitute the road traffic environment, with human factors consistently being the decisive factor in road traffic accidents. European statistics on traffic accidents show that 80%–90% of traffic accidents are caused by human factors related to the driver, approximately 10% are caused by vehicle factors, and only about 5% are caused by environmental factors. Drivers are the information processors, decision-makers, and executors of vehicles on the road.

[0004] According to surveys, two major causes of traffic accidents are psychological factors (such as road rage) and physiological factors (drowsy driving). When drivers are psychologically stressed or fatigued, their ability to perceive and judge their surroundings and their control over the vehicle are significantly reduced, making them highly susceptible to traffic accidents. Currently, preventative measures against driver fatigue typically involve image monitoring of the driver's driving process and determining whether the driver is drowsy based on the images. However, simple drowsiness monitoring only reflects the result of fatigue driving and cannot accurately predict the driver's fatigue and psychological state in advance, thus its effectiveness in preventing dangerous driving is limited. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-dimensional data mining method for driver safety based on driver physiological and psychological monitoring. This invention can solve the problem that the existing technology of simple image monitoring cannot accurately determine the driver's physical state and has a poor effect on preventing dangerous driving.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] A multi-dimensional data mining method for driver safety based on driver physiological and psychological monitoring includes the following steps:

[0008] (i) The physiological information of the driver is monitored by the physiological monitoring module and transmitted to the gimbal. The physiological information includes: the driver's facial video image, voice information, heart rate information, blood pressure information, sitting posture information and body temperature information.

[0009] (ii) The gimbal processes physiological information to determine the driver's physiological state;

[0010] (III) The gimbal infers the driver's psychological state based on the driver's physiological state;

[0011] (iv) Combine the driver's physiological and psychological state to obtain the driver's driving status. If there is any abnormality, the staff will remind and guide the driver to drive the bus safely to a safe area to rest through the gimbal.

[0012] (v) Make short-term and long-term predictions of the driver's physiological and psychological state to determine whether the driver can continue driving.

[0013] The physiological monitoring module includes two near-field infrared cameras, a microphone, a smart wristband, and a smart seat cushion. The first near-field infrared camera is installed at the front of the dashboard in the driver's cab, and the second near-field infrared camera is installed above the driver's head in the cab. The resolution of the near-field infrared cameras is 1280*960, and the frame rate is set to 50 frames per second. The horizontal distance between the first near-field infrared camera and the driver is 70-100cm, and the vertical distance between the second near-field infrared camera and the driver's eyes is 30-40cm. The microphones are all located in the driver's cab. The smart wristband is worn on the driver's wrist and collects the driver's heart rate and blood pressure information. The smart seat cushion is placed on the driver's seat and is equipped with a six-axis attitude sensor integrating a three-axis accelerometer and a three-axis gyroscope, a half-bridge pressure sensor, and a human infrared temperature sensor. A data storage device is installed on the bus, and the two near-field infrared cameras, microphone, smart wristband, and smart seat cushion are all connected to the data storage device. The data storage device is connected to the gimbal signal via a wireless network.

[0014] Step (1) is as follows: Two near-field infrared cameras capture the driver's driving status in real time and transmit the acquired facial video images of the driver to the data storage. The microphone acquires the driver's voice information in real time and transmits the voice information to the data storage. The smart bracelet collects the driver's heart rate and blood pressure information in real time through the equipped ECG signal measurement sensor and photoelectric sensor and transmits the heart rate and blood pressure information to the data storage. The smart seat accurately monitors the driver's sitting posture information in real time through the equipped six-axis posture sensor and half-bridge pressure sensor. The smart seat monitors the driver's body temperature information in real time through the equipped human infrared body temperature sensor. The smart seat transmits the monitored sitting posture information and body temperature information to the data storage. The data storage then transmits the collected facial video images, voice information, heart rate information, blood pressure information, sitting posture information and body temperature information of the driver to the gimbal via the 5G network.

[0015] Step (II) specifically involves: The gimbal processes the facial video images to perform facial expression recognition and eye fatigue recognition for the driver. To avoid altering the original image pixel values ​​and ensure image clarity before and after rotation, the gimbal rotates each image by an angle θ around the origin. After the rotation transformation, the pixel coordinates are:

[0016]

[0017] Where (x, y) are the pixel coordinates of the original image; (x1, y1) are the coordinates of the image after the corresponding pixel (x, y) has been rotated. According to equation (1), each image is rotated by 90°, 180° and 270° respectively, that is, θ takes 90°, 180° and 270° respectively. The expanded dataset is 3 times the size of the original dataset.

[0018] Facial expression recognition is performed in three steps:

[0019] (I) Image preprocessing: Due to factors such as bumps that may occur during bus driving, the acquired facial images may have defects such as inconsistent focal length and contrast. Therefore, proper image preprocessing is essential. Image preprocessing includes face alignment, data augmentation, and image normalization:

[0020] The face alignment module of the gimbal uses supervised descent to detect facial landmarks, such as eyebrows, nose, eyes, mouth and facial contours, for face recognition.

[0021] The gimbal's data augmentation module generates new training samples, improving the generalization ability and robustness of the convolutional neural network model in the facial feature extraction module, avoiding network overfitting, providing a large amount of training data for subsequent expression classification, and improving the accuracy of the convolutional neural network model in recognition.

[0022] Since the acquired images are RGB three-channel color images, with each channel ranging from 0 to 255, the computational load is too large. Therefore, grayscale normalization is necessary. Furthermore, the size of face images may vary depending on different road conditions. Images that are too large may contain redundant information, while images that are too small may overlook key features, reducing recognition accuracy. Additionally, machine learning algorithms generally require input images of a fixed size. Therefore, image size normalization is required. The gimbal's image normalization module performs grayscale and size normalization to ensure images have the same size and grayscale value range. Grayscale normalization uses the average method, maximum method, and weighted average method. The grayscale formula is as follows:

[0023] Gray=0.3R(x,y)+0.59G(x,y)+0.11B(x,y) (2)

[0024] In equation (2), Gray represents grayscale, R(x,y) refers to the red channel value, G(x,y) refers to the green channel value, and B(x,y) refers to the blue channel value.

[0025] Size normalization employs nearest neighbor interpolation and bilinear interpolation algorithms;

[0026] (II) Facial Feature Extraction: The gimbal's facial feature extraction module uses a convolutional neural network (CNN) model to automatically extract features from the input facial video images, achieving end-to-end learning. This CNN model consists of convolutional layers, activation layers, pooling layers, and fully connected layers. It is an improvement on the classic LeNet-5 model. The structure includes three convolutional layers (C1, C2, C3), three pooling layers (P1, P2, P3), where P1 is a max pooling layer, and P2 and P3 are average pooling layers, and one fully connected layer (F1). The pooling operation combines max pooling and average pooling. The parameters of each layer in this CNN model are shown in Table 1.

[0027] Table 1. Parameters of each layer in the convolutional neural network model

[0028] layer type Feature map convolution kernel Step length 0 enter 48×48 ———— ———— C1 convolution 48×48×32 5×5 1 P1 Max pooling 24×24×32 ———— 2 C2 convolution 24×24×64 3×3 1 P2 Average pooling 12×12×64 ———— 2 C3 convolution 12×12×64 3×3 1 P3 Average pooling 6×6×64 ———— 2 F1 Fully connected layer 1×1024 ———— ———— ;

[0029] After multiple experiments and tests, the outputs of each network layer after multi-scale pooling were selected to be feature matrices of different scales: 1×1×r, 2×2×r, and 3×3×r, where r is the number of feature maps and the value of r ranges from 1 to 100. The three feature matrices are arranged in columns to form a (13×r)×1 column vector. Finally, after feature fusion, a P-dimensional facial expression feature column vector with multiple scales and attributes is formed as the input vector x of the ELM classifier for facial expression classification.

[0030] (III) Facial Expression Classification: The facial expression classification module of the gimbal uses an ELM classifier for expression recognition. The ELM classifier consists of an input layer, a hidden layer, and an output layer. During network training, the ELM classifier only needs to calculate the weight matrix between the hidden layer and the output layer. The weight matrix and bias values ​​between its input layer and the hidden layer are randomly generated and do not require calculation or iterative updates. The number of neurons in the input layer, hidden layer, and output layer are d, l, and m, respectively.

[0031] The output of the i-th hidden node is:

[0032] g(x, w) i ,b i )=g(xw i +b i (3)

[0033] In equation (3), x is a column vector of facial expression features in P-dimensional space, g is the activation function, and w i Let b be the input weight vector between the i-th hidden node and all input nodes. i Let be the bias of the i-th hidden node, i = 1, 2, ..., l; g is the ReLU function, i.e.:

[0034] g(x, w) i b i ) = max(0, xw i +b i (4)

[0035] The connection between the input layer and the hidden layer is a mapping process. Because the input vector x is a feature vector in P-dimensional space, this connection is a process of mapping from P-dimensional space to l-dimensional space. The mapped feature vector of the input vector x is:

[0036] h(x)=[g(x,w1,b1),g(x,w2,b2),...,g(x,w i b i (5)

[0037] The output layer has m output nodes, where m is the number of different expressions, and each output node corresponds to one expression. The output weight between the i-th hidden node and the j-th output node is denoted as β. ij Where j = 1, 2, ..., m; therefore, the value of the j-th output node is:

[0038]

[0039] Therefore, the output vector of the input sample x in the hidden layer can be represented as:

[0040] f(x)=f1(x), f2(x),...,f m (x)=h(x)β (7)

[0041] In equation (7),

[0042]

[0043] During the testing phase, the input test sample test-x corresponds to the following expression category:

[0044] label(text-x) = arg j=1,2,...,m maxf j (text-x) (9)

[0045] The ELM classifier categorizes drivers' facial expressions into pleasure, focus, boredom, confusion, and frustration based on facial features extracted by the convolutional neural network model, as shown in Table 2.

[0046] Table 2 Classification of Basic Facial Expressions

[0047]

[0048]

[0049] Because when a person is fatigued, they will blink more frequently, keep their eyes closed for longer, yawn more, and in severe cases, even doze off. Under normal circumstances, a person blinks 10 to 25 times per minute, and the duration of eye closure for one blink is about 0.2 seconds. Based on this phenomenon, the three eye indicators that best represent the fatigue state are selected as the characteristic parameters of eye fatigue identification. The three eye indicators are: continuous eye closure time, eye closure frame rate, and blink frequency.

[0050] The face landmark localization algorithm based on cascaded regression trees is used for human eye localization. The extracted left and right eye regions are represented by left and right rectangular bounding boxes. The localization calculation rules are as follows:

[0051] W = 1.6 × W e H = 3 × H e (10)

[0052] In Equation (10): We is the horizontal distance between human eye feature points 36 and 39, He is the average of the vertical distances between human eye feature points 37 and 41 and between 38 and 40, and W and H are the width and height of the located eye region;

[0053] To accurately and quickly identify the open / closed state of the eyes, we calculated the aspect ratio between the height and width of the eyes. The aspect ratio when the eyes are open varies very little between individuals and remains completely unchanged regardless of uniform image scaling or facial rotation. The formula for calculating the eye aspect ratio is:

[0054]

[0055] In formula (11): P1 is the center point of the left side of the left eye, P2 is the center point of the upper left of the left eye, P3 is the center point of the upper right of the left eye, P4 is the center point of the right side of the left eye, P5 is the center point of the lower right of the left eye, and P6 is the center point of the lower left of the left eye.

[0056] To accurately identify eye status, the average EAR of both eyes is taken as the feature for eye opening and closing recognition:

[0057] EAR = Mean(EAR) left EAR right (12)

[0058] In formula (12): EAR left The aspect ratio of the left eye, EAR right The aspect ratio of the right eye;

[0059] When the human eye is closed, although it may be affected by dark areas such as eyelashes and eyelids, the largest dark area will not appear in the pupil area. Therefore, compared with the open eye, the number of black pixels in the binary image is reduced when the eye is closed. Therefore, an adaptive threshold method is used to accumulate the difference and define two states: "state 0" and "state 1". When the difference of black pixels in the binary image of the human eye area is less than 0, it changes from "state 0" to "state 1". In "state 1", if the difference is less than the threshold T(t), the difference is accumulated and the state remains unchanged; if the difference is greater than or equal to the threshold T(t), the difference is not accumulated and the state changes to "state 0".

[0060] Based on the above conclusions, during a blink, the EAR value first decreases until it approaches 0, and then gradually increases to the value of a normal open eye state. Let E be the threshold of EAR. When the EAR threshold is less than E, the eyes begin to close. When the EAR threshold is close to the normal open eye state, that is, greater than E, the eyes open.

[0061] The gimbal processes the voice information to complete the driver's speech recognition. The gimbal's speech extraction module converts the raw speech signal into a digital signal through sampling quantization, determines the start and end of the speech through endpoint detection, adjusts the mid, high, and low frequency amplitudes appropriately through pre-emphasis processing, obtains frame-level speech sequences through frame segmentation and windowing, and finally extracts sound features, representing them in the form of feature vectors.

[0062] Quantization sampling of speech signals involves measuring the information of the original speech signal in a digital system and converting it into a digital signal that a computer can recognize. This means discretizing the analog signal in terms of amplitude and time. Sampling involves sequentially capturing audio signals at fixed time intervals from a continuous analog signal and replacing the continuous analog signal with the amplitude value of the audio signal at the moment of sampling. The sampling frequency is set to 22.05kHz. Through sampling, the speech analog signal is transformed into a discrete-time signal. The amplitude value of the discrete-time signal is then quantized in stages, that is, the signal amplitude is divided into multiple intervals, and the sampled amplitude value within the same interval is represented by the same quantization value.

[0063] Endpoint detection involves setting thresholds for energy and zero-crossing rate to detect endpoints in speech. With a frame length of 32ms, endpoint detection can pinpoint the start and end of valid speech, removing quiet and noise components. This not only shortens the speech time series but also improves speech emotion recognition after removing influencing factors, and helps distinguish between valid and invalid speech regions. Pre-emphasis is used because the speech signal loses high-frequency energy through radiation from the lips and nose of the vocal organs, leading to energy attenuation. To protect speech information, reduce loss, and prevent waveform distortion, a first-order digital filter is used to pre-emphasize the speech signal, compensating for the loss of certain components in the high-frequency part. The transfer function of the first-order digital filter is:

[0064] H(z) = 1 - μz -1 (13)

[0065] In equation (13), μ is the strengthening coefficient, with a value range of [0.9-1], and z -1 This indicates a delay of one unit of time;

[0066] The sound feature extraction employs the Mel-frequency cepstral coefficient extraction method. Mel frequencies, based on human auditory characteristics, describe subjective pitch and exhibit a non-linear relationship with the objective pitch frequency f. A speech sample undergoes a series of preprocessing steps to obtain a frame of speech, which is then subjected to Fourier transform to obtain the spectral energy distribution. The Mel filter bank consists of a set of Mel-scale bandpass triangular filters. These triangular filters are not of equal bandwidth; rather, their bandwidth increases from low frequencies, resulting in decreasing frequency resolution. The input sound frequency domain signal is multiplied and added, and the output signal of each individual filter serves as a fundamental feature of the sound signal. This feature resides in the Mel spectral domain, and the perception of sound is linear. M triangular filters are used to smooth the spectrum and eliminate harmonics. Let the center frequency f(m) be given, and the transfer function H of the Mel filter bank be... r (k) is represented as:

[0067]

[0068] in,

[0069] The power spectrum is obtained by taking the modulus of the spectrum of the frame speech signal and squaring it. The logarithmic energy of the output after passing through the filter bank is calculated. Finally, the MFCC coefficients are calculated by discrete cosine transform and then filtered.

[0070] The gimbal's speech recognition module uses a pattern matching method based on vector quantization (VQ) to recognize the extracted sound features. The pattern matching process consists of two steps:

[0071] The first step is to train the VQ model using voiceprint feature parameters to generate a voiceprint template library;

[0072] The second step is to input the voiceprint features of multiple speakers and perform matching operations with the pre-trained voiceprint model to determine whether it is a speaker in the template.

[0073] For the VQ model, during the training phase, the speaker's MFCC feature vectors are clustered based on the principle of minimum mean square error to generate a codebook composed of different codewords. During the recognition phase, the distortion distance between the MFCC feature vector of the speech to be recognized and the codebook in the template is calculated. The mean square error and absolute error are used as the measures, and the sentence with the minimum distortion distance is selected as the successfully matched MFCC feature vector.

[0074] The gimbal's speech-text module automatically converts the recognized speech into text, and then performs text segmentation, which is a very important step in natural language processing. In order to facilitate the subsequent selection of data features and make the analysis simpler, text segmentation is used to break down long sentences or paragraphs into data structures based on words.

[0075] Word segmentation can be divided into three types: ① dictionary matching segmentation, which segments the text according to specific rules and then matches it with words in the dictionary to determine whether segmentation is necessary; ② statistical segmentation methods, commonly using algorithms such as Hidden Markov Models and Conditional Random Fields; ③ deep learning-based segmentation.

[0076] Word segmentation tools include Hanlp, Stanford Tokenizer, THULAC, and Jieba. The word segmentation steps are as follows:

[0077] (1) First, efficient word graph scanning based on prefix dictionary is adopted to generate a directed acyclic graph (DAG) composed of all possible word combinations of Chinese characters in the sentence;

[0078] (2) Use dynamic programming to find the maximum probability path and find the maximum segmentation combination based on word frequency;

[0079] (3) For new words, an HMM model based on the word-forming ability of Chinese characters is adopted, and the Viterbi algorithm is used;

[0080] Before performing word frequency statistics on the word list, additional processing is often required. Because during the word segmentation process, a large number of function words such as "de" and "le", numeral-classifiers, and punctuation marks are separated into individual words. These words are called stop words. Stop words have a high frequency of occurrence but no application value, and these stop words should be removed when constructing a word library by counting word frequencies;

[0081] The cloud platform processes the collected heart rate information and blood pressure information of the driver, judges the physical health status of the driver, and at the same time analyzes and processes the sitting posture information and effectively monitors the driving situation of the driver by combining real-time video images. The cloud platform processes the body temperature information to determine whether the driver's body temperature is normal.

[0082] Step (iii) is specifically as follows: The cloud platform inputs the facial expression information, voice and text information, heart rate information, blood pressure information, and body temperature information obtained in step (ii) into the SVM model, and the SVM model can infer the mental state of the driver.

[0083] Step (iv) is specifically as follows: The staff obtains the physiological state and mental state of the driver through the cloud platform, combines the physiological state and mental state to obtain the driving state of the driver. If it is found that the current driving state of the driver is not conducive to driving, the staff dials the phone equipped in the cab through the cloud platform and the 5G network to communicate with the driver, and reminds and guides the driver to drive the bus safely to a safe area for rest.

[0084] Step (v) is specifically as follows: Body temperature monitoring should be carried out before the driver drives the bus every day, compared with historical body temperature data, to judge whether the driver's body temperature data has changed due to not sleeping well the previous day or weather reasons, and to make short-term and long-term predictions on the driver's driving state, and judge whether the driver can continue to drive:

[0085] The short-term prediction is made based on the physiological state and mental state monitored by the driver on the same day;

[0086] The long-term prediction is made based on the short-term prediction situations of the driver in the past seven times and the actual driving situations in the past seven times;

[0087] Compare the daily short-term prediction situation with the actual driving situation on the same day for scoring. When the daily score is greater than 80%, it means that the short-term prediction accuracy on the same day is high. When the average value of the scores in the past seven times is greater than 80%, this long-term prediction is available.

[0088] Because each driver's physical condition and individual circumstances are different, drivers may exhibit behaviors such as frequent blinking, speaking too fast, blood pressure exceeding medical standards, or varying heart rates. Therefore, pre-adjustment processing is required for each driver's specific driving situation: When a driver's blinking frequency exceeds the average blinking frequency, the blinking frequency should be multiplied by an adjustment coefficient k1, reducing the weight of the average blinking frequency; when a driver speaks too fast, the speaking speed should be multiplied by an adjustment coefficient k2, reducing the weight of the speaking speed; when a driver's heart rate is too fast under normal conditions, the heart rate should be multiplied by an adjustment coefficient k3, reducing the weight of the driver's heart rate; and when a driver's blood pressure is too high, the speaking speed should be multiplied by an adjustment coefficient k4, reducing the weight of the driver's blood pressure. These weighting coefficients are adjusted by staff on the gimbal.

[0089] This invention possesses significant substantive features and substantial advancements compared to existing technologies. Specifically, it employs a near-field infrared camera to acquire facial video images. A gimbal processes these images to recognize the driver's facial expressions and eye fatigue. Simultaneously, it collects the driver's voice information via a microphone, which the gimbal then performs voice recognition. A smart bracelet collects the driver's heart rate and blood pressure information, while a smart seat monitors the driver's posture and body temperature. Based on facial expression information, eye fatigue information, voice and text information, heart rate information, blood pressure information, and body temperature information, the driver's physiological state is determined. The gimbal then transmits the facial expression information... Information such as eye fatigue, voice and text, heart rate, blood pressure, and body temperature is input into the SVM model, which can then infer the driver's psychological state. Staff members obtain the driver's physiological and psychological states through the gimbal, and combine these to determine the driver's driving status. If the driver's current driving status is found to be unfavorable, staff members can use the gimbal and 5G network to contact the driver via the telephone located in the driver's cab, reminding and guiding the driver to safely drive the bus to a safe area to rest. This invention can also make short-term and long-term predictions of the driver's driving status to determine whether the driver can continue driving. Attached Figure Description

[0090] Figure 1 This is a block diagram showing the connections between the various components of the present invention.

[0091] Figure 2 This is a schematic diagram illustrating the acquisition of the driver's psychological state according to the present invention.

[0092] Figure 3 This is an extraction image of the key points of the driver's eye contour in this invention. Detailed Implementation

[0093] The embodiments of the present invention are further described below with reference to the accompanying drawings.

[0094] like Figure 1-3 As shown, a multi-dimensional data mining method for driver safety based on driver physiological and psychological monitoring includes the following steps:

[0095] (i) The physiological information of the driver is monitored by the physiological monitoring module and transmitted to the gimbal 1. The physiological information includes: the driver's facial video image, voice information, heart rate information, blood pressure information, sitting posture information and body temperature information.

[0096] (ii) The gimbal 1 processes physiological information to determine the driver's physiological state;

[0097] (III) The gimbal 1 infers the driver's psychological state based on the driver's physiological state;

[0098] (iv) The driver’s driving status is obtained by combining the physiological and psychological states. If there is any abnormality, the staff will remind and guide the driver to drive the bus safely to a safe area to rest through the gimbal 1.

[0099] (v) Make short-term and long-term predictions of the driver's physiological and psychological state to determine whether the driver can continue driving.

[0100] The physiological monitoring module includes two near-field infrared cameras 2, a microphone 3, a smart bracelet 4, and a smart seat cushion 5. The first near-field infrared camera 2 is installed at the front of the dashboard in the driver's cab, and the second near-field infrared camera 2 is installed above the driver's head in the driver's cab. The resolution of the near-field infrared cameras 2 is 1280*960, and the frame rate is set to 50 frames / second. The horizontal distance between the first near-field infrared camera 2 and the driver is 70-100cm, and the vertical distance between the second near-field infrared camera 2 and the driver's eyes is 30-40cm. The microphone 3 is located in the driver's cab. The smart bracelet 4 is worn on the driver's wrist and collects the driver's heart rate and blood pressure information. The smart seat cushion 5 is placed on the driver's seat. The smart seat cushion 5 is equipped with a six-axis attitude sensor integrating a three-axis accelerometer and a three-axis gyroscope, a half-bridge pressure sensor, and a human infrared temperature sensor. A data storage device 6 is installed on the bus. The two near-field infrared cameras 2, microphone 3, smart bracelet 4, and smart seat cushion 5 are all connected to the data storage device 6. The data storage device 6 is connected to the gimbal 1 via a wireless network.

[0101] Step (I) is as follows: Two near-field infrared cameras 2 capture the driver's driving status in real time and transmit the acquired facial video images of the driver to the data storage 6. Microphone 3 acquires the driver's voice information in real time and transmits the voice information to the data storage 6. Smart bracelet 4 collects the driver's heart rate and blood pressure information in real time through the equipped electrocardiogram signal measurement sensor and photoelectric sensor and transmits the heart rate and blood pressure information to the data storage 6. Smart seat 5 monitors the driver's sitting posture information in real time and accurately through the equipped six-axis posture sensor and half-bridge pressure sensor. Smart seat 5 monitors the driver's body temperature information in real time through the equipped human infrared body temperature sensor. Smart seat 5 transmits the monitored sitting posture information and body temperature information to the data storage 6. The data storage 6 then transmits the collected facial video images, voice information, heart rate information, blood pressure information, sitting posture information and body temperature information of the driver to the gimbal 1 through the 5G network.

[0102] Step (II) specifically involves: Gimbal 1 processing the facial video images to perform facial expression recognition and eye fatigue recognition for the driver. To avoid altering the original image pixel values ​​and ensure image clarity before and after rotation, Gimbal 1 rotates each image by an angle θ around the origin. After the rotation transformation, the pixel coordinates are:

[0103]

[0104] Where (x, y) are the pixel coordinates of the original image; (x1, y1) are the coordinates of the image after the corresponding pixel (x, y) has been rotated. According to equation (1), each image is rotated by 90°, 180° and 270° respectively, that is, θ takes 90°, 180° and 270° respectively. The expanded dataset is 3 times the size of the original dataset.

[0105] Facial expression recognition is performed in three steps:

[0106] (I) Image preprocessing: Due to factors such as bumps that may occur during bus driving, the acquired facial images may have defects such as inconsistent focal length and contrast. Therefore, proper image preprocessing is essential. Image preprocessing includes face alignment, data augmentation, and image normalization:

[0107] The face alignment module of the gimbal 1 uses supervised descent to detect facial landmarks, such as eyebrows, nose, eyes, mouth and facial contours, for face recognition.

[0108] The data augmentation module of GMT1 generates new training samples, improves the generalization ability and robustness of the convolutional neural network model of the facial feature extraction module, avoids network overfitting, provides a large amount of training data for subsequent expression classification, and improves the accuracy of the convolutional neural network model recognition.

[0109] Since the acquired images are RGB three-channel color images, with each channel ranging from 0 to 255, the computational load is too large. Therefore, grayscale normalization is necessary. Furthermore, the size of face images may vary depending on road conditions. Images that are too large may contain redundant information, while images that are too small may miss key features, reducing recognition accuracy. Additionally, machine learning algorithms generally require input images of a fixed size. Therefore, image size normalization is required. The image normalization module of the gimbal 1 performs grayscale and size normalization to ensure images have the same size and grayscale value range. Grayscale normalization uses the average method, maximum method, and weighted average method. The grayscale formula is as follows:

[0110] Gray=0.3R(x,y)+0.59G(x,y)+0.11B(x,y) (2)

[0111] In equation (2), Gray represents grayscale, R(x,y) refers to the red channel value, G(x,y) refers to the green channel value, and B(x,y) refers to the blue channel value.

[0112] Size normalization employs nearest neighbor interpolation and bilinear interpolation algorithms;

[0113] (II) Facial Feature Extraction: The facial feature extraction module of the gimbal 1 uses a convolutional neural network model to automatically extract features from the input facial video images, achieving "end-to-end" learning. The network layers of this convolutional neural network model are divided into convolutional layers, activation layers, pooling layers, and fully connected layers. This convolutional neural network model is an improvement based on the classic LeNet-5 model. The structure of this convolutional neural network model includes three convolutional layers: C1, C2, and C3; three pooling layers: P1, P2, and P3, where P1 is a max pooling layer, and P2 and P3 are average pooling layers; and one fully connected layer: F1. The pooling operation uses a combination of max pooling and average pooling. The network layer parameters of each layer of this convolutional neural network model are shown in Table 1.

[0114] Table 1. Parameters of each layer in the convolutional neural network model

[0115] layer type Feature map convolution kernel Step length 0 enter 48×48 ———— ———— C1 convolution 48×48×32 5×5 1 P1 Max pooling 24×24×32 ———— 2 C2 convolution 24×24×64 3×3 1 P2 Average pooling 12×12×64 ———— 2 C3 convolution 12×12×64 3×3 1 P3 Average pooling 6×6×64 ———— 2 F1 Fully connected layer 1×1024 ———— ———— ;

[0116] After multiple experiments and tests, it was finally determined that the output of each network layer after multi-scale pooling operation is three feature matrices of different scales, namely 1×1×r, 2×2×r, and 3×3×r, where r is the number of feature maps and the value of r ranges from 1 to 100. The three feature matrices are arranged in columns to form a (13×r)×1 column vector. Finally, after feature fusion, a P-dimensional facial expression feature column vector with multiple scales and attributes is formed as the input vector x of the ELM classifier for facial expression classification.

[0117] (III) Facial Expression Classification: The facial expression classification module of the gimbal 1 uses an ELM classifier for expression recognition. The ELM classifier consists of an input layer, a hidden layer, and an output layer. During network training, the ELM classifier only needs to calculate the weight matrix between the hidden layer and the output layer. The weight matrix and bias values ​​between its input layer and the hidden layer are randomly generated and do not require calculation or iterative updates. The number of neurons in the input layer, hidden layer, and output layer are d, l, and m, respectively.

[0118] The output of the i-th hidden node is:

[0119] g(x, w) i b i )=g(xw i +b i (3)

[0120] In equation (3), x is a column vector of facial expression features in P-dimensional space, g is the activation function, and w i Let b be the input weight vector between the i-th hidden node and all input nodes. i Let be the bias of the i-th hidden node, i = 1, 2, ..., l; g is the ReLU function, i.e.:

[0121] g(x, w) i b i ) = max(0, xw i +b i (4)

[0122] The connection between the input layer and the hidden layer is a mapping process. Because the input vector x is a feature vector in P-dimensional space, this connection is a process of mapping from P-dimensional space to l-dimensional space. The mapped feature vector of the input vector x is:

[0123] h(x)=[g(x,w1,b1),g(x,w2,b2),...,g(x,w i b i (5)

[0124] The output layer has m output nodes, where m is the number of different expressions, and each output node corresponds to one expression. The output weight between the i-th hidden node and the j-th output node is denoted as β. ij Where j = 1, 2, ..., m; therefore, the value of the j-th output node is:

[0125]

[0126] Therefore, the output vector of the input sample x in the hidden layer can be represented as:

[0127] f(x)=f1(x), f2(x),...,f m (x)=h(x)β (7)

[0128] In equation (7),

[0129]

[0130] During the testing phase, the input test sample test-x corresponds to the following expression category:

[0131] label(text-x) = arg j=1,2,...,m maxf j (text-x) (9)

[0132] The ELM classifier categorizes drivers' facial expressions into pleasure, focus, boredom, confusion, and frustration based on facial features extracted by the convolutional neural network model, as shown in Table 2.

[0133] Table 2 Classification of Basic Facial Expressions

[0134]

[0135] Because when a person is fatigued, they will blink more frequently, keep their eyes closed for longer, yawn more, and in severe cases, even doze off. Under normal circumstances, a person blinks 10 to 25 times per minute, and the duration of eye closure for one blink is about 0.2 seconds. Based on this phenomenon, the three eye indicators that best represent the fatigue state are selected as the characteristic parameters of eye fatigue identification. The three eye indicators are: continuous eye closure time, eye closure frame rate, and blink frequency.

[0136] The face landmark localization algorithm based on cascaded regression trees is used for human eye localization. The extracted left and right eye regions are represented by left and right rectangular bounding boxes. The localization calculation rules are as follows:

[0137] W = 1.6 × W e H = 3 × H e (10)

[0138] In Equation (10): We is the horizontal distance between human eye feature points 36 and 39, He is the average of the vertical distances between human eye feature points 37 and 41 and between 38 and 40, and W and H are the width and height of the located eye region;

[0139] To accurately and quickly identify the open / closed state of the eyes, we calculated the aspect ratio between the height and width of the eyes. The aspect ratio when the eyes are open varies very little between individuals and remains completely unchanged regardless of uniform image scaling or facial rotation. The formula for calculating the eye aspect ratio is:

[0140]

[0141] In formula (11): P1 is the center point of the left left eye, P2 is the center point of the upper left left eye, P3 is the center point of the upper right left eye, P4 is the center point of the right left eye, P5 is the center point of the lower right left eye, and P6 is the center point of the lower left left eye. See Figure 3 As shown;

[0142] To accurately identify eye status, the average EAR of both eyes is taken as the feature for eye opening and closing recognition:

[0143] EAR = Mean(EAR) left EAR right (12)

[0144] In formula (12): EAR left The aspect ratio of the left eye, EAR right The aspect ratio of the right eye;

[0145] When the human eye is closed, although it may be affected by dark areas such as eyelashes and eyelids, the largest dark area will not appear in the pupil area. Therefore, compared with the open eye, the number of black pixels in the binary image is reduced when the eye is closed. Therefore, an adaptive threshold method is used to accumulate the difference and define two states: "state 0" and "state 1". When the difference of black pixels in the binary image of the human eye area is less than 0, it changes from "state 0" to "state 1". In "state 1", if the difference is less than the threshold T(t), the difference is accumulated and the state remains unchanged; if the difference is greater than or equal to the threshold T(t), the difference is not accumulated and the state changes to "state 0".

[0146] Based on the above conclusions, during a blink, the EAR value first decreases until it approaches 0, and then gradually increases to the value of a normal open eye state. Let E be the threshold of EAR. When the EAR threshold is less than E, the eyes begin to close. When the EAR threshold is close to the normal open eye state, that is, greater than E, the eyes open.

[0147] The gimbal 1 processes the voice information to complete the driver's voice recognition. The voice extraction module of the gimbal 1 converts the original voice signal into a digital signal through sampling quantization, determines the beginning and end of the voice through endpoint detection, adjusts the mid, high and low frequency amplitudes appropriately through pre-emphasis processing, obtains the frame-level voice sequence through frame segmentation and windowing, and finally extracts the sound features, representing them in the form of feature vectors.

[0148] Quantization sampling of speech signals involves measuring the information of the original speech signal in a digital system and converting it into a digital signal that a computer can recognize. This means discretizing the analog signal in terms of amplitude and time. Sampling involves sequentially capturing audio signals at fixed time intervals from a continuous analog signal and replacing the continuous analog signal with the amplitude value of the audio signal at the moment of sampling. The sampling frequency is set to 22.05kHz. Through sampling, the speech analog signal is transformed into a discrete-time signal. The amplitude value of the discrete-time signal is then quantized in stages, that is, the signal amplitude is divided into multiple intervals, and the sampled amplitude value within the same interval is represented by the same quantization value.

[0149] Endpoint detection involves setting thresholds for energy and zero-crossing rate to detect endpoints in speech. With a frame length of 32ms, endpoint detection can pinpoint the start and end of valid speech, removing quiet and noise components. This not only shortens the speech time series but also improves speech emotion recognition after removing influencing factors, and helps distinguish between valid and invalid speech regions. Pre-emphasis is used because the speech signal loses high-frequency energy through radiation from the lips and nose of the vocal organs, leading to energy attenuation. To protect speech information, reduce loss, and prevent waveform distortion, a first-order digital filter is used to pre-emphasize the speech signal, compensating for the loss of certain components in the high-frequency part. The transfer function of the first-order digital filter is:

[0150] H(z) = 1 - μz -1 (13)

[0151] In equation (13), μ is the strengthening coefficient, with a value range of [0.9-1], and z -1 This indicates a delay of one unit of time;

[0152] The sound feature extraction employs the Mel-frequency cepstral coefficient extraction method. Mel frequencies, based on human auditory characteristics, describe subjective pitch and exhibit a non-linear relationship with the objective pitch frequency f. A speech sample undergoes a series of preprocessing steps to obtain a frame of speech, which is then subjected to Fourier transform to obtain the spectral energy distribution. The Mel filter bank consists of a set of Mel-scale bandpass triangular filters. These triangular filters are not of equal bandwidth; rather, their bandwidth increases from low frequencies, resulting in decreasing frequency resolution. The input sound frequency domain signal is multiplied and added, and the output signal of each individual filter serves as a fundamental feature of the sound signal. This feature resides in the Mel spectral domain, and the perception of sound is linear. M triangular filters are used to smooth the spectrum and eliminate harmonics. Let the center frequency f(m) be given, and the transfer function H of the Mel filter bank be... r (k) is represented as:

[0153]

[0154] in,

[0155] The power spectrum is obtained by taking the modulus of the spectrum of the frame speech signal and squaring it. The logarithmic energy of the output after passing through the filter bank is calculated. Finally, the MFCC coefficients are calculated by discrete cosine transform and then filtered.

[0156] The speech recognition module of the PTZ 1 uses a pattern matching method based on vector quantization (VQ) to recognize the extracted sound features. The pattern matching is divided into two steps:

[0157] The first step is to train the VQ model using voiceprint feature parameters to generate a voiceprint template library;

[0158] The second step is to input the voiceprint features of multiple speakers and perform matching operations with the pre-trained voiceprint model to determine whether it is a speaker in the template.

[0159] For the VQ model, during the training phase, the speaker's MFCC feature vectors are clustered based on the principle of minimum mean square error to generate a codebook composed of different codewords. During the recognition phase, the distortion distance between the MFCC feature vector of the speech to be recognized and the codebook in the template is calculated. The mean square error and absolute error are used as the measures, and the sentence with the minimum distortion distance is selected as the successfully matched MFCC feature vector.

[0160] The speech-text module of the gimbal 1 automatically converts the recognized speech into text, and then performs text segmentation, which is a very important step in natural language processing. In order to facilitate the subsequent selection of data features and make the analysis simpler, long sentences or paragraphs are broken down into data structures based on words.

[0161] Word segmentation can be divided into three methods: ① The word segmentation method of dictionary matching, which segments the text according to specific rules and then matches it with the words in the dictionary to determine whether to segment the words; ② The word segmentation method based on statistics, and the commonly used algorithms are Hidden Markov Model and Conditional Random Field; ③ The method based on deep learning;

[0162] Word segmentation tools include Hanlp, Stanford Word Segmentation, THULAC, and Jieba. This patent preferably uses the word segmentation tool "Jieba" for word segmentation. The word segmentation steps are as follows:

[0163] (1) First, use a prefix dictionary to achieve efficient word graph scanning, and generate a directed acyclic graph DAG composed of all possible word formation situations of Chinese characters in the sentence;

[0164] (2) Use dynamic programming to find the maximum probability path and find the maximum segmentation combination based on word frequency;

[0165] (3) For new words, use the HMM model based on the word formation ability of Chinese characters and use the Viterbi algorithm;

[0166] Before performing word frequency statistics on the word list, additional processing is often required because during the word segmentation process, a large number of function words such as "de" and "le", numeral-classifiers, and punctuation marks will be segmented into words separately. These words are called stop words. Stop words have a high frequency of occurrence but no application value, and these stop words should be removed when constructing a word library by counting word frequencies;

[0167] The pan-tilt 1 processes the collected heart rate information and blood pressure information of the driver to judge the driver's physical health condition. At the same time, it analyzes and processes the sitting posture information and effectively monitors the driver's driving situation in combination with the real-time video image. The pan-tilt 1 processes the body temperature information to determine whether the driver's body temperature is normal.

[0168] Step (three) is specifically: The pan-tilt 1 inputs the facial expression information, voice and text information, heart rate information, blood pressure information, and body temperature information obtained in step (two) into the SVM model, and the SVM model can infer the driver's mental state.

[0169] Step (four) is specifically: The staff obtains the driver's physiological state and mental state through the pan-tilt 1, combines the physiological state and mental state to obtain the driver's driving state. If it is found that the driver's current driving state is not conducive to driving, the staff dials the phone equipped in the cab through the pan-tilt 1 and the 5G network to communicate with the driver, and reminds and guides the driver to drive the bus safely to a safe area for rest.

[0170] Step (5) specifically involves: Drivers' body temperature should be monitored before each day's bus driving, compared with historical temperature data, to determine if the change was due to insufficient sleep the previous day or weather conditions. Short-term and long-term predictions of the driver's driving condition should also be made to determine whether the driver can continue driving.

[0171] Short-term forecasts are based on the driver's physiological and psychological state monitored on the same day.

[0172] Long-term forecasts are based on the driver's seven most recent short-term forecasts and seven most recent actual driving situations.

[0173] The daily short-term forecast is scored by comparing it with the actual driving conditions of the day. When the daily score is greater than 80%, it indicates that the short-term forecast is highly accurate. When the average score of the last seven times is greater than 80%, the long-term forecast is usable.

[0174] Because each driver's physical condition and individual circumstances are different, drivers may exhibit behaviors such as frequent blinking, speaking too fast, blood pressure exceeding medical standards, or varying heart rates. Therefore, pre-adjustment processing is required for each driver's specific driving situation: When a driver's blinking frequency exceeds the average blinking frequency, the blinking frequency should be multiplied by an adjustment coefficient k1 to reduce the weight of the average blinking frequency; when a driver speaks too fast, the speaking speed should be multiplied by an adjustment coefficient k2 to reduce the weight of the speaking speed; when a driver's heart rate is too fast under normal conditions, the heart rate should be multiplied by an adjustment coefficient k3 to reduce the weight of the heart rate; and when a driver's blood pressure is too high, the speaking speed should be multiplied by an adjustment coefficient k4 to reduce the weight of the blood pressure. These weighting coefficients are adjusted by staff on the gimbal 1.

[0175] This invention uses a near-field infrared camera 2 to acquire facial video images. A gimbal 1 processes these images to recognize the driver's facial expressions and eye fatigue. Simultaneously, a microphone 3 collects the driver's voice information, which the gimbal 1 then recognizes. A smart bracelet 4 collects the driver's heart rate and blood pressure information, and a smart seat cushion 5 monitors the driver's posture and body temperature. Based on facial expressions, eye fatigue, voice and text information, heart rate, blood pressure, and body temperature, the invention determines the driver's physiological state. The gimbal 1 then transmits the facial expressions, eye fatigue, voice, and text information to the gimbal. Textual information, heart rate information, blood pressure information, and body temperature information are input into the SVM model, which can then infer the driver's psychological state. Staff members use the gimbal 1 to obtain the driver's physiological and psychological states, and combine these states to determine the driver's driving status. If the driver's current driving status is found to be unfavorable, staff members use the gimbal 1 and the 5G network to contact the driver via the telephone located in the driver's cab, reminding and guiding the driver to safely drive the bus to a safe area to rest. This invention makes short-term and long-term predictions of the driver's driving status to determine whether the driver can continue driving.

[0176] The above embodiments are only used to illustrate and not limit the technical solutions of the present invention. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the present invention without departing from the spirit and scope of the present invention. Any modifications or partial substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring, characterized in that: Specifically, the following steps are included: (i) The physiological information of the driver is monitored by the physiological monitoring module and transmitted to the gimbal. The physiological information includes: the driver's facial video image, voice information, heart rate information, blood pressure information, sitting posture information and body temperature information. (ii) The gimbal processes physiological information to determine the driver's physiological state; (III) The gimbal infers the driver's psychological state based on the driver's physiological state; (iv) The driver’s driving status is obtained by combining the physiological and psychological states. If there is any abnormality, the staff will remind and guide the driver to drive the bus safely to a safe area to rest through the gimbal. (v) Make short-term and long-term predictions of the driver's physiological and psychological state to determine whether the driver can continue driving.

2. The method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring according to claim 1, characterized in that: The physiological monitoring module includes two near-field infrared cameras, a microphone, a smart wristband, and a smart seat cushion. The first near-field infrared camera is installed at the front of the dashboard in the driver's cab, and the second near-field infrared camera is installed above the driver's head in the cab. The resolution of the near-field infrared cameras is 1280*960, and the frame rate is set to 50 frames per second. The horizontal distance between the first near-field infrared camera and the driver is 70-100cm, and the vertical distance between the second near-field infrared camera and the driver's eyes is 30-40cm. The microphones are all located in the driver's cab. The smart wristband is worn on the driver's wrist and collects the driver's heart rate and blood pressure information. The smart seat cushion is placed on the driver's seat and is equipped with a six-axis attitude sensor integrating a three-axis accelerometer and a three-axis gyroscope, a half-bridge pressure sensor, and a human infrared temperature sensor. A data storage device is installed on the bus, and the two near-field infrared cameras, microphone, smart wristband, and smart seat cushion are all connected to the data storage device. The data storage device is connected to the gimbal signal via a wireless network.

3. The method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring according to claim 2, characterized in that: Step (1) is as follows: Two near-field infrared cameras capture the driver's driving status in real time and transmit the acquired facial video images of the driver to the data storage. The microphone acquires the driver's voice information in real time and transmits the voice information to the data storage. The smart bracelet collects the driver's heart rate and blood pressure information in real time through the equipped ECG signal measurement sensor and photoelectric sensor and transmits the heart rate and blood pressure information to the data storage. The smart seat accurately monitors the driver's sitting posture information in real time through the equipped six-axis posture sensor and half-bridge pressure sensor. The smart seat monitors the driver's body temperature information in real time through the equipped human infrared body temperature sensor. The smart seat transmits the monitored sitting posture information and body temperature information to the data storage. The data storage then transmits the collected facial video images, voice information, heart rate information, blood pressure information, sitting posture information and body temperature information of the driver to the gimbal via the 5G network.

4. The method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring according to claim 3, characterized in that: Step (II) specifically involves: The gimbal processes the facial video images to perform facial expression recognition and eye fatigue recognition for the driver. To avoid altering the original image pixel values ​​and ensure image clarity before and after rotation, the gimbal rotates each image by an angle θ around the origin. After the rotation transformation, the pixel coordinates are: Where (x, y) are the pixel coordinates of the original image; (x1, y1) are the coordinates of the image after the corresponding pixel (x, y) has been rotated. According to equation (1), each image is rotated by 90°, 180° and 270° respectively, that is, θ takes 90°, 180° and 270° respectively. The expanded dataset is 3 times the size of the original dataset. Facial expression recognition is performed in three steps: (I) Image preprocessing: Due to factors such as bumps during bus driving, the acquired facial images may have defects such as inconsistent focal length and contrast. Therefore, proper image preprocessing is essential. Image preprocessing includes face alignment, data augmentation, and image normalization: The face alignment module of the gimbal uses supervised descent to detect facial landmarks, such as eyebrows, nose, eyes, mouth and facial contours, for face recognition. The gimbal's data augmentation module generates new training samples, improving the generalization ability and robustness of the convolutional neural network model in the facial feature extraction module, avoiding network overfitting, providing a large amount of training data for subsequent expression classification, and improving the accuracy of the convolutional neural network model in recognition. Since the acquired images are RGB three-channel color images, with each channel ranging from 0 to 255, the computational load is too large. Therefore, grayscale normalization is necessary. Furthermore, the size of face images may vary depending on different road conditions. Images that are too large may contain redundant information, while images that are too small may overlook key features, reducing recognition accuracy. Additionally, machine learning algorithms generally require input images of a fixed size. Therefore, image size normalization is required. The gimbal's image normalization module performs grayscale and size normalization to ensure images have the same size and grayscale value range. Grayscale normalization uses the average method, maximum method, and weighted average method. The grayscale formula is as follows: Gray=0.3R(x,y)+0.59G(x,y)+0.11B(x,y) (2) In equation (2), Gray represents grayscale, R(x,y) refers to the red channel value, G(x,y) refers to the green channel value, and B(x,y) refers to the blue channel value. Size normalization employs nearest neighbor interpolation and bilinear interpolation algorithms; (II) Facial Feature Extraction: The gimbal's facial feature extraction module uses a convolutional neural network (CNN) model to automatically extract features from the input facial video images, achieving "end-to-end" learning. This CNN model consists of convolutional layers, activation layers, pooling layers, and fully connected layers. It is an improvement on the classic LeNet-5 model. The structure includes three convolutional layers (C1, C2, C3), three pooling layers (P1, P2, P3), where P1 is a max pooling layer, and P2 and P3 are average pooling layers. It also includes a fully connected layer (F1). The pooling operation combines max pooling and average pooling. The parameters of each layer in this CNN model are shown in Table 1. Table 1. Parameters of each layer in the convolutional neural network model ; After multiple experiments and tests, the outputs of each network layer after multi-scale pooling were selected to be feature matrices of different scales: 1×1×r, 2×2×r, and 3×3×r, where r is the number of feature maps and the value of r ranges from 1 to 100. The three feature matrices are arranged in columns to form a (13×r)×1 column vector. Finally, after feature fusion, a P-dimensional facial expression feature column vector with multiple scales and attributes is formed as the input vector x of the ELM classifier for facial expression classification. (III) Facial Expression Classification: The facial expression classification module of the gimbal uses an ELM classifier for expression recognition. The ELM classifier consists of an input layer, a hidden layer, and an output layer. During network training, the ELM classifier only needs to calculate the weight matrix between the hidden layer and the output layer. The weight matrix and bias values ​​between its input layer and the hidden layer are randomly generated and do not need to be calculated or iteratively updated. The number of neurons in the input layer, hidden layer, and output layer are d, l, and m, respectively. The output of the i-th hidden node is: g(x,w i ,b i )=g(xw i +b i ) (3) In equation (3), x is a column vector of facial expression features in P-dimensional space, g is the activation function, and w i Let b be the input weight vector between the i-th hidden node and all input nodes. i Let be the bias of the i-th hidden node, i = 1, 2, ..., l; g is the ReLU function, i.e.: g(x,w i ,b i )=max(0,xw i +b i ) (4) The connection between the input layer and the hidden layer is a mapping process. Because the input vector x is a feature vector in P-dimensional space, this connection is a process of mapping from P-dimensional space to l-dimensional space. The mapped feature vector of the input vector x is: h(x)=[g(x,w1,b1),g(x,w2,b2),...,g(x,w i ,b i )] (5) The output layer has m output nodes, where m is the number of different expressions, and each output node corresponds to one expression. The output weight between the i-th hidden node and the j-th output node is denoted as β. ij Where j = 1, 2, ..., m; therefore, the value of the j-th output node is: Therefore, the output vector of the input sample x in the hidden layer can be represented as: f(x)=f1(x),f2(x),...,f m (x)=h(x)β (7) In equation (7), During the testing phase, the input test sample test-x corresponds to the following expression category: label(text-x)=arg j=1,2,...,m maxf j (text-x) (9) The ELM classifier categorizes drivers' facial expressions into pleasure, focus, boredom, confusion, and frustration based on facial features extracted by the convolutional neural network model, as shown in Table 2. Table 2 Classification of Basic Facial Expressions Because when a person is fatigued, they will blink more frequently, keep their eyes closed for longer, yawn more, and in severe cases, even doze off. Under normal circumstances, a person blinks 10 to 25 times per minute, and the duration of eye closure for one blink is about 0.2 seconds. Based on this phenomenon, the three eye indicators that best represent the fatigue state are selected as the characteristic parameters of eye fatigue identification. The three eye indicators are: continuous eye closure time, eye closure frame rate, and blink frequency. The face landmark localization algorithm based on cascaded regression trees is used for human eye localization. The extracted left and right eye regions are represented by left and right rectangular bounding boxes. The localization calculation rules are as follows: W=1.6×W e ,H=3×H e (10) In Equation (10): We is the horizontal distance between human eye feature points 36 and 39, He is the average of the vertical distances between human eye feature points 37 and 41 and between 38 and 40, and W and H are the width and height of the located eye region; To accurately and quickly identify the open / closed state of the eyes, we calculated the aspect ratio between the height and width of the eyes. The aspect ratio when the eyes are open varies very little between individuals and remains completely unchanged regardless of uniform image scaling or facial rotation. The formula for calculating the eye aspect ratio is: In formula (11): P1 is the center point of the left side of the left eye, P2 is the center point of the upper left of the left eye, P3 is the center point of the upper right of the left eye, P4 is the center point of the right side of the left eye, P5 is the center point of the lower right of the left eye, and P6 is the center point of the lower left of the left eye. To accurately identify eye status, the average EAR of both eyes is taken as the feature for eye opening and closing recognition: EAR=Mean(EAR left ,EAR right ) (12) In formula (12): EAR left The aspect ratio of the left eye, EAR right The aspect ratio of the right eye; When the human eye is closed, although it may be affected by dark areas such as eyelashes and eyelids, the largest dark area will not appear in the pupil area. Therefore, compared with the open eye, the number of black pixels in the binary image is reduced when the eye is closed. Therefore, an adaptive threshold method is used to accumulate the difference and define two states, "state 0" and "state 1". When the difference of black pixels in the binary image of the human eye area is less than 0, it changes from "state 0" to "state 1". In "state 1", if the difference is less than the threshold T(t), the difference is accumulated and the state remains unchanged; if the difference is greater than or equal to the threshold T(t), the difference is not accumulated and the state changes to "state 0". Based on the above conclusions, during a blink, the EAR value first decreases until it approaches 0, and then gradually increases to the value of a normal open eye state. Let E be the threshold of EAR. When the EAR threshold is less than E, the eyes begin to close. When the EAR threshold is close to the normal open eye state, that is, greater than E, the eyes open. The gimbal processes the voice information to complete the driver's speech recognition. The gimbal's speech extraction module converts the raw speech signal into a digital signal through sampling quantization, determines the start and end of the speech through endpoint detection, adjusts the mid, high, and low frequency amplitudes appropriately through pre-emphasis processing, obtains frame-level speech sequences through frame segmentation and windowing, and finally extracts sound features, representing them in the form of feature vectors. Quantization sampling of speech signals involves measuring the information of the original speech signal in a digital system and converting it into a digital signal that a computer can recognize. This means discretizing the analog signal in terms of amplitude and time. Sampling involves sequentially capturing audio signals at fixed time intervals from a continuous analog signal and replacing the continuous analog signal with the amplitude value of the audio signal at the moment of sampling. The sampling frequency is set to 22.05kHz. Through sampling, the speech analog signal is transformed into a discrete-time signal. The amplitude value of the discrete-time signal is then quantized in stages, that is, the signal amplitude is divided into multiple intervals, and the sampled amplitude value within the same interval is represented by the same quantization value. Endpoint detection involves setting thresholds for energy and zero-crossing rate to detect endpoints in speech. With a frame length of 32ms, endpoint detection can pinpoint the start and end of valid speech, removing quiet and noise components. This not only shortens the speech time series but also improves speech emotion recognition after removing influencing factors, and helps distinguish between valid and invalid speech regions. Pre-emphasis is used because the speech signal loses high-frequency energy through radiation from the lips and nose of the vocal organs, leading to energy attenuation. To protect speech information, reduce loss, and prevent waveform distortion, a first-order digital filter is used to pre-emphasize the speech signal, compensating for the loss of certain components in the high-frequency part. The transfer function of the first-order digital filter is: H(z)=1-μz -1 (13) In equation (13), μ is the strengthening coefficient, with a value range of [0.9-1], and z -1 This indicates a delay of one unit of time; The sound feature extraction employs the Mel-frequency cepstral coefficient extraction method. Mel frequencies, based on human auditory characteristics, describe subjective pitch and exhibit a non-linear relationship with the objective pitch frequency f. A speech sample undergoes a series of preprocessing steps to obtain a frame of speech, which is then subjected to Fourier transform to obtain the spectral energy distribution. The Mel filter bank consists of a set of Mel-scale bandpass triangular filters. These triangular filters are not of equal bandwidth; rather, their bandwidth increases from low frequencies, resulting in decreasing frequency resolution. The input sound frequency domain signal is multiplied and added, and the output signal of each individual filter serves as a fundamental feature of the sound signal. This feature resides in the Mel spectral domain, and the perception of sound is linear. M triangular filters are used to smooth the spectrum and eliminate harmonics. Let the center frequency f(m) be given, and the transfer function H of the Mel filter bank be... r (k) is represented as: in, The power spectrum is obtained by taking the modulus of the spectrum of the frame speech signal and squaring it. The logarithmic energy of the output after passing through the filter bank is calculated. Finally, the MFCC coefficients are calculated by discrete cosine transform and then filtered. The voice recognition module of the pan-tilt uses a pattern matching method based on vector quantization (VQ) to recognize the extracted voice features. The pattern matching is divided into two steps: In the first step, a VQ model is trained with voiceprint feature parameters to generate a voiceprint template library; In the second step, the voiceprint features of multiple speakers' voices are input and subjected to matching operations with the already trained voiceprint model to determine whether they are the speakers in the template; For the VQ model, during the training stage, the MFCC feature vectors of the speakers are clustered based on the principle of minimum mean square error to generate a codebook composed of different codewords; during the recognition stage, the MFCC feature vectors of the voice to be recognized are used to calculate the distortion distance with the codebook in the template, and the mean square error and absolute value error are used for measurement. The sentence with the minimum distortion distance is selected as the MFCC feature vector for successful matching; The voice text module of the pan-tilt automatically converts the recognized voice into text and then performs text word segmentation. Text word segmentation is a very important step in natural language processing; to facilitate the subsequent selection of data features and make the analysis simpler, the text word segmentation method is used to decompose long sentences or paragraphs into a data structure with words as units; Word segmentation can be divided into three methods: ① The word segmentation method of dictionary matching, where the text is segmented according to specific rules and then matched with the words in the dictionary to determine whether word segmentation is correct; ② The word segmentation method based on statistics, with common algorithms such as the hidden Markov model and conditional random field; ③ The method based on deep learning; Word segmentation tools include Hanlp, Stanford Word Segmentation, THULAC, and Jieba. The word segmentation steps are as follows: (1) First, an efficient word graph scan is implemented based on the prefix dictionary to generate a directed acyclic graph (DAG) composed of all possible word formation situations of Chinese characters in the sentence; (2) Dynamic programming is used to find the maximum probability path to find the maximum segmentation combination based on word frequency; (3) For new words, an HMM model based on the word formation ability of Chinese characters is used, and the Viterbi algorithm is applied; Additional processing is often required before performing word frequency statistics on the word list because during the word segmentation process, a large number of function words such as "de" and "le", numeral-classifiers, and punctuation marks will be segmented into individual words. These words are called stop words. Stop words have a high occurrence frequency but no application value and should be removed when constructing the word frequency database for word frequency statistics; The pan-tilt processes the collected heart rate information and blood pressure information of the driver to judge the driver's physical health status. At the same time, it analyzes and processes the sitting posture information and effectively monitors the driver's driving situation in combination with real-time video images. The pan-tilt processes the body temperature information to determine whether the driver's body temperature is normal.

5. The method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring according to claim 4, characterized in that: Step (3) is specifically as follows: The pan-tilt inputs the facial expression information, voice and text information, heart rate information, blood pressure information, and body temperature information obtained in step (2) into the SVM model, and the SVM model can then infer the driver's mental state.

6. The method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring according to claim 5, characterized in that: Step (four) is as follows: Staff members obtain the driver's physiological and psychological state through the gimbal, and combine the physiological and psychological state to obtain the driver's driving status. If it is found that the driver's current driving status is not conducive to driving, the staff members use the gimbal and 5G network to make a phone call to the driver in the driver's cab to communicate with the driver, remind and guide the driver to drive the bus safely to a safe area to rest.

7. The method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring according to claim 6, characterized in that: Step (5) specifically involves: Drivers' body temperature should be monitored before each day's bus driving, compared with historical temperature data, to determine if the change was due to insufficient sleep the previous day or weather conditions. Short-term and long-term predictions of the driver's driving condition should also be made to determine whether the driver can continue driving. Short-term forecasts are based on the driver's physiological and psychological state monitored on the same day. Long-term forecasts are based on the driver's seven most recent short-term forecasts and seven most recent actual driving situations. The daily short-term forecast is scored by comparing it with the actual driving conditions of the day. When the daily score is greater than 80%, it indicates that the short-term forecast is highly accurate. When the average score of the last seven times is greater than 80%, the long-term forecast is usable.

8. The method for multi-dimensional data mining of driving safety based on driver physiological and psychological monitoring according to claim 6, characterized in that: Because each driver's physical condition and individual circumstances are different, drivers may have issues such as frequent blinking, speaking too fast, blood pressure exceeding medical standards, or varying heart rate frequencies. Therefore, pre-adjustment processing is needed for each driver's specific driving situation: when a driver's blinking frequency exceeds the average blinking frequency, the driver's blinking frequency should be multiplied by an adjustment factor k1, reducing the weight of the blinking frequency to the average blinking frequency; when a driver speaks too fast, the driver's speaking speed should be multiplied by an adjustment factor k2, reducing the weight of the speaking speed to the average speaking speed; when a driver's heart rate is too fast under normal circumstances, the driver's heart rate speed should be multiplied by an adjustment factor k3, reducing the weight of the driver's heart rate. When a driver's blood pressure is too high, the driver's speaking speed should be multiplied by an adjustment factor k4 to reduce the weight of the driver's blood pressure. This weighting factor is adjusted by staff on the gimbal.