Training Method for Fatigue Driving Detection Model, Fatigue Driving Detection Method and Device

Through a deep convolutional neural network combining multimodal data of infrared images and audio signals for feature extraction and model optimization, the accuracy problem of fatigue driving detection in complex environments in the prior art is solved, and more efficient and reliable fatigue driving detection is achieved.

CN119475259BActive Publication Date: 2025-06-24中电信数字城市科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510072987.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-24
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing fatigue driving detection methods have reduced detection accuracy in night or in environments with strong light interference, and detection based on voice signals is difficult to cover all fatigue manifestations, especially in environments with high noise interference.

Method used

A deep convolutional neural network is used to combine multimodal data of infrared images and audio signals for feature extraction. By constructing an undirected graph and loss function optimization model, the accuracy and reliability of fatigue driving detection are improved.

Benefits of technology

It improves the accuracy and reliability of fatigue driving detection, especially in complex environments, and enhances the performance of the detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119475259B_ABST
    Figure CN119475259B_ABST
Patent Text Reader

Abstract

The present invention provides a training method for a fatigue driving detection model, a fatigue driving detection method and device, which relate to the technical field of machine learning, and include: obtaining training sample data; wherein, the training sample data includes: infrared images of a driver, audio signals and driver states, and the driver states include: fatigue states and non-fatigue states; extracting feature vectors of the infrared images and audio signals based on a deep convolutional neural network to obtain infrared image samples and spectrogram samples; training a fatigue driving detection model based on the similarity of the infrared image samples and spectrogram samples and the driver states to obtain a trained fatigue driving detection model. The present invention improves the accuracy and reliability of fatigue driving detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and in particular, to a training method for a fatigue driving detection model, a fatigue driving detection method, and a device. Background Art

[0002] With the rapid development of China's economy, automobiles have become the most commonly used means of transportation for citizens. In recent years, the actual traffic flow on the road has been increasing rapidly. While automobiles provide convenience to life, they also bring a series of problems, such as frequent traffic accidents. According to a report, among the many causes of traffic accidents, fatigue driving is one of the factors with the largest number of fatalities. Therefore, fatigue driving has always been a problem that traffic management departments and drivers attach great importance to. If the driver can be reminded to rest in time when signs of fatigue driving occur, the occurrence of traffic accidents caused by driver fatigue driving can be effectively reduced.

[0003] In traditional fatigue detection methods, such as detection methods based on visual data or physiological signals, the detection accuracy will drop significantly in an environment with night or strong light interference. The fatigue driving detection technology based on voice signals is difficult to cover all fatigue manifestations, especially in a driving environment with large noise interference, and the reliability of detection will be affected. In summary, the existing fatigue driving detection methods all have certain defects. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a training method for a fatigue driving detection model, a fatigue driving detection method, and a device, so as to improve the accuracy and reliability of fatigue driving detection.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] In the first aspect, the present invention provides a training method for a fatigue driving detection model, including: obtaining training sample data; wherein, the training sample data includes: infrared images of the driver, audio signals, and the driver's state, and the driver's state includes: fatigue state and non-fatigue state; extracting feature vectors of the infrared images and audio signals based on a deep convolutional neural network to obtain infrared image samples and spectrogram samples; training the fatigue driving detection model based on the similarity of the infrared image samples and spectrogram samples and the driver's state to obtain a trained fatigue driving detection model.

[0007] Optionally, obtaining training sample data includes: collecting infrared images of the driver through an infrared camera at a preset sampling rate; collecting audio signals of the driver through a microphone array; synchronizing the infrared images and audio signals, and obtaining the driver's state corresponding to the synchronized infrared images and audio signals.

[0008] Optionally, training the fatigue driving detection model based on the similarity between the infrared image samples and the spectrogram samples and the driver state includes: calculating the similarity between the infrared image samples and the spectrogram samples based on the feature vectors of the infrared images and the feature vectors of the audio signals, and constructing an undirected graph based on the similarity between the infrared image samples and the spectrogram samples; wherein, the nodes of the undirected graph are infrared image samples or spectrogram samples, and the edges of the undirected graph are the similarities between the infrared image samples or spectrogram samples with the same driver state; determining the initial node feature matrix of the undirected graph, and determining the node feature matrix of each layer of the undirected graph based on the initial node feature matrix; constructing a first loss function based on the node feature matrix; constructing a second loss function based on the similarity between the infrared image samples and the spectrogram samples; determining the sum of the first loss function and the second loss function as the target loss function, and training the fatigue driving detection model by minimizing the target loss function.

[0009] Optionally, constructing a first loss function based on the node feature matrix includes: constructing the first loss function according to the following formula:

[0010]

[0011]

[0012]

[0013]

[0014] Wherein, represents the first loss function, represents the node i in the node feature matrix of the last layer of the undirected graph, represents the node i and the node j the weight between, is a parameter for enhancing the similarity between the infrared image samples and the spectrogram samples, is a parameter for reducing the similarity between the infrared image samples or the spectrogram samples, represents the boundary threshold, and are trade-off parameters.

[0015] Optionally, constructing a second loss function based on the similarity between the infrared image samples and the spectrogram samples includes: determining the spectrogram sample with the highest similarity and the same driver state as the infrared image sample as the positive sample, and determining the spectrogram sample with the highest similarity and a different driver state from the infrared image sample as the negative sample; constructing a second loss function based on the similarity between the infrared image sample and the positive sample and the negative sample; wherein, the second loss function is:

[0016]

[0017] Among them, represents the i th infrared image sample, represents the positive sample in the spectrogram sample corresponding to the i th infrared image sample, represents the negative sample in the spectrogram sample corresponding to the i th infrared image sample, N represents the number of samples, sim() represents the calculation of cosine similarity, represents the temperature coefficient, which is used to control the smoothness of the cosine similarity.

[0018] In a second aspect, the present invention provides a fatigue driving detection method, including: acquiring a real-time infrared image and a real-time audio signal of a driver; extracting feature vectors of the real-time infrared image and the real-time audio signal based on a deep convolutional neural network; inputting the feature vectors of the real-time infrared image and the real-time audio signal into a pre-trained fatigue driving detection model to obtain the driver's state; wherein, the fatigue driving detection model is trained based on the training method of any one of the fatigue driving detection models provided in the first aspect, and the driver's state includes: a fatigue state and a non-fatigue state.

[0019] In a third aspect, the present invention provides a training device for a fatigue driving detection model, including: a sample acquisition module, configured to acquire training sample data; wherein, the training sample data includes: an infrared image, an audio signal of a driver, and the driver's state, and the driver's state includes: a fatigue state and a non-fatigue state; a first feature extraction module, configured to extract feature vectors of the infrared image and the audio signal based on a deep convolutional neural network to obtain infrared image samples and spectrogram samples; a model training module, configured to train the fatigue driving detection model based on the similarity of the infrared image samples and the spectrogram samples and the driver's state to obtain a trained fatigue driving detection model.

[0020] In a fourth aspect, the present invention provides a fatigue driving detection device, including: a data acquisition module, configured to acquire a real-time infrared image and a real-time audio signal of a driver; a second feature extraction module, configured to extract feature vectors of the real-time infrared image and the real-time audio signal based on a deep convolutional neural network; a fatigue driving detection module, configured to input the feature vectors of the real-time infrared image and the real-time audio signal into a pre-trained fatigue driving detection model to obtain the driver's state; wherein, the fatigue driving detection model is trained based on the training method of any one of the fatigue driving detection models provided in the first aspect, and the driver's state includes: a fatigue state and a non-fatigue state.

[0021] Fifth aspect, the present invention provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any method provided in the above first aspect or the second aspect.

[0022] Sixth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of any method provided in the above first aspect or the second aspect.

[0023] The present invention brings the following beneficial effects:

[0024] The training method, fatigue driving detection method and device of the fatigue driving detection model provided by the present invention first obtain training sample data; wherein, the training sample data includes: infrared images of the driver, audio signals and driver states, and the driver states include: fatigue state and non-fatigue state; then based on a deep convolutional neural network, feature vectors of the infrared images and audio signals are extracted to obtain infrared image samples and spectrogram samples; finally, based on the similarity of the infrared image samples and spectrogram samples and the driver state, the fatigue driving detection model is trained to obtain a trained fatigue driving detection model. In the above method, by combining infrared images and spectrogram data and using a deep convolutional neural network for feature extraction of multi-modal data, the accuracy of fatigue driving detection is improved; at the same time, the above method can enhance the similarity of infrared image samples and spectrogram samples and reduce the similarity of infrared image samples and spectrogram samples in different driver states, thereby improving the performance of the fatigue driving detection model and further improving the accuracy and reliability of fatigue driving detection.

[0025] Other features and advantages of the present invention will be described in the following specification, and part of them will be obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims and drawings.

[0026] To make the above objectives, features and advantages of the present invention more obvious and understandable, the following specific embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings

[0027] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 Flow chart of a method for training a fatigue driving detection model provided by an embodiment of the present invention;

[0029] Figure 2 Detection flow chart of a fatigue driving detection provided by an embodiment of the present invention;

[0030] Figure 3 Flow chart of a fatigue driving detection method provided by an embodiment of the present invention;

[0031] Figure 4 Structural schematic diagram of a training device for a fatigue driving detection model provided by an embodiment of the present invention;

[0032] Figure 5 Structural schematic diagram of a fatigue driving detection device provided by an embodiment of the present invention;

[0033] Figure 6 Structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0034] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] Currently, in traditional fatigue detection methods, such as detection methods based on visual data or physiological signals, including facial feature analysis, eye movement detection, etc., a facial image of the driver is obtained through a camera, and features such as the eye closing frequency and the number of yawns are analyzed to determine the fatigue state. However, such methods are very sensitive to lighting conditions, and the detection accuracy will drop significantly in a night or strong light interference environment. Infrared technology can capture clear images in a completely dark or low-light environment, making up for the deficiencies of visible light, but there are still problems of information loss in the detection methods of infrared technology.

[0036] The fatigue driving detection technology based on voice signals can capture the differences in voice features of the driver in different fatigue states by analyzing the frequency and energy distribution changes of voice signals. For example, the voice in a fatigue state may show features such as a slower speech rate and a narrower frequency range. Although spectrograms show certain advantages in analyzing voice fatigue features, this method is also difficult to cover all fatigue manifestations, especially in a driving environment with large noise interference, and the reliability of detection will be affected.

[0037] Based on this, a training method, a fatigue driving detection method, and a device for a fatigue driving detection model provided by an embodiment of the present invention can improve the accuracy and reliability of fatigue driving detection.

[0038] For the convenience of understanding this embodiment, first, a training method for a fatigue driving detection model disclosed in an embodiment of the present invention will be introduced in detail. This method can be executed by an electronic device, such as a smartphone, a computer, a tablet computer, etc. Refer to Figure 1 The flowchart of a training method for a fatigue driving detection model shown in the figure schematically shows that this method mainly includes the following steps S101 to step S103:

[0039] Step S101: Obtain training sample data.

[0040] Among them, the training sample data includes: infrared images of the driver, audio signals, and the driver's state. The driver's state includes: a fatigue state and a non-fatigue state.

[0041] In one implementation, first, infrared images of the driver are collected by an infrared camera at a preset sampling rate; then, audio signals of the driver are collected by a microphone array; finally, the infrared images and audio signals are synchronously processed, and the driver's state corresponding to the synchronized infrared images and audio signals is obtained.

[0042] In specific implementation, to ensure the reliability of the data, in this embodiment, the time-domain information of the microphone array and the infrared camera is synchronized. Based on the sampling rate of the infrared camera of 1 frame / second (i.e., the preset sampling rate), every time the infrared camera collects one frame of data, the audio signals collected by the microphone array from t - 1s to the current moment t are intercepted at the same time. That is, one frame of infrared image collected at the current moment t and the 1s audio signal collected by the corresponding microphone array from t - 1s to the current moment t are taken as a set of data; then, the driver's state (fatigue state or non-fatigue state) corresponding to each set of infrared images and audio signals is determined.

[0043] Step S102: Extract feature vectors of the infrared images and audio signals based on a deep convolutional neural network to obtain infrared image samples and spectrogram samples.

[0044] In one implementation, for infrared images, first, the infrared images are decoded and framed; then, preprocessing operations such as filtering, noise reduction, and cropping are performed on the decoded and framed infrared images; finally, a deep convolutional neural network is used to extract the feature vectors of the preprocessed infrared images to obtain the embedding of infrared image modality A (i.e., infrared image samples) , where, i =1,2,…, N represents the i th infrared image sample, dis the dimension of the embedding space.

[0045] For an audio signal, first, the audio signal is framed, windowed, and subjected to a short-time discrete Fourier transform to obtain the spectrogram of the audio signal; then, a deep convolutional neural network is used to extract the feature vectors of the spectrogram, and the embedding of spectrogram modality B (i.e., the spectrogram sample) is obtained. , where i = 1, 2, …, N represents the i th spectrogram sample, d is the dimension of the embedding space.

[0046] Step S103: Train the fatigue driving detection model based on the similarity between the infrared image samples and the spectrogram samples and the driver state to obtain a trained fatigue driving detection model.

[0047] In one embodiment, in order to further improve the fatigue driving detection accuracy of the fatigue driving detection model in two different modalities, namely the infrared modality and the audio modality, an adaptive joint graph embedding optimization method is proposed in this embodiment. An undirected graph is constructed through the similarity between the infrared image samples and the spectrogram samples and the driver state, and the embedding of the undirected graph nodes is optimized by minimizing the loss function. In addition, a negative sample discovery strategy is provided in this embodiment to optimize the difficult-to-correctly-detect fatigue states. Specifically, the positive and negative samples are determined based on the similarity between the infrared image samples and the spectrogram samples, and a loss function is constructed according to the similarity between the positive and negative samples, and the model is optimized by minimizing the loss function.

[0048] In the embodiment of the present invention, by using the infrared image samples, the spectrogram samples, and their corresponding driver states, the trained fatigue driving detection model can be obtained by optimizing and training the fatigue driving detection model through the aforementioned adaptive joint graph embedding optimization method and negative sample discovery strategy.

[0049] The training method of the fatigue driving detection model provided by the present invention combines infrared images and spectrogram data, uses a deep convolutional neural network to extract the features of multi-modal data, and improves the accuracy of fatigue driving detection. At the same time, the above method can enhance the similarity between the infrared image samples and the spectrogram samples, reduce the similarity between the infrared image samples and the spectrogram samples under different driver states, thereby improving the performance of the fatigue driving detection model and further improving the accuracy and reliability of fatigue driving detection.

[0050] In one embodiment, for the aforementioned step S103, that is, when training the fatigue driving detection model based on the similarity between the infrared image samples and the spectrogram samples and the driver state, the following methods may be included but are not limited to, mainly including the following steps (1) to (5):

[0051] Step (1): Calculate the similarity between the infrared image samples and spectrogram samples based on the feature vectors of the infrared images and the audio signals, and construct an undirected graph based on the similarity between the infrared image samples and the spectrogram samples.

[0052] In specific implementation, first construct an undirected graph based on the similarity between the spectrogram samples and the infrared image samples , where the nodes of the undirected graph are infrared image samples or spectrogram samples, and the edges of the undirected graph are the similarities between the infrared image samples or spectrogram samples of the same driver state.

[0053] Specifically, each node in the undirected graph corresponds to a spectrogram sample or an infrared image sample, and the embedding vector (i.e., the feature vector) generated by the spectrogram or the infrared image is denoted as , for two nodes and , if the spectrogram samples or infrared image samples they correspond to belong to the same category (i.e., the driver states are the same, both belong to the fatigue state or both belong to the non-fatigue state), then an edge and is established between them, and the cosine similarity is used to calculate the weight of the edge . .

[0054]

[0055] Among them, represents the feature vector of node , represents the feature vector of node , represents the weight of the edge between nodes and .

[0056] Step (2): Determine the initial node feature matrix of the undirected graph, and determine the node feature matrix of each layer of the undirected graph based on the initial node feature matrix.

[0057] In specific implementation, assume that is the node feature matrix of the l -th layer, represents the initial node feature matrix, then the node feature matrix of the l + 1-th layer can be expressed as:

[0058]

[0059] Among them, Denote the l node feature matrix of the +1-th layer, l denote the node feature matrix of the -th layer, L , where denotes the number of layers of the undirected graph, , denote the l edge weight matrix of the -th layer, denote the degree matrix,

[0060] Step (3), construct the first loss function based on the node feature matrix.

[0061] In specific implementation, the first loss function can be constructed according to the following formula:

[0062]

[0063]

[0064]

[0065]

[0066] where denotes the first loss function, denotes the node i feature matrix of the node at the last layer of the undirected graph, denotes the node i and the node j weight between them, is a parameter used to enhance the similarity between infrared image samples and spectrogram samples, is a parameter used to reduce the similarity between infrared image samples or spectrogram samples, denotes the boundary threshold, and are trade-off parameters.

[0067] In the embodiments of the present invention, the embedding representation of the nodes can be optimized by minimizing the loss function of the graph convolutional network, enhancing the discrimination between different samples (audio modality or infrared modality), and improving the effectiveness and robustness of the fatigue driving detection model for fatigue state detection.

[0068] Step (4), construct the second loss function based on the similarity between infrared image samples and spectrogram samples.

[0069] In specific implementation, in order to enable the fatigue driving detection model to better detect the fatigue state at night or in low-light environments, in the cross-modal (audio modality and infrared modality) contrast learning of the embodiments of the present invention, a negative sample discovery strategy is designed, which can be specifically optimized for the fatigue state that is difficult to correctly detect. This method can significantly improve the detection ability of the model for fatigue driving, and can better identify the characteristics of the fatigue state especially in complex environments such as at night or in low light, thereby improving the overall performance. Its calculation method is as follows:

[0070] First, use cosine similarity to calculate the similarity between the samples in the infrared image modality A and all the samples in the spectrogram modality B :

[0071]

[0072] Then, determine the spectrogram sample with the same driver state as the infrared image sample and the highest similarity as the positive sample, and determine the spectrogram sample with a different driver state from the infrared image sample and the highest similarity as the negative sample.

[0073] Specifically, select the sample that is most similar but belongs to a different driver state (such as "fatigue" or "non-fatigue") as the negative sample (that is, determine the spectrogram sample with a different driver state from the infrared image sample and the highest similarity as the negative sample):

[0074]

[0075] At the same time, select the sample that is most similar but belongs to the same driver state as the positive sample.

[0076] Finally, construct a second loss function based on the similarity between the infrared image sample and the positive and negative samples; among them, the second loss function is:

[0077]

[0078] Among them, represents the i th infrared image sample, represents the positive sample in the spectrogram sample corresponding to the i th infrared image sample, represents the negative sample in the spectrogram sample corresponding to the i th infrared image sample, N represents the number of samples, sim() represents cosine similarity calculation, Represents the temperature coefficient, which is used to control the smoothness of the cosine similarity.

[0079] Step (5), determining the sum of the first loss function and the second loss function as the target loss function, and training the fatigue driving detection model by minimizing the target loss function.

[0080] In specific implementation, the target loss function is:

[0081]

[0082] where, is the trade-off parameter.

[0083] In the embodiment of the present invention, the final target loss function combines the negative sample mining loss and the cross-modal contrast loss, which can achieve a comprehensive optimization of the fatigue driving detection model.

[0084] The above method provided by the embodiment of the present invention, by combining infrared images and spectrogram data, uses a deep learning model to perform feature extraction and fusion analysis of multi-modal data, realizes fast and accurate fatigue detection, overcomes the limitations of single-modal methods in complex driving environments, and at the same time avoids the dependence on external physical sensors, improving the user experience. Through this multi-modal fusion detection method, an efficient, reliable and environment-adaptive solution is provided for fatigue driving monitoring.

[0085] For ease of understanding, the embodiment of the present invention also provides a detection process for fatigue driving detection, as shown in Figure 2 shown, mainly including the following steps 1 to step 7:

[0086] Step 1: Collect infrared images through an infrared camera and collect audio signals through a microphone array.

[0087] Step 2: In the image preprocessing module, first decode and extract frames from the infrared image, and then perform preprocessing operations such as filtering, noise reduction, and cropping.

[0088] Step 3: In the audio signal preprocessing module, perform frame division, windowing, and short-time discrete Fourier transform processing on the audio signal to obtain the spectrogram of the audio signal.

[0089] Step 4: Use a deep convolutional neural network in the image feature extraction module to extract the feature vector of the infrared image, and obtain the embedding of the infrared image modality A .

[0090] Step 5: Use a deep convolutional neural network in the audio feature extraction module to extract the feature vector of the spectrogram, and obtain the embedding of the spectrogram modality B .

[0091] Step 6: Input the feature vectors of the spectrogram and the infrared image into the joint embedding space module. Perform adaptive joint graph embedding optimization and negative sample discovery in the joint embedding space module respectively.

[0092] Specifically, the process of adaptive joint graph embedding optimization refers to the aforementioned steps (1) to (3), and the process of negative sample discovery refers to the aforementioned step (4), which will not be elaborated here.

[0093] Step 7: Output the driver's fatigue state.

[0094] In the above method provided by the embodiments of the present invention, through the adaptive joint graph embedding optimization method, the fatigue state can be detected more efficiently in the shared joint embedding space. Through this method, the similarity between the spectrogram (audio modality) and the infrared image (infrared modality) is enhanced, especially in the same fatigue state (such as "fatigue" or "non-fatigue"), while effectively reducing the similarity of cross-modal (infrared modality and audio modality) samples in different states, thereby improving the cross-modal learning ability of the model and significantly improving the detection accuracy and robustness of the fatigue state. Through the negative sample discovery strategy, the fatigue state that is difficult to detect correctly can be optimized specifically. This method can significantly improve the model's detection ability for fatigue driving, especially in complex environments such as night or low light, where the characteristics of the fatigue state can be better identified, thereby improving the overall performance.

[0095] The present invention also provides a fatigue driving detection method, see Figure 3 The flowchart of a fatigue driving detection method shown, which shows that the method mainly includes the following steps S301 to step S303:

[0096] Step S301: Obtain the driver's real-time infrared image and real-time audio signal.

[0097] Step S302: Extract the feature vectors of the real-time infrared image and the real-time audio signal based on the deep convolutional neural network.

[0098] Step S303: Input the feature vectors of the real-time infrared image and the real-time audio signal into a pre-trained fatigue driving detection model to obtain the driver's state; wherein, the fatigue driving detection model is trained by the training method of the fatigue driving detection model according to any one of the aforementioned embodiments, and the driver's state includes: fatigue state and non-fatigue state.

[0099] The above-mentioned fatigue driving detection method provided by the present invention uses a fatigue driving detection model to detect fatigue driving. By combining infrared images and spectrogram data, stable detection can be achieved under various lighting conditions (especially at night or in low-light environments); at the same time, spectrograms are used to extract fatigue-related features from the driver's speech, analyze the changes in speech frequency and energy, and infrared images can capture the driver's facial temperature and expression changes, thereby providing more comprehensive physiological and behavioral information and improving the accuracy of fatigue state recognition.

[0100] For the training method of the fatigue driving detection model provided in the foregoing embodiment, the present invention also provides a training device for the fatigue driving detection model. Refer to Figure 4 the structural schematic diagram of a training device for a fatigue driving detection model shown in

[0101] A sample acquisition module 401, configured to acquire training sample data; wherein, the training sample data includes: infrared images of the driver, audio signals, and driver states, and the driver states include: fatigue states and non-fatigue states;

[0102] A first feature extraction module 402, configured to extract feature vectors of the infrared image and the audio signal based on a deep convolutional neural network to obtain infrared image samples and spectrogram samples;

[0103] A model training module 403, configured to train the fatigue driving detection model based on the similarity between the infrared image samples and the spectrogram samples and the driver states to obtain a trained fatigue driving detection model.

[0104] The above-mentioned training device for the fatigue driving detection model provided by the present invention combines infrared image and spectrogram data, uses a deep convolutional neural network to extract features of multi-modal data, and improves the accuracy of fatigue driving detection; at the same time, the above device can enhance the similarity between infrared image samples and spectrogram samples, reduce the similarity between infrared image samples and spectrogram samples under different driver states, thereby improving the performance of the fatigue driving detection model and further improving the accuracy and reliability of fatigue driving detection.

[0105] For the fatigue driving detection method provided in the foregoing embodiment, the embodiment of the present invention also provides a fatigue driving detection device. Refer to Figure 5 the structural schematic diagram of a fatigue driving detection device shown in

[0106] A data acquisition module 501, configured to acquire real-time infrared images and real-time audio signals of the driver;

[0107] The second feature extraction module 502 is configured to extract feature vectors of the real-time infrared image and the real-time audio signal based on a deep convolutional neural network;

[0108] The fatigue driving detection module 503 is configured to input the feature vectors of the real-time infrared image and the real-time audio signal into a pre-trained fatigue driving detection model to obtain the driver's state.

[0109] Wherein, the fatigue driving detection model is trained by using the training method of the fatigue driving detection model according to any one of the foregoing embodiments, and the driver's state includes: a fatigue state and a non-fatigue state.

[0110] The above-mentioned fatigue driving detection device provided by the present invention uses a fatigue driving detection model to detect fatigue driving. By combining infrared images and spectrogram data, stable detection can be achieved under various lighting conditions (especially at night or in low-light environments); at the same time, the spectrogram is used to extract fatigue-related features from the driver's voice, analyze the changes in voice frequency and energy, and the infrared image can capture the facial temperature and expression changes of the driver, thereby providing more comprehensive physiological and behavioral information and improving the accuracy of fatigue state recognition.

[0111] It should be noted that the device provided by the embodiments of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding contents in the foregoing method embodiments.

[0112] The embodiments of the present invention further provide an electronic device. Specifically, the electronic device includes a processor and a storage device; a computer program is stored on the storage device, and when the computer program is run by the processor, it executes the method according to any one of the above embodiments.

[0113] Figure 6 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes: a processor 60, a memory 61, a bus 62, and a communication interface 63. The processor 60, the communication interface 63, and the memory 61 are connected through the bus 62; the processor 60 is configured to execute an executable module stored in the memory 61, such as a computer program.

[0114] Wherein, the memory 61 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 63 (which may be wired or wireless), a communication connection between the system network element and at least one other network element is realized, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.

[0115] The bus 62 can be an ISA bus, a PCI bus, an EISA bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 it is only represented by a single bidirectional arrow in the figure, but it does not mean that there is only one bus or one type of bus.

[0116] Among them, the memory 61 is used to store a program. After receiving an execution instruction, the processor 60 executes the program. The method executed by the device defined by the flow process disclosed in any embodiment of the foregoing embodiments of the present invention can be applied to the processor 60 or implemented by the processor 60.

[0117] The processor 60 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 60 or by instructions in the form of software. The above-mentioned processor 60 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 61, and the processor 60 reads the information in the memory 61 and combines its hardware to complete the steps of the above method.

[0118] The computer program product of the readable storage medium provided by the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the foregoing method embodiments, which will not be elaborated herein.

[0119] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0120] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A training method for a fatigue driving detection model, characterized in that: include: Acquire training sample data; wherein the training sample data includes: an infrared image of the driver, an audio signal and a driver state, wherein the driver state includes: a fatigue state and a non-fatigue state; Extracting feature vectors of the infrared image and the audio signal based on a deep convolutional neural network to obtain infrared image samples and spectrogram samples; The fatigue driving detection model is trained based on the similarity between the infrared image sample and the spectrogram sample and the driver state to obtain a trained fatigue driving detection model; wherein the training process includes: calculating the similarity between the infrared image sample and the spectrogram sample based on the feature vector of the infrared image and the feature vector of the audio signal, and constructing an undirected graph based on the similarity between the infrared image sample and the spectrogram sample; wherein the nodes of the undirected graph are infrared image samples or spectrogram samples, and the edges of the undirected graph are the similarities between infrared image samples or spectrogram samples of the same driver state; determining an initial node feature matrix of the undirected graph, and determining a node feature matrix of each layer of the undirected graph based on the initial node feature matrix; constructing a first loss function based on the node feature matrix; constructing a second loss function based on the similarity between the infrared image sample and the spectrogram sample; determining the sum of the first loss function and the second loss function as the target loss function, and training the fatigue driving detection model by minimizing the target loss function; Constructing a first loss function based on the node feature matrix includes: constructing the first loss function according to the following formula: in, represents the first loss function, Representation Node i The node feature matrix at the last layer of the undirected graph, Representation Node i and nodes j The weight between is a parameter for improving the similarity between the infrared image sample and the spectrogram sample, is a parameter for reducing the similarity between the infrared image samples or the spectrogram samples, represents the boundary threshold, and To balance the parameters, Representation Node j The node feature matrix at the last level of the undirected graph.

2. The training method according to claim 1, characterized in that: Get training sample data, including: The infrared image of the driver is collected by an infrared camera at a preset sampling rate; collecting an audio signal of the driver through a microphone array; The infrared image and the audio signal are processed synchronously, and the driver status corresponding to the synchronized infrared image and the audio signal is obtained.

3. The training method according to claim 1, characterized in that: Constructing a second loss function based on the similarity between the infrared image sample and the spectrogram sample, comprising: Determine the sound spectrogram sample with the same driver status as the infrared image sample and the highest similarity as a positive sample, and determine the sound spectrogram sample with a different driver status as the infrared image sample and the highest similarity as a negative sample; A second loss function is constructed based on the similarity between the infrared image sample and the positive sample and the negative sample; wherein the second loss function is: in, Indicates i Infrared image samples, Indicates i The positive samples in the spectrogram samples corresponding to the infrared image samples, Indicates i Negative samples in the spectrogram samples corresponding to infrared image samples, N Indicates the number of samples, sim() indicates the cosine similarity calculation, Represents the temperature coefficient, which is used to control the smoothness of the cosine similarity.

4. A method for detecting fatigue driving, characterized in that: include: Acquire the driver's real-time infrared image and real-time audio signal; Extracting feature vectors of the real-time infrared image and the real-time audio signal based on a deep convolutional neural network; The feature vectors of the real-time infrared image and the real-time audio signal are input into a pre-trained fatigue driving detection model to obtain the driver's state; wherein the fatigue driving detection model is trained based on the training method of the fatigue driving detection model according to any one of claims 1 to 3, and the driver's state includes: fatigue state and non-fatigue state.

5. A training device for a fatigue driving detection model, characterized in that: include: A sample acquisition module, used to acquire training sample data; wherein the training sample data includes: an infrared image of the driver, an audio signal and a driver state, wherein the driver state includes: a fatigue state and a non-fatigue state; A first feature extraction module, used to extract feature vectors of the infrared image and the audio signal based on a deep convolutional neural network to obtain infrared image samples and spectrogram samples; A model training module, used for training a fatigue driving detection model based on the similarity between the infrared image sample and the spectrogram sample and the driver state to obtain a trained fatigue driving detection model; wherein the training process includes: calculating the similarity between the infrared image sample and the spectrogram sample based on the feature vector of the infrared image and the feature vector of the audio signal, and constructing an undirected graph based on the similarity between the infrared image sample and the spectrogram sample; wherein the nodes of the undirected graph are infrared image samples or spectrogram samples, and the edges of the undirected graph are the similarities between infrared image samples or spectrogram samples of the same driver state; determining an initial node feature matrix of the undirected graph, and determining a node feature matrix of each layer of the undirected graph based on the initial node feature matrix; constructing a first loss function based on the node feature matrix; constructing a second loss function based on the similarity between the infrared image sample and the spectrogram sample; determining the sum of the first loss function and the second loss function as a target loss function, and training the fatigue driving detection model by minimizing the target loss function; Constructing a first loss function based on the node feature matrix includes: constructing the first loss function according to the following formula: in, represents the first loss function, Representation Node i The node feature matrix at the last layer of the undirected graph, Representation Node i and nodes j The weight between is a parameter for improving the similarity between the infrared image sample and the spectrogram sample, is a parameter for reducing the similarity between the infrared image samples or the spectrogram samples, represents the boundary threshold, and To balance the parameters, Representation Node j The node feature matrix at the last level of the undirected graph.

6. A fatigue driving detection device, characterized in that: include: A data acquisition module, used to acquire the driver's real-time infrared image and real-time audio signal; A second feature extraction module, used for extracting feature vectors of the real-time infrared image and the real-time audio signal based on a deep convolutional neural network; A fatigue driving detection module is used to input the feature vectors of the real-time infrared image and the real-time audio signal into a pre-trained fatigue driving detection model to obtain the driver state of the driver; wherein the fatigue driving detection model is trained based on the training method of the fatigue driving detection model according to any one of claims 1 to 3, and the driver state includes: a fatigue state and a non-fatigue state.

7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 3 or the steps of the method according to claim 4.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 or the steps of the method according to claim 4 are executed.

Citation Information

Patent Citations

  • Driving takeover reminding device and method based on video recognition and steering wheel sensor

    CN111619580A

  • Fatigue driving identification method and system, electronic equipment and storage medium

    CN115331204A