Traffic management-oriented multi-mode driving state evaluation system
By integrating visual, voice, and physiological signal detection modules, the multimodal driving state assessment system enables collaborative perception and cross-validation of multiple risk factors. This solves the problem of low accuracy of assessment results in existing technologies, reduces the rate of missed and false alarms, and promptly detects complex dangerous states of drivers, thus assisting in traffic management.
Patent Information
- Application Number
- CN202511275240.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies rely on single-modal perception, which cannot achieve collaborative perception and cross-validation of multiple risk factors. This results in low accuracy of assessment results, high rates of missed and false alarms, and difficulty in timely detection of drivers' complex dangerous situations in complex scenarios.
A multimodal driving state assessment system is adopted, which integrates visual, voice and physiological signal detection modules. It uses lightweight deep learning networks and machine learning algorithms to collect and analyze multi-dimensional data, and achieves collaborative perception and cross-validation of multiple risk factors.
It enables multi-dimensional capture of complex dangerous situations, reduces the rate of missed alarms and false alarms, and can promptly detect dangerous situations of drivers in complex scenarios, assisting traffic management and reducing safety hazards.
Smart Images

Figure CN120997994A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of driving state assessment systems, specifically relating to a multimodal driving state assessment system for traffic management. Background Technology
[0002] Traffic management personnel's supervision of driving status refers to the use of relevant technical means by traffic management personnel to monitor, evaluate and manage the driving status of drivers in order to ensure road traffic safety.
[0003] However, existing technologies have the following shortcomings: they rely heavily on manual visual inspection and breathalyzers, and can only detect single driving risk dimensions such as fatigue, drunk driving, or physical condition in isolation, failing to achieve collaborative perception and cross-verification of multiple risk factors, and making it difficult to comprehensively capture complex dangerous situations; existing technologies rely on single-modal perception methods, which are easily affected by factors such as changes in ambient light, individual differences in drivers, sensor malfunctions, lighting, and weather, resulting in low accuracy of assessment results and high rates of missed and false alarms; in addition, in complex scenarios such as long-distance driving and peak traffic hours, it is difficult to detect driver fatigue, drunk driving, and other issues in a timely manner, posing safety hazards.
[0004] Therefore, a new method is urgently needed. Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal driving state assessment system for traffic management. This system enables collaborative perception and cross-validation of multiple risk factors, captures complex dangerous states from multiple dimensions, reduces the rate of missed alarms and false alarms, and promptly detects dangerous states in complex scenarios to assist traffic management and reduce safety hazards.
[0006] To achieve the above objectives, the present invention provides a multimodal driving state assessment system for traffic management, comprising a hardware subsystem and a host computer subsystem; the hardware subsystem includes a detection module connected to a central processing unit, the central processing unit being connected to a storage module and a communication and control module, and the communication and control module being connected to the host computer subsystem.
[0007] Preferably, the detection module includes a visual detection module, a voice detection module, and a physiological signal detection module.
[0008] Preferably, the host computer subsystem includes a communication and data receiving module, a machine learning and analysis module, and a visualization and decision-making module.
[0009] The preferred detection process of the visual inspection module is as follows: Relying on a lightweight deep learning network, the system processes the real-time video stream captured by the camera frame by frame and outputs face region data with key point coordinates. The trained facial feature extraction model is invoked to analyze the pupil diameter change sequence; combined with blinking frequency and the proportion of time with eyes closed, it helps to judge fatigue status; at the same time, by observing eye movement trajectory, it identifies inattention and distraction tendency. It integrates thermal imaging technology to capture facial features; analyzes muscle stiffness manifestations, and outputs preliminary judgment labels; The face localization coordinates, key point tracking trajectory, and pupil / thermal imaging analysis results are structured and integrated. After scheduling by the central processing unit, the integrated data is stored in the storage module and marked with the "Visual Detection Module - To be Fusion" label, waiting for multimodal fusion decision to call. After the data is stored in the storage module, the visual inspection module sends a "data ready" signal to the central processing unit and enters a low-power waiting state until it receives a data retrieval instruction from the fusion decision module. Then, it processes the real-time video stream captured by the camera frame by frame and outputs face region data with key point coordinates.
[0010] The preferred detection process for the voice detection module is as follows: For the acquired raw speech signal, Mel frequency cepstral coefficients, short-time energy, zero-crossing rate, speech rate, pitch variation, and pause interval acoustic features are extracted, and the acoustic features are constructed into feature vectors. By comparing with preset thresholds, it can identify slow speech and flat tone; A pre-trained acoustic model is used to detect slurred speech, voice tremor, repetition, and delayed response. After each analysis, the anomaly level and corresponding anomaly type label are output and temporarily stored with a timestamp. The original feature parameters and segmented anomaly judgment results are integrated, scheduled by the central processing unit, stored in the storage module, marked with the "speech detection module - to be fused" label, and await multimodal fusion call; After data storage is completed, a "data ready" signal is sent to the central processing unit, and the module enters a low-power listening state, waiting to receive data retrieval instructions from the fusion decision module.
[0011] The preferred detection process for the physiological signal detection module is as follows: Collect fingertip pulse wave signals to obtain a complete waveform including the rising edge of systole, the falling edge of diastole, and the dicrotic wave; perform noise reduction, anomaly repair, and feature extraction on the collected signals to obtain the preprocessed signal; The original waveform features, physiological parameters, and contact status markers are integrated, scheduled by the central processing unit, stored in the storage module, and marked with the "physiological signal module - to be fused" identifier, waiting for multimodal fusion to be called; After data storage is completed, a "data ready" signal is sent back to the central processing unit, and the module enters low-power standby mode until it receives a retrieval instruction from the fusion decision module.
[0012] Preferably, the hardware subsystem and the host computer subsystem are integrated and deployed on a portable embedded terminal, which is installed in temporary detection points, road enforcement stations, and portable traffic police detection equipment.
[0013] This invention also provides a multimodal driving state assessment method for traffic management, comprising the following steps: S1. Complete hardware driver loading and communication link establishment to provide a runtime environment for subsequent processes. All modules are ready to start. S2. Acquire facial images, extract features, output status labels and confidence scores, and synchronize them to the modality fusion queue; S3. Acquire speech signals, extract acoustic features and determine the state, and synchronize the results to the modality fusion queue; S4. Acquire pulse wave signals, analyze physiological characteristics and output results, and synchronize them to the modality fusion queue; S5. The communication and control module reads the data from S2-S4 in the modal fusion queue, performs weighted voting according to preset weights, and integrates the comprehensive driving state judgment result. S6 integrates the judgment results and raw data from S5, encrypts and packages them, and uploads them to the host computer. After the communication and data receiving module parses and distributes them, the machine learning module builds a model based on the data, identifies risks and predicts trends, and triggers early warnings based on state thresholds.
[0014] Therefore, the present invention employs the aforementioned multimodal driving state assessment system for traffic management. Compared with the prior art, the technical solution of the present invention has the following beneficial effects: (1) This invention uses the technical means of multimodal fusion of vision, speech and physiological signals to achieve collaborative perception and cross-verification of multiple risk factors (fatigue, drunk driving, abnormal emotions, abnormal conditions, etc.), thus overcoming the problem of single perception dimension and inability to fully capture complex dangerous states in traditional technology, and thus achieving the technical effect of capturing complex dangerous states in multiple dimensions. (2) The present invention adopts a multimodal sensing method to overcome the problems of weak anti-interference ability and susceptibility to environmental and individual differences in traditional technology, thereby achieving the technical effect of reducing interference and lowering the rate of missed alarms and false alarms. (3) The present invention adopts the technical means of real-time acquisition and analysis of multimodal data, which overcomes the problem that traditional technologies are difficult to detect dangerous situations in a timely manner in complex scenarios such as long-distance driving and peak traffic periods, thereby achieving the technical effect of timely detection of dangerous situations, assisting traffic management, and reducing safety hazards.
[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0016] Figure 1This is an architecture diagram of an embodiment of a multimodal driving state assessment system for traffic management according to the present invention; Figure 2 This is a visual detection flowchart of an embodiment of a multimodal driving state assessment system for traffic management according to the present invention; Figure 3 This is a flowchart illustrating the voice detection process of an embodiment of a multimodal driving state assessment system for traffic management according to the present invention. Figure 4 This is a flowchart illustrating the physiological signal detection process of an embodiment of a multimodal driving state assessment system for traffic management according to the present invention. Figure 5 This is a flowchart illustrating the multimodal fusion process of an embodiment of a multimodal driving state assessment system for traffic management according to the present invention. Figure 6 This is a flowchart illustrating the fusion decision-making process of an embodiment of a multimodal driving state assessment system for traffic management according to the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used in the present invention should have the ordinary meaning understood by those skilled in the art.
[0018] Example 1 like Figure 1 As shown, the present invention provides a multimodal driving state assessment system for traffic management, comprising a hardware subsystem and a host computer subsystem. The hardware subsystem includes a detection module, a central processing unit, a storage module, and a communication and control module. The detection module is connected to the central processing unit, the central processing unit is connected to the storage module and the communication and control module, and the communication and control module is connected to the host computer subsystem. The detection module adopts a miniaturized and low-power design, and can be integrated into portable embedded terminals (such as small industrial computers and edge computing devices) to meet the space constraints of temporary detection points and portable traffic police equipment; the detection module includes a visual detection module, a voice detection module and a physiological signal detection module. The visual detection module includes a multispectral near-infrared camera, an infrared illumination array, and a focusable optical lens, which can acquire high-contrast images in backlight / nighttime conditions. The acquired images are analyzed through a lightweight deep learning network to output fatigue / drunk driving / abnormal state judgment results, including facial features, thermal imaging anomalies, and pupil change data, which are then stored in the storage module. like Figure 2 As shown, the detection process of the visual inspection module includes face localization and tracking, specifically as follows: Relying on a lightweight deep learning network, the system processes the real-time video stream captured by the camera frame by frame. First, it performs driver face localization, accurately extracts multiple facial key points such as eyes, eyebrows, bridge of the nose, lips, and jaw, and automatically tracks the facial region based on the key point displacement vectors to cope with small head movements and posture changes of the driver, outputting facial region data with key point coordinates to lay the foundation for subsequent feature analysis; Fatigue and abnormality detection, specifically: The trained facial feature extraction model is invoked to analyze the pupil diameter change sequence, such as calculating the pupil contraction / dilation frequency and the difference between the maximum and minimum diameters per unit time; combined with blinking frequency and the proportion of time with eyes closed, it helps to judge fatigue status; at the same time, by observing the eye movement trajectory, it can identify inattention and distraction tendencies. By integrating thermal imaging technology, facial temperature distribution data is collected to identify typical thermal characteristics of drunk driving, such as facial flushing and abnormally elevated body temperature on the forehead / behind the ears. By linking with facial key point models, abnormal manifestations such as lip trembling and facial muscle stiffness (such as facial loss of control caused by the influence of medicinal wine) are analyzed, and preliminary judgment labels such as "fatigue tendency", "suspected drunk driving" and "abnormal attention" are output. Data integration and storage, specifically: The face localization coordinates, key point tracking trajectory, and pupil / thermal imaging analysis results are structured and integrated. After scheduling by the central processing unit, the integrated data is stored in the storage module and marked with the "Visual Detection Module - To be Fusion" label, waiting for multimodal fusion decision to call. Awaiting fusion readiness, specifically: After the data is stored in the storage module, the vision inspection module sends a "data ready" signal to the central processing unit and enters a low-power waiting state until it receives a data retrieval instruction from the fusion decision module or a new round of image acquisition trigger signal.
[0019] Through localization, feature analysis, structured storage, and collaborative waiting, the visual detection module completes multi-dimensional visual feature extraction, providing a dual-dimensional input of "fine-grained facial features + intelligent judgment labels" for multimodal fusion.
[0020] like Figure 3As shown, the voice detection module includes a microphone array and an audio preprocessing circuit. It adopts a passive audio acquisition mode, which does not require active wake-up and silently listens to audio in driving scenarios, adapting to complex in-vehicle noise environments. The speech detection module's detection process includes acquiring acoustic feature parameters, specifically: The original speech signal (divided into 25ms / frame and 10ms frame shift) was used to extract acoustic features such as Mel frequency cepstral coefficients, short-time energy, zero-crossing rate, speech rate, pitch variation, and pause interval. These acoustic features were then used to construct a feature vector to prepare for subsequent analysis. Segmented analysis and anomaly detection are as follows: By comparing with preset thresholds, such as normal speech rate range and tone fluctuation range, it can identify slow speech rate (below 30% of the threshold) and flat tone (fundamental frequency standard deviation < 5Hz). A pre-trained acoustic model is used to detect slurred speech, voice tremor, repetition, and delayed response. After each analysis, the "abnormality level" (0 - no abnormality, 1 - mild abnormality, 2 - severe abnormality) and the corresponding abnormality type label are output and temporarily stored with a timestamp. Data integration and storage, specifically: The original feature parameters and segmented anomaly judgment results are integrated, scheduled by the central processing unit, stored in the storage module, marked with the "speech detection module - to be fused" label, and await multimodal fusion call; Awaiting fusion readiness, specifically: After data storage is completed, a "data ready" signal is sent to the central processing unit. The module then enters a low-power listening state, waiting to receive data retrieval instructions from the fusion decision module or for a new audio segment to trigger the acquisition process.
[0021] By passively collecting data, extracting features in multiple dimensions, diagnosing segmented anomalies, and storing them in a structured manner, the speech detection module achieves non-intrusive speech state monitoring. This avoids the risk of active interaction interfering with driving and, through the dual judgment of "rules + models," identifies speech anomalies caused by drunk driving or fatigue, providing complementary inputs of "acoustic features + abnormal behavior labels" for multimodal fusion.
[0022] like Figure 4 As shown, the physiological signal detection module is based on the finger touch pulse sensor of the MAX86150 chip, integrates photoelectric volumetric recording (PPG) circuit, supports fingertip pulse wave acquisition, and has automatic contact detection and baseline drift suppression functions. The detection process of the physiological signal detection module includes pulse acquisition, specifically as follows: Collect fingertip pulse wave signals to obtain a complete waveform including the rising edge of systole, the falling edge of diastole, and the dicrotic wave; perform noise reduction, anomaly repair, and feature extraction on the collected signals to obtain the preprocessed signal; Data integration and storage, specifically: The original waveform features, physiological parameters, and contact status markers are integrated, scheduled by the central processing unit, stored in the storage module, and marked with the "physiological signal module - to be fused" identifier, waiting for multimodal fusion to be called; Awaiting fusion readiness, specifically: After data storage is completed, the module sends a "data ready" signal to the central processing unit and enters low-power standby mode until it receives a retrieval instruction from the fusion decision module or detects a change in finger contact status.
[0023] Through touch detection, intelligent calibration, precise data collection, and physiological feature mining, the physiological signal detection module enables non-invasive driving status monitoring. It utilizes the cardiovascular system status information contained in pulse waves to help identify driving risks such as fatigue, drunk driving, and psychogenic distraction.
[0024] The communication and control module implements data packaging, compression encryption, and network transmission optimization in the main control chip. It does not rely on large servers and can be embedded in the local processing unit of portable terminals to ensure that local data fusion and preliminary judgment can still be completed at temporary detection points without stable networks. It also establishes a connection with the communication and data receiving module of the host computer subsystem through network links (such as Ethernet, 5G, etc.) to realize data transmission. The host computer subsystem can be adapted to the fixed terminal display of the road enforcement station, and can also be presented on portable devices through lightweight clients (such as tablets and law enforcement recorders) to meet the needs of traffic police to view results and enforce the law quickly at temporary inspection points. It includes: communication and data receiving module, machine learning and analysis module, and visualization and decision-making module. The communication and data receiving module is responsible for receiving the recognition results and raw data after the fusion of visual, speech and physiological signals. It supports functions such as multi-terminal concurrent access, data caching, and structured storage, providing a data foundation for subsequent analysis.
[0025] Based on received historical data, the machine learning and analysis module uses deep learning algorithms to build driver behavior models, identify risk patterns, predict future driving trends, and support adaptive model training and strategy optimization. It optimizes performance through strategies such as computational acceleration, memory optimization, and pipelined parallelism to improve the system's intelligence level.
[0026] The visualization and decision support module is embedded in the decision support system. When the driving condition exceeds the set range, the system will issue a warning. When the driving condition does not exceed the set range, different decision suggestions will be provided. For example, when the fatigue condition is close to the set threshold, the system will output a decision to the staff and push a yellow alert to remind the driver to reduce driving time.
[0027] like Figures 5-6As shown, the present invention also provides a multimodal driving state assessment method for traffic management, comprising the following steps: S1. Automatically load the drivers for sub-modules such as camera module, microphone array, and pulse sensor, and check the connection status of each module to ensure normal function; at the same time, establish a communication link with the traffic monitoring platform and complete the preparation work before data upload; each hardware module is in a standby state, and the communication and control module and storage module are ready to provide the operating environment, data storage and transmission support. S2. Using a multispectral near-infrared camera, infrared illumination array, and adjustable focus optical lens, facial images of the driver are acquired with voice and host computer prompts. Infrared illumination ensures high-contrast imaging under nighttime or backlight conditions. A lightweight neural network model based on the MobileNetV2 backbone network combined with the BlazeFace structure is loaded to complete face localization and track key facial points such as eyes and eyebrows. Features such as the percentage of eye closure frames, yawning frequency, and nodding amplitude are extracted per unit time and input into the CNN-LSTM network to output fatigue, normal, and other state labels and confidence scores. The output time-stamped judgment results and confidence vectors are synchronized to the modality fusion queue. S3. A passive voice acquisition strategy is adopted. Under the guidance of voice and the host computer page, the driver is prompted intermittently to say specified sentences such as "I am currently in a normal state". The voice signal is captured through a microphone array. The system calls the voice feature extraction module to perform frame segmentation on the raw voice (25ms per frame, frame shift 10ms). After simultaneously filtering out wind noise and other noise, acoustic features such as MFCC, speech rate, and tone are extracted. After standardization, the features are input into the neural network model to identify abnormal behaviors such as slurred speech and voice tremor, and to determine states such as drunk driving and abnormal emotions. The output judgment results with timestamps and confidence vectors are synchronized to the modality fusion queue. S4. Using a physiological signal detection module equipped with a MAX86150 chip, the system guides the driver to place their finger in the detection area to collect pulse wave signals, guided by voice and computer prompts. The system automatically identifies signal stability and performs baseline calibration, preprocesses and corrects anomalies in the collected signals, dynamically calculates reference thresholds through multiple samplings, identifies peaks and feature points (the position where the signal rises to half the amplitude), analyzes physiological characteristics such as pulse rise time and heart rate variability, and constructs physiological feature vectors. These feature vectors are then input into a locally deployed lightweight model. The output timestamped judgment results and confidence vectors are synchronized to the modality fusion queue. The S5 communication and control module reads the multi-dimensional status labels and confidence levels output by S2-S4. In the main control chip, a decision-level fusion strategy is employed, performing weighted voting calculations according to preset modal weights. The system integrates the judgment results and confidence levels of fatigue, drunk driving, and abnormal states output by each module, calculates the fusion score of the three states, and finally outputs the comprehensive driving state judgment result. For example, if the comprehensive score is: fatigue 60 points > normal 30 points, then "fatigue" is judged; and the fused decision result is output. The S6 communication and control module compresses, encrypts, and packages the fused judgment results and raw data from S5, and transmits them to the host computer subsystem via the network. The host computer's communication and data receiving module parses, converts, and distributes the received data, providing data support for subsequent analysis. The machine learning and analysis module uses historical data and deep learning algorithms to build a driver behavior model, optimizes performance through strategies such as computational acceleration, identifies risk patterns, and predicts driving state trends. When the driving condition exceeds the set threshold, the system triggers real-time local warnings such as buzzers and voice broadcasts; and converts the warning information into structured data packets and uploads them to the traffic management platform.
[0028] Therefore, the present invention adopts the above-mentioned multimodal driving state assessment system for traffic management. This system realizes collaborative perception and cross-verification of multiple risk factors, can capture complex dangerous states from multiple dimensions, reduce the rate of missed alarms and false alarms, and can detect dangerous states in a timely manner in complex scenarios, assisting traffic management and reducing safety hazards.
[0029] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multimodal driving state assessment system for traffic management, characterized in that, It includes a hardware subsystem and a host computer subsystem; the hardware subsystem includes a detection module, which is connected to a central processing unit, which is connected to a storage module and a communication and control module, respectively, and the communication and control module is connected to the host computer subsystem.
2. The multimodal driving state assessment system for traffic management according to claim 1, characterized in that, The detection module includes a visual detection module, a voice detection module, and a physiological signal detection module.
3. The multimodal driving state assessment system for traffic management according to claim 1, characterized in that, The host computer subsystem includes a communication and data receiving module, a machine learning and analysis module, and a visualization and decision-making module.
4. The multimodal driving state assessment system for traffic management according to claim 1, characterized in that, The detection process of the visual inspection module is as follows: Relying on a lightweight deep learning network, the system processes the real-time video stream captured by the camera frame by frame and outputs face region data with key point coordinates. The trained facial feature extraction model is invoked to analyze the pupil diameter change sequence; combined with blinking frequency and the proportion of time with eyes closed, it helps to judge fatigue status; at the same time, by observing eye movement trajectory, it identifies inattention and distraction tendency. Fusion thermal imaging technology to capture facial features; Analyze muscle stiffness symptoms and output preliminary assessment labels; The face localization coordinates, key point tracking trajectory, and pupil / thermal imaging analysis results are structured and integrated. After scheduling by the central processing unit, the integrated data is stored in the storage module and marked with the "Visual Detection Module - To be Fusion" label, waiting for multimodal fusion decision to call. After the data is stored in the storage module, the visual inspection module sends a "data ready" signal to the central processing unit and enters a low-power waiting state until it receives a data retrieval instruction from the fusion decision module, and processes the real-time video stream captured by the camera frame by frame. Output face region data with key point coordinates.
5. A multimodal driving state assessment system for traffic management according to claim 1, characterized in that, The detection process of the voice detection module is as follows: For the acquired raw speech signal, Mel frequency cepstral coefficients, short-time energy, zero-crossing rate, speech rate, pitch variation, and pause interval acoustic features are extracted, and the acoustic features are constructed into feature vectors. By comparing with preset thresholds, it can identify slow speech and flat tone; A pre-trained acoustic model is used to detect slurred speech, voice tremor, repetition, and delayed response. After each analysis, the anomaly level and corresponding anomaly type label are output and temporarily stored with a timestamp. The original feature parameters and segmented anomaly detection results are integrated, scheduled by the central processing unit, stored in the storage module, marked with "speech detection module - to be fused", and await multimodal fusion call; After data storage is completed, a "data ready" signal is sent to the central processing unit, and the module enters a low-power listening state, waiting to receive data retrieval instructions from the fusion decision module.
6. The multimodal driving state assessment system for traffic management according to claim 1, characterized in that, The detection process of the physiological signal detection module is as follows: Collect fingertip pulse wave signals to obtain a complete waveform including the rising edge of systole, the falling edge of diastole, and the dicrotic wave; and perform noise reduction, anomaly repair, and feature extraction operations on the collected signals. Obtain the preprocessed signal; The original waveform features, physiological parameters, and contact status markers are integrated, scheduled by the central processing unit, stored in the storage module, and marked with the "physiological signal module - to be fused" identifier, waiting for multimodal fusion to be called; After data storage is completed, a "data ready" signal is sent back to the central processing unit, and the module enters low-power standby mode until it receives a retrieval instruction from the fusion decision module.
7. A multimodal driving state assessment system for traffic management according to claim 1, characterized in that, The hardware subsystem and the host computer subsystem are integrated and deployed on a portable embedded terminal, which is installed in temporary detection points, road enforcement stations, and portable traffic police detection equipment.
8. A multimodal driving state assessment method for traffic management, applied to a multimodal driving state assessment system for traffic management as described in any one of claims 1-7, characterized in that, Includes the following steps: S1. Complete hardware driver loading and communication link establishment to provide a runtime environment for subsequent processes. All modules are ready to start. S2. Acquire facial images, extract features, output status labels and confidence scores, and synchronize them to the modality fusion queue; S3. Acquire speech signals, extract acoustic features and determine the state, and synchronize the results to the modality fusion queue; S4. Acquire pulse wave signals, analyze physiological characteristics and output results, and synchronize them to the modality fusion queue; S5. The communication and control module reads the data from S2-S4 in the modal fusion queue, performs weighted voting according to preset weights, and integrates the comprehensive driving state judgment result. S6 integrates the judgment results and raw data from S5, encrypts and packages them, and uploads them to the host computer. After the communication and data receiving module parses and distributes them, the machine learning module builds a model based on the data, identifies risks, and predicts trends. And trigger an alert based on the status threshold.