Driver high-risk behavior dynamic identification system based on computer vision

By combining a multi-camera layout and an improved convolutional neural network model with dynamic time warping algorithm and analytic hierarchy process, the problems of insufficient multi-dimensional feature capture and real-time performance in existing driver high-risk behavior recognition systems are solved, achieving accurate recognition and risk assessment of driver behavior.

CN120853145APending Publication Date: 2025-10-28TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510937587.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing high-risk driver behavior recognition systems have shortcomings in multi-camera layout, feature extraction algorithms, risk assessment models, and real-time optimization, resulting in low recognition accuracy in complex scenarios, susceptibility to background noise interference, low computational efficiency, and inability to dynamically adapt to different driving environments and individual differences.

Method used

By employing a multi-camera layout, an improved convolutional neural network model to enhance attention mechanisms, a behavior pattern database, and a three-layer neural network architecture, combined with dynamic time warping algorithms and analytic hierarchy process, accurate extraction of driver behavior and risk assessment can be achieved.

Benefits of technology

It improves the efficiency and accuracy of feature extraction, enabling accurate identification and scientific judgment of high-risk behaviors of drivers, timely judgment of dangerous behaviors such as fatigue and distraction, and dynamic adaptation in different environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853145A_ABST
    Figure CN120853145A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent traffic and computer vision, in particular to a driver high-risk behavior dynamic recognition system based on computer vision, which comprises an image acquisition module, a feature extraction module, a behavior analysis module, a risk assessment module and an early warning output module, the image acquisition module, the feature extraction module, the behavior analysis module, the risk assessment module and the early warning output module are connected through a high-speed data bus; the image acquisition module is connected with the feature extraction module and transmits acquired image data to the feature extraction module; the feature extraction module processes the image data and then transmits the image data to the behavior analysis module; an analysis result of the behavior analysis module is transmitted to the risk assessment module; the risk assessment module carries out risk grade determination according to the analysis result and transmits the determination result to the early warning output module. According to the method, an improved CNN algorithm, an attention mechanism algorithm and a DTW algorithm are fused, a three-layer neural network is combined, and efficient recognition and risk level scientific evaluation of driver behaviors are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent transportation and computer vision technology, and in particular to a dynamic recognition system for high-risk driver behaviors based on computer vision. Background Technology

[0002] With the development of intelligent transportation technology, driver high-risk behavior recognition systems have gradually evolved from early single-modal detection based on physical sensors to multi-modal recognition combining computer vision technology. Early technologies collected driver operation data through onboard sensors, but their ability to capture behavioral features was limited. Computer vision-based solutions, on the other hand, use cameras to acquire image information and combine it with traditional machine learning algorithms (such as support vector machines and hidden Markov models) to perform behavior analysis. In recent years, with the popularization of deep learning technology, models such as convolutional neural networks have been introduced into the feature extraction stage, promoting the intelligent development of recognition systems.

[0003] However, existing technologies still have significant shortcomings in practical applications: First, most systems use a single camera or a single in-vehicle perspective layout, making it difficult to comprehensively capture the multi-dimensional features of the driver's facial expressions, body postures, and driving environment, resulting in insufficient recognition accuracy in complex scenarios; Second, feature extraction algorithms lack a focusing mechanism for key information, and traditional convolutional neural networks are easily affected by background noise when processing driver behavior data, and have high feature dimensions and low computational efficiency; Third, risk assessment models mostly use fixed weights or a single neural network architecture, which cannot dynamically adapt to different driving environments (such as rainy days, strong light) and individual driver differences, and lack a real-time update mechanism, making the model accuracy prone to decay after long-term operation; In addition, the real-time optimization of existing systems is insufficient, and when faced with high frame rate image data, the serial computing architecture is difficult to meet the timeliness requirements of dynamic recognition. Summary of the Invention

[0004] The purpose of this invention is to overcome the above-mentioned problems and provide a dynamic recognition system for high-risk driver behaviors based on computer vision. To achieve the above objective, this invention adopts the following technical solution:

[0005] A computer vision-based dynamic recognition system for high-risk driver behaviors includes an image acquisition module, a feature extraction module, a behavior analysis module, a risk assessment module, and a warning output module. These modules are connected via a high-speed data bus. The image acquisition module is connected to the feature extraction module, transmitting the acquired image data to it. The feature extraction module processes the image data and then transmits it to the behavior analysis module. The behavior analysis module transmits its analysis results to the risk assessment module. The risk assessment module determines the risk level based on the analysis results and transmits the determination result to the warning output module.

[0006] Furthermore, the image acquisition module includes no fewer than three cameras: the first camera is positioned directly in front of the steering wheel, the second camera is positioned above the driver's left door, and the third camera is positioned above the driver's right door. Each camera uses a CMOS image sensor of the same specifications to acquire image data at a constant frame rate and resolution, and transmits the image data to the feature extraction module through a high-speed data interface.

[0007] Furthermore, the feature extraction module employs an improved convolutional neural network model, adding an attention mechanism module to the original network structure. The convolutional layer calculation formula is as follows:

[0008]

[0009] in, For the j-th feature map in the l-th layer, M j Given the set of input feature maps, For the i-th feature map of the (l-1)th layer, The convolution kernel in the l-th layer connects the feature maps i and j. Here, f is the bias term, and f is the activation function;

[0010] The feature extraction module normalizes the image data and extracts the driver's facial expression features, body posture features, and driving environment features sequentially through convolutional layers, pooling layers, and attention mechanism modules. After compressing the dimensions of the extracted features, they are transmitted to the behavior analysis module in the form of feature vectors with pre-defined dimensions.

[0011] Furthermore, the behavior pattern database constructed by the behavior analysis module is stored in a non-volatile storage medium using an embedded database management system; the database stores data on fatigue driving behavior patterns, distracted driving behavior patterns, and violation operation behavior patterns; the behavior analysis module employs an optimized version of the dynamic time warping algorithm, and its formula for calculating the similarity between the feature vector and each behavior pattern is as follows:

[0012] D(i,j)=d(x i ,y j )+min{D(i-1,j),D(i,j-1),D(i-1,j-1)}

[0013] Where D(i,j) represents the cumulative distance between the i-th and j-th feature points, d(x i ,y j This represents the distance between two feature points calculated using a distance metric algorithm.

[0014] The behavior analysis module compares the calculated similarity with a preset threshold. When the similarity is greater than the threshold, it is determined that the match is successful, and the matching result is transmitted to the risk assessment module in binary code form.

[0015] Furthermore, the risk assessment model established by the risk assessment module is based on a three-layer neural network architecture, including an input layer, a hidden layer, and an output layer. The number of nodes in the input layer corresponds to the number of behavior categories output by the behavior analysis module, the number of nodes in the hidden layer is the integer part of the arithmetic mean of the number of nodes in the input layer and the number of nodes in the output layer, and the number of nodes in the output layer is three. The weights of each high-risk behavior factor are determined using the analytic hierarchy process (AHP), and the calculation formula is as follows:

[0016]

[0017] Among them, W i Let a be the weight of the i-th factor. ij To determine the elements of the matrix, the judgment matrix is ​​constructed based on expert experience and historical data, where n is the number of factors;

[0018] After receiving the results from the behavior analysis module, the risk assessment module performs a weighted calculation on the behavior category code and its corresponding weight, outputs the risk level probability distribution through the Softmax function, takes the level with the highest probability as the final judgment result, and transmits it to the early warning output module in the form of an ASCII string.

[0019] Furthermore, the warning output module includes an audible and visual alarm unit, a variable frequency sound alarm device, and an information push unit; the audible and visual alarm unit uses a programmable LED light source and switches between different colors of light through pulse width modulation technology; the variable frequency sound alarm device outputs different frequency prompt sounds by changing the frequency of the drive signal; the information push unit integrates a wireless communication module, uses the TCP / IP protocol to send risk level information to the vehicle central control system terminal in JSON format, and sends brief risk level information to the driver's mobile terminal through the short message service protocol.

[0020] Furthermore, the feature extraction module adopts a multi-core computing platform; through the CUDA parallel computing programming model, the convolutional and pooling layer operations in the improved convolutional neural network model are mapped to thread blocks and thread grids for parallel processing; CUDA streaming technology is used to distribute data preprocessing, feature extraction, and result post-processing operations to different streams to achieve multi-task pipeline processing; frequently accessed data is stored through shared memory and access paths are optimized using memory merging technology; and double buffering technology is used to achieve overlapping execution of the current data block computation and the next data block loading.

[0021] Furthermore, the behavior analysis module includes an online learning unit that employs an incremental learning algorithm. When the system runs continuously for a preset duration and collects a preset number of new unmatched image data, the online learning mechanism is triggered. The new data is labeled, and the model parameters in the behavior pattern database are updated using stochastic gradient descent. Each update selects a preset number of image data as a training batch, sets a preset learning rate, and recalculates the similarity threshold of each behavior pattern in the behavior pattern database after the update is completed.

[0022] Furthermore, the risk assessment module includes an environmental sensor data acquisition unit, comprising a temperature sensor, a humidity sensor, a raindrop sensor, and a light sensor; when any environmental parameter detected by any environmental sensor exceeds a preset range, the risk assessment module initiates an adaptive adjustment program, adjusts the weights of each risk factor according to a preset environmental parameter and weight adjustment mapping table, and recalculates the risk level determination threshold.

[0023] The advantages of the present invention are:

[0024] 1. This invention employs an improved convolutional neural network model and adds an attention mechanism module in the feature extraction module. After normalizing the image data, it extracts the driver's facial expressions, body postures, and driving environment features, achieving accurate extraction of driver behavior features. This improves the efficiency and accuracy of feature extraction and lays a reliable foundation for subsequent behavior analysis.

[0025] 2. This invention constructs a behavior pattern database by using a behavior analysis module to store data on fatigued driving, distracted driving, and illegal operation behavior patterns. It uses an optimized version of the dynamic time warping algorithm to calculate the similarity between feature vectors and each behavior pattern and compares them with a threshold, thereby achieving accurate identification of high-risk behaviors of drivers and enabling timely judgment of whether drivers are fatigued, distracted, or engaging in other dangerous behaviors.

[0026] 3. This invention establishes a risk assessment model based on a three-layer neural network architecture through a risk assessment module, determines the weights of each high-risk behavioral factor by combining the analytic hierarchy process, and outputs the probability distribution of risk levels through weighted calculation and the Softmax function. This enables a scientific determination of the risk level of driver behavior, accurately identifies the risk level, and provides a basis for early warning. Attached Figure Description

[0027] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application.

[0028] In the attached diagram:

[0029] Figure 1 This is a system framework diagram of the computer vision-based dynamic recognition system for high-risk driver behaviors in Example 1. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0031] The present invention will now be described in detail and specifically through specific embodiments to enable a better understanding of the invention. However, the following embodiments do not limit the scope of protection of the present invention.

[0032] Example 1

[0033] A computer vision-based dynamic recognition system for high-risk driver behaviors includes an image acquisition module, a feature extraction module, a behavior analysis module, a risk assessment module, and a warning output module. These modules are connected via a high-speed data bus. The image acquisition module is connected to the feature extraction module, transmitting the acquired image data to it. The feature extraction module processes the image data and then transmits it to the behavior analysis module. The behavior analysis module transmits its analysis results to the risk assessment module. The risk assessment module determines the risk level based on the analysis results and transmits the determination result to the warning output module.

[0034] In a specific embodiment, the image acquisition module employs three Basler acA2000-50gm CMOS cameras. The first camera is fixed to the center console directly in front of the steering wheel, while the second and third cameras are mounted above the driver's left and right doors, respectively. Each camera is equipped with an ON Semiconductor AR0234 image sensor, acquiring images at 50fps and 1920×1080 resolution, and transmitting them to the feature extraction module via a USB 3.1 Gen2 high-speed data interface. The feature extraction module utilizes the NVIDIA Jetson AGX Orin computing platform and is connected to the behavior analysis module, risk assessment module, and warning output module via a PCIe 4.0 bus to achieve real-time data transmission.

[0035] Furthermore, the image acquisition module includes no fewer than three cameras: the first camera is positioned directly in front of the steering wheel, the second camera is positioned above the driver's left door, and the third camera is positioned above the driver's right door. Each camera uses a CMOS image sensor of the same specifications to acquire image data at a constant frame rate and resolution, and transmits the image data to the feature extraction module through a high-speed data interface.

[0036] In this specific embodiment, all three cameras use Sony IMX377 CMOS image sensors. The first camera is perpendicular to the steering wheel's central axis with a 12° pitch angle; the second camera is mounted above the left door trim panel with a 40° horizontal tilt angle and an 18° pitch angle; and the third camera is mounted above the right door trim panel with a -40° horizontal tilt angle and an 18° pitch angle. Each camera transmits images via a 4-channel CSI-2 interface, achieving a total transmission rate of 11.6GB / s, ensuring that 1920×1080@50fps images are delivered to the feature extraction module without delay.

[0037] Furthermore, the feature extraction module employs an improved convolutional neural network model, adding an attention mechanism module to the original network structure. The convolutional layer calculation formula is as follows:

[0038]

[0039] in, For the j-th feature map in the l-th layer, M j Given the set of input feature maps, For the i-th feature map of the (l-1)th layer, The convolution kernel in the l-th layer connects the feature maps i and j. Here, f is the bias term, and f is the activation function;

[0040] The feature extraction module normalizes the image data and extracts the driver's facial expression features, body posture features, and driving environment features sequentially through convolutional layers, pooling layers, and attention mechanism modules. After compressing the dimensions of the extracted features, they are transmitted to the behavior analysis module in the form of feature vectors with pre-defined dimensions.

[0041] In a specific embodiment, based on the improved ResNet-50 network architecture, a CBAM attention mechanism module is embedded after each bottleneck block. In the convolutional layer computation, the input layer l-1 is a 32×32×256 feature map set M. j The j-th feature map in layer l is generated by a 3×3 convolution kernel K. l ij and bias term b lConvolution is performed with j=0.12, where the weight matrix is ​​[[0.15,0.2,-0.1],[0.3,0.45,0.2],[-0.15,0.1,0.35]]), and the activation function is Leaky ReLU. After normalization, the image data is sequentially passed through a 5-layer convolution-max pooling-attention module to compress the driver's facial expressions, body postures, and other features into a 256-dimensional feature vector, which is then transmitted to the behavior analysis module.

[0042] Furthermore, the behavior pattern database constructed by the behavior analysis module is stored in a non-volatile storage medium using an embedded database management system; the database stores data on fatigue driving behavior patterns, distracted driving behavior patterns, and violation operation behavior patterns; the behavior analysis module employs an optimized version of the dynamic time warping algorithm, and its formula for calculating the similarity between the feature vector and each behavior pattern is as follows:

[0043] D(i,j)=d(x i ,y j )+min{D(i-1,j),D(i,j-1),D(i-1,j-1)}

[0044] Where D(i,j) represents the cumulative distance between the i-th and j-th feature points, d(x i ,y j This represents the distance between two feature points calculated using a distance metric algorithm.

[0045] The behavior analysis module compares the calculated similarity with a preset threshold. When the similarity is greater than the threshold, it is determined that the match is successful, and the matching result is transmitted to the risk assessment module in binary code form.

[0046] In a specific embodiment, the behavior pattern database uses an embedded SQLite 3.39.4 database, stored in a Kingston UV500 240GB SSD, and includes three types of behavior patterns: fatigued driving, such as eye closure duration >180ms and blinking frequency <5 times / minute; distracted driving, such as head deflection angle 228° and duration >1.2s; and illegal operation, such as hands off the steering wheel >2.5s. In the optimized DTW algorithm, the input feature vector X = [0.78, 0.85, 0.62, 0.91, 0.73], behavior pattern Y = [0.75, 0.88, 0.65, 0.89, 0.71], when calculating the cumulative distance D(i,j), The recursive formula is: D(5,5)=d(x5,y5)+min{D(4,5),D(5,4),D(4,4)} where D(4,4)=0.12, and finally D(5,5)=0.02+0.12=0.14. The matching threshold is set to 0.2. When the similarity 1-0.14=0.86>the threshold, the binary code is output to the risk assessment module.

[0047] Furthermore, the risk assessment model established by the risk assessment module is based on a three-layer neural network architecture, including an input layer, a hidden layer, and an output layer. The number of nodes in the input layer corresponds to the number of behavior categories output by the behavior analysis module, the number of nodes in the hidden layer is the integer part of the arithmetic mean of the number of nodes in the input layer and the number of nodes in the output layer, and the number of nodes in the output layer is three. The weights of each high-risk behavior factor are determined using the analytic hierarchy process (AHP), and the calculation formula is as follows:

[0048]

[0049] Among them, W i Let a be the weight of the i-th factor. ij To determine the elements of the matrix, the judgment matrix is ​​constructed based on expert experience and historical data, where n is the number of factors;

[0050] After receiving the results from the behavior analysis module, the risk assessment module performs a weighted calculation on the behavior category code and its corresponding weight, outputs the risk level probability distribution through the Softmax function, takes the level with the highest probability as the final judgment result, and transmits it to the early warning output module in the form of an ASCII string.

[0051] In a specific embodiment, in the three-layer neural network architecture, the input layer has 3 nodes, the hidden layer has (3+3) / 2 = 3 nodes, and the output layer has 3 nodes. The risk levels are low, medium, and high. A 5×5 judgment matrix A is constructed using the analytic hierarchy process, including assessments of fatigued driving, distracted driving, and violations of operating procedures.

[0052]

[0053] The weight W1 is calculated as follows: W1 = (1 + 1 / 3 + 1 / 5) / (1 + 1 / 3 + 1 / 5 + 3 + 1 + 1 / 3 + 5 + 3 + 1) = (23 / 15) / (64 / 5) = 23 / 192 ≈ 0.12. Distracted driving, weight 0.35, is input as the behavior code, and after weighted calculation, it is input into the Softmax function.

[0054]

[0055] Assuming the input vector is [1.2, 2.5, 0.8] and the output probability distribution is [0.13, 0.72, 0.15], it is determined to be of medium risk and transmitted to the early warning module as an ASCII string.

[0056] Furthermore, the warning output module includes an audible and visual alarm unit, a variable frequency sound alarm device, and an information push unit; the audible and visual alarm unit uses a programmable LED light source and switches between different colors of light through pulse width modulation technology; the variable frequency sound alarm device outputs different frequency prompt sounds by changing the frequency of the drive signal; the information push unit integrates a wireless communication module, uses the TCP / IP protocol to send risk level information to the vehicle central control system terminal in JSON format, and sends brief risk level information to the driver's mobile terminal through the short message service protocol.

[0057] In a specific embodiment, the audible and visual alarm unit uses a WS2812B programmable LED light, which switches the light using PWM technology: red with an 80% duty cycle for high risk and yellow with a 60% duty cycle for medium risk. The variable frequency audible alarm device uses an HMB12A05 buzzer, which outputs a 750Hz square wave signal for low risk and a 1300Hz square wave signal for high risk. The information push unit integrates a SIM7600CE wireless module, which sends JSON data to the vehicle's central control system via TCP / IP protocol, and simultaneously sends an "Emergency Warning: Current High-Risk Distracted Driving Status" message to the driver's mobile phone via SMS protocol.

[0058] Furthermore, the feature extraction module adopts a multi-core computing platform; through the CUDA parallel computing programming model, the convolutional and pooling layer operations in the improved convolutional neural network model are mapped to thread blocks and thread grids for parallel processing; CUDA streaming technology is used to distribute data preprocessing, feature extraction, and result post-processing operations to different streams to achieve multi-task pipeline processing; frequently accessed data is stored through shared memory and access paths are optimized using memory merging technology; and double buffering technology is used to achieve overlapping execution of the current data block computation and the next data block loading.

[0059] In this specific embodiment, an NVIDIA A100 GPU is used. The convolutional layers are divided into 64×64 thread blocks using the CUDA 11.8 programming model, with each grid containing 16×16 thread blocks. Five CUDA streams are used for processing: stream 0 handles image data preprocessing, streams 1-3 handle different stages of convolution-pooling-attention computation, and stream 4 handles feature vector post-processing. Shared memory is set to 96KB to store convolutional kernel parameters, and access paths are optimized using memory merging techniques. Double buffering is employed, using two 256MB buffers, bufferA and bufferB. While the GPU processes data in bufferA, the CPU simultaneously loads the next batch of data into bufferB, achieving overlapping computation and loading, improving processing efficiency to 200 FPS.

[0060] Furthermore, the behavior analysis module includes an online learning unit that employs an incremental learning algorithm. When the system runs continuously for a preset duration and collects a preset number of new unmatched image data, the online learning mechanism is triggered. The new data is labeled, and the model parameters in the behavior pattern database are updated using stochastic gradient descent. Each update selects a preset number of image data as a training batch, sets a preset learning rate, and recalculates the similarity threshold of each behavior pattern in the behavior pattern database after the update is completed.

[0061] In a specific embodiment, the online learning mechanism is triggered when the system runs continuously for 6 hours and collects 300 frames of unmatched image data. New data is manually labeled using the LabelImg tool, categorized as fatigued driving, distracted driving, traffic violations, and normal driving. The model parameters are updated using stochastic gradient descent with a learning rate of 0.005, a training batch of 64 frames, and 200 iterations. After the update, the similarity threshold is recalculated; for example, the fatigued driving threshold is adjusted from 0.75 to 0.78. Cross-validation is used to ensure an improvement in model accuracy of ≥3%.

[0062] Furthermore, the risk assessment module includes an environmental sensor data acquisition unit, comprising a temperature sensor, a humidity sensor, a raindrop sensor, and a light sensor; when any environmental parameter detected by any environmental sensor exceeds a preset range, the risk assessment module initiates an adaptive adjustment program, adjusts the weights of each risk factor according to a preset environmental parameter and weight adjustment mapping table, and recalculates the risk level determination threshold.

[0063] In a specific embodiment, the environmental sensor data acquisition unit includes: a DHT22 temperature sensor with a measurement range of -40℃ to 80℃ and an accuracy of ±0.5℃, installed near the air conditioning vent on the center console; a HIH-4000 humidity sensor with a measurement range of 0-100%RH and an accuracy of ±3%RH, installed inside the instrument panel; a raindrop sensor (model RY802) that detects rainfall ≥0.5mm / min, attached to the inside of the windshield; and a BH1750 light sensor with a measurement range of 1-65535lx and an accuracy of ±20%, installed on the rearview mirror base inside the vehicle. When the DHT22 detects a temperature of 42℃, exceeding the preset upper limit of 40℃, the risk assessment module, based on the pre-stored mapping table, increases the distracted driving weight by 5% for every 1℃ increase in temperature, adjusting the distracted driving weight from 0.3 to 0.36, and recalculating the risk level judgment thresholds: the original high-risk threshold of 0.7 is adjusted to 0.67, and the medium-risk threshold of 0.4-0.7 is adjusted to 0.35-0.67, ensuring the accuracy of risk assessment in high-temperature environments.

[0064] The specific embodiments of the present invention have been described in detail above, but they are merely examples, and the present invention is not equivalent to the specific embodiments described above. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, all equivalent transformations and modifications made without departing from the spirit and scope of the present invention should be covered within the scope of the present invention.

Claims

1. A computer vision-based dynamic recognition system for high-risk driver behaviors, characterized in that, It includes an image acquisition module, a feature extraction module, a behavior analysis module, a risk assessment module, and an early warning output module; the image acquisition module, feature extraction module, behavior analysis module, risk assessment module, and early warning output module are connected via a high-speed data bus; the image acquisition module is connected to the feature extraction module, transmitting the acquired image data to the feature extraction module; The feature extraction module processes the image data and transmits it to the behavior analysis module; the behavior analysis module transmits the analysis results to the risk assessment module; the risk assessment module determines the risk level based on the analysis results and transmits the determination results to the early warning output module.

2. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 1, characterized in that, The image acquisition module includes no fewer than three cameras: the first camera is positioned directly in front of the steering wheel, the second camera is positioned above the driver's left door, and the third camera is positioned above the driver's right door. Each camera uses a CMOS image sensor of the same specifications to acquire image data at a constant frame rate and resolution, and transmits the image data to the feature extraction module through a high-speed data interface.

3. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 2, characterized in that, The feature extraction module employs an improved convolutional neural network model, adding an attention mechanism module to the original network structure. The convolutional layer calculation formula is as follows: in, For the j-th feature map in the l-th layer, M j Given the set of input feature maps, For the i-th feature map of the (l-1)th layer, The convolution kernel in the l-th layer connects the feature maps i and j. Here, f is the bias term, and f is the activation function; The feature extraction module normalizes the image data and extracts the driver's facial expression features, body posture features, and driving environment features sequentially through convolutional layers, pooling layers, and attention mechanism modules. After compressing the dimensions of the extracted features, they are transmitted to the behavior analysis module in the form of feature vectors with pre-defined dimensions.

4. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 3, characterized in that, The behavior pattern database constructed by the behavior analysis module is stored in a non-volatile storage medium and uses an embedded database management system. The database stores data on fatigued driving behavior patterns, distracted driving behavior patterns, and illegal operation behavior patterns. The behavior analysis module uses an optimized version of the dynamic time warping algorithm, and its formula for calculating the similarity between the feature vector and each behavior pattern is as follows: D(i,j)=d(x i ,y j )+min{D(i-1,j),D(i,j-1),D(i-1,j-1)} Where D(i,j) represents the cumulative distance between the i-th and j-th feature points, d(x i ,y j This represents the distance between two feature points calculated using a distance metric algorithm. The behavior analysis module compares the calculated similarity with a preset threshold. When the similarity is greater than the threshold, it is determined that the match is successful, and the matching result is transmitted to the risk assessment module in binary code form.

5. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 4, characterized in that, The risk assessment module establishes a risk assessment model based on a three-layer neural network architecture, including an input layer, a hidden layer, and an output layer. The number of nodes in the input layer corresponds to the number of behavior categories output by the behavior analysis module. The number of nodes in the hidden layer is the integer part of the arithmetic mean of the number of nodes in the input layer and the number of nodes in the output layer. The number of nodes in the output layer is three. The weights of each high-risk behavior factor are determined using the analytic hierarchy process (AHP), and the calculation formula is as follows: Among them, W i Let a be the weight of the i-th factor. ij To determine the elements of the matrix, the judgment matrix is ​​constructed based on expert experience and historical data, where n is the number of factors; After receiving the results from the behavior analysis module, the risk assessment module performs a weighted calculation on the behavior category code and its corresponding weight, outputs the risk level probability distribution through the Softmax function, takes the level with the highest probability as the final judgment result, and transmits it to the early warning output module in the form of an ASCII string.

6. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 5, characterized in that, The warning output module includes an audible and visual alarm unit, a variable frequency sound alarm device, and an information push unit. The audible and visual alarm unit uses a programmable LED light source and switches between different colors of light using pulse width modulation technology. The variable frequency sound alarm device outputs different frequency prompt sounds by changing the frequency of the drive signal. The information push unit integrates a wireless communication module, uses the TCP / IP protocol to send risk level information in JSON format data packets to the vehicle central control system terminal, and sends brief risk level information to the driver's mobile terminal via a short message service protocol.

7. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 6, characterized in that, The feature extraction module adopts a multi-core computing platform; the convolutional and pooling layer operations in the improved convolutional neural network model are mapped to thread blocks and thread grids for parallel processing through the CUDA parallel computing programming model; CUDA streaming technology is used to distribute data preprocessing, feature extraction, and result post-processing operations to different streams to achieve multi-task pipeline processing; frequently accessed data is stored through shared memory and access paths are optimized using memory merging technology; and double buffering technology is used to achieve overlapping execution of the current data block computation and the next data block loading.

8. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 7, characterized in that, The behavior analysis module includes an online learning unit that uses an incremental learning algorithm. When the system runs continuously for a preset duration and collects a preset number of new unmatched image data, the online learning mechanism is triggered. New data is labeled, and the model parameters in the behavior pattern database are updated using stochastic gradient descent. Each update selects a pre-set number of image data as a training batch and sets a pre-set learning rate. After the update is completed, the similarity threshold of each behavior pattern in the behavior pattern database is recalculated.

9. The computer vision-based dynamic recognition system for high-risk driver behavior according to claim 8, characterized in that, The risk assessment module includes an environmental sensor data acquisition unit, comprising a temperature sensor, a humidity sensor, a raindrop sensor, and a light sensor. When any environmental parameter detected by any environmental sensor exceeds a preset range, the risk assessment module initiates an adaptive adjustment program, adjusts the weights of each risk factor according to a preset environmental parameter and weight adjustment mapping table, and recalculates the risk level determination threshold.