A driver behavior rapid identification method

CN117953474BActive Publication Date: 2026-08-18XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410058211.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2026-08-18
Estimated Expiration
2044-01-15

AI Technical Summary

Technical Problem

CPU在神经网络实现中面临着计算效率低、功耗高、体积大、实时性受限等多重缺点

Benefits of technology

[0040]This invention utilizes a target recognition algorithm to identify driver behavior, avoiding the need for medical equipment to collect drivers' physiological indicators and thus preventing potential negative impacts on drivers. Furthermore, because the recognition algorithm achieves a more accurate and comprehensive assessment of driver behavior classification, capturing subtle movements and details, it provides a more objective evaluation of driver behavior compared to traditional physiological indicator detection. Moreover, compared to CPU or GPU-based deep learning network frameworks, FPGAs possess hardware-level parallel computing capabilities. Therefore, this invention utilizes a processing system (PS) for flow control of the recognition algorithm and programmable logic (PL) for accelerating the computation of recognition algorithm data, achieving more efficient processing and improving computational efficiency, greatly meeting the real-time requirements of automotive applications. Additionally, due to the low-power design of FPGAs, this invention effectively reduces the overall system's energy consumption, thus helping to improve the vehicle system's endurance and meeting the high energy management standards of the automotive environment. Furthermore, the miniaturized design of FPGAs makes them more suitable for limited vehicle space, overcoming the size inconvenience of traditional GPUs, making them easier to deploy and more widely applicable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117953474B_ABST
    Figure CN117953474B_ABST
Patent Text Reader

Abstract

The application discloses a driver behavior quick recognition method, which comprises the following steps: sending a to-be-recognized image to a PS end; processing the to-be-recognized image by the PS end using a trained recognition network and combining prior frame information of the trained recognition network, and sending a convolution processing task or a pooling processing task encountered currently to a PL end to assist the PL end in processing during the processing; executing the processing task sent by the PS end by the PL end and sending a calculation result obtained by the PL end to the PS end; continuing the processing by the PS end according to the calculation result, continuously sending a convolution processing task or a pooling processing task newly encountered to the PL end to assist the PL end in processing, and obtaining a three-dimensional processing result of the to-be-recognized image by the PS end; and performing non-maximum suppression processing on the three-dimensional processing result after decoding processing to obtain a recognition result of the to-be-recognized image. The application has the advantages of fast recognition speed, easy deployment, wide application range and low power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent recognition technology, specifically relating to a method for rapid recognition of driver behavior. Background Technology

[0002] With the continuous development of artificial intelligence and FPGA image processing technologies, unprecedented conditions have been provided for real-time portrait processing, leading to a surge of image processing applications based on neural network algorithms in military, commercial, and medical fields. These applications can quickly remove redundant information from images and efficiently process target recognition through image extraction technology, effectively improving the performance of machine vision. Modern electronic information technology is gradually integrating into the field of intelligent driving, playing a crucial role in upgrading traditional driving services. Compared to single-chip microcomputer embedded systems, FPGAs offer superior performance in terms of response speed and power consumption.

[0003] Currently, mainstream image recognition methods primarily rely on artificial neural network technology. Researchers have invested significant time and effort in improving and optimizing algorithm performance. Today, compared to traditional image recognition techniques, neural networks offer significant advantages in speed and accuracy; while the parallel processing capabilities of FPGAs enable pipelined computation, providing faster response times for image processing applications—a feature lacking in traditional microprocessors. Therefore, FPGAs are often chosen as processing chips in devices with high real-time requirements. With the increasing availability of internal hardware resources within FPGAs, the advantages of using FPGAs for image processing technology development are becoming increasingly apparent.

[0004] Driver behavior detection schemes based on physiological indicators (such as electroencephalogram signals) can monitor the driver's driving status in real time. By comparing the driver's actual physiological data with physiological data obtained under normal driving conditions, it can effectively reflect whether the driver is experiencing tension, anxiety, or fatigue.

[0005] The current image recognition-based solution integrates theories from computer vision and artificial intelligence. This technology relies on high-efficiency image sensors to accurately extract multi-dimensional driver status information, including head movements, facial features, expression changes, and other details related to behavior and human characteristics. This method achieves contactless acquisition of driver feature information. This technology has a significant advantage in its high sensitivity to subtle movements and micro-expressions, capturing subtle changes in driver behavior and thus revealing a more comprehensive picture of their current driving state. By comprehensively analyzing this feature information, the system can accurately determine the driver's psychological and physiological state in real time, providing a deeper insight into driving safety.

[0006] However, detection schemes based on physiological indicators require the assistance of medical devices to collect the driver's physiological information, which increases cost and operational complexity. Secondly, this method interferes with the driver's actual driving process, as the driver may react differently even when aware that physiological indicators are being detected, thus affecting the correctness and accuracy of the detection results. Although image recognition-based methods are contactless, current convolutional neural network-based driver abnormal behavior detection schemes rely on CPU or GPU hardware platforms. CPUs face multiple drawbacks in neural network implementations, including low computational efficiency, high power consumption, large size, and limited real-time performance. This may particularly limit performance in automotive applications. While GPUs excel in parallel computing, they also suffer from high power consumption, large size, and relatively high cost. In the vehicle environment, especially for small in-vehicle devices, these issues may limit the widespread use of GPUs in automotive applications. Considering the requirements of automotive applications for small size, low power consumption, and high real-time performance, traditional CPUs and GPUs struggle to meet the stringent computational performance requirements of vehicle intelligent systems. In other words, there is currently a lack of a fast, easily deployable, widely applicable, and low-power-consumption-required method for rapid driver behavior recognition. Summary of the Invention

[0007] To address the aforementioned problems in the existing technology, the present invention provides a method for rapid driver behavior recognition.

[0008] The technical problem to be solved by this invention is achieved through the following technical solution:

[0009] A method for rapid driver behavior recognition includes:

[0010] Acquire the image to be identified and send the image to the PS terminal;

[0011] The PS end uses a trained recognition network and combines the prior bounding box information of the trained recognition network to process the image to be recognized. During the processing, the currently encountered convolution processing task or pooling processing task is sent to the PL end so that the PL end can assist in the processing.

[0012] The PL terminal executes the processing task sent by the PS terminal and sends the obtained calculation result to the PS terminal;

[0013] The PS end continues to process the calculation results and sends newly encountered convolution or pooling tasks to the PL end for the PL end to assist in the processing until the PS end obtains the three-dimensional processing result of the image to be recognized.

[0014] The PS terminal decodes the 3D processing result and then performs non-maximum suppression processing to obtain the recognition result of the image to be recognized.

[0015] In some embodiments, the step of the PS end sending the currently encountered convolution processing task to the PL end includes:

[0016] The PS terminal divides the current input feature map and the convolution kernel used to perform convolution processing on the current input feature map into blocks, and sends the resulting input feature map blocks and convolution kernel blocks to the PL terminal in real time; wherein, the current input feature map refers to the feature map that needs to be convolved.

[0017] In some embodiments, the PS terminal divides the current input feature map and the convolution kernel used to perform convolution processing on the current input feature map into blocks, and sends the resulting input feature map blocks and convolution kernel blocks to the PL terminal in real time, including:

[0018] The PS terminal divides the number of channels of each convolution kernel used for convolution processing of the current input feature map according to the first preset division parameters, with the number of channels as the block dimension, and sends the divided convolution kernel blocks to the PL terminal in real time.

[0019] The PS end divides the current output feature map according to the second preset division parameters and the current output feature map, and determines the size of each output feature map block obtained when dividing the current output feature map into blocks; the current output feature map refers to the feature map obtained after convolution processing of the current input feature map using a convolution kernel used to convolve the current input feature map;

[0020] The PS terminal determines the size of each input feature map block obtained when dividing the current input feature map into blocks based on the width and height of each output feature map block, as well as the size, stride, and padding of each convolutional kernel block.

[0021] The PS terminal divides the current input feature map into blocks according to the size of each divided input feature map block, and sends the divided input feature map blocks to the PL terminal in real time.

[0022] In some embodiments, the PL terminal executes the processing task sent by the PS terminal and sends the obtained calculation result to the PS terminal, including:

[0023] The PL terminal receives the convolution kernel block and input feature map block sent by the PS terminal in real time, and uses the received convolution kernel block to perform convolution processing on the received input feature map block, and then performs accumulation and concatenation to obtain the corresponding output feature map block, and sends the obtained output feature map block to the PS terminal in real time.

[0024] In some embodiments, the PS end continues processing based on the calculation result, and sends newly encountered convolution processing tasks or pooling processing tasks to the PL end for the PL end to assist in processing, until the PS end obtains the three-dimensional processing result of the image to be recognized, including:

[0025] The PS end will stitch together the output feature map blocks received in real time from the PL end until the current output feature map is obtained.

[0026] When the PS terminal needs to perform convolution or pooling operations on the current output feature map, the PS terminal generates a new convolution or pooling task based on the current output feature map and continues to send it to the PL terminal so that the PL terminal can assist in the processing until the PS terminal obtains the three-dimensional processing result of the image to be recognized.

[0027] When the PS terminal needs to perform operations other than convolution and pooling on the current output feature map, the PS terminal performs the operation on the current output feature map. When it needs to perform convolution or pooling on the result of the operation, the PS terminal generates a new convolution or pooling task based on the result of the operation and sends it to the PL terminal to assist in the processing until the PS terminal obtains the three-dimensional processing result of the image to be recognized.

[0028] In some embodiments, the step of the PS end sending the currently encountered pooling processing task to the PL end includes:

[0029] The PS terminal will divide the current input feature map into blocks and send the pooling kernel used to pool the current input feature map, as well as the resulting input feature map blocks, to the PL terminal in real time; wherein, the current input feature map refers to the feature map that needs to be pooled.

[0030] Accordingly, the PL terminal executes the processing task sent by the PS terminal and sends the obtained calculation result to the PS terminal, including:

[0031] The PL terminal receives the input feature map blocks and pooling kernels sent by the PS terminal in real time, and performs pooling processing on each received input feature map block using the received pooling kernel, and then sends the pooling result to the PS terminal in real time.

[0032] In some embodiments, the PS end continues processing based on the calculation result, and sends newly encountered convolution processing tasks or pooling processing tasks to the PL end for the PL end to assist in processing, until the PS end obtains the three-dimensional processing result of the image to be recognized, including:

[0033] The PS end will concatenate the pooling results received in real time from the PL end until the complete pooling result of the current input feature map is obtained.

[0034] When the PS terminal needs to perform convolution or pooling operations on the complete pooling result of the current input feature map, the PS terminal generates a new convolution or pooling processing task based on the complete pooling result of the current input feature map, and continues to send it to the PL terminal so that the PL terminal can assist in the processing, until the PS terminal obtains the three-dimensional processing result of the image to be recognized.

[0035] When the PS terminal needs to perform operations other than convolution and pooling on the complete pooling result of the current input feature map, the PS terminal performs operations on the complete pooling result of the current input feature map. When it needs to perform convolution or pooling operations on the processed result, the PS terminal generates a new convolution or pooling processing task based on the processed result and the corresponding pooling kernel, and continues to send it to the PL terminal for the PL terminal to assist in processing, until the PS terminal obtains the three-dimensional processing result of the image to be recognized.

[0036] In some embodiments, the three-dimensional processing result includes: a first three-dimensional processing result and a second three-dimensional processing result; wherein, the first three-dimensional processing result is used to detect large targets, and the second three-dimensional processing result is used to detect small targets; both the first three-dimensional processing result and the second three-dimensional processing result contain K×(5+C), where K represents the number of prior boxes of the trained recognition network, 5 represents that each bounding box has five pieces of information: the x-coordinate of the bounding box center point, the y-coordinate of the bounding box center point, the width of the bounding box, the height of the bounding box, and the confidence level of the bounding box, and C represents the number of classification categories.

[0037] In some embodiments, the first three-dimensional processing result is (13,13,K×(5+C)), and the second three-dimensional processing result is (26,26,K×(5+C)); wherein, the two 13s in (13,13,K×(5+C)) represent dividing the image into 13×13 grid units, and the two 26s in (26,26,K×(5+C)) represent dividing the image into 26×26 grid units.

[0038] In some embodiments, the trained recognition network is a trained Tiny YOLOv4 network, and the recognition result includes: the driver's position in the image to be recognized, and the driver's behavior category.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] This invention utilizes a target recognition algorithm to identify driver behavior, avoiding the need for medical equipment to collect drivers' physiological indicators and thus preventing potential negative impacts on drivers. Furthermore, because the recognition algorithm achieves a more accurate and comprehensive assessment of driver behavior classification, capturing subtle movements and details, it provides a more objective evaluation of driver behavior compared to traditional physiological indicator detection. Moreover, compared to CPU or GPU-based deep learning network frameworks, FPGAs possess hardware-level parallel computing capabilities. Therefore, this invention utilizes a processing system (PS) for flow control of the recognition algorithm and programmable logic (PL) for accelerating the computation of recognition algorithm data, achieving more efficient processing and improving computational efficiency, greatly meeting the real-time requirements of automotive applications. Additionally, due to the low-power design of FPGAs, this invention effectively reduces the overall system's energy consumption, thus helping to improve the vehicle system's endurance and meeting the high energy management standards of the automotive environment. Furthermore, the miniaturized design of FPGAs makes them more suitable for limited vehicle space, overcoming the size inconvenience of traditional GPUs, making them easier to deploy and more widely applicable.

[0041] The present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0042] Figure 1 This is a flowchart illustrating a method for rapid driver behavior recognition provided in an embodiment of the present invention;

[0043] Figure 2 This is an overall block diagram of an exemplary system for performing the rapid driver behavior recognition method provided in an embodiment of the present invention;

[0044] Figure 3 This is a schematic diagram of an exemplary network structure of Tiny YOLOv4 provided in an embodiment of the present invention;

[0045] Figure 4 This is an exemplary classification diagram of driver behavior categories provided in an embodiment of the present invention;

[0046] Figure 5 This is an exemplary flowchart for rapidly identifying driver behavior provided by an embodiment of the present invention.

[0047] Figure 6 This is a schematic diagram of an exemplary PS and PL end jointly performing split convolution processing according to an embodiment of the present invention;

[0048] Figure 7 This is a schematic diagram of a joint implementation of convolution acceleration using the PS and PL ends provided in an embodiment of the present invention;

[0049] Figure 8 This is a schematic diagram of a joint pooling acceleration implementation provided by the PS and PL ends in an embodiment of the present invention. Detailed Implementation

[0050] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0051] Figure 1 This is a flowchart illustrating a method for rapid driver behavior recognition provided in an embodiment of the present invention. The method includes:

[0052] S101. Obtain the image to be recognized and send it to the PS terminal.

[0053] S102 and PS end use the trained recognition network and combine the prior bounding box information of the trained recognition network to process the image to be recognized. During the processing, the currently encountered convolution processing task or pooling processing task is sent to PL end so that PL end can assist in the processing.

[0054] S103, the PL terminal executes the processing task sent by the PS terminal and sends the obtained calculation results to the PS terminal.

[0055] S104. The PS end continues to process the data based on the calculation results, and sends any new convolution or pooling tasks to the PL end for assistance, until the PS end obtains the 3D processing result of the image to be recognized.

[0056] The S105 and PS terminals decode the 3D processing results and then perform non-maximum suppression processing to obtain the recognition result of the image to be recognized.

[0057] For example, Figure 2 This is an overall block diagram of the system used to implement this rapid driver behavior recognition method. (See diagram below.) Figure 2As shown, the system includes a camera, a display, and an AX7350 development board. The AX7350 development board has a PS (Power Positioner) and a PL (Power Positioner) terminal, which transmit data via the AXIDMA protocol. The PL terminal includes a convolution acceleration module and a pooling acceleration module. When the camera captures an image to be recognized, it sends the image to the PS terminal. The PS terminal uses a trained recognition network and its prior bounding boxes to process the image. During processing, if a convolution task is encountered, the PS terminal sends the convolution task to the PL terminal in real-time using the AXIDMA protocol. The PL terminal uses its convolution acceleration module to process the convolution task and returns the processing result to the PS terminal in real-time using the AXIDMA protocol. For convolution processing tasks, the PS end uses the AXIDMA protocol to send the pooling processing task to the PL end in real time. The PL end uses the pooling acceleration module to process the pooling processing task and uses the AXIDMA protocol to return the processing result to the PS end in real time. The PS end performs subsequent processing based on the received processing result, and when it encounters a new convolution processing task or a new pooling processing task, it continues to send it to the PL end for accelerated processing until the PS end obtains the 3D processing result of the image. Then, the PS end decodes the 3D processing result of the image to convert the network output into the actual bounding box position and category information. After that, non-maximum suppression processing is performed to finally obtain the detection result (i.e., the recognition result) of the image, and displays the detection result on the display.

[0058] Here, the image to be recognized is an RGB image. The PS end converts the RGB image into an input feature map of size 416×416×3 (where 416 means dividing the image into 416×416 grid units, all colors can be obtained by using the three primary colors of RGB with corresponding weights, and 3 refers to the size of the three primary colors), and then performs subsequent processing on the input feature map.

[0059] In some embodiments, the trained recognition network can be employed as follows: Figure 2 The system shown is trained on the PS terminal. In other embodiments, the trained recognition network may also be pre-trained on other devices and then imported into the PS terminal. This invention does not limit this.

[0060] Here, the trained recognition network can be a trained Tiny YOLOv4 network; the prior box information can be the number of prior boxes and the size of each prior box, for example, the number of prior boxes is 6.

[0061] Tiny YOLOv4's single forward propagation design enables it to quickly and accurately perform target detection and classification tasks in driver behavior recognition. Furthermore, its global context understanding and multi-scale feature extraction enhance its adaptability to complex driving scenarios. Its bounding box regression technology achieves high-precision target localization, providing a reliable foundation for accurate driver behavior analysis. To prevent overfitting, this invention employs regularization methods and adjustments to the learning rate for optimization. The network structure of Tiny YOLOv4 is as follows: Figure 3 As shown, the main components include the backbone network, the Feature Pyramid Network (FPN), and the YOLO Head. The backbone network uses CSPDarknet53-tiny, which contains a series of convolutional layers, residual blocks, and CSP connections to learn image features. Its input image dimensions are Win×Hin×Cin(416, 416, 3). The FPN is used to introduce multi-scale feature information, mainly involving lateral connections and upsampling operations. In Tiny YOLOv4, the FPN connects the feature maps of the bottom and upper layers to form a pyramid structure. The dimensions of the bottom layer feature map are Wout1×Hout1×Cout1(13, 13, 128), while the dimensions of the upper layer feature map are Wout2×Hout2×Cout2(26, 26, 256). The dimensions of the concatenated feature map are Wout2×Hout2×(Cout1+Cout2)(26, 26, 128+256). The YOLO Head is the detection head of the model, responsible for generating object detection predictions. It processes the upsampled feature map through a series of convolutional operations to extract higher-level feature representations. The output size of the YOLO Head is 13×13×(K×(5+C)) and 26×26×(K×(5+C)), the former for detecting large objects and the latter for detecting small objects. K represents the number of prior boxes, and C represents the number of classification categories. Each grid cell contains K×(5+C) bounding boxes, and each bounding box predicts 5+C pieces of information, including: the bounding box's position information (b... x b y b h b w The bounding box's confidence score (conf) and the probability that the bounding box's target belongs to each of the C categories, b. x The x-coordinate of the center point of the bounding box, b y Indicates the ordinate of the center point of the bounding box, b h Indicates the bounding box width, b w Indicates the height of the bounding box

[0062] Specifically, the network structure of Tiny YOLOv4 is shown in Table 1.

[0063] Table 1. Tiny YOLOv4 Network Structure

[0064]

[0065]

[0066] In Table 1, kernel shape represents the size of the convolutional kernel or pooling kernel, Cout represents the number of output channels, Cin represents the number of input channels, H represents the height of the convolutional kernel, and W represents the width of the convolutional kernel; Stride indicates the step size; Padding indicates the number of rows and columns of zero-padding in the input feature map; Input Feature Map indicates the size of the input feature map, H represents the height of the feature map, W represents the width of the feature map, and C1 represents the number of channels in the feature map; Output Feature Map indicates the size of the output feature map. In Table 1, the number of output channels for headP4.conv2 and headP5.conv2 is K×(5+C)).

[0067] Tiny YOLOv4 draws inspiration from Faster R-CNN in its design, replacing fully connected layers with convolutional layers to obtain higher resolution output feature maps. To enable the network to locate target positions more quickly and accurately, it can automatically generate candidate box sizes suitable for target detection based on training data. Therefore, this invention uses K-Means clustering to initialize bounding box regression. The traditional K-Means algorithm uses Euclidean distance to measure the similarity between sample points and cluster centers. However, due to the influence of bounding box size, large-sized anchors have larger errors than small-sized anchors, making large bounding boxes more susceptible to errors and thus causing some deviation in the output results. To eliminate the influence of bounding box size on errors, Tiny YOLOv4 introduces the overlap ratio (IoU) between predicted and true bounding boxes as a metric to measure similarity during the clustering process. The IoU calculation process is shown in Equation (1). The metric function is shown in Equation (2) to achieve highly robust and reliable bounding box initialization, where, , A represents the ground truth box, B represents the predicted box; d(box,centroid)=1-IoU(box,centroid)(2), box represents the bounding box to be tested, centroid represents the center of the prior box, and the prior box is a set of boxes with different aspect ratios obtained by the Kmeans algorithm before training.

[0068] The Tiny YOLOv4 network predicts the offsets (Offset(b)) of the bounding box relative to the top-left corner of the corresponding grid cell at four predicted positions. x ,b y ,b w ,b h The bounding box position, width, height, and predicted value are calculated using the following formula (3): (the formula is missing from the original text.) σ represents the sigmoid function, t x t y It is the offset of the center point of the Anchor (prior bounding box) relative to the top-left corner of its grid cell, t w t h c represents the height and width of the bounding box of the object in the real box. x c y Let p be the center coordinates of the Anchor. w p h Let σ(t) represent the width and height of the anchor. o ) represents the confidence level, and pr(object) = 1 when there is an object (target) and pr(object) = 0 when there is no object. IoU(b,object) represents the overlap rate between the predicted box and the ground truth box.

[0069] To train the bounding box to closely approximate the size and location of the ground truth bounding box, a loss function is introduced that can be divided into three parts. The first is the error introduced by the class, i.e., the classification loss; the second is the error between the predicted bounding box and the ground truth bounding box, i.e., the localization loss introduced by the bounding box; and the third is the confidence loss due to confidence error. When a target is detected, the classification loss of the bounding box is the Cross-Entropy Loss of the conditional class probabilities for each class.

[0070] The classification loss function is shown in equation (4):

[0071]

[0072] Among them, S 2 B represents traversing all grid cells in the image, and B represents traversing all predicted bounding boxes. p indicates whether the current sample is a positive sample. i (c) represents the probability that the i-th grid cell belongs to category c. This represents the actual value of the category to which the labeled box belongs.

[0073] The localization loss prevents weighted absolute errors due to different dimensions by predicting the square roots of the width and height of the bounding box. The localization loss function is shown in Equation (5):

[0074]

[0075] To improve the accuracy of the bounding box, the loss is multiplied by λ. coord =2-w i ×h i w i h i It is the width and height of the prediction box, t x t y It is the offset of the center point of the anchor relative to the top-left corner of its grid cell, t w t h This represents the height and width of the bounding box of the target within the actual frame.

[0076] The confidence loss is divided into two cases: with and without a target. If a target is detected in the bounding box, the confidence loss function is as shown in formula (6):

[0077]

[0078] Among them, C i This represents the score probability that the predicted bounding box contains the target object. This represents the true value, i.e., the target value predicted by the i-th grid cell. otherwise

[0079] If no target is detected in the bounding box, the confidence loss function is as shown in Equation (7), and there is no classification or localization loss, Equation (7) is:

[0080]

[0081] Therefore, the total loss of Tiny YOLOv4 is as shown in Equation (8): L=L1+L2+L co +L cno (8).

[0082] Here, when training the Tiny YOLOv4 network, the StateFarm-distracted-driver-detection dataset can be selected. The images in this dataset are derived from in-vehicle video recordings, processed into 2D images with a size of 640×480, and these images have similar cropping angles. Since the environmental backgrounds of all images in the dataset are essentially the same, this may lead to overfitting during training. Therefore, in the preprocessing stage, this invention can employ randomization processes such as scaling and translation. By introducing these processes, the diversity of the data is increased, which can avoid potential overfitting during training. Specifically, 9000 images were extracted as the training set, and 1000 images were selected as the test set. Since this dataset does not provide annotation information, this invention uses the labelImg annotation tool for manual annotation, according to... Figure 4 The images in the dataset are classified into 10 classes (c1 to c9), and the target regions in the training dataset are manually labeled for object detection. Specifically, for example... Figure 4 As shown, from images c1 to c9, these 10 categories specifically represent: "normal driving," "texting-right," "talking on the phone-right," "texting-left," "talking on the phone-left," "operating the radio," "drinking water," "reaching behind," "covering hair and face," and "talking to passenger." To adapt to the Tiny YOLOv4 network, this invention annotated the target area according to its annotation format, ultimately generating a new annotation file.

[0083] In this invention, when the trained recognition network is a trained Tiny YOLOv4 network, the three-dimensional processing result of an image to be recognized is the first three-dimensional processing result (13,13,K×(5+C)) and the second three-dimensional processing result is (26,26,K×(5+C)). The first three-dimensional processing result is used to detect large targets, and the second three-dimensional processing result is used to detect small targets. In (13,13,K×(5+C)), the two 13s indicate that the image to be recognized is divided into 13×13 grid units, and the two 26s in (26,26,K×(5+C)) indicate that the image to be recognized is divided into 26×26 grid units. Accordingly, the recognition result of an image to be recognized is: the position of the driver in the image to be recognized, and the driver's behavior category.

[0084] For example, Figure 5 A flowchart for rapid driver behavior identification. (e.g.) Figure 5 As shown, the image to be recognized is input into the PS terminal. The PS terminal uses a pre-trained Tiny YOLOv4 network and combines the prior bounding box information of the pre-trained Tiny YOLOv4 network to process the image to be recognized, obtaining the network output results (13, 13, K×(5+C)) and (26, 26, K×(5+C)). Each grid cell in the network output results corresponds to K×(5+C) bounding boxes, and each bounding box predicts 5+C kinds of information: the position information of the bounding box (b... x ,b y b h b w The bounding box's confidence score (conf) and the probability that the bounding box's target belongs to each of the C categories, b. x The x-coordinate of the center point of the bounding box, b y Indicates the ordinate of the center point of the bounding box, b h Indicates the bounding box width, b w The height of the bounding box is represented. By decoding the network output, multiple bounding boxes and 5+C types of information for each bounding box can be obtained. Then, non-maximum suppression is performed on the bounding boxes with confidence scores greater than a threshold obtained from the decoded bounding boxes. Finally, a bounding box representing the driver's position in the image and the category of the driver's behavior contained in the bounding box are obtained, and the recognition process ends.

[0085] Because the feature maps of the Tiny YOLOv4 network have a large spatial dimension, the increased network depth and number of convolutional kernels result in a large amount of resources required for the complete feature maps and convolutional weight parameters of a single-layer network. Therefore, this invention splits the feature map data and convolutional kernels during the computation process into several data blocks, which are then sent to the PL end for computation. This avoids the problem that the complete feature maps and convolutional weight parameters of a single-layer network may not be able to be directly accommodated in the BRAM of the PL end.

[0086] In some embodiments, in S102 above, the step of the PS end sending the currently encountered convolution processing task to the PL end is as follows: the PS end divides the current input feature map and the convolution kernel used to perform convolution processing on the current input feature map into blocks, and sends the block-sized input feature map blocks and convolution kernel blocks to the PL end in real time; wherein, the current input feature map refers to the feature map that needs to be convolved.

[0087] Specifically, the PS end divides the number of channels of each convolutional kernel used for convolution processing of the current input feature map according to the first preset partitioning parameter, with the number of channels as the partitioning dimension, and sends the partitioned convolutional kernel blocks to the PL end in real time; and the PS end divides the current output feature map according to the second preset partitioning parameter and the current output feature map, and determines the size of each output feature map block obtained when the current output feature map is partitioned, where the current output feature map refers to the feature map obtained after convolution processing of the current input feature map using the convolutional kernel used for convolution processing of the current input feature map; then, the PS end determines the size of each input feature map block obtained when the current input feature map is partitioned according to the width and height of each output feature map block, as well as the size, stride and padding of each convolutional kernel block; the PS end can then partition the current input feature map according to the size of each input feature map block, and send the partitioned input feature map blocks to the PL end in real time.

[0088] Here, the first preset partitioning parameter can be the preset number of channels for each convolutional kernel block. Therefore, based on the preset number of channels for each convolutional kernel block and the number of channels for the convolutional kernel to be partitioned, each convolutional kernel to be partitioned can be divided into multiple convolutional kernel blocks. For each convolutional kernel used to perform convolution processing on the current input feature map, the number of convolutional kernel blocks obtained after partitioning each convolutional kernel is the same. The second preset partitioning parameter can be the preset number of output feature map blocks. When partitioning the current input feature map, the size of each output feature map block obtained when partitioning the current output feature map can be determined first based on the size of the current output feature map corresponding to the current input feature map and the second preset partitioning parameter. Then, based on the size of each output feature map block and the size, stride, and padding of each convolutional kernel block obtained from the above partitioning, the size and number of channels of each input feature map block obtained after partitioning the current input feature map can be deduced. Finally, the current input feature map can be partitioned based on the deduced size and number of channels of each input feature map block.

[0089] For example, when the size of an input feature map patch is Hin×Win×Cin, the size of the resulting convolutional block is K'×K'×Cout, the stride of the kernel block is S, the padding is P, and the size of an output feature map patch is Hout×Wout×Cout', then according to the convolution operation, we can obtain... By working backwards That is, the width Win of the input feature map block is (Wout-1)×S-2×P+K', the height Hin is (Hout-1)×S-2×P+K', and the number of channels is Cout.

[0090] In some embodiments, S103 specifically includes: the PL end receiving the convolution kernel block and the input feature map block sent by the PS end in real time, and using the received convolution kernel block to perform convolution processing on the received input feature map block, and then accumulating and splicing it to obtain the corresponding output feature map block, and sending the obtained output feature map block to the PS end in real time.

[0091] In practical applications, due to limitations in the loading capabilities of some PS terminals, it is impossible to load the current input feature map and each convolutional kernel used for convolution processing of the current input feature map all at once. Therefore, the PS terminal can load a portion of the input feature map and a portion of the convolutional kernel each time, and then divide the loaded portion of the input feature map and the portion of the convolutional kernel into blocks and send them to the PL terminal for processing. After that, it can load a portion of the input feature map and a portion of the convolutional kernel again, divide them into blocks, and continue to send them to the PL terminal for processing. At the same time, after receiving a portion of the input feature map block and a portion of the convolutional kernel block, the PL terminal uses the received convolutional kernel block to perform convolution processing on the corresponding received feature map block, and returns the resulting output feature map block to the PS terminal in real time.

[0092] In some embodiments, S104 specifically includes:

[0093] S1041. The PS end will stitch together the output feature map blocks received in real time from the PL end until the current output feature map is obtained.

[0094] S1042. When the PS terminal needs to perform convolution or pooling operations on the current output feature map, the PS terminal generates a new convolution or pooling task based on the current output feature map and continues to send it to the PL terminal to assist in the processing until the PS terminal obtains the three-dimensional processing result of the image to be recognized.

[0095] S1043. When the PS end needs to perform operations other than convolution and pooling on the current output feature map, the PS end performs the operation on the current output feature map. When it needs to perform convolution or pooling on the result of the operation, the PS end generates a new convolution or pooling task based on the result of the operation and continues to send it to the PL end for the PL end to assist in the processing until the PS end obtains the three-dimensional processing result of the image to be recognized.

[0096] For example, Figure 6 This is a schematic diagram illustrating the principle of joint split convolution processing by the PS and PL ends. Figure 6In the diagram, R and C represent the number of rows and columns of the output feature map, respectively. CHout represents the number of channels in the output feature map, which is the same as the number of convolutional kernels used to convolve the input feature map. K1 is the size of each of the CHout convolutional kernels, and the number of channels in each CHout convolutional kernel is the same as the number of channels in the input feature map. Tn is the number of channels in each convolutional kernel block obtained after dividing each of the CHout convolutional kernels into blocks, and the number of channels in each convolutional kernel block is the same as the number of channels in each input feature map block obtained after dividing the input feature map into blocks. Tr and Tc represent the number of rows and columns of each output feature map block calculated by the PL end, Tm represents the number of channels in each output feature map block calculated by the PL end, S represents the stride of each of the CHout convolutional kernels, and P represents the padding value of each of the CHout convolutional kernels. Figure 6 As shown, if the PS end currently segments an input feature map block of ((Tc-1)×S-2×P+K1)×((Tc-1)×S-2×P+K1)×Tn from the input feature map, and segments Tm K1×K1×Tn convolutional kernel blocks from CHout convolutional kernels and sends them to the PL end, the PL end uses Tm K1×K1×Tn convolutional kernel blocks to perform convolution processing on an input feature map block of ((Tc-1)×S-2×P+K1)×((Tc-1)×S-2×P+K1)×Tn, and obtains a Tr×Tc×Tm output feature map block in the R×C×CHout output feature map.

[0097] For example, Figure 7 This is a schematic diagram illustrating the principle of jointly accelerating convolution using the PS and PL ends. (Example) Figure 7 As shown, for the current input feature map and the convolution kernel used to perform convolution processing on the current input feature map, the PS end divides the input feature map and the convolution kernel into blocks, and then sends the block-based input feature map blocks and convolution kernel blocks to the PL end using the AXIDMA protocol. The PL end first buffers the received input feature map blocks and convolution kernel blocks, and then inputs the buffered input feature map blocks and convolution kernel blocks into its own convolution acceleration module. The convolution acceleration module outputs the calculated output feature map blocks. The PL end sends the output feature map blocks to the PS end using the AXIDMA protocol. The PS end accumulates and concatenates the received output feature map blocks to finally obtain the output feature map corresponding to the current input feature map.

[0098] In some embodiments, in S102 above, the step of the PS end sending the currently encountered pooling processing task to the PL end is as follows: the PS end divides the current input feature map into blocks, and sends the pooling kernel used to pool the current input feature map, as well as the input feature map blocks obtained from the block division, to the PL end in real time; wherein, the current input feature map refers to the feature map that needs to be pooled; correspondingly, S103 above specifically includes: the PL end receives the input feature map blocks and pooling kernel sent by the PS end in real time, and performs pooling processing on each received input feature map block using the received pooling kernel, and then sends the pooling result to the PS end in real time.

[0099] Here, pooling can be max pooling, average pooling, etc., and this invention does not limit it.

[0100] In some embodiments, S104 specifically includes:

[0101] S1044: The PS end will concatenate the pooling results received from the PL end in real time until the complete pooling result of the current input feature map is obtained.

[0102] S1045. When the PS end needs to perform convolution or pooling operations on the complete pooling result of the current input feature map, the PS end generates a new convolution or pooling processing task based on the complete pooling result of the current input feature map, and continues to send it to the PL end for the PL end to assist in processing, until the PS end obtains the three-dimensional processing result of the image to be recognized.

[0103] S1046. When the PS terminal needs to perform operations other than convolution and pooling on the complete pooling result of the current input feature map, the PS terminal performs operations on the complete pooling result of the current input feature map. When it needs to perform convolution or pooling operations on the processed result, the PS terminal generates a new convolution or pooling processing task based on the processed result and the corresponding pooling kernel, and continues to send it to the PL terminal for the PL terminal to assist in processing until the PS terminal obtains the three-dimensional processing result of the image to be recognized.

[0104] For example, Figure 8 This is a schematic diagram illustrating the principle of pooling acceleration achieved jointly by the PS and PL ends. (Example) Figure 8As shown, for the current input feature map, the PS end divides the input feature map into blocks and then sends the divided input feature map blocks and corresponding pooling kernels to the PL end using the AXIDMA protocol. The PL end first buffers the received input feature map blocks and pooling kernels, and then inputs the buffered input feature map blocks and pooling kernels into its own pooling acceleration module. The pooling acceleration module outputs the calculated pooling result, and the PL end sends the pooling result to the PS end using the AXIDMA protocol. The PS end concatenates the received pooling results to finally obtain the complete pooling result corresponding to the current input feature map.

[0105] This invention utilizes a target recognition algorithm to identify driver behavior, avoiding the need for medical equipment to collect drivers' physiological indicators and thus preventing potential negative impacts on drivers. Furthermore, because the recognition algorithm achieves a more accurate and comprehensive assessment of driver behavior classification, capturing subtle movements and details, it provides a more objective evaluation of driver behavior compared to traditional physiological indicator detection. Moreover, compared to CPU or GPU-based deep learning network frameworks, FPGAs possess hardware-level parallel computing capabilities. Therefore, this invention utilizes a processing system (PS) for flow control of the recognition algorithm and programmable logic (PL) for accelerating the computation of recognition algorithm data, achieving more efficient processing and improving computational efficiency, greatly meeting the real-time requirements of automotive applications. Additionally, due to the low-power design of FPGAs, this invention effectively reduces the overall system's energy consumption, thus helping to improve the vehicle system's endurance and meeting the high energy management standards of the automotive environment. Furthermore, the miniaturized design of FPGAs makes them more suitable for limited vehicle space, overcoming the size inconvenience of traditional GPUs, making them easier to deploy and more widely applicable.

[0106] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0107] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0108] In this specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple instances. While different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce a good effect.

[0109] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A method for rapid identification of driver behavior, characterized in that, include: Acquire the image to be identified and send the image to the PS terminal; The PS end uses a trained recognition network and combines the prior bounding box information of the trained recognition network to process the image to be recognized. During the processing, the currently encountered convolution processing task or pooling processing task is sent to the PL end so that the PL end can assist in the processing. The PL terminal executes the processing task sent by the PS terminal and sends the obtained calculation result to the PS terminal; The PS end continues to process the calculation results and sends newly encountered convolution or pooling tasks to the PL end for the PL end to assist in the processing until the PS end obtains the three-dimensional processing result of the image to be recognized. The PS terminal decodes the 3D processing result and then performs non-maximum suppression processing to obtain the recognition result of the image to be recognized. The step of the PS end sending the currently encountered pooling processing task to the PL end includes: The PS terminal divides the current input feature map into blocks and sends the pooling kernel used to pool the current input feature map, as well as the resulting input feature map blocks, to the PL terminal in real time; wherein, the current input feature map refers to the feature map that needs to be pooled. Accordingly, the PL terminal executes the processing task sent by the PS terminal and sends the obtained calculation result to the PS terminal, including: The PL terminal receives the input feature map blocks and pooling kernels sent by the PS terminal in real time, and performs pooling processing on each received input feature map block using the received pooling kernel, and then sends the pooling result to the PS terminal in real time. Accordingly, the PS end continues processing based on the calculation result, and sends newly encountered convolution processing tasks or pooling processing tasks to the PL end for the PL end to assist in processing, until the PS end obtains the three-dimensional processing result of the image to be recognized, including: The PS end will concatenate the pooling results received in real time from the PL end until the complete pooling result of the current input feature map is obtained. When the PS terminal needs to perform convolution or pooling operations on the complete pooling result of the current input feature map, the PS terminal generates a new convolution or pooling processing task based on the complete pooling result of the current input feature map, and continues to send it to the PL terminal so that the PL terminal can assist in the processing, until the PS terminal obtains the three-dimensional processing result of the image to be recognized. When the PS terminal needs to perform operations other than convolution and pooling on the complete pooling result of the current input feature map, the PS terminal performs operations on the complete pooling result of the current input feature map. When it needs to perform convolution or pooling operations on the processed result, the PS terminal generates a new convolution or pooling task based on the processed result and the corresponding pooling kernel, and continues to send it to the PL terminal so that the PL terminal can assist in the processing until the PS terminal obtains the three-dimensional processing result of the image to be recognized. The PL terminal executes the processing task sent by the PS terminal and sends the obtained calculation result to the PS terminal, and also includes: The PL terminal receives the convolution kernel block and input feature map block sent by the PS terminal in real time, and uses the received convolution kernel block to perform convolution processing on the received input feature map block, and then performs accumulation and concatenation to obtain the corresponding output feature map block, and sends the obtained output feature map block to the PS terminal in real time. Accordingly, the PS end continues processing based on the calculation result, and sends newly encountered convolution processing tasks or pooling processing tasks to the PL end for the PL end to assist in processing, until the PS end obtains the three-dimensional processing result of the image to be recognized, including: The PS end will stitch together the output feature map blocks received in real time from the PL end until the current output feature map is obtained; When the PS terminal needs to perform convolution or pooling operations on the current output feature map, the PS terminal generates a new convolution or pooling task based on the current output feature map and continues to send it to the PL terminal so that the PL terminal can assist in the processing until the PS terminal obtains the three-dimensional processing result of the image to be recognized. When the PS terminal needs to perform operations other than convolution and pooling on the current output feature map, the PS terminal performs the operation on the current output feature map. When it needs to perform convolution or pooling on the result of the operation, the PS terminal generates a new convolution or pooling task based on the result of the operation and sends it to the PL terminal to assist in the processing until the PS terminal obtains the three-dimensional processing result of the image to be recognized.

2. The method of claim 1, wherein The steps by which the PS end sends the currently encountered convolution processing task to the PL end include: The PS terminal divides the current input feature map and the convolution kernel used to perform convolution processing on the current input feature map into blocks, and sends the resulting input feature map blocks and convolution kernel blocks to the PL terminal in real time; wherein, the current input feature map refers to the feature map that needs to be convolved.

3. The method of claim 2, wherein The PS terminal divides the current input feature map and the convolution kernel used to perform convolution processing on the current input feature map into blocks, and sends the resulting input feature map blocks and convolution kernel blocks to the PL terminal in real time, including: The PS terminal divides the number of channels of each convolution kernel used for convolution processing of the current input feature map according to the first preset division parameters, with the number of channels as the block dimension, and sends the divided convolution kernel blocks to the PL terminal in real time. The PS end divides the current output feature map according to the second preset division parameters and the current output feature map, and determines the size of each output feature map block obtained when dividing the current output feature map into blocks; the current output feature map refers to the feature map obtained after convolution processing of the current input feature map using a convolution kernel used to convolve the current input feature map; The PS terminal determines the size of each input feature map block obtained when dividing the current input feature map into blocks based on the width and height of each output feature map block, as well as the size, stride, and padding of each convolutional kernel; The PS terminal divides the current input feature map into blocks according to the size of each input feature map block, and sends the resulting input feature map blocks to the PL terminal in real time.

4. The method for rapid driver behavior recognition according to claim 1, characterized in that, The three-dimensional processing results include: a first three-dimensional processing result and a second three-dimensional processing result; wherein, the first three-dimensional processing result is used to detect large targets, and the second three-dimensional processing result is used to detect small targets; both the first three-dimensional processing result and the second three-dimensional processing result contain K×(5+C), where K represents the number of prior boxes of the trained recognition network, 5 represents that each bounding box has five pieces of information: the x-coordinate of the bounding box center point, the y-coordinate of the bounding box center point, the width of the bounding box, the height of the bounding box, and the confidence level of the bounding box, and C represents the number of classification categories.

5. The method for rapid driver behavior recognition according to claim 4, characterized in that, The first 3D processing result is (13, 13, K×(5+C)), and the second 3D processing result is (26, 26, K×(5+C)); where the two 13s in (13, 13, K×(5+C)) indicate that the image is divided into 13 parts. The image is divided into 13 grid cells, where the two 26s in (26, 26, K×(5+C)) represent dividing the image into 26 grid cells. 26 grid cells.

6. The method for rapid driver behavior recognition according to claim 1, characterized in that, The trained recognition network is a trained Tiny YOLOv4 network, and the recognition results include: the driver's position in the image to be recognized, and the driver's behavior category.

Citation Information

Patent Citations

  • MobileNet-SSD target detection device and method based on FPGA acceleration

    CN113051216A

  • Handwritten numeral recognition implementation method

    CN114299514A