Multi - perspective Data Fusion Driver Behavior Prediction Method and Device
Through the multi-view data fusion method, the multi-angle image data is obtained using three cameras to extract and fusion features, solving the problems of inaccurate and occlusion of driver behavior prediction in the prior art, and achieving higher precision driver behavior monitoring and prediction.
Patent Information
- Application Number
- CN202410967440.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-07-18
AI Technical Summary
When monitoring driver abnormal behavior, it is difficult to predict driver behavior comprehensively and accurately, and the limited viewing angle of a single camera leads to occlusion problems, which is poorly robust.
Using a multi-view data fusion method, the driver's first side image, second side image and front image are obtained through three cameras, frame sampling and feature extraction are performed, action feature vectors and eye feature vectors are determined, feature fusion is performed, and the fully connected network layer is input to determine the driver's behavior results.
It effectively avoids occlusion problems, improves prediction accuracy, and can conduct comprehensive and accurate monitoring and prediction of drivers' behavior.
Smart Images

Figure CN118918568B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of data processing, and particularly to a multi-view data fusion driver behavior prediction method. Background Art
[0002] Abnormal behaviors of drivers in the vehicle include: smoking, making phone calls, eating, drinking water, and fatigue driving, etc. Any of the above behaviors will have an adverse impact on the safe driving of the vehicle. Especially fatigue driving will greatly extend the reaction time of the driver. When encountering an emergency, it is difficult for the driver to make a standardized operation in time, thus leading to traffic accidents. Therefore, in the field of assisted driving, real-time monitoring of drivers' abnormal behaviors is an important means to ensure the travel safety of drivers. At present, the methods for monitoring drivers' abnormal behaviors mainly include physiological signal-based monitoring methods and vision-based monitoring methods. The physiological signal-based monitoring technology mainly monitors physiological signals such as electrocardiogram (ECG), electroencephalogram (EEG), and skin conductance response in real time, and can identify the fatigue and distraction states of drivers. The vision-based monitoring technology mainly captures the facial expressions, eye movements, and body movements of drivers through cameras, and uses convolutional neural networks (CNNs) and other deep learning models for analysis to provide abnormal behavior predictions.
[0003] The methods for physiological signal monitoring often require wearing instrument devices and are difficult to apply in actual scenarios; while most of the existing vision-based solutions only focus on the first four or the last one of the abnormal behaviors such as smoking, making phone calls, eating, drinking water, and fatigue driving, that is, they only focus on the driver's actions or only focus on whether the driver has fatigue driving behavior. However, such methods cannot comprehensively and accurately predict the driver's behavior; at the same time, most of the existing methods use a single camera for image data collection. Therefore, due to the limited perspective, it is easy to cause occlusion and inaccurate behavior detection; at the same time, due to the complexity of the actual situation, the driver may wear glasses, sunglasses, etc. The robustness of the existing eye key point localization methods is poor. Thus, there is an urgent need for a better solution. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a multi-view data fusion driver behavior prediction method. One or more embodiments of this specification also relate to a multi-view data fusion driver behavior prediction device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a multi-view data fusion driver behavior prediction method is provided, including:
[0006] Obtain image data; wherein, the image data includes a first side image, a second side image, and a front image;
[0007] Perform frame sampling based on the image data to determine frame data;
[0008] Determine a multi-view feature vector based on the frame data;
[0009] Perform feature fusion based on the multi-view feature vector to determine a fusion feature, and input the fusion feature into a fully connected network layer to determine the driver behavior result.
[0010] In a possible implementation, obtaining the image data includes:
[0011] Obtain the first side image of the driver through a first camera;
[0012] Obtain the second side image of the driver through a second camera;
[0013] Obtain the front image of the driver through a third camera;
[0014] Determine the image data based on the first side image, the second side image, and the front image.
[0015] In a possible implementation, performing frame sampling based on the image data to determine frame data includes:
[0016] Determine a first interval and a second interval;
[0017] Perform frame sampling based on the first side image, the second side image, and the first interval to determine action frame data;
[0018] Perform frame sampling based on the front image and the second interval to determine eye frame data.
[0019] In a possible implementation, determining a multi-view feature vector based on the frame data includes:
[0020] Perform action analysis on the action frame data to determine an action feature vector;
[0021] Perform eye feature extraction on the eye frame data to determine an eye feature vector.
[0022] In a possible implementation, performing action analysis on the action frame data to determine an action feature vector includes:
[0023] Extract skeleton sequence data based on the action frame data;
[0024] Determine angle data based on the skeleton sequence data, and input the angle data into a convolutional model to determine an initial action feature;
[0025] Based on the initial action features corresponding to the first side image and the second side image, perform feature fusion to determine the action feature vector.
[0026] In a possible implementation, perform eye feature extraction on the eye frame data to determine the eye feature vector, including:
[0027] Perform feature extraction based on the eye frame data to determine the initial eye features;
[0028] Perform attention calculation based on the initial eye features to determine the attention weights;
[0029] Perform feature fusion based on the attention weights and the original feature map corresponding to the eye frame data to obtain the target eye features;
[0030] Perform eye state evaluation based on the target eye features to determine the eye feature vector.
[0031] In a possible implementation, perform feature fusion based on the multi-view feature vectors to determine the fusion features, including:
[0032] Determine the eye feature weights and the action feature weights;
[0033] Perform feature fusion based on the eye feature weights, the action feature weights, the action feature vector, and the eye feature vector to determine the fusion features.
[0034] According to the second aspect of the embodiments of the present specification, there is provided a multi-view data fusion driver behavior prediction device, including:
[0035] A data acquisition module, configured to acquire image data; wherein, the image data includes a first side image, a second side image, and a front image;
[0036] A data sampling module, configured to perform frame sampling based on the image data to determine the frame data;
[0037] A feature determination module, configured to determine the multi-view feature vectors based on the frame data;
[0038] A feature fusion module, configured to perform feature fusion based on the multi-view feature vectors to determine the fusion features, and input the fusion features into a fully connected network layer to determine the driver behavior result.
[0039] According to the third aspect of the embodiments of the present specification, there is provided a computing device, including:
[0040] A memory and a processor;
[0041] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above multi-view data fusion driver behavior prediction method are implemented.
[0042] According to the fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above multi-view data fusion driver behavior prediction method are implemented.
[0043] According to the fifth aspect of the embodiments of the present specification, a computer program is provided. When the computer program is executed on a computer, the computer is made to execute the steps of the above multi-view data fusion driver behavior prediction method.
[0044] The embodiments of the present specification provide a multi-view data fusion driver behavior prediction method and device. The multi-view data fusion driver behavior prediction method includes: acquiring image data; wherein the image data includes a first side image, a second side image, and a front image; performing frame sampling based on the image data to determine frame data; determining a multi-view feature vector based on the frame data; performing feature fusion based on the multi-view feature vector to determine a fusion feature, and inputting the fusion feature into a fully connected network layer to determine the driver behavior result. By using three cameras and analyzing the driver behavior from two aspects of the overall action and eye state, the occlusion problem can be effectively avoided and the prediction accuracy can be increased. Description of the Drawings
[0045] Figure 1 is a flowchart of a multi-view data fusion driver behavior prediction method provided by an embodiment of the present specification;
[0046] Figure 2 is a schematic diagram of the principle of a multi-view data fusion driver behavior prediction method provided by an embodiment of the present specification;
[0047] Figure 3 is a schematic diagram of the scene of a multi-view data fusion driver behavior prediction method provided by an embodiment of the present specification;
[0048] Figure 4 is a schematic diagram of action feature extraction of a multi-view data fusion driver behavior prediction method provided by an embodiment of the present specification;
[0049] Figure 5 is a schematic diagram of human key points of a multi-view data fusion driver behavior prediction method provided by an embodiment of the present specification;
[0050] Figure 6 is a network architecture diagram of feature extraction of a multi-view data fusion driver behavior prediction method provided by an embodiment of the present specification;
[0051] Figure 7 It is a schematic diagram of eye key point detection for a multi - perspective data fusion driver behavior prediction method provided by an embodiment of this specification;
[0052] Figure 8 It is a schematic diagram of the positions of eye key points for a multi - perspective data fusion driver behavior prediction method provided by an embodiment of this specification;
[0053] Figure 9 It is a schematic structural diagram of a multi - perspective data fusion driver behavior prediction device provided by an embodiment of this specification;
[0054] Figure 10 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0055] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0056] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0057] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first can also be referred to as the second, and similarly, the second can also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining".
[0058] In this specification, a multi - perspective data fusion driver behavior prediction method is provided. This specification also relates to a multi - perspective data fusion driver behavior prediction device, a computing device, and a computer - readable storage medium, which will be described in detail one by one in the following embodiments.
[0059] SeeFigure 1 , Figure 1 shows a flowchart of a multi-view data fusion driver behavior prediction method provided according to an embodiment of this specification, which specifically includes the following steps.
[0060] Step 101: Obtain image data; wherein, the image data includes a first side image, a second side image, and a front image.
[0061] See Figure 2 , considering the particularity of the in-vehicle scenario and the characteristics of drivers' abnormal behaviors in the embodiments of this specification, three cameras are used to collect data, and multi-view data based on the drivers' limb movements and eye features is used to comprehensively detect drivers' abnormal behaviors. Finally, the safety level of the driver's behavior is output through decision fusion, which can not only effectively reduce the problems caused by picture occlusion, but also comprehensively monitor abnormal behaviors, timely give early warnings to the driver's dangerous driving behaviors, and thus ensure travel safety.
[0062] In a possible implementation manner, obtaining image data includes: obtaining a first side image of the driver through a first camera; obtaining a second side image of the driver through a second camera; obtaining a front image of the driver through a third camera; and determining the image data based on the first side image, the second side image, and the front image.
[0063] In practical applications, see Figure 3 , first, multi-view data is obtained. Three cameras are used to collect data in the vehicle. The first and second cameras are on the left and right sides of the driver and are responsible for collecting action information, and the third camera is directly facing the driver and is responsible for collecting face and eye information.
[0064] The embodiments of this specification use three cameras and analyze the driver's behavior from both the overall action and eye state aspects, which can effectively avoid occlusion problems and increase the prediction accuracy.
[0065] Step 102: Perform frame sampling based on the image data to determine frame data.
[0066] In a possible implementation manner, performing frame sampling based on the image data to determine frame data includes: determining a first interval and a second interval; performing frame sampling based on the first side image, the second side image, and the first interval to determine action frame data; and performing frame sampling based on the front image and the second interval to determine eye frame data.
[0067] In practical applications, according to different processing information, this method adopts an equal-frame sampling method with different intervals for different cameras. The video information collected by the camera is tentatively set at 30 frames per second. A video segment is a unit for extracting feature information. The time for capturing actions such as drinking water, smoking, eating, and nodding off is approximately 0.5 seconds to 3 seconds, and the time for capturing fatigue states such as blinking is approximately 0.25 seconds. Therefore, the frame sampling module of this method uses different numbers of sampling segments in different processing branches. In action recognition, the sampling interval is 2 frames, and in eye feature recognition, the interval is 1 frame. 36 frames and 18 frames are respectively used as a sampling segment in the two branches for sampling. The sampling order is from first to last, which ensures the timing of network sampling, reduces the sampling task volume, and improves the system efficiency.
[0068] In the embodiment of this specification, during frame sampling, the sampling strategy is to adopt an equal-frame sampling method with different intervals for different cameras according to different processing information, which can improve the system efficiency.
[0069] Step 103: Determine the multi-view feature vector based on the frame data.
[0070] In practical applications, referring to Figure 3 , after the data collected by the first and second cameras are respectively frame-sampled and then merged and input into the action feature extraction module to obtain intermediate features f1 and f2, the two-view features are then fused to obtain the action feature F1; the data collected by the third camera is frame-sampled and then input into the eye key point localization module to localize the key points, and then the EAR value is calculated to obtain the eye feature F2.
[0071] In a possible implementation manner, determining the multi-view feature vector based on the frame data includes: performing action analysis on the action frame data to determine the action feature vector; performing eye feature extraction on the eye frame data to determine the eye feature vector.
[0072] Specifically, referring to Figure 4 , performing action analysis on the action frame data to determine the action feature vector includes: extracting skeleton sequence data based on the action frame data; determining angle data based on the skeleton sequence data, and inputting the angle data into a convolutional model to determine the initial action feature; performing feature fusion on the initial action features corresponding to the first side image and the second side image to determine the action feature vector.
[0073] For example, due to the high reliability of Openpose, it can be selected as the pre-network of the system to process continuous video streams. Since the video data needs to be processed in real time, and considering the driving scenario, the recognition of the driver's actions mainly involves the skeleton sequence of the driver's upper body. Therefore, in order to save computational overhead as much as possible while ensuring accuracy, 7 key points can be selected as the main human skeleton, such as Figure 5 shown.
[0074] Table 1
[0075] Annotation order Joint name 0 Nose 1 Left elbow 2 Left wrist 3 Right elbow 4 Right wrist 5 Neck 6 Bottom of spine
[0076] Furthermore, referring to Table 1, after passing through the OpenPose model, a human body key point heat map can be obtained. Applying NMS (Non-Maximum Suppression) to the human body key point heat map gives the scattered coordinates of the human body key points. Then, combining the PAF (Part Affinity Field) information, the affinity between key point pairs is calculated through 8-point discrete sampling line integral. Finally, the KM algorithm is applied to assemble the human skeleton sequence required by this method for use by subsequent modules.
[0077] After outputting the human skeleton sequence, these skeleton sequences will be used as the input of the action recognition network for detection. Since the recognition of actions involves many factors, for example, during driving, slowly turning the head may be that the driver is actively relaxing the shoulders and neck, but if the head position changes too quickly, it may be a momentary nodding off due to fatigue driving. Therefore, this technology uses three branches of angle, angular velocity, and angular acceleration for action recognition. After convolution, the results are fused. The structural schematic diagram is as Figure 6 shown.
[0078] The network input size is 6 * 36 * 3, which is stacked by 3 independent data, namely the angle sequence, the angular velocity sequence, and the angular acceleration sequence. They describe the motion information from 3 dimensions for use by subsequent networks. Combining the posture characteristics during the driver's driving process, the defined human skeleton consists of 7 key points, and the 7 key points can form 6 segments of bones, minimizing the system cost as much as possible. The essence of an action is the movement of the skeleton, so the skeleton sequence theoretically contains all the information required to understand the action. For the driver, the transformation of the human body posture is mainly reflected in the change of the angle between each segment of the bone and the reference plane. Therefore, we first calculate the angle between each segment of the bone and the horizontal plane according to the skeleton information of each frame to form a 6-dimensional feature vector:
[0079] f = [θ1, θ2, θ3, θ4, θ5, θ6]
[0080] where the angle of each segment of the bone is calculated by the arctangent formula:
[0081]
[0082] In actual operation, the direction of the bone vector should be judged first, and the angle value should be corrected according to the quadrant where it is located, so that the angles of all bones are within [0, 2π). We formed a 6D feature vector by manually extracting features from the skeleton of each frame. In this technology, we set the sliding window size for each recognition to 36 frames, that is, each time 6 * 36 data is sent to the first branch of the subsequent convolutional module for calculation and intermediate features are output.
[0083]
[0084] We can obtain the angular velocity and displacement velocity sequences by taking the first-order difference of the obtained angle and centroid position sequences:
[0085] The range of the bone included angle defined in the previous text is [0, 2π), and the counterclockwise direction is the positive direction. Therefore, a positive angular velocity represents that the movement trend of the bone is counterclockwise, and a negative one is the opposite. Similarly, we can continue to take the difference of the angular velocity and displacement velocity sequences to obtain the acceleration sequence:
[0086]
[0087] The above operations decompose the obtained skeleton sequence into three sequences of angles, angular velocities, and angular accelerations required by the subsequent network, providing relatively rich action semantic information for the network from different dimensions. Then, the data of the three branches are input into the convolutional model. The convolutional kernel size is 6 * 4, the stride is 4, and convolution is performed along the frame length direction. Finally, the 6 * 9 feature data output by each branch are subjected to feature fusion, converting the three channels into a single channel, reducing the dimension of the extracted features, and then fusing the features f1 and f2 of the two cameras to obtain the driver action recognition feature F1(18 * 1).
[0088] In the action recognition of the embodiments of this specification, this design only selects 7 bone key points in combination with the driver's action characteristics, which can effectively reduce the computational cost while basically meeting the actual needs, and is suitable for application in real-time systems.
[0089] Specifically, for the eye frame data, eye feature extraction is performed to determine the eye feature vector, including: performing feature extraction based on the eye frame data to determine the initial eye features; performing attention calculation based on the initial eye features to determine the attention weights; performing feature fusion based on the attention weights and the original feature map corresponding to the eye frame data to obtain the target eye features; and performing eye state evaluation based on the target eye features to determine the eye feature vector.
[0090] In practical applications, the self-attention mechanism is a technique that can dynamically adjust the attention distribution based on the internal relationships of the input data. In this mechanism, each input position is assigned a relationship weight with other positions, and then based on these weights, the model's attention to the input is adjusted. The eye key-point detection method based on the self-attention mechanism can dynamically adjust the model's attention to the eye key-points according to the features of the eye region in the input image, thereby improving the accuracy and robustness of eye key-point detection.
[0091] See Figure 7 , first, the input driver face image is passed through a convolutional neural network (CNN) to extract a feature map. For each position in the feature map, it is mapped to a low-dimensional feature vector. For the feature vector of each position, a fully connected network is used to calculate the relationship weights with other positions. Then, according to the calculated attention weights, the feature vectors of each position in the feature map are weighted and summed to obtain a weighted feature representation. The weighted feature representation is fused with the original feature map to obtain the final feature representation.
[0092] The self-attention network consists of three main parts: a feature extraction module, an attention calculation module, and a feature fusion module.
[0093] Feature extraction module: A convolutional neural network (CNN) is used to extract the features of the input image.
[0094] Attention calculation module: A fully connected network is used. The input of this fully connected network is the feature vector of each position in the feature map, and the output is the attention weight corresponding to the position. Through this fully connected network, the model can learn the importance of each position in the entire input image and adjust the model's attention to the input according to these weights.
[0095] Feature fusion module: Apply the calculated attention weights to the feature map and fuse the weighted features with the original feature map.
[0096] After obtaining the coordinates of the eye key-points, the EAR value is calculated. The EAR value of the eye refers to the eye aspect ratio. It is an index used to evaluate the eye state.
[0097]
[0098] Among them, p1, p2, p3, p4, p5, p6 respectively represent the coordinates of the eye key-points, and the positions are as Figure 8 shown. Calculate the EAR values of 18 frames for both eyes each time, and average the EAR values of both eyes to obtain the EAR vector F2(18*1).
[0099] This embodiment of the specification proposes a brand-new eye key point detection method. The eye key point detection method based on the self-attention mechanism can dynamically adjust the model's attention degree to the eye key points according to the features of the eye region in the input image, thereby improving the accuracy and robustness of eye key point detection.
[0100] Step 104: Determine the fusion feature based on the multi-view feature vectors, and input the fusion feature into the fully connected network layer to determine the driver behavior result.
[0101] In practical applications, the action feature and the eye feature are weighted and fused, and finally the danger level of the driver behavior from 1 to 10 is output through the fully connected network.
[0102] In a possible implementation manner, determining the fusion feature based on the multi-view feature vectors includes: determining the eye feature weight and the action feature weight; performing feature fusion based on the eye feature weight, the action feature weight, the action feature vector, and the eye feature vector to determine the fusion feature.
[0103] In practical applications, the action recognition feature and the eye fatigue feature obtained in the above steps are weighted and fused, and different weights are assigned to the F1 feature and the F2 feature.
[0104] Finally, the danger level of the driver action from 0 to 10 is finally output through the fully connected network, where 0 represents normal driving. Then, a reward and punishment mechanism can be designed according to the action danger level to ensure the driving safety of the driver.
[0105] This embodiment of the specification provides a multi-view data fusion driver behavior prediction method and apparatus. The multi-view data fusion driver behavior prediction method includes: acquiring image data; wherein, the image data includes a first side image, a second side image, and a front image; performing frame sampling based on the image data to determine frame data; determining multi-view feature vectors based on the frame data; determining the fusion feature based on the multi-view feature vectors, and inputting the fusion feature into the fully connected network layer to determine the driver behavior result. By using three cameras and analyzing the driver behavior from both the overall action and the eye state aspects, the occlusion problem can be effectively avoided and the prediction accuracy can be increased.
[0106] Corresponding to the above method embodiment, this specification also provides an embodiment of a multi-view data fusion driver behavior prediction apparatus. Figure 9 The structural schematic diagram of a multi-view data fusion driver behavior prediction apparatus provided by an embodiment of this specification is shown. As Figure 9 shown, the apparatus includes:
[0107] The data acquisition module 901 is configured to acquire image data; wherein, the image data includes a first side image, a second side image, and a front image;
[0108] The data sampling module 902 is configured to perform frame sampling based on the image data to determine frame data;
[0109] The feature determination module 903 is configured to determine a multi-view feature vector based on the frame data;
[0110] The feature fusion module 904 is configured to perform feature fusion based on the multi-view feature vector to determine a fusion feature, and input the fusion feature into a fully connected network layer to determine the driver behavior result.
[0111] In a possible implementation, the data acquisition module 901 is further configured to:
[0112] Acquire the first side image of the driver through a first camera;
[0113] Acquire the second side image of the driver through a second camera;
[0114] Acquire the front image of the driver through a third camera;
[0115] Determine the image data based on the first side image, the second side image, and the front image.
[0116] In a possible implementation, the data sampling module 902 is further configured to:
[0117] Determine a first interval and a second interval;
[0118] Perform frame sampling based on the first side image, the second side image, and the first interval to determine action frame data;
[0119] Perform frame sampling based on the front image and the second interval to determine eye frame data.
[0120] In a possible implementation, the feature determination module 903 is further configured to:
[0121] Perform action analysis on the action frame data to determine an action feature vector;
[0122] Perform eye feature extraction on the eye frame data to determine an eye feature vector.
[0123] In a possible implementation, the feature determination module 903 is further configured to:
[0124] Extract skeleton sequence data based on the action frame data;
[0125] Determine angle data based on the skeleton sequence data, and input the angle data into a convolutional model to determine an initial action feature;
[0126] Based on the initial action features corresponding to the first side image and the second side image, perform feature fusion to determine the action feature vector.
[0127] In a possible implementation, the feature determination module 903 is further configured to:
[0128] Extract features based on the eye frame data to determine the initial eye features;
[0129] Calculate the attention based on the initial eye features to determine the attention weights;
[0130] Perform feature fusion based on the attention weights and the original feature map corresponding to the eye frame data to obtain the target eye features;
[0131] Evaluate the eye state based on the target eye features to determine the eye feature vector.
[0132] In a possible implementation, the feature fusion module 904 is further configured to:
[0133] Determine the eye feature weights and the action feature weights;
[0134] Perform feature fusion based on the eye feature weights, the action feature weights, the action feature vector, and the eye feature vector to determine the fusion features.
[0135] The embodiments of this specification provide a multi-view data fusion driver behavior prediction method and apparatus. The multi-view data fusion driver behavior prediction apparatus includes: obtaining image data, where the image data includes a first side image, a second side image, and a front image; performing frame sampling based on the image data to determine frame data; determining multi-view feature vectors based on the frame data; performing feature fusion based on the multi-view feature vectors to determine fusion features, and inputting the fusion features into a fully connected network layer to determine the driver behavior result. By using three cameras and analyzing the driver behavior from both the overall action and the eye state aspects, the occlusion problem can be effectively avoided and the prediction accuracy can be increased.
[0136] The above is a schematic solution of a multi-view data fusion driver behavior prediction apparatus in this embodiment. It should be noted that the technical solution of this multi-view data fusion driver behavior prediction apparatus and the technical solution of the above multi-view data fusion driver behavior prediction method belong to the same concept. For the details not described in the technical solution of the multi-view data fusion driver behavior prediction apparatus, reference can be made to the description of the technical solution of the above multi-view data fusion driver behavior prediction method.
[0137] Figure 10The structural block diagram of a computing device 1000 provided according to an embodiment of this specification is shown. The components of the computing device 1000 include but are not limited to a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 via a bus 1030, and a database 1050 is used to store data.
[0138] The computing device 1000 further includes an access device 1040, and the access device 1040 enables the computing device 1000 to communicate via one or more networks 1060. Examples of these networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1040 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0139] In an embodiment of this specification, the above components of the computing device 1000 and Figure 10 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 10 the shown structural block diagram of the computing device is only for illustrative purposes and is not a limitation on the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0140] The computing device 1000 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). The computing device 1000 can also be a mobile or stationary server.
[0141] Among them, the processor 1020 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above multi-view data fusion driver behavior prediction method are implemented. The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above multi-view data fusion driver behavior prediction method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above multi-view data fusion driver behavior prediction method.
[0142] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the steps of the above multi-view data fusion driver behavior prediction method are implemented.
[0143] The above is a schematic solution of a computer-readable storage medium in this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above multi-view data fusion driver behavior prediction method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above multi-view data fusion driver behavior prediction method.
[0144] An embodiment of this specification also provides a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above multi-view data fusion driver behavior prediction method.
[0145] The above is a schematic solution of a computer program in this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above multi-view data fusion driver behavior prediction method belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the description of the technical solution of the above multi-view data fusion driver behavior prediction method.
[0146] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0147] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a removable hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0148] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0149] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0150] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to only the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A multi-view data fusion driver behavior prediction method, characterized in that: include: Acquire image data; wherein the image data includes a first side image, a second side image and a front image; Perform frame sampling based on the image data to determine frame data; Determine a multi-view feature vector based on the frame data; Performing feature fusion based on the multi-view feature vectors to determine fusion features, and inputting the fusion features into a fully connected network layer to determine the driver behavior result; The performing frame sampling based on the image data to determine the frame data includes: determining a first interval and a second interval; Perform frame sampling based on the first side image, the second side image and the first interval to determine action frame data; Perform frame sampling based on the frontal image and the second interval to determine eye frame data; The determining of the multi-view feature vector based on the frame data comprises: Performing motion analysis on the motion frame data to determine a motion feature vector; Extracting eye features from the eye frame data to determine an eye feature vector; The performing motion analysis on the motion frame data to determine the motion feature vector includes: Extracting skeleton sequence data based on the action frame data; Determine angle data based on the skeleton sequence data, and input the angle data into a convolution model to determine initial motion features; Perform feature fusion based on the initial action features corresponding to the first side image and the second side image to determine an action feature vector; The step of extracting eye features from the eye frame data to determine an eye feature vector includes: Performing feature extraction based on the eye frame data to determine initial eye features; Performing attention calculation based on the initial eye features to determine an attention weight; Performing feature fusion based on the attention weight and the original feature map corresponding to the eye frame data to obtain a target eye feature; Performing eye state assessment based on the target eye feature to determine an eye feature vector; The step of performing feature fusion based on the multi-view feature vectors to determine fusion features includes: Determine eye feature weights and motion feature weights; Feature fusion is performed based on the eye feature weight, the action feature weight, the action feature vector and the eye feature vector to determine a fusion feature.
2. The method according to claim 1, characterized in that The acquiring of image data comprises: Acquire the first side image of the driver by a first camera; acquiring the second side image of the driver by a second camera; acquiring the front image of the driver by a third camera; Image data is determined based on the first lateral image, the second lateral image, and the frontal image.
3. A multi-view data fusion driver behavior prediction device, used to implement the steps of the multi-view data fusion driver behavior prediction method according to any one of claims 1 to 2, characterized in that: include: A data acquisition module, configured to acquire image data; wherein the image data includes a first side image, a second side image and a front image; A data sampling module is configured to perform frame sampling based on the image data to determine frame data; A feature determination module, configured to determine a multi-view feature vector based on the frame data; The feature fusion module is configured to perform feature fusion based on the multi-view feature vector to determine a fusion feature, and input the fusion feature into a fully connected network layer to determine a driver behavior result.
4. A computing device, characterized in that include: Memory and processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the multi-view data fusion driver behavior prediction method described in any one of claims 1 to 2 are implemented.
5. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the multi-view data fusion driver behavior prediction method described in any one of claims 1 to 2.
Citation Information
Patent Citations
Child ADHD screening evaluation system based on a multi-modal deep learning technology
CN111528859A
Driving behavior detection method based on multiple view angles
CN118015600A