Human Factor Intelligent Driving Behavior Prediction Method, System, Terminal Device and Storage Medium
Through multimodal fusion of driver physiological signals and vehicle scene videos and three-dimensional backbone network analysis, the adaptability, accuracy and real-time problems of Transformer model in intelligent driving behavior prediction are solved, and the overall effect of driving behavior prediction is improved.
Patent Information
- Application Number
- CN202310987155.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-08-07
AI Technical Summary
The existing Transformer model has problems such as high adaptability but low accuracy, poor interpretability, low real-time and high data processing time consumption in human intelligent driving behavior prediction.
By obtaining the driver's physiological signals, fast Fourier transform and multi-period decomposition are performed, combining vehicle road scene video frame prediction, multimodal synchronous data fusion and three-dimensional backbone network layer for feature analysis, and driving behavior interpretation and inference layers are introduced to improve the real-time and accuracy of prediction.
It improves the real-time and accuracy of driving behavior prediction, reduces the time consumption of task submodules, and enhances the explanatory and generalization capabilities of the model.
Smart Images

Figure CN117022305B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent cockpits, and particularly to a human factor intelligent driving behavior prediction method, system, terminal device, and storage medium. Background Art
[0002] Human factor intelligent driving behavior prediction is a technology that predicts the possible future behaviors or actions of a vehicle by analyzing and learning historical driving data. Such predictions may include behaviors such as vehicle steering, acceleration, deceleration, and lane change. In autonomous driving technology, human factor intelligent driving behavior prediction is particularly important. By predicting the behaviors of other vehicles and pedestrians, the autonomous driving system can make decisions in advance to avoid possible collisions and improve driving safety. Human factor intelligent driving behavior prediction usually relies on machine learning and artificial intelligence technologies, including but not limited to algorithms such as deep learning and reinforcement learning. These algorithms learn from a large amount of driving data to capture the patterns of driving behaviors, thereby realizing the prediction of future behaviors.
[0003] The Transformer model is a deep learning model based on the self-attention mechanism and is widely used in the field of natural language processing. However, due to its powerful sequence modeling ability, the Transformer model is also used to process other types of sequence data, including human factor intelligent driving behavior prediction. In the application of human factor intelligent driving behavior prediction, the Transformer model can effectively process the trajectory data of the vehicle, which is essentially a time series. Information such as position, speed, and acceleration at each moment can be regarded as an element in the sequence, and the Transformer model can capture the dependencies between these elements to predict the future behavior of the vehicle.
[0004] In practical applications, traditional Transformer models have the following defects in the application scenario of human factor intelligent driving behavior prediction: (1) When the end-to-end intelligent cockpit model based on the Transformer model is used for driving behavior prediction, it has high adaptability but low accuracy; (2) The interpretability of the end-to-end autonomous driving based on the Transformer is poor, which hinders its practical application; (3) The end-to-end intelligent cockpit model based on the Transformer will design task sub-modules according to the data characteristics of different sensors, and then perform feature extraction separately. Each sub-task also often uses deep learning models. For example, a convolutional neural network model is used for visual data, and a multi-layer perceptron model such as a BP neural network is used for regression tasks. Each task sub-module also consumes a lot of time, resulting in its use only being possible for intelligent cockpit vehicles at very low speeds; (4) Each prediction task of the Transformer model uses a complete segment of data. However, in the intelligent cockpit environment, if predictions are made after all the data is collected, the real-time performance of the intelligent cockpit model will be reduced; thus, the above-mentioned defects will lead to a deterioration in the prediction effect of driving behavior. Summary of the Invention
[0005] In order to improve the prediction effect of driving behavior, the present application provides a human factor intelligent driving behavior prediction method, system, terminal device and storage medium.
[0006] In the first aspect, the present application provides a human factor intelligent driving behavior prediction method, including the following steps:
[0007] Obtain the physiological signals of the driver;
[0008] Perform a fast Fourier transform on the physiological signals to generate corresponding amplitude-frequency characteristics, and obtain the acquisition frequencies in the amplitude-frequency characteristics that meet the preset amplitude-frequency selection criteria;
[0009] Perform multi-period decomposition on the physiological signals according to the periods of the acquisition frequencies to generate corresponding data decomposition result samples; perform two-dimensional spatial expansion on the data decomposition result samples respectively according to the multi-source time series data coding layer to generate corresponding two-dimensional spatial data;
[0010] Predict the target continuous frames corresponding to the vehicle road scene video according to the vehicle road scene video frame prediction layer to generate corresponding iterative predicted future frames;
[0011] Perform a merging operation on the two-dimensional spatial data, the target continuous frames, and the iterative predicted future frames according to the multi-modal synchronous data fusion layer to generate corresponding multi-scale three-dimensional features;
[0012] Perform feature analysis and processing on the multi-scale 3D features according to the 3D backbone network layer to generate corresponding target output features; perform analysis and processing on the target output features according to the driving behavior interpretation layer and the driving behavior reasoning layer respectively to generate corresponding driving behavior description information and driving behavior reasoning information.
[0013] By adopting the above technical solution, physiological data of vehicle drivers is collected and analyzed, and at the same time, combined with the predicted analysis data of vehicle road scene video frames, that is, a vehicle road scene video frame prediction layer is introduced, which can be used for event prediction in the vehicle driving environment, and human factor intelligent driving behavior prediction is carried out under the condition that the event has not occurred, rather than waiting for the event to occur and then performing classification and prediction. Furthermore, the real-time performance of the overall vehicle behavior prediction is improved. Then, after fusing the physiological state of the above driver and the multi-modal synchronous data corresponding to the vehicle road prediction, feature extraction of the multi-modal synchronous data is performed through the 3D backbone network layer to obtain corresponding compressed target output features. Different from the traditional end-to-end intelligent cockpit model, the time consumption brought by each task sub-module is reduced. Secondly, by introducing the human factor intelligent driving behavior interpretation layer and the human factor intelligent driving behavior reasoning layer to analyze the obtained target output features, the behavior of the vehicle on the road can be explained, and the reason behind the behavior adopted by the vehicle can be explained. Through the data extraction and analysis of the above algorithm logic layer, the overall prediction effect of driving behavior is thus improved.
[0014] Optionally, the performing feature analysis and processing on the multi-scale 3D features according to the 3D backbone network layer to generate corresponding target output features includes the following steps:
[0015] Obtain the corresponding 3D feature segmentation rule in the 3D backbone network layer;
[0016] According to the 3D feature segmentation rule, divide the multi-scale 3D features into H / 4 × W / 4 × ((2 + N + 5×3) / 6) sub-features.
[0017] By adopting the above technical solution, through linear mapping, the dimension of the sub-features can be reduced or increased. Dimension reduction can reduce the number of model parameters and computational complexity, improve computational efficiency, and at the same time, dimension increase can introduce more feature expression dimensions and improve the expression ability of the model.
[0018] Optionally, after dividing the multi-scale 3D features into H / 4 × W / 4 × ((2 + N + 5×3) / 6) sub-features according to the 3D feature segmentation rule, the following steps are further included:
[0019] Obtain the corresponding linear encoding rule in the 3D backbone network layer;
[0020] According to the linear coding rule, each of the sub - features is linearly mapped to a vector C, where the vector C can be of any dimension.
[0021] By adopting the above - mentioned technical solution, mapping high - dimensional features to a low - dimensional vector C helps reduce computational complexity, thereby improving the analysis and calculation efficiency of the data model.
[0022] Optionally, the three - dimensional backbone network layer includes a self - attention coding rule. The steps for performing feature analysis and processing on the multi - scale three - dimensional features by the three - dimensional backbone network layer to generate corresponding target output features are as follows:
[0023] S1. Perform a spatial sampling on the multi - scale three - dimensional features and output the corresponding first target feature;
[0024] S2. Perform a Video Swin Transformer blocks operation on the multi - scale three - dimensional features and output the corresponding second target feature. The MPL layer in the Video Swin Transformer blocks operation corresponds to a 1×1 convolutional layer in the model, and the number of convolutional kernels is equal to the dimension of the sub - features input to the model;
[0025] S3. Repeat steps S1 and S2;
[0026] S4. Repeat step S3 for K times, where K is a preset positive integer.
[0027] By adopting the above - mentioned technical solution, performing spatial sampling can reduce the size of the multi - scale three - dimensional features to half of the original, expand the number of channels of the multi - scale three - dimensional features to twice the original, and performing the Video Swin Transformer operation can reduce the model parameters, thereby improving the inference speed of the model.
[0028] Optionally, the size of the multi - scale three - dimensional features is H×W×(2 + N + 5×3), and the number of channels is (2 + N + 5×3), where H is the height of the corresponding feature map in the multi - scale three - dimensional features, and W is the width of the corresponding feature map in the multi - scale three - dimensional features.
[0029] By adopting the above - mentioned technical solution, the number of channels is increased according to the operation of the multi - modal synchronous data fusion layer, enabling the model to learn more types of features. Thus, when facing different input data, the model can maintain good recognition ability and improve the generalization ability of the model.
[0030] Optionally, according to a preset selection rule, when selecting N in the number of channels, the number of channels is selected as an integer multiple of 6.
[0031] By adopting the above technical solution, during the deep learning algorithm process, if the number of channels is set to be divisible by 6, then the advantage of parallel computing can be utilized to better use the network architecture in the swin Transform, thereby improving the computing efficiency.
[0032] Optionally, the multi-scale three-dimensional feature is decomposed into H×W×3×((2 + N + 5×3) / 3), where the first three dimensions redefine each frame in the multi-scale three-dimensional feature, and each frame contains H×W×3 pixels.
[0033] By adopting the above technical solution, data can be better organized and processed. Decomposing the multi-scale three-dimensional feature into smaller data blocks can improve the data processing efficiency.
[0034] In a second aspect, the present application provides a human factor intelligent driving behavior prediction system, including:
[0035] A physiological signal acquisition module for acquiring the physiological signals of the driver;
[0036] A transformation module for performing a fast Fourier transform on the physiological signal to generate corresponding amplitude-frequency characteristics and obtaining the acquisition frequencies that meet the preset amplitude-frequency selection criteria in the amplitude-frequency characteristics;
[0037] A multi-period decomposition module for performing multi-period decomposition on the physiological signal according to the period of the acquisition frequency to generate corresponding data decomposition result samples;
[0038] A spatial expansion module for respectively performing two-dimensional spatial expansion on the data decomposition result samples according to the multi-source time series data coding layer to generate corresponding two-dimensional spatial data;
[0039] A prediction module for predicting the target consecutive frames corresponding to the vehicle road scene video according to the vehicle road scene video frame prediction layer to generate corresponding iterative predicted future frames;
[0040] A data fusion module for performing a merging operation on the two-dimensional spatial data, the target consecutive frames, and the iterative predicted future frames according to the multi-modal synchronous data fusion layer to generate corresponding multi-scale three-dimensional features;
[0041] A feature analysis module for performing feature analysis processing on the multi-scale three-dimensional feature according to the three-dimensional backbone network layer to generate corresponding target output features;
[0042] A behavior interpretation and reasoning module for respectively performing analysis processing on the target output features according to the driving behavior interpretation layer and the driving behavior reasoning layer to generate corresponding driving behavior description information and driving behavior reasoning information.
[0043] By adopting the above technical solutions, the physiological data of the vehicle driver can be collected and analyzed by the physiological signal acquisition module, transformation module, multi-period decomposition module and spatial expansion module, so as to more accurately understand the physiological state of the driver. At the same time, through the prediction module combined with the prediction analysis data of the vehicle road scene video frames, that is, introducing the vehicle road scene video frame prediction layer, it can be used for event prediction in the vehicle driving environment, and conduct human factor intelligent driving behavior prediction under the condition that the event has not occurred, rather than waiting for the event to occur and then conducting classification and prediction, thereby improving the real-time performance of the overall vehicle behavior prediction. Then, after fusing the above physiological state of the driver and the multi-modal synchronous data corresponding to the vehicle road prediction through the data fusion module, the three-dimensional backbone network layer in the feature analysis module extracts the features of the multi-modal synchronous data to obtain the corresponding compressed target output features. Different from the traditional end-to-end intelligent cockpit model, it reduces the time consumption brought by each task sub-module. Secondly, through the behavior interpretation and reasoning module, introducing the human factor intelligent driving behavior interpretation layer and the human factor intelligent driving behavior reasoning layer to analyze the obtained target output features, it can explain the behavior of the vehicle on the road and explain the reasons behind the behavior taken by the vehicle. Through the data extraction and analysis of the above algorithm logic layer, the overall prediction effect of the driving behavior is improved.
[0044] In a third aspect, the present application provides a terminal device, adopting the following technical solutions:
[0045] A terminal device includes a memory and a processor. Computer instructions capable of running on the processor are stored in the memory. When the processor loads and executes the computer instructions, the above-mentioned human factor intelligent driving behavior prediction method is adopted.
[0046] By adopting the above technical solutions, by generating computer instructions for the above-mentioned human factor intelligent driving behavior prediction method and storing them in the memory to be loaded and executed by the processor, thus, a terminal device is manufactured according to the memory and the processor, which is convenient to use.
[0047] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solutions:
[0048] A computer-readable storage medium stores computer instructions. When the computer instructions are loaded and executed by a processor, the above-mentioned human factor intelligent driving behavior prediction method is adopted.
[0049] By adopting the above technical solutions, by generating computer instructions for the above-mentioned human factor intelligent driving behavior prediction method and storing them in the computer-readable storage medium to be loaded and executed by the processor, through the computer-readable storage medium, it is convenient for the computer instructions to be readable and stored.
[0050] In summary, the present application includes at least one of the following beneficial technical effects: By collecting and analyzing the physiological data of vehicle drivers, the physiological state of the driver can be understood more accurately. At the same time, combined with the predictive analysis data of vehicle road scene video frames, that is, introducing a vehicle road scene video frame prediction layer, it can be used for event prediction in the vehicle driving environment, and human factor intelligent driving behavior prediction can be carried out under the condition that the event has not occurred, instead of waiting for the event to occur and then classifying and predicting, thereby improving the real-time performance of the overall vehicle behavior prediction. Then, after fusing the above-mentioned physiological state of the driver and the multi-modal synchronous data corresponding to the vehicle road prediction, feature extraction of the multi-modal synchronous data is performed through a three-dimensional backbone network layer to obtain corresponding compressed target output features. Different from the traditional end-to-end intelligent cockpit model, it reduces the time consumption brought by each task sub-module. Secondly, introducing a human factor intelligent driving behavior interpretation layer and a human factor intelligent driving behavior reasoning layer to analyze the obtained target output features can explain the behavior of the vehicle on the road and explain the reasons behind the behavior adopted by the vehicle. Through the data extraction and analysis of the above algorithm logic layer, the overall prediction effect of driving behavior is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is a schematic flowchart of steps S101 to S108 in a human factor intelligent driving behavior prediction method of the present application.
[0052] Figure 2 is a logic block diagram of functional modules in a human factor intelligent driving behavior prediction method of the present application.
[0053] Figure 3 is a schematic flowchart of steps S201 to S202 in a human factor intelligent driving behavior prediction method of the present application.
[0054] Figure 4 is a schematic flowchart of steps S301 to S302 in a human factor intelligent driving behavior prediction method of the present application.
[0055] Figure 5 is a schematic flowchart of steps S1 to S4 in a human factor intelligent driving behavior prediction method of the present application.
[0056] Figure 6 is a schematic diagram of modules of a human factor intelligent driving behavior prediction system of the present application.
[0057] DESCRIPTION OF REFERENCE NUMERALS:
[0058] 1. Physiological signal acquisition module; 2. Transformation module; 3. Multi-period decomposition module; 4. Spatial expansion module; 5. Prediction module; 6. Data fusion module; 7. Feature analysis module; 8. Behavior interpretation and reasoning module. Detailed implementation manners
[0059] The following further describes the present application in detail with reference to the appended Figure 1-6 drawings.
[0060] An embodiment of the present application discloses a human factor intelligent driving behavior prediction method. As Figure 1 shown, it includes the following steps:
[0061] S101. Obtain the physiological signals of the driver;
[0062] S102. Perform a fast Fourier transform on the physiological signals to generate corresponding amplitude-frequency characteristics, and obtain the acquisition frequencies that meet the preset amplitude-frequency selection criteria in the amplitude-frequency characteristics;
[0063] S103. Perform multi-period decomposition on the physiological signals according to the period of the acquisition frequencies to generate corresponding data decomposition result samples;
[0064] S104. Perform two-dimensional spatial expansion on the data decomposition result samples respectively according to the multi-source time-series data encoding layer to generate corresponding two-dimensional spatial data;
[0065] S105. Predict the target continuous frames corresponding to the vehicle road scene video according to the vehicle road scene video frame prediction layer to generate corresponding iterative predicted future frames;
[0066] S106. Perform a merging operation on the two-dimensional spatial data, the target continuous frames, and the iterative predicted future frames according to the multi-modal synchronization data fusion layer to generate corresponding multi-scale three-dimensional features;
[0067] S107. Perform feature analysis processing on the multi-scale three-dimensional features according to the three-dimensional backbone network layer to generate corresponding target output features; S108. Perform analysis processing on the target output features respectively according to the driving behavior interpretation layer and the driving behavior reasoning layer to generate corresponding driving behavior description information and driving behavior reasoning information.
[0068] In step S101, the physiological signals include electrocardiogram signals, electromyogram signals, and electroencephalogram signals. The electrocardiogram signal can be marked as X1, the electromyogram signal as X2, and the electroencephalogram signal as X3. Among them, X3 includes N acquisition channels, and the acquisition frequencies and dimensions of the N acquisition channels are the same as those of X1, where N≥1.
[0069] Specifically, the electrocardiogram signal X1 is the electrical activity signal of the driver's heart collected by electrodes, which can provide information about heart rhythm, heart rate variability, etc., and is of great significance for evaluating the driver's cardiovascular health status and mental state.
[0070] Secondly, the electromyogram signal is the electrical activity signal of the driver's muscles collected by electrodes. It can reflect the contraction and relaxation of the driver's muscles and plays an important role in studying the driver's muscle fatigue, fine motor control, etc. The electroencephalogram signal X3 is the electrical activity signal of the driver's brain collected by electrodes. The electroencephalogram signal X3 can provide information about the driver's cognition, attention, emotion, etc., and is of great significance for studying the driver's cognitive state and workload.
[0071] Among them, by setting the acquisition frequencies and dimensions of the N acquisition channels of X3 to be the same as those of X1, the electroencephalogram signal and the electrocardiogram signal can be conveniently compared and fused. In this way, the association and similarity between the two signals can be observed more directly, which helps to reveal the mechanism of the driver's heart-brain interaction.
[0072] In step S102, the above-mentioned X1, X2, and X3 can be respectively subjected to fast Fourier transform to generate the amplitude-frequency characteristics corresponding to X1, X2, and X3 respectively, and the acquisition frequencies that meet the preset amplitude-frequency selection criteria in the amplitude-frequency characteristics corresponding to X1 are marked as K 1m , the acquisition frequencies that meet the preset amplitude-frequency selection criteria in the amplitude-frequency characteristics corresponding to X2 are marked as K 2m , the acquisition frequencies that meet the preset amplitude-frequency selection criteria in the amplitude-frequency characteristics corresponding to X3 are marked as K 3Nm , m≥1;
[0073] Specifically, for the above-mentioned given X1, X2, and X3, the fast Fourier transform (FFT) can be respectively performed on each of their channels to obtain the corresponding amplitude-frequency characteristics. The amplitude-frequency characteristics refer to the situation where the amplitude of the signal changes with frequency in the frequency domain.
[0074] Secondly, the preset amplitude-frequency selection criteria for X1 are to select the first M frequencies with the largest amplitudes corresponding to X1, which are respectively denoted as K 11 , K 12 ... K 1m , the preset amplitude-frequency selection criteria for X2 are to select the first M frequencies with the largest amplitudes corresponding to X2, which are respectively denoted as K 21 , K 22 ... K 2m , the preset amplitude-frequency selection criteria for X3 are to select the first M frequencies with the largest amplitudes corresponding to each component of X3, which are respectively denoted as: including the first M acquisition frequencies K corresponding to the first acquisition channel 311 , K 312 ... K 31m , until the first M acquisition frequencies K corresponding to the last channel 3N1 , K 3N2 ... K 3Nm .
[0075] Specifically, the criteria for the first M frequencies with the largest amplitudes is based on the spectrum amplitude obtained after the FFT transform, and then the spectrum obtained after the FFT transform is sorted by amplitude, and then the first M frequencies with the largest amplitudes are selected. Among them, after the FFT transform, the amplitude value of the spectrum represents the amplitude of each frequency component in the signal. By sorting the spectrum by amplitude, the first M frequencies with the largest amplitudes can be found, which means that these frequency components have the largest amplitudes in the signal.
[0076] Furthermore, the frequency with the largest amplitude selected by the above preset amplitude frequency selection standard can provide information about the main components of the signal in the frequency domain. By analyzing these frequencies, the frequency characteristics and main amplitude distribution of the signal can be understood, further revealing the characteristics and properties of the signal. It should be noted that the number M of the first M frequencies selected can be set according to actual needs to adapt to the analysis of specific application scenarios.
[0077] For example, the spectrum of X1 obtained by FFT transformation is as follows: frequency: 10Hz, 20Hz, 30Hz, 40Hz and 50Hz, amplitude: 20, 30, 15, 25, 10, among which the frequency and amplitude correspond one to one according to the sorting, and then the corresponding frequencies are sorted according to the size of the above amplitude to obtain 20Hz, 40Hz, 10Hz, 30Hz, 50Hz. If the setting value of M in the preset amplitude frequency selection standard of X1 is 3, 20Hz, 40Hz, and 10Hz are selected as the target acquisition frequencies.
[0078] In step S103, the multi-period decomposition specifically includes: 1m The corresponding period T 1m Perform multi-period decomposition on X1 and generate the corresponding data decomposition result sample R X1 , according to K 2m The corresponding period T 2m Perform multi-period decomposition on X2 and generate the corresponding data decomposition result as R X2 , according to K 3Nm The corresponding period T 3Nm Perform multi-period decomposition on X3 and generate the corresponding data decomposition result as R X3 .
[0079] Specifically, the data decomposition result sample R X1 , R X2 and R X3 It is obtained by decomposing the original signals X1, X2 and X3 into components of different periods. The data in each sample represents the contribution of the signal component of the corresponding period in the original signal. Therefore, the sample R X1 , R X2 and RX3 Help to understand the characteristics and contributions of different periodic components in the original signal. They provide perspectives for different periodic analyses of the signal and can be used to identify and analyze the periodic behavior and periodic components in the signal.
[0080] For example, the electrocardiogram signal X1 contains components of multiple periods. Select a period of T 1m to perform multi-period decomposition on X1 and generate a corresponding data decomposition result sample R X1 . Among them, the period of X1 is 10 sampling points, and T 1m = 5 sampling points of K 1m is selected for decomposition. Through multi-period decomposition, 5 data decomposition result samples R X1 are obtained, corresponding to 5 different periodic components respectively.
[0081] Furthermore, the data decomposition result sample R X1 is specifically as follows. R X1 -1: [0, 0, 1, 1, 0], R X1 -2: [0, 1, 1, 0, 0], R X1 -3: [1, 1, 0, 0, 0], R X1 -4: [1, 0, 0, 0, 1], R X1 -5: [0, 0, 0, 1, 1]. These data decomposition result samples R X1 represent the characteristics of the signal X1 under different periodic components. The data in each sample represents the contribution of the corresponding periodic component in the original signal. By observing these samples, the presence and characteristics of different periodic components in the signal X1 can be understood.
[0082] In step S104, according to the multivariate time series data encoding layer, the above-mentioned R X1 , R X2 and R X3 are respectively extended in two-dimensional space to generate and label the two-dimensional space data corresponding to R X1 as P1, the two-dimensional space data corresponding to R X2 as P2, and the two-dimensional space data corresponding to R X3 as P 3N .
[0083] Specifically, through the multivariate time series data encoding layer, the obtained data decomposition result samples R X1 , R X2 and R X3 are subjected to data dimensionality increase, that is, each of the above data decomposition result samples is extended from one-dimensional space to two-dimensional space.
[0084] Among them, the data decomposition result samples R X1 , R X2and R X3 For each element value of X3 , corresponding its position in the two - dimensional plane to the time axis as height or brightness information in the two - dimensional space.
[0085] For example, for the data decomposition result sample R X1 , expand it into a two - dimensional space data P1, where each element represents the presence or absence of the original signal X1 in the corresponding period. Then, P1 can be plotted as a two - dimensional image, and different colors or brightness are used on this image to represent the presence or absence of the signal.
[0086] Secondly, by observing the images of P1, P2 and P 3N , it is possible to intuitively understand the presence or absence of signals X1, X2 and X3 in different periods, as well as their timing characteristics. At the same time, by marking different colors or brightness, different periodic components can be distinguished and explained.
[0087] In step S105, if the input of the current model is only two consecutive frames I i and I i+1 , then the target consecutive frames I i and I i+1 corresponding to the vehicle road scene video can be predicted according to the vehicle road scene video frame prediction layer, and the corresponding iterative predicted future frame I i+n is generated, where i≥1, n>1, and the video prediction model is: I i+2 =F θ (I i ,I i+1 ), θ is the set of all trainable model parameters, and the vehicle road scene video frame prediction model makes the difference between the predicted next frame I i+2 and the actually existing next - frame video frame I i+2 in the dataset reach the minimum value.
[0088] Specifically, the vehicle road scene video frame prediction layer is used to predict the vehicle road scene video, and the video signal of the road in front of the vehicle can be collected by using a camera. Among them, when using the current two video frames I i and I i+1 as the input, the predicted next frame I i+2 =F θ (I i ,I i+1 ) can be calculated through the video prediction model to obtain the predicted next frame I i+2 , and the future frames I i+3 and I i+4 are iteratively predicted.
[0089] Among them, in the above prediction process, in order to satisfy the condition that by adjusting the model parameter θ, the difference between the predicted next frame Ii+2 and the next video frame Ii+2 that actually exists in the dataset reaches the minimum value, the real next video frame I in the training dataset can be used i+2 and the predicted next video frame I i+2 The difference between them is used as the loss function for optimization.
[0090] Secondly, by minimizing the loss function, the model parameter θ can be adjusted to make the prediction result closer to the real next video frame. Through continuous iterative training and optimization, the accuracy and generalization ability of the prediction model can be gradually improved, minimizing the difference between the predicted next video frame Ii+2 and the real next video frame.
[0091] In step S106, according to the multi-modal synchronous data fusion layer, the above P1, P2, P 3N , the target consecutive frames I i and I i+1 as well as the iterative predicted future frame I i+n Perform a merging operation to generate corresponding multi-scale three-dimensional features.
[0092] Specifically, according to the data types and characteristics of different vehicle driver physiological signals and road scene video prediction signals, different merging methods can be selected. For example, for two-dimensional spatial data such as P1, P2, P 3N , they can be superimposed in the two-dimensional space to generate a new two-dimensional space data. For the target consecutive frames I i , I i+1 as well as the iterative predicted future frame I i+n , they can be superimposed on the time axis to generate a new time series data.
[0093] Furthermore, by fusing data of different types and characteristics in the above way, a multi-scale three-dimensional feature is generated. The multi-scale three-dimensional feature contains the spatial, temporal, and modal information of the original data and can represent the characteristics of the original data more comprehensively.
[0094] In practical applications, the obtained multi-scale three-dimensional feature can be used as the input for subsequent data analysis and processing. For example, the multi-scale three-dimensional feature can be used for machine learning or deep learning training to extract deeper features and information.
[0095] In step S107, the three-dimensional Transformer backbone network layer can be used to perform feature analysis processing on the obtained multi-scale three-dimensional feature to generate corresponding target output features.
[0096] Specifically, the 3D Transformer backbone network layer is a deep learning model for processing 3D data. It is an extension based on the Transformer model and is particularly suitable for processing 3D data with spatio-temporal structures.
[0097] In practical applications, the 3D Transformer first encodes the above-mentioned input multi-scale 3D features through the self-attention mechanism to capture the dependencies between features. Then, the encoded features are further processed through multiple layers of fully connected networks and normalization layers.
[0098] Among them, through the feature extraction of multi-modal synchronous data by the 3D Transformer backbone network layer, corresponding compressed features are obtained. Compared with the traditional end-to-end intelligent cockpit model, this reduces the time consumption brought by each task sub-module.
[0099] For example, in an intelligent cockpit system, the process of data processing and analysis usually involves multiple task sub-modules, such as object detection, trajectory prediction, behavior classification, etc. The traditional end-to-end intelligent cockpit model usually needs to perform feature extraction separately for each task sub-module, which will increase the time consumption of the entire system to a certain extent. In contrast to the above processing and analysis, using the 3D Transformer backbone network layer can perform unified feature extraction on data from different modalities (such as image, radar, lidar data, etc.), enabling different task sub-modules to share the same feature representation and reducing the time consumption of performing multiple feature extractions on the same data.
[0100] Secondly, the 3D Transformer backbone network layer can also compress the extracted features, that is, reduce the dimension of the features while retaining key information. This can further reduce the computational amount of subsequent task sub-modules, thereby reducing the time overhead of the entire system.
[0101] Furthermore, the result of feature analysis and processing is a target output feature, which contains the deep information of the original input features. This target output feature can be used for subsequent tasks, such as classification, regression, or prediction.
[0102] In step S108, for the driving behavior interpretation layer, its main objective is to interpret the above-mentioned target output feature and generate corresponding driving behavior description information. This includes mapping the driver's behaviors (such as accelerating, decelerating, steering, etc.) to specific behavior features.
[0103] Among them, the previous process involves a series of data processing and analysis operations, including feature extraction, feature selection, feature mapping, etc. Through these operations, the driving behavior interpretation layer can extract meaningful information from the original behavior features, so as to better understand the driver's behavior.
[0104] Secondly, the main goal of the driving behavior reasoning layer is to reason about the target output features and generate corresponding driving behavior reasoning information. This includes predicting the driver's future behavior, such as the driver may accelerate, decelerate or turn next. This process may involve a series of complex data processing and analysis operations, including feature extraction, feature selection, feature mapping, model training, model prediction, etc. Through these operations, the driving behavior reasoning layer can predict the driver's future behavior from the original behavior features, so as to better predict the driver's behavior.
[0105] For example, the target output features are the physiological signals of the driver and the multi-scale three-dimensional features predicted by the road scene video, which are input to the driving behavior interpretation layer. Among them, the physiological signal shows that the driver's heart rate is rising, and the features predicted by the road scene video indicate that there is an emergency braking situation ahead. Combining the above information, the driving behavior interpretation layer immediately outputs the corresponding driving behavior description information as "the driver may have seen the emergency ahead and felt nervous".
[0106] Another example is that the target output features are the physiological signals of the driver and the multi-scale three-dimensional features predicted by the road scene video, which are input to the driving behavior reasoning layer. Among them, the video prediction features show that there is an emergency braking situation ahead and the driver's heart rate is rising. Then the driving behavior reasoning layer may reason that "the driver may make an emergency brake".
[0107] Specifically, as Figure 2 shown, it is the functional module logic block diagram of the solution of this application.
[0108] The human - factor intelligent driving behavior prediction method provided in this embodiment collects and analyzes the physiological data of vehicle drivers, which can more accurately understand the physiological state of the drivers. At the same time, combined with the prediction and analysis data of vehicle road scene video frames, that is, introducing a vehicle road scene video frame prediction layer, it can be used for event prediction in the vehicle driving environment. Human - factor intelligent driving behavior prediction is carried out under the condition that the event has not occurred, rather than waiting for the event to occur and then carrying out classification and prediction. Furthermore, the real - time performance of the overall vehicle behavior prediction is improved. Then, after fusing the physiological state of the above - mentioned driver and the multi - modal synchronous data corresponding to the vehicle road prediction, feature extraction of the multi - modal synchronous data is carried out through a three - dimensional backbone network layer to obtain corresponding compressed target output features. Different from the traditional end - to - end intelligent cockpit model, it reduces the time consumption brought by each task sub - module. Secondly, introducing a human - factor intelligent driving behavior explanation layer and a human - factor intelligent driving behavior reasoning layer to analyze the obtained target output features can explain the behavior of the vehicle on the road and explain the reasons behind the behavior taken by the vehicle. Through the data extraction and analysis of the above - mentioned algorithm logic layer, the overall prediction effect of driving behavior is improved.
[0109] In one implementation manner of this embodiment, the size of the multi - scale three - dimensional feature is H×W×(2 + N + 5×3), and the number of channels is (2 + N + 5×3), where H is the height of the corresponding feature map in the multi - scale three - dimensional feature, and W is the width of the corresponding feature map in the multi - scale three - dimensional feature.
[0110] Among them, in the number of channels (2 + N + 5×3), 2 can represent two frames of images I i and I i+1 applicable to the vehicle road scene video frame prediction layer, N can represent two - dimensional spatial data P1, P2, …, P N , and 5×3 can represent the iteratively generated images I i+2 、I i+3 and I i+4 .
[0111] Specifically, generated through the final setting of the size of the above - mentioned multi - scale three - dimensional feature, it can fuse information from different sources to obtain a richer and more complete feature representation, which helps the model better understand and interpret the input data.
[0112] The human - factor intelligent driving behavior prediction method provided in this implementation manner increases the number of channels according to the operation of the multi - modal synchronous data fusion layer, enabling the model to learn more types of features. Thus, when facing different input data, it can maintain good recognition ability and improve the generalization ability of the model.
[0113] In one implementation of this embodiment, according to the preset selection rule, when selecting the number of channels N, the number of channels is selected as an integer multiple of 6. Among them, the preset selection rule is a specific data selection rule set when designing or constructing the multi-temporal data encoding layer, and is used to determine the values of certain parameters or settings.
[0114] Specifically, since each group of features, that is, two-dimensional spatial data or iteratively generated images, consists of 6 channels or parameters, every 6 channels correspond to a specific feature set, such as position, speed, acceleration, direction, etc. Setting the number of channels to a multiple of 6 according to the preset selection rule can ensure that each feature set is completely included in the output features, and the boundaries between each feature set are clear, so as to facilitate separation and parsing, thereby improving the analysis and calculation efficiency of the output data.
[0115] For the human factor intelligent driving behavior prediction method provided by this implementation, in the deep learning algorithm process, if the number of channels is set to be divisible by 6, then the advantage of parallel computing can be utilized to better use the network architecture in the swin Transform, thereby improving the computing efficiency.
[0116] In one implementation of this embodiment, the multi-scale three-dimensional feature is decomposed into H×W×3×((2 + N + 5×3) / 3), where the first three dimensions redefine each frame in the multi-scale three-dimensional feature, and each frame contains H×W×3 pixels.
[0117] Specifically, the first three dimensions H×W×3 define the size and format of each frame, that is, each frame contains H×W pixels, and each pixel consists of three channels (RGB), and thus can intuitively represent the spatial structure and color information of the image. The last dimension (2 + N + 5×3) / 3 represents the number of frames. In this way, the original multi-scale three-dimensional feature can be expanded into a series of frames, and each frame is an H×W×3 image.
[0118] Among them, through the above decomposition operation, it helps to better understand and analyze the feature space. In particular, the spatial structure and color information of each frame can be intuitively seen, so that it is easier to understand and use these features. In addition, this decomposition operation can also facilitate subsequent processing and operations. For example, various processing algorithms for two-dimensional images (such as convolution, pooling, etc.) can be directly used to process each frame.
[0119] The human factor intelligent driving behavior prediction method provided by this implementation can better organize and process data. Decomposing the multi-scale three-dimensional feature into smaller data blocks can improve the data processing efficiency.
[0120] In one implementation of this embodiment, as Figure 3As shown, step S107 is to perform feature analysis processing on the multi-scale three-dimensional features according to the three-dimensional backbone network layer to generate corresponding target output features, including the following steps:
[0121] S201. Obtain the corresponding three-dimensional feature segmentation rules in the three-dimensional backbone network layer;
[0122] S202. According to the three-dimensional feature segmentation rules, divide the multi-scale three-dimensional features into H / 4×W / 4×((2 + N + 5×3) / 6) sub-features.
[0123] In steps S201 to S202, as can be seen from the above, the three-dimensional backbone network layer can be specifically set as a three-dimensional Transformer backbone network layer, and the three-dimensional Transformer backbone network layer includes operations corresponding to the three-dimensional feature segmentation rules. Among them, the three-dimensional feature segmentation operation is a method of segmenting features in three-dimensional space, which is usually used in applications in the fields of computer vision and machine learning. The goal of this operation is to divide a three-dimensional feature space (for example, a three-dimensional point cloud captured by a depth camera, or a three-dimensional feature map generated by a certain machine learning model) into multiple independent parts or regions, and each part or region usually represents an independent object or a part of a scene.
[0124] Specifically, the above process specifically includes feature extraction: extracting useful features from the original data (such as three-dimensional point cloud or image). These features may include color, texture, shape, depth, etc.; feature space construction: constructing a three-dimensional feature space according to the extracted features. In this feature space, the position of each point represents its corresponding feature value; feature segmentation: in the feature space, dividing the feature space into multiple parts or regions according to a certain criterion (such as the similarity or continuity of feature values).
[0125] In this embodiment, the above three-dimensional feature segmentation operation defines a three-dimensional block with a size of 4×4×3×2 as a sub-feature. Therefore, the multi-scale three-dimensional features can be divided into H / 4×W / 4×((2 + N + 5×3) / 6) sub-features through the three-dimensional feature segmentation operation.
[0126] Specifically, H / 4×W / 4 means that in the spatial dimensions (i.e., height H and width W), each dimension is divided into 1 / 4 of the original. This means that each original feature is now divided into 16 (= 4×4) sub-features. This operation can improve the spatial resolution and thus be able to describe each object or scene more meticulously. Secondly, (2 + N + 5×3) / 6 means that in the channel dimension, each original feature is divided into (2 + N + 5×3) / 6 sub-features. This means that more attributes or features can be described for each object or scene. Among them, the dimension of each sub-feature is 4×4×3×2 = 96.
[0127] The human factor intelligent driving behavior prediction method provided in this embodiment can reduce or increase the dimension of the sub-features through linear mapping. Dimension reduction can reduce the number of model parameters and computational complexity, improve computational efficiency, and dimension increase can introduce more feature expression dimensions and improve the expression ability of the model.
[0128] In one implementation of this embodiment, Figure 4 As shown, after step S202, that is, dividing the multi-scale three-dimensional features into H / 4×W / 4×((2+N+5×3) / 6) sub-features according to the three-dimensional feature segmentation rule, the following steps are also included:
[0129] S301. Obtain the corresponding linear coding rules in the three-dimensional backbone network layer;
[0130] S302. According to the linear coding rule, each sub-feature is linearly mapped to a vector C, where the vector C has any dimension.
[0131] In step S301 to step S302, the three-dimensional Transformer backbone network layer also includes related operations corresponding to the linear coding rule. Among them, the linear coding operation is a dimensionality reduction or compression operation, which can convert the input high-dimensional features into low-dimensional features while retaining the effective information in the original features as much as possible. This operation is usually implemented by one or more linear transformations (for example, fully connected layers, convolutional layers, etc.).
[0132] In this embodiment, each token obtained after the three-dimensional feature segmentation operation is linearly mapped to a vector C according to the above linear encoding operation, and the dimension of the vector C can be any dimension. Specifically, after the three-dimensional feature segmentation operation, a series of sub-features can be obtained, and then through the linear encoding operation, these sub-features can be linearly mapped to a new vector space, which is called vector C.
[0133] Among them, vector C is the result of the linear encoding operation, which represents the information of the original sub-feature in the new vector space. The dimension of vector C can be arbitrary, depending on the output dimension selected when designing the linear encoding operation. Choosing an appropriate dimension can reduce computational complexity and memory consumption while retaining sufficient information.
[0134] Secondly, linear mapping is an operation that maps an input vector to an output vector, which satisfies the distributive law of addition and scalar multiplication. In this process, the original high-dimensional sub-features are mapped to a low-dimensional vector C. This mapping can be achieved through a fully connected layer (or other linear transformations).
[0135] The human factor intelligent driving behavior prediction method provided by this embodiment maps high-dimensional features to a low-dimensional vector C, which helps reduce computational complexity and thus improve the analysis and calculation efficiency of the data model.
[0136] In an implementation manner provided by this embodiment, as Figure 5 shown, the three-dimensional backbone network layer includes a self-attention encoding rule. Step S107, that is, according to the three-dimensional backbone network layer, performing feature analysis and processing on multi-scale three-dimensional features to generate corresponding target output features includes the following steps:
[0137] S1. Perform a spatial sampling on the multi-scale three-dimensional features and output the corresponding first target feature;
[0138] S2. Perform a Video Swin Transformer blocks operation on the multi-scale three-dimensional features and output the corresponding second target feature. The MPL layer corresponding to the Video Swin Transformer blocks operation in the model is a 1×1 convolutional layer, and the number of convolutional kernels is equal to the dimension of the sub-features input to the model;
[0139] S3. Repeat S1 and S2;
[0140] S4. Repeat S3, and the number of repetitions is K times, where K is a preset positive integer.
[0141] In step S1, the spatial sampling can be regarded as a downsampling operation of data, which can reduce the complexity of data by reducing the spatial resolution of the data.
[0142] In step S2, Video Swin Transformer is a Transformer-based model that can process video data. Here, the MPL layer in the Video Swin Transformer blocks operation is a 1×1 convolutional layer, and the number of convolutional kernels is equal to the dimension of the input sub-features, which means that each input feature will have a corresponding convolutional kernel to process it.
[0143] In step S3, repeat S1 and S2. In this way, the three-dimensional features can be processed at different spatial scales, so as to obtain a richer and more refined feature representation.
[0144] Specifically, by performing S1 and S2 in each iteration, the model can extract information from the original features multiple times, and the information extracted each time may be different, which can increase the information acquisition ability of the model.
[0145] In step S4, this step can be regarded as an iterative process. Each iteration processes the three-dimensional features at different spatial scales. Through multiple iterations, the model can learn richer and more complex feature representations from the three-dimensional features at different scales. Among them, the value of the execution times K depends on the specific model design and application scenario settings.
[0146] Specifically, each time step S3 is repeatedly executed, the model performs spatial sampling and VideoSwin Transformer blocks operations on the input three-dimensional features, which helps the model extract richer and more complex features from different angles and scales. In addition, the execution times K actually also determines the depth of the model. In deep learning, the depth of the model is usually proportional to the complexity and abstraction level of the features it can learn. Therefore, increasing the value of K can enable the model to learn more complex features, thereby improving the performance of the model.
[0147] This embodiment provides a human factor intelligent driving behavior prediction method. Performing spatial sampling can change the size of the multi-scale three-dimensional features to half of the original, and at the same time expand the number of channels of the multi-scale three-dimensional features to twice the original. And performing Video Swin Transformer operations can reduce the model parameters, thereby improving the inference speed of the model.
[0148] This application embodiment discloses a human factor intelligent driving behavior prediction system, as Figure 6 shown, including:
[0149] A physiological signal acquisition module 1, configured to acquire the physiological signals of the driver;
[0150] A transformation module 2, configured to perform a fast Fourier transform on the physiological signals to generate corresponding amplitude-frequency characteristics, and acquire the acquisition frequencies that meet the preset amplitude-frequency selection criteria in the amplitude-frequency characteristics;
[0151] A multi-period decomposition module 3, configured to perform multi-period decomposition on the physiological signals according to the periods of the acquisition frequencies to generate corresponding data decomposition result samples;
[0152] A spatial expansion module 4, configured to perform two-dimensional spatial expansion on the data decomposition result samples respectively according to the multivariate time series data encoding layer to generate corresponding two-dimensional spatial data;
[0153] A prediction module 5, configured to predict the target continuous frames corresponding to the vehicle road scene video according to the vehicle road scene video frame prediction layer to generate corresponding iterative predicted future frames;
[0154] A data fusion module 6, configured to perform a merging operation on the two-dimensional spatial data, the target continuous frames, and the iterative predicted future frames according to the multi-modal synchronous data fusion layer to generate corresponding multi-scale three-dimensional features;
[0155] A feature analysis module 7, configured to perform feature analysis processing on multi-scale three-dimensional features according to a three-dimensional backbone network layer to generate corresponding target output features;
[0156] A behavior interpretation and reasoning module 8, configured to perform analysis processing on the target output features according to a driving behavior interpretation layer and a driving behavior reasoning layer respectively to generate corresponding driving behavior description information and driving behavior reasoning information.
[0157] The human factor intelligent driving behavior prediction system provided in this embodiment can collect and analyze the physiological data of vehicle drivers more accurately according to the physiological signal acquisition module 1, the transformation module 2, the multi-period decomposition module 3, and the spatial expansion module 4. At the same time, through the prediction module 5 combined with the prediction analysis data of vehicle road scene video frames, that is, introducing a vehicle road scene video frame prediction layer, it can be used for event prediction in the vehicle driving environment, and perform human factor intelligent driving behavior prediction under the condition that the event has not occurred, rather than waiting for the event to occur and then performing classification and prediction, thereby improving the real-time performance of the overall vehicle behavior prediction. Then, after fusing the above-mentioned physiological state of the driver and the multi-modal synchronization data corresponding to the vehicle road prediction through the data fusion module 6, the feature extraction of the multi-modal synchronization data is performed through the three-dimensional backbone network layer in the feature analysis module 7 to obtain corresponding compressed target output features. Different from the traditional end-to-end intelligent cockpit model, it reduces the time consumption brought by each task sub-module. Secondly, by introducing a human factor intelligent driving behavior interpretation layer and a human factor intelligent driving behavior reasoning layer in the behavior interpretation and reasoning module 8 to analyze the obtained target output features, it can explain the behavior of the vehicle on the road and explain the reasons behind the behavior taken by the vehicle. Through the data extraction and analysis of the above algorithm logic layer, the overall prediction effect of driving behavior is improved.
[0158] It should be noted that the human factor intelligent driving behavior prediction system provided in the embodiment of the present application further includes each module and / or corresponding sub-module corresponding to the logical functions or logical steps of any of the above human factor intelligent driving behavior prediction methods, achieving the same effect as each logical function or logical step, and will not be repeated here specifically.
[0159] The embodiment of the present application also discloses a terminal device, including a memory, a processor, and computer instructions stored in the memory and capable of running on the processor. When the processor executes the computer instructions, it adopts any of the human factor intelligent driving behavior prediction methods in the above embodiment.
[0160] Among them, the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server. Moreover, the terminal device includes, but is not limited to, a processor and a memory. For example, the terminal device may further include input / output devices, network access devices, and a bus, etc.
[0161] Among them, the processor can adopt a central processing unit (CPU). Of course, according to the actual usage situation, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc. This application does not make any restrictions on this.
[0162] Among them, the memory can be an internal storage unit of the terminal device. For example, the hard disk or memory of the terminal device. It can also be an external storage device of the terminal device. For example, a plug-in hard disk, a smart media card (SMC), a secure digital card (SD), or a flash card (FC), etc. equipped on the terminal device. Moreover, the memory can also be a combination of the internal storage unit and the external storage device of the terminal device. The memory is used to store computer instructions and other instructions and data required by the terminal device. The memory can also be used to temporarily store the data that has been output or will be output. This application does not make any restrictions on this.
[0163] Among them, through this terminal device, any one of the human factor intelligent driving behavior prediction methods in the above embodiments is stored in the memory of the terminal device, and is loaded and executed on the processor of the terminal device, which is convenient for use.
[0164] This application embodiment also discloses a computer-readable storage medium. Moreover, the computer-readable storage medium stores computer instructions. Among them, when the computer instructions are executed by the processor, any one of the human factor intelligent driving behavior prediction methods in the above embodiments is adopted.
[0165] Among them, the computer instructions can be stored in the computer-readable medium. The computer instructions include computer instruction codes. The computer instruction codes can be in the form of source code, object code, executable files, or some middleware forms, etc. The computer-readable medium includes any entity or device, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. that can carry the computer instruction codes. It should be noted that the computer-readable medium includes, but is not limited to, the above components.
[0166] Among them, through this computer-readable storage medium, any one of the human factor intelligent driving behavior prediction methods in the above embodiments is stored in the computer-readable storage medium, and is loaded and executed on a processor to facilitate the storage and application of the above methods.
[0167] The above are all the preferred embodiments of this application, and the protection scope of this application is not limited accordingly. Therefore, any equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.
Claims
1. A human factor intelligent driving behavior prediction method, characterized in that, It includes the following steps: Obtain the physiological signals of the driver; Perform a fast Fourier transform on the physiological signals to generate corresponding amplitude-frequency characteristics, and obtain the acquisition frequencies that meet the preset amplitude-frequency selection criteria in the amplitude-frequency characteristics; Perform multi-period decomposition on the physiological signals according to the period of the acquisition frequencies to generate corresponding data decomposition result samples; Perform two-dimensional spatial expansion on the data decomposition result samples respectively according to the multi-variable time series data encoding layer to generate corresponding two-dimensional spatial data; Predict the target consecutive frames corresponding to the vehicle road scene video according to the vehicle road scene video frame prediction layer to generate corresponding iterative predicted future frames; Perform a merging operation on the two-dimensional spatial data, the target consecutive frames, and the iterative predicted future frames according to the multi-modal synchronous data fusion layer to generate corresponding multi-scale three-dimensional features; Perform feature analysis processing on the multi-scale three-dimensional features according to the three-dimensional backbone network layer to generate corresponding target output features; Perform analysis processing on the target output features respectively according to the human factor intelligent driving behavior interpretation layer and the human factor intelligent driving behavior reasoning layer to generate corresponding human factor intelligent driving behavior description information and human factor intelligent driving behavior reasoning information.
2. The human factor intelligent driving behavior prediction method according to claim 1, wherein The performing feature analysis processing on the multi-scale three-dimensional features according to the three-dimensional backbone network layer to generate corresponding target output features includes the following steps: Obtain the corresponding three-dimensional feature segmentation rules in the three-dimensional backbone network layer; Divide the multi-scale three-dimensional features into H / 4×W / 4×((2 + N + 5×3) / 6) sub-features according to the three-dimensional feature segmentation rules.
3. The human factor intelligent driving behavior prediction method according to claim 2, wherein, After dividing the multi-scale three-dimensional features into H / 4×W / 4×((2 + N + 5×3) / 6) sub-features according to the three-dimensional feature segmentation rules, the following steps are further included: Obtain the corresponding linear encoding rules in the three-dimensional backbone network layer; Linearly map each sub-feature to a vector C according to the linear encoding rules, and the vector C has any dimension.
4. The human factor intelligent driving behavior prediction method according to claim 2, wherein The three-dimensional backbone network layer includes self-attention encoding rules. The performing feature analysis processing on the multi-scale three-dimensional features according to the three-dimensional backbone network layer to generate corresponding target output features includes the following steps: S1. Perform one spatial sampling on the multi-scale three-dimensional features and output corresponding first target features; S2. Perform one Video Swin Transformer blocks operation on the multi-scale three-dimensional features and output corresponding second target features. The MPL layer corresponding to the Video Swin Transformer blocks operation in the model is a 1×1 convolutional layer, and the number of convolution kernels is equal to the dimension of the sub-features input to the model; S3. Repeat S1 and S2; S4. Repeat S3, and the number of repetitions is K times, where K is a preset positive integer.
5. A human factor intelligent driving behavior prediction method according to claim 1, characterized in that, The size of the multi-scale three-dimensional features is H×W×(2 + N + 5×3), and the number of channels is (2 + N + 5×3), where H is the height of the feature map corresponding to the multi-scale three-dimensional features, and W is the width of the feature map corresponding to the multi-scale three-dimensional features.
6. The human factor intelligent driving behavior prediction method according to claim 5, wherein, According to the preset selection rule, when selecting N among the number of channels, the number of channels is selected as an integer multiple of 6.
7. The human factor intelligent driving behavior prediction method according to claim 5, characterized in that, The multi-scale three-dimensional feature is decomposed into H×W×3×((2 + N + 5×3) / 3), where the first three dimensions redefine each frame in the multi-scale three-dimensional feature, and each frame contains H×W×3 pixels.
8. A human factor intelligent driving behavior prediction system, characterized in that, Including: A physiological signal acquisition module (1) for acquiring the physiological signal of the driver; A transformation module (2) for performing a fast Fourier transform on the physiological signal to generate a corresponding amplitude-frequency characteristic, and acquiring the acquisition frequency that meets the preset amplitude-frequency selection criterion in the amplitude-frequency characteristic; A multi-period decomposition module (3) for performing multi-period decomposition on the physiological signal according to the period of the acquisition frequency to generate a corresponding data decomposition result sample; A spatial expansion module (4) for performing two-dimensional spatial expansion on the data decomposition result sample respectively according to the multi-source time-series data coding layer to generate a corresponding two-dimensional spatial data; A prediction module (5) for predicting the target consecutive frames corresponding to the vehicle road scene video according to the vehicle road scene video frame prediction layer to generate a corresponding iterative predicted future frame; A data fusion module (6) for performing a merging operation on the two-dimensional spatial data, the target consecutive frames, and the iterative predicted future frames according to the multi-modal synchronous data fusion layer to generate a corresponding multi-scale three-dimensional feature; A feature analysis module (7) for performing feature analysis processing on the multi-scale three-dimensional feature according to the three-dimensional backbone network layer to generate a corresponding target output feature; A behavior interpretation and reasoning module (8) for performing analysis processing on the target output feature respectively according to the human factor intelligent driving behavior interpretation layer and the human factor intelligent driving behavior reasoning layer to generate corresponding human factor intelligent driving behavior description information and human factor intelligent driving behavior reasoning information.
9. A terminal device, comprising a memory and a processor, characterized in that, The memory stores computer instructions that can run on the processor. When the processor loads and executes the computer instructions, a human factor intelligent driving behavior prediction method as described in any one of claims 1 to 7 is adopted.
10. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are loaded and executed by the processor, a human factor intelligent driving behavior prediction method as described in any one of claims 1 to 7 is adopted.
Citation Information
Patent Citations
Dangerous driving prediction method and device, terminal equipment and storage medium
CN114670848A
Multi-mode human factor intelligent state recognition model creation and real-time state monitoring method and system and storage medium
CN115758097A
Cited By
Human-factor intelligent driving behavior prediction method and system, and terminal device and storage medium
EP4691877A1