An identification method and system for human-vehicle interaction
Through multi-sensor fusion technology and deep neural network architecture, the misidentification and delay problems of human-vehicle interaction recognition system in complex environments are solved, high-precision and efficient environmental perception are achieved, and the safety and intelligence of the vehicle are enhanced.
Patent Information
- Application Number
- CN202411350048.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-09-26
AI Technical Summary
In complex environments, the human-vehicle interaction recognition system has problems with misidentification or misidentification, difficulty in multi-target tracking and delays in dynamic scenarios, resulting in vehicle response deviations or accidents.
Multi-sensor fusion technology is adopted, including cameras, lidar, millimeter-wave radar and infrared sensors, and feature-level fusion is performed through deep neural network architecture to obtain the outside environment state, and improve data accuracy through data alignment and coordinate conversion.
It significantly improves the accuracy and real-time nature of object recognition, reduces the cognitive burden of drivers, enhances the safety of the vehicle, and improves the robustness and intelligence level of the system.
Smart Images

Figure CN119312276B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle - human interaction, and more specifically, it relates to a recognition method and system for vehicle - human interaction. Background Art
[0002] In vehicle - human interaction recognition, the recognition accuracy problem is a key challenge. Especially in complex environments, there are problems such as false recognition or missed recognition, difficulty in multi - target tracking, and latency in dynamic scenarios.
[0003] For false recognition or missed recognition, the system may mistake background objects for pedestrians or vehicles, causing the vehicle to make wrong reactions. For example, it may suddenly brake when there is no pedestrian, affecting traffic flow. If the system fails to recognize a real pedestrian or vehicle, it may fail to make timely braking or avoidance actions, thus leading to accidents, especially when a pedestrian suddenly appears in front of the vehicle. Difficulty in multi - target tracking means that the system may lose the position of a certain pedestrian or confuse two different targets, resulting in deviation in the vehicle's reaction. For example, the system may think that a pedestrian has left the lane when in fact the pedestrian is still in a dangerous area, or it may ignore some key targets, especially those that are occluded or small, such as children or bicycles, which may cause the vehicle to fail to avoid or brake in time. In addition, the acquisition times and coordinate systems of different sensors may be inconsistent, resulting in inaccurate information during data fusion. Therefore, how to design the recognition for vehicle - human interaction is an urgent problem to be solved. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a recognition method for vehicle - human interaction to achieve intelligent interaction between vehicles and humans.
[0005] To achieve the above - mentioned purpose, the present invention provides the following technical solutions: The recognition method for vehicle - human interaction includes:
[0006] Step S1: Collect the original data of b_c types of sensors;
[0007] Step S2: Align the original data of b_c types of sensors to obtain pre - processed data after alignment;
[0008] Step S3: Perform feature - level fusion on the pre - processed data of each sensor according to the deep neural network architecture to obtain the current environmental state outside the vehicle;
[0009] Step S4: Collect the current state data of the vehicle body, and perform warning and reminder display according to the current environmental state outside the vehicle;
[0010] Preferably, the b_c type sensors include cameras, lidars, millimeter-wave radars, and infrared sensors installed on the vehicle body. The cameras are used to provide high-resolution 2D images, the lidars are used to provide high-precision 3D point cloud data, the millimeter-wave radars are used to provide information on the distance, speed, and relative motion of objects, and the infrared sensors are used to provide temperature or thermal imaging data.
[0011] Preferably, the method for aligning the raw data of the b_c type sensors includes:
[0012] Step A1: Synchronize the time of the raw data of the b_c type sensors using the timestamp alignment method;
[0013] Step A2: Synchronize the space of the raw data of the b_c type sensors using coordinate transformation.
[0014] Preferably, the method for synchronizing the time of the raw data of the b_c type sensors using the timestamp alignment method includes:
[0015] Select one of the sensors as the base sensor, calculate the time delay between the base sensor and any other sensor, set the cycle time T and the delay hypothesis value set, the delay hypothesis value set includes preset time delay values, and form two time series A[T] and B[T] from the raw data of the base sensor and other sensors within the cycle time T;
[0016] Calculate the signal means μA and μB of A[T] and B[T], and decentralize the time series A[T] and B[T] to obtain A 1 [T] and B 1 [T], and calculate the cross-correlation function Assign τ to the values in the delay hypothesis value set and calculate. Select the value of τ when the cross-correlation function R AB (τ) is the largest as the actual delay time, and use the actual delay time to correct the time of the other sensors compared with the base sensor to obtain the corrected time series as B[T-τ];
[0017] Successively obtain the corrected time series of each sensor, and use interpolation, discarding, or repeating data to align the data of all sensors at the same time point.
[0018] Preferably, the method for synchronizing the space of the raw data of the b_c type sensors using coordinate transformation includes:
[0019] Set the base coordinate system, and the base coordinate system is the vehicle coordinate system or the world coordinate system;
[0020] Convert the camera coordinate system to the base coordinate system, obtain the 2D image captured by the camera, and add a coordinate system to form a 2D image coordinate system. Pixel points are mapped to the base coordinate system through the camera's internal parameter matrix and external parameter matrix;
[0021] Among them, the internal parameter matrix of the camera Among them, (f x , f y ) is the focal length, (c x , c y ) is the position of the optical center. The external parameter matrix of the camera is calibrated using a calibration board. By taking z_p calibration board images at different angles and using the camera calibration function to calculate the relationship between the pixel coordinates of the corner points in the calibration board image and the base coordinates, the external parameter matrix of the camera can be obtained;
[0022] The method for converting the infrared sensor coordinate system to the base coordinate system is the same as that of the camera;
[0023] Convert the lidar coordinate system to the base coordinate system, and convert the lidar data to the base coordinate system through the external parameter matrix and translation vector of the lidar;
[0024] Among them, the method for obtaining the external parameter matrix and translation vector of the lidar is as follows: Using a circular dot array as the calibration target, at different angles and positions, simultaneously collect data using the lidar and the camera, use point cloud registration technology to align the lidar point cloud data with the camera image data, obtain the external parameter matrix and translation vector, and optimize the external parameter matrix through the nonlinear least squares method;
[0025] The method for converting the millimeter-wave radar coordinate system to the base coordinate system is the same as that of the lidar.
[0026] Preferably, the method for performing feature-level fusion on the preprocessed data of each sensor includes:
[0027] Set a comprehensive sample set, convert the images captured by the camera into RGB images or grayscale images of standard size and perform normalization processing; Voxelize the lidar point cloud data to generate a bird's-eye view or depth map; Convert the millimeter-wave radar data into an image form; Convert the infrared sensor data into a heat map or grayscale image and perform normalization processing; All the collected data forms a comprehensive sample set;
[0028] Build a deep neural network architecture for multi-sensor fusion. Set the input of the input layer as the comprehensive sample set. The data of the camera and infrared sensor are used for feature extraction by the CNN network, and the data of the lidar and millimeter-wave radar are used for feature extraction by the 3D convolutional neural network. Then, the feature vectors of each sensor are obtained, and the weighted sum is used to fuse the feature vectors into a large overall vector. The large overall vector is input into the fully connected layer, and the output layer is set to output u_p classification tasks and u_k regression tasks. The classification tasks are used to identify objects of u_p categories, and the regression task outputs are used to predict the continuous attributes of the objects;
[0029] Integrate the output of the deep neural network architecture to obtain the current external environment state;
[0030] The classification tasks are activated using the Softmax function, and the regression tasks are activated using the linear activation function. The combined loss function The classification tasks use cross-entropy loss, and the regression tasks use mean squared error loss. Among them, ε(j) represents the weight of the jth classification task, CE_loss(j) represents the cross-entropy loss of the jth classification task, θ(u) represents the weight of the u-th classification task, MSE_loss(u) represents the mean squared error loss of the u-th regression task, and the Adam or SGD is used to optimize the model.
[0031] Preferably, the method for using the CNN network to extract features from the data of the camera and infrared sensor includes:
[0032] Set the input as the data of the camera and infrared sensor. The size of the original image is (H, W, C), where H is the height, W is the width, and C is the number of channels;
[0033] Use multiple convolutional layers to extract local features Use the ReLU activation function to introduce non-linearity to obtain the transformed output F_relu(k, h, m). Downsample the feature map through the pooling layer, and repeat the convolution and pooling operations to obtain the low-resolution feature map (H 11 , W 11 , C);
[0034] Among them, H 11 is the height, W 11 is the width, F_con(k, h, m) represents the feature value of the mth channel at the position (k, h) in the output feature map, I(k + p, h + r, c) is the pixel value of the input image at the position (k, h) and (p, r) on the channel c, W_m(p, r, c) represents the weight of the mth convolutional kernel at the position (p, r) and channel c, Bm represents the bias term of the mth convolutional kernel, and K_P represents the size of the convolutional kernel;
[0035] Use bilinear interpolation to restore the low-resolution feature map to the resolution of the original image, obtain the feature map F_up(k, h, m), perform a skip connection on the high-resolution feature map F_con(k, h, m) of the early convolutional layer and the upsampled feature map F_up(k, h, m) to obtain the feature map F_sk(k, h, m), perform global average on the feature map F_sk(k, h, m), and traverse each feature map to generate a fixed-length feature vector.
[0036] Preferably, the method for extracting features from the data of the lidar and millimeter-wave radar using a 3D convolutional neural network includes:
[0037] Normalize the size and normalize the image data of the lidar or millimeter-wave radar, and stack continuous S_c frames of images to form a four-dimensional tensor;
[0038] Design l_h convolutional layers, introduce adaptive pooling after each convolutional layer, use four-dimensional convolutional kernels to perform convolution on the input four-dimensional tensor, extract spatio-temporal features, output a feature map, add a temporal modeling layer after the convolutional layer, the temporal modeling layer is based on LSTM or GRU, set the output of the temporal modeling layer as the motion information of the target, and design a fully connected layer after the temporal modeling layer, and the output of the fully connected layer is a feature vector representing the comprehensive feature information of the target.
[0039] Preferably, the current state data of the vehicle body includes the driving state of the vehicle and various control parameters. The current environmental state outside the vehicle is displayed through the display screen in the vehicle or a voice broadcast is used for warning and reminding, and the driver controls the vehicle body according to the display and warning.
[0040] A recognition method for vehicle-human interaction includes:
[0041] Data acquisition module: used to acquire the original data of b_c types of sensors;
[0042] Data processing module: used to align the original data of each sensor in time and space to obtain the preprocessed data after data alignment;
[0043] Vehicle body state fitting and prediction module: build a deep neural network architecture, and perform feature-level fusion on the preprocessed data of each sensor through the deep neural network architecture to obtain the current environmental state outside the vehicle;
[0044] Control output module: used to interact with the driver, and give warnings and reminder displays to the driver according to the current environmental state outside the vehicle.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] In the present invention, by combining data from multiple sensors, high-precision real-time environmental perception is provided, and the accuracy of identifying objects (such as pedestrians, vehicles, obstacles, etc.) is significantly improved. The driving state of the vehicle and the surrounding environmental state are used to generate intelligent warnings and reminders, reducing the driver's cognitive burden and enhancing safety. By using deep learning technology, a large amount of sensor data is automatically processed and analyzed to generate environmental information and alarms in real time, improving the intelligence level of the vehicle. The system can automatically adjust the alarm content according to real-time environmental changes, making the interaction between people and vehicles more natural and enhancing the user experience;
[0047] By building a deep neural network architecture, the preprocessed data of each sensor is subjected to feature-level fusion through the deep neural network architecture. For the data of cameras and infrared sensors, a CNN network is used to extract features based on image seams, increasing the data processing efficiency and quality. For the data of lidar and millimeter-wave radar, a 3D convolutional neural network is used to extract features based on time series and adaptive image processing size, efficiently and accurately capturing the dynamic features of objects and making more accurate predictions of objects;
[0048] An adaptive regulation and optimization mechanism continuously improves the overall efficiency of intelligent analysis. Multiple modules cooperate with each other to enable comprehensive data collection, processing, and analysis, providing real-time and accurate analysis of detection data and achieving efficient coordinated control of multiple articulated arms in a spider crane. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 FIG. is a schematic structural diagram of an identification method for vehicle-human interaction proposed by the present invention;
[0050] Figure 2 FIG. is a schematic diagram of the method applied to the identification method for vehicle-human interaction in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] Hereinafter, exemplary embodiments according to the present application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.
[0052] Embodiment 1
[0053] Referring to Figure 1 and Figure 2 , Embodiment 1 further illustrates an identification method for vehicle-human interaction proposed by the present invention.
[0054] In vehicle-human interaction identification, the issue of identification accuracy is a key challenge, especially in complex environments, where there are problems such as misidentification or missed identification, difficulty in multi-target tracking, and delays in dynamic scenarios.
[0055] For misidentification or missed identification, the system may mistake background objects for pedestrians or vehicles, causing the vehicle to make incorrect responses. For example, it may suddenly brake when there are no pedestrians, affecting traffic flow. If the system fails to identify a real pedestrian or vehicle, it may lead to a failure to brake or avoid in a timely manner, thus triggering an accident, especially when a pedestrian suddenly appears in front of the vehicle. The reasons for this include:
[0056] 1. In low-light environments such as at night or in tunnels, insufficient light significantly reduces the visual recognition ability of the camera. Although infrared cameras or lidar can provide assistance, these sensors also have their limitations, such as lower resolution or limited detection range.
[0057] 2. In severe weather conditions such as heavy rain or heavy snow, rain or snowflakes will interfere with the camera's field of view and affect image clarity. In addition, raindrops or snowflakes may accumulate on the sensor surface, causing the sensor to fail. The signals of radar and lidar may also be distorted due to the reflection or scattering of rain and snow.
[0058] 3. When sunlight directly shines on the camera, strong light may cause overexposure, making the camera unable to correctly identify pedestrians or vehicles. In addition, strong light may produce reflections or shadows, which will also interfere with the system's recognition.
[0059] 4. In urban environments, the background is often very complex, including moving people, vehicles, billboards, trees, buildings, etc. Cameras or sensors may have difficulty distinguishing pedestrians, vehicles, and background objects. When the background is similar in color and texture to pedestrians or vehicles, the probability of misidentification increases.
[0060] Difficult multi-object tracking. The system may lose the position of a certain pedestrian or confuse two different targets, resulting in a deviation in the vehicle's response. For example, the system may think that a pedestrian has left the lane when in fact the pedestrian is still in a dangerous area. It may also ignore some key targets, especially those that are occluded or small, such as children or bicycles, which may cause the vehicle to fail to avoid or brake in a timely manner. The reasons for this include:
[0061] 1. On busy urban streets or in crowded areas, there are many targets such as pedestrians, vehicles, and bicycles, and the speeds, directions, and behaviors of these targets are different. In this case, the camera or sensor needs to detect and track multiple targets simultaneously, but the algorithm's ability is limited and it may be difficult to handle too many targets.
[0062] 2. When multiple targets are occluded from each other (for example, a pedestrian walks beside a car, or a vehicle drives behind a pedestrian), the system may have difficulty tracking the occluded target. This kind of occlusion is especially common during peak traffic hours or on narrow streets.
[0063] 3. When the appearance, size, color and other characteristics of pedestrians or vehicles are too similar, the system may have difficulty distinguishing different targets. This can cause the system to merge multiple targets into one, or misidentify one target as multiple targets between different frames.
[0064] These algorithms require a large amount of computing resources, and the computing time will increase significantly, resulting in delays. Similarly, there may be delays during the data transmission process, especially when multiple sensors are fused. The data update frequencies of each sensor are different, and the transmission paths and data processing speeds are also inconsistent, which easily leads to time delays. Especially in complex scenarios, it may cause delays in decision-making. Due to the limitations of computing and sensor processing capabilities, it may not be able to respond in a timely manner when facing rapidly changing scenarios. Delayed decisions may cause the vehicle to be unable to avoid dangers in time, increasing the risk of accidents.
[0065] The following is a technical solution to solve the above problems, and the specific content is as follows:
[0066] The b_c types of sensors include cameras, lidars, millimeter-wave radars, and infrared sensors installed on the vehicle body. The camera is used to provide high-resolution 2D images, the lidar is used to provide high-precision 3D point cloud data, the millimeter-wave radar is used to provide information on the distance, speed, and relative motion of objects, and the infrared sensor is used to provide temperature or thermal imaging data.
[0067] The camera is used to provide high-resolution visual information, which can capture detailed images and videos and is suitable for use in daytime and well-lit environments. It is mainly used for visual features such as color, texture, and edge detection. The lidar (LiDAR) is used to provide three-dimensional point cloud data of the environment, which can accurately perceive the shape and relative distance of objects, detect the geometric shape, size, and distance of objects, and is suitable for target detection in complex backgrounds and dynamic scenarios. The millimeter-wave radar has the characteristic of strong penetration and can work normally in bad weather (such as heavy rain, heavy snow, haze, etc.), and is especially suitable for detecting fast-moving targets. The infrared sensor can effectively detect pedestrians and vehicles at night or in low-light environments through thermal imaging technology and can detect objects in low-light or night environments.
[0068] Multi-sensor fusion can effectively improve the robustness of the system in various complex environments and overcome the limitations of a single sensor by integrating data from different types of sensors.
[0069] The methods for data alignment of the raw data of the b_c types of sensors include:
[0070] Step A1: Use the timestamp alignment method to synchronize the time of the raw data of the b_c types of sensors;
[0071] Step A2: Use coordinate transformation to perform spatial synchronization on the raw data of b_c types of sensors;
[0072] The method for performing time synchronization on the raw data of b_c types of sensors using the timestamp alignment method includes:
[0073] Select one of the sensors as the base sensor (the base sensor is generally selected to be collected at the same time point, and the sensor with the shortest raw data transmission time, that is, the fastest transmission arrival, is selected). Calculate the time delay between the base sensor and any other sensor, set the cycle time T and the delay hypothesis value set. The delay hypothesis value set includes preset time delay values, which can be set according to experimental data analysis or experience. Form two time series A[T] and B[T] from the raw data of the base sensor and other sensors within the cycle time T;
[0074] Calculate the signal means μA and μB of A[T] and B[T], where N represents the number of raw data within the cycle time, i represents the time index, a_t(i) represents the raw data value at time i in the time series A[T], b_t(i) represents the raw data value at time i in the time series B[T]. Decentralize the time series A[T] and B[T] to obtain A 1 [T] and B 1 [T], A 1 [T] = A[T] - μA, B 1 [T] = B[T] - μB, calculate the cross-correlation function where a 1 _t(i) represents the data value at time i in A 1 [T], b 1 _t(i + τ) represents the data value after delaying τ time at time i. Assign τ to traverse the values in the delay hypothesis value set and perform calculations. Select the value of τ when the cross-correlation function R AB (τ) is the largest as the actual delay time. Use the actual delay time to correct the time of other sensors compared with the base sensor, and obtain the corrected time series as B[T - τ];
[0075] Successively obtain the corrected time series of each sensor, and use interpolation, discarding, or repeating data to align the data of all sensors at the same time point;
[0076] For example: If the camera captures 30 frames of images per second, the LiDAR captures 10 frames of 3D point clouds per second, and the millimeter-wave radar obtains 20 speed samples per second, then these data need to be discarded or repeated and aligned according to the time stamps to ensure that there is valid data of each sensor at the same time point.
[0077] Data from different sensors are usually in their respective coordinate systems, so coordinate transformation is needed to map all data into the same reference coordinate system.
[0078] The method of spatially synchronizing the raw data of b_c types of sensors using coordinate transformation includes:
[0079] Set a base coordinate system, which can be the vehicle coordinate system or the world coordinate system;
[0080] Convert the camera coordinate system into the base coordinate system, obtain the 2D image collected by the camera, and add a coordinate system to form a 2D image coordinate system. Pixel points are mapped into the base coordinate system through the camera's internal parameter matrix and external parameter matrix;
[0081] Among them, the internal parameter matrix of the camera Among them, (f x , f y ) is the focal length, (c x , c y ) is the position of the optical center. The external parameter matrix of the camera is calibrated using a calibration board (such as a checkerboard). By taking z_p images of the calibration board at different angles and using a camera calibration function (such as OpenCV) to calculate the relationship between the pixel coordinates of the corner points in the calibration board image and the base coordinates, the external parameter matrix of the camera can be obtained;
[0082] The method of converting the infrared sensor coordinate system into the base coordinate system is the same as that of the camera;
[0083] Convert the LiDAR coordinate system into the base coordinate system. LiDAR directly provides 3D point cloud data. The LiDAR data is converted into the base coordinate system through the external parameter matrix and translation vector of the LiDAR;
[0084] Among them, the method of obtaining the external parameter matrix and translation vector of the LiDAR is as follows: Use a known circular dot array to ensure that it can be captured by the LiDAR and the camera simultaneously at different positions. Use the circular dot array as the calibration target, collect data simultaneously with the LiDAR and the camera at different angles and positions, use point cloud registration technology to align the point cloud data of the LiDAR with the image data of the camera, obtain the external parameter matrix and translation vector, and optimize the external parameter matrix through the nonlinear least squares method (such as the Levenberg - Marquardt algorithm). It can be implemented using open - source libraries (such as PCL, Open3D).
[0085] The point cloud registration technology can be ICP (Iterative Closest Point) or feature matching. ICP calculates rotation and translation by iteratively minimizing the distance between two sets of point clouds. Feature matching extracts feature points from the LiDAR and camera data for matching.
[0086] The method of converting the millimeter-wave radar coordinate system to the base coordinate system is the same as that of the lidar.
[0087] The extrinsic parameter matrix and the translation vector are calibrated through extrinsic calibration: It is necessary to obtain the extrinsic parameters (position and rotation relationship) of sensors such as cameras, LiDARs, millimeter-wave radars, and infrared sensors relative to the base coordinate system through the calibration process. Calibration can be completed by collecting calibration targets (images and point clouds of calibration plates or known markers). After that, the data of LiDAR, millimeter-wave radar, and camera are projected into the base coordinate system using the calibration results.
[0088] For example: Suppose the camera is located 1 meter in front of the vehicle, and the LiDAR is on the top of the vehicle. By calibration, their relative positions to the base coordinate (extrinsic parameter matrix) are obtained, and then the 3D point cloud coordinates of the LiDAR and the 2D images of the camera are projected into the base coordinate system for easy fusion. Furthermore, the coordinate systems of all sensors are converted into the same coordinate system to achieve spatial unity and alignment.
[0089] The multi-modal deep learning model can fuse and process data from different sensors, thereby improving the accuracy of multi-object detection and recognition.
[0090] The methods for feature-level fusion of the preprocessed data of each sensor include:
[0091] Set a comprehensive sample set, convert the images captured by the camera into standard-sized RGB images or grayscale images, and perform standardization processing (such as normalization); voxelize the lidar point cloud data to generate a bird's-eye view or depth map; convert the millimeter-wave radar data into an image form (such as a radar reflection intensity map); convert the infrared sensor data into a heat map or grayscale image, and perform standardization processing; all the collected data form a comprehensive sample set;
[0092] Build a deep neural network architecture for multi-sensor fusion, set the input of the input layer as the comprehensive sample set, design independent feature extraction networks for each sensor, use the CNN network to extract features from the data of the camera and infrared sensor, use the 3D convolutional neural network to extract features from the data of the lidar and millimeter-wave radar, and then obtain the feature vectors of each sensor. Use weighted summation to fuse the feature vectors into a large overall vector, input the large overall vector into the fully connected layer, and set the output layer as the output of u_p classification tasks and the output of u_k regression tasks. The classification tasks are used to identify objects of u_p categories, and the regression task output is used to predict the continuous attributes of the objects, such as distance, speed, and direction;
[0093] Integrate the output of the deep neural network architecture to obtain the current environmental state outside the vehicle;
[0094] Among them, the weights for weighted summation of each feature vector are obtained according to the Self-attention mechanism; during the fusion process, the Self-attention mechanism allows the feature vector of each sensor to interact with its own and the feature vectors of other sensors, dynamically learning the correlation of different sensor features.
[0095] The specific steps are as follows:
[0096] Set the input:
[0097] Suppose we have feature vectors from different sensors:
[0098] Camera feature: F_ca, with dimension (B, D);
[0099] LiDAR feature: F_LiD, with dimension (B, D);
[0100] Millimeter-wave radar feature: F_ra, with dimension (B, D);
[0101] Infrared feature: F_inf, with dimension (B, D);
[0102] Among them, B is the batch size and D is the dimension of the feature vector.
[0103] Feature matrix construction:
[0104] Combine these feature vectors into a matrix F, F = [F_ca, F_LiD, F_ra, F_inf], with shape (B, N, D), where N is the number of sensors (e.g., 4 sensors) and D is the feature dimension of each sensor:
[0105] Generate query (Q), key (K), and value (V):
[0106] Perform a linear transformation on the input feature matrix F to obtain the query matrix Q, key matrix K, and value matrix V. The dimension of each matrix is (B, N, D), indicating that each sensor generates independent query, key, and value vectors, Q = W_q * F, Q = W_k * F, V = W_v * F, where W_q, W_k, and W_v are learned weight matrices, usually of size (D, D).
[0107] Attention weight calculation:
[0108] Calculate the dot product of the query and the key, and normalize the result through Softmax to obtain the attention score matrix Here, K T is the transpose of the key matrix, is the scaling factor to ensure numerical stability.
[0109] Weighted summation:
[0110] Use the attention weight matrix A to perform weighted summation on the value matrix V, and calculate the output feature matrix \(F_y = A\cdot V\);
[0111] Output:
[0112] Take the feature matrix \(F_y\) after attention as the fused feature vector, and input it through a fully connected layer or directly as the input of subsequent layers.
[0113] The classification task is activated using the Softmax function, and the regression task is activated using a linear activation function. A combined loss function is used For the classification task, cross-entropy loss is used, and for the regression task, mean squared error loss is used. Here, \(\epsilon(j)\) represents the weight of the \(j\)-th classification task, \(CE\_loss(j)\) represents the cross-entropy loss of the \(j\)-th classification task, \(\theta(u)\) represents the weight of the \(u\)-th classification task, \(MSE\_loss(u)\) represents the mean squared error loss of the \(u\)-th regression task. The Adam or SGD optimizer is used to optimize the model.
[0114] The use of image segmentation technology can effectively separate the background and the target, thereby reducing the misidentification of objects in the background. Integrate the image segmentation technology into multi-sensor CNN feature extraction to output feature vectors. The methods for feature extraction of the data from the camera and the infrared sensor using the CNN network include:
[0115] Set the input as the data from the camera and the infrared sensor. The size of the original image is \((H, W, C)\), where \(H\) is the height, \(W\) is the width, and \(C\) is the number of channels;
[0116] Use multiple convolutional layers to extract local features Use the ReLU activation function to introduce non-linearity, and obtain the transformed output \(F\_relu(k, h, m)\). Downsample the feature map through the pooling layer, and repeat the convolution and pooling operations to obtain the low-resolution feature map \((H\) 11 , W 11 , C);
[0117] Among them, \(H\) 11 is the height, \(W\) 11 is the width, \(F\_con(k, h, m)\) represents the feature value of the \(m\)-th channel at the position \((k, h)\) in the output feature map, \(I(k + p, h + r, c)\) is the pixel value of the input image at the position \((k, h)\) and \((p, r)\) on the channel \(c\), \(W_m(p, r, c)\) represents the weight of the \(m\)-th convolutional kernel at the position \((p, r)\) and channel \(c\), \(Bm\) represents the bias term of the \(m\)-th convolutional kernel, and \(K_P\) represents the size of the convolutional kernel;
[0118] Restore the low-resolution feature map to the resolution of the original image using bilinear interpolation to obtain the feature map F_up(k, h, m). Perform a skip connection between the high-resolution feature map F_con(k, h, m) of the early convolutional layer and the upsampled feature map F_up(k, h, m) to obtain the feature map F_sk(k, h, m). Perform global average on the feature map F_sk(k, h, m) and traverse each feature map to generate a fixed-length feature vector.
[0119] By integrating image segmentation technology into multi-sensor CNN feature extraction, technical problems such as background and target separation, feature extraction complexity, and spatial information loss are solved, achieving efficient and accurate target recognition ability and enhancing the robustness and feature representation ability of the model.
[0120] The method for feature extraction of lidar and millimeter-wave radar data using a 3D convolutional neural network includes:
[0121] Normalize the size and normalize the image data of lidar or millimeter-wave radar, and stack continuous S_c frames of images to form a four-dimensional tensor.
[0122] Design l_h convolutional layers, introduce adaptive pooling after each convolutional layer, use four-dimensional convolutional kernels to perform convolution on the input four-dimensional tensor to extract spatio-temporal features and output feature maps. Add a temporal modeling layer after the convolutional layer. The temporal modeling layer is based on LSTM or GRU and is used to capture the dynamic changes of time series data, enhance the understanding of target state changes, set the output of the temporal modeling layer as the motion information of the target, and design a fully connected layer after the temporal modeling layer. The output of the fully connected layer is a feature vector representing the comprehensive feature information of the target.
[0123] By introducing temporal modeling and adaptive pooling, dynamic changes can be better captured, the model's understanding ability of the target can be improved, and the integrity of feature information can be maintained at the same time.
[0124] The current state data of the vehicle body includes the driving state of the vehicle and various control parameters, such as driving speed, direction, and other device parameters that can control the vehicle. The current environmental state outside the vehicle is displayed through the in-vehicle display screen or given a warning reminder through voice broadcast. The driver controls the vehicle body based on the display and warning.
[0125] Embodiment 2
[0126] Refer to Figure 1 and Figure 2 , Embodiment 2 further illustrates an identification method for vehicle-human interaction proposed by the present invention.
[0127] The identification method for vehicle-human interaction includes the following steps:
[0128] Step S1: Collect the raw data of b_c types of sensors;
[0129] Step S2: Align the raw data of b_c types of sensors to obtain the preprocessed data after alignment;
[0130] Step S3: Perform feature-level fusion on the preprocessed data of each sensor according to the deep neural network architecture to obtain the current external environment state of the vehicle;
[0131] Step S4: Collect the current state data of the vehicle body, and perform warning and reminder display according to the current external environment state of the vehicle.
[0132] The described recognition system for human-vehicle interaction, which is applied to a recognition method for human-vehicle interaction, includes:
[0133] Data acquisition module: used to collect the raw data of b_c types of sensors;
[0134] Data processing module: used to align the raw data of each sensor in terms of time and space to obtain the preprocessed data after data alignment;
[0135] Vehicle body state fitting and prediction module: Build a deep neural network architecture, and perform feature-level fusion on the preprocessed data of each sensor through the deep neural network architecture to obtain the current external environment state of the vehicle;
[0136] Control output module: used to interact with the driver, and perform warning and reminder display on the driver according to the current external environment state of the vehicle;
[0137] Each module is connected in a wired and / or wireless manner.
[0138] In addition, according to the embodiments of the present application, the process described in the accompanying drawings of a recognition method for human-vehicle interaction can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium, and the non-transitory machine-readable storage medium stores machine-readable instructions, and the machine-readable instructions can be run by a processor to execute instructions corresponding to the method steps provided by the present application. Of course, the architecture shown in the accompanying drawings of a recognition method for human-vehicle interaction is only exemplary. When implementing different devices, adaptive selection or adjustment can be made according to actual needs.
[0139] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the real situation. The preset parameters and threshold selection in the formulas are set by those skilled in the art according to the actual situation.
[0140] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. A recognition method for human-vehicle interaction, characterized in that: The human-vehicle interaction recognition method comprises: Step S1: Collecting raw data from b_c types of sensors; Step S2: aligning the raw data of b_c types of sensors to obtain aligned pre-processed data; Step S3: perform feature-level fusion on the pre-processed data of each sensor according to the deep neural network architecture to obtain the current environment status outside the vehicle; The method for fusing the preprocessed data of each sensor at the feature level includes: Set up a comprehensive sample set, convert the images captured by the camera into RGB images or grayscale images of standard size, and perform standardization processing; voxelize the laser radar point cloud data to generate a bird's-eye view or depth map; convert the millimeter wave radar data into image form; convert the infrared sensor data into a thermal map or grayscale map, and perform standardization processing; all collected data form a comprehensive sample set; Build a deep neural network architecture for multi-sensor fusion, set the input layer to a comprehensive sample set, use CNN network to extract features from camera and infrared sensor data, use 3D convolutional neural network to extract features from lidar and millimeter wave radar data, and then obtain the feature vector of each sensor. Use weighted summation to fuse each feature vector into a large overall vector, input the large overall vector into the fully connected layer, and set the output layer to u_p classification task outputs and u_k regression task outputs. The classification task is used to identify objects of u_p categories, and the regression task output is used to predict the continuous attributes of the object. The output of the integrated deep neural network architecture is used to obtain the current state of the environment outside the vehicle; The classification task uses the Softmax function for activation, the regression task uses the linear activation function for activation, and the combined loss function is used The classification task uses cross entropy loss, and the regression task uses mean square error loss, where ε(j) represents the weight of the j-th classification task, CE_loss(j) represents the cross entropy loss of the j-th classification task, θ(u) represents the weight of the u-th regression task, and MSE_loss(u) represents the mean square error loss of the u-th regression task. Use Adam or SGD to optimize the model; Step S4: Collect the current status data of the vehicle body, and display warnings and reminders according to the current environmental status outside the vehicle.
2. The human-vehicle interaction recognition method according to claim 1, characterized in that: The b_c types of sensors include cameras, lidars, millimeter-wave radars and infrared sensors installed on the vehicle body. The cameras are used to provide high-resolution 2D images, the lidars are used to provide high-precision 3D point cloud data, the millimeter-wave radars are used to provide distance, speed and relative motion information of objects, and the infrared sensors are used to provide temperature or thermal imaging data.
3. The human-vehicle interaction recognition method according to claim 2, characterized in that: The method for aligning the raw data of b_c types of sensors includes: Step A1: Use the timestamp alignment method to synchronize the raw data of b_c types of sensors; Step A2: Use coordinate transformation to spatially synchronize the raw data of b_c types of sensors.
4. The human-vehicle interaction recognition method according to claim 3, characterized in that: The method for performing time synchronization on raw data of b_c types of sensors using the timestamp alignment method includes: One of the sensors is selected as the base sensor, the time delay between the base sensor and any other sensors is calculated, the cycle time T and the delay assumption value set are set, the delay assumption value set includes a preset time delay value, and the raw data of the base sensor and other sensors within the cycle time T are formed into two sets of time series A[T] and B[T]; Calculate the signal means μA and μB of A[T] and B[T], and decentralize the time series A[T] and B[T] to obtain A 1 [T] and B 1 [T], calculate the cross-correlation function Assign τ to the values in the delay hypothesis value set and perform calculations, and select the cross-correlation function R AB The value of τ when (τ) is the maximum is taken as the actual delay time. The actual delay time is used to perform time correction on other sensors compared with the base sensor, and the corrected time series is obtained as B[T-τ]; Obtain the corrected time series of each sensor in turn, and use interpolation, discarding or repeating data to align the data of all sensors at the same time point; Where N represents the number of raw data in the cycle time, i represents the time index, and a 1 _t(i) represents A 1 [T] is the data value at time i, b 1 _t(i+τ) represents the data value after time i is delayed by τ.
5. The human-vehicle interaction recognition method according to claim 4, characterized in that: The method for spatially synchronizing raw data of b_c types of sensors using coordinate transformation includes: Set the base coordinate system, which can be the vehicle coordinate system or the world coordinate system; Convert the camera coordinate system to the base coordinate system, obtain the 2D image captured by the camera, and add the coordinate system to form a 2D image coordinate system. The pixel points are mapped to the base coordinate system through the camera intrinsic parameter matrix and extrinsic parameter matrix. Among them, the camera's internal parameter matrix Among them, (f x , f y ) is the focal length, (c x , c y ) is the optical center position. The camera's external parameter matrix is calibrated using a calibration plate. By taking z_p calibration plate images at different angles and using the camera calibration function to calculate the relationship between the pixel coordinates of the corner points in the calibration plate image and the base coordinates, the camera's external parameter matrix can be obtained. The method of converting the infrared sensor coordinate system to the base coordinate system is the same as that of the camera; Convert the laser radar coordinate system to the base coordinate system, and convert the laser radar data to the base coordinate system through the laser radar's external parameter matrix and translation vector; The method for obtaining the external parameter matrix and translation vector of the laser radar is as follows: using the dot array as the calibration target, using the laser radar and the camera to collect data at different angles and positions, using the point cloud registration technology to align the point cloud data of the laser radar with the image data of the camera, obtaining the external parameter matrix and translation vector, and optimizing the external parameter matrix by the nonlinear least squares method; The method of converting the millimeter wave radar coordinate system into the base coordinate system is the same as that of the laser radar.
6. The human-vehicle interaction recognition method according to claim 5, characterized in that: The method for extracting features from the camera and infrared sensor data using a CNN network includes: The input is set to the data of the camera and infrared sensor, and the size of the original image is (H, W, C), where H is the height, W is the width, and C is the number of channels; Use multiple convolutional layers to extract local features Use the ReLU activation function to introduce nonlinearity, obtain the transformed output F_relu(k,h,m), downsample the feature map through the pooling layer, and repeat the convolution and pooling operations to obtain a low-resolution feature map (H 11 ,W 11 ,C); Among them, H 11 is the height, W 11 is the width, F_con(k,h,m) represents the feature value of the mth channel at position (k,h) in the output feature map, I(k+p,h+r,c) is the pixel value of the input image at position (k,h) and (p,r) on channel c, W_m(p,r,c) represents the weight of the mth convolution kernel at position (p,r) and channel c, Bm represents the bias term of the mth convolution kernel, and K_P represents the size of the convolution kernel; Use bilinear interpolation to restore the low-resolution feature map to the resolution of the original image to obtain the feature map F_up(k,h,m), jump connect the high-resolution feature map F_con(k,h,m) of the early convolutional layer with the upsampled feature map F_up(k,h,m) to obtain the feature map F_sk(k,h,m), perform global averaging on the feature map F_sk(k,h,m), and traverse each feature map to generate a feature vector of fixed length.
7. The human-vehicle interaction recognition method according to claim 6, characterized in that: The method for extracting features from the laser radar and millimeter wave radar data using a 3D convolutional neural network includes: The image data of the laser radar or millimeter wave radar is normalized and processed in size, and the continuous S_c frame images are stacked to form a four-dimensional tensor; Design l_h convolutional layers, introduce adaptive pooling after each convolutional layer, use four-dimensional convolution kernel to convolve the input four-dimensional tensor, extract spatiotemporal features, output feature maps, add a timing modeling layer after the convolutional layer, the timing modeling layer is based on LSTM or GRU, set the output of the timing modeling layer as the motion information of the target, design a fully connected layer after the timing modeling layer, and the output of the fully connected layer is a feature vector representing the comprehensive feature information of the target.
8. The human-vehicle interaction recognition method according to claim 7, characterized in that: The current status data of the vehicle body includes the driving status of the vehicle and various control parameters. The current environmental status outside the vehicle is displayed on the display screen inside the vehicle, or a voice broadcast is used for warning reminders. The driver controls the vehicle body according to the display and warning.
9. A human-vehicle interaction recognition system, applied to the human-vehicle interaction recognition method according to any one of claims 1 to 8, characterized in that: The human-vehicle interaction recognition system comprises: Data acquisition module: used to collect raw data from b_c types of sensors; Data processing module: used to align the raw data of each sensor in time and space, and obtain pre-processed data after data alignment; Vehicle state fitting prediction module: Build a deep neural network architecture, use the deep neural network architecture to fuse the pre-processed data of each sensor at the feature level, and obtain the current environment state outside the vehicle; Control output module: used to interact with the driver and provide warnings and reminders to the driver based on the current environmental status outside the vehicle.
Citation Information
Patent Citations
Road data acquisition and simulation scene establishment integrated system and method
CN112307594A
Target detection method based on fusion of prior positioning of millimeter-wave radar and visual feature
US20220198806A1