Feature fusion method based on Transform multi-source sensor
By combining multi-view cameras and LiDAR sensors with a multimodal self-attention mechanism for feature fusion, the problem of sensor fusion methods being unable to capture global features is solved, thereby improving the performance and safety of autonomous driving systems.
Patent Information
- Application Number
- CN202511084701.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-07
AI Technical Summary
Existing sensor fusion methods are mostly limited to the integration of local perception information, making it difficult to accurately capture global feature interactions in complex traffic environments and effectively characterize the motion uncertainty of traffic participants, thus affecting the robustness and safety of autonomous driving decisions.
Image data is acquired and features are extracted using a multi-view camera mounted on the target vehicle, and point cloud data is acquired using a lidar sensor. Multi-scale feature fusion is performed through a multimodal self-attention mechanism to obtain a global feature sequence, and the target vehicle is controlled based on the global feature sequence.
It enables the acquisition of global scene information in autonomous driving scenarios, improves the performance and safety of autonomous driving systems, and reduces driving decision errors and collision risks.
Smart Images

Figure CN120913022A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic driving, and in particular to a feature fusion method and device based on a Transformer multi-source sensor, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of automatic driving and intelligent transportation, vehicle-mounted sensors and computer vision technology play an increasingly important role in vehicle environment perception. Common sensors in automatic driving technology include RGB cameras and LiDARs, both of which have complementary advantages.
[0003] At present, the academic and industrial circles have carried out a large amount of research and practice around intelligent vehicle environment perception, but there are still the following three problems that have not been solved in the existing environment perception related technologies: 1. The existing sensor fusion methods (such as geometric projection methods and local feature aggregation methods of convolutional neural networks) are mostly limited to the integration of local perception information, and it is difficult to accurately capture the global feature interaction in complex traffic environments, such as the mutual influence and cooperative relationship among traffic signals, vehicles, pedestrians and other elements in the intersection. 2. The existing technologies fail to effectively consider the motion uncertainty of traffic participants in complex traffic environments, such as sudden acceleration and deceleration of vehicles, random deviation of trajectories, so as to accurately describe the uncertain interaction coupling relationship between the ego vehicle and other traffic participants with strong randomness, and also difficult to effectively depict the dynamic evolution law of the real traffic situation. 3. The existing environment perception technology has deficiencies in the quantitative modeling of the uncertainty of the motion state of other traffic participants, especially the randomness of kinematic information such as acceleration mutation and trajectory change, so as to affect the robustness and safety of automatic driving decision.
[0004] Therefore, how to capture the global features in complex traffic environments is a problem that needs to be solved by the technical personnel in this field. SUMMARY
[0005] The present application provides a feature fusion method and device based on a Transformer multi-source sensor, an electronic device and a storage medium, to solve the problem that the existing sensor fusion methods are mostly limited to the integration of local perception information, and it is difficult to accurately capture the global feature interaction in complex traffic environments.
[0006] According to an aspect of the present application, a feature fusion method based on a Transformer multi-source sensor is provided, which comprises:
[0007] acquiring image data around the target vehicle by a multi-camera carried by the target vehicle, and performing image feature extraction on the image data;
[0008] The laser radar sensor carried by the target vehicle is used to obtain point cloud data around the target vehicle, and point cloud feature extraction is performed on the point cloud data; wherein, the image data and the point cloud data both contain environmental information around the target vehicle;
[0009] A multi-modal self-attention mechanism is used to perform multi-scale feature fusion on the extracted image features and point cloud features, to obtain a global feature sequence;
[0010] The global feature sequence is used to predict a waypoint sequence of the target vehicle at a next time, and the target vehicle is controlled according to the predicted waypoint sequence.
[0011] According to another aspect of the present application, a feature fusion device based on a Transformer multi-source sensor is provided, and the device comprises:
[0012] An image feature extraction module is configured to use a multi-view camera carried by a target vehicle to obtain image data around the target vehicle, and perform image feature extraction on the image data;
[0013] A point cloud feature extraction module is configured to use a laser radar sensor carried by the target vehicle to obtain point cloud data around the target vehicle, and perform point cloud feature extraction on the point cloud data; wherein, the image data and the point cloud data both contain environmental information around the target vehicle;
[0014] A global feature fusion module is configured to use a multi-modal self-attention mechanism to perform multi-scale feature fusion on the extracted image features and point cloud features, to obtain a global feature sequence;
[0015] A vehicle control module is configured to use the global feature sequence to predict a waypoint sequence of the target vehicle at a next time, and control the target vehicle according to the predicted waypoint sequence.
[0016] According to another aspect of the present application, an electronic device is provided, and the electronic device comprises:
[0017] at least one processor; and
[0018] a memory in communication with the at least one processor; wherein,
[0019] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the feature fusion method based on a Transformer multi-source sensor according to any one of the embodiments of the present application.
[0020] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for causing a processor to implement the method for feature fusion based on a Transformer multi-source sensor according to any of the embodiments of the present application when executed.
[0021] According to another aspect of the present application, a computer program product is provided, which comprises a computer program for implementing the method for feature fusion based on a Transformer multi-source sensor according to any of the embodiments of the present application when executed by a processor.
[0022] The technical scheme of the embodiment of the present application acquires image data around the target vehicle by using the multi-view camera carried by the target vehicle, and extracts image features from the image data; acquires point cloud data around the target vehicle by using the laser radar sensor carried by the target vehicle, and extracts point cloud features from the point cloud data; wherein the image data and the point cloud data both contain environmental information around the target vehicle; performs multi-scale feature fusion on the extracted image features and point cloud features by using a multi-modal self-attention mechanism to obtain a global feature sequence; predicts a waypoint sequence of the target vehicle at the next moment according to the global feature sequence, and controls the target vehicle according to the predicted waypoint sequence. The problem that existing sensor fusion methods are mostly limited to the integration of local perception information and are difficult to accurately capture global feature interactions in complex traffic environments is solved. The different spatial information from images and point clouds is efficiently fused by using a multi-modal self-attention mechanism to obtain global scene information in an autonomous driving scenario, and the beneficial effect of improving the performance of an autonomous driving system is achieved.
[0023] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0025] Figure 1 is a flowchart of a method for feature fusion based on a Transformer multi-source sensor according to an embodiment of the present application;
[0026] Figure 2is a flow chart of a feature fusion method based on a Transformer multi-source sensor according to Embodiment Two of the present application;
[0027] Figure 3 is a flow chart of another feature fusion method based on a Transformer multi-source sensor according to Embodiment Two of the present application;
[0028] Figure 4 is a flow chart of a feature fusion method according to Embodiment Two of the present application;
[0029] Figure 5 is a detailed structure flow chart of an autoregressive waypoint prediction network according to Embodiment Two of the present application;
[0030] Figure 6 is a CARLA driving simulator schematic diagram according to Embodiment Two of the present application;
[0031] Figure 7 is a structural schematic diagram of a feature fusion device based on a Transformer multi-source sensor according to Embodiment Three of the present application;
[0032] Figure 8 is a structural schematic diagram of an electronic device according to Embodiment Four of the present application. DETAILED DESCRIPTION
[0033] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.
[0034] Wherein, the acquisition, storage, use and processing of data in the technical scheme of the present application comply with the relevant provisions of laws and regulations. It should be noted that the terms "first", "second", "target", "original" and the like in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include", "equal" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] Embodiment one
[0036] Figure 1 A flowchart of a feature fusion method based on a Transformer multi-source sensor is provided for the first embodiment of the present application. The present embodiment can be applicable to the case of feature fusion based on a Transformer for features obtained by a multi-source sensor. The method can be executed by a feature fusion device based on a Transformer multi-source sensor, which can be realized in the form of hardware and / or software. The feature fusion device based on a Transformer multi-source sensor can be configured in any electronic device with network communication function. As shown in the figure, the method comprises: Figure 1
[0037] S110, acquiring image data around the target vehicle by a multi-view camera carried by the target vehicle, and performing image feature extraction on the image data.
[0038] Wherein, the multi-view camera can refer to a camera system that simultaneously acquires scene information by using multiple cameras or sensors, and realizes omnidirectional perception and understanding of the scene through analysis and fusion of multiple images. In the present embodiment, the image data around the target vehicle is acquired by the front-view camera, rear-view camera and side-view camera carried by the target vehicle.
[0039] The image data is the acquired RGB image data around the target vehicle, and the image data contains the environmental information around the target vehicle. The environmental information includes but is not limited to the road information and the traffic participant information around the target vehicle.
[0040] The image feature extraction can refer to extracting spatial semantic features of the acquired image data around the target vehicle to describe semantic information and detail information of the road or traffic participants around the target vehicle.
[0041] In S120, point cloud data around the target vehicle is acquired by using a laser radar sensor carried by the target vehicle, and point cloud feature extraction is performed on the point cloud data.
[0042] The laser radar sensor can refer to acquiring distance, speed, direction and shape information of a target object by emitting laser and receiving its echo signal. In the embodiment of the application, the laser radar sensor is used to acquire distance, speed, direction and shape information of traffic participants or traffic signs on the road around the target vehicle.
[0043] The point cloud data is a data set of spatial points scanned by a three-dimensional laser radar sensor, each point containing three-dimensional coordinate information (X, Y, Z), and some also containing color, reflectivity, echo times and other information. In the embodiment of the application, the point cloud data is used to determine the environmental information around the target vehicle.
[0044] The point cloud feature extraction can refer to converting the acquired three-dimensional point cloud data into a two-dimensional pseudo-image representation and further encoding to obtain a high-dimensional point cloud feature map. The point cloud features are extracted to obtain the global spatial context information around the target vehicle.
[0045] In S130, multi-scale feature fusion is performed on the extracted image features and point cloud features by using a multi-modal self-attention mechanism to obtain a global feature sequence.
[0046] The multi-modal self-attention mechanism is a core component of the Transformer model, which maps the input sequence to multiple different representation spaces and calculates the attention weights between these representations, so that the model can learn the information in different subspaces of the sequence. In the embodiment of the application, the multi-modal self-attention mechanism included in the Transformer model is used to fuse the image features and point cloud features, and calculate the attention weights of the image features and point cloud features respectively, so as to obtain the global features around the target vehicle and strengthen the interaction between the global context and the local detail information.
[0047] Multi-scale feature fusion can refer to integrating feature information of different scales to improve the understanding ability of image content. In the embodiment of the application, the image features and point cloud features of different scales are fused to improve the understanding of the global context information around the target vehicle.
[0048] The image features and the point cloud features are fused through the multi-modal self-attention mechanism in the embodiment of the application to obtain the global scene information and the features of the target vehicle around, so that the understanding of the global scene is improved.
[0049] S140, predict a waypoint sequence of the target vehicle at a next time according to the global feature sequence, and control the target vehicle according to the predicted waypoint sequence.
[0050] The waypoint sequence can be an ordered arrangement of a series of waypoints and is commonly used to guide a vehicle to sail along a specific path to achieve a specific transportation or operation task. The waypoint sequence in the embodiment of the application can be a waypoint trajectory sequence, that is, the sailing trajectory of the target vehicle at the next time is predicted according to the global feature sequence, and the sequence composed of the position points of each position in the sailing trajectory is called the waypoint sequence. The predicted waypoint sequence is input into the controller of the target vehicle to intelligently control the target vehicle.
[0051] The embodiment of the application provides a feature fusion method based on a Transformer multi-source sensor, image data around a target vehicle is acquired by using a multi-camera carried by the target vehicle, and image features are extracted from the image data; point cloud data around the target vehicle is acquired by using a laser radar sensor carried by the target vehicle, and point cloud features are extracted from the point cloud data; the image data and the point cloud data both contain environmental information around the target vehicle; multi-scale feature fusion of the extracted image features and the point cloud features is performed by using a multi-modal self-attention mechanism to obtain a global feature sequence; a waypoint sequence of the target vehicle at a next time is predicted according to the global feature sequence, and the target vehicle is controlled according to the predicted waypoint sequence. The technical scheme of the embodiment of the application realizes efficient fusion of different spatial information from images and point clouds through a multi-modal self-attention mechanism to obtain global scene information in an automatic driving scene and improve the performance of an automatic driving system.
[0052] Embodiment two
[0053] Figure 2 A flowchart of a feature fusion method based on a Transformer multi-source sensor is provided for the second embodiment of the application, the foregoing embodiment is further optimized on the basis of the above-mentioned embodiment, and the embodiment of the application can be combined with each optional scheme in one or more of the above-mentioned embodiments. As shown in the figure, the method comprises the following steps. Figure 2
[0054] S210, image data around a target vehicle is acquired by using a multi-camera carried by the target vehicle, and image features are extracted from the image data.
[0055] Wherein, referring to Figure 3 image data is acquired by a multi-view camera (including front, side and rear view cameras) mounted on the target vehicle, and the image data is taken as input for deep neural network feature extraction.
[0056] As an optional but non-limiting implementation, the image data around the target vehicle is acquired by the multi-view camera mounted on the target vehicle, and image feature extraction on the image data includes but is not limited to steps A1-A3:
[0057] Step A1: image data around the target vehicle is acquired by the multi-view camera mounted on the target vehicle, and the image data is subjected to image cropping and normalization processing to obtain pre-processed image data.
[0058] Step A2: the pre-processed image data is input into a preset convolutional neural network to obtain a multi-scale feature map.
[0059] Step A3: the multi-scale feature map is input into a spatial pyramid pooling layer to perform feature fusion on the multi-scale feature map to obtain a unified high-dimensional image feature vector; wherein the high-dimensional image feature vector is used to describe semantic information and detailed information of the road scene around the target vehicle.
[0060] Wherein, the multi-view camera mounted on the target vehicle is used to acquire image data around the target vehicle, and the image data is RGB image data. The RGB image data is subjected to field of view cropping and normalization processing to generate image input for deep neural network feature extraction, and the resolution of the pre-processed image data is 24x224. The pre-processed image data is input into a preset convolutional neural network to extract spatial semantic features of each camera image. The preset convolutional neural network can be a pre-trained ResNet-50 convolutional neural network.
[0061] Specifically, the image feature extraction module uses the multi-scale feature maps output by the Conv1, Conv2, Conv3 and Conv4 layers in the pre-trained ResNet-50 convolutional neural network, and after spatial pyramid pooling (SPP), the feature maps of different scales are fused into a unified high-dimensional feature vector to describe the semantic information and detailed information of the road scene. The feature extraction operation can be represented as:
[0062] F img =SPP(ResNet50(I RGB ))
[0063] Wherein, F img represents the extracted image feature, and I RGBRGB image data, SPP represents a spatial pyramid pooling layer, and ResNet50 represents a pre-trained ResNet-50 convolutional neural network.
[0064] S220, acquiring point cloud data around the target vehicle by using a laser radar sensor carried by the target vehicle, and performing point cloud feature extraction on the point cloud data.
[0065] The point cloud data around the target vehicle is acquired by using a laser radar sensor carried by the target vehicle, and the point cloud data is taken as an input for deep neural network feature extraction.
[0066] As an optional but non-limiting implementation manner, the acquiring of the point cloud data around the target vehicle by using the laser radar sensor carried by the target vehicle and the point cloud feature extraction on the point cloud data include but are not limited to steps B1-B2:
[0067] Step B1: acquiring three-dimensional point cloud data around the target vehicle by using a laser radar sensor carried by the target vehicle, and performing voxelization processing on the three-dimensional point cloud data to convert the three-dimensional point cloud data into a two-dimensional feature map.
[0068] Step B2: inputting the two-dimensional feature map into a preset convolutional neural network to perform feature extraction and spatial context information extraction, and obtaining a high-dimensional bird's eye view feature vector.
[0069] The three-dimensional point cloud data around the target vehicle is acquired by using a laser radar sensor carried by the target vehicle, voxelization processing is performed on the three-dimensional point cloud data, a voxelization method is used to project original point cloud to a two-dimensional pseudo image representation under a bird's eye view (BEV) perspective, and finally a convolutional neural network is used to further extract features and spatial context information from the two-dimensional pseudo image representation, and output a BEV feature tensor with a size of 128x128x64. The feature extraction process is represented as:
[0070] F bev =CNN(Voxelization(P LiDAR ))
[0071] F bev represents the extracted point cloud feature, P LiDAR represents input point cloud data, CNN represents a convolutional neural network, and Voxelization represents voxelization processing.
[0072] In the embodiment of the present application, image data and point cloud data are acquired, and features of the image data and the point cloud data are extracted, so as to fully capture global spatial context information in a traffic environment, and effectively solve the problem that the existing sensor fusion method based on local geometric projection is difficult to accurately capture global scene information in a complex urban driving scene, and is prone to cause driving decision errors.
[0073] S230, flatten the extracted image features into a first mark sequence, and add a learnable first spatial position code in the first mark sequence to obtain an image sequence.
[0074] Wherein, referring to Figure 4 After the image features and the point cloud features are acquired, the image features and the point cloud features need to be fused. Specifically, the feature map of each modality is first flattened into a mark sequence, and a learnable spatial position code is added to form an input form suitable for the Transformer. The interaction relationship between the two modality feature sequences is calculated through the multi-head self-attention mechanism in the Transformer structure, so as to realize feature fusion and obtain a fusion feature sequence carrying rich scene context. The fusion feature sequence is further processed by multiple Transformer modules, and is connected in residual with the original modality feature map at different spatial scales, so as to strengthen the interaction of global context and local detail information.
[0075] Optionally, the embodiment of the present application constructs an image feature processing branch to perform feature dimension reduction and spatial position information embedding on the RGB image data collected by the camera modality. Specifically, the dimension reduction mapping of the image features is:
[0076] X′ img =W img X img +b img
[0077] Wherein: X img is a 2048-dimensional feature vector output by image feature extraction; W img is a weight matrix of image feature linear mapping; b img is an image feature bias term; X′ img is a 512-dimensional feature vector after image feature dimension reduction.
[0078] First, the image feature processing branch receives the high-dimensional image features output by the image feature extraction (ResNet-50), and the feature vector dimension is 2048. Then, the 2048-dimensional high-dimensional image feature vector is mapped to a unified 512-dimensional feature space through a linear projection layer (Linear Projection Layer) for dimension reduction, so as to facilitate subsequent processing of the Transformer structure. Then, the 512-dimensional image feature sequence after dimension reduction is input into the positional encoding layer (Positional Encoding Layer), and the positional encoding information is added to the image feature sequence, so as to explicitly represent the spatial position information of the image spatial features, so that the Transformer module can perceive the difference in spatial position information between the image features in subsequent processing. After the above steps, the output of the image feature processing branch is an N img token sequence with a length of N
[0079] S240, flatten the extracted point cloud features into a second token sequence, and add a learnable second spatial position encoding to the second token sequence to obtain a point cloud sequence.
[0080] In the embodiment of the present application, a point cloud feature processing branch is also provided to project the three-dimensional point cloud data collected by the laser radar modality into a BEV (Bird's Eye View) feature, reconstruct the spatial feature, and encode the position. Specifically, the following processing steps are included: first, the point cloud (LiDAR) feature processing branch receives the BEV spatial feature tensor output by the LiDAR feature extraction module, and the dimension of the feature tensor is 128x128x64. The spatial resolution is 128x128, and the feature channel dimension of each spatial position is 64. Second, the BEV spatial feature tensor is subjected to a spatial dimension flattening (Spatial Flattening) operation, and the three-dimensional tensor with a dimension of 128x128x64 is flattened into a feature token sequence with a length of 16384, and each token has an initial dimension of 64. Then, the 64-dimensional feature tokens are uniformly mapped to a 512-dimensional feature space through a linear projection layer (Linear Projection Layer), so as to realize dimension matching with the image feature branch and facilitate subsequent unified feature fusion. The formula for feature mapping is:
[0081] X′ lidar =W lidar X lidar +b lidar
[0082] wherein: X lidar is the initial 64-dimensional LiDAR feature sequence after spatial flattening, Wlidar is a weight matrix for linear mapping of point cloud features; b lidar is a bias term for point cloud features; X' lidar is a 512-dimensional feature vector after dimension reduction of point cloud features.
[0083] Then, the 512-dimensional feature token after uniform mapping is input into a positional encoding layer (Positional Encoding Layer), and position encoding information is introduced into the BEV feature token sequence of the LiDAR modality, so that the Transformer fusion unit can fully capture the spatial position relationship information of the LiDAR feature. After the above steps, the output of the LiDAR feature processing branch is a sequence of 16384 features tokens, each token having a dimension of 512.
[0084] S250, using a multi-modal self-attention mechanism to perform multi-scale feature fusion on the image sequence and the point cloud sequence to obtain a global feature sequence.
[0085] Among them, the multi-modal self-attention mechanism in the Transformer architecture is used to perform multi-scale feature fusion on the image sequence and the point cloud sequence to obtain a global feature sequence.
[0086] As an optional but non-limiting implementation manner, the multi-modal self-attention mechanism is used to perform multi-scale feature fusion on the image sequence and the point cloud sequence to obtain a global feature sequence, including but not limited to steps C1-C2:
[0087] Step C1: concatenating the image sequence and the point cloud sequence into a multi-modal feature sequence.
[0088] Step C2: taking the multi-modal feature sequence as the input of the Transformer encoder, and performing cross-modal information interaction and fusion through the multi-modal self-attention mechanism to obtain a global feature sequence.
[0089] Among them, the fusion unit performs multi-scale feature fusion based on the self-attention mechanism (Self-Attention Mechanism) in the Transformer architecture to obtain a global feature sequence. Specifically, the fusion unit concatenates the 512-dimensional image feature token sequence output by the image feature processing branch and the 512-dimensional LiDAR feature token sequence output by the LiDAR feature processing branch into a unified multi-modal feature sequence, with a length of (N img+16384) tokens, each token has a dimension of 512. The multi-modal feature sequence is input into the Transformer encoder, and the multi-modal information interaction and fusion are realized through the multi-head self-attention mechanism. In this process, each modality feature token not only interacts with other tokens in the same modality, but also interacts with the feature tokens from another modality, so as to realize effective cross-modal information exchange and feature complementation. The multi-head self-attention calculation method is as follows:
[0090]
[0091] Wherein: Attention represents the multi-head self-attention mechanism, softmax represents the normalization function, Q is the query matrix, K is the key matrix, V is the value matrix, d k is the dimension of the key vector.
[0092] The embodiment of the application respectively constructs a double-channel convolutional neural network for processing image data and LiDAR bird's eye view (BEV) data, and introduces a Transformer module in the network multi-layer feature extraction stage, realizes multi-scale feature fusion between the two modal data through cross-modal attention mechanism, significantly improves the effectiveness of sensor fusion, and solves the problem that the traditional fusion method is difficult to balance the fusion accuracy of different resolution features of multi-modal data.
[0093] S260, according to the global feature sequence, the target vehicle's waypoint sequence at the next time is predicted, and the target vehicle is controlled according to the predicted waypoint sequence.
[0094] Wherein, the global feature sequence fused by the multi-modal Transformer is further compressed in feature dimension through global average pooling operation, and a high-dimensional scene feature vector is obtained. The formula of feature compression is as follows:
[0095]
[0096] Wherein: X global is the compressed scene feature vector, N is the length of the fused feature sequence; is the i-th feature token of the fused feature sequence.
[0097] The compressed feature vector is reduced in dimension through a multilayer perceptron (MLP, also known as a multilayer neural network), and the MLP network is represented as:
[0098]
[0099] Wherein: Xmlp is the output feature of the MLP network; is the weight matrix of the MLP; is the bias term; and σ is the ReLU activation function.
[0100] The scene feature vector is sent to a Gated Recurrent Unit (GRU) network for autoregressive waypoint prediction after dimension reduction and feature enhancement via a multi-layer neural network (MLP). Specifically, the GRU receives the current feature state and the navigation target position (GPS coordinates) as inputs, and predicts the two-dimensional waypoint positions of the vehicle at future consecutive time steps to realize end-to-end autonomous driving decision-making. The specific recursive formula is:
[0101] h t = GRU(h t-1 , [p t-1 ; g])
[0102] where h t is the hidden state at the t-th step, p t-1 is the predicted waypoint at the (t-1)-th step, and g is the navigation target position vector.
[0103] The target vehicle takes the predicted future waypoint sequence as input and realizes steering, throttle, and brake control of the actual vehicle through a pre-set PID controller. Specifically, steering control is realized by minimizing the angle between the predicted waypoint and the current vehicle direction through the PID controller; throttle and brake control is dynamically adjusted according to the distance change between predicted waypoints to ensure smooth and safe completion of the driving task.
[0104] Optionally, referring to Figure 5 , the autoregressive waypoint prediction network receives the output feature from the multi-modal Transformer fusion module, i.e., the fused global feature vector, denoted as fusion feature vector f. The dimension of the fusion feature vector f is 512. In order to improve the subsequent calculation efficiency and reduce the prediction complexity, the feature vector f is dimensionally compressed through a multi-layer perceptron (MLP):
[0105] f' = σ(W2·σ(W1f+b1)+b2)
[0106] where: represents the input fusion feature; is the weight matrix of the MLP; is the bias vector; and σ(·) represents the ReLU nonlinear activation function; is the output feature vector of the MLP.
[0107] First, a multi-layer perceptron (MLP) is used to compress the feature dimension of the fusion feature vector f from 512 dimensions to 64 dimensions, to improve the efficiency of feature utilization and reduce the complexity of subsequent prediction calculations. Specifically, the MLP consists of two hidden layers, the first hidden layer contains 256 neurons, and the second hidden layer contains 128 neurons. Each hidden layer uses a ReLU activation function to introduce non-linear characteristics and improve the expression ability for complex driving scenarios. The 64-dimensional feature vector f' after MLP processing is used to initialize the hidden state of a single-layer gated recurrent unit (GRU). The GRU is responsible for completing the recursive prediction of the future series of waypoints.
[0108] As an optional but non-limiting implementation, the prediction of the target vehicle's waypoint sequence at the next time according to the global feature sequence, and the control of the target vehicle according to the predicted waypoint sequence, include but are not limited to steps D1-D4:
[0109] Step D1: The global feature sequence is compressed in feature dimension by a global average pooling operation to obtain a first feature vector.
[0110] Step D2: The first feature vector is processed by a multi-layer perceptron for dimension reduction to obtain a second feature vector.
[0111] Step D3: The second feature vector and the navigation target position of the target vehicle are input into a gated recurrent unit network to predict a two-dimensional waypoint sequence of the target vehicle at the next time.
[0112] Step D4: The two-dimensional waypoint sequence is input into a preset controller to control the target vehicle; wherein the control includes steering, throttle and brake control of the target vehicle.
[0113] The input vector of the GRU network includes a position input vector at the current prediction time and a target position vector. The position input vector is the relative position in the coordinate system centered on the current time of the ego vehicle, and the target position vector is provided by a high-level path planning system (GPS global path planning) to guide the network to generate the correct local prediction direction. Specifically, at the first prediction time, the position input vector is initialized to (0, 0), indicating the current coordinate position of the ego vehicle. The target position vector is a two-dimensional vector representing the relative target point position in the ego vehicle coordinate system.
[0114] In the recursive prediction process, each prediction step GRU unit takes the hidden state of the previous prediction step and the input vector of the current step as input, and generates the hidden state of the next prediction step and the waypoint offset (δw). Specifically: in the first step of prediction, the GRU hidden state is initialized with the fusion feature vector output by the MLP, and (0, 0) is used as input to predict the first waypoint offset. In subsequent steps, the hidden state update of each GRU unit is represented by the following formula:
[0115] z t =σ(W z x t +U z h t-1 +b z )
[0116] r t =σ(W r x t +U r h t-1 +b r )
[0117]
[0118] where x t is the input vector of the current time step, for example (0, 0) or the previous waypoint prediction value; h t-1 represents the hidden state of the previous time step; z t represents the update gate, which determines how much of the previous memory to retain in the current state; r t represents the reset gate, which determines the degree of "forgetting" of previous memory by the current input; represents the candidate hidden state, which is generated by the current input and part of the historical information; h t represents the hidden state of the current time step (i.e. the output of the GRU); ⊙ represents element-wise multiplication; W * , U * represent weight matrices, b * represents a bias term, * includes z, r, h; σ() represents the Sigmoid function.
[0119] The GRU network is followed by a linear output layer, which is responsible for converting the hidden state output by the GRU into a two-dimensional waypoint offset δwt, and the final waypoint prediction result is obtained through the following formula:
[0120] δw t =W o h t +b o
[0121] where δw t represents the waypoint offset corresponding to the time step; W oa weight matrix representing a linear layer; b o a bias term.
[0122] The GRU recurrent prediction network proceeds with 4 prediction steps to generate the future 4 waypoints prediction results for vehicle control, i.e., {w1, w2, w3, w4}. The four predicted waypoints constitute the trajectory points of the vehicle planning path in the future period of time, so that the downstream controller can generate the steering, throttle and brake control instructions of the vehicle according to the waypoints.
[0123] Two PID controllers are further used in the embodiment of the present application to implement low-level control of the vehicle, so as to realize actual steering control and speed control of the vehicle. The lateral PID controller utilizes the forward direction determined by the predicted waypoints to realize vehicle steering angle control.
[0124]
[0125] wherein u steer (t) represents the steering angle control amount at the current time; e(t) represents the lateral error between the current position and the target direction (determined by the predicted waypoints); K p represents the proportional gain, which controls the response to the current lateral error; K i represents the integral gain, which suppresses the lateral steady-state error; K d represents the derivative gain, which improves the response to the trend of change in the lateral error.
[0126] The longitudinal PID controller utilizes the distance between the consecutive predicted waypoints to calculate the speed control instruction, thereby realizing acceleration and deceleration control of the vehicle.
[0127]
[0128] wherein u thrpttle / brake (t) represents the speed control amount at the current time; d(t) represents the longitudinal error between the current position and the target direction (determined by the predicted waypoints); K' p represents the proportional gain, which controls the response to the longitudinal error; K' i represents the integral gain, which suppresses the longitudinal steady-state error; K' d represents the derivative gain, which improves the response to the trend of change in the longitudinal error.
[0129] The PID controller parameters adopt the typical values provided by the CARLA simulator officially, and optionally, the lateral PID controller parameters can be respectively set as K p = 1.25, K i = 0.75, K d = 0.3, and the longitudinal PID controller parameters can be respectively set as K' p = 5.0, K' i= 0.5, K' = 1.0 d The detailed parameters and control method of the PID controller can be adjusted according to actual conditions to adapt to different driving scenes, and are not specifically limited in the embodiment of the present application.
[0130] The embodiment of the present application proposes an end-to-end autoregressive waypoint prediction network, which performs time series processing on the global feature representation after multi-modal fusion to output the future navigation trajectory of the target vehicle at multiple time steps, and realizes vehicle control decision in combination with the PID controller. The technical solution significantly improves the driving stability and safety of the automatic driving system in dense traffic scenes, reduces the risk of vehicle collision, and effectively solves the problem of high accident risk of existing end-to-end driving methods in complex scenes such as multi-lane lane changing and unprotected turning.
[0131] As an optional but non-limiting implementation, the method further includes but is not limited to steps E1-E3:
[0132] Step E1: determining the state of the target vehicle when controlling the target vehicle.
[0133] Step E2: if the target vehicle is stationary for more than a preset time threshold, performing a crawling operation on the target vehicle through a preset controller; wherein the crawling operation refers to configuring a crawling speed for the target vehicle through the preset controller.
[0134] Step E3: detecting whether there is an obstacle in front of the target vehicle in real time according to the point cloud data, and terminating the crawling operation when an obstacle is detected.
[0135] In order to further improve the safety and reliability of the autoregressive waypoint prediction network in actual driving, the system sets the following safety auxiliary strategy: when the vehicle is stationary for more than a preset threshold time (55 seconds), a smaller target speed (4 m / s) will be temporarily set for the PID controller to avoid permanent parking of the vehicle due to inertia. In order to avoid collision during the crawling process, the LiDAR data or the vehicle detection result based on CenterNet is detected to monitor whether there is an obstacle in front of the vehicle within a certain range in real time; once a potential obstacle is detected, the crawling strategy will be terminated immediately to ensure driving safety.
[0136] The embodiment of the present application proposes an active safety heuristic mechanism for model safe driving, which monitors the dynamic obstacle information of the area in front of the vehicle in real time, triggers the active safety strategy in time when there is a potential collision risk, effectively reduces the safety accidents caused by the uncertainty of automatic driving decision, and solves the problem of insufficient safety redundancy of the existing end-to-end automatic driving method.
[0137] Optionally, the autoregressive waypoint prediction network of the embodiment of the present application is trained using a supervised learning method, and the loss function is trained using the Euclidean distance loss (L2 loss) between the predicted waypoints and the expert true waypoints to improve the overall generalization performance. The training of the autoregressive waypoint prediction network uses the AdamW optimizer, the initial learning rate is 1x10-4, the training round is set to about 40 epochs, 4 NVIDIA RTX 4080 GPUs are used for training, and the batch size is 12. The specific parameters can be adjusted according to the model size and data size to obtain the best performance, which is not specifically limited in the embodiment of the present application. The autoregressive waypoint prediction network described in the present application effectively realizes high-precision prediction from fused features to multi-step waypoints, effectively improves the safety and reliability of the autonomous driving system, and significantly reduces the risk of collision and violation during driving.
[0138] In an optional scheme of the embodiment of the present application, referring to Figure 6 The embodiment of the present application evaluates and verifies the performance in the CARLA driving simulator environment. The CARLA simulator can provide a high-fidelity virtual city driving scene, including dense traffic conditions, intersections, turning sections, signal-controlled intersections, and other complex traffic environments. In order to fully evaluate the effectiveness of the autonomous driving system proposed in the present application, the present application uses a variety of challenging driving tasks for experimental verification, including but not limited to unprotected left turn, straight driving at intersections, turning through, and passing through dense vehicle interaction sections.
[0139] Specifically, the present application uses the provided 100 evaluation routes, as well as the self-designed high-difficulty route set Longest6 to comprehensively and objectively evaluate the system. The evaluation indicators include driving success rate (Success Rate, SR), route completion (Route Completion, RC), and system comprehensive driving score (Driving Score, DS). The driving success rate is defined as the percentage of vehicles successfully reaching the target without collision or violation events on a given route; the route completion is defined as the percentage of the vehicle safely driving and successfully completing the route mileage in the total mileage; and the comprehensive driving score is the overall performance indicator after comprehensively considering the route completion and safety, and its definition formula is:
[0140]
[0141] In the formula: DS represents the overall driving score; RC represents the route completion rate; Collision represents the number of collision events during vehicle driving; Infraction represents the number of violations (such as running red lights, crossing lines, etc.) during driving; α and β represent the penalty coefficients for collision events and violations, respectively. This invention adopts the default recommended values of CARLA, namely (α = 0.50) and (β = 0.30).
[0142] Experimental results show that the multimodal Transformer fusion architecture proposed in this invention significantly improves the system's performance on the aforementioned evaluation metrics compared to traditional multimodal fusion methods and image-only input methods. Specifically, in the CARLA leaderboard evaluation, the overall driving score (DS) of the proposed method is significantly higher than that of existing technologies, demonstrating a clear technological advantage. This is mainly attributed to the fact that the multimodal Transformer fusion method proposed in this invention effectively captures global contextual information in complex traffic scenarios, enhances the interactive fusion capability of different modal data, and thus significantly improves the system's decision-making ability and safe driving performance in complex traffic scenarios.
[0143] Furthermore, in order to improve the understanding and generalization performance of the multimodal fusion architecture of the present invention on complex scene information, the present invention also introduces several auxiliary supervised tasks, specifically including: image branch auxiliary tasks and LiDAR branch auxiliary tasks.
[0144] The image branching auxiliary task includes semantic segmentation and depth prediction tasks, and its loss function is defined as follows:
[0145]
[0146] Where: L seg This represents the semantic segmentation cross-entropy loss; Let represent the label of the true class (c) of the (i)th pixel; i and c represent the probabilities of the predicted classes by the model.
[0147] The LiDAR-assisted tasks include high-definition map (HD Map) prediction and vehicle detection. The loss function for the vehicle detection task can be defined as:
[0148] L det =L cls (p,p * )+l reg (b,b * )
[0149] Among them, L det Indicates loss in vehicle inspection tasks; L cls (p,p* ) represents the classification loss of the target class, which measures the difference between the model predicted class probability p and the actual class label p * reg (b,b * ) represents the regression loss of the bounding box parameters, which measures the distance or overlap between the predicted bounding box b and the real bounding box b * The above loss terms are implemented by cross-entropy loss and IoU loss.
[0150] The introduction of these auxiliary tasks significantly improves the generalization ability and robustness of the system, effectively reducing the collision rate and the occurrence rate of violation events in complex urban environments.
[0151] In addition, the Transformer fusion structure used by the system also significantly solves the bottleneck problem that the existing technology based on geometric projection or late fusion method cannot effectively extract cross-modal global scene information. Specifically, through the self-attention mechanism of the Transformer, the global context information between different modal feature sequences is modeled, effectively capturing the mutual relationship between features in the traffic scene, especially the interaction information of key traffic elements such as long-distance vehicles and traffic lights.
[0152] The automatic driving method described in the application adopts a multi-task learning framework, and introduces deep prediction, semantic segmentation, high-definition map prediction, and target vehicle detection and other auxiliary supervision signals in the model training stage, further enhancing the representation ability of intermediate features, significantly improving the robustness and generalization performance of the driving model, and solving the problem that single task supervision in the existing method cannot meet the model generalization ability requirement in complex dynamic driving scenes.
[0153] The embodiment of the application provides a feature fusion method based on a Transformer multi-source sensor, which effectively improves the perception ability and decision accuracy of the automatic driving system in complex traffic scenes through a multi-scale Transformer cross-modal feature fusion technology, an autoregressive waypoint prediction strategy, and a multi-task auxiliary supervision learning framework, significantly improves the vehicle driving safety and driving stability, and can effectively meet the automatic driving application requirements of the real traffic environment.
[0154] Embodiment three
[0155] Figure 7 A structural schematic diagram of a feature fusion device based on a Transformer multi-source sensor provided by the third embodiment of the application. As shown in the figure, the device comprises: Figure 7
[0156] The image feature extraction module 710 is configured to acquire image data around the target vehicle by using a multi-view camera carried by the target vehicle, and perform image feature extraction on the image data.
[0157] The point cloud feature extraction module 720 is configured to acquire point cloud data around the target vehicle by using a laser radar sensor carried by the target vehicle, and perform point cloud feature extraction on the point cloud data; wherein the image data and the point cloud data both contain environmental information around the target vehicle.
[0158] The global feature fusion module 730 is configured to perform multi-scale feature fusion on the extracted image features and point cloud features by using a multi-modal self-attention mechanism, to obtain a global feature sequence.
[0159] The vehicle control module 740 is configured to predict a waypoint sequence of the target vehicle at a next time according to the global feature sequence, and control the target vehicle according to the predicted waypoint sequence.
[0160] Optionally, the image feature extraction module is specifically configured to:
[0161] acquire image data around the target vehicle by using a multi-view camera carried by the target vehicle, and perform image cropping and normalization processing on the image data to obtain preprocessed image data;
[0162] input the preprocessed image data into a preset convolutional neural network to obtain a multi-scale feature map;
[0163] input the multi-scale feature map into a spatial pyramid pooling layer to perform feature fusion on the multi-scale feature map, to obtain a unified high-dimensional image feature vector; wherein the high-dimensional image feature vector is used to describe semantic information and detailed information of a road scene around the target vehicle.
[0164] Optionally, the point cloud feature extraction module is specifically configured to:
[0165] acquire three-dimensional point cloud data around the target vehicle by using a laser radar sensor carried by the target vehicle, and perform voxelization processing on the three-dimensional point cloud data to convert the three-dimensional point cloud data into a two-dimensional feature map;
[0166] input the two-dimensional feature map into a preset convolutional neural network to perform feature extraction and spatial context information extraction, to obtain a high-dimensional bird's eye view feature vector.
[0167] Optionally, the global feature fusion module is specifically configured to:
[0168] flatten the extracted image features into a first mark sequence, and add a learnable first spatial position encoding in the first mark sequence to obtain an image sequence.
[0169] The extracted point cloud features are flattened into a second token sequence, and a learnable second spatial position encoding is added in the second token sequence to obtain a point cloud sequence;
[0170] The image sequence and the point cloud sequence are fused by a multi-modal self-attention mechanism to obtain a global feature sequence.
[0171] Optionally, the global feature fusion module is specifically configured to:
[0172] The image sequence and the point cloud sequence are spliced into a multi-modal feature sequence;
[0173] The multi-modal feature sequence is taken as an input of a Transformer encoder, and cross-modal information interaction and fusion are performed by a multi-modal self-attention mechanism to obtain a global feature sequence.
[0174] Optionally, the vehicle control module is specifically configured to:
[0175] The global feature sequence is compressed in feature dimension by a global average pooling operation to obtain a first feature vector;
[0176] The first feature vector is processed by a multi-layer perceptron for dimension reduction to obtain a second feature vector;
[0177] The second feature vector and a navigation target position of the target vehicle are input into a gated recurrent unit network to predict a two-dimensional waypoint sequence of the target vehicle at a next time;
[0178] The two-dimensional waypoint sequence is input into a preset controller to control the target vehicle; wherein the control includes steering, throttle and brake control of the target vehicle.
[0179] Optionally, the device further comprises a vehicle safety auxiliary module, which is specifically configured to:
[0180] When the target vehicle is controlled, the state of the target vehicle is determined;
[0181] If the target vehicle is stationary for more than a preset time threshold, the target vehicle is controlled by the preset controller to perform a crawling operation; wherein the crawling operation refers to configuring a crawling speed for the target vehicle by the preset controller;
[0182] According to the point cloud data, it is detected in real time whether there is an obstacle in front of the target vehicle, and the crawling operation is terminated when it is detected that there is an obstacle.
[0183] The device for feature fusion based on the Transformer multi-source sensor provided in the embodiments of the present application can execute the method for feature fusion based on the Transformer multi-source sensor provided in any of the embodiments of the present application, has the corresponding functions and advantages of executing the method for feature fusion based on the Transformer multi-source sensor, and the detailed process is described in the foregoing method for feature fusion based on the Transformer multi-source sensor.
[0184] Embodiment four
[0185] Figure 8 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0186] As shown in Figure 8 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0187] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0188] The processor 11 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the Transformer-based multi-source sensor feature fusion method.
[0189] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-described functions defined in the methods of embodiments of the present application are performed.
[0190] In some embodiments, the Transformer-based multi-source sensor feature fusion method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the Transformer-based multi-source sensor feature fusion method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the Transformer-based multi-source sensor feature fusion method by any other appropriate means, for example, by means of firmware.
[0191] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a special-purpose standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0192] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package, and partially on a machine or entirely on a remote machine or server.
[0193] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0194] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0195] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0196] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0197] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.
[0198] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above, but only by the scope of the appended claims.
Claims
1. A method for feature fusion based on a Transformer multi-source sensor, characterized in that, The method comprises the following steps: acquiring image data around the target vehicle by a multi-camera carried by the target vehicle, and performing image feature extraction on the image data; acquiring point cloud data around the target vehicle by a laser radar sensor carried by the target vehicle, and performing point cloud feature extraction on the point cloud data; wherein the image data and the point cloud data both contain environmental information around the target vehicle; performing multi-scale feature fusion on the extracted image features and point cloud features by a multi-modal self-attention mechanism to obtain a global feature sequence; predicting a waypoint sequence of the target vehicle at a next time according to the global feature sequence, and controlling the target vehicle according to the predicted waypoint sequence.
2. The method of claim 1, wherein, The method of acquiring image data around the target vehicle by a multi-camera carried by the target vehicle, and performing image feature extraction on the image data comprises the following steps: acquiring image data around the target vehicle by a multi-camera carried by the target vehicle, and performing image cropping and normalization processing on the image data to obtain preprocessed image data; inputting the preprocessed image data into a preset convolutional neural network to obtain a multi-scale feature map; inputting the multi-scale feature map into a spatial pyramid pooling layer to perform feature fusion on the multi-scale feature map to obtain a unified high-dimensional image feature vector; wherein the high-dimensional image feature vector is used to describe semantic information and detailed information of a road scene around the target vehicle.
3. The method of claim 1, wherein, The method of acquiring point cloud data around the target vehicle by a laser radar sensor carried by the target vehicle, and performing point cloud feature extraction on the point cloud data comprises the following steps: acquiring three-dimensional point cloud data around the target vehicle by a laser radar sensor carried by the target vehicle, and performing voxelization processing on the three-dimensional point cloud data to convert the three-dimensional point cloud data into a two-dimensional feature map; inputting the two-dimensional feature map into a preset convolutional neural network to extract features and spatial context information, and obtaining a high-dimensional bird's eye view feature vector.
4. The method of claim 1, wherein, The method of performing multi-scale feature fusion on the extracted image features and point cloud features by a multi-modal self-attention mechanism to obtain a global feature sequence comprises the following steps: flattening the extracted image features into a first token sequence, and adding a learnable first spatial position encoding in the first token sequence to obtain an image sequence; flattening the extracted point cloud features into a second token sequence, and adding a learnable second spatial position encoding in the second token sequence to obtain a point cloud sequence; performing multi-scale feature fusion on the image sequence and the point cloud sequence by a multi-modal self-attention mechanism to obtain a global feature sequence.
5. The method of claim 4, wherein, The method of performing multi-scale feature fusion on the image sequence and the point cloud sequence by a multi-modal self-attention mechanism to obtain a global feature sequence comprises the following steps: splicing the image sequence and the point cloud sequence into a multi-modal feature sequence; inputting the multi-modal feature sequence as an input of a Transformer encoder, and performing cross-modal information interaction and fusion by a multi-modal self-attention mechanism to obtain a global feature sequence.
6. The method of claim 1, wherein, The method comprises the following steps: The global feature sequence is compressed in feature dimension by a global average pooling operation to obtain a first feature vector; The first feature vector is processed by a multilayer perceptron to obtain a second feature vector; The second feature vector and the navigation target position of the target vehicle are input into a gated recurrent unit network to predict a two-dimensional waypoint sequence of the target vehicle at the next moment; The two-dimensional waypoint sequence is input into a preset controller to control the target vehicle; wherein the control includes steering, throttle and brake control of the target vehicle.
7. The method of claim 1, wherein, The method further comprises: When the target vehicle is controlled, the state of the target vehicle is determined; If the target vehicle is stationary for more than a preset time threshold, the target vehicle is controlled by the preset controller to perform a crawling operation; wherein the crawling operation refers to configuring a crawling speed for the target vehicle by the preset controller; The presence of obstacles in front of the target vehicle is detected in real time according to the point cloud data, and the crawling operation is terminated when obstacles are detected.
8. A feature fusion device based on a Transformer multi-source sensor, characterized in that, It comprises: An image feature extraction module is used to obtain image data around the target vehicle by a multi-camera carried by the target vehicle, and image features are extracted from the image data; A point cloud feature extraction module is used to obtain point cloud data around the target vehicle by a laser radar sensor carried by the target vehicle, and point cloud features are extracted from the point cloud data; wherein the image data and the point cloud data both contain environmental information around the target vehicle; A global feature fusion module is used to perform multi-scale feature fusion on the extracted image features and point cloud features by a multi-modal self-attention mechanism to obtain a global feature sequence; A vehicle control module is used to predict a waypoint sequence of the target vehicle at the next moment according to the global feature sequence, and to control the target vehicle according to the predicted waypoint sequence.
9. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the feature fusion method based on the Transformer multi-source sensor according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the feature fusion method based on the Transformer multi-source sensor according to any one of claims 1-7 when executed.