Laser radar multi-task real-time sensing method combined with Transform coding

By combining Transformer coding with a multi-task real-time perception method for LiDAR, the problems of low hardware efficiency and poor detection accuracy of radar target detection in complex scenarios are solved. This method enables multi-task collaborative detection, improves the robustness and accuracy of dynamic object recognition and edge region segmentation, and is suitable for applications such as maritime inspection and intelligent driving.

CN120953753APending Publication Date: 2025-11-14STATE POWER INVESTMENT CORP JIANGSU OFFSHORE WIND POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510955972.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing radar target detection technologies suffer from low hardware efficiency and poor detection accuracy in complex scenarios, making it difficult to achieve multi-task collaborative detection. Furthermore, they involve a large amount of computation, making it difficult to implement real-time processing on resource-limited embedded devices.

Method used

A multi-task real-time perception method for LiDAR combining Transformer encoding is adopted. By collecting and preprocessing LiDAR 3D point cloud data, a two-dimensional network BEV feature map is generated. The Transformer encoder is used for encoding and decoding, and semantic segmentation, target detection and motion segmentation are processed separately. Multi-task joint prediction is performed by combining cross-level feature pyramid fusion and cross-task semantic weighting mechanism.

Benefits of technology

It significantly improves the robustness and detection accuracy of dynamic object recognition and edge region segmentation in complex scenarios, and achieves high efficiency and accuracy of real-time perception for multi-task applications, making it suitable for fields such as maritime inspection and intelligent driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953753A_ABST
    Figure CN120953753A_ABST
Patent Text Reader

Abstract

The invention discloses a laser radar multi-task real-time sensing method combined with Transform coding, and relates to the technical field of radar target detection, and the method comprises the steps: collecting laser radar 3D point cloud image data, converting the data into BEV aerial view representation, and constructing a two-dimensional grid BEV feature map; constructing a Transform encoder to perform multi-scale feature extraction on the BEV feature map, decoding multi-level semantic features output by the encoder, and generating initial prediction of target detection, semantic segmentation and motion segmentation tasks; and meanwhile, an SWAG semantic weighting guide module is introduced to carry out cross-task feature fusion on semantic segmentation features and target detection features, so that the information interaction between tasks is enhanced. According to the method, the limitations of low hardware efficiency and poor detection precision of a traditional CNN architecture in a complex scene are solved, and the robustness and detection precision of dynamic object recognition and edge region segmentation in the complex scene are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of radar target detection technology, and in particular to a multi-task real-time sensing method for lidar that incorporates Transformer coding. Background Technology

[0002] Radar target detection technology is a key detection technology that integrates environmental perception, target recognition, and motion state analysis. In modern military and civilian fields, the demand for target detection in complex scenarios is increasing. Besides conventional target detection and tracking tasks, accurately distinguishing different types of targets, identifying their motion characteristics, and predicting their trajectories are fundamental to ensuring the system's detection effectiveness and decision-making accuracy. Therefore, achieving high-precision, real-time, and robust target detection within the monitoring area has become an important research direction.

[0003] While optical sensors are widely used in target recognition tasks due to their intuitive imaging and rich information, they are highly dependent on weather conditions and visibility, especially at night or in adverse weather conditions, where their detection performance significantly decreases. Radar, as the core sensor in a target detection system, possesses advantages such as strong all-weather operation and accurate ranging and velocity measurement, providing stable yet structured echo information, effectively compensating for the shortcomings of optical sensors.

[0004] In practical radar target detection tasks, researchers face numerous technical challenges. First, radar echo signals themselves are subject to noise interference and exhibit significant signal-to-noise ratio fluctuations, with signal quality deteriorating further at long distances or in complex environments, making it difficult to accurately identify small targets or low-reflectivity objects. Second, existing research largely focuses on single-task implementations, such as target classification or trajectory prediction, lacking systematic exploration of multi-task collaborative detection. Furthermore, different tasks typically require training data from different scenarios, and discrepancies in data distribution and inconsistent labeling standards between tasks significantly complicate joint modeling. Moreover, complex detection algorithms often involve enormous computational demands, making them difficult to deploy on resource-constrained embedded devices for real-time processing.

[0005] In recent years, with the development of hardware computing power and algorithm technology, end-to-end detection methods based on multi-task learning have gradually attracted attention. Multi-task learning can effectively improve the overall performance and adaptability of detection systems by mining the feature correlations and complementarities between tasks. However, existing methods are still mainly based on optical or fusion detection, and multi-task detection schemes for radar are still relatively limited. Summary of the Invention

[0006] The problem to be solved by this invention is to provide a multi-task real-time perception method for LiDAR that combines Transformer encoding, in order to overcome the limitations of low hardware efficiency and poor detection accuracy of traditional CNN architecture in complex scenes, and improve the robustness and detection accuracy of dynamic object recognition and edge region segmentation in complex scenes.

[0007] This invention adopts the following technical solution: a multi-task real-time sensing method for lidar combining Transformer coding, comprising the following steps:

[0008] S1. Collect and preprocess data: Collect and preprocess 3D point cloud data from LiDAR, and use the BEV conversion method to generate a two-dimensional network BEV feature map.

[0009] S2. Construct a Transformer encoder to encode the BEV feature map of the two-dimensional network, forming feature data at multiple different levels;

[0010] S3. Divide the decoding process into different detection tasks, including: semantic segmentation task, object detection task, and motion segmentation task. Construct specific Transformer decoders for different tasks to decode the feature data processed in S2 and generate feature layer data for different tasks.

[0011] S4. Input the feature layer data of different tasks into the segmentation head of different tasks to obtain the detection results of different tasks. Obtain semantic labels based on the semantic segmentation task, obtain predicted detection boxes containing target center key points, target direction and target size based on the target detection task, and obtain motion segmentation masks based on the motion segmentation task.

[0012] S5. Based on the feature data of different levels formed in step S2 and the feature layer data of different tasks formed in step S3, establish a real-time prediction model for LiDAR multi-tasks that combines Transformer encoding. Combining the long-term dependency and parallel computing capabilities of Transformer, and through cross-level feature pyramid fusion and cross-task semantic weighting mechanism, perform multi-task joint prediction on the LiDAR point cloud scene to obtain real-time perception results including semantic segmentation labels, target detection boxes and motion segmentation masks.

[0013] Preferably, step S1 collects and preprocesses the data, including the following sub-steps;

[0014] S1.1 Collect and preprocess 3D point cloud data from LiDAR to form the raw dataset;

[0015] S1.2. Use distance-based point cloud densification method to enhance 3D point cloud data, preserve object outlines, increase point cloud density, and improve the clarity of distant object boundaries.

[0016] S1.3. Use the BEV conversion method to convert 3D point cloud data into two-dimensional network BEV feature maps.

[0017] Preferably, in step S2, the Transformer encoder encoding process includes the following sub-steps:

[0018] S2.1. Convert the 2D network BEV feature map into sequence elements and use them as input to the Transformer encoder;

[0019] S2.2. For each sequence element of the BEV feature map, perform feature embedding and position embedding to obtain the input embedding vector X;

[0020] S2.3. Perform multi-head attention calculations on the input embedding vector X to obtain the corresponding attention weights, and then perform a weighted sum to obtain the output result of each head; the multi-head self-attention mechanism has h parallel heads, and the outputs of the h heads are concatenated and processed by a linear transformation matrix W. O This yields the final output Z of the multi-head self-attention mechanism.

[0021] S2.4 Feed the output result Z into the feedforward neural network module to transform and extract features, thereby enhancing the model's expressive power;

[0022] S2.5. Use layer normalization to normalize the data and use residual connections to directly add the input and output data of the module to alleviate the gradient vanishing problem.

[0023] Preferably, in step S3, a specific Transformer decoder is constructed for different tasks to decode the feature data processed in S2, including the following sub-steps:

[0024] S3.1 Reconstruct the sequence features output by the Transformer encoder into 2D network data, including: fixed-position encoding and sequence feature reconstruction;

[0025] S3.2 Divide the decoding process into different detection tasks, including: semantic segmentation task, object detection task, and motion segmentation task. Use a specific decoder for each detection task to generate the corresponding output.

[0026] Preferably, in step S3.1, the fixed position encoding involves adding a predefined two-dimensional position encoding P to map the input sequence feature Z to the spatial position of the original image, thereby obtaining the sequence feature Z′.

[0027] The sequence feature reshaping process reshapes the sequence feature Z′ into a two-dimensional feature map F. reshaped .

[0028] Preferably, in step S3.2, a specific decoder is used for different detection tasks to generate corresponding outputs, as follows:

[0029] S3.2.1 Semantic segmentation task decoding: The low-resolution features output by the encoder are gradually restored to spatial resolution through a series of upsampling blocks to generate multi-level semantic features;

[0030] S3.2.2 Target Detection Task Decoding: A key-point-based detection framework is adopted, combined with a semantic guidance mechanism, to generate corresponding multi-level fusion features;

[0031] S3.2.3 Motion segmentation task decoding: Capture motion information between each frame, use a motion decoder to gradually restore spatial feature resolution, upsample the features to the original size, and predict the motion segmentation mask.

[0032] Preferably, the semantic segmentation task decoding includes the following sub-steps:

[0033] Upsampling: Upsampling is performed using transposed convolution to improve the low-resolution feature map F. low Upsampled to the target size, feature F is obtained. up ;

[0034] Feature concatenation: The upsampled features F... up Features F of the layer corresponding to the encoder skip By concatenating along the channel dimension, we obtain the concatenated feature F. concat ;

[0035] Convolutional processing: Semantic segmentation decoding convolutional processing is adopted. The concatenated features are fused through a series of convolutional layers to reduce the number of channels and extract contextual information, resulting in the output feature F. out ;

[0036] Repeated upsampling: Repeat the above process to gradually restore the resolution.

[0037] Preferably, the target detection task decoding includes the following sub-steps:

[0038] Semantic segmentation branch decoding: Process Transformer features to generate semantic information S;

[0039] Object detection branch decoding: directly processes Transformer features to generate object detection features F. det ;

[0040] SWAG semantic weighting and guidance: Semantic information S is fused into the object detection branch and projected into the joint embedding space; matching vectors are calculated for the two dimensions in the embedding space to obtain matching vector m; corresponding weight vectors are calculated for matching vector m, and the detection features and weight vectors are weighted and fused to obtain the fused feature F. fusion .

[0041] Preferably, the motion segmentation task decoding includes the following sub-steps:

[0042] Encoder feature extraction: Obtaining multi-scale features from different levels of the Transformer;

[0043] Upsampling and final feature fusion: Upsampling is performed through multiple transposed convolutions and then concatenated with the shallowest features of the encoder to obtain the final dense feature map F. final ;

[0044] Segmentation head prediction: Map features to column space and use Softmax to generate probability distribution.

[0045] Preferably, in step S4, the feature layer data processed in S3 is input into the segmentation head of different tasks to obtain the detection results of different tasks, including the following sub-steps:

[0046] S4.1, Semantic Segmentation Task Output:

[0047] For the final output dense feature map F final The input is fed into the pixel-level classification head, and the Softmax classifier in the classification head performs classification prediction on each pixel to obtain the semantic label P;

[0048] S4.2, Target Detection Task Output:

[0049] S4.2.1 Target Center Key Point Prediction: Based on the final fusion feature F fusion Generate key point heatmaps and predict the center location of each category c;

[0050] S4.2.2 Target size prediction: Perform offset prediction to compensate for the loss of coordinate accuracy caused by downsampling, and then perform size prediction to predict the width and length of each detection box;

[0051] S4.2.3 Target direction prediction: The direction is discretized into several bins, and the direction angle θ is determined by the argmax function after the Softmax function.

[0052] S4.3, Motion Segmentation Task Output:

[0053] Features are mapped to the class space using 1×1 convolution, and a probability distribution is generated using the Softmax function. Probability calculations are then performed to obtain the motion segmentation mask M.

[0054] Preferably, in step S5, establishing a multi-task real-time prediction model for LiDAR combined with Transformer encoding includes the following sub-steps:

[0055] S5.1, based on the different levels of feature data formed by S2 (including pyramid-shaped feature levels containing low-resolution high semantic features and high-resolution detail features) and the different task feature layer data formed by S3 (intermediate features for semantic segmentation, object detection, and motion segmentation), construct a "cross-level-cross-task" dual-path fusion architecture to perform end-to-end multi-task joint prediction.

[0056] S5.2. By capturing global contextual relationships through the long-term dependency modeling capability of Transformer, improving inference efficiency by leveraging parallel computing characteristics, and combining multi-task feature fusion mechanism, the robustness of dynamic object recognition and edge region segmentation in complex scenarios is improved.

[0057] S5.3 outputs multi-dimensional perception results, which can directly serve fields such as maritime inspection, intelligent driving, and industrial security, providing efficient and accurate solutions for real-time environmental perception and decision-making.

[0058] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:

[0059] 1. The present invention provides a multi-task real-time perception method for LiDAR, which constructs a two-dimensional gridded BEV feature map based on the acquired LiDAR point cloud map; then divides the BEV feature map into multiple image blocks, maps them to serialized input through a linear embedding layer, divides the feature map into local windows, performs attention calculation on the features within the local windows, and achieves cross-region feature interaction through window offset in adjacent stages. In each stage, the feature map is downsampled by a factor of two and the number of channels is expanded through block merging operation, thereby gradually constructing a pyramid-shaped feature containing low-resolution high semantic information and high-resolution detailed features. Finally, five feature layers of different scales are output, realizing the effective capture of semantic features from local geometric structure to global scene.

[0060] 2. The multi-task real-time perception method of LiDAR in this invention decodes the multi-level semantic features output by the encoder through a task-aware Transformer decoder to generate features for different tasks. In the feature layer output after decoding, a SWAG-based semantic weighted guidance module is used to perform cross-task feature fusion of semantic segmentation features and target detection features as the input for the final target detection task, thereby enhancing the ability to locate and classify targets in object detection.

[0061] 3. The method of this invention utilizes a Transformer-based LiDAR detection network architecture to handle multi-task real-time detection problems, which can significantly improve the robustness and detection accuracy of dynamic object recognition and edge region segmentation in complex scenes. It is the first to realize multi-task real-time perception of LiDAR based on pure Transformer, providing a perception solution that combines high efficiency, accuracy and deployability for multi-task real-time detection problems. Attached Figure Description

[0062] Figure 1 This is a flowchart of the multi-task real-time sensing method for lidar of the present invention;

[0063] Figure 2 This is a structural diagram of the Transformer encoder of the present invention;

[0064] Figure 3 This is a schematic diagram of the multi-head self-attention mechanism of the present invention;

[0065] Figure 4 This is a structural diagram of the semantic segmentation decoder of the present invention;

[0066] Figure 5 This is a structural diagram of the target detection decoder of the present invention;

[0067] Figure 6 This is a structural diagram of the motion segmentation decoder of the present invention. Detailed Implementation

[0068] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the application will be further described in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in this invention. All non-innovative embodiments based on these embodiments by other researchers in the art are within the protection scope of this invention. Furthermore, the step numbers in the embodiments of this invention are only set for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0069] In one embodiment of the present invention, a multi-task real-time sensing method for lidar combining Transformer coding is provided, such as... Figure 1 As shown, it includes the following steps:

[0070] Step S1: Collect and preprocess the data to generate a two-dimensional network BEV feature map.

[0071] Specifically, in this embodiment, the preprocessing process includes:

[0072] S11: Collect and preprocess LiDAR 3D point cloud map data and related data to form the raw dataset;

[0073] S12: Data Augmentation. To further improve the density of the BEV representation and long-range detection capability, a distance-based point cloud densification method is used to augment the BEV feature map, preserving the object outline, increasing the point cloud density, and improving the clarity of distant object boundaries. The calculation formula is as follows:

[0074]

[0075] x′=(r+Δr)cosθsinφ, y′=(r+Δr)cosθcosφ, z′=(r+Δr)sinθ

[0076] Where (x,y,z) represents the 3D point cloud coordinates, (r,φ,θ) represents the spherical coordinates, the expanded points are distributed along the depth direction (r axis), the object outline is preserved while the point cloud density is increased, especially the boundary clarity of distant objects is improved, Δr represents the depth expansion amount, and (x′,y′,z′) represents the enhanced 3D point cloud coordinates.

[0077] S13: Use the BEV transformation method to convert 3D point cloud data into a 2D network BEV feature map. For a point P = (x, y, z) in the 3D point cloud, map it to the coordinates (u, v) of the BEV mesh using the following formula:

[0078]

[0079] Among them, (x min ,y min (r) represents the lower limit of the BEV range. x ,r y The ) indicates the resolution of the BEV cell. This indicates a floor operation, ensuring that coordinates are mapped to the grid.

[0080] Step S2: The data processed in step S1 is encoded into the BEV feature map using the constructed Transformer encoder, forming feature data at multiple different levels, such as... Figure 2 As shown.

[0081] Specifically, in this embodiment, the Transformer encoder encoding process includes:

[0082] S21: Convert the 2D network BEV feature map into sequence elements and use them as input to the Transformer encoder;

[0083] S22: For each sequence element of the BEV data, perform feature embedding and position embedding to obtain the feature vector E and position vector P respectively. Then, add them together to obtain the final input embedding vector X. The calculation formula is as follows:

[0084]

[0085]

[0086] X = E + P

[0087] Where E is the feature vector, P is the position vector, [f1,f2,...,f C ] is the feature vector of each cell, W emb Here, pos is the feature embedding matrix, and d is the position index. model Let i be the hidden layer dimension of the Transformer model, and i is the dimension index.

[0088] S23: Perform multi-head attention calculations on the input embedding vectors to obtain the corresponding weights, and then perform weighted summation based on the attention weights to obtain the output results of each head.

[0089] First, we will elaborate on the core computational module of the attention mechanism in the Transformer model, such as... Figure 3 As shown, the input vector X is a BEV feature map, which is first processed by the weight matrix W. Q W K W V Project X into a query, key, and value vector.

[0090] Then, perform dot product scaling (divided by) This is to prevent the high-dimensional dot product value from becoming too large, which would cause the softmax gradient to vanish, where d is the vector dimension.

[0091] Next, Softmax normalization is performed to generate an attention weight matrix, highlighting semantically salient regions.

[0092] Finally, the weighted summation outputs a context-aware output matrix by weighting the value vector v.

[0093] Preferably, in this embodiment, the multi-head self-attention mechanism has h parallel heads, and the outputs of the h heads are concatenated and processed by a linear transformation matrix W. O The final output Z of the multi-head self-attention is obtained, and the calculation formula is as follows:

[0094] Q = XW Q K = XW K V = XW V

[0095]

[0096] Z = W O [Z1;Z2;...;Z h ]

[0097] The query vector Q, key vector K, and value vector V are derived from the input embedding vector X through three nonlinear transformation matrices W. Q W K W V get, It is the dimension of K, Z i This is the output of the i-th attention head.

[0098] S24: The output is fed into the feedforward neural network (FFN) to further transform and extract features, enhancing the model's expressive power. The calculation formula is as follows:

[0099] FFN(Z)=W2(ReLU(W1Z+b1))+b2

[0100] Where W1 and W2 are the weight matrices of the first and second fully connected layers, respectively, b1 and b2 are the bias vectors of the first and second fully connected layers, respectively, ReLU is the non-linear activation function, and FFN(Z) is the final feature vector.

[0101] S25: Data is normalized using layer normalization, and residual connections are used to directly add the module's input and output, thus alleviating the vanishing gradient problem and making the model easier to train. The calculation formula is as follows:

[0102]

[0103] Z out1 =Z + LN(Z)

[0104] Z out2 =FFN(Z out1 )+LN(FFN(Z out1 ))

[0105] Where μ is the mean, σ 2 Z is the variance, ∈ is a small constant to prevent division by zero, γ and β are learnable parameters, LN(Z) is the layer normalization function, and Z out1 Z out2 These are the outputs of the multi-head self-attention module and the feedforward neural network FFN, respectively.

[0106] Step S3: Construct specific decoders for different tasks to decode the feature data processed in Step 2, generating feature layers for three different tasks: semantic segmentation, object detection, and motion segmentation.

[0107] Specifically, in this embodiment, constructing a specific Transformer decoder includes the following sub-steps:

[0108] S31. Reconstruct the sequence features output by the Transformer encoder into 2D network data, including: fixed-position encoding and sequence feature reconstruction;

[0109] Fixed position encoding: Add a predefined two-dimensional position encoding P, and map it to the spatial position of the original image to obtain the sequence feature Z′;

[0110] Sequence feature reshaping: Reshaping the sequence feature Z′ into a two-dimensional feature map F reshaped .

[0111] The calculation formula is as follows:

[0112] Z′=Z+P

[0113]

[0114] Where H, W, and D are the height, width, and number of channels of the feature map, respectively, and Reshape is the reshaping function.

[0115] S32. Divide the decoding process into different detection tasks, including semantic segmentation, object detection, and motion segmentation. Use a specific decoder for each detection task to generate the corresponding output.

[0116] Specifically, in this embodiment, the semantic segmentation task decoding involves gradually restoring the spatial resolution of the low-resolution features output by the encoder through a series of upsampling blocks to generate multi-level semantic features; the object detection task decoding involves using a key-point-based detection framework combined with semantic guidance to generate corresponding multi-level fusion features; and the motion segmentation task decoding involves capturing motion information between each frame, gradually restoring the spatial feature resolution using a motion decoder, upsampling the features to their original size, and predicting the motion segmentation mask.

[0117] Furthermore, semantic segmentation tasks such as decoding Figure 4 As shown.

[0118] First, upsampling is performed using transposed convolution to upsample the low-resolution feature map F. low The formula for upsampling to the target size is as follows:

[0119]

[0120] Where s represents the upsampling factor, H target W target C low These represent the target height of the upsampled feature map, the target width of the upsampled feature map, and the number of channels in the low-resolution feature map, respectively. TransposeConv is the upsampling function.

[0121] Subsequently, semantic segmentation and decoding features are concatenated, and the upsampled features F are then... up Features F of the layer corresponding to the encoder skip When concatenating along the channel dimension, the calculation formula is as follows:

[0122]

[0123] Among them, C skip The number of channels for the encoder's skip connection feature is denoted by , and Concat is the concatenation function.

[0124] Next, semantic segmentation decoding convolution processing is used. The concatenated features are fused through a series of convolutional layers to reduce the number of channels and extract contextual information. The calculation formula is as follows:

[0125]

[0126] ConvBlock contains multiple 3×3 convolutions + BatchNorm + ReLU activations.

[0127] Finally, repeat the above process M times to restore the resolution.

[0128] Furthermore, the object detection task decodes as follows: Figure 5 As shown.

[0129] First, perform semantic segmentation branch decoding, process Transformer features, and generate a semantic segmentation map:

[0130]

[0131] Here, SemanticDecoder is the semantic segmentation decoding function, and K represents the number of semantic categories.

[0132] Then, the object detection branch is decoded, directly processing the Transformer features to generate the object detection-related feature map F. det :

[0133]

[0134] Where DetDecoder is the branch decoding function, C det The number of channels in the feature map of the target detection branch.

[0135] Subsequently, the semantic feature S and the detection feature F are combined using the SWAG semantic weighting and guidance module. det Projecting onto the joint embedding space enhances the guiding role of semantic content in object detection tasks; its calculation formula is as follows:

[0136]

[0137] Where σ is the activation function, and K′ is the projection dimension. These are the semantic feature projection weight matrix and the detection feature projection weight matrix, respectively. These are the semantic feature projection bias vector and the detection feature projection bias vector, respectively.

[0138] For two dimensions h in the embedding space sem with h det Perform matching vector calculation to obtain matching vector m:

[0139]

[0140] The generated matching vector m is subjected to corresponding weight vector calculation, and the detected features and weight vectors are then weighted and fused to obtain the final fused feature F. fusion :

[0141]

[0142]

[0143] In this context, ⊙ represents element-wise multiplication, and Conv represents convolution.

[0144] Furthermore, motion segmentation task decoding, such as Figure 6 As shown.

[0145] First, multi-scale two-dimensional features F are obtained from different layers of the Transformer. Then, partial resolution is recovered through multiple transposed convolutions and upsampling. The calculation formula is as follows:

[0146]

[0147] Then, the upsampled features F up Features F of the layer corresponding to the encoder skip The features F are obtained by splicing them together. merge The calculation formula is as follows:

[0148]

[0149] Subsequently, the stitched features are fused using standard convolutional blocks, and the above operation is repeated until the resolution is restored. The calculation formula is as follows:

[0150]

[0151] Finally, repeat the above operations until the resolution is restored, and then compare the final features upsampled to the final resolution with the shallowest features F from the encoder. skip The formula for concatenating and convolving is:

[0152]

[0153] Step S4: Input the feature layer data processed in step S3 into the segmentation head of different tasks to obtain the detection results of different tasks.

[0154] Specifically, this embodiment includes the following sub-steps:

[0155] S41: Output of semantic segmentation task:

[0156] For the final output dense feature map F final The input is fed into a pixel-level classification head, where the Softmax classifier in the classification head performs classification prediction on each pixel to obtain the semantic label P. The calculation formula is as follows:

[0157]

[0158] Where K is the number of categories, P i,j,k Conv represents the probability that pixel (i,j) belongs to class K. 1×1 It is a 1×1 convolution;

[0159] S42: Output of the object detection task:

[0160] First, based on the final fusion feature F fusion Generate a keypoint heatmap and predict the center position of each category c. The calculation formula is as follows:

[0161]

[0162] Among them, H c It is a heatmap of category c, activated by sigmoid, Conv heatmap It is a convolution function used to generate key point heatmaps;

[0163] Then, offset prediction is performed to compensate for the loss of coordinate accuracy caused by downsampling. The calculation formula is as follows:

[0164]

[0165] Among them, O i,j =(o x ,o y ) represents the offset of point (i,j), Conv offset It is a convolution function used to predict the offset of the target center coordinates;

[0166] Size prediction is performed, predicting the width and length of each detection box. The calculation formula is as follows:

[0167]

[0168] Among them, S i,j This represents the size of the object corresponding to point (i,j).

[0169] Subsequently, direction prediction is performed. In this embodiment, the direction is discretized into 36 bins, each bin being 5°. The direction angle θ is determined by the argmax function following the softmax function, and its calculation formula is as follows:

[0170]

[0171] Where Δθ is the discretized direction angle, Conv dir It is a convolution function used to predict the target orientation angle;

[0172] Specifically, in this embodiment, Δθ = 180° / N.

[0173] S43: Output of motion segmentation task:

[0174] First, the features are mapped to the class space using 1×1 convolution, and the probability distribution is generated using the Softmax function, the calculation formula of which is as follows:

[0175]

[0176] Where K∈{1,2} is the category index, representing motion or static state, Conv 1×1 The weight is Bias is

[0177] Then, the probability is calculated, and the formula is as follows:

[0178]

[0179] Where (i,j) represents the pixel coordinates.

[0180] Step S5: Based on the data processed in steps S2 and S3, a multi-task real-time prediction model for lidar that combines Transformer encoding is established, which combines the long-term dependency modeling and parallel computing capabilities of Transformer.

[0181] In summary, the LiDAR multi-task real-time perception method proposed in this invention, which combines Transformer coding, demonstrates higher accuracy and stability in multi-task real-time monitoring, providing a more reliable prediction tool for practical applications.

[0182] The method of this invention can simultaneously realize tasks such as marine target detection, obstacle classification and dynamic trajectory prediction on a shipborne embedded platform. It improves inspection accuracy and computational efficiency through multi-task feature fusion and global attention mechanism. It is applicable to scenarios such as ship inspection, obstacle collision avoidance and dynamic target tracking in offshore wind farms, and can significantly improve real-time navigation safety and operation and maintenance efficiency in complex sea conditions.

[0183] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-task real-time sensing method for lidar combining Transformer coding, characterized in that, Includes the following steps: S1. Collect and preprocess data: Collect and preprocess 3D point cloud data from LiDAR, and use the BEV conversion method to generate a two-dimensional network BEV feature map. S2. Construct a Transformer encoder to encode the BEV feature map of the two-dimensional network, forming feature data at multiple different levels; S3. Divide the decoding process into different detection tasks, including: semantic segmentation task, object detection task, and motion segmentation task. Construct specific Transformer decoders for different tasks to decode the feature data processed in S2 and generate feature layer data for different tasks. S4. Input the feature layer data of different tasks into the segmentation head of different tasks to obtain the detection results of different tasks. Obtain semantic labels based on the semantic segmentation task, obtain predicted detection boxes containing target center key points, target direction and target size based on the target detection task, and obtain motion segmentation masks based on the motion segmentation task. S5. Based on the feature data of different levels formed in step S2 and the feature layer data of different tasks formed in step S3, establish a real-time prediction model for LiDAR multi-tasks that combines Transformer encoding. Combining the long-term dependency and parallel computing capabilities of Transformer, and through cross-level feature pyramid fusion and cross-task semantic weighting mechanism, perform multi-task joint prediction on the LiDAR point cloud scene to obtain real-time perception results including semantic segmentation labels, target detection boxes and motion segmentation masks.

2. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 1, characterized in that, In step S1, the collection and preprocessing of data includes the following sub-steps: S1.1 Collect and preprocess 3D point cloud data from LiDAR to form the raw dataset; S1.

2. Use distance-based point cloud densification method to enhance 3D point cloud data, preserve object outlines, increase point cloud density, and improve the clarity of distant object boundaries. S1.

3. Use the BEV transformation method to convert 3D point cloud data into a two-dimensional network BEV feature map, mapping the 3D point cloud data (x,y,z) to BEV grid coordinates (u,v): Among them, (x min ,y min ) represents the lower limit of the BEV range, (r x ,r y () indicates the resolution of the BEV cell. This indicates the floor function.

3. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 1, characterized in that, In step S2, the Transformer encoder encoding process includes the following sub-steps: S2.

1. Convert the 2D network BEV feature map into sequence elements and use them as input to the Transformer encoder; S2.

2. For each sequence element of the BEV feature map, perform feature embedding and position embedding to obtain the input embedding vector X; S2.

3. Perform multi-head attention calculations on the input embedding vector X to obtain the corresponding attention weights, and then perform a weighted sum to obtain the output of each head; the multi-head self-attention mechanism has h parallel heads, and the outputs of the h heads are concatenated and processed by a linear transformation matrix W. O This yields the final output Z of the multi-head self-attention mechanism. S2.4 Feed the output result Z into the feedforward neural network module to transform and extract features, thereby enhancing the model's expressive power; S2.

5. Use layer normalization to normalize the data and use residual connections to directly add the input and output data to alleviate gradient vanishing.

4. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 3, characterized in that, In step S3, a specific Transformer decoder is constructed, including the following sub-steps: S3.1 Reconstruct the sequence features output by the Transformer encoder into 2D network data, including: fixed-position encoding and sequence feature reconstruction; S3.2 Divide the decoding process into different detection tasks, including semantic segmentation, object detection, and motion segmentation. Use a specific decoder for each detection task to generate the corresponding output.

5. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 4, characterized in that, The fixed-position encoding adds a predefined two-dimensional position code P, mapping the input sequence features Z to the spatial position of the original image. The calculation formula is as follows: Z′=Z+P The sequence feature reshaping process reshapes the sequence feature Z′ into a two-dimensional feature map F. reshaped : Where H, W, and D are the height, width, and number of channels of the feature map, respectively, and Reshape is the reshaping function.

6. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 5, characterized in that, Specific decoders are used for different detection tasks to generate corresponding outputs, as follows: S3.2.1 Semantic Segmentation Task Decoding: The low-resolution features output by the encoder are gradually restored to spatial resolution through a series of upsampling blocks to generate multi-level semantic features; S3.2.2 Target Detection Task Decoding: A key-point-based detection framework is adopted, combined with semantic guidance, to generate corresponding multi-level fusion features; S3.2.3 Motion segmentation task decoding: Capture motion information between each frame, use a motion decoder to gradually restore spatial feature resolution, upsample the features to the original size, and predict the motion segmentation mask.

7. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 6, characterized in that, The semantic segmentation task decoding includes the following sub-steps: Upsampling: Upsampling is performed using transposed convolution to upsample the low-resolution feature map F. low Upsampled to the target size, feature F is obtained. up : Where s represents the upsampling factor, H target W target C low These represent the target height of the upsampled feature map, the target width of the upsampled feature map, and the number of channels in the low-resolution feature map, respectively; TransposeConv is the upsampling function. Feature concatenation: The upsampled features F... up Features F of the layer corresponding to the encoder skip By concatenating along the channel dimension, we obtain the concatenated feature F. concat : Among them, C skip The number of channels for the encoder's skip connection features; Concat is the concatenation function. Convolutional processing: Semantic segmentation decoding convolutional processing is employed, and the data is fused and concatenated through a series of convolutional layers to reduce the number of channels and extract contextual information, resulting in the output feature F. out : Among them, C out This represents the number of channels in the output feature map after convolution processing during semantic segmentation decoding, and ConvBlock is the convolution function.

8. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 7, characterized in that, The target detection task decoding includes the following sub-steps: Semantic segmentation branch decoding: Processing Transformer features to generate semantic information S: Where SemanticDecoder is the semantic segmentation decoding function, and K represents the number of semantic categories; Object detection branch decoding: directly processes Transformer features to generate object detection features F. det : Where DetDecoder is the branch decoding function, C det The number of channels in the feature map of the target detection branch; SWAG semantic weighting and guidance: fusing semantic information S into the object detection branch and projecting it into the joint embedding space: Where σ is the activation function, and K′ is the projection dimension. These are the semantic feature projection weight matrix and the detection feature projection weight matrix, respectively. These are the semantic feature projection bias vector and the detection feature projection bias vector, respectively. For two dimensions h in the embedding space sem with h det Perform matching vector calculation to obtain matching vector m: The matching vector m is weighted accordingly, and the detected features and weight vectors are weighted and fused to obtain the fused feature F. fusion : Where ⊙ represents element-wise multiplication, and Conv represents convolution.

9. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 8, characterized in that, The motion segmentation task decoding includes the following sub-steps: Encoder feature extraction: Multi-scale features are obtained from different layers of the Transformer, and partial resolution is recovered through multiple transposed convolutions and upsampling. Upsampling and final feature fusion: The upsampled features F up Features F of the layer corresponding to the encoder skip The features F are obtained by splicing them together. merge : The stitched features are fused using standard convolutional blocks, and this process is repeated until the final resolution is restored. The features F upsampled to the final resolution up With the shallowest layer features F of the encoder skip The concatenation and convolution processes yield the final dense feature map F. final : Where C″ is the number of channels for the intermediate upsampling feature, C final This represents the number of channels in the final feature map. Segmentation head prediction: using feature F final Map to the column space and use Softmax to generate the probability distribution.

10. The multi-task real-time sensing method for lidar combined with Transformer coding according to claim 9, characterized in that, In step S4, the feature layer data processed in S3 is input into the segmentation heads of different tasks to obtain the detection results of different tasks, including the following sub-steps: S4.1 Semantic Segmentation Task Output: The final output dense feature map F final The input is fed into a pixel-level classification header, where the Softmax classifier performs classification prediction on each pixel to obtain the semantic label P: Where K is the number of categories, P i,j,k Conv represents the probability that pixel (i,j) belongs to class K. 1×1 It is a 1×1 convolution; S4.2, Target Detection Task Output: S4.2.1 Target Center Key Point Prediction: Based on Fusion Feature F fusion Generate a keypoint heatmap and predict the center location of each category c: Among them, H c It is a heatmap of category c, activated by the sigmoid function, Conv heatmap It is a convolution function used to generate key point heatmaps; S4.2.2 Target Size Prediction: Perform offset prediction to compensate for the loss of coordinate accuracy caused by downsampling: Among them, o i,j =(o x ,o y ) represents the offset of point (i,j), Conv offset It is a convolution function used to predict the offset of the target center coordinates; Perform size prediction, predicting the width and length of each detection box: Among them, S i,j It is the size of the object corresponding to point (i,j), Conv size It is a convolution function used to predict target size; S4.2.3 Target Direction Prediction: Direction prediction is performed by discretizing the direction. The direction angle θ is determined by the argmax function following the softmax function. θ i,j =argmax(Softmax(Θ i,j ))×Δθ Where Δθ is the discretized direction angle, Conv dir It is a convolution function used to predict the target orientation angle; S4.3, Motion Segmentation Task Output: via Conv 1×1 The features are mapped to the class space, and the probability distribution is generated using the Softmax function: Probability calculations are performed to obtain the motion segmentation mask M: Where K∈{1, 2+}, represents the category index, W seg For Conv 1×1 The weight, b seg For Conv 1×1 The bias.