A Method for Extracting and Identifying Multi-Perspective Features of Radar Images
By constructing a multi-view feature extraction and identification network, using Swin Transformer block and loss function optimization, the problem of insufficient information utilization in radar image multi-view recognition is solved, and high-precision radar target classification is achieved.
Patent Information
- Application Number
- CN202310901621.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-21
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-07-21
AI Technical Summary
The existing radar image multi-view recognition method fails to fully utilize the relevant information between different viewing angles, resulting in insufficient recognition performance, especially in EOC experiments, the recognition accuracy needs to be improved.
By constructing a multi-view feature extraction and identification network, multi-view feature is extracted using Swin Transformer block, and network parameters are optimized by combining cross entropy loss and triple loss to achieve effective classification of multi-view feature.
When using a small number of original data sets, the accuracy of radar image classification is significantly improved and the performance of radar target automatic identification system is improved.
Smart Images

Figure CN116778341B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of automatic target recognition of radar images, and particularly relates to a method for extracting and discriminating multi-view features of radar images. Background Art
[0002] Synthetic aperture radar in radar has been widely used in many civilian and military fields due to its all-weather, all-day and high-resolution imaging capabilities. However, due to the speckle noise and complex features in radar images, it is usually difficult to interpret and understand intuitively. Automatic target recognition is the key to synthetic aperture radar image interpretation. In recent years, with the development of machine learning, deep learning-based methods have greatly improved the recognition accuracy and efficiency of radar images. Currently, most automatic target recognition methods for radar images are proposed for single-view input. However, multi-view radar images contain richer classification features. To further improve the performance of radar target automatic recognition systems, it is necessary to extract and discriminate effective features from multi-view radar images.
[0003] In practice, modern radars can obtain radar images from different views, which include richer classification features than a single view. Therefore, some studies on multi-view modes have been proposed in recent years and some promising results have been achieved. The literature "Zhang, F.; Hu, C.; Yin, Q.; Li, W.; Li, H.; Hong, W. Multi-Aspect-Aware Bidirectional LSTM Networks for Synthetic Aperture Radar Target Recognition. IEEE Access 2017, 5, 26880–26891." proposed a bidirectional long short-term memory recurrent neural network structure based on learning of spatially varying scattering information to achieve the extraction of spatial scattering features. However, this method still needs to use a large number of radar images and does not fully utilize the correlation information between different multi-view images. The literature "Pei, J.; Huang, Y.; Huo, W.; Zhang, Y.; Yang, J.; Yeo, T. SAR Automatic Target Recognition Based on Multiview Deep Learning Framework. IEEE Trans. Geosci. Remote Sens. 2018, 56, 2196–2210." proposed a deep learning radar target automatic recognition framework based on multi-view, adopting a multi-input parallel network topology to extract and fuse features of radar images input from different views layer by layer. However, the recognition performance of this method needs to be improved, especially in the EOC experiment. Summary of the Invention
[0004] The object of the present invention is to overcome the deficiencies of the prior art and provide a method for extracting and discriminating multi-view features of radar images. By extracting and discriminating multi-view features, the accuracy of radar image classification can be effectively improved, and the performance of the radar target automatic recognition system can be enhanced.
[0005] The object of the present invention is achieved by the following technical solutions: A method for extracting and discriminating multi-view features of radar images, comprising the following steps:
[0006] S1. The radar platform collects ground target image samples: The radar platform obtains multi-view images of a given ground target at different elevation angles and azimuth angles within different visual ranges.
[0007] S2. Preprocess the collected radar image samples; including the following sub-steps:
[0008] S21. Rotate by azimuth angle: All radar images will be rotated by a specific azimuth angle to align them to the same azimuth.
[0009] S22. Central cropping and normalization: Use the method of central cropping to crop the collected radar image samples into slices of the same size with the target located at the center, and perform normalization processing on the slices.
[0010] S23. Use a gray-scale enhancement method based on a power function to perform gray-scale enhancement processing on the image.
[0011] S3. Construct a multi-view image combination data set: Obtain a data set by arranging and combining multi-view radar images of the target within the same view interval.
[0012] S4. Build a multi-view image combination feature extraction network: The smallest unit of the picture is transformed from pixels into blocks of a preset size through a slice division layer, and the pixel values in a block are combined into a vector; the generated vector will sequentially pass through three consecutive stages, stage 1, stage 2, and stage 3. Stage 1 consists of a linear embedding layer and a Swin Transformer block; stage 2 and stage 3 consist of a slice merging layer and a Swin Transformer block.
[0013] S5. Build a multi-view image combination feature discrimination network: Input the multi-view features into the global average pooling layer and the feature dimensionality reduction module respectively. The multi-view features input into the global average pooling layer pass through the fully connected layer to obtain the predicted label, which is used as the discrimination result of the multi-view image combination feature; and calculate the distance between the probability distributions of the predicted label and the true label, which is used as the cross-entropy loss l CE ;
[0014] The feature dimensionality reduction module reduces the dimensionality of the input multi-view features and classifies the multi-view features into three categories: anchor, positive, and negative. Anchor is a sample randomly selected from the training dataset. Positive represents a sample of the same category as the anchor, and negative represents a sample of a different category. The triplet loss describes reducing the distance between the positive and the anchor and increasing the distance between the negative and the anchor, which is expressed as:
[0015]
[0016] where and represent the i-th samples in the anchor, positive, and negative respectively. N represents the total number of N samples. represents the L2 norm, and m requires that the difference between the distance between the anchor and the negative and the distance between the anchor and the positive is greater than m.
[0017] The final combined loss function l constructed by the feature discrimination network part Joint is expressed as:
[0018] minimize l Joint = minimize(λl CE + μl Triplet )
[0019] where λ and μ are hyperparameters, representing the weights of the cross-entropy loss and the triplet loss respectively. According to the combined loss function, the parameters of the multi-view image combination feature extraction network are optimized using the backpropagation algorithm.
[0020] S6. Input the dataset obtained in S3 into the multi-view image combination feature extraction network and the multi-view image combination feature discrimination network for training, and use the trained network to identify unknown radar images.
[0021] The specific implementation method of the step S3 is as follows: Assume that Y (raw) = {Y1, Y2, …, Y C} represents the radar original image set. The image set belongs to the i-th target category, and their corresponding azimuth angles are represents the target category label, C represents the number of target categories, and n i represents the total number of images of the i-th target category. For the given number of views k, obtain all view combinations of a class of radar images, and the number of combinations is Then, each combination The images in or Finally, the multi-view radar images of the target within the same viewing angle range θ are arranged and combined, that is to obtain the dataset of the i-th target category.
[0022] The Swin Transformer block includes two consecutive sub-blocks, which extract local and global features by calculating self-attention in local and cross windows respectively; the first sub-block sequentially includes a normalization layer, a window-based multi-head self-attention mechanism, a normalization layer, and a multi-layer perceptron, and the second sub-block sequentially includes a normalization layer, a sliding window-based multi-head self-attention mechanism, a normalization layer, and a multi-layer perceptron, and a residual structure is used for connection after the window-based multi-head self-attention mechanism, the sliding window-based multi-head self-attention mechanism, and the multi-layer perceptron.
[0023] The beneficial effects of the present invention are: compared with the prior art, the present invention utilizes the multi-view feature extraction part and the feature discrimination part, and can effectively extract multi-view features from the input radar image, and group the same kind together and separate different kinds, so as to realize the effective classification of radar image targets. Compared with the existing radar image deep network classification methods, the method of the present invention can still achieve excellent classification performance when only using a small amount of original datasets, can effectively improve the accuracy of radar image classification, and improve the performance of the radar automatic target recognition system. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flowchart of the method of the present invention;
[0025] Figure 2 is a schematic diagram of the geometric model for collecting ground target radar images adopted by the present invention;
[0026] Figure 3 is a schematic diagram of generating nine 3-view radar image combinations from six original radar images of the present invention;
[0027] Figure 4 is a schematic diagram of the multi-view feature extraction and discrimination network structure of the present invention;
[0028] Figure 5 is a schematic diagram of the Swin Transformer block structure of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0029] The present invention uses simulation experiments to verify all the proposed steps and conclusions, and the simulation experiments are verified correctly on the platforms of pytorch1.12.0, python3.7 and windows10 operating systems. To facilitate the understanding of the technical content of the present invention by those skilled in the art, the content of the present invention will be further elaborated below with reference to the accompanying drawings.
[0030] As Figure 1 shown, a method for extracting and discriminating multi-view features of radar images according to the present invention includes the following steps:
[0031] S1. The radar platform collects ground target image samples: In the multi-view synthetic aperture radar signal acquisition, the radar platform obtains multi-view images of a given ground target at different elevation angles and azimuth angles within different viewing distances; the geometric model of this embodiment is as Figure 2 shown. For the sake of easy analysis, only the change of azimuth angle is considered here. For a given viewing angle interval θ and the number of viewing angles K (K>1), the radar platform sequentially collects radar images of the original ground target (Target) with the same resolution from the azimuth angles of 0 to 360° (View1, View2, View3,..., View k).
[0032] S2. Preprocess the collected radar image samples; it includes the following sub-steps:
[0033] S21. Rotate by azimuth angle: Radar images are usually sensitive to views or azimuth heights. In order to reduce the sensitivity of the azimuth difference and maintain the electromagnetic scattering information of the target from multiple perspectives at the same time, all radar images will be rotated by a specific azimuth angle to align them to the same azimuth; all radar image samples are transformed through a rotation matrix, and the matrix is:
[0034]
[0035] where is the angle of radar image rotation relative to the given coordinate axis, [u v] T is the coordinate of the transformed radar image, and [p q] T is the original coordinate of the radar image.
[0036] S22. Central cropping and normalization: Using the method of central cropping, the collected radar image samples are cropped into slices of the same size with the target located at the center, and the slices are normalized; the expression of the normalization process is:
[0037]
[0038] Where X is the image before normalization, and X′ represents the image after normalization. X(i,j) represents the pixel value at the (m,n) position of the image, min[X] represents the minimum pixel value in image X, and max[X] represents its maximum value.
[0039] S23. Perform gray-scale enhancement processing on the image using a gray-scale enhancement method based on a power function, and its expression is:
[0040] x′(u,v)=[x(u,v)] β
[0041] where β is the enhancement factor.
[0042] S3. Construct a multi-view image combination dataset: Obtain the dataset by arranging and combining multi-view radar images of the target within the same view angle range; the specific implementation method is: Assume Y (raw) ={Y1,Y2,…,Y C} represents the set of original radar images. The image set belongs to the i-th target category, and their corresponding azimuth angles are represents the target category label, C represents the number of target categories, and n i represents the total number of images of the i-th target category; for a given number of view angles k, obtain all view angle combinations of a class of radar images, and the number of combinations is Then, the images in each combination are arranged in ascending order according to their azimuth angles, that is or Finally, arrange and combine the multi-view radar images of the target within the same view angle range θ, that is to obtain the dataset of the i-th target category.
[0043] An example of the arrangement and combination method is as Figure 3 shown. In each view angle interval θ, the number of view angles k = 3, and nine three-view radar image combinations for training can be obtained from only six original radar images, as Figure 3 shown.
[0044] And as θ and k increase, for a given number of original radar images, more training data can be obtained. Therefore, for each original radar target category, we can obtain sufficient multi-view radar image combinations from a small number of original radar images to train the network.
[0045] Use the publicly available measured radar ground moving and stationary target MSTAR (moving and stationary target acquisition and recognition) dataset. For the training dataset, in the case of 2 viewpoints with k = 2, only about 50% of the original dataset is used to construct the multi-view radar image combination. For the 3-view input with k = 3, only about 33% is used, and for the 4-view with k = 4, only about 20% is used. The specific usage quantity of each type of target is shown in Table 1 and Table 2. Table 1 shows the dataset situation under the SOC (standard operating condition), and Table 2 shows the dataset situation under the EOC-C (extended operating condition - configuration variant). Through the multi-view combination of the original radar images, 21,834, 48,764, and 43,533 multi-view combinations can be formed for the training dataset under the SOC condition for 2 viewpoints, 3 viewpoints, and 4 viewpoints respectively. Under the EOC-C condition, 7,160, 14,445, and 11,380 multi-view combinations can be formed for the training dataset for 2 viewpoints, 3 viewpoints, and 4 viewpoints respectively. For the test dataset, all the original radar images will be used to form the multi-view combination, but for each type of target, we only randomly extract 2,000 samples from the formed multi-view combination. That is, the size of the test dataset is 20,000 under the 10-class target situation of SOC and 14,000 under the 7-target situation of EOC-C.
[0046] Table 1 Quantity of original radar images used in the training and test datasets under the SOC condition
[0047]
[0048]
[0049] Table 2 Quantity of original radar images used in the training and test datasets under the EOC-C condition
[0050]
[0051] S4. Build a multi-view image combination feature extraction network: The feature extraction part is one of the key components of the proposed method, and its network structure is as Figure 4As shown in the upper part. After the multi-view radar image is read in, it is represented as a pixel matrix. First, through the patch partition layer, the smallest unit of the picture is changed from a pixel to a block of a preset size (4×4), that is, the pixel matrix is segmented by blocks containing 4×4 pixels, and the pixel values in a block are combined into a vector; then the generated vector will sequentially pass through three consecutive stages, stage 1, stage 2, and stage 3. Stage 1 consists of a Linear Embedding layer and a Swin Transformer block; the Linear Embedding layer converts the size of the input vector into a preset value that the Swin Transformer block can adapt to. Then, through the Patch Merging layer, the network is constructed into a hierarchical structure, so that multi-scale features can be obtained, and the number of vectors will gradually decrease during the deepening of the network, which is similar to the pooling layer in a convolutional neural network. Stage 2 and stage 3 consist of a Patch Merging layer and a Swin Transformer block.
[0052] The core element of the multi-view feature extraction part is the Swin Transformer block, and its specific structure is as Figure 5 shown. The Swin Transformer block includes two consecutive sub-blocks, which extract local and global features by calculating self-attention in local and cross windows respectively; sub-block one sequentially includes a layer normalization (LN), a window based multi-head self-attention mechanism (W-MSA), a normalization layer, and a multi-layer perceptron (MLP). Sub-block two sequentially includes a normalization layer, a shifted window based multi-head self-attention mechanism (SW-MSA), a normalization layer, and a multi-layer perceptron, and a residual structure is used for connection after the window based multi-head self-attention mechanism, the shifted window based multi-head self-attention mechanism, and the multi-layer perceptron;
[0053] The forward process of the Swin Transformer block is shown by the following formula:
[0054]
[0055] where is the feature obtained by adding the output of the window based multi-head self-attention mechanism to the original input; z l is the output of the multi-layer perceptron and The added feature is also the output feature of Sub-block 1; is the output of the multi-head self-attention mechanism based on the sliding window and the added feature; z l+1 is the output of the multi-layer perceptron and the added feature, which is also the output feature of Sub-block 2; l represents the l-th Swin Transformer block.
[0056] The window-based multi-head self-attention mechanism (W-MSA) means dividing the features into small windows and performing multi-head self-attention mechanism (MSA) calculations within each small window. The sliding window-based multi-head self-attention mechanism (SW-MSA) is because the window-based multi-head self-attention mechanism only calculates within each window, and there is no information transfer between windows. If the window is offset and then the multi-head self-attention mechanism is calculated, this problem can be avoided. The offset method adopted can be understood as the window being offset by half a pixel of the window size to the right and down respectively from the upper left corner of the feature map. The parts that are extra due to the offset in the lower and right sides are respectively filled into the parts that are vacant due to the offset in the upper and left sides.
[0057] The calculation expression of the multi-head self-attention mechanism is:
[0058]
[0059] where and d K = d V = d model / n; d model represents the dimension of the network model, W i Q 、W i K 、W i V and W O all represent the weight matrices corresponding to the superscripts; n represents the number of heads of the multi-head self-attention mechanism, that is, how many times the self-attention mechanism is calculated, i ∈ [1, n].
[0060] The calculation expression of the self-attention mechanism is:
[0061]
[0062] where Q, K, and V respectively represent the matrices formed by packing a series of queries, keys, and values together, d K represents the dimension of matrix K, and the softmax function represents the normalized exponential function. By using the softmax function, the output values of multi-classification are converted into a probability distribution with a sum of 1 and a range in [0, 1];
[0063] S5. Build a multi-view image combined feature discrimination network; the feature discrimination part combines cross-entropy loss (CE loss) and triplet loss. Its network structure is as shown in the lower part of Figure 4 The multi-view features are respectively input into the global average pooling layer and the feature dimensionality reduction module. The multi-view features input into the global average pooling layer obtain the predicted labels through the fully connected layer, which are used as the discrimination results of the multi-view image combined features; and calculate the distance between the predicted labels and the probability distributions of the true labels as the cross-entropy loss l CE ;
[0064] The feature dimensionality reduction module reduces the dimensionality of the input multi-view features, and divides the multi-view features into three categories: anchor, positive, and negative; anchor is a sample randomly selected from the training dataset, positive represents a sample of the same category as anchor, and negative represents a sample of a different category; Triplet loss describes reducing the distance between positive and anchor, and expanding the distance between negative and anchor, expressed as:
[0065]
[0066] where and respectively represent the i-th sample in anchor, positive, and negative, N represents a total of N samples, represents the L2 norm, and m requires that the difference between the distance between anchor and negative and the distance between anchor and positive is greater than m.
[0067] The final combined loss function l constructed by the feature discrimination network part Joint is expressed as:
[0068] minimizel Joint =minimize(λl CE +μl Triplet )
[0069] where λ and μ are hyperparameters, representing the weights of cross-entropy loss and triplet loss respectively; according to the combined loss function, use the backpropagation algorithm to optimize the parameters of the multi-view image combined feature extraction network.
[0070] S6. Input the dataset obtained in S3 into the multi-view image combined feature extraction network and the multi-view image combined feature discrimination network for training, and use the trained networks to identify unknown radar images. Use the multi-view combined dataset construction method described in S3 to construct the network training dataset and the test dataset respectively, and then train the network. Stop training when the accuracy rate of the test dataset stabilizes and no longer increases to obtain the final multi-view image combined feature extraction network. Considering the trade-off between data acquisition cost and network training cost, the viewing angle interval θ in the multi-view training and test experiments is set to 45°. During the training process, the initial learning rate is set to 0.0001, the batch size is set to 16, the window size is set to 4×4, and the Adam optimizer is used for training acceleration optimization. Through the automatic learning rate adjustment and the training method of resuming training from breakpoints of the Adam optimizer, continuously improve the recognition rate of the network for ground target categories, so that the designed network obtains better feature extraction and discrimination capabilities. In addition, for the enhancement factor β of gray-scale enhancement in the radar image preprocessing in S2, it is set to 0.4, and the size of the radar image after central cropping is set to 96×96.
[0071] Table 3 is the confusion matrix of the 4-view classification results of the method of the present invention under the SOC condition. Table 4 is the table of the number of samples and the recognition rate of the present invention under the SOC condition. The recognition rate for 2 views is 99.45%, the recognition rate for 3 views is 99.61%, and the recognition rate for 4 views is 99.67%. Table 5 is the table of the number of samples and the recognition rate of the present invention under the EOC-C condition. The recognition rate can reach 99.89% for 4-view input, the recognition rate for 2 views is 99.29%, and the recognition rate for 3 views is 99.37%. Moreover, the radar image combination method can obtain a large number of multi-view radar image combinations. Therefore, for the 2-view input situation, only about 50% of the original dataset is used to construct the multi-view radar image combination, only about 33% for 3-view input, and only about 20% for 4-view input. However, the recognition rates for different views are all higher than 99%. It can be seen that the method of the present invention can still achieve excellent classification performance when only using a small amount of the original dataset.
[0072] Table 3
[0073]
[0074] Table 4 Table of the number of samples and the recognition rate under the SOC condition
[0075]
[0076]
[0077] Table 5 Table of the number of samples and the recognition rate under the EOC-C condition
[0078] Number of original radar images Number of training samples generated Recognition rate 2 viewpoints 499 7160 99.29% 3 viewpoints 334 14445 99.37% 4 viewpoints 251 11380 99.89%
[0079] Those of ordinary skill in the art will realize that the embodiments described herein are to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.
Claims
1. A method for extracting and discriminating multi-view features of radar images, characterized in that It includes the following steps: S1. The radar platform collects ground target image samples: The radar platform obtains multi-view images of a given ground target at different elevation angles and azimuth angles within different line-of-sight distances; S2. Preprocess the collected radar image samples; It includes the following sub-steps: S21. Rotation by azimuth angle: All radar images will be rotated by a specific azimuth angle to align them to the same azimuth; S22. Central cropping and normalization: Using the method of central cropping, the collected radar image samples are cropped into slices of the same size with the target at the center, and the slices are normalized; S23. Use the gray-scale enhancement method based on the power function to perform gray-scale enhancement processing on the images; S3. Construct a multi-view image combination dataset: The multi-view radar images of the target within the same view interval are combined through permutation and combination to obtain a dataset; S4. Build a multi-view image combination feature extraction network: The smallest unit of the picture is changed from pixels to blocks of a preset size through slice partitioning layers, and the pixel values in a block are combined into a vector; The generated vector will sequentially pass through three consecutive stages, stage 1, stage 2, and stage 3. Stage 1 consists of a linear embedding layer and a Swin Transformer block; Stage 2 and stage 3 consist of a slice merging layer and a Swin Transformer block; S5. Build a multi-view image combined feature discrimination network: Input the multi-view features into the global average pooling layer and the feature dimensionality reduction module respectively. The multi-view features input into the global average pooling layer pass through the fully connected layer to obtain the predicted labels, which are used as the discrimination results of the multi-view image combined features; and calculate the distance between the probability distributions of the predicted labels and the true labels as the cross-entropy loss l CE ; The feature dimension reduction module reduces the dimension of the input multi-view features and classifies the multi-view features into three categories: anchor, positive, and negative; Anchor is a sample randomly selected from the training dataset, positive represents a sample of the same category as the anchor, and negative represents a sample of a different category; The triplet loss describes reducing the distance between the positive and the anchor and expanding the distance between the negative and the anchor, expressed as: wherein and respectively represent the i-th samples in anchor, positive and negative, N represents there are N samples in total, represents the two-norm, and m requires that the difference between the distance between anchor and negative and the distance between anchor and positive needs to be greater than m; The final combined loss function l constructed for the feature discrimination network part Joint is expressed as: minimize l Joint = minimize(λl CE + μl Triplet ) where λ and μ are hyperparameters, representing the weights of the cross-entropy loss and the triplet loss respectively; According to the joint loss function, the parameters of the multi-view image combination feature extraction network are optimized using the backpropagation algorithm; S6. Input the dataset obtained in S3 into the multi-view image combination feature extraction network and the multi-view image combination feature discrimination network for training, and use the trained network to identify unknown radar images.
2. The method for extracting and identifying multi-view features of radar images according to claim 1, characterized in that, The specific implementation method of step S3 is as follows: Assume Y (raw) ={Y1, Y2, …, Y C} represents the set of original radar images. The image set belongs to the i-th target category, and their corresponding azimuth angles are represents the target category label, C represents the number of target categories, and n i represents the total number of images of the i-th target category. For a given number of viewpoints k, all viewpoint combinations of a class of radar images are obtained, and the number of combinations is Then, the images in each combination are arranged in ascending order according to their azimuth angles, that is Or Finally, the multi-view radar images of the target within the same viewpoint interval θ are arranged and combined, that is The dataset of the i-th target category is obtained.
3. A method for extracting and identifying multi - perspective features of radar images according to claim 1, characterized in that, The Swin Transformer block includes two consecutive sub-blocks, which extract local and global features by calculating self-attention in local and cross windows respectively; Sub-block one sequentially includes a normalization layer, a window-based multi-head self-attention mechanism, a normalization layer, and a multi-layer perceptron. Sub-block two sequentially includes a normalization layer, a sliding window-based multi-head self-attention mechanism, a normalization layer, and a multi-layer perceptron, and a residual structure is used for connection after the window-based multi-head self-attention mechanism, the sliding window-based multi-head self-attention mechanism, and the multi-layer perceptron.
Citation Information
Patent Citations
Radar automatic target identification method based on multi-view variable convolutional neural network
CN113505833A
BEV-based image detection model training and target detection method and device
CN116188893A